Skip to main content

Automated Data Classification and Semantic Types

Data Classification and Semantic Types enable easy categorization of your data and help you ensure compliance with various information protection acts. If Classification is added to an object, the relevant badge is displayed on all its representations within your repository and can be seen by all users.

Badge

Dataedo comes with a wide range of tools to help you with automatic Classification and with fine-tuning the Semantic Types and Classification protocols to your organization's needs.

Predefined Data Classification​

Dataedo ships with predefined Data Classifications, matching various information privacy protection acts. These are:

CCPA - California Consumer Privacy Act
Identify and classify data related to CCPA compliance.
FERPA - Family Educational Rights and Privacy Act
Find and manage student data in compliance with FERPA.
GDPR - General Data Protection Regulation
Detect columns containing personal data as defined by GDPR.
HIPAA - Health Insurance Portability and Accountability Act
Classify and protect health-related data under HIPAA.
PCI - Payment Card Industry
Identify and secure payment card data for PCI compliance.
PII - Personally Identifiable Information
Find and classify PII to enhance data protection.
PIPEDA - Personal Information Protection and Electronic Documents Act
Find and classify PIPEDA to enhance data protection.
caution

Please note that the above built-in Dataedo Classifications should be treated as a starting point and help with fulfilling these policies. We do not track the most recent changes in them and we cannot guarantee that they are up to date with the current regulation status.

Semantic Types​

Semantic Types identify major classes of your data. Using data samples or column data (depending on your configuration), Dataedo can detect major categories the data in your column falls into (like names, postal addresses, identity card numbers, and many more).

Dataedo ships with over 80 ready-to-use Semantic Types. You can edit them and add new ones in Catalog Settings.

Semantic Type classification is the basis of our Data Classification process.

How it works​

  1. Dataedo checks if the connector used for your Metadata Import supports Data-based Classification and if Data Access is enabled
  2. For supported connectors — Dataedo extracts a data sample of the first 1000 values from each column. The samples are then tested against all existing Semantic Types and their rules. A match percentage is calculated for each Semantic Type
  3. If Data Access is disabled or a connector is not supported — column names are compared against Semantic Types with column-name rules. Match information is retained
  4. If a unique match with a percentage value over 70 exists, a Semantic Type is assigned to a column. If multiple matches exist, the one with the highest percentage is selected. That Semantic Type is then used to determine Data Classification. If there are multiple matches with the same, highest percentage, none will be selected
  5. In Steward Hub, you can review columns where a Semantic Type could not be assigned automatically

Data Access​

By default, Data Classification and Semantic Type detection use a sample of the data stored in your databases. If you want to restrict Dataedo's access and ensure that no actual data is read, you can disable Data Access in settings.

Navigate to Settings>System settings and click the General tab. There, disable the Enable Data-Based Classification toggle. Then click Save to confirm your choice.

Bar

This means that metadata (schema information) is used for your classification exclusively — only column name-based rules will be used to detect Semantic Types. All other Semantic Type features (like type-based Classification, manual Semantic Type assignment or dashboards) will still be usable.

We recommend leaving Data Access on. Without it, Data Classification is less accurate and might lead to an increase in false positives.

warning

If you have run Classification with Data Access on in the past and then disable it, column-based matching will not be possible on already classified columns. Only past Classification data will be taken into account.

What is classified​

Data Classification works on structures that can contain data fitting one Semantic Type category. This means that larger data structures like tables or entire databases cannot be classified — a table can contain multiple columns, each corresponding to a different Semantic Type.

On the other hand, structures like columns or arrays can (and often should) hold data that corresponds to a single Semantic Type. As such, Dataedo's Data Classification can target those structures.

Therefore, the following objects can be classified:

  • columns of a table
  • columns of a dataset
  • columns of a view

Quick Start​

This section shows a simple, no-additional-configuration-needed process to run your first Data Classification.

Step 1 — activate Data Classification​

Before any Data Classification is possible, you first have to choose which Classifications to use in your repository. Head to Settings>Catalog settings and click the Classifications tab. Choose the Classification you wish to use and click its Active toggle. You can now use it for Data Classification. You can have multiple Data Classifications active at once.

settings

If none of the available Classifications match your needs, you can define your own.

Step 2 — run Metadata Import and Data Classification​

Active Data Classifications run automatically after a Metadata Import unless disabled. Wait for your next scheduled import or trigger it directly from the Schedule tab.

Bar

Running a Metadata Import automatically schedules a Data Classification task immediately after it if conditions are met.

Step 3 — explore results​

After Data Classification finishes, badges with the summary of assigned Semantic Types and Data Classification will appear directly on objects. Hovering over the badge shows a full breakdown of objects that have been classified.

Badge

You can also check a full overview of your Classification in Data Governance>Classifications.

Badge

In that dashboard you can switch between global, repository-wide classification widgets and specific Classifications (a). You can also filter based on data sources and Domains (b).

filter

The view offers statistics regarding Classification per data source and Domain, as well as sources where many columns are still not classified.

filter

Steward Hub​

When it is not possible to assign a Semantic Type with enough confidence, the type and classification suggestions have to be manually confirmed by users. Steward Hub will show the objects requiring extra attention in the Semantic Types section. Learn more here.

Badge

Metadata Sync​

Dataedo can track diverging Data Classification changes between your data source and your Dataedo repository and optionally propagate them in either direction. This lets you keep classification consistent across both systems and keep track of all conflicts that can arise when changing information either in Dataedo or your original data source.

Learn more here

Supported Sources​

The connectors listed below support Data-based Classification. Other connectors can still benefit from classification, but based exclusively on schema information.

Dataedo is an end-to-end data governance solution for mid-sized organizations.
Data Lineage • Data Quality • Data Catalog