Sensitive Data Catalog Classification Using Enrichment Labels
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data catalog systems face inefficiencies in manually classifying large volumes of data for sensitivity, leading to time-consuming and error-prone processes that can result in data leakages.
Innovation Solution
A data catalog system that automatically identifies and classifies sensitive information using multiple sensitive data discovery techniques, computing confidence scores and determining enrichment labels to enhance accuracy and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual classification of sensitive data is performed by subject matter experts, then classification accuracy can be maintained, but the process becomes time-consuming and error-prone
Solution Approach 1:
The system enables automatic self-classification of data objects by computing sensitivity scores based on multiple discovery techniques and confidence scores, eliminating the need for manual classification by subject matter experts while maintaining accuracy through automated enrichment labels
Solution Approach 2:
The patent replaces the mechanical manual process of classification with an automated computational system that uses multiple sensitive data discovery techniques, confidence score computation, and weighted scoring to automatically classify data objects
2Reliability
If multiple standalone tools are used for sensitive metadata discovery, then comprehensive detection can be achieved, but the process becomes computationally expensive and time-consuming
Solution Approach 1:
The patent merges multiple sensitive data discovery techniques into a single integrated system that processes data objects through multiple techniques simultaneously, combining their results through confidence score computation and weighted scoring to achieve comprehensive detection without the overhead of separate standalone tools
Solution Approach 2:
The system creates a universal data catalog system that performs multiple functions including sensitive data discovery, confidence score computation, sensitivity score calculation, and automatic classification within a single platform, eliminating the need for multiple specialized tools
Data Source
AI summary
A data catalog system is described that includes capabilities for automatically identifying and classifying sensitive information stored in data objects associated with various data sources. The data catalog system identifies a data object associated with a data asset stored in a data catalog metadata repository and computes a sensitivity score for the data object based on a set of one or more sensitive data identification techniques. The system determines a set of enrichment labels for the data object based on the sensitivity score computed for the data object. The enrichment labels are used to further qualify, enrich, or classify the data objects identified as containing sensitive information. For instance, the enrichment labels may identify a set of custom properties to be assigned to a data object, identify glossary terms to be applied to the data object or the enrichment labels may identify tags to be assigned to the data object.


