Sensitive Data Discovery via Metadata Machine Learning Pipelines
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large organizations face challenges in efficiently and comprehensively identifying sensitive data across vast databases due to the impracticality of deep profiling tools and inconsistencies in manual certification methods, leading to 'dark pools' of data and compliance risks.
Innovation Solution
A system utilizing metadata processing pipelines with machine learning techniques to automatically discover sensitive data by enriching, predicting, and mapping metadata categories, applying policy rules for risk classification and tagging, thereby reducing the dataset size and processing time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If deep profiling tools are used to comprehensively process large volumes of data, then measurement precision of sensitive data is improved, but loss of time increases significantly (taking over a decade to process petabytes of data)
Solution Approach 1:
The patent extracts and processes only metadata from data sources rather than profiling the entire dataset. The metadata processing pipeline harvests, enriches, and analyzes metadata to predict sensitive data categories, thereby avoiding the time-consuming task of processing petabytes of actual data while maintaining identification accuracy.
Solution Approach 2:
The patent segments the data processing task into two parts: (1) processing metadata to predict sensitive data categories and locations, and (2) using these predictions to guide selective profiling of only the identified sensitive data portions. This segmentation dramatically reduces processing time while maintaining measurement precision.
2Ease of operation
If manual certification methods are used by application owners, then ease of operation is improved, but reliability of sensitive data identification deteriorates due to inconsistencies and compliance risks
Solution Approach 1:
The system enables automated self-service for sensitive data identification through the metadata processing pipeline and machine learning models. The pipeline automatically harvests metadata, enriches it with additional context, predicts sensitive data categories, and generates compliance reports without requiring manual intervention, thereby ensuring consistency and reliability while maintaining ease of operation.
Solution Approach 2:
The system incorporates feedback mechanisms where the automated identification results can be reviewed and refined. The metadata enrichment process includes gathering feedback from multiple sources (data catalogs, data lakes, cloud storage) to continuously improve the accuracy and reliability of sensitive data identification while keeping the process automated and easy to operate.
3Productivity
If metadata processing pipelines with machine learning techniques are used, then productivity of sensitive data discovery is improved, but device complexity increases
Solution Approach 1:
The metadata processing pipeline is designed as a universal system that can handle multiple data sources (data catalogs, data lakes, cloud storage) and perform multiple functions (harvesting, enrichment, prediction, classification) through a single integrated architecture. This multi-functionality improves productivity without proportionally increasing complexity, as the same infrastructure serves multiple purposes.
Solution Approach 2:
The patent introduces metadata as an intermediary layer between the raw data sources and the analysis system. Instead of directly processing complex datasets, the system processes metadata that serves as a simplified representation, thereby improving productivity while managing device complexity through this intermediary abstraction layer.
4Loss of time
If automated metadata processing is implemented, then loss of time is reduced, but loss of information may increase due to potential inaccuracies in predicted categories
Solution Approach 1:
The system performs preliminary actions by processing metadata first to predict sensitive data categories and locations before conducting any actual data profiling. This preliminary metadata analysis guides subsequent targeted profiling, reducing overall processing time while maintaining accuracy through a two-stage approach that validates predictions against actual data characteristics.
Solution Approach 2:
The metadata enrichment process incorporates feedback from multiple data sources and validation mechanisms. The system cross-references metadata predictions with actual data characteristics, allowing for correction and refinement of predicted categories. This feedback loop ensures that time-saving automation does not compromise information accuracy, as inaccuracies can be identified and corrected through the feedback mechanism.
Data Source
AI summary
A method for auto discovery of sensitive data may include: (1) receiving, at data enrichment computer program in a metadata processing pipeline, raw metadata from a plurality of different data sources; (2) enriching, by the data enrichment computer program, the raw metadata; (3) converting, by the data enrichment computer program, the raw metadata and the enhanced raw metadata into a sentence structure; (4) predicting, by a category prediction computer program in the metadata processing pipeline, a predicted category for the sentence structure; (5) identifying, by a sensitive data mapping computer program, a sensitive data category that is mapped to the predicted category based on a policy mapping rule; (6) determining, by the sensitive data mapping computer program, a risk classification rating for the predicted category; and (7) tagging, by the sensitive data mapping computer program, the data source associated with the metadata based on the risk classification rating.


