Metadata Labeling for Sensitive Data Lineage Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data integration systems lack efficient methods to identify, label, and track sensitive personal information across disparate computing networks and applications, particularly in cloud-based environments, which poses challenges for compliance with regulations like GDPR and increases the risk of data infringement.
Innovation Solution
A data integration protection assistance system that searches code instructions and metadata to identify sensitive data model field values, applies labels, and generates fieldname lineage maps to track the movement of sensitive information across different applications and locations, ensuring appropriate security measures are applied and compliance reports are generated.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual identification and tracking of sensitive data is performed, then data security can be maintained, but productivity and compliance efficiency deteriorate due to the complexity and scale of data integration processes
Solution Approach 1:
The system enables self-service by automatically identifying, labeling, and tracking sensitive data through machine learning models that autonomously analyze metadata and code instructions, eliminating the need for manual security monitoring while maintaining high compliance efficiency
Solution Approach 2:
The patent replaces manual mechanical processes with automated computational systems, using machine learning algorithms to substitute human analysts in identifying and tracking sensitive information across data integration processes, thereby improving both security reliability and compliance productivity
2Measurement precision
If comprehensive tracking of sensitive information is implemented across all data integration processes, then compliance accuracy is improved, but device complexity and implementation difficulty worsen
Solution Approach 1:
The system performs preliminary actions by pre-labeling data fields and establishing tracking mechanisms before data integration processes begin, using machine learning models to proactively identify sensitive information and set up monitoring workflows in advance, which simplifies subsequent compliance tracking
Solution Approach 2:
The patent introduces intermediary components such as metadata labels and tracking identifiers that mediate between raw data and compliance monitoring systems, enabling accurate compliance measurement without requiring direct complex analysis of all data streams
3Productivity
If automated machine learning models are used to identify sensitive data, then productivity and compliance efficiency are improved, but measurement precision deteriorates due to challenges in identifying non-obvious sensitive information
Solution Approach 1:
The system employs dynamic machine learning models that continuously learn and adapt from labeled data, improving their identification accuracy over time through retraining cycles, allowing the models to evolve their measurement precision while maintaining high productivity
Solution Approach 2:
The patent implements feedback mechanisms where identified sensitive data is reviewed, labeled, and used to retrain the machine learning models, creating a continuous improvement loop that enhances identification accuracy while maintaining automated high-speed processing capability
Data Source
AI summary
An information handling system operating a data integration protection assistance system may comprise a processor linking first and second data set field names identified within a data integration process for transferring a data set field value identified by the first data field name at a source location to a destination location for storage under the second data field name. The processor may receive a user instruction to label data set field names incorporating a search term as sensitive private individual data, determine the first data set field name incorporates the search term and the second data set field name does not incorporate the search term, and label both the first and second data set field names as sensitive private individual data. A graphical user interface may display the first and second data set field names, to track migration of data set field values containing sensitive personal information, despite renaming.


