Metadata Labeling for Sensitive Data Lineage Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data integration systems lack efficient methods to identify, label, and track sensitive personal information across disparate computing networks and applications, particularly in cloud-based environments, which poses challenges for compliance with regulations like GDPR and increases the risk of data infringement.

Innovation Solution

A data integration protection assistance system that searches code instructions and metadata to identify sensitive data model field values, applies labels, and generates fieldname lineage maps to track the movement of sensitive information across different applications and locations, ensuring appropriate security measures are applied and compliance reports are generated.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual identification and tracking of sensitive data is performed, then data security can be maintained, but productivity and compliance efficiency deteriorate due to the complexity and scale of data integration processes

Engineering Contradiction:
Improvedata securityVSAvoidcompliance efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system enables self-service by automatically identifying, labeling, and tracking sensitive data through machine learning models that autonomously analyze metadata and code instructions, eliminating the need for manual security monitoring while maintaining high compliance efficiency

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical processes with automated computational systems, using machine learning algorithms to substitute human analysts in identifying and tracking sensitive information across data integration processes, thereby improving both security reliability and compliance productivity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If comprehensive tracking of sensitive information is implemented across all data integration processes, then compliance accuracy is improved, but device complexity and implementation difficulty worsen

Engineering Contradiction:
Improvecompliance accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-labeling data fields and establishing tracking mechanisms before data integration processes begin, using machine learning models to proactively identify sensitive information and set up monitoring workflows in advance, which simplifies subsequent compliance tracking

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces intermediary components such as metadata labels and tracking identifiers that mediate between raw data and compliance monitoring systems, enabling accurate compliance measurement without requiring direct complex analysis of all data streams

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If automated machine learning models are used to identify sensitive data, then productivity and compliance efficiency are improved, but measurement precision deteriorates due to challenges in identifying non-obvious sensitive information

Engineering Contradiction:
Improvecompliance efficiencyVSAvoididentification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system employs dynamic machine learning models that continuously learn and adapt from labeled data, improving their identification accuracy over time through retraining cycles, allowing the models to evolve their measurement precision while maintaining high productivity

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent implements feedback mechanisms where identified sensitive data is reviewed, labeled, and used to retrain the machine learning models, creating a continuous improvement loop that enhances identification accuracy while maintaining automated high-speed processing capability

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20240354297A1System and method of intelligent translation of metadata label names and mapping to natural language understanding
Publication Date: 2024.10.24 BOOMI LP
  • US20240354297A1 patent drawing
  • US20240354297A1 patent drawing
  • US20240354297A1 patent drawing

AI summary

An information handling system operating a data integration protection assistance system may comprise a processor linking first and second data set field names identified within a data integration process for transferring a data set field value identified by the first data field name at a source location to a destination location for storage under the second data field name. The processor may receive a user instruction to label data set field names incorporating a search term as sensitive private individual data, determine the first data set field name incorporates the search term and the second data set field name does not incorporate the search term, and label both the first and second data set field names as sensitive private individual data. A graphical user interface may display the first and second data set field names, to track migration of data set field values containing sensitive personal information, despite renaming.