Sensitive Data Discovery via Metadata Machine Learning Pipelines

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large organizations face challenges in efficiently and comprehensively identifying sensitive data across vast databases due to the impracticality of deep profiling tools and inconsistencies in manual certification methods, leading to 'dark pools' of data and compliance risks.

Innovation Solution

A system utilizing metadata processing pipelines with machine learning techniques to automatically discover sensitive data by enriching, predicting, and mapping metadata categories, applying policy rules for risk classification and tagging, thereby reducing the dataset size and processing time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If deep profiling tools are used to comprehensively process large volumes of data, then measurement precision of sensitive data is improved, but loss of time increases significantly (taking over a decade to process petabytes of data)

Engineering Contradiction:
Improvesensitive data identification accuracyVSAvoiddata processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts and processes only metadata from data sources rather than profiling the entire dataset. The metadata processing pipeline harvests, enriches, and analyzes metadata to predict sensitive data categories, thereby avoiding the time-consuming task of processing petabytes of actual data while maintaining identification accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the data processing task into two parts: (1) processing metadata to predict sensitive data categories and locations, and (2) using these predictions to guide selective profiling of only the identified sensitive data portions. This segmentation dramatically reduces processing time while maintaining measurement precision.

Inventive Principle:
Principle #1Segmentation

2Ease of operation

If manual certification methods are used by application owners, then ease of operation is improved, but reliability of sensitive data identification deteriorates due to inconsistencies and compliance risks

Engineering Contradiction:
Improvedata certification simplicityVSAvoidsensitive data identification consistency
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system enables automated self-service for sensitive data identification through the metadata processing pipeline and machine learning models. The pipeline automatically harvests metadata, enriches it with additional context, predicts sensitive data categories, and generates compliance reports without requiring manual intervention, thereby ensuring consistency and reliability while maintaining ease of operation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system incorporates feedback mechanisms where the automated identification results can be reviewed and refined. The metadata enrichment process includes gathering feedback from multiple sources (data catalogs, data lakes, cloud storage) to continuously improve the accuracy and reliability of sensitive data identification while keeping the process automated and easy to operate.

Inventive Principle:
Principle #23Feedback

3Productivity

If metadata processing pipelines with machine learning techniques are used, then productivity of sensitive data discovery is improved, but device complexity increases

Engineering Contradiction:
Improvesensitive data discovery speedVSAvoidsystem architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The metadata processing pipeline is designed as a universal system that can handle multiple data sources (data catalogs, data lakes, cloud storage) and perform multiple functions (harvesting, enrichment, prediction, classification) through a single integrated architecture. This multi-functionality improves productivity without proportionally increasing complexity, as the same infrastructure serves multiple purposes.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces metadata as an intermediary layer between the raw data sources and the analysis system. Instead of directly processing complex datasets, the system processes metadata that serves as a simplified representation, thereby improving productivity while managing device complexity through this intermediary abstraction layer.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Loss of time

If automated metadata processing is implemented, then loss of time is reduced, but loss of information may increase due to potential inaccuracies in predicted categories

Engineering Contradiction:
Improvedata processing durationVSAvoiddata classification accuracy
Core Design Contradiction:
Loss of timeVSLoss of information

Solution Approach 1:

The system performs preliminary actions by processing metadata first to predict sensitive data categories and locations before conducting any actual data profiling. This preliminary metadata analysis guides subsequent targeted profiling, reducing overall processing time while maintaining accuracy through a two-stage approach that validates predictions against actual data characteristics.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The metadata enrichment process incorporates feedback from multiple data sources and validation mechanisms. The system cross-references metadata predictions with actual data characteristics, allowing for correction and refinement of predicted categories. This feedback loop ensures that time-saving automation does not compromise information accuracy, as inaccuracies can be identified and corrected through the feedback mechanism.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12373592B2Systems and methods for auto discovery of sensitive data in applications or databases using metadata via machine learning techniques
Publication Date: 2025.07.29 JPMORGAN CHASE BANK NA
  • US12373592B2 patent drawing
  • US12373592B2 patent drawing
  • US12373592B2 patent drawing

AI summary

A method for auto discovery of sensitive data may include: (1) receiving, at data enrichment computer program in a metadata processing pipeline, raw metadata from a plurality of different data sources; (2) enriching, by the data enrichment computer program, the raw metadata; (3) converting, by the data enrichment computer program, the raw metadata and the enhanced raw metadata into a sentence structure; (4) predicting, by a category prediction computer program in the metadata processing pipeline, a predicted category for the sentence structure; (5) identifying, by a sensitive data mapping computer program, a sensitive data category that is mapped to the predicted category based on a policy mapping rule; (6) determining, by the sensitive data mapping computer program, a risk classification rating for the predicted category; and (7) tagging, by the sensitive data mapping computer program, the data source associated with the metadata based on the risk classification rating.