Metadata-Driven Data Minimization Without Reading Data Content
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data minimization methods require reading the actual content of large data pools, making them time-consuming, expensive, and inefficient, while also posing challenges in managing privacy and security in compliance with regulations like GDPR and CCPA.
Innovation Solution
A computing system and method that utilize a Machine Learning (ML) and Natural Language Processing (NLP) model to determine data characteristics, generate minimization parameters, and perform data minimization operations without reading the data content, using metadata and prestored rules to minimize datasets based on privacy regulations and business requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If selective cleaning is used to minimize data by reading actual content, then data privacy and security are ensured, but time consumption and cost increase significantly
Solution Approach 1:
The patent extracts and utilizes metadata from data pools without reading the actual data content. The system identifies and processes only the metadata portions that contain information about data characteristics, purposes, and sensitivity levels, thereby achieving data minimization goals while avoiding the time-consuming task of reading entire data sets.
Solution Approach 2:
The system performs preliminary classification and tagging of data during the metadata extraction phase, organizing data into categories based on sensitivity levels and purposes before any minimization operations are applied. This preliminary organization enables efficient subsequent processing without requiring repeated data reads.
2Measurement precision
If selective cleaning is used to minimize data by reading actual content, then accurate identification of sensitive data is achieved, but computational cost and resources increase
Solution Approach 1:
The patent introduces metadata as an intermediary layer between the data storage system and the minimization processing system. By analyzing metadata rather than actual data content, the system achieves accurate identification of sensitive data types and purposes while consuming significantly fewer computational resources.
3Productivity
If data minimization is performed without reading data content, then time and resources are saved, but ability to identify sensitive data may be compromised
Solution Approach 1:
The system performs preliminary classification and tagging of data during metadata extraction, organizing data into sensitivity categories and purpose groups before minimization operations. This advance organization ensures accurate identification of sensitive data while maintaining high processing speed.
Solution Approach 2:
The patent replaces the mechanical approach of reading and analyzing actual data content with a metadata-based analysis system. The metadata contains structured information about data characteristics, purposes, and sensitivity levels, enabling accurate classification without the need to read actual data content.
Data Source
AI summary
A system and method for performing data minimization without reading data content is disclosed. The method includes receiving a request from a user to perform data minimization and retrieving metadata associated with plurality of datasets based on the request. The method further includes determining one or more characteristics of the retrieved metadata based on one or more data parameters and one or more derived data parameters and generating one or more minimization parameters and one or more data sensitivity parameters for each of the plurality of datasets by using a trained data minimization based ML and NLP model. The method includes determining portions of the plurality of datasets based on the one or more minimization parameters, the one or more data sensitivity parameters, privacy regulations and business requirements and performing one or more minimizing operations on the determined portions of the plurality of datasets based on prestored rules.


