Obfuscated Data Generation for Privacy-Preserving Model Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data processing methods for training machine learning models face challenges in maintaining data utility while ensuring privacy and security, particularly with sensitive information, as traditional anonymization techniques often degrade data quality and fail to prevent re-identification.
Innovation Solution
An apparatus and method using generative machine learning models to generate obfuscated data elements by clustering and selecting subsets based on distance measures and clustering algorithms, ensuring data utility and compliance with privacy standards.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional anonymization techniques are used to protect privacy, then security and confidentiality are improved, but data utility deteriorates and machine learning model performance is impaired
Solution Approach 1:
The patent generates synthetic data copies that replicate the statistical properties and patterns of original data without containing actual sensitive information. These synthetic copies maintain data utility for machine learning training while eliminating privacy risks associated with real data
Solution Approach 2:
The system transforms data by changing its representation parameters through generative models, converting real data into synthetic data with different statistical parameters that preserve structural relationships but eliminate identifiable information
2Ease of manufacture
If traditional anonymization techniques are applied to complex data such as images or audio, then processing is simpler, but re-identification becomes possible due to extensive information content
Solution Approach 1:
For complex data types like images and audio, the system creates synthetic copies that capture essential patterns and relationships without preserving identifiable features. The generative models learn from original data to produce realistic-looking synthetic data that cannot be traced back to specific individuals
Solution Approach 2:
The patent transforms complex high-dimensional data into synthetic representations by learning underlying probability distributions. This dimensional transformation allows the data to be processed and analyzed while maintaining privacy through the generative model's abstraction layer
3Loss of information
If real-world data is used directly for training machine learning models, then data quality and diversity are maximized, but privacy breaches and unauthorized access occur
Solution Approach 1:
The system replaces real data with synthetic data copies that maintain the quality and diversity needed for effective machine learning training. These copies preserve statistical properties, relationships, and patterns while eliminating the privacy breach risks of using actual sensitive data
Solution Approach 2:
The generative model acts as an intermediary between the original sensitive data and the machine learning training process. It receives real data for learning patterns but outputs only synthetic data for training, mediating the interaction to protect privacy while maintaining data utility
Data Source
AI summary
An apparatus for generating obfuscated data within a computing environment, comprising a processor and a memory containing instructions configuring the processor to access a database containing a plurality of private data elements belonging to at least a private record, generate a set of obfuscated data elements, representative of the at least a private record, as a function of the plurality of private data elements using an generative model, determine a first distance measure between at least an obfuscated data element within the set of obfuscated data elements and at least a private data element of the plurality of private data elements within the database, and verify the first distance measure is within a distance range, wherein a minimum threshold of the distance range is determined as a function of a deidentification parameter and a maximum threshold of the distance range is determined as a function of an obfuscation parameter.


