Data Obfuscation Preserving Predictive Relationships
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems for protecting personally identifiable information (PII) fail to preserve the structure necessary for certain machines or prediction models to utilize the data effectively, limiting the ability to safely transmit and train on such data.
Innovation Solution
A system and method that involves retrieving PII datasets, identifying security parameters, applying a random permutation, and generating obfuscated data to maintain local predictive relationships, allowing safe transmission and training of machine learning algorithms while protecting PII.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional encryption and data protection techniques are applied to personally identifiable information, then security and privacy protection are improved, but the data structure and predictive relationships are lost making the data unusable for machine learning
Solution Approach 1:
The patent segments the data protection process into multiple stages: original data is divided into training data and validation data, then transformed into obfuscated versions while preserving statistical properties. This segmentation allows different portions of data to serve different purposes - some for maintaining security, others for preserving predictive relationships.
Solution Approach 2:
The patent applies parameter changes by transforming the data through obfuscation techniques that modify specific parameters while preserving others. The obfuscation process changes the raw data parameters but maintains statistical parameters like mean, variance, and correlation structures, enabling machine learning while protecting privacy.
2Reliability
If data is obfuscated to protect PII, then privacy protection is improved, but the ability of prediction models to utilize the data effectively deteriorates
Solution Approach 1:
The patent creates obfuscated copies of the original data that preserve the essential statistical properties and predictive relationships needed for machine learning. These copies serve as substitutes for the original PII data, allowing third parties to train models without accessing actual personal information.
Solution Approach 2:
The obfuscation system serves multiple functions simultaneously: it protects privacy by obscuring PII, preserves predictive relationships for machine learning, and enables third-party access to data. This multi-functionality resolves the contradiction between privacy protection and model training effectiveness.
3Reliability
If strict data protection protocols are implemented, then security compliance is improved, but data transmission and sharing with third parties is restricted
Solution Approach 1:
The patent introduces obfuscated data as an intermediary between the original PII data and third-party prediction services. This intermediary form maintains security compliance by not exposing actual PII while still enabling data sharing and transmission to external parties for model training.
Data Source
AI summary
The invention relates to obfuscating data while maintaining local predictive relationships. An embodiment of the present invention is directed to cryptographically obfuscating a data set in a manner that hides personally identifiable information (PII) while allowing third parties to train classes of machine learning algorithms effectively. According to an embodiment of the present invention, the obfuscation acts as a symmetric encryption so that the original obfuscating party may relate the predictions on the obfuscated data to the original PII. The various features of the present invention enable third party prediction services to safely interact with PII.


