Deidentified Content Hashing for Privacy-Compliant Model Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing medical data privacy laws and contractual limitations restrict the long-term storage and usage of user data for training machine learning models, necessitating a secure method to deidentify personal information while allowing model training.
Innovation Solution
A two-step hashing process using a first hashing algorithm to generate deidentified content and a second hashing algorithm to create surrogates, enabling secure sharing and training of machine learning models without revealing personal information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If personal information is stored and used for training machine learning models, then model training accuracy is improved, but data privacy compliance deteriorates
Solution Approach 1:
The patent introduces hash values as an intermediary representation between personal information and machine learning models. The hashing function transforms personal information into hash values that preserve statistical properties for model training while preventing direct access to original data, thus serving as a mediator that satisfies both training accuracy and privacy compliance requirements
Solution Approach 2:
The patent creates a copied representation of personal information through hashing. The hash values are mathematical transformations that replicate the statistical characteristics needed for model training without being the original personal information, allowing models to learn from data patterns while the original sensitive information remains protected
2Reliability
If personal information is deidentified using hashing, then data privacy compliance is improved, but model training effectiveness deteriorates
Solution Approach 1:
The patent changes the parameter representation of personal information by transforming it into hash values through a mathematical function. This parameter transformation maintains the statistical properties necessary for model training while ensuring privacy compliance, as the hash values retain distribution characteristics without exposing original data
3Measurement precision
If user data is stored long-term for model training, then model accuracy is improved, but data security risks worsen
Solution Approach 1:
The patent uses hash values as an intermediary storage form instead of storing original personal information. This intermediary representation allows long-term storage for model training purposes while minimizing security risks, as the hashed data cannot be easily reversed to obtain original sensitive information even if storage systems are compromised
Data Source
AI summary
A method, computer program product, and computing system for processing raw content to identify personal information; replacing the personal information within the raw content with a first mathematical representation of the personal information generated using a first hashing algorithm, thus defining deidentified content; and defining a selected surrogate for the first mathematical representation, wherein the selected surrogate has a second mathematical representation generated using a second hashing algorithm that is equivalent to the first mathematical representation of the personal information that was generated using the first hashing algorithm.


