Synthetic PHI Generation for Machine Learning Privacy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for de-identifying and anonymizing private health information (PHI) often strip away crucial elements, rendering the data useless for clinical or functional applications, while public computing environments used by machine learning models pose security risks for PHI.
Innovation Solution
Generating synthetic PHI parameters that are statistically similar to original PHI, allowing these parameters to be used in user requests to machine learning models, ensuring the original data remains confidential and the de-identified data remains useful.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If existing de-identification techniques are applied to PHI, then confidentiality is improved, but clinical applicability deteriorates
Solution Approach 1:
The patent creates synthetic copies of PHI data that replicate the statistical properties and clinical characteristics of original data without containing actual personal information. These synthetic datasets serve as substitutes for real PHI in machine learning training, maintaining clinical utility while eliminating privacy risks associated with the original data
Solution Approach 2:
The system transforms original PHI parameters into synthetic parameters by modifying key attributes such as replacing actual patient identifiers with synthetic ones, while preserving the statistical distributions and relationships between variables. This parameter transformation maintains the clinical relevance of the data for training purposes
2Measurement precision
If original PHI is used for training machine learning models, then model accuracy is improved, but security risks worsen
Solution Approach 1:
The patent introduces synthetic data as an intermediary between the original PHI and the machine learning model. This intermediary layer allows the model to learn from data that closely resembles real PHI in terms of statistical properties and clinical relationships, while the actual sensitive information never directly contacts the model or public computing environments
Solution Approach 2:
Synthetic copies of PHI are generated to serve as training data, replacing the need to use original sensitive data. These copies maintain the essential patterns and relationships needed for accurate model training while eliminating the security vulnerabilities associated with handling real PHI in public or cloud-based computing environments
3Reliability
If PHI is fully anonymized, then privacy protection is improved, but data usefulness deteriorates
Solution Approach 1:
The system applies selective parameter changes to PHI, transforming only the identifying elements while preserving the clinical and statistical parameters that give the data its usefulness. This partial transformation approach maintains privacy protection while retaining the data's adaptability for clinical research and machine learning applications
Data Source
AI summary
Described herein are techniques for performing a task using a trained external machine learning model. In some embodiments, a user request comprising a task and one or more protected health information (PHI) parameters associated with a patient may be received. Using one or more trained machine learning models, the PHI parameters may be extracted and a revised user request may be generated by replacing the PHI parameters with one or more synthetic PHI parameters. The revised user request may be provided to the trained external machine learning model and a response to the task comprising the synthetic PHI parameters may be received. The trained machine learning models may generate a revised response to the task by replacing the synthetic PHI parameters with the PHI parameters of the user request.


