Synthetic PHI Generation for Machine Learning Privacy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for de-identifying and anonymizing private health information (PHI) often strip away crucial elements, rendering the data useless for clinical or functional applications, while public computing environments used by machine learning models pose security risks for PHI.

Innovation Solution

Generating synthetic PHI parameters that are statistically similar to original PHI, allowing these parameters to be used in user requests to machine learning models, ensuring the original data remains confidential and the de-identified data remains useful.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If existing de-identification techniques are applied to PHI, then confidentiality is improved, but clinical applicability deteriorates

Engineering Contradiction:
ImproveconfidentialityVSAvoidclinical applicability
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent creates synthetic copies of PHI data that replicate the statistical properties and clinical characteristics of original data without containing actual personal information. These synthetic datasets serve as substitutes for real PHI in machine learning training, maintaining clinical utility while eliminating privacy risks associated with the original data

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system transforms original PHI parameters into synthetic parameters by modifying key attributes such as replacing actual patient identifiers with synthetic ones, while preserving the statistical distributions and relationships between variables. This parameter transformation maintains the clinical relevance of the data for training purposes

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If original PHI is used for training machine learning models, then model accuracy is improved, but security risks worsen

Engineering Contradiction:
Improvemodel accuracyVSAvoidsecurity risks
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent introduces synthetic data as an intermediary between the original PHI and the machine learning model. This intermediary layer allows the model to learn from data that closely resembles real PHI in terms of statistical properties and clinical relationships, while the actual sensitive information never directly contacts the model or public computing environments

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

Synthetic copies of PHI are generated to serve as training data, replacing the need to use original sensitive data. These copies maintain the essential patterns and relationships needed for accurate model training while eliminating the security vulnerabilities associated with handling real PHI in public or cloud-based computing environments

Inventive Principle:
Principle #26Copying

3Reliability

If PHI is fully anonymized, then privacy protection is improved, but data usefulness deteriorates

Engineering Contradiction:
Improveprivacy protectionVSAvoiddata usefulness
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system applies selective parameter changes to PHI, transforming only the identifying elements while preserving the clinical and statistical parameters that give the data its usefulness. This partial transformation approach maintains privacy protection while retaining the data's adaptability for clinical research and machine learning applications

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240362363A1Systems and methods for anonymizing private data for use in machine learning models
Publication Date: 2024.10.31 SYNTHPOP INC
  • US20240362363A1 patent drawing
  • US20240362363A1 patent drawing
  • US20240362363A1 patent drawing

AI summary

Described herein are techniques for performing a task using a trained external machine learning model. In some embodiments, a user request comprising a task and one or more protected health information (PHI) parameters associated with a patient may be received. Using one or more trained machine learning models, the PHI parameters may be extracted and a revised user request may be generated by replacing the PHI parameters with one or more synthetic PHI parameters. The revised user request may be provided to the trained external machine learning model and a response to the task comprising the synthetic PHI parameters may be received. The trained machine learning models may generate a revised response to the task by replacing the synthetic PHI parameters with the PHI parameters of the user request.