Synthetic Medical Data Generation for AI Development
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The exchange of medical data across organizational boundaries is restricted due to data protection regulations, making it difficult to develop and validate artificial intelligence systems, as anonymization is time-consuming and risky, and access to valuable datasets is limited, hindering research and development.
Innovation Solution
A method for creating a synthetic dataset from a medical dataset using a sampling function that replaces original values, allowing for the local creation and transfer of synthetic datasets across facilities while maintaining data protection compliance, enabling the utilization of medical data for AI development without revealing personal information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If medical data is exchanged across organizational boundaries, then AI system development and validation can proceed, but data protection regulations are violated and patient privacy is compromised
Solution Approach 1:
The patent creates synthetic copies of medical data that replicate the statistical characteristics and patterns of real patient data without containing actual personal information. These synthetic datasets can be freely exchanged and used for AI development while maintaining data protection compliance, as they are artificial reproductions rather than copies of sensitive personal data.
Solution Approach 2:
The patent introduces synthetic data as an intermediary medium between the original medical data and the AI development process. This intermediary allows information to be transferred for research purposes without directly exposing sensitive patient data, thus mediating between the need for data access and the requirement for privacy protection.
2Reliability
If real medical data is anonymized before exchange, then data protection is improved, but the anonymization process is time-consuming and costly
Solution Approach 1:
The patent performs anonymization in advance by generating synthetic data that is inherently anonymized. Instead of taking real data and removing personal information through time-consuming processes, the system pre-creates synthetic representations that never contained personal information to begin with, eliminating the need for subsequent anonymization efforts.
Solution Approach 2:
Rather than modifying real data through anonymization, the patent creates synthetic copies that capture the essential statistical properties without containing any personal identifiers. This copying approach bypasses the time-consuming anonymization process entirely while achieving the same data protection goals.
3Reliability
If real medical data is pseudonymized before exchange, then data protection is improved, but the risk of re-identification remains and significant legal and financial consequences can occur
Solution Approach 1:
The patent creates synthetic copies of medical data that replicate patterns and statistical characteristics without preserving any link to real individuals. Since these are artificial reproductions generated through sampling functions rather than modified versions of real data, the re-identification risk that plagues pseudonymization is completely eliminated.
Solution Approach 2:
The patent transforms the limitation of not having access to real data into a benefit by generating synthetic data that provides all the necessary statistical properties for AI development while inherently eliminating the re-identification risk. The very act of creating synthetic rather than modified data becomes the solution to the re-identification problem.
4Reliability
If access to medical data is limited to the duration of collaboration, then data security is improved, but sustained development and validation of AI systems becomes difficult
Solution Approach 1:
The patent creates synthetic datasets that can be freely distributed and retained by multiple parties without time limitations. These synthetic copies serve as permanent, reusable resources for AI development and validation, eliminating the need for ongoing access agreements or collaboration duration restrictions while maintaining data security.
Solution Approach 2:
The patent enables parties to discard the need for continuous access to original data by retaining synthetic datasets locally. The synthetic data can be stored indefinitely and reused multiple times for different development and validation activities without requiring continued access to the source facility or ongoing collaboration agreements.
Data Source
AI summary
Methods and apparatuses are for a medical dataset stored locally within a first facility and including a number of original individual datasets assigned to real existing patients and including original values for one or more higher-ranking variables. An embodiment of the method includes creation of a synthetic dataset based on the medical dataset, the synthetic dataset including a number of synthetic individual datasets including synthetic values for the same higher-ranking variables as the medical dataset, not relatable to an original existing patient, the creation being undertaken locally within the first facility by application of a sampling function to the medical data; and transfer of the synthetic dataset from the first facility to a central unit outside the first facility. The synthetic dataset is utilizable within the central unit.


