Synthetic Medical Data Generation for AI Development

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The exchange of medical data across organizational boundaries is restricted due to data protection regulations, making it difficult to develop and validate artificial intelligence systems, as anonymization is time-consuming and risky, and access to valuable datasets is limited, hindering research and development.

Innovation Solution

A method for creating a synthetic dataset from a medical dataset using a sampling function that replaces original values, allowing for the local creation and transfer of synthetic datasets across facilities while maintaining data protection compliance, enabling the utilization of medical data for AI development without revealing personal information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If medical data is exchanged across organizational boundaries, then AI system development and validation can proceed, but data protection regulations are violated and patient privacy is compromised

Engineering Contradiction:
ImproveAI system developmentVSAvoiddata protection compliance
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent creates synthetic copies of medical data that replicate the statistical characteristics and patterns of real patient data without containing actual personal information. These synthetic datasets can be freely exchanged and used for AI development while maintaining data protection compliance, as they are artificial reproductions rather than copies of sensitive personal data.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces synthetic data as an intermediary medium between the original medical data and the AI development process. This intermediary allows information to be transferred for research purposes without directly exposing sensitive patient data, thus mediating between the need for data access and the requirement for privacy protection.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If real medical data is anonymized before exchange, then data protection is improved, but the anonymization process is time-consuming and costly

Engineering Contradiction:
Improvedata protectionVSAvoidanonymization process
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs anonymization in advance by generating synthetic data that is inherently anonymized. Instead of taking real data and removing personal information through time-consuming processes, the system pre-creates synthetic representations that never contained personal information to begin with, eliminating the need for subsequent anonymization efforts.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Rather than modifying real data through anonymization, the patent creates synthetic copies that capture the essential statistical properties without containing any personal identifiers. This copying approach bypasses the time-consuming anonymization process entirely while achieving the same data protection goals.

Inventive Principle:
Principle #26Copying

3Reliability

If real medical data is pseudonymized before exchange, then data protection is improved, but the risk of re-identification remains and significant legal and financial consequences can occur

Engineering Contradiction:
Improvedata protectionVSAvoidre-identification risk
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent creates synthetic copies of medical data that replicate patterns and statistical characteristics without preserving any link to real individuals. Since these are artificial reproductions generated through sampling functions rather than modified versions of real data, the re-identification risk that plagues pseudonymization is completely eliminated.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms the limitation of not having access to real data into a benefit by generating synthetic data that provides all the necessary statistical properties for AI development while inherently eliminating the re-identification risk. The very act of creating synthetic rather than modified data becomes the solution to the re-identification problem.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

4Reliability

If access to medical data is limited to the duration of collaboration, then data security is improved, but sustained development and validation of AI systems becomes difficult

Engineering Contradiction:
Improvedata securityVSAvoiddata availability
Core Design Contradiction:
ReliabilityVSDuration of action of stationary object

Solution Approach 1:

The patent creates synthetic datasets that can be freely distributed and retained by multiple parties without time limitations. These synthetic copies serve as permanent, reusable resources for AI development and validation, eliminating the need for ongoing access agreements or collaboration duration restrictions while maintaining data security.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent enables parties to discard the need for continuous access to original data by retaining synthetic datasets locally. The synthetic data can be stored indefinitely and reused multiple times for different development and validation activities without requiring continued access to the source facility or ongoing collaboration agreements.

Inventive Principle:
Principle #34Discarding and recovering

Data Source

PatentUS20220068446A1Utilization of medical data across organizational boundaries
Publication Date: 2022.03.03 SIEMENS HEALTHINEERS AG
  • US20220068446A1 patent drawing
  • US20220068446A1 patent drawing
  • US20220068446A1 patent drawing

AI summary

Methods and apparatuses are for a medical dataset stored locally within a first facility and including a number of original individual datasets assigned to real existing patients and including original values for one or more higher-ranking variables. An embodiment of the method includes creation of a synthetic dataset based on the medical dataset, the synthetic dataset including a number of synthetic individual datasets including synthetic values for the same higher-ranking variables as the medical dataset, not relatable to an original existing patient, the creation being undertaken locally within the first facility by application of a sampling function to the medical data; and transfer of the synthetic dataset from the first facility to a central unit outside the first facility. The synthetic dataset is utilizable within the central unit.