Synthetic EHR Data Generation via Variational Autoencoder

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing EHR data analytics face challenges in utilizing time series data due to privacy concerns and limited data size, requiring synthetic data that maintains fidelity and expands in size while protecting original data privacy.

Innovation Solution

A method involving a variational autoencoder (VAE) that captures an original EHR dataset, reduces its dimensionality, and applies a stochastic process prior to generate synthetic data, ensuring it approximates the original while expanding its size, combined with a differential privacy mechanism to protect original data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If original EHR data is used for analytics, then data fidelity is maintained, but privacy concerns arise and data size remains limited

Engineering Contradiction:
Improvedata fidelityVSAvoidprivacy concerns
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent creates synthetic copies of EHR data through a generative model that learns the underlying distribution of original data. The variational autoencoder encodes original data into latent representations, which are then decoded to generate synthetic data copies that maintain statistical properties and fidelity while being detached from original patient identities, thus resolving the privacy concern while preserving data reliability

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms data parameters by converting original EHR data into latent space representations through encoding, then generating new data instances by sampling from the learned distribution. This parameter transformation allows the system to maintain data fidelity through distributional consistency while producing unlimited synthetic variations that address both privacy and data size limitations

Inventive Principle:
Principle #35Parameter changes

2Reliability

If original EHR data is used for analytics, then data quality is maintained, but data size is limited

Engineering Contradiction:
Improvedata qualityVSAvoiddata size
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent performs preliminary learning of the data distribution from the original EHR dataset before generation. The variational autoencoder is trained offline to capture the underlying patterns and relationships in the data, storing this knowledge in the latent space. This preliminary action enables subsequent unlimited generation of high-quality synthetic data without requiring additional original data, thus resolving the data size limitation while maintaining quality

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system generates multiple synthetic copies of the original data by sampling from the learned latent distribution. Each synthetic data instance is a unique copy that preserves the statistical properties and relationships of the original data, enabling researchers to obtain unlimited quantities of high-quality data for analysis while the original dataset remains unchanged

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12014293B2Electronic health record data synthesization
Publication Date: 2024.06.18 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12014293B2 patent drawing
  • US12014293B2 patent drawing
  • US12014293B2 patent drawing

AI summary

The present disclosure relates to a method, system and computer program product for electronic health record (EHR) data synthetization. According to the method, an original EHR dataset X is captured. A latent space Z is generated from the original EHR dataset X, wherein dimensionality of Z is lower than that of X. A stochastic process prior module is applied to the latent space Z. Synthetic EHR dataset X′ is reconstructed from the latent space Z after being applied with the stochastic process prior.