Synthetic EHR Generation via Hierarchical Autoregressive Language Model

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for generating synthetic Electronic Health Records (EHRs) face challenges in producing high-dimensional, realistic, and privacy-preserving data, particularly due to limitations in existing generative AI models that struggle with high-dimensional and sparse EHR data, and the scarcity of labeled data for rare diseases and small demographic groups.

Innovation Solution

The Hierarchical Autoregressive Language Model (HALO) platform generates high-dimensional synthetic EHRs by using a multi-granularity approach to model binary sequences of over a million variables, preserving statistical properties and enabling the creation of labeled datasets, while ensuring privacy through the generation of uncorrelated data that cannot be re-identified.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional generative AI models are used to generate synthetic EHR data, then privacy is preserved through de-identification, but the data becomes low-dimensional and loses statistical properties of real high-dimensional EHR data

Engineering Contradiction:
Improveprivacy preservationVSAvoiddata dimensionality and statistical fidelity
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

The patent creates synthetic EHR data that copies the statistical properties and high-dimensional structure of real EHR data while being uncorrelated to any individual patient. The generative model learns the underlying data distribution and generates new records that mirror real data characteristics without copying actual patient information, thus achieving both privacy preservation and statistical fidelity.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms the approach by changing how EHR data is represented and generated. Instead of aggregating codes or removing rare codes to simplify data, the model maintains the full high-dimensional structure with over a million variables, preserving rare codes and detailed medical information while generating synthetic records that match the parameter distribution of real data.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If rule-based approaches are used to generate synthetic EHR data, then privacy is maintained through de-identification methods, but the capacity for realism and utility is limited

Engineering Contradiction:
Improveprivacy protectionVSAvoidrealism and utility of synthetic data
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent replaces rule-based mechanical approaches with a deep learning generative model. Instead of applying predefined rules for de-identification and data synthesis, the system uses a neural network trained on real EHR data to learn complex patterns and generate realistic synthetic records, significantly improving the adaptability and utility of the generated data while maintaining privacy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Manufacturing precision

If GAN-based methods are used to generate EHR data, then realistic data can be produced, but the data is aggregated into one time step and loses longitudinal structure

Engineering Contradiction:
Improverealism of generated dataVSAvoidlongitudinal temporal structure
Core Design Contradiction:
Manufacturing precisionVSStability of the object's composition

Solution Approach 1:

The patent introduces temporal dynamics into the generative model by processing EHR data as sequential visits with time connections. The model generates synthetic records that maintain the longitudinal structure of patient journeys across multiple visits, capturing how health conditions evolve over time rather than treating each visit as an isolated static snapshot.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent adds the temporal dimension back into synthetic EHR generation. While GANs typically produce single-time-step outputs, this system generates multi-visit longitudinal records that preserve the time-based structure of healthcare data, enabling realistic simulation of patient trajectories across multiple encounters with the healthcare system.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Reliability

If EHR data is de-identified to preserve privacy, then patient confidentiality is protected, but the utility of the data for machine learning development is reduced

Engineering Contradiction:
Improvepatient confidentialityVSAvoiddata utility for ML training
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent creates synthetic copies of EHR data that preserve all statistical properties and medical code structures needed for machine learning training. These synthetic records serve as faithful replicas for ML development purposes while being completely uncorrelated to any real patient, eliminating the need to de-identify actual patient data and thus preserving full data utility.

Inventive Principle:
Principle #26Copying

5Ease of manufacture

If rare codes are removed or aggregated to simplify EHR data, then data processing becomes easier, but the high-dimensional nature and sparsity of real EHR data is lost

Engineering Contradiction:
Improvedata processing simplicityVSAvoidhigh-dimensional structure preservation
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent changes the approach to handling rare medical codes instead of removing or aggregating them. The generative model is trained to preserve the full high-dimensional structure including rare codes, maintaining the sparsity pattern characteristic of real EHR data. This allows the system to process complex high-dimensional data while accurately reflecting the true nature of medical records.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240105292A1Platform for synthesizing high-dimensional longitudinal electronic health records using a deep learning language model
Publication Date: 2024.03.28 MEDISYN INC
  • US20240105292A1 patent drawing
  • US20240105292A1 patent drawing
  • US20240105292A1 patent drawing

AI summary

In one aspect, the present disclosure relates to a platform for creating synthetic electronic health records, the platform being configured to perform operations including receiving EHR data and encoding the received EHR data as a plurality of fixed length vectors to form a fixed-length matrix. The platform provides the fixed-length matrix to a machine learning model as input to produce a plurality of visit history representations. For one or more particular visit history representations, of the plurality of visit history representations, the platform applies code information associated with the particular visit history. One or more appended visit histories are provided to one or more masked linear layers to produce a probability matrix comprising probabilities for each code for each visit. The platform produces one or more synthetic EHRs based on repeated sequential generation of and sampling from the probability matrix.