Synthetic EHR Generation via Hierarchical Autoregressive Language Model
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for generating synthetic Electronic Health Records (EHRs) face challenges in producing high-dimensional, realistic, and privacy-preserving data, particularly due to limitations in existing generative AI models that struggle with high-dimensional and sparse EHR data, and the scarcity of labeled data for rare diseases and small demographic groups.
Innovation Solution
The Hierarchical Autoregressive Language Model (HALO) platform generates high-dimensional synthetic EHRs by using a multi-granularity approach to model binary sequences of over a million variables, preserving statistical properties and enabling the creation of labeled datasets, while ensuring privacy through the generation of uncorrelated data that cannot be re-identified.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional generative AI models are used to generate synthetic EHR data, then privacy is preserved through de-identification, but the data becomes low-dimensional and loses statistical properties of real high-dimensional EHR data
Solution Approach 1:
The patent creates synthetic EHR data that copies the statistical properties and high-dimensional structure of real EHR data while being uncorrelated to any individual patient. The generative model learns the underlying data distribution and generates new records that mirror real data characteristics without copying actual patient information, thus achieving both privacy preservation and statistical fidelity.
Solution Approach 2:
The patent transforms the approach by changing how EHR data is represented and generated. Instead of aggregating codes or removing rare codes to simplify data, the model maintains the full high-dimensional structure with over a million variables, preserving rare codes and detailed medical information while generating synthetic records that match the parameter distribution of real data.
2Reliability
If rule-based approaches are used to generate synthetic EHR data, then privacy is maintained through de-identification methods, but the capacity for realism and utility is limited
Solution Approach 1:
The patent replaces rule-based mechanical approaches with a deep learning generative model. Instead of applying predefined rules for de-identification and data synthesis, the system uses a neural network trained on real EHR data to learn complex patterns and generate realistic synthetic records, significantly improving the adaptability and utility of the generated data while maintaining privacy.
3Manufacturing precision
If GAN-based methods are used to generate EHR data, then realistic data can be produced, but the data is aggregated into one time step and loses longitudinal structure
Solution Approach 1:
The patent introduces temporal dynamics into the generative model by processing EHR data as sequential visits with time connections. The model generates synthetic records that maintain the longitudinal structure of patient journeys across multiple visits, capturing how health conditions evolve over time rather than treating each visit as an isolated static snapshot.
Solution Approach 2:
The patent adds the temporal dimension back into synthetic EHR generation. While GANs typically produce single-time-step outputs, this system generates multi-visit longitudinal records that preserve the time-based structure of healthcare data, enabling realistic simulation of patient trajectories across multiple encounters with the healthcare system.
4Reliability
If EHR data is de-identified to preserve privacy, then patient confidentiality is protected, but the utility of the data for machine learning development is reduced
Solution Approach 1:
The patent creates synthetic copies of EHR data that preserve all statistical properties and medical code structures needed for machine learning training. These synthetic records serve as faithful replicas for ML development purposes while being completely uncorrelated to any real patient, eliminating the need to de-identify actual patient data and thus preserving full data utility.
5Ease of manufacture
If rare codes are removed or aggregated to simplify EHR data, then data processing becomes easier, but the high-dimensional nature and sparsity of real EHR data is lost
Solution Approach 1:
The patent changes the approach to handling rare medical codes instead of removing or aggregating them. The generative model is trained to preserve the full high-dimensional structure including rare codes, maintaining the sparsity pattern characteristic of real EHR data. This allows the system to process complex high-dimensional data while accurately reflecting the true nature of medical records.
Data Source
AI summary
In one aspect, the present disclosure relates to a platform for creating synthetic electronic health records, the platform being configured to perform operations including receiving EHR data and encoding the received EHR data as a plurality of fixed length vectors to form a fixed-length matrix. The platform provides the fixed-length matrix to a machine learning model as input to produce a plurality of visit history representations. For one or more particular visit history representations, of the plurality of visit history representations, the platform applies code information associated with the particular visit history. One or more appended visit histories are provided to one or more masked linear layers to produce a probability matrix comprising probabilities for each code for each visit. The platform produces one or more synthetic EHRs based on repeated sequential generation of and sampling from the probability matrix.


