Synthetic Data Augmentation for Medical Record ML Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The scarcity of parallel training datasets for machine learning models that generate electronic health records (EHRs) from transcriptions, due to privacy concerns and legal reasons, limits the effectiveness of these models.

Innovation Solution

The implementation of data augmentation techniques to generate synthetic training data from a limited existing dataset, using two 'teacher' machine learning models to create synthetic transcripts and medical records, which are then combined with the natural dataset to form a larger synthetic training dataset.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data augmentation is used to generate synthetic training data, then the quantity of training data increases, but the complexity of the training process increases

Engineering Contradiction:
Improvequantity of training dataVSAvoidcomplexity of training process
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-generating synthetic training data using teacher models before the actual student model training. The synthetic data is created in advance through data augmentation techniques, including generating synthetic transcriptions from medical records and synthetic medical records from transcriptions. This preliminary data preparation reduces the burden during the actual training phase, as the student model can directly learn from the pre-prepared augmented dataset without requiring complex real-time data generation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces teacher models as intermediaries between the limited real data and the student model. These teacher models generate synthetic training data that acts as an intermediary representation, bridging the gap between scarce real medical data and the large-scale training data needed for effective student model learning. The teacher models transform real medical records and transcriptions into synthetic counterparts that preserve medical knowledge while enabling data augmentation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Use of energy by moving object

If knowledge distillation is used to create a smaller student model, then computational resources required decrease, but training accuracy may be compromised

Engineering Contradiction:
Improvecomputational resources requiredVSAvoidtraining accuracy
Core Design Contradiction:
Use of energy by moving objectVSReliability

Solution Approach 1:

The patent applies copying by creating a student model that replicates the knowledge and functionality of larger teacher models. The student model is trained using knowledge distillation, where it learns from the softened probability distributions and predictions of the teacher models. This copying process enables the smaller student model to approximate the performance of larger models while requiring fewer computational resources for inference and deployment.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent employs parameter changes by modifying the training objectives and loss functions during knowledge distillation. The student model is trained with a combination of standard cross-entropy loss and distillation loss, which guides it to match the teacher model's output distributions. This parameter adjustment in the training process enables the smaller model to capture the essential knowledge from larger models, maintaining accuracy while reducing model size and computational requirements.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12266434B2System and method for inverse summarization of medical records with data augmentation and knowledge distillation
Publication Date: 2025.04.01 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12266434B2 patent drawing
  • US12266434B2 patent drawing
  • US12266434B2 patent drawing

AI summary

A method, computer program product, and computing system for generating a first synthetic dataset including a synthetic transcription and a corresponding natural dictation record using a first machine learning model trained to generate transcriptions from medical records. A second synthetic dataset including a synthetic medical record and a corresponding natural transcription is generated using a second machine learning model trained to generate medical records from transcriptions. The first synthetic dataset and the second synthetic dataset are combined with a natural dataset into a synthetic training dataset.