Synthetic Training Data Generation for ML Generalizability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning models face limitations in generalizability due to inadequate training data quality, volume, variety, and velocity, leading to poor performance in real-world scenarios.

Innovation Solution

A system that generates synthetic training data by augmenting annotated source images through element insertion, modality variation, and geometric transformations, creating diverse and voluminous training images to enhance model generalizability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If real-world annotated training data is collected and used, then model training accuracy is improved, but data volume and variety are limited

Engineering Contradiction:
Improvetraining data qualityVSAvoidtraining data volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system creates synthetic copies of annotated training data by generating new images with inserted elements of interest and background elements. These synthetic copies preserve the annotation quality of original data while multiplying the available training volume through automated generation processes

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system varies multiple parameters in generated training data including element positions, sizes, orientations, background configurations, and image characteristics. This parameter variation maintains data quality while exponentially increasing data variety and volume through controlled transformations

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If real-world annotated training data is collected, then data variety is improved, but data collection time and cost increase

Engineering Contradiction:
Improvetraining data varietyVSAvoiddata collection time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary data preparation by pre-defining libraries of elements of interest and background elements with various parameters. This preliminary setup enables rapid generation of diverse training data without time-consuming real-world collection and annotation processes

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system replicates diverse training scenarios through synthetic generation rather than physical collection. By copying and transforming existing annotated data with varied parameter combinations, the system achieves high data variety instantaneously without extended collection timelines

Inventive Principle:
Principle #26Copying

3Reliability

If more training data is generated through augmentation, then model generalizability is improved, but system complexity increases

Engineering Contradiction:
Improvemodel generalizabilityVSAvoiddata generation system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system segments the data generation process into distinct modular components: element insertion module, background generation module, parameter variation module, and annotation preservation module. This segmentation manages complexity by making each component independent and interchangeable while collectively achieving high model generalizability

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11720647B2Synthetic training data generation for improved machine learning model generalizability
Publication Date: 2023.08.08 GE PRECISION HEALTHCARE LLC
  • US11720647B2 patent drawing
  • US11720647B2 patent drawing
  • US11720647B2 patent drawing

AI summary

Systems and techniques that facilitate synthetic training data generation for improved machine learning generalizability are provided. In various embodiments, an element augmentation component can generate a set of preliminary annotated training images based on an annotated source image. In various aspects, a preliminary annotated training image can be formed by inserting at least one element of interest or at least one background element into the annotated source image. In various instances, a modality augmentation component can generate a set of intermediate annotated training images based on the set of preliminary annotated training images. In various cases, an intermediate annotated training image can be formed by varying at least one modality-based characteristic of a preliminary annotated training image. In various aspects, a geometry augmentation component can generate a set of deployable annotated training images based on the set of intermediate annotated training images. In various instances, a deployable annotated training image can be formed by varying at least one geometric characteristic of an intermediate annotated training image. In various embodiments, a training component can train a machine learning model on the set of deployable annotated training images.