Synthetic X-ray Data Generation for Deep Learning Bias Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning models for processing medical imagery, particularly those with deep architectures, face issues such as overfitting and bias due to limited labeled clinical training cases and biased datasets, leading to poor generalization and prediction inaccuracies.

Innovation Solution

A training data modification system that modifies medical training X-ray imagery by adding or removing image structures and simulating X-ray properties, incorporating models of medical tools and anatomical structures, to create a more diverse and unbiased dataset, which is then used to train machine learning models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If more labeled clinical training cases are assembled to train machine learning models, then model generalization improves, but the process becomes laborious and expensive

Engineering Contradiction:
Improvemodel generalizationVSAvoiddata assembly time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent uses synthetic data generation to create artificial copies of medical imaging data that mimic real clinical cases. Instead of manually assembling labeled training cases, the system generates synthetic images with known ground truth labels, eliminating the need for time-consuming manual data collection and annotation while providing sufficient training samples for model generalization

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary data preparation by pre-generating large datasets of synthetic medical images with embedded ground truth information before actual model training begins. This preliminary action creates a ready-to-use training corpus that eliminates the need for time-consuming data assembly during the modeling process

Inventive Principle:
Principle #10Preliminary action

2Quantity of substance

If training data is collected from existing clinical cases, then data availability improves, but dataset bias increases leading to prediction inaccuracies

Engineering Contradiction:
Improvetraining data availabilityVSAvoidprediction accuracy
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent creates synthetic copies of medical imaging data that replicate the statistical properties and variability of real clinical data without inheriting its biases. These synthetic datasets provide diverse training examples that are not constrained by the limitations of existing clinical case collections, enabling more reliable and generalizable predictions

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent systematically varies parameters in synthetic data generation to create diverse training scenarios that explicitly address underrepresented cases. By controlling generation parameters, the system can balance class distributions and ensure adequate representation of rare conditions, thereby reducing dataset bias and improving prediction accuracy across all patient populations

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12254677B2System and methods for augmenting X-ray images for training of deep neural networks
Publication Date: 2025.03.18 KONINKLIJKE PHILIPS NV
  • US12254677B2 patent drawing
  • US12254677B2 patent drawing
  • US12254677B2 patent drawing

AI summary

A training data modification system (TDM) for machine learning and related methods. The system comprises a data modifier (DM) configured to perform a modification operation to modify medical training X-ray imagery of a patient. The modification operation causes image structures in the modified medical training imagery. The image structure is representative of a property of i) a medical procedure, ii) an image acquisition operation by an X-ray-based medical imaging apparatus (IA), iii) an anatomy of the patient.