Synthetic Contrasted X-Ray Data Generation for Model Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The limited availability and diversity of contrasted x-ray-based image data, particularly in computed tomography angiography, pose challenges for training machine learning algorithms and diagnostic tools, leading to issues like overfitting and domain gap vulnerabilities.
Innovation Solution
A method involving a trained generative adversarial neural network synthesizes combined x-ray-based image data by segmenting non-contrasted images, generating a region-of-interest mask, and applying a modelled anatomical structure to create contrasted and non-contrasted regions, leveraging non-contrasted images as a high-quality background.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If contrasted x-ray-based image data are acquired and used for training machine learning algorithms, then the model can learn from real pathological cases, but the amount of available data is limited and lacks diversity
Solution Approach 1:
The patent creates synthetic copies of contrasted image data by combining non-contrasted images with segmented anatomical structures and contrast agent simulations. This generates artificial training data that mimics real contrasted images without requiring additional actual medical scans, thereby increasing data quantity while maintaining training quality
Solution Approach 2:
The patent performs preliminary segmentation of anatomical structures and preparation of contrast agent models before generating the final synthetic images. This pre-processing enables efficient batch generation of diverse training data with controlled pathological variations, addressing both data quantity and diversity requirements
2Productivity
If a limited dataset is used for training, then the training process is faster and requires less computational resources, but the model suffers from overfitting and poor generalization
Solution Approach 1:
By generating multiple synthetic copies of training data with varied pathological presentations and anatomical variations, the patent creates a larger effective dataset that improves model generalization without requiring proportionally more computational resources, as the synthesis process is more efficient than acquiring and processing additional real medical images
Solution Approach 2:
The patent varies parameters such as contrast agent concentration, anatomical structure dimensions, and pathological features in the synthetic data generation process. This parameter diversification enables the model to learn more robust features and improve generalization while keeping the base dataset size manageable
3Adaptability or versatility
If diverse pathological cases are included in the training data, then the model becomes more aware of biological and pathological diversity, but the data curation becomes more complex and time-consuming
Solution Approach 1:
The patent pre-segments anatomical structures and pre-processes contrast agent models before synthesis, creating reusable components that can be combined in various ways to generate diverse pathological cases. This preliminary preparation significantly reduces the complexity of curating diverse training data compared to manually collecting and annotating real diverse cases
Solution Approach 2:
The patent segments anatomical structures and contrast agent distributions into separate manageable components that can be independently manipulated and recombined. This segmentation enables systematic generation of diverse pathological variations without requiring complex end-to-end data curation processes
4Reliability
If images from multiple vendors and clinical sites are used for training, then the model becomes more robust to domain gaps, but the data acquisition and harmonization becomes more difficult
Solution Approach 1:
The patent creates synthetic images that can mimic the characteristics of different vendor devices and clinical sites through parameter adjustment, without requiring actual acquisition from multiple sources. This copying approach achieves domain robustness while avoiding the complexities of multi-source data acquisition and harmonization
Solution Approach 2:
The patent adjusts parameters in the synthetic data generation process to simulate different imaging conditions, vendor characteristics, and clinical site variations. This parameter-based control enables efficient creation of domain-diverse training data without the logistical challenges of collecting data from multiple real-world sources
Data Source
Figure 1
Figure 2
AI summary
A computer-implemented method for synthesizing combined x-ray-based image data comprising both contrasted and non-contrasted image data, the method comprising the following steps: - receiving a non-contrasted x-ray-based image (20); - segmenting of at least one target anatomical structure on the x-ray-based image to generate a segmentation (3) of the at least one target anatomical structure; - based on the segmentation (3), generating a region-of-interest mask for the non-contrasted x-ray-based image (20); - generating a modelled anatomical structure (8) that comprises a contrasted anatomical structure (9) and corresponds to the at least one target anatomical structure; - applying a trained generative machine learning algorithm on the non-contrasted x-ray-based image (20) with the region-of-interest mask and on the modelled anatomical structure (8) to generate synthetic combined x-ray-based image data (13) comprising both contrasted and non-contrasted image regions, such that an area defined by the region-of-interest mask is changed based on the modelled anatomical structure (8) to be contrasted; - providing the synthetic combined x-ray-based image data (13).