Synthetic CT Data Generation via Autoencoder Transfer Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge in tomographic reconstruction lies in the scarcity of clinical data due to patient privacy and proprietary concerns, coupled with the unavailability of ground truth data for deep learning-based reconstructions, particularly in areas like 4D cardiac imaging, where data is either sparse or expensive to label, and unsupervised learning is immature.
Innovation Solution
A method and system that integrate simulation, emulation, and transfer learning using auto encoder networks and generative artificial neural networks to generate synthesized data sets from CT sinograms and images, leveraging a combination of simulator and emulator features to produce realistic anatomical and physical data for deep learning-based reconstructions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If deep learning techniques are used for tomographic reconstruction, then reconstruction quality is improved, but data availability deteriorates due to patient privacy and proprietary considerations
Solution Approach 1:
The patent creates synthetic CT data copies through simulation and transfer learning that mimic real clinical data characteristics without using actual patient data. The autoencoder network learns from limited real data and generates synthetic sinograms and images that preserve anatomical and physical properties, enabling deep learning training while maintaining patient privacy.
Solution Approach 2:
The patent performs preliminary data preparation by pre-processing real CT scans to extract anatomical structures and physical properties before generating synthetic data. This preliminary action includes creating ground truth labels, segmenting organs, and storing physical properties that will be transferred to synthetic data generation, enabling subsequent realistic synthetic data creation.
2Measurement precision
If ground truth data is obtained through expert labeling, then data accuracy is improved, but cost and time consumption increase
Solution Approach 1:
The patent implements self-service through automated ground truth generation using the autoencoder network. The system automatically extracts anatomical structures, segment organs, and generate labels without requiring expert manual annotation. The network learns from pre-processed real data and autonomously creates accurate ground truth for synthetic data generation, eliminating time-consuming expert labeling.
3Reliability
If more real clinical data is collected for training, then model performance is improved, but patient privacy risks increase
Solution Approach 1:
The patent replaces real clinical data with synthetic data copies generated through simulation and transfer learning. The autoencoder network creates artificial sinograms and CT images that replicate the statistical properties, anatomical variations, and physical characteristics of real data without containing any actual patient information, thus maintaining model performance while eliminating privacy risks.
Solution Approach 2:
The patent transforms real data parameters into synthetic equivalents by learning the underlying distributions and relationships. The system captures anatomical parameters, physical properties, and imaging characteristics from real data, then generates synthetic data with matched parameter distributions, ensuring model performance without using sensitive patient information.
Data Source
AI summary
In some embodiments, a method of machine learning includes identifying, by an auto encoder network, a simulator feature based, at least in part, on a received first simulator data set and an emulator feature based, at least in part, on a received first emulator data set. The method further includes determining, by a synthesis control circuitry, a synthesized feature based, at least in part, on the simulator feature and based, at least in part, on the emulator feature; and generating, by the auto encoder network, an intermediate data set based, at least in part, on a second simulator data set and including the synthesized feature. Some embodiments of the method further include determining, by a generative artificial neural network, a synthesized data set based, at least in part, on the intermediate data set and based, at least in part, on an objective function.


