Pseudo Data Generation for Medical Model Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In the medical field, collecting a large number of diverse medical data sets for machine learning, such as deep neural networks, is challenging due to privacy concerns, leading to insufficient data for training and low accuracy in model performance.
Innovation Solution
A pseudo data generation apparatus that collects and processes data sets to generate pseudo physical parameters, which are then used in magnetic resonance simulations to create pseudo collection data, simulating MR signals, thereby augmenting the available data for training machine learning models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If actual medical data is collected for training machine learning models, then data diversity and authenticity are improved, but privacy protection requirements worsen and data collection becomes difficult
Solution Approach 1:
The patent generates pseudo data that copies the essential characteristics and statistical properties of actual medical data without containing real patient information. The pseudo data is created by transforming input data through a generative model that preserves data distribution patterns while eliminating identifiable information, thus providing sufficient training data quantity without violating privacy constraints
Solution Approach 2:
The patent uses synthetic pseudo data that can be generated on-demand without requiring collection, storage, or protection of sensitive patient information. This disposable pseudo data approach allows unlimited generation of training samples without the ongoing privacy management costs and constraints associated with real medical data
2Measurement precision
If a large number of diverse medical data sets are collected, then machine learning model accuracy is improved, but data collection difficulty increases due to privacy concerns
Solution Approach 1:
The patent implements a self-service data generation system where the generative model automatically creates pseudo data samples based on input data distributions. The system serves itself by continuously generating training data without requiring external data collection efforts, thereby maintaining high model accuracy while eliminating data collection difficulty
Solution Approach 2:
The patent transforms input data parameters through a generative model to create pseudo data with preserved statistical properties. By changing the data representation parameters while maintaining distribution characteristics, the system generates diverse training samples that improve model accuracy without the collection difficulties associated with real data
3Reliability
If real medical data is used for training, then data authenticity is improved, but the quantity of available data becomes insufficient
Solution Approach 1:
The patent segments the data generation process into multiple stages where input data is processed through transformations to create pseudo data. This segmentation allows the system to maintain authenticity through structured transformation while expanding quantity by generating multiple pseudo samples from each input data point
Solution Approach 2:
The patent creates copies of real data characteristics through generative models that replicate statistical properties, distribution patterns, and structural features. These copies maintain authenticity for training purposes while enabling unlimited quantity generation without the constraints of actual data availability
Data Source
AI summary
According to one embodiment, a pseudo data generation apparatus comprising processing circuitry. The processing circuitry collects a data set including data values of one or more dimensions. The processing circuitry performs conversion of the data values of the one or more dimensions included in the data set. The processing circuitry generates a pseudo physical parameter relating to each of one or more physical amounts.


