Pseudo Data Generation for Medical Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In the medical field, collecting a large number of diverse medical data sets for machine learning, such as deep neural networks, is challenging due to privacy concerns, leading to insufficient data for training and low accuracy in model performance.

Innovation Solution

A pseudo data generation apparatus that collects and processes data sets to generate pseudo physical parameters, which are then used in magnetic resonance simulations to create pseudo collection data, simulating MR signals, thereby augmenting the available data for training machine learning models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If actual medical data is collected for training machine learning models, then data diversity and authenticity are improved, but privacy protection requirements worsen and data collection becomes difficult

Engineering Contradiction:
Improvenumber of training dataVSAvoidprivacy protection constraints
Core Design Contradiction:
Quantity of substanceVSObject-affected harmful factors

Solution Approach 1:

The patent generates pseudo data that copies the essential characteristics and statistical properties of actual medical data without containing real patient information. The pseudo data is created by transforming input data through a generative model that preserves data distribution patterns while eliminating identifiable information, thus providing sufficient training data quantity without violating privacy constraints

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent uses synthetic pseudo data that can be generated on-demand without requiring collection, storage, or protection of sensitive patient information. This disposable pseudo data approach allows unlimited generation of training samples without the ongoing privacy management costs and constraints associated with real medical data

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

2Measurement precision

If a large number of diverse medical data sets are collected, then machine learning model accuracy is improved, but data collection difficulty increases due to privacy concerns

Engineering Contradiction:
Improvemodel accuracyVSAvoiddata collection difficulty
Core Design Contradiction:
Measurement precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The patent implements a self-service data generation system where the generative model automatically creates pseudo data samples based on input data distributions. The system serves itself by continuously generating training data without requiring external data collection efforts, thereby maintaining high model accuracy while eliminating data collection difficulty

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent transforms input data parameters through a generative model to create pseudo data with preserved statistical properties. By changing the data representation parameters while maintaining distribution characteristics, the system generates diverse training samples that improve model accuracy without the collection difficulties associated with real data

Inventive Principle:
Principle #35Parameter changes

3Reliability

If real medical data is used for training, then data authenticity is improved, but the quantity of available data becomes insufficient

Engineering Contradiction:
Improvedata authenticityVSAvoiddata quantity
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent segments the data generation process into multiple stages where input data is processed through transformations to create pseudo data. This segmentation allows the system to maintain authenticity through structured transformation while expanding quantity by generating multiple pseudo samples from each input data point

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates copies of real data characteristics through generative models that replicate statistical properties, distribution patterns, and structural features. These copies maintain authenticity for training purposes while enabling unlimited quantity generation without the constraints of actual data availability

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20220343634A1Pseudo data generation apparatus, pseudo data generation method, and non-transitory storage medium
Publication Date: 2022.10.27 CANON KK
  • US20220343634A1 patent drawing
  • US20220343634A1 patent drawing
  • US20220343634A1 patent drawing

AI summary

According to one embodiment, a pseudo data generation apparatus comprising processing circuitry. The processing circuitry collects a data set including data values of one or more dimensions. The processing circuitry performs conversion of the data values of the one or more dimensions included in the data set. The processing circuitry generates a pseudo physical parameter relating to each of one or more physical amounts.