Pseudo-labeling for ML Model Training with Weakly-Labeled Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In the field of healthcare, machine learning models for physiological parameter prediction face challenges due to the scarcity of high-quality labeled data, leading to overfitting and poor generalization, especially when using invasive methods for data collection and requiring expert annotation for medical images.

Innovation Solution

A training method and system that leverages weakly-labeled data to generate pseudo-labeled data, using initial machine learning models trained on labeled data to improve data representations, and applying transformations to enhance model robustness and prediction accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If machine learning models are trained using high-quality labeled data, then prediction accuracy can be improved, but the scarcity of labeled data leads to overfitting and poor generalization

Engineering Contradiction:
Improveprediction accuracyVSAvoidgeneralization ability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent creates pseudo-labeled data by copying the labeling process through an initial model's predictions. The initial model trained on limited labeled data generates predictions for unlabeled data, creating synthetic labeled datasets that expand the training data without requiring additional expert annotations, thus improving generalization while maintaining accuracy

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary training with labeled data to create an initial model, which then serves as a tool to generate pseudo-labeled data before the final training stage. This preliminary action enables the system to leverage unlabeled data effectively in the subsequent fine-tuning phase

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If invasive surgeries are performed to obtain FFR measurements for training data, then data quality can be improved, but the process becomes challenging and time-consuming

Engineering Contradiction:
Improvedata qualityVSAvoiddata collection time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

Instead of performing invasive surgeries to obtain ground truth FFR measurements for every training case, the patent copies the essential information from a small set of surgically-obtained measurements and uses the initial model to generate pseudo-labels for the remaining training data, dramatically reducing the time and risk associated with invasive data collection

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces an initial model as an intermediary that translates limited high-quality surgical measurements into expanded training data through pseudo-labeling, bridging the gap between scarce invasive measurements and the need for large-scale training datasets

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If expert annotation is performed for medical image data, then labeling quality can be improved, but annotation time and cost increase significantly

Engineering Contradiction:
Improvelabeling qualityVSAvoidannotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent copies the expert annotation process by using an initial model trained on expert-annotated data to generate predictions that serve as pseudo-labels for unlabeled data, effectively replicating expert-level labeling quality without requiring actual expert time for each annotation

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs self-service labeling where the initial model automatically generates pseudo-labels for the training data without requiring external expert intervention, enabling the system to scale labeling capacity independently of expert availability

Inventive Principle:
Principle #25Self-service

4Reliability

If early stopping is used to avoid overfitting, then model robustness can be improved, but the approach ignores the challenges of weakly-labeled data

Engineering Contradiction:
Improvemodel robustnessVSAvoidhandling weakly-labeled data
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent performs preliminary training on labeled data to create a robust initial model before transitioning to weakly-labeled pseudo-data for fine-tuning. This staged approach allows early stopping to be effectively applied during the initial phase while enabling adaptation to weakly-labeled data in the subsequent phase

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the training parameters by transitioning from labeled data with strict validation criteria to pseudo-labeled data with different training dynamics. This parameter change enables the model to adapt from a regime where early stopping is critical to one where leveraging large amounts of weakly-labeled data is prioritized

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20220215958A1System and method for training machine learning models with unlabeled or weakly-labeled data and applying the same for physiological analysis
Publication Date: 2022.07.07 SHENZHEN KEYA MEDICAL TECH CORP
  • US20220215958A1 patent drawing
  • US20220215958A1 patent drawing
  • US20220215958A1 patent drawing

AI summary

The present disclosure relates to training methods for a machine learning model for physiological analysis. The training method may include receiving training data including a first dataset of labeled data of a physiological-related parameter and a second dataset of weakly-labeled data of the physiological-related parameter. The training method further includes training, by at least one processor, an initial machine learning model using the first dataset, and applying, by the at least one processor, the initial machine learning model to the second dataset to generate a third dataset of pseudo-labeled data of the physiological-related parameter. The training method also includes training, by the at least one processor, the machine learning model based on the first dataset and the third dataset, and providing the trained machine learning model for predicting the physiological-related parameter. Thereby, the weakly-labeled dataset may be sufficiently utilized in training of the machine learning model and improve ts p iformance.