Pseudo-labeling for ML Model Training with Weakly-Labeled Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In the field of healthcare, machine learning models for physiological parameter prediction face challenges due to the scarcity of high-quality labeled data, leading to overfitting and poor generalization, especially when using invasive methods for data collection and requiring expert annotation for medical images.
Innovation Solution
A training method and system that leverages weakly-labeled data to generate pseudo-labeled data, using initial machine learning models trained on labeled data to improve data representations, and applying transformations to enhance model robustness and prediction accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If machine learning models are trained using high-quality labeled data, then prediction accuracy can be improved, but the scarcity of labeled data leads to overfitting and poor generalization
Solution Approach 1:
The patent creates pseudo-labeled data by copying the labeling process through an initial model's predictions. The initial model trained on limited labeled data generates predictions for unlabeled data, creating synthetic labeled datasets that expand the training data without requiring additional expert annotations, thus improving generalization while maintaining accuracy
Solution Approach 2:
The patent performs preliminary training with labeled data to create an initial model, which then serves as a tool to generate pseudo-labeled data before the final training stage. This preliminary action enables the system to leverage unlabeled data effectively in the subsequent fine-tuning phase
2Measurement precision
If invasive surgeries are performed to obtain FFR measurements for training data, then data quality can be improved, but the process becomes challenging and time-consuming
Solution Approach 1:
Instead of performing invasive surgeries to obtain ground truth FFR measurements for every training case, the patent copies the essential information from a small set of surgically-obtained measurements and uses the initial model to generate pseudo-labels for the remaining training data, dramatically reducing the time and risk associated with invasive data collection
Solution Approach 2:
The patent introduces an initial model as an intermediary that translates limited high-quality surgical measurements into expanded training data through pseudo-labeling, bridging the gap between scarce invasive measurements and the need for large-scale training datasets
3Measurement precision
If expert annotation is performed for medical image data, then labeling quality can be improved, but annotation time and cost increase significantly
Solution Approach 1:
The patent copies the expert annotation process by using an initial model trained on expert-annotated data to generate predictions that serve as pseudo-labels for unlabeled data, effectively replicating expert-level labeling quality without requiring actual expert time for each annotation
Solution Approach 2:
The system performs self-service labeling where the initial model automatically generates pseudo-labels for the training data without requiring external expert intervention, enabling the system to scale labeling capacity independently of expert availability
4Reliability
If early stopping is used to avoid overfitting, then model robustness can be improved, but the approach ignores the challenges of weakly-labeled data
Solution Approach 1:
The patent performs preliminary training on labeled data to create a robust initial model before transitioning to weakly-labeled pseudo-data for fine-tuning. This staged approach allows early stopping to be effectively applied during the initial phase while enabling adaptation to weakly-labeled data in the subsequent phase
Solution Approach 2:
The patent changes the training parameters by transitioning from labeled data with strict validation criteria to pseudo-labeled data with different training dynamics. This parameter change enables the model to adapt from a regime where early stopping is critical to one where leveraging large amounts of weakly-labeled data is prioritized
Data Source
AI summary
The present disclosure relates to training methods for a machine learning model for physiological analysis. The training method may include receiving training data including a first dataset of labeled data of a physiological-related parameter and a second dataset of weakly-labeled data of the physiological-related parameter. The training method further includes training, by at least one processor, an initial machine learning model using the first dataset, and applying, by the at least one processor, the initial machine learning model to the second dataset to generate a third dataset of pseudo-labeled data of the physiological-related parameter. The training method also includes training, by the at least one processor, the machine learning model based on the first dataset and the third dataset, and providing the trained machine learning model for predicting the physiological-related parameter. Thereby, the weakly-labeled dataset may be sufficiently utilized in training of the machine learning model and improve ts p iformance.


