Remote PPG Model Training via Self-Supervised Contrastive Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current remote photoplethysmography (PPG) systems require labeled video clips for training, which can be unavailable or difficult to develop, limiting their effectiveness in estimating heart rate from unlabeled video data.
Innovation Solution
A system and method for training a remote PPG model using self-supervised contrastive learning with unlabeled video clips, allowing the model to output a subject PPG signal based on a subject video clip, and optionally providing a saliency signal indicating the strongest PPG signal location, without the need for labeled data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If supervised deep learning approaches are used to train remote PPG models, then measurement precision of heart rate estimation is improved, but loss of time and resources increase due to the need for expensive labeled video data
Solution Approach 1:
The system employs self-supervised learning where the model trains itself using unlabeled video data by creating synthetic supervision signals from the data itself. The contrastive learning framework generates positive and negative samples from the unlabeled video clips, allowing the model to learn meaningful representations without requiring manually annotated heart rate data, thus eliminating the time-consuming data labeling process while maintaining estimation accuracy
2Measurement precision
If supervised deep learning approaches are used to train remote PPG models, then measurement precision of heart rate estimation is improved, but loss of substance increases due to expensive annotated data requirements
Solution Approach 1:
The model performs self-supervised learning by generating its own training signals from unlabeled video data through contrastive learning. By creating positive samples (temporally adjacent video clips) and negative samples (temporally distant video clips) automatically, the system eliminates the need for expensive annotated datasets, reducing resource consumption while achieving comparable or superior performance to supervised methods
Solution Approach 2:
The system creates synthetic copies and transformations of the original unlabeled video data to generate diverse training samples. Through temporal copying and transformation of video clips into positive and negative pairs, the model learns robust representations without requiring external annotated data, effectively multiplying the utility of available unlabeled data
3Ease of operation
If self-supervised contrastive learning is used to train remote PPG models, then ease of operation is improved by using unlabeled video data, but measurement precision may worsen compared to supervised methods
Solution Approach 1:
The contrastive learning framework implements a feedback mechanism where the model's own predictions and representations are used to generate training signals. By computing similarity between positive samples and comparing against negative samples, the system creates an automatic feedback loop that guides learning without external annotations. This feedback-driven approach maintains measurement precision while dramatically improving ease of operation
Solution Approach 2:
The system changes the learning paradigm from supervised to self-supervised, fundamentally altering the training parameters and objectives. By modifying the loss function to use contrastive objectives instead of supervised classification losses, and by changing how training samples are constructed (temporal relationships instead of labeled pairs), the system achieves both ease of operation and maintains precision through different computational pathways
Data Source
AI summary
Systems and methods for training remote photoplethysmography (“PPG”) models that outputs a subject PPG signal based on a subject video clip of a subject are described herein. The system may have a processor and a memory in communication with the processor. The memory may include a training module having instructions that, when executed by the processor, cause the processor to train the remote PPG model in a self-supervised contrastive learning manner using an unlabeled video clip having a sequence of images of a face of a person.


