Unsupervised learning of physiological signals from video
The unsupervised, non-contrastive learning framework for rPPG effectively extracts physiological signals from video streams, addressing data scarcity and synchronization challenges, achieving superior performance and robustness in vital sign estimation.
Patent Information
- Application Number
- JP2025526779
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-11
- Filing Date
- 2023-11-13
- Publication Date
- 2025-11-14
AI Technical Summary
Existing video-based vital sign estimation methods face challenges such as data scarcity, privacy concerns, and synchronization issues, leading to limited performance and robustness, especially in unsupervised learning scenarios.
An unsupervised, non-contrastive learning framework for remote photoplethysmography (rPPG) that uses a combination of visible light, near-infrared, and thermal video streams, synchronized and processed to extract physiological signals like pulse rate and waveform without ground truth data, employing a 3DCNN and loss functions to constrain frequency bands and encourage variance.
The framework achieves accurate and robust estimation of physiological signals, outperforming supervised methods on benchmark datasets and demonstrating resilience to environmental variations, while reducing the need for labeled data.
Smart Images

Figure 2025537281000001_ABST
Abstract
Description
[Technical Field]
[0001] Priority information This application claims the benefit of U.S. Provisional Patent Application No. 63 / 424,606, filed November 11, 2022, which is incorporated herein by reference in its entirety.
[0002] FIELD OF THE INVENTION Embodiments of the present invention relate generally to the use of biometrics, and more particularly to unsupervised learning of physiological signals from video. [Background technology]
[0003] In general, biometrics can be used to track vital signs that provide indicators regarding a subject's physical condition, which can be used in a variety of ways. As an example, for border security or health surveillance, vital signs can be used to screen for health risks (e.g., body temperature) or detect deception (e.g., changes in pulse rate or pupil diameter). While body temperature detection is a well-developed technology, collecting other useful and accurate vital signs, such as pulse rate (i.e., heart rate or beats per minute) and pulse waveform, has required the attachment of physical devices to the subject. The desire to perform biometric measurements without physical contact has led to several video-based technologies.
[0004] Reliable pulse rate or waveform estimation from a camera sensor is more difficult than with a contact plethysmograph for several reasons: Changes in reflected light from the skin surface due to light absorption by blood are much smaller than those due to changes in lighting; subject movement, even in ambient lighting environments, dramatically changes the reflected light, overwhelming the pulse signal.
[0005] Camera-based vital sign estimation is a rapidly growing field that enables contactless health monitoring in a variety of environments. While the number of successful approaches is growing rapidly, the size of benchmark video datasets containing concurrent recordings of vital signs remains relatively stagnant. It is well known in the machine learning community that increasing the amount and diversity of training data is an effective strategy for improving performance.
[0006] Collecting remote physiological data is challenging for several reasons. First, recording hours of high-quality video results in bulky and unmanageable data volumes. Second, recording diverse subjects with associated medical data is challenging due to privacy concerns. Furthermore, synchronizing contact measurements and video recordings across various environments is highly dependent on the investigator's hardware infrastructure and lab environment. Even contact measurements used for ground truth contain noise, making data curation difficult. These challenges, which contribute to data scarcity, hinder the scaling and robustness of models.
[0007] Few publicly available sources of simultaneous video and physiological recordings make it difficult to gather data for supervised training. Fortunately, recent studies have shown that unsupervised training of remote photoplethysmography (rPPG) is effective.
[0008] The dominant class of approaches for remote pulse estimation has shifted over the past decade from blind source separation, through linear color transforms, to training supervised deep learning-based models. While color transforms generalize well across many datasets, deep learning-based models exhibit better accuracy when tested on data from a similar distribution as the training set. For this reason, deep learning research has focused on optimizing neural architectures for robust spatial and temporal feature extraction from limited benchmark datasets.
[0009] To avoid data bottlenecks, large-scale synthetic physiological datasets have recently become popular. The SCAMPS dataset contains videos of 2,800 synthetic avatars in various environments with various corresponding labels, such as PPG, ECG, respiration, and facial action units. The UCLA-synthetic dataset contains 480 videos and shows that training models with both real and synthetic data yields the best results. Another strength of synthetic datasets is their ability to cover a wide range of skin tones, which can be difficult when collecting real data.
[0010] Another potential solution to the lack of physiological training data is unsupervised learning, which requires only a large set of images and periodic priors for the output signal.
[0011] Self-supervised learning has made rapid progress for image representation learning. Recently, two major classes of approaches have competed: contrastive learning and non-contrastive learning (or regularized learning). Contrastive approaches prescribe a criterion for distinguishing whether two samples are the same or different, and compare embeddings to either pull or push predictions. Non-contrastive methods force variance in predictions across batches to expand positive pairs and avoid collapse, where the model embedding resides in a small subspace of the feature space rather than across a larger or entire embedding space. Distillation methods avoid collapse by using only positive samples and applying moving average and stopping gradient operators. Another class of approaches maximizes the information content of the embedding.
[0012] All existing unsupervised rPPG approaches are contrastive. In a contrastive framework, pairs of input videos are passed as inputs to the same model, and predictions on similar videos are attracted, while predictions from dissimilar videos are rejected. Although the method for selecting negative samples varies across previous literature, the underlying contrastive framework is similar.
[0013] Accordingly, the present inventors have developed systems, devices, methods, and non-transitory computer-readable instructions that enable unsupervised, non-contrastive learning of physiological signals from video. Summary of the Invention
[0014] Accordingly, the present invention is directed to unsupervised, non-contrastive learning of physiological signals from video that substantially avoids one or more of the problems due to limitations and drawbacks of the related art.
[0015] Additional features and advantages of the invention will be set forth in the description which follows, and in part will be obvious from the description, or may be learned by the practice of the invention. The objectives and other advantages of the invention will be realized and attained by the structure particularly pointed out in the accompanying drawings and the written description and claims thereof.
[0016] To achieve these and other advantages, and in accordance with the purpose of the present invention, as embodied and broadly described, embodiments include systems, apparatus, methods, and non-transitory computer-readable instructions for uncontrolled unsupervised learning of physiological signals from video streams, the computer-implemented method including capturing a video stream of a subject, the video stream including a sequence of frames, processing each frame of the video stream to update a physiological signal detection function, determining a physiological signal from the video stream, and applying the updated physiological signal detection function to subsequent video streams.
[0017] In connection with any of the various embodiments, the media or video stream includes one or more of a visible light video stream, a near-infrared video stream, a long-wavelength infrared video stream, a thermal video stream, and an audio stream of the subject.
[0018] In accordance with any of the various embodiments, the physiological signal includes at least one of a pulse rate, a blood pressure, and an eye blink rate.
[0019] In accordance with any of the various embodiments, the physiological signal includes at least one of a pulse rate or an audio frequency.
[0020] In connection with any of the various embodiments, cropping each frame of the media stream to encapsulate a region of interest including one or more of the face, cheeks, forehead, and eyes.
[0021] In accordance with any of the various embodiments, the region of interest includes two or more body parts.
[0022] In connection with any of the various embodiments, combining at least two of the visible light video stream, the near-infrared video stream, and the thermal video stream into a fused video stream.
[0023] In connection with any of the various embodiments, the visible light video stream, the near-infrared video stream, and / or the thermal video stream are combined according to a synchronizer.
[0024] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory and are intended to provide further explanation of the invention as claimed. [Brief explanation of the drawings]
[0025] The accompanying drawings, which are included to provide a further understanding of the invention and are incorporated in and constitute a part of this specification, illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.
[0026] [Figure 1] FIG. 1 is a diagram showing a pulse waveform estimation system.
[0027] [Figure 2] Schematic of the non-contrastive unsupervised learning (NCUL) framework for remote photoplethysmography (rPPG) compared to traditional supervised and unsupervised learning.
[0028] [Figure 3] Figure 1 shows the results of training using only the bandwidth loss Lb.
[0029] [Figure 4] Figure showing the results of models trained and tested on subject-disjoint partitions from the same dataset.
[0030] [Figure 5] Figure showing intra-dataset waveform predictions for the baseline dataset by the end-to-end unsupervised model over an 8-second window.
[0031] [Figure 6] Figure showing the results of NCUL and supervised training on the same architecture.
[0032] [Figure 7] Figure showing training and testing results for UBFC-rPPG.
[0033] [Figure 8] FIG. 1 illustrates a computer-implemented method for unsupervised, non-contrastive learning of physiological signals from video streams. DETAILED DESCRIPTION OF THE INVENTION
[0034] Reference will now be made in detail to the embodiments of the present invention, which are illustrated in the accompanying drawings, and wherever possible, like reference numerals are used to refer to like elements.
[0035] Embodiments of a user interface and associated methods for using the device are described. However, it should be understood that the user interface and associated methods are applicable to numerous device types, including portable communication devices such as tablets and mobile phones. Portable communication devices can support a variety of applications, such as wired and wireless communication. Various applications that can be stored (in non-transitory memory) and executed (by a processor) on the device can use at least one common physical user interface device, such as a touchscreen. One or more features of the touchscreen and corresponding information displayed on the device can be adjusted and / or varied from one application to another and / or within each application. In this manner, the common physical architecture of the device can support a variety of applications with intuitive and transparent user interfaces.
[0036] Embodiments of the present invention provide systems, devices, methods, and non-transitory computer-readable instructions for measuring one or more biometrics, including heart rate, pulse waveform, and / or respiration, without physical contact with a subject. Other biometrics may include pulse, gaze, eye blink, pupillary measurement, facial temperature, oxygen level, blood pressure, voice, tone and / or frequency of voice, micro-expressions, etc. In various embodiments, the systems, devices, methods, and instructions collect, process, and analyze video captured in one or more modalities (e.g., visible light, near-infrared, long-wave infrared, thermal) to provide unsupervised learning of physiological signals from video signals or video data (e.g., MP4).
[0037] To expand the possibilities for addressing the challenges of remote human monitoring, additional biometric sensors can be used: in various embodiments, changes in the subject's gaze, eye blink rate, pupil diameter, speech, facial temperature, and micro-expressions can also be used.
[0038] As described herein, the pulse or pulse waveform relative to a subject's heart rate may be used as a biometric input to establish characteristics of the subject's physical state and how they change over an observation period (e.g., during interrogation or other activities). Remote photoplethysmography (rPPG) is the monitoring of blood volume pulse waves from a camera located at a distance. rPPG allows blood volume pulses to be detected from video footage located at a distance from the skin surface. The disclosure of U.S. Application No. 17 / 591,929, filed February 3, 2022, entitled "VIDEO BASED DETECTION OF PULSE WAVEFORM," is incorporated herein by reference in its entirety.
[0039] 1 shows a system 100 for pulse waveform estimation. The system 100 includes an optical sensor system 1, a video I / O system 6, and a video processing system 101.
[0040] Optical sensor system 1 includes one or more camera sensors, each configured to capture a video stream including a sequence of frames. For example, optical sensor system 1 may include a visible light camera 2, a near-infrared camera 3, a thermal camera 4, or any combination thereof. When multiple camera sensors are utilized (e.g., single-modality or multi-modality), the resulting multiple video streams may be synchronized according to a synchronizer 5. Alternatively, or in addition, one or more video analytics techniques may be utilized to synchronize the video streams. While a visible light camera 2, a near-infrared camera 3, and a thermal camera 4 are enumerated, other media devices, such as a speech recorder, may also be used.
[0041] The video I / O system 6 receives one or more captured video streams. For example, the video I / O system 6 is configured to receive a raw visible light video stream 7, a near-infrared video stream 8, and a thermal video stream 9 from the optical sensor system 1. Here, the received video streams may be stored according to a known digital format. When multiple video streams are received (e.g., single-modality or multi-modality), the fusion processor 10 is configured to combine the received video streams. For example, the fusion processor 10 may combine the visible light video stream 7, the near-infrared video stream 8, and / or the thermal video stream 9 into a fused video stream 11. Here, each stream may be synchronized according to an output (e.g., a clock signal) from the synchronizer 5.
[0042] In the video processing system 101, the region of interest detector 12 detects (i.e., spatially locates) one or more spatial regions of interest (ROIs) within each video frame. The ROIs may be the face, another body part (e.g., hand, arm, leg, neck, etc.), or any combination of body parts. First, the region of interest detector 12 determines one or more coarse spatial ROIs within each video frame. The region of interest detector 12 is robust to strong occlusions of the face due to face masks or other head garments. The frame pre-processor 13 then crops the frames to encapsulate the one or more ROIs. In some embodiments, the cropping involves downsizing each frame using bicubic interpolation to reduce the number of image pixels to be processed. Alternatively, or additionally, the cropped frames may be further resized to smaller images.
[0043] The sequence preparation system 14 aggregates a batch of ordered sequences or subsequences of frames to be processed from the frame preprocessor 13. A three-dimensional convolutional neural network (3DCNN) 15 then receives the sequences or subsequences of frames from the sequence preparation system 14. The 3DCNN 15 processes the sequences or subsequences of frames with a three-dimensional convolutional neural network to determine the spatial and temporal dimensions of each frame in the sequence or subsequence of frames and generate pulse waveform points for each frame in the sequence of frames. The 3DCNN 15 applies a series of three-dimensional convolutions, averaging, pooling, and nonlinearities to generate a one-dimensional signal that approximates the pulse waveform 16 of the input sequence or subsequence.
[0044] In some configurations, the pulse aggregation system 17 combines any number of pulse waveforms 16 from a sequence or subsequence of frames into an aggregated pulse waveform 18 that represents the entire video stream. The diagnostic extractor 19 is configured to calculate heart rate and heart rate variability from the aggregated pulse waveform 18. The calculated heart rates of various subsequences may be compared to identify heart rate variability. The display device 20 receives real-time or near-real-time updates from the diagnostic extractor 19 and displays the aggregated pulse waveform 18, heart rate, and heart rate variability to an operator. The storage device 21 is configured to store the aggregated pulse waveform 18, heart rate, and heart rate variability associated with the subject.
[0045] Additionally or alternatively, a sequence of frames may be divided into partially overlapping subsequences within the sequence preparation system 14, where a first subsequence of frames overlaps with a second subsequence of frames. Frame overlap between subsequences prevents edge effects. The pulse aggregation system 17 may then apply a Hann function to each subsequence and sum the overlapping subsequences to generate an aggregate pulse waveform 18 with the same number of samples as the frames in the original video stream. In some configurations, each subsequence is passed individually to the 3DCNN 15, which performs a series of operations to generate a pulse waveform for each subsequence 16. Each pulse waveform output from the 3DCNN 15 is a time series with real values for each video frame. Each subsequence is processed individually by the 3DCNN 15, and then recombined.
[0046] In some embodiments, one or more filters may be applied to the region of interest. For example, one or more wavelengths of LED light may be filtered. The LED may be illuminated over the entire region of interest and its surrounding surface or a portion thereof. Additionally or alternatively, the time signal of non-skin regions may be further processed. For example, analyzing the eyebrows or the sclera (white of the eye) may identify changes that are strongly correlated with movement, but may not necessarily correlate with plethysmography. Detection of a periodic signal consistent with a pulse on a surface other than the skin may indicate a non-existent subject or an attempted security breach.
[0047] Although illustrated as a single system, the functionality of system 100 may be implemented as a distributed system. While system 100 determines heart rate, other distributed components track, for example, changes in the subject's gaze, eye blink rate, pupil diameter, speech, facial temperature, and micro-expressions. Furthermore, the functionality disclosed herein may be implemented on separate servers or devices, which may be coupled via a network, such as a security kiosk coupled to a back-end server. Furthermore, one or more components of system 100 may not be included. For example, system 100 may be a smartphone or tablet device that includes a processor, memory, and display, but may not include one or more of the other components shown in FIG. 1 . This embodiment may be implemented using a variety of processing devices and memory storage devices. For example, a CPU and / or GPU may be used in the processing system to reduce execution time and calculate pulse rate in near real time. System 100 may be part of a larger system. Accordingly, system 100 may include one or more additional functional modules.
[0048] Subtle quasi-periodic physiological signals, such as blood volume pulse and respiration, may be extracted from RGB video, enabling remote health monitoring and other applications. Advances in remote pulse estimation, i.e., remote photoplethysmography (rPPG), are currently driven by supervised deep learning solutions. However, current approaches are trained and evaluated on limited benchmark datasets recorded in ground truth from contact PPG sensors.
[0049] The present embodiment provides the first uncontrolled unsupervised learning framework for signal regression to reduce and / or remove the constraint of labeled video data. With minimal assumptions of periodicity and finite bandwidth, the present embodiment identifies blood volume pulses directly from unlabeled video. Encouraging sparse power spectra within desired bandwidth limits and variance over batches of power spectra is sufficient to learn the visual features of periodic signals. The use of unlabeled video data not specifically created for rPPG has been validated to train a robust pulse rate estimator. Given the limited induced bias and positive empirical results, the present embodiment can be easily applied to other periodic signals from video, enabling multiple physiological measurements without the need for ground truth.
[0050] Embodiments provide unsupervised, non-contrastive learning of physiological signals from video. A model can be trained to extract periodic signals from video using a non-contrastive formulation. Given a video of a face and instructed to identify periodic signals between 40 and 180 beats per minute, the model successfully learns to estimate the blood volume pulse, despite subtleties in visual facial features that provide information about blood pulse.
[0051] The embodiments provide a framework for unsupervised learning by leveraging periodic signal priors. The embodiments also provide the first unsupervised learning method for camera-based vital sign measurement. Furthermore, the embodiments enable training models using non-rPPG-specific video datasets without ground truth vital signs.
[0052] Figure 2 provides an overview of our non-contrastive unsupervised learning (NCUL) framework for remote photoplethysmography (rPPG) compared to traditional supervised and unsupervised learning. The supervised and controlled losses use distance metrics to the ground truth or other samples. The framework applies the loss directly to prediction by shaping the frequency spectrum and encouraging variance across batches of inputs. Power outside the band limit is penalized to learn invariance to irrelevant frequencies. Power within the band limit is encouraged to be sparsely distributed around the maximum frequency.
[0053] First, we formulate the general setup for signal regression from video. Given a video sample x sampled from a dataset D, i ∈R T×W×H×C consists of T images of size W × H pixels across C channels, captured uniformly over time. State-of-the-art techniques use a waveform R of the same length as the video. T ∋y i =f(x i ) has been proposed. Recently, a spatiotemporal neural network has been used as the model f to effectively model the task end-to-end. While previous work has focused on supervised minimization of the loss for contact pulse measurements, here non-contrastive learning uses only the model estimate.
[0054] The estimated pulse can be subjected to strong priors (e.g., frequency range or strong periodicity). One or more constraints can then be implemented in the frequency domain rather than the time domain. Thus, the waveform prediction is passed through an FFT before computing the loss.
[0055] loss One advantage of unsupervised learning for periodic signals is that the solution space can be heavily constrained. For physiological signals like respiration and blood volume pulse, sound upper and lower bounds on the frequency in breaths and beats per minute are known. The extracted signal is relatively sparse in the frequency domain, and the model filters out noise signals present in the video. These constraints simplify the problem of finding good features for the desired signal in the data.
[0056] One of the most useful constraints on a model is frequency band limiting. Previous unsupervised methods have used the Independent Power Ratio (IPR) as a validation metric for model selection. The use of frequency band limiting is also useful when training a model. If the lower limit of the band limit is a and the upper limit is b, then the bandwidth loss is:
number
[0057] Figure 3 shows the results of training with only the bandwidth loss Lb. Each column shows predictions from a model trained on UBFC-rPPG for 20 epochs using one or each loss. The first two rows show samples in the time and frequency domains, respectively. The last row shows the distribution of signal power across the validation set. The bandwidth loss penalizes signal power outside predefined band limits (e.g., 40–180 bpm), constraining the output space. The last row also shows that the model restricts signal power between the band limits. The sparsity loss encourages narrow spectra with strong periodic components. By itself, the model learns solutions characterized by very low frequencies. The variance loss encourages diverse power spectra across a batch, preventing the model from converging to a narrow bandwidth. Combined, this allows the model to learn to estimate periodic signals within the desired band limits.
[0058] Blood volume pulsations contain dominant frequencies, and pulse rate is the most common physiological marker used in practice. By penalizing broadband predictions, the model can be strengthened. This also simplifies the true signal the model must identify by penalizing visual dynamics that may not exhibit periodicity. A similar formulation to IPR can be used, but the frequency boundaries are chosen around the maximum predicted frequency. Specifically, the sparsity loss L S is the same as equation (1) when the band limits are replaced by -∞ and ∞, and a = argmax(F) - ΔF and b = argmax(F) + ΔF are used. The experiment was conducted with ΔF set to 6 beats per minute. Figure 3 shows the results of training using only the sparsity loss in the second column. For a single sample, the power spectrum is very sparse. For the entire dataset, completely ignoring high frequencies and learning visual features corresponding to low frequencies makes it easier to predict a sparse solution.
[0059] One risk of asymmetric methods is that the model collapses to a trivial solution, making predictions independent of the input features. In regularization methods like VICReg, a hinge loss for batch variance of predictions is used to enforce diverse outputs. A similar strategy can avoid the collapse of the model, but instead spread the variance of the power spectral density to a uniform distribution across the supported frequencies.
[0060] The variance loss operates on a uniform prior distribution P over d frequencies and a batch of n spectral densities, F = [v1,…,vn], where each vector is a d-dimensional frequency decomposition of the predicted waveform. A normalized sum Q of the densities over the batch is computed. We also define the variance loss as the squared Wasserstein distance to the uniform prior:
number
[0061] In summary, the training loss function is the sum of the losses mentioned above:
number
[0062] Although it was possible to weight certain loss components more than others, the loss was formulated to scale between 0 and 1. Experiments showed that the unweighted sum gave good performance. The combined loss function encourages the model to explore across the supported frequencies to discover visual features of strong periodic signals. Surprisingly, this framework is sufficient for learning to regress blood volume in videos, as shown in the last column of Figure 3.
[0063] Unlike known approaches that apply only frequency augmentation, multiple augmentations are applied in both spatial and temporal dimensions, extending the model to identify invariance to noise that may be encountered in real environments.
[0064] Image Augmentation Each pixel location in the clip is added with random Gaussian noise with mean 0 and standard deviation 2 on the original image scale from 0 to 255. Illumination is then boosted by adding a constant sampled from a Gaussian distribution with mean 0 and standard deviation 10 to each pixel in the clip, darkening or brightening the image.
[0065] spatial enhancement The video clip is randomly flipped horizontally with a 50% chance. The spatial dimensions of the clip are randomly cropped to a square between half its original length and its original length. The cropped clip is linearly interpolated back to its original spatial dimensions.
[0066] Temporal enhancement Based on the general assumption that the desired signal is strongly periodic and sparsely represented in the Fourier domain, the video clip is randomly reversed along the time dimension with a 50% probability. Note that the Fourier decomposition of a time-reversed sine wave is the same as the original sine wave.
[0067] Frequency Boost Perhaps the most important augmentation is frequency resampling, where video is linearly interpolated to different frame rates. This augmentation is particularly interesting for rPPG because it transforms the video input and target signal equivalently along the time dimension, making them equivariant. Given the invariant transformations described above, τ(·)~T, the equivariant frequency resampling operation, φ(·)Φ, and the model f(·) for inferring waveforms from video, we have:
number
[0068] This is a powerful augmentation because it allows us to expand the target distribution along with the video input. In our experiments, the input clips were randomly resampled by a factor c~U(0.6, 1.4). After applying the resampling augmentation, we then scaled the band limit by c to avoid penalizing the model if the augmentation pushed the original pulse frequency outside the original band limit.
[0069] We conducted several experiments to evaluate intra- and inter-dataset performance. PURE, UBFC-rPPG, and DDPM were used as benchmark rPPG datasets for both training and testing, while the CelebV-HQ dataset and HKBU-MARs were used for unsupervised training only.
[0070] Deception Detection and Physiological Monitoring (DDPM) was conducted in the form of interviews with 86 subjects, who attempted to answer questions deceptively. The interviews were recorded at 90 frames per second for an average of over 10 minutes. The naturalistic speech and frequent head pose changes make this a challenging and under-constrained rPPG dataset.
[0071] PURE is a benchmark rPPG dataset consisting of 10 subjects recorded over six sessions. Each session lasted approximately 1 minute, and raw footage was recorded at 30 fps. For each subject, the six sessions consisted of (1) stationary, (2) speaking, (3) slow head movement, (4) fast head movement, (5) small head rotation, and (6) medium head rotation. Pulse rates were at or near the subjects' resting pulse rates.
[0072] UBFC-rPPG contains 42 one-minute videos recorded at 30 fps, in which subjects play a timed mathematical game to increase their heart rate while restricting head movement during the recording.
[0073] The HKBU 3D Mask Attack with Real-World Variation (HKBU-MARs) consists of 12 subjects filmed with seven different cameras in six different lighting configurations, with 504 10-second videos each. The diverse lighting and camera sensors make this a valuable dataset for unsupervised training. Version 2 of HKBU-MARs was used, which contains footage of both realistic 3D masks and unmasked subjects.
[0074] The High-Quality Celebrity Video Dataset (CelebVHQ) is a set of processed YouTube® videos containing 35,666 facial videos from over 15,000 IDs. The videos vary dramatically in length, lighting conditions, emotions, motion, skin tone, and camera sensors. Given sufficient methods for unsupervised learning, the CelebV-HQ dataset appears to be a good candidate for training a robust model for pulse estimation. The biggest challenge in using online videos is the low quality of the images, which may be compressed before uploading and by the video provider. Compression is a known challenge for rPPG, as blood volume pulses are highly optically subtle.
[0075] Training, data preprocessing To prepare video clips for the spatiotemporal deep learning model, 68 facial landmarks were extracted using OpenFace. A bounding box was defined at the minimum and maximum (x, y) positions of each frame, with a 5% horizontal crop extension to ensure the presence of cheeks and chin. The top and bottom were extended by 30% and 5% of the bounding box height, respectively, to include the forehead and chin. The shorter of the two axes was extended to the length of the other axis to form a square. Cropped frames were resized to 64 × 64 pixels using bicubic interpolation. To speed up processing of the large CelebV-HQ dataset, we used MediaPipe Face Mesh for the landmarks.
[0076] Model architecture A 3D-CNN may be the architecture used. A temporal kernel width of 5 was used, and the default zero padding was replaced by repeating edges. Zero padding along the temporal dimension can introduce edge effects that add artificial frequencies to the predictions. Experiments show that internal temporal dilation can cause aliasing and reduce the model's bandwidth for certain frequencies. The loss and framework may be applied to any task and architecture with dense predictions along one or more dimensions. However, common rPPG architectures such as DeepPhys and MTTS-CAN may not be suitable for this approach because they consume a very small number of frames and the number of time points must be large enough to provide sufficient frequency resolution in the FFT.
[0077] Supervised training To properly compare this embodiment with the supervised embodiment, we used the same model architecture and trained it with negative Pearson loss between the predicted waveform and the ground truth of the contact sensor. During training, the same extensions were applied, except for time reversal. Models were trained for 200 epochs for PURE and UBFC-rPPG, and 40 epochs for DDPM. The model from the epoch with the lowest loss on the validation set was selected for testing.
[0078] Unsupervised Training Unsupervised models are trained for the same number of epochs as in the supervised setting for both PURE and UBFCrPPG, but for DDPM, they are trained for an additional 40 epochs due to the significantly more challenging nature of this dataset.Unlike previous unsupervised approaches, the validation set was utilized for model selection by selecting the model with the lowest sum of bandpass loss and sparsity loss on unseen examples.
[0079] evaluation Pulse rate is calculated as the highest spectral peak between 0.66 Hz and 3 Hz (corresponding to 40 bpm and 180 bpm) in a 10-second sliding window. For reliable evaluation, the same procedure is applied to the ground truth waveform. Common error metrics such as mean absolute error (MAE), root mean square error (RMSE), and Pearson correlation coefficient between pulse rates are applied. Five-fold cross-validation was performed for both PURE and UBFC, using the predefined dataset split of DDPM. These three models had different initial settings, resulting in training 15 models each for PURE and UBFC, and 3 models for DDPM. The mean and standard deviation of the resulting errors are shown.
[0080] Results, Dataset Test Figure 4 shows the results of models trained and tested on subject-disjoint partitions from the same dataset. For PURE and UBFC, MAEs below 1 bpm were achieved, outperforming or comparable to all traditional and supervised learning approaches. For PURE, the embodiment achieves the lowest MAE and a Pearson r of nearly 1. For DDPM, performance declines due to the overall difficulty of the dataset and noisy segments caused by fingertip oximeter movement. Nevertheless, the embodiment outperforms existing unsupervised approaches and traditional methods, and is only surpassed by supervised deep learning approaches.
[0081] Compared to other unsupervised methods, Contrast-Phys provides the most competitive performance on all datasets except DDPM. Note that the embodiment provides the lowest MAE on all datasets, but has a higher RMSE. This is likely due to the use of harmonic removal as a post-processing step in estimating pulse rate, which is not described in the publicly available code but is available in the publicly available code.
[0082] Figure 5 shows the intra-dataset waveform predictions of the baseline dataset over an 8-second window by the end-to-end unsupervised model. The model predictions are surprisingly periodic, even without filtering. Note that phase is not considered during training, so each model learns its own phase.
[0083] Dataset cross-testing In addition to traditional within-dataset experiments, we conducted cross-dataset tests to analyze whether our approach is robust to changes in lighting, camera sensors, pulse rate distribution, and motion. Figure 6 shows the results of NCUL and supervised training on the same architecture. We found that, in general, the performance of supervised and unsupervised approaches is comparable when transferred to different data sources. When transferred to UBFC-rPPG and DDPM, training on PURE alone yields relatively poor results due to the small pulse rate variability and lack of motion. Training on DDPM yields the best overall results, as it is the largest dataset and captures more subject motion compared to the other datasets.
[0084] CelebV-HQ Video Training Given the abundance of face videos publicly available online, we trained our model on faces from the CelebV-HQ dataset. After processing the downloadable videos with Mediapipe and resampling the video clips to 30 fps, the unlabeled dataset consisted of 34,029 videos. The model was trained for 23 epochs, at which point training was manually stopped as the validation loss reached a plateau. Unfortunately, we found that the model was unable to converge to the true blood volume pulse. This failure was likely due to poor data quality caused by compression. Although the videos were downloaded at the highest available quality, the videos were likely compressed multiple times, effectively removing the pulse signal entirely. The MAE (bpm) for UBFC-rPPG, PURE, and DDPM were 19.22, 24.83, and 27.41, respectively.
[0085] HKBU-MARs video training Although the HKBU-MARs dataset was designed for face-presentation attack detection, the model was trained on "real" video sessions within the dataset. The bottom panel of Figure 6 shows the results of training on HKBUMARs alone, followed by testing on the benchmark rPPG dataset. Training on HKBU-MARs outperforms all training sets except DDPM, which is an order of magnitude larger and contains motion artifacts, when transferred to UBFC-rPPG and PURE. This is the first successful experiment, to our knowledge, demonstrating that videos other than rPPG can be used to train robust models, even without ground-truth pulse labels.
[0086] Ablation studies on losses To analyze the contribution of the loss components, we trained a model using a combination of all loss components. Figure 7 shows the training and testing results on UBFC-rPPG. The bandpass loss is important for discovering the true blood volume pulse, while the sparsity loss and variance loss alone do not learn the desired signal. Surprisingly, combining the bandpass loss with only one of the sparsity loss or variance loss performs worse than the bandpass loss alone. However, when all three components are combined, the model achieves improved results.
[0087] Consideration It was initially surprising that unsupervised training could produce comparable or improved rPPG estimation models compared to supervised training. However, unsupervised training has several potential advantages. From a hardware perspective, one of the challenges with supervised training is aligning the contact pulse waveform with the video frames. There can be a time lag between the pulse sensor and the camera, effectively feeding the model out-of-phase targets during training. An advantage of unsupervised training is that it gives the model the freedom to learn phase directly from the video.
[0088] Because the PPG is highly sensitive to motion, the contact pulse signal can also be noisy. Because motion can occur simultaneously on the face and fingertips, motion noise can appear on the target and mislead the model into learning visual features that should be invariant to them.
[0089] From a physiological standpoint, the pulse observed optically at the fingertip with a contact sensor is out of phase with the facial pulse because blood travels a different path before reaching the peripheral microvasculature, making alignment nearly impossible without shifting the target on rPPG estimates from existing methods.
[0090] Furthermore, the morphological shape of the contact PPG waveform depends on numerous factors, including the wavelength of light (and corresponding tissue penetration depth), external pressure from the oximeter clip, and vasodilation at the measurement site. These external factors mean that the morphology and phase of the target PPG waveform are likely to differ from the observed rPPG waveform. Training a model to predict a proxy target introduced unnecessary artifacts and did not learn optimal features for accurate pulse measurement.
[0091] The success of the proposed asymmetric approach depends on the specific characteristics of the data, the model, and how the two interact. The limited capacity of the model is actually a strength, as it forces the features it discovers to generalize across input values. A network with infinite capacity may discover spurious signals in the training data and fail to generalize. By constraining the model's predictions to have specific periodic properties, a limited-capacity model should be able to find a common set of features to generate a signal present in all training samples, i.e., the blood volume pulse in the dataset.
[0092] As a beneficial side effect, the model inherently learns to ignore common noise contributors such as lighting, rigid motion, non-rigid motion (e.g., speaking, smiling), and camera noise, because they can result in signals outside the specified band limits or with a uniform power spectrum. Even if the noise exhibits periodic trends within the band limits in some samples, those features will produce poor signals in other samples. Therefore, end-to-end unsupervised approaches are particularly well-suited for periodic problems.
[0093] FIG. 8 shows a computer-implemented method for unsupervised learning of physiological signals from video streams.
[0094] At 810, the method captures a media stream (e.g., a video stream or an audio stream) of a subject, the video stream including a sequence of frames. The video stream may include one or more of a visible light video stream, a near-infrared video stream, and a thermal video stream of the subject. In some examples, the method may combine at least two of the visible light video stream, the near-infrared video stream, and / or the thermal video stream into a fused video stream that is processed. The visible light video stream, the near-infrared video stream, and / or the thermal video stream are combined according to a synchronizer and / or one or more video analysis techniques.
[0095] Next, at 820, the method processes each frame of the media stream to update the physiological signal detection function (eg, update the function depicted as f(·) in FIG. 2 or its weights).
[0096] At 830, the method determines physiological signals from the media stream. For example, the physiological signals may include multiple biometrics, including heart rate, pulse waveform, and / or respiration. Other examples may include pulse, gaze, eye blinking, pupil measurement, facial temperature, oxygen level, blood pressure, voice, tone and / or frequency of voice, micro-expressions, etc.) Although not shown, the updated physiological signal detection function can be easily applied to subsequent video streams.
[0097] Aside from the general performance gains, NCUL is end-to-end unsupervised, meaning it does not require subsequent fine-tuning from learned representations and can instead infer waveforms directly. The extensions are also simpler to implement and require less processing during training.
[0098] Therefore, this embodiment introduces a novel non-contrastive learning approach for end-to-end unsupervised signal regression, with concrete experiments on blood volume pulse estimation from face videos. A simple loss formulation requiring loose frequency constraints is shown to be effective in learning powerful visual features.
[0099] It will be apparent to those skilled in the art that various modifications and variations can be made in the unsupervised learning of physiological signals from video of the present invention without departing from the spirit or scope of the present invention. Therefore, it is intended that the present invention cover the modifications and variations of the present invention provided they come within the scope of the appended claims and their equivalents.
Claims
1. 1. A computer-implemented method for unsupervised learning of physiological signals from video streams, comprising: capturing a video stream of a subject, the video stream including a sequence of frames; processing each frame of the video stream to update a physiological signal detection function; determining the physiological signals of the subject from the video stream; applying the updated physiological signal detection function to a subsequent video stream.
2. 10. The computer-implemented method of claim 1, wherein the video stream comprises one or more of a visible light video stream, a near-infrared video stream, a long-wave infrared video stream, a thermal video stream, and an audio stream of the subject.
3. The computer-implemented method of claim 1 , wherein the physiological signal comprises at least one of a pulse rate, a blood pressure, and an eye blink rate.
4. The computer-implemented method of claim 1 , wherein the physiological signal includes at least one of pulse rate or audio frequency.
5. The computer-implemented method of claim 1 , further comprising cropping each frame of the media stream to encapsulate an area of interest including one or more of a face, cheeks, forehead, and eyes.
6. The computer-implemented method of claim 5 , wherein the region of interest includes two or more body parts.
7. 10. The computer-implemented method of claim 1, further comprising combining at least two of the visible light video stream, the near-infrared video stream, and the thermal video stream into a fused video stream.
8. The computer-implemented method of claim 7 , wherein the visible light video stream, the near-infrared video stream, and / or the thermal video stream are combined according to a synchronizer.
9. 1. A system for unsupervised learning of physiological signals from video streams, comprising: a processor; a memory that stores one or more programs for execution by the processor; The one or more programs: capturing a video stream of a subject, the video stream including a sequence of frames; processing each frame of the video stream to update a physiological signal detection function; determining the physiological signals of the subject from the video stream; applying the updated physiological signal detection function to a subsequent video stream.
10. 10. The system of claim 9, wherein the media stream comprises one or more of a visible light video stream, a near-infrared video stream, a long-wave infrared video stream, a thermal video stream, and an audio stream of the subject.
11. The system of claim 9 , wherein the physiological signal includes at least one of pulse rate, blood pressure, and eye blink rate.
12. The system of claim 9 , wherein the physiological signal includes at least one of pulse rate or audio frequency.
13. The system of claim 9 , further comprising cropping each frame of the media stream to encapsulate an area of interest including one or more of a face, cheeks, forehead, and eyes.
14. The system of claim 13 , wherein the region of interest includes two or more body parts.
15. 10. The system of claim 9, further comprising combining at least two of the visible light video stream, the near-infrared video stream, and the thermal video stream into a fused video stream.
16. The system of claim 15 , wherein the visible light video stream, the near-infrared video stream, and / or the thermal video stream are combined according to a synchronizer.