Conditional diffusion model-based non-contact long-term heart rate variability measurement method
By combining a conditional diffusion model with a signal reconstruction method based on chromaticity and Gaussian noise, the shortcomings of existing non-contact HRV measurement methods in signal stability and long-term modeling capability are solved. This enables high-quality HRV analysis in complex environments and is suitable for non-contact long-term physiological parameter monitoring.
Patent Information
- Application Number
- CN202511105529.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-14
AI Technical Summary
Existing non-contact HRV measurement methods have shortcomings in signal stability, waveform fidelity, and long-term modeling capabilities. In particular, they are difficult to generate high-quality, long-term rPPG signals in complex interference environments, which affects the accuracy of HRV analysis.
A conditional diffusion model-based approach is adopted, which simulates real-world noise disturbances by combining chromatic signals and Gaussian noise through forward noise addition and backward noise reduction processes. Time-frequency ridges are used as conditional guides to construct a hybrid neural network for signal reconstruction, thereby improving the robustness and fidelity of the signal.
Generating high-quality, long-duration rPPG signals under complex interference environments improves the accuracy and stability of HRV analysis, making it suitable for non-contact long-term physiological parameter monitoring and possessing high precision and wide application value.
Smart Images

Figure CN120938384A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of non-contact physiological signal detection and analysis technology, specifically to a non-contact long-term heart rate variability measurement method based on a conditional diffusion model. Background Technology
[0002] Heart rate variability (HRV) is an important physiological indicator for assessing the function of the autonomic nervous system and has significant applications in clinical diagnosis, health monitoring, and emotion recognition. Traditional HRV measurement methods rely on contact electrode devices, such as electrocardiogram (ECG) machines and pulse oximeters. These devices may cause discomfort with prolonged use and are not suitable for special populations such as infants and burn patients.
[0003] In recent years, remote photoplethysmography (rPPG) based on video analytics has emerged as a potential solution for HRV monitoring due to its non-contact nature, convenience, and ease of implementation. rPPG technology captures subtle color changes in facial skin caused by heartbeats using a regular camera, thereby extracting blood volume pulse (BVP) signals. However, high-quality HRV analysis requires obtaining high-fidelity BVP signals over extended periods, a challenge that current rPPG technologies face. In practical applications, rPPG signals are susceptible to interference from factors such as ambient lighting changes, facial motion artifacts, and individual skin color differences, leading to unstable signal quality and consequently affecting the accuracy of subsequent inter-heart beat interval (IBI) sequence extraction.
[0004] Traditional rPPG methods, such as blind source separation and model-based methods, have achieved good results in mean heart rate estimation, but they still have significant shortcomings in maintaining BVP waveform fidelity and stably extracting high-quality IBI sequences. In recent years, the application of deep learning technology has significantly improved the quality of BVP signals, but most deep learning methods still focus on short-term heart rate estimation or HRV calculation within a limited time window, lacking the ability to model long-term BVP signals, and thus failing to meet the practical needs of continuous HRV monitoring.
[0005] rPPG signal extraction is essentially a problem of reconstructing complex physiological waveforms. In recent years, generative models have gained increasing attention in the rPPG field, and Generative Adversarial Networks (GANs) have achieved initial results in several rPPG tasks. However, GANs suffer from problems related to training stability and pattern collapse. In contrast, diffusion models, with their stable training process and superior target distribution modeling capabilities, have become a potential new direction for rPPG signal reconstruction. Diffusion models recover signals step by step through multi-step iterations, maintaining high signal fidelity. In particular, their flexible conditional modeling capabilities provide new opportunities for developing rPPG enhancement methods oriented towards HRV analysis. However, existing diffusion-based rPPG methods typically assume that the noise during the diffusion process follows a simple Gaussian distribution. This cannot accurately simulate the common non-Gaussian, non-stationary, and structurally complex interference patterns in rPPG signals, limiting the realism and generalization ability of the generated results. Summary of the Invention
[0006] This invention aims to overcome the shortcomings of existing non-contact HRV measurement methods in terms of signal stability, waveform fidelity, and long-term modeling capabilities. It proposes a non-contact long-term HRV measurement method based on a conditional diffusion model, which aims to stably generate high-quality, long-term rPPG signals from videos under complex interference environments, so as to further accurately estimate long-term HRV indicators. This provides a high-precision, robust, and universal solution for non-contact long-term physiological parameter monitoring.
[0007] The present invention adopts the following solution to solve the technical problem: The present invention provides a non-contact long-term heart rate variability measurement method based on a conditional diffusion model, characterized by the following steps: Step 1: Acquire face video data and preprocess it to obtain the chroma signal y in the video; Step 2: Gradually add noise to the acquired reference BVP signal through the forward noise addition process in the conditional diffusion model. By adding Gaussian noise and the chromaticity signal y, the noise signal at time step t is obtained. ; Step 3: Design a hybrid neural network consisting of a feature fusion module and multiple Mamba modules, and then... The time-frequency ridge h of y and time step t are processed to obtain the noise residual estimate at time step t. ; Step 4: Construct the loss function for the hybrid neural network This model is used to train the hybrid neural network, obtaining the optimal hybrid neural model, which is then used to output the optimal noise residual estimate. ; Step 5: Estimating the optimal noise residual Through the backward denoising process in the conditional diffusion model, and guided by the chrominance signal y and the time-frequency ridge h extracted by its wavelet synchronous compression transform, a high-quality reconstructed signal is gradually recovered. .
[0008] The non-contact long-term heart rate variability measurement method based on the conditional diffusion model described in this invention is also characterized in that step one includes: Step 1.1 Process the faces in each frame of the face video data to construct the facial region of interest; Step 1.2 Extract the original BVP signal from each facial region of interest using a chroma algorithm, and perform weighted fusion of the original BVP signals from all facial regions of interest to obtain the fused chroma signal; Step 1.3 Perform detrending processing and bandpass filtering on the fused chroma signal to obtain the preprocessed chroma signal y.
[0009] Furthermore, in step two, the noise signal at time step t in the forward noise addition process is obtained using equation (1). : (1) In equation (1), Indicates the reference BVP signal. Indicates the first t The noise signal at the time step, where y represents the preprocessed chroma signal. For the first t The interpolation function for the time step, and 0 ≤ ≤ 1; For the first step of the forward noise addition process t The product coefficient accumulated over time steps, For the first step of the forward noise addition process t The variance term of the time step, where I is the identity matrix. Indicates a Gaussian distribution. This represents the probability distribution during the forward noise addition process.
[0010] Furthermore, step three includes: Step 3.1 Time-encode the nth time step and then match the encoded nth time step with... The frequency domain features h of y are dimensionally aligned and fused to generate a fused multi-source feature representation. Step 3.2 Input the fused multi-source feature representation into the backbone network composed of multiple stacked Mamba modules for state-space modeling, and output the noise residual estimate at time step t. .
[0011] Furthermore, in step four, the loss function is constructed using equation (2). : (2) In equation (2), and There are two constant terms. for Gaussian noise in Let represent the target noise term at time step t, and we have: (3). Furthermore, in step five, the noise signal at time step t-1 in the backward denoising process is obtained using equation (4). Ultimately, high-quality reconstruction signals are gradually obtained. : (4) In equation (4), This represents the noise signal at time step t-1. For the first step in the reverse denoising process t The variance term of the time step, ℎ represents the time-frequency ridge extracted from y. This represents the probability distribution during the reverse denoising process; Let represent the mean function at time step t in the conditional diffusion model, and we have: (5) In equation (5), , , For the first t The three weighting coefficients corresponding to the time step.
[0012] The present invention provides an electronic device, comprising a memory and a processor, wherein the memory is used to store a program that supports the processor in executing the non-contact long-term heart rate variability measurement method, and the processor is configured to execute the program stored in the memory.
[0013] The present invention discloses a computer-readable storage medium on which a computer program is stored, wherein the computer program, when executed by a processor, performs the steps of the non-contact long-term heart rate variability measurement method.
[0014] Compared with existing technologies, the beneficial effects of this invention are reflected in: 1. In the forward diffusion process of this invention, chroma signals extracted from the video and Gaussian noise are added to simulate noise disturbances in real scenes, which enhances the robustness of the model to complex disturbances such as motion and lighting. 2. In the reverse denoising process of this invention, the time-frequency ridge line extracted by the chroma signal and its wavelet synchronous compression transform (WSST) is used as a conditional guide, which effectively captures the rhythmic changes in the rPPG signal and further improves the temporal consistency and waveform fidelity of the signal.
[0015] 3. The model used in this invention can use short time-term samples during training and can directly output continuous rPPG signals of arbitrary length during inference, avoiding the timing distortion caused by segment splicing. 4. This invention provides a high-precision, stable, and non-contact long-term HRV measurement method for clinical health monitoring, stress assessment, and remote electrocardiogram analysis, and has broad application value. Attached Figure Description
[0016] Figure 1 This is a flowchart of the method of the present invention; Figure 2 The method of this invention uses WSST to extract the time-frequency ridge plot of the chroma signal; Figure 3 This is a schematic diagram of the LTHRV dataset acquisition scenario using the method of this invention; Figure 4 The following is a comparison of IBI sequences on the LTHRV dataset according to the present invention: (a) normal, (b) apnea. Detailed Implementation
[0017] In this embodiment, the non-contact long-term heart rate variability measurement method based on the conditional diffusion model includes the following steps: Step 1, such as Figure 1 The flowchart of BVP signal extraction and preprocessing in part a is shown. The face video data is acquired and preprocessed to obtain the chrominance signal y in the video. Step 1.1 Process the face in each frame of the face video data, using the MediaPipe facial landmark detection algorithm to extract 145 landmarks, covering areas such as the forehead, cheekbones, and sides of the neck, to more comprehensively capture blood flow signal features. Then, construct a region of interest (ROI) centered on each landmark.
[0018] Step 1.2 Extract the original BVP signal from each facial region of interest using a chroma algorithm, and perform weighted fusion of the original BVP signals from all facial regions of interest to obtain the fused chroma signal; Step 1.3 Perform detrending processing and sixth-order Butterworth bandpass filtering (frequency band 0.65–2.5 Hz) on the fused chroma signal to obtain the preprocessed chroma signal y.
[0019] Step 2: Gradually add noise to the acquired reference BVP signal through the forward noise addition process in the conditional diffusion model. By adding Gaussian noise and the chromaticity signal y, the noise signal at time step t is obtained. ; In this example, such as Figure 1 The flowcharts for the forward diffusion process and the reverse denoising process of the conditional diffusion model in part b are shown. First, the reference BVP signal is... The signal is weighted and mixed with the chrominance signal y, and Gaussian noise is introduced. The signal is then subjected to forward diffusion processing, and the noise signal at time step t in the forward diffusion process is obtained using equation (1). : (1) In equation (1), Indicates the reference BVP signal. Indicates the first t The noise signal at the time step, where y represents the preprocessed chroma signal. For the first t The interpolation function for the time step, and 0 ≤ ≤ 1; For the first step of the forward noise addition process t The product coefficient accumulated over time steps, For the first step of the forward noise addition process t The variance term of the time step, where I is the identity matrix. Indicates a Gaussian distribution. This represents the probability distribution during the forward noise addition process. In this example, in addition to injecting Gaussian noise, the chromaticity signal y is also introduced during the forward noise addition process to construct a conditional distribution. Interpolation function. The mixing ratio of the reference signal and the chroma signal is controlled and gradually increased from 0 to 1, so that the early noise addition stage is closer to the reference signal, and the chroma features are gradually incorporated in the later stage, realizing the transition modeling from clean signal to real rPPG noise interference.
[0020] Step 3: Design a hybrid neural network consisting of a feature fusion module and multiple Mamba modules, and then... The time-frequency ridge h of y and time step t are processed to obtain the noise residual estimate at time step t. ; Step 3.1 In this example, time-encode the nth time step and then compare the encoded nth time step with... The frequency domain features h of y are dimensionally aligned and fused to generate a fused multi-source feature representation. The frequency domain features h are obtained by applying wavelet synchronous compression transform (WSST) to the chrominance signal y, representing the main ridge information and the most concentrated heart rate frequency trajectory in the signal. Figure 2 As shown, it stably tracks instantaneous frequency changes, which is crucial for capturing rhythmic characteristics.
[0021] Step 3.2 In this example, as Figure 1 The diagram of the one-dimensional convolutional and Mamba hybrid network structure in part c shows that the fused multi-source feature representation is input into the backbone network consisting of 8 stacked Mamba modules for state space modeling, and the noise residual estimate at time step t is output. .
[0022] Specifically, after performing layer normalization on the input multi-source fusion features, the data is processed by the Mamba module, which includes a state-space modeling mechanism. This module constructs a continuous-time dynamic system model to model the temporal changes between different moments in the feature sequence, extracting long-range dependent features reflecting the rhythm evolution, thereby capturing the key dynamics of heart rate over time. During the modeling process, the input time-series data undergoes multi-scale segmentation, dividing the time axis into segments of 1 / 2, 1 / 4, and 1 / 8 respectively, and then inputting each segment into the Mamba module for sub-interval modeling, resulting in multi-scale local state representations. After fusing these results, residual connection and normalization operations are performed to obtain the enhanced state representation.
[0023] Finally, a nonlinear transformation and channel compression are performed on the above enhanced state representation to obtain the noise residual estimation result at time step t. This provides crucial prediction information for the subsequent reverse denoising process.
[0024] Step 4: Construct the loss function for the hybrid neural network This model is used to train the hybrid neural network, obtaining the optimal hybrid neural model, which is then used to output the optimal noise residual estimate. The loss function is constructed using equation (2). : (2) In equation (2), and There are two constant terms. for Gaussian noise in Let represent the target noise term at time step t, and we have: (3) The loss function is derived from the Evidence Lower Bound (ELBO) in variational inference, and its goal is to minimize the noise residual of the network output. residuals of the actual constructed target Mean squared error (MSE) between the target residuals. It also considers the difference between the reference signal and the chromaticity signal, as well as the Gaussian noise term, to more realistically simulate the diffusion disturbance.
[0025] Step 5, in this example, as Figure 1 The flowcharts for the forward diffusion process and reverse denoising process of the conditional diffusion model in part b are shown, based on the optimal noise residual estimation. Through the backward denoising process in the conditional diffusion model, and guided by the chrominance signal y and the time-frequency ridge h extracted by its wavelet synchronous compression transform, a high-quality reconstructed signal is gradually recovered. .
[0026] The noise signal at time step t-1 in the backward denoising process is obtained using equation (4). Ultimately, high-quality reconstruction signals are gradually obtained. : (4) In equation (4), This represents the noise signal at time step t-1. For the first step in the reverse denoising process t The variance term of the time step, ℎ represents the time-frequency ridge extracted from y. This represents the probability distribution during the reverse denoising process; Let represent the mean function at time step t in the conditional diffusion model, and we have: (5) In equation (5), , , For the first t The three weighting coefficients corresponding to the time step.
[0027] This invention is implemented based on the PyTorch framework and trained on an NVIDIA A6000 GPU. During model training, a linear noise scheduling strategy is used, with a total number of steps set to 50 and a noise range of [1×10⁻⁻⁴]. 4 [0.35], the batch size during training is set to 64, and the initial learning rate is 2×10⁻ 4 A total of 60,000 training iterations were performed. During the inference phase, a fast sampling strategy was used to process the test signal, and the sampling noise sequence used was: [1×10⁻ 4, 1×10⁻³, 1×10⁻², 5×10⁻², 0.2, 0.35).
[0028] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the above-described method, and the processor is configured to execute the program stored in the memory.
[0029] In this embodiment, a computer-readable storage medium stores a computer program, which is executed by a processor to perform the steps of the above method.
[0030] To verify the effectiveness of the method of this invention, we collected the only currently available long-term video HRV dataset, LTHRV. The collection scenario is as follows: Figure 3As shown, the dataset contains 36 videos from 18 subjects (12 males and 6 females), each 15 minutes long, covering two typical physiological states: normal breathing and apnea. Tables 1 and 2 show the performance comparison of the HRVFusion method of this invention with several existing representative deep learning methods in the HRV measurement task, covering the mean absolute error (MAE) and Pearson correlation coefficient (r) of time and frequency domain indicators.The comparison methods include: TSCAN, from the paper "Multi-task temporal shift attention networks for on-device contactless vitals measurement" published at the 2020 Advances in Neural Information Processing Systems, which models the dynamic relationship between video and pulse signals based on spatiotemporal convolution; PhysFormer, from the paper "Physformer: Facial video-based physiological measurement with temporal difference transformer" published at the 2022 IEEE International Conference on Computer Vision, which enhances rhythm modeling capabilities by introducing a temporal difference mechanism and a Transformer structure; DeepPhys, from the paper "Deepphys: Video-based physiological measurement using convolutional attention networks" published at the 2018 European Conference on Computer Vision, which uses convolutional attention networks to model facial regions and regresses physiological fluctuations based on video signals; and EfficientPhys, from the paper "Efficientphys: Enabling simple, fast and accurate camera-based..." The paper "cardiac measurement" was published at the 2023 IEEE Winter Conference on Applications of Computer Vision, highlighting a lightweight and efficient network structure to improve measurement accuracy under low-resource conditions. "RhythmFormer" comes from the paper "Rhythmformer: Extracting patterned rppgsignals based on periodic sparse attention" published in the journal Pattern Recognition in 2025. It integrates a frequency modulation modeling mechanism to capture more stable heart rate rhythm changes.The following HRV indicators were used for comparative evaluation: SDNN measures the standard deviation of overall heart rate fluctuations, reflecting the overall balance of the sympathetic-parasympathetic nervous system; RMSSD is the average of the square root of the difference between consecutive heartbeat intervals, mainly reflecting parasympathetic nerve activity; pNN50 is the proportion of adjacent RR interval differences greater than 50ms, measuring the intensity of heart rate rhythm changes; LF and HF correspond to the frequency domain characteristics of sympathetic / parasympathetic activity, respectively; the LF / HF ratio measures the sympathetic-parasympathetic balance in the autonomic nervous system.
[0031] Table 1: HRV assessment results in normal breathing scenarios on the LTHRV dataset Table 2: HRV assessment results for apnea scenarios on the LTHRV dataset As can be seen from Tables 1 and 2, under both normal breathing and breath-holding scenarios, the method of the present invention outperforms other methods in all indicators in the time and frequency domains, especially in key indicators reflecting heart rate variability such as SDNN and RMSSD, demonstrating the good robustness and physiological consistency of the method of the present invention in long-term HRV measurement tasks.
[0032] also, Figure 4 The reconstruction results of the 5-minute IBI sequence under two respiratory conditions are presented, demonstrating that the method of the present invention can effectively reproduce the heart rate fluctuation details in the reference signal, showing high reliability and physiological relevance in HRV measurement.
Claims
1. A non-contact long-term heart rate variability measurement method based on a conditional diffusion model, characterized in that, Includes the following steps: Step 1: Acquire face video data and preprocess it to obtain the chroma signal y in the video; Step 2: Gradually add noise to the acquired reference BVP signal through the forward noise addition process in the conditional diffusion model. By adding Gaussian noise and the chromaticity signal y, the noise signal at time step t is obtained. ; Step 3: Design a hybrid neural network consisting of a feature fusion module and multiple Mamba modules, and then... The time-frequency ridge h of y and time step t are processed to obtain the noise residual estimate at time step t. ; Step 4: Construct the loss function for the hybrid neural network This model is used to train the hybrid neural network, obtaining the optimal hybrid neural model, which is then used to output the optimal noise residual estimate. ; Step 5: Estimating the optimal noise residual Through the backward denoising process in the conditional diffusion model, and guided by the chrominance signal y and the time-frequency ridge h extracted by its wavelet synchronous compression transform, a high-quality reconstructed signal is gradually recovered. .
2. The non-contact long-term heart rate variability measurement method based on the conditional diffusion model according to claim 1, characterized in that, Step one includes: Step 1.1 Process the faces in each frame of the face video data to construct the facial region of interest; Step 1.2 Extract the original BVP signal from each facial region of interest using a chroma algorithm, and perform weighted fusion of the original BVP signals from all facial regions of interest to obtain the fused chroma signal; Step 1.3 Perform detrending processing and bandpass filtering on the fused chroma signal to obtain the preprocessed chroma signal y.
3. The non-contact long-term heart rate variability measurement method based on the conditional diffusion model according to claim 2, characterized in that, In step two, the noise signal at time step t in the forward noise addition process is obtained using equation (1). : (1) In equation (1), Indicates the reference BVP signal. Indicates the first t The noise signal at the time step, where y represents the preprocessed chroma signal. For the first t The interpolation function for the time step, and 0 ≤ ≤ 1; For the first step of the forward noise addition process t The product coefficient accumulated over time steps, For the first step of the forward noise addition process t The variance term of the time step, where I is the identity matrix. Indicates a Gaussian distribution. This represents the probability distribution during the forward noise addition process.
4. The non-contact long-term heart rate variability measurement method based on the conditional diffusion model according to claim 3, characterized in that, Step three includes: Step 3.1 Time-encode the nth time step and then match the encoded nth time step with... The frequency domain features h of y are dimensionally aligned and fused to generate a fused multi-source feature representation. Step 3.2 Input the fused multi-source feature representation into the backbone network composed of multiple stacked Mamba modules for state-space modeling, and output the noise residual estimate at time step t. .
5. The non-contact long-term heart rate variability measurement method based on the conditional diffusion model according to claim 4, characterized in that, In step four, the loss function is constructed using equation (2). : (2) In equation (2), and There are two constant terms. for Gaussian noise in Let represent the target noise term at time step t, and we have: (3)。 6. The non-contact long-term heart rate variability measurement method based on the conditional diffusion model according to claim 5, characterized in that, In step five, the noise signal at time step t-1 in the backward denoising process is obtained using equation (4). Ultimately, high-quality reconstruction signals are gradually obtained. : (4) In equation (4), This represents the noise signal at time step t-1. For the first step in the reverse denoising process t The variance term of the time step, ℎ represents the time-frequency ridge extracted from y. This represents the probability distribution during the reverse denoising process; Let represent the mean function at time step t in the conditional diffusion model, and we have: (5) In equation (5), , , For the first t The three weighting coefficients corresponding to the time step.
7. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports the processor in executing the non-contact long-term heart rate variability measurement method according to any one of claims 1-6, the processor being configured to execute the program stored in the memory.
8. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is run by the processor, it performs the steps of the non-contact long-term heart rate variability measurement method according to any one of claims 1-6.