Heart rate variability estimation method and system based on remote face video

Through the cascade structure and dynamic phase correlation loss function optimization of the convolutional layer and the convolutional layer in space-time separation, the motion artifact and lighting interference problems in the estimation of center rate variability of remote face videos are solved, and efficient and accurate heart rate variability measurement is achieved.

CN120267264APending Publication Date: 2025-07-08QINGDAO UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510568879.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art has problems of motion artifacts and light interference in the estimation of center rate variability of remote face videos, resulting in signal waveform distortion, making it difficult to achieve robust contactless heart rate variability measurement.

Method used

The rPPG signal extraction model based on the cascade structure of the spatiotemporal separation convolution layer and the timing variability convolution layer is adopted, combined with the optimization of dynamic phase correlation loss function, and the influence of motion artifacts and illumination changes are suppressed through the 3D convolution layer, the spatiotemporal feature extraction module, the 3D transposed convolution layer and the maximum pooling layer.

Benefits of technology

It improves the accuracy of heart rate variability estimation, can accurately identify slight changes in heartbeat intervals in complex environments, effectively suppresses noise interference, and improves the accuracy of signal waveform evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120267264A_ABST
    Figure CN120267264A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of heart rate variability, and provides a heart rate variability estimation method and system based on a remote face video, and the method comprises the steps: collecting remote face video data, and carrying out the preprocessing of the remote face video data; real PPG signals are collected in a synchronous contact mode; constructing an rPPG signal extraction model, wherein the rPPG signal extraction model comprises a 3D convolution layer, a spatial-temporal feature extraction module, a 3D transpose convolution layer and a maximum pooling layer; the space-time feature extraction module is of a cascade structure of a space-time separation convolution layer and a time sequence variable convolution layer; extracting the model from the input rPPG signal to obtain an rPPG signal; and performing signal processing on the rPPG signal to obtain an estimated value of the heart rate variability related index. And optimizing the rPPG signal extraction model through a dynamic phase correlation loss function. Robust rPPG signal extraction resistant to motion artifacts and illumination interference is remotely achieved from complex video data, and the accuracy of heart rate variability estimation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of heart rate variability, and particularly to a method and system for estimating heart rate variability based on remote face video. Background Art

[0002] The statements in this section merely provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] Among various physiological signals, heart rate variability (HRV), as an important indicator reflecting the health status of the human body, has always received extensive attention in medical research and clinical practice. Heart rate variability refers to the variation in the time intervals between heartbeats. It is not only an important indicator of the function of the autonomic nervous system but also an important physiological marker for evaluating individual health and predicting disease risks. For example, lower heart rate variability is usually associated with diseases such as cardiovascular diseases, depression, and anxiety disorders, while higher heart rate variability usually indicates good autonomic nerve regulation ability and health status. Therefore, through the real-time monitoring of heart rate variability, it is possible to provide strong support for individual health assessment, disease prevention, treatment effect evaluation, etc.

[0004] The measurement of heart rate variability usually relies on contact devices, such as electrocardiogram (ECG), photoplethysmogram (PPG), etc. ECG measurement has high accuracy and stability and is particularly suitable for professional diagnosis in a medical environment. PPG measures the blood flow changes in the skin through an optical sensor (such as an LED lamp and a photodetector) and then estimates the heart rate and heart rate variability. Although ECG and PPG are currently the most common contact physiological signal measurement technologies, they have obvious deficiencies in terms of comfort, convenience, adaptability, etc. Especially in the scenario of continuous health monitoring in daily life, wearing electrodes or sensors may not only affect the user's freedom of movement but also cause discomfort in some cases. Therefore, how to achieve non-contact, comfortable, and efficient heart rate variability measurement has become an urgent problem to be solved in the current field of intelligent health monitoring.

[0005] To break through the physical contact barrier, remote photoplethysmography (rPPG) emerged as the times require. Based on the principle of optical imaging, this technology extracts hemodynamic signals caused by heartbeats without contact by analyzing the periodic minute color changes on the surface of the human skin in the video. Its physical mechanism stems from the pulsed increase in subcutaneous capillary blood volume during heart contraction, resulting in enhanced absorption of light with a specific wavelength by the skin, and the light intensity reflected to the camera then fluctuates rhythmically. Since the Verkruysse team first extracted remote PPG signals (i.e., rPPG signals) from face videos in 2008, this technology has been extended to the measurement of multiple parameters such as respiratory rate and blood oxygen saturation, and is regarded as a revolutionary breakthrough in intelligent health monitoring.

[0006] However, regardless of which physiological parameter is measured, a robust and accurate rPPG signal needs to be extracted from the face video. Although rPPG, as a non-contact method, solves the contact problem, there are still challenges in practical applications. For example, during the face video acquisition process:

[0007] (1) Voluntary movements of the subject (such as head rotation, facial expression changes) and environmental movements (such as camera jitter) will introduce severe motion artifacts, resulting in signal waveform distortion;

[0008] (2) Sudden changes in environmental light intensity (such as shadow switching, screen reflection) and heterogeneous light sources (such as mixed lighting with multiple color temperatures) will mask weak blood flow signals. Summary of the Invention

[0009] The purpose of the present invention is to provide a method and system for estimating heart rate variability based on remote face videos, aiming to remotely extract a robust rPPG signal with anti-motion artifact and anti-light interference from complex video data and improve the accuracy of heart rate variability estimation.

[0010] To achieve the above purpose, the present invention adopts the following technical solutions:

[0011] The first aspect of the present invention provides a method for estimating heart rate variability based on remote face videos, including:

[0012] Collect remote face video data and preprocess the remote face video data; synchronously collect real PPG signals by contact;

[0013] Construct an rPPG signal extraction model;

[0014] Input the preprocessed remote face video data into the rPPG signal extraction model to obtain an rPPG signal;

[0015] Process the rPPG signal to obtain an estimated value of the heart rate variability-related index;

[0016] Among them, the rPPG signal extraction model includes a 3D convolutional layer, a spatio-temporal feature extraction module, a 3D transposed convolutional layer, and a max pooling layer connected in sequence; the spatio-temporal feature extraction module is a cascaded structure of a spatio-temporal separable convolutional layer and a temporal variability convolutional layer.

[0017] The second aspect of the present invention provides a heart rate variability estimation system based on a remote face video, including:

[0018] A data acquisition module, configured to: acquire remote face video data, preprocess the remote face video data; synchronously acquire a real PPG signal in a contact manner;

[0019] A model construction module, configured to: construct an rPPG signal extraction model;

[0020] A signal recovery module, configured to: input the preprocessed remote face video data into the rPPG signal extraction model to obtain an rPPG signal;

[0021] A signal processing module, configured to: process the rPPG signal to obtain an estimated value of a heart rate variability related index;

[0022] Among them, the rPPG signal extraction model includes a 3D convolutional layer, a spatio-temporal feature extraction module, a 3D transposed convolutional layer, and a max pooling layer connected in sequence; the spatio-temporal feature extraction module is a cascaded structure of a spatio-temporal separable convolutional layer and a temporal variability convolutional layer.

[0023] The third aspect of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the steps in a heart rate variability estimation method based on a remote face video as described in the first aspect of the present invention are implemented.

[0024] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor, and when the processor executes the program, the steps in a heart rate variability estimation method based on a remote face video as described in the first aspect of the present invention are implemented.

[0025] The technical solution of the present invention has the following beneficial effects:

[0026] (1) Through the cascaded structure of the spatio-temporal separable convolutional layer and the temporal variability convolutional layer, the present invention allows the model to capture local spatial patterns and time dependencies separately, so as to more efficiently learn the spatio-temporal features in the rPPG signal.

[0027] (2) The time-varying deformable convolutional layer of the present invention can improve the time-series modeling ability of rPPG signals without increasing significant computational complexity, enabling the model to accurately identify minute changes in the heart rate interval and effectively suppressing time-series noise interference caused by motion artifacts, light changes, etc.

[0028] (3) The present invention jointly optimizes the dynamic time warping loss and the negative Pearson correlation loss with phase constraints through the dynamic phase correlation loss. The dynamic time warping loss is used to calculate the optimal alignment path between the predicted rPPG signal and the true PPG signal. The negative Pearson correlation loss introduces a phase penalty term based on cross-correlation analysis to reduce the phase error between the predicted rPPG signal and the true PPG signal, further solving the problems of motion artifacts and sudden changes in the lighting environment and more accurately evaluating the signal waveform.

[0029] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0031] Figure 1 It is a flowchart of the heart rate variability estimation method in Embodiment 1 of the present invention;

[0032] Figure 2 (a) It is a schematic diagram of the 3D convolutional layer structure of the rPPG signal extraction model in Embodiment 1 of the present invention;

[0033] Figure 2 (b) It is a schematic diagram of the spatio-temporal feature extraction module structure of the rPPG signal extraction model in Embodiment 1 of the present invention;

[0034] Figure 2 (c) It is a schematic diagram of the 3D transposed convolutional layer structure of the rPPG signal extraction model in Embodiment 1 of the present invention;

[0035] Figure 3 It is a schematic diagram of the steps of the spatio-temporal separable convolutional layer in Embodiment 1 of the present invention;

[0036] Figure 4 It is a schematic diagram of the steps of the time-varying deformable convolutional layer in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0038] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0039] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0040] Embodiment 1

[0041] As Figure 1 shown, this embodiment discloses a method for estimating heart rate variability based on remote face video; including:

[0042] Step 1: Collect remote face video data, preprocess the remote face video data, and divide it into a training set, a validation set, and a test set. Synchronously collect real PPG signals in a contact way.

[0043] The preprocessing includes locating the face area, cropping, selecting the resolution of the face area, determining the sample length, and slicing to increase the number of samples. Specifically:

[0044] First, determine the facial region of interest through the 68 facial feature point detection method, select 72×72 pixels as the resolution size of the facial region of interest, effectively retain the blood flow characteristics in the data and suppress noise interference; while ensuring the capture of a complete heartbeat cycle, balance the computational complexity, and determine 128 frames as the optimal video sample length; after determining the video sample length, perform slicing processing on the original video, and adopt the method of overlapping sliding windows to increase the number of samples; the preprocessed data is divided into a training set, a validation set, and a test set according to 6:2:2.

[0045] Step 2: Build an rPPG signal extraction model and train it through the training set and the validation set.

[0046] As Figure 1 、 Figure 2 shown, this embodiment proposes an rPPG signal extraction model DST-rPPGNet (Deformable Spatio Temporal rPPG Network) based on spatio-temporal dynamic modeling, which realizes efficient feature decoupling through a spatio-temporal feature extraction module.

[0047] The structure of the rPPG signal extraction model DST-rPPGNet is as follows:

[0048] Based on the 3DCNN encoder-decoder structure, the training set and the validation set sequentially pass through a 3D convolutional layer (Conv3D), multiple spatio-temporal feature extraction modules (ST Block), multiple 3D transposed convolutional layers (ConvTrans3D), and a max pooling layer (MaxPool) to fully model the time-space information and extract high-quality rPPG signals. Specifically,

[0049] 3D convolutional layer (Conv3D): It includes a convolutional layer, a 3D batch normalization layer, and a ReLu function.

[0050] In the encoding stage, the convolutional layer extracts the local spatio-temporal features of the facial video data through a sliding spatio-temporal convolutional kernel; the 3D batch normalization layer normalizes the channels of the local spatio-temporal features to accelerate the training convergence; the ReLu function applies non-linear activation to enhance the model's expression ability. The three cooperate in sequence to form the core module of the encoder and jointly optimize the spatio-temporal feature extraction.

[0051] That is, the training set and the validation set are input into the 3D convolutional layer to extract spatio-temporal correlation, and the local spatio-temporal features are output to the spatio-temporal feature extraction module;

[0052] Spatio-temporal feature extraction module (ST Block): After the local spatio-temporal features are input into the spatio-temporal feature extraction module, the time and space features are decoupled, and after dynamically adjusting the spatio-temporal dimensions, the deep spatio-temporal fusion features are output to the 3D transposed convolutional layer;

[0053] It should be noted that the spatio-temporal feature extraction module, as the core module of the rPPG signal extraction model DST-rPPGNet, is a cascaded structure of a spatio-temporal separation convolutional layer and a temporal variability convolutional layer.

[0054] As Figure 3 shown, spatio-temporal separation convolutional layer:

[0055] First, a 1×3×3 depthwise separable spatial convolution is used to extract local features of the output of the 3D convolutional layer in the spatial dimension (H×W) while keeping the time dimension unchanged; then a 3×1×1 depthwise separable temporal convolution is used to perform sequence modeling in the time dimension (T) while keeping the spatial dimension unchanged.

[0056] The spatio-temporal separation convolutional layer decouples the spatial and time features, outputs the feature tensor [C, T, H, W], and inputs it into the temporal variability convolutional layer; where C is the number of channels, T is the time dimension, H is the height, and W is the width.

[0057] This approach allows the model to capture local spatial patterns and temporal dependencies separately, reducing computational redundancy and enhancing the model's independent modeling ability for static details and dynamic changes, thus enabling more efficient learning of spatio-temporal features in rPPG signals.

[0058] As Figure 4 shown, the temporal deformable convolutional layer:

[0059] In this embodiment, the temporal deformable convolution consists of two core parts: temporal offset prediction and dynamic feature sampling. The specific steps are as follows:

[0060] (1) Initialize the feature tensor [C, T, H, W] output by the spatio-temporal separation convolutional layer.

[0061] (2) Input the feature tensor [C, T, H, W] into the offset convolutional layer to predict the offset for each time step, and output the offset tensor [kernel_size, T(δ1, δ2, δ3, ……, δT), H, W]. Here, kernel_size is the convolutional kernel size; δ1, δ2, δ3, ……, δT are the offsets. In this embodiment, this offset convolutional layer is an ordinary 3D convolutional layer used to obtain the offsets, with a convolutional kernel of 3×1×1.

[0062] This step enables the learning of offsets to act only on the time dimension without affecting spatial features, thus ensuring that the model focuses on modeling temporal dependencies.

[0063] (3) Introduce an index mapping mechanism to dynamically sample temporal features. Specifically:

[0064] (3.1) Generate an index matrix [kernel_size, T].

[0065] (3.2) For each time step of the feature tensor, index with the central time point t (the current processing position during the sampling process) as the reference, and dynamically adjust the index according to the predicted offset δT to ensure that the convolutional kernel can adaptively capture the most relevant temporal information.

[0066] The prediction of offsets enables the model to adaptively adjust the sampling position according to the dynamic changes of the input features, rather than performing sliding convolution within a fixed time window.

[0067] (3.3) Dynamically sample and traverse each convolutional kernel position. To efficiently implement the sampling of features, based on the index matrix, directly extract the dynamically adjusted time step features in the time dimension instead of using an explicit loop to traverse each data point of each time step.

[0068] (3.4) Concatenate all the features sampled by indexing in the channel dimension, finally forming an extended feature tensor [C×kernel_size, T, H, W]. That is, at each time step, features at multiple different time offsets are fused, enhancing the model's ability to capture short-term and long-term time dependencies.

[0069] (4) After dynamically sampling the temporal features, the extended feature tensor [C×kernel_size, T, H, W] is input into the fusion convolutional layer for feature fusion, and deep spatio-temporal fusion features are output; the number of input channels of the fusion convolutional layer is equal to the original channel number C multiplied by the kernel size, ensuring that the information in the time dimension is fully utilized.

[0070] 3D Transposed Convolutional Layer (ConvTrans3D): It includes a transposed convolutional layer, a 3D batch normalization layer, and a ReLu function.

[0071] In the decoding stage, the transposed convolutional layer expands the temporal features of the deep spatio-temporal fusion features through deconvolution operations; the 3D batch normalization layer normalizes the temporal features across spatio-temporal dimensions to stabilize the training process; the ReLu function enhances the feature expression ability through non-linear activation and suppresses invalid negative values. The three components form the core module of the decoder, gradually restoring high-resolution spatio-temporal features, and finally supporting the reconstruction of high-fidelity rPPG signals from the compressed features, providing an accurate spatio-temporal information basis for heart rate variability analysis.

[0072] That is, the deep spatio-temporal fusion features are input into the 3D transposed convolutional layer, and the time dimension is expanded through deconvolution, and high-resolution spatio-temporal features are output to the max pooling layer;

[0073] Max Pooling Layer (MaxPool): As the last layer of the model, the max pooling layer has the key functions of spatio-temporal feature dimensionality reduction and signal compression. Global max pooling is used to take the maximum value in the spatial dimensions (H and W), eliminating spatial redundancy and retaining the most significant features at each time point. In this way, the four-dimensional tensor is converted into a one-dimensional time series.

[0074] That is, the high-resolution spatio-temporal features are input into the max pooling layer, and the high-dimensional features are converted into a one-dimensional time feature sequence through spatio-temporal feature dimensionality compression.

[0075] Through the above steps, the temporal deformable convolutional layer can improve the temporal modeling ability of the rPPG signal without increasing significant computational complexity, enabling the model to accurately identify small changes in the heart rate interval and effectively suppress temporal noise interference caused by motion artifacts, light changes, etc.

[0076] Step 3: Optimize the rPPG signal extraction model through a dynamic phase correlation loss function.

[0077] During the model construction process, the dynamic phase correlation loss was optimized for the heart rate variability estimation task based on face videos.

[0078] The dynamic phase correlation loss jointly optimizes the dynamic time warping (DTW) loss and the negative Pearson correlation loss with phase constraints. Its mathematical form is defined as follows:

[0079] L = αL DTW + γL Pearson

[0080] where α and γ are weight hyperparameters that respectively control the relative importance of the dynamic time warping loss and the negative Pearson loss with phase constraints during the optimization process. In this embodiment, α = 0.3 and γ = 1.

[0081] Specifically, the dynamic time warping (DTW) loss:

[0082] Dynamic time warping is a method for measuring the similarity between two time series. It allows the time axis to be non-linearly stretched or compressed to find the best time alignment. Due to delays or motion artifacts during the rPPG extraction process, the predicted signal may have a slight time shift at the peak position, although the overall trend is the same. At this time, traditional MSE or Pearson coefficients may not be able to effectively capture this local time difference, while DTW can better evaluate the waveform correspondence through dynamic alignment. The calculation method is as follows:

[0083] Let the predicted rPPG signal be p i = (p1, p2,..., p n ), and the true PPG signal be t j = (t1, t2,..., t n ), n is the number of samples, and DTW calculates the optimal alignment path by constructing a cumulative cost matrix D:

[0084] d(p i , t j ) = (p i - t j ) 2

[0085]

[0086] where d(p i , t j ) is the Euclidean distance between signal points.

[0087] Finally, the dynamic time warping loss is defined as the cumulative cost along the optimal path P:

[0088]

[0089] Among them, N is the total number of points on the alignment path, and the index variable (from 1 to N) for path traversal; i and j are the position indices of the predicted signal and the true signal corresponding to the k-th point on the path respectively; P is the alignment path calculated by fastdtw, and this method uses a constraint radius r = 10 for approximate calculation, reducing the computational complexity to O(N), which is suitable for long-time series modeling.

[0090] Specifically, the negative Pearson loss with phase constraint:

[0091] In the matching process of the predicted rPPG signal and the true PPG signal, although dynamic time warping can align the time scales of the signals, due to the dynamic changes of physiological signals, sampling errors, and individual differences, there may still be a phase shift between the predicted rPPG signal and the true PPG signal. The phase shift will cause a time misalignment of the heartbeat characteristics, making the signals that should be highly correlated show a low similarity, thus affecting the learning effect of the model. Therefore, in this embodiment, a phase penalty term based on cross-correlation analysis is introduced to reduce the phase error that may still exist after the rPPG signal is time-aligned.

[0092] Cross-correlation (CC) is a mathematical tool used to measure the similarity between two signals under different time lags. For the predicted signal p j and the target signal t j , its cross-correlation function is defined as follows:

[0093]

[0094] Among them, τ represents different phase shifts, and C(τ) calculates the matching degree between it and τ when the predicted signal p is shifted forward or backward by τ time steps.

[0095] In actual calculations, the cross-correlation can be efficiently calculated using the fast Fourier transform:

[0096] C = F -1 (F(p) · F(t) * )

[0097] Among them, F() represents the Fourier transform, F() * represents the complex conjugate transform, and F -1 () is the inverse Fourier transform.

[0098] The optimal phase shift Δφ of the predicted rPPG signal relative to the true PPG signal is estimated by calculating the position argmaxC of the cross-correlation maximum:

[0099]

[0100] where m is the signal length, and Δφ is normalized to a phase shift between [0, 1].

[0101] To reduce the influence of the phase shift during the optimization process, this embodiment proposes a phase penalty term λ based on the Gaussian decay function:

[0102] λ = exp(-βΔφ 2 )

[0103] where β is a hyperparameter, and β = 0.3 is taken to control the intensity of the phase penalty. When the phase shift between the predicted rPPG signal and the true PPG signal is larger, λ decays rapidly, making the model tend to reduce the phase error during training, thereby enhancing the temporal alignment between the predicted signal and the target signal.

[0104] Finally, the negative Pearson correlation loss is corrected to:

[0105] L Pearson = 1 - λρ(p, t)

[0106] where ρ() is the Pearson correlation coefficient between the predicted rPPG signal value and the true PPG signal value.

[0107] This improvement enables the model to not only maximize the correlation between the rPPG signal and the true PPG signal, but also effectively suppress the interference of the phase shift on the optimization objective, improving the accuracy and stability of signal matching.

[0108] Step Four: Input the test set into the optimized rPPG signal extraction model to obtain the rPPG signal.

[0109] Step Five: Process the rPPG signal to obtain the estimated values of the heart rate variability related indicators.

[0110] During the process of calculating the heart rate variability based on the rPPG signal, optimization operations such as filtering, interpolation, and dynamic threshold peak localization are taken. The specific operations are as follows:

[0111] First, detrending filtering and FIR band-pass filtering (0.5 - 3.5 Hz) are used to eliminate baseline drift and high-frequency noise. Subsequently, the signal is upsampled to 256 Hz through cubic spline interpolation to improve the time resolution. Finally, dynamic threshold peak detection is combined to locate the peaks and calculate the time-domain and frequency-domain indicators. This process ensures the robustness of the indicator calculation in a complex noise environment through eliminating abnormal intervals and power spectrum normalization processing.

[0112] Step 6: Compare the estimated value with the accurate value of the heart rate variability related index calculated from the ground truth PPG signal, calculate the error, and verify the effectiveness of the model.

[0113] In the estimation of heart rate variability, it is necessary to evaluate the estimation accuracy of time domain and frequency domain indexes.

[0114] The time domain indexes evaluate the overall variability of heart rate variability by statistically analyzing the fluctuation characteristics of the RR interval (adjacent heartbeat intervals) sequence:

[0115] Average of NN intervals (AVNN):

[0116]

[0117] where L is the number of heartbeat intervals, and NNI i is the i-th heartbeat interval time.

[0118] AVNN reflects the change of average heart rate. A small AVNN error also means that the rPPG predicted signal can not only accurately reflect the average heart rate, but also better depict its fluctuation trend.

[0119] Root Mean Square of Successive Differences (RMSSD):

[0120]

[0121] RMSSD is the square root mean of the change of adjacent NNI i intervals, mainly reflecting the high-frequency (parasympathetic nerve) component, and is usually used to evaluate short-term heart rate variability. The smaller the RMSSD error, the stronger the ability of the rPPG signal to capture short-term heart rate changes and parasympathetic nerve activities.

[0122] Standard Deviation of NN intervals (SDNN):

[0123]

[0124] where is the average value of the heartbeat intervals.

[0125] SDNN reflects the overall variability of the entire RR interval sequence and is usually used to evaluate the long-term fluctuation of heart rate changes. The smaller the SDNN error is, the more accurate the rPPG's estimation of the long-term fluctuation of heart rate variability.

[0126] Frequency domain indices quantify heart rate variability through spectral analysis of the RR interval. There are three commonly used frequency domain evaluation indices, namely, the curve area of low frequency (LF), the curve area of high frequency (HF), and their ratio LF / HF. The frequency range of low frequency is 0.04 - 0.15 Hz, and the frequency range of high frequency is 0.15 - 0.4 Hz. In actual measurement, it is usually necessary to normalize the values of LF and HF to obtain the normalized low frequency component LFnu and the normalized high frequency component HFnu, and their definitions are as follows:

[0127]

[0128]

[0129] For the error of the above evaluation criteria, in this embodiment, the mean absolute error is used as the time domain index, and the root mean square error and the Pearson correlation coefficient are used as the frequency domain indices.

[0130] Embodiment 2

[0131] This embodiment discloses a heart rate variability estimation system based on remote face video, including:

[0132] A data acquisition module, configured to: acquire remote face video data, preprocess the remote face video data; synchronously acquire real PPG signals in a contact manner;

[0133] A model construction module, configured to: construct an rPPG signal extraction model;

[0134] A signal recovery module, configured to: input the preprocessed remote face video data into the rPPG signal extraction model to obtain an rPPG signal;

[0135] A signal processing module, configured to: perform signal processing on the rPPG signal to obtain an estimated value of the heart rate variability related index;

[0136] Wherein, the rPPG signal extraction model includes a 3D convolutional layer, a spatio-temporal feature extraction module, a 3D transposed convolutional layer, and a max pooling layer connected in sequence; the spatio-temporal feature extraction module is a cascaded structure of a spatio-temporal separation convolutional layer and a temporal variability convolutional layer.

[0137] Embodiment 3

[0138] The purpose of this embodiment is to provide a computer-readable storage medium. A computer-readable storage medium stores a computer program thereon, and when the program is executed by a processor, it implements the steps in a heart rate variability estimation method based on remote face video as described in Embodiment 1 of the present disclosure.

[0139] Embodiment 4

[0140] The purpose of this embodiment is to provide an electronic device. An electronic device includes a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, the steps in a method for estimating heart rate variability based on remote face video as described in Embodiment 1 of the present disclosure are implemented.

[0141] The steps involved in the devices in the above Embodiments 2, 3, and 4 correspond to those in Method Embodiment 1. For specific implementation manners, reference may be made to the relevant description part of Embodiment 1. The term "computer-readable storage medium" should be understood to include a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0142] Those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0143] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, they do not limit the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solutions of the present invention, various modifications or deformations that can be made without creative efforts by those skilled in the art are still within the protection scope of the present invention.

Claims

1. A method for estimating heart rate variability based on remote face video, characterized in that, Including: Collecting remote face video data and preprocessing the remote face video data; Synchronously collecting real PPG signals in a contact manner; Constructing an rPPG signal extraction model; Inputting the preprocessed remote face video data into the rPPG signal extraction model to obtain an rPPG signal; Performing signal processing on the rPPG signal to obtain an estimated value of a heart rate variability related index; Wherein, the rPPG signal extraction model includes a 3D convolutional layer, a spatio-temporal feature extraction module, a 3D transposed convolutional layer, and a max pooling layer connected in sequence; the spatio-temporal feature extraction module is a cascaded structure of a spatio-temporal separation convolutional layer and a temporal variability convolutional layer.

2. The method for estimating heart rate variability based on remote face video according to claim 1, wherein The preprocessing includes locating the face area, cropping, selecting the face area resolution, determining the sample length, and slicing to increase the number of samples.

3. The method for estimating heart rate variability based on remote facial video according to claim 1, wherein The specific structure of the rPPG signal extraction model is as follows: the remote face video data is input into the 3D convolutional layer to extract spatio-temporal correlation, and the output local spatio-temporal features are sent to the spatio-temporal feature extraction module; the spatio-temporal feature extraction module decouples the time and space features, dynamically adjusts the spatio-temporal dimensions, and then outputs deep spatio-temporal fusion features to the 3D transposed convolutional layer; the 3D transposed convolutional layer expands the time dimension through deconvolution and outputs high-resolution spatio-temporal features to the max pooling layer; the max pooling layer converts the high-dimensional features into a one-dimensional time feature sequence through spatio-temporal feature dimension compression.

4. The method for estimating heart rate variability based on remote face video according to claim 1, wherein The spatio-temporal separation convolutional layer includes: First, using a 1×3×3 depthwise separable spatial convolution to perform local feature extraction on the output of the 3D convolutional layer in the spatial dimension while keeping the time dimension unchanged; then using a 3×1×1 depthwise separable temporal convolution to perform sequence modeling in the time dimension while keeping the spatial dimension unchanged; The spatio-temporal separation convolutional layer decouples the spatial and time features, outputs a feature tensor [C, T, H, W], and inputs it into the temporal variability convolutional layer; where C is the number of channels, T is the time dimension, H is the height, and W is the width.

5. The method for estimating heart rate variability based on remote face video according to claim 4, characterized in that, The temporal variability convolutional layer includes: 1) Initializing the feature tensor [C, T, H, W] output by the spatio-temporal separation convolutional layer; 2) Inputting the feature tensor [C, T, H, W] into an offset convolutional layer to predict the offset amount at each time step, and outputting an offset tensor [kernel_size, T(δ1, δ2, δ3, ……, δT), H, W]; where kernel_size is the convolutional kernel size; δ1, δ2, δ3, ……, δT are the offset amounts; 3) Introducing an index mapping mechanism to dynamically sample temporal features; specifically: generating an index matrix [kernel_size, T]; for each time step of the feature tensor, indexing is performed based on the central time point t, and the index is dynamically adjusted according to the predicted offset amount δT; dynamically sampling traverses each convolutional kernel position, and directly extracts the dynamically adjusted time step features in the time dimension; all the features sampled through indexing are concatenated in the channel dimension, and finally an extended feature tensor [C×kernel_size, T, H, W] is formed. 4) After dynamically sampling the temporal features, the extended feature tensor [C×kernel_size,T,H,W] is input into the fusion convolutional layer for feature fusion, and the deep spatio-temporal fusion features are output; the number of input channels of the fusion convolutional layer is equal to the original number of channels C multiplied by the kernel size.

6. The method for estimating heart rate variability based on remote face video according to claim 1, wherein, Optimize the rPPG signal extraction model through the dynamic phase correlation loss; specifically: the dynamic phase correlation loss combines the dynamic time warping loss and the negative Pearson correlation loss with phase constraints; The dynamic time warping loss L DTW By constructing an accumulated cost matrix, calculate the optimal alignment path between the predicted rPPG signal and the true PPG signal; The negative Pearson correlation loss L Pearson Introduce a phase penalty term based on cross-correlation analysis to reduce the phase error between the predicted rPPG signal and the true PPG signal; The dynamic phase correlation loss L is defined as follows: L = αL DTW + γL Pearson where a and γ are weight hyperparameters.

7. The method for estimating heart rate variability based on remote face video according to claim 1, wherein The rPPG signal is processed to obtain an estimated value of the heart rate variability related index, specifically: Detrend filtering and FIR band-pass filtering are used to eliminate baseline drift and high-frequency noise; the rPPG signal is upsampled to 256Hz by cubic spline interpolation to improve the time resolution; the dynamic threshold peak detection is combined to locate the peak and calculate the time-domain and frequency-domain indexes.

8. A heart rate variability estimation system based on remote face video, characterized in that, Including: A data acquisition module, configured to: acquire remote face video data and preprocess the remote face video data; Synchronously collect real PPG signals in a contact way; A model construction module, configured to: construct an rPPG signal extraction model; A signal recovery module, configured to: input the preprocessed remote face video data into the rPPG signal extraction model to obtain an rPPG signal; A signal processing module, configured to: process the rPPG signal to obtain an estimated value of the heart rate variability related index; wherein, the rPPG signal extraction model includes a 3D convolutional layer, a spatio-temporal feature extraction module, a 3D transposed convolutional layer, and a max pooling layer connected in sequence; the spatio-temporal feature extraction module is a cascaded structure of a spatio-temporal separation convolutional layer and a temporal variability convolutional layer.

9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in a method for estimating heart rate variability based on remote face video according to any one of claims 1-7.

10. An electronic device, comprising a memory, a processor, and a program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in a method for estimating heart rate variability based on remote face video according to any one of claims 1-7.