Long time sequence rPPG signal heart rate variability estimation method and system
Through the long-term rPPG signal heart rate variability estimation method, the AttenHRVNet model and the spatiotemporal attention mechanism are used to solve the problem of signal quality in non-contact HRV measurement and insufficient data set acquisition time in contactless HRV measurement, and high accuracy and stability HRV monitoring is achieved.
Patent Information
- Application Number
- CN202510068588.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has problems of signal quality in non-contact heart rate variability (HRV) measurements, and the traditional data sets have short collection time, which cannot fully reflect the time and frequency domain characteristics of HRV.
The heart rate variability estimation method of long-term rPPG signal was used to collect video information through the camera, and the AttenHRVNet model combined with the spatiotemporal attention mechanism was used to extract rich spatiotemporal features, and HRV indexes were optimized and extracted through peak detection, cosine function fitting and pseudo-peak removal strategies.
It significantly improves the accuracy and stability of HRV monitoring, enhances the HRV monitoring capabilities of the model in complex environments, and can more comprehensively reflect the time and frequency domain characteristics of HRV.
Smart Images

Figure CN119969986A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of heart rate detection, and in particular relates to a method and system for estimating heart rate variability of long-time series rPPG signals. Background Art
[0002] Heart rate variability (HRV) is a key physiological indicator for measuring the regulatory ability of the autonomic nervous system. It is evaluated by analyzing the temporal changes in the intervals between consecutive heart beats (RR intervals). Generally speaking, a higher HRV means good neural regulation, while a lower HRV may indicate heart disease, psychological stress or other health problems. Therefore, it has important clinical significance in the fields of cardiovascular disease prediction, psychological stress assessment and exercise rehabilitation [3].
[0003] At present, common methods for measuring HRV include ECG (electrocardiogram) and BVP (blood volume pulse) signals. Although these methods have high accuracy, they are contact- and invasive-related due to the need for direct contact with the skin, making them difficult to meet the needs of daily health monitoring, especially for widespread use in daily life scenarios. Therefore, non-contact HRV measurement has become an important research direction. However, the accuracy of non-contact methods depends on data quality. These methods usually use optical sensors to capture signals, which are easily disturbed by changes in ambient light and background noise, resulting in unstable signal quality. In addition, motion interference (such as changes in facial expressions and head rotation) may also affect the accuracy of the signal. Many traditional HRV data sets are collected for a short period of time, usually only a few minutes, which limits the data's comprehensive reflection of HRV time and frequency domain characteristics. Therefore, in order to improve the accuracy of HRV analysis, it is usually required that the data collection time be at least 5 minutes to avoid deviations caused by short-term collection.
[0004] With the rapid development of deep learning technology, heart rate variability (HRV) analysis has gradually been automated, especially 3D convolutional neural networks (3D-CNN) have shown significant advantages in capturing spatiotemporal features. 3D-CNN can effectively process the spatiotemporal dynamics in HRV data, especially in complex environments, and can adapt to interference such as lighting changes and facial movements, thereby improving the stability and accuracy of monitoring. Combined with the spatiotemporal attention mechanism, the model can focus on key time and space areas, further improving the accuracy and robustness of feature extraction.
[0005] Although traditional short-term data sets provide limited HRV information, they cannot fully reflect its time and frequency domain characteristics, which may affect the stability of the model, especially in complex environments. Summary of the invention
[0006] The technical problem to be solved by the present invention is to overcome the shortcomings of the above-mentioned prior art and provide a method and system for estimating heart rate variability of long-time series rPPG signals.
[0007] The technical solution adopted to solve the above technical problem is: a method for estimating heart rate variability of long-time series rPPG signals, comprising the following specific steps:
[0008] Step 1: Collect video information of the person to be detected through the camera;
[0009] Step 2: Preprocess the video information to generate an image sequence with the same frame size.
[0010] Step 3: The image frame is input into the AttenHRVNet model and the preliminary rPPG signal is output;
[0011] Step 4: Optimize and extract clean rPPG signals through Swart filtering, normalization and smoothing data enhancement operations;
[0012] Step 5: Use the peak detection algorithm to locate the R wave position, calculate the RR interval, and further extract the HRV index.
[0013] Furthermore, it includes a face detection module, a data preprocessing module, an encoder module, a decoder module and an evaluation module. The face detection module is responsible for reducing the interference of irrelevant background in the video on the analysis, uniformly processing the resolution of the video and adjusting it to 128×128 pixels. The multi-task cascade convolutional neural network (MTCNN) algorithm is adopted. Through its cascade structure and multi-task learning strategy, accurate face detection and alignment are achieved. In the first frame of the video, MTCNN is used to detect the face coordinates and maintain the same coordinate area in subsequent frames to ensure the consistency and accuracy of the detection;
[0014] In order to ensure the accuracy of HRV analysis, the data preprocessing module first normalizes the data to reduce interference, and extracts the peak through peak detection and cosine curve fitting. The data is first transposed and normalized to the range of [0,1]. Then, the data is Gaussian filtered through a smoothing filter to reduce the impact of noise and ensure signal quality. Peak detection determines the peak position by comparing the size relationship between the signal value and the adjacent points. Specifically, if a point f(t) is greater than the value of its adjacent points before and after, the point is considered to be a peak, that is:
[0015] f(t)>f(t-1)and f(t)>f(t+1)
[0016] Multiple constraints were introduced for screening, including height constraint, which filters out peaks with smaller amplitudes by setting the minimum peak height; interval constraint, which specifies the minimum interval between peaks to avoid too dense peaks; width constraint, which defines the width of the peak as the time interval from both sides of the peak to the half-maximum amplitude, and sets the minimum and maximum widths;
[0017] After extracting the effective peaks, calculate the interval between adjacent peaks, called the RR interval, record the maximum RR interval, and fit the signal based on this value using the cosine function. The cosine function can effectively model the periodic fluctuation of heart rate, and its mathematical expression is:
[0018]
[0019] Where A is the amplitude, which indicates the amplitude of the signal peak, T max is the period of the cosine function, the corresponding maximum peak interval is the phase offset, and is the DC component. The above cosine function is fitted by the least squares method to optimize these parameters so that the fitting curve is closer to the original signal. The objective function of the least squares method is:
[0020]
[0021] where t i is the position of the peak. By solving this optimization problem, the optimal parameters can be obtained, and finally the reconstructed signal, that is, the smoothed signal, is obtained. In order to further improve the accuracy of heart rate variability analysis, a pseudo peak removal strategy is adopted to remove those peaks whose RR intervals deviate significantly from the normal range, and 30% of the outliers are used as the threshold.
[0022] Through the above technical solutions, we can successfully extract the periodic characteristics of the heart rate signal through peak detection, maximum peak interval extraction and cosine function fitting, and improve the accuracy of the data through pseudo peak removal. This process provides stable and reliable label data for subsequent HRV analysis, which helps the network model to better fit the curve and improve the accuracy and reliability of the analysis results.
[0023] Furthermore, the encoder module consists of a series of 3D convolutional blocks, which aims to extract rich spatiotemporal features from the video input. Each convolutional block includes a convolutional layer, a batch normalization layer, and an ELU activation function to effectively capture the local and global features of the input data. In addition, the encoder module also introduces two pooling operations: spatial pooling (MaxpoolSpa), which gradually reduces the spatial resolution; spatiotemporal pooling (MaxpoolSpaTem), which further compresses the spatiotemporal feature dimensions and retains key multi-scale information;
[0024] The encoder module combines spatial attention and temporal attention mechanisms, where:
[0025] Spatial Attention:
[0026] This mechanism focuses on the key areas of the input features through weighted allocation, helping the model to automatically identify and strengthen the representation of salient areas;
[0027] Time attention:
[0028] This mechanism focuses on the timing characteristics of the signal and can capture the periodic characteristics of the signal changing over time;
[0029] Input spatiotemporal characteristics First, three convolutional layers are used to extract features related to space and time. The global feature A is obtained through the convolutional layer ConvA. The spatial attention map and the temporal attention map are calculated through the convolutional layers ConvB and ConvV respectively. The spatial attention map B is normalized by the softmax operation to obtain the weight S. att :
[0030]
[0031] These weights reflect the importance of different spatial locations;
[0032] The temporal attention map V normalizes the time dimension and highlights the importance of different time steps to obtain the weight T att :
[0033]
[0034] Then, by calculating the temporary global feature A and the spatial attention map S att Weighted, get the spatial weighted feature F S :
[0035] F S =S att ·A
[0036] Then the global descriptor F S With the temporal attention map T att Combined, the final feature representation F is obtained by weighting T :
[0037] F T =T att ·F S
[0038] Finally, the weighted features are subjected to the convolutional layer. T Reconstruct, restore to the same shape as the input
[0039] X'=Conv(F T )
[0040] By independently weighting space and time, the most important spatiotemporal information in the input data can be effectively captured.
[0041] The above technical solution helps to extract meaningful physiological features in different signal areas in HRV analysis, especially in facial areas or other cases where input signals may be interfered by noise. It can effectively locate and enhance important features. Temporal attention helps the model better understand and model the periodic fluctuations in heart rate variability, thereby enhancing the ability to model the laws of heart rate changes.
[0042] Furthermore, the decoder module is responsible for recovering high-resolution spatiotemporal signals from the low-dimensional features generated by the encoder and finally predicting the rPPG signal. Each upsampling layer consists of a deconvolution layer (ConvTranspose3d), a batch normalization layer (Batch Normalization) and an ELU activation function.
[0043] First, the decoder upsamples the features through the deconvolution layer to gradually restore its spatial resolution. After each upsampling, the network will concatenate the features of the current decoder layer with the corresponding high-resolution features in the encoder. The output after upsampling and feature fusion will be adjusted in shape through adaptive pooling to adapt to the shape of the rPPG signal generated by the last convolution block;
[0044] Next, the data is input into the LSTM module. LSTM captures the long-term dependencies in the input sequence through its long short-term memory units, thereby modeling the temporal information in the video sequence.
[0045] Through the above technical solutions, the encoder-decoder architecture of AttenHRVNet effectively combines spatiotemporal feature extraction, attention mechanism, upsampling and time series modeling technology. The encoder part extracts spatiotemporal features through 3D convolution and strengthens the expression of key features through the dual attention mechanism; the decoder part gradually restores the spatial resolution through upsampling and feature concatenation, and uses the LSTM module to model the time series dependency, and finally outputs accurate rPPG signals. This encoder-decoder architecture performs well in processing complex spatiotemporal video data tasks, especially in physiological signal extraction tasks.
[0046] Furthermore, the negative Pearson correlation coefficient (L ρ ) is used as the main loss function. In order to further reduce the amplitude error between the predicted rPPG signal and the true PPG signal, (L ρ) is combined with the mean absolute error (MAE, L1) to form a combined loss function:
[0047]
[0048] Loss = α*L ρ +(1-α)*L1
[0049] Among them, N represents the length of the input, x and y represent the predicted rPPG signal and the true PPG signal, respectively, and the optimal value of α is 0.7.
[0050] Furthermore, after obtaining the complete rPPG signal, the evaluation module first extracts the RR interval through peak detection. Based on the RR interval, the time domain index of HRV can be further calculated. In terms of frequency domain index, the main focus is on low-frequency (LF) and high-frequency (HF) components, which reflect the regulatory effects of the sympathetic and parasympathetic nerves on the heart, respectively;
[0051] In addition, the LF / HF ratio was used to represent the balance between the sympathetic and parasympathetic nerves, and the LF and HF were standardized and converted into the sum. The frequency domain indices were evaluated by the root mean square error (RMSE) and Pearson correlation coefficient (r) to comprehensively analyze the individual's cardiac health status and autonomic nervous system function.
[0052] Furthermore, Mean RR (mean RR interval) is the average value of consecutive heartbeat intervals, reflecting the overall level of heart rate, with a normal range of 800-1200ms. The calculation formula is:
[0053]
[0054] Where is the i-th RR interval, N is the total number of RR intervals;
[0055] SDNN (overall standard deviation) is the standard deviation of the RR interval, reflecting the overall variability of heart rate. The normal range is 102-180ms. The calculation formula is:
[0056]
[0057] RMSSD (mean square of difference) is the root mean square of the difference between adjacent RR intervals, reflecting the fast-changing components in HRV. The normal range is 15-39ms. The calculation formula is:
[0058]
[0059] SDSD (Successive Difference of Successive Differences) refers to the standard deviation of the difference between two adjacent RR intervals, which can reflect the effect of parasympathetic nerve activity on heart rate. The normal range is 20-50ms. The calculation formula is:
[0060]
[0061] PNN50 (Percentage of Successive RR Intervals Differing by More Than50ms) indicates the ratio of adjacent RR intervals differing by more than 50 milliseconds within a period of time, which reflects the degree of fluctuation between heartbeats, and the normal range is 10%-30%;
[0062] The following evaluation indicators are used to evaluate the system:
[0063] MAE (mean absolute error) is a commonly used indicator to measure the average absolute difference between the predicted value and the actual value. It can reflect the average absolute difference between the predicted value and the actual value. MAE is defined as:
[0064]
[0065] Where N is the number of data points, y i is the actual value, is the predicted value.
[0066] RMSE (root mean square error) is another commonly used indicator to measure the difference between the predicted value and the actual value. It can reflect the standard deviation of the prediction error. A larger error will have a greater impact on RMSE. RMSE is defined as:
[0067]
[0068] The Pearson Correlation Coefficient (r) is an indicator to measure the linear correlation between the predicted value and the actual value. It reflects the strength and direction of the relationship between the two variables. The value range is between 1 and -1. The closer to 1, the stronger the positive correlation, and the closer to -1, the stronger the negative correlation. The Pearson correlation coefficient is defined as:
[0069]
[0070] Evaluation is carried out through the above indicators, so as to make targeted adjustments.
[0071] The beneficial effects of the present invention are as follows: the present invention enhances the HRV monitoring capability of the model in complex environments by introducing a 10-minute HRV dataset, and significantly improves the accuracy and stability of HRV monitoring by combining 3D-CNN and long-term datasets for spatiotemporal feature extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 It is the overall flow chart of the present invention;
[0073] Figure 2 It is the overall algorithm framework and network architecture AttenHRVNet diagram of the present invention;
[0074] Figure 3 It is a structural diagram of the encoder module of the present invention;
[0075] Figure 4 It is a comparison diagram of the real BVP signal and the rPPG signal on the UBFC_rPPG data set of the present invention;
[0076] Figure 5 This is an experimental diagram of the present invention on the UBFC_rPPG dataset;
[0077] Figure 6 This is an experimental diagram of the present invention on the COHFACE dataset;
[0078] Figure 7 This is an experimental diagram of the present invention on the ZJXU_HRV dataset. DETAILED DESCRIPTION
[0079] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0080] like Figure 1-Figure 3 As shown, a method and system for estimating heart rate variability of a long-time series rPPG signal in this embodiment includes the following specific steps:
[0081] Step 1: Collect video information of the person to be detected through the camera;
[0082] Step 2: Preprocess the video information to generate an image sequence with the same frame size.
[0083] Step 3: The image frame is input into the AttenHRVNet model and the preliminary rPPG signal is output;
[0084] Step 4: Optimize and extract clean rPPG signals through Swart filtering, normalization and smoothing data enhancement operations;
[0085] Step 5: Use the peak detection algorithm to locate the R wave position, calculate the RR interval, and further extract the HRV index.
[0086] It includes face detection module, data preprocessing module, encoder module, decoder module and evaluation module. The face detection module is responsible for reducing the interference of irrelevant background in the video on the analysis, uniformly processing the resolution of the video and adjusting it to 128×128 pixels. It adopts the multi-task cascade convolutional neural network (MTCNN) algorithm. Through its cascade structure and multi-task learning strategy, it realizes accurate face detection and alignment. In the first frame of the video, MTCNN is used to detect the face coordinates and maintain the same coordinate area in subsequent frames to ensure the consistency and accuracy of detection.
[0087] In order to ensure the accuracy of HRV analysis, the data preprocessing module first normalizes the data to reduce interference, and extracts the peak through peak detection and cosine curve fitting. The data is first transposed and normalized to the range of [0,1]. Then, the data is Gaussian filtered through a smoothing filter to reduce the impact of noise and ensure signal quality. Peak detection determines the peak position by comparing the size relationship between the signal value and the adjacent points. Specifically, if a point f(t) is greater than the value of its adjacent points before and after, the point is considered to be a peak, that is:
[0088] f(t)>f(t-1)and f(t)>f(t+1)
[0089] Multiple constraints were introduced for screening, including height constraint, which filters out peaks with smaller amplitudes by setting the minimum peak height; interval constraint, which specifies the minimum interval between peaks to avoid too dense peaks; width constraint, which defines the width of the peak as the time interval from both sides of the peak to the half-maximum amplitude, and sets the minimum and maximum widths;
[0090] After extracting the effective peaks, calculate the interval between adjacent peaks, called the RR interval, record the maximum RR interval, and fit the signal based on this value using the cosine function. The cosine function can effectively model the periodic fluctuation of heart rate, and its mathematical expression is:
[0091]
[0092] Where A is the amplitude, which indicates the amplitude of the signal peak, T max is the period of the cosine function, the corresponding maximum peak interval is the phase offset, and is the DC component. The above cosine function is fitted by the least squares method to optimize these parameters so that the fitting curve is closer to the original signal. The objective function of the least squares method is:
[0093]
[0094] where t i is the position of the peak. By solving this optimization problem, the optimal parameters can be obtained, and finally the reconstructed signal, that is, the smoothed signal, is obtained. In order to further improve the accuracy of heart rate variability analysis, a pseudo peak removal strategy is adopted to remove those peaks whose RR intervals deviate significantly from the normal range, and 30% of the outliers are used as the threshold.
[0095] The encoder module consists of a series of 3D convolutional blocks, which are designed to extract rich spatiotemporal features from the video input. Each convolutional block includes a convolutional layer, a batch normalization layer, and an ELU activation function to effectively capture the local and global features of the input data. In addition, the encoder module also introduces two pooling operations: spatial pooling (MaxpoolSpa), which gradually reduces the spatial resolution; spatiotemporal pooling (MaxpoolSpaTem), which further compresses the spatiotemporal feature dimensions and retains key multi-scale information.
[0096] The encoder module combines spatial attention and temporal attention mechanisms, where:
[0097] Spatial Attention:
[0098] This mechanism focuses on the key areas of the input features through weighted allocation, helping the model to automatically identify and strengthen the representation of salient areas;
[0099] Time attention:
[0100] This mechanism focuses on the timing characteristics of the signal and can capture the periodic characteristics of the signal changing over time;
[0101] Input spatiotemporal characteristics First, three convolutional layers are used to extract features related to space and time. The global feature A is obtained through the convolutional layer ConvA. The spatial attention map and the temporal attention map are calculated through the convolutional layers ConvB and ConvV respectively. The spatial attention map B is normalized by the softmax operation to obtain the weight S. att :
[0102]
[0103] These weights reflect the importance of different spatial locations;
[0104] The temporal attention map V normalizes the time dimension and highlights the importance of different time steps to obtain the weight T att :
[0105]
[0106] Then, by calculating the temporary global feature A and the spatial attention map S att Weighted, get the spatial weighted feature F S :
[0107] F S =S att ·A
[0108] Then the global descriptor F S With the temporal attention map T att Combined, the final feature representation F is obtained by weighting T :
[0109] F T =T att ·F S
[0110] Finally, the weighted features are subjected to the convolutional layer. T Reconstruct, restore to the same shape as the input
[0111] X'=Conv(F T )
[0112] By independently weighting space and time, the most important spatiotemporal information in the input data can be effectively captured.
[0113] The decoder module is responsible for recovering high-resolution spatiotemporal signals from the low-dimensional features generated by the encoder and finally predicting the rPPG signal. Each upsampling layer consists of a deconvolution layer (ConvTranspose3d), a batch normalization layer (BatchNormalization), and an ELU activation function.
[0114] First, the decoder upsamples the features through the deconvolution layer to gradually restore its spatial resolution. After each upsampling, the network will concatenate the features of the current decoder layer with the corresponding high-resolution features in the encoder. The output after upsampling and feature fusion will be adjusted in shape through adaptive pooling to adapt to the shape of the rPPG signal generated by the last convolution block;
[0115] Next, the data is input into the LSTM module. LSTM captures the long-term dependencies in the input sequence through its long short-term memory units, thereby modeling the temporal information in the video sequence.
[0116] The negative Pearson correlation coefficient (L ρ ) is used as the main loss function. In order to further reduce the amplitude error between the predicted rPPG signal and the true PPG signal, (L ρ ) is combined with the mean absolute error (MAE, L1) to form a combined loss function:
[0117]
[0118] Loss = α*L ρ +(1-α)*L1
[0119] Among them, N represents the length of the input, x and y represent the predicted rPPG signal and the true PPG signal, respectively, and the optimal value of α is 0.7.
[0120] After obtaining the complete rPPG signal, the evaluation module first extracts the RR interval through peak detection. Based on the RR interval, the time domain index of HRV can be further calculated. In terms of frequency domain index, the main focus is on low frequency (LF) and high frequency (HF) components, which reflect the regulatory effects of the sympathetic and parasympathetic nerves on the heart respectively.
[0121] In addition, the LF / HF ratio was used to represent the balance between the sympathetic and parasympathetic nerves, and the LF and HF were standardized and converted into the sum. The frequency domain indices were evaluated by the root mean square error (RMSE) and Pearson correlation coefficient (r) to comprehensively analyze the individual's cardiac health status and autonomic nervous system function.
[0122] Mean RR (mean RR interval) is the average value of consecutive heartbeat intervals, reflecting the overall level of heart rate. The normal range is 800-1200ms. The calculation formula is:
[0123]
[0124] Where is the i-th RR interval, N is the total number of RR intervals;
[0125] SDNN (overall standard deviation) is the standard deviation of the RR interval, reflecting the overall variability of heart rate. The normal range is 102-180ms. The calculation formula is:
[0126]
[0127] RMSSD (mean square of difference) is the root mean square of the difference between adjacent RR intervals, reflecting the fast-changing components in HRV. The normal range is 15-39ms. The calculation formula is:
[0128]
[0129] SDSD (Successive Difference of Successive Differences) refers to the standard deviation of the difference between two adjacent RR intervals, which can reflect the effect of parasympathetic nerve activity on heart rate. The normal range is 20-50ms. The calculation formula is:
[0130]
[0131] PNN50 (Percentage of Successive RR Intervals Differing by More Than50ms) indicates the ratio of adjacent RR intervals differing by more than 50 milliseconds within a period of time, which reflects the degree of fluctuation between heartbeats, and the normal range is 10%-30%;
[0132] The following evaluation indicators are used to evaluate the system:
[0133] MAE (mean absolute error) is a commonly used indicator to measure the average absolute difference between the predicted value and the actual value. It can reflect the average absolute difference between the predicted value and the actual value. MAE is defined as:
[0134]
[0135] Where N is the number of data points, y i is the actual value, is the predicted value.
[0136] RMSE (root mean square error) is another commonly used indicator to measure the difference between the predicted value and the actual value. It can reflect the standard deviation of the prediction error. A larger error will have a greater impact on RMSE. RMSE is defined as:
[0137]
[0138] The Pearson Correlation Coefficient (r) is an indicator to measure the linear correlation between the predicted value and the actual value. It reflects the strength and direction of the relationship between the two variables. The value range is between 1 and -1. The closer to 1, the stronger the positive correlation, and the closer to -1, the stronger the negative correlation. The Pearson correlation coefficient is defined as:
[0139]
[0140] Evaluation is carried out through the above indicators, so as to make targeted adjustments.
[0141] like Figure 4 and Figure 5 As shown, experiments on the UBFC_rPPG dataset:
[0142] As can be seen from Table 1, the performance of traditional methods is usually inferior to that of deep learning-based methods. This is because traditional methods rely on prior assumptions and are easily disturbed by illumination changes and motion noise. Among traditional methods, POS performs worse than CHROM in time domain indicators, while CHROM performs better than POS in frequency domain indicators. This is mainly because CHROM uses the ratio of color components to suppress the influence of some illumination changes, while POS is less robust under complex conditions. In contrast, deep learning-based methods significantly improve performance, among which the effect of unlabeled methods is significantly lower than that of labeled methods. The AttenHRVNet proposed in the present invention achieves the optimal MAE of 5.22ms and 10.1ms in the time domain indicators SDNN and RMSSD, respectively, and its Pearson correlation coefficient is only slightly lower than that of the Shuffle-rPPGNet framework in frequency domain indicators.
[0143] Table 1 HRV performance evaluation of different methods on UBFC_rPPG dataset
[0144]
[0145]
[0146] In order to further analyze the relationship between the predicted HRV index and the true HRV index, the present invention selected 40 samples from the test data set for comparison.
[0147] Figure 5 The red dashed line in the figure represents the mean, and the two blue dashed lines represent the 95% confidence interval. The solid line in the regression graph represents the standard line where the predicted HR is completely consistent with the true HR. It can be seen that the standard deviation of the proposed method is smaller, and the predicted HRV is more consistent with the true HRV. This shows that the AttenHRVNet method has better prediction accuracy and robustness.
[0148] like Figure 6 As shown, experiments on the COHFACE dataset:
[0149] The experimental results on the COHFACE dataset are shown in Table 2. The experimental results are provided by the research of DAS et al. As can be seen from the table, our proposed method has achieved the best results in key indicators such as PNN50 and SDNN, showing the significant advantages of the network in these indicators. Although there is still a certain gap in some indicators such as MeanNN and RMSSD, this is mainly due to the errors caused by the high compression degree of the COHFACE dataset and the limited data quality.
[0150] However, despite these errors, Figure 6The Bland-Altman plot and regression plot in the figure still prove the feasibility and reliability of the network proposed in this invention in HRV estimation. Specifically, the Bland-Altman plot shows the deviation distribution between the predicted results and the true values, verifying the consistency of the network on different data sets; the regression plot shows that a strong linear correlation is maintained between the predicted values and the true values, further confirming the robustness and practicality of the network for HRV estimation.
[0151] Table 2 HRV index performance evaluation of different methods on the COHFACE dataset
[0152]
[0153]
[0154] First, the UBFC-rPPG dataset was used to verify the effectiveness of the network on this dataset, demonstrating its ability to extract rPPG signals and the accuracy of heart rate variability analysis under different experimental conditions. Then, the COHFACE dataset was further tested. Although the compression effect of this dataset added challenges, the network still showed strong robustness and was able to effectively extract rPPG signals from compressed videos, supporting its potential in real-world applications. However, the short time length of existing datasets limits the investigation of long-term monitoring. Therefore, in order to verify the performance of the network on longer data, a 10-minute dataset was introduced, and corresponding experiments were conducted to further verify its stability and performance in long-term data processing.
[0155] like Figure 7 As shown, experiments on the ZJXU_HRV dataset:
[0156] The ZJXU_HRV dataset is divided into three time periods of 1 minute, 5 minutes, and 10 minutes, and the HRV index is calculated in these time periods. Through this segmentation method, we can evaluate the impact of different time windows on the stability and reliability of HRV indicators, thereby providing a more valuable reference for HRV analysis based on long-term data. In particular, the use of longer time windows can more comprehensively reflect the changing characteristics of heart rate variability at different time scales.
[0157] Table 3 Comparison of calculation results of HRV indicators in different time windows
[0158]
[0159] Table 3 shows the mean square error (RMSE), standard deviation (SD), correlation coefficient (r) and other performances of the main HRV indices such as MeanNN, SDNN, RMSSD, SDSD, PNN50, LF and LF / HF in different time windows. The results show that as the time window increases (from 1 minute to 10 minutes), the RMSE and SD of the HRV indices gradually decrease. For example, the RMSE of SDNN decreases from 27.47 in 1 minute to 16.89 in 10 minutes, and the SD of PNN50 decreases from 0.21 to 0.06. This shows that longer time windows can more effectively smooth instantaneous fluctuations and noise, thereby improving the calculation stability of HRV indices. In addition, the correlation coefficient (r) increases significantly with the increase of the time window, from 0.38 in 1 minute to 0.71 in 10 minutes, indicating that the HRV indices under longer time windows have higher reliability and consistency. This trend is particularly significant in time domain indices such as SDNN and RMSSD, and similar regularity is also shown in frequency domain indices (such as LF and LF / HF).
[0160] Figure 7 The distribution of heart rate variability (HRV) indicators predicted by the non-contact vision-based method in 1-minute, 5-minute, and 10-minute time periods is shown. Each sub-figure represents a specific HRV indicator, and the statistical characteristics of the predicted values in different time periods are shown through box plots, including the median, interquartile range, and outliers. It can be observed from the figure that the predicted values in the 10-minute time period show higher stability than those in the 1-minute time period. This is reflected in the smaller interquartile range of the 10-minute data and the fewer outliers, indicating that the HRV indicator predictions for longer time periods are more consistent and stable. In contrast, the 1-minute data has larger fluctuations, which may be affected by noise or unstable factors in a short period of time. This phenomenon is particularly evident in frequency domain indicators. Frequency domain indicators, such as low frequency (LF), high frequency (HF), and LF / HF ratio, are usually greatly affected by signal noise and short-term fluctuations, so the prediction results for shorter time periods fluctuate significantly. As the time period increases, the prediction results of these frequency domain indicators tend to be stable, the interquartile range becomes narrower, and the outliers decrease, showing higher reliability. This trend indicates that data collection over a longer period of time can more effectively reduce noise interference and improve the accuracy and consistency of frequency domain indicators.
[0161] Experimental results show that although short time windows can capture higher-resolution heart rate variability information, long time windows show more significant advantages in reliability and indicator consistency.
[0162] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention.
Claims
1. A method for estimating heart rate variability of long-time rPPG signals, characterized in that: The specific steps include: Step 1: Collect video information of the person to be detected through the camera; Step 2: Preprocess the video information to generate an image sequence with the same frame size. Step 3: The image frame is input into the AttenHRVNet model and the preliminary rPPG signal is output; Step 4: Optimize and extract clean rPPG signals through Swart filtering, normalization and smoothing data enhancement operations; Step 5: Use the peak detection algorithm to locate the R wave position, calculate the RR interval, and further extract the HRV index.
2. A long-time series rPPG signal heart rate variability estimation system according to claim 1, characterized in that: It includes face detection module, data preprocessing module, encoder module, decoder module and evaluation module. The face detection module is responsible for reducing the interference of irrelevant background in the video on the analysis, uniformly processing the resolution of the video and adjusting it to 128×128 pixels. It adopts the multi-task cascade convolutional neural network (MTCNN) algorithm. Through its cascade structure and multi-task learning strategy, it realizes accurate face detection and alignment. In the first frame of the video, MTCNN is used to detect the face coordinates and maintain the same coordinate area in subsequent frames to ensure the consistency and accuracy of detection. In order to ensure the accuracy of HRV analysis, the data preprocessing module first normalizes the data to reduce interference, and extracts the peak through peak detection and cosine curve fitting. The data is first transposed and normalized to the range of [0,1]. Then, the data is Gaussian filtered through a smoothing filter to reduce the impact of noise and ensure signal quality. Peak detection determines the peak position by comparing the size relationship between the signal value and the adjacent points. Specifically, if a point f(t) is greater than the value of its adjacent points before and after, the point is considered to be a peak, that is: f(t)>f(t-1)and f(t)>f(t+1) Multiple constraints were introduced for screening, including height constraint, which filters out peaks with smaller amplitudes by setting the minimum peak height; interval constraint, which specifies the minimum interval between peaks to avoid too dense peaks; width constraint, which defines the width of the peak as the time interval from both sides of the peak to the half-maximum amplitude, and sets the minimum and maximum widths; After extracting the effective peaks, calculate the interval between adjacent peaks, called the RR interval, record the maximum RR interval, and fit the signal based on this value using the cosine function. The cosine function can effectively model the periodic fluctuation of heart rate, and its mathematical expression is: Where A is the amplitude, which indicates the amplitude of the signal peak, T max is the period of the cosine function, the corresponding maximum peak interval is the phase offset, and is the DC component. The above cosine function is fitted by the least squares method to optimize these parameters so that the fitting curve is closer to the original signal. The objective function of the least squares method is: where t i is the position of the peak. By solving this optimization problem, the optimal parameters can be obtained, and finally the reconstructed signal, that is, the smoothed signal, is obtained. In order to further improve the accuracy of heart rate variability analysis, a pseudo peak removal strategy is adopted to remove those peaks whose RR intervals deviate significantly from the normal range, and 30% of the outliers are used as the threshold.
3. A long-time series rPPG signal heart rate variability estimation system according to claim 2, characterized in that: The encoder module consists of a series of 3D convolutional blocks, which are designed to extract rich spatiotemporal features from the video input. Each convolutional block includes a convolutional layer, a batch normalization layer, and an ELU activation function to effectively capture the local and global features of the input data. In addition, the encoder module also introduces two pooling operations: spatial pooling (MaxpoolSpa), which gradually reduces the spatial resolution; spatiotemporal pooling (MaxpoolSpaTem), which further compresses the spatiotemporal feature dimensions and retains key multi-scale information. The encoder module combines spatial attention and temporal attention mechanisms, where: Spatial Attention: This mechanism focuses on the key areas of the input features through weighted allocation, helping the model to automatically identify and strengthen the representation of salient areas; Time attention: This mechanism focuses on the timing characteristics of the signal and can capture the periodic characteristics of the signal changing over time; Input spatiotemporal characteristics First, three convolutional layers are used to extract features related to space and time. The global feature A is obtained through the convolutional layer ConvA. The spatial attention map and the temporal attention map are calculated through the convolutional layers ConvB and ConvV respectively. The spatial attention map B is normalized by the softmax operation to obtain the weight S. att : These weights reflect the importance of different spatial locations; The temporal attention map V normalizes the time dimension and highlights the importance of different time steps to obtain the weight T att : Then, by calculating the temporary global feature A and the spatial attention map S att Weighted, get the spatial weighted feature F S : F S =S att ·A Then the global descriptor F S With the temporal attention map T att Combined, the final feature representation F is obtained by weighting T : F T =T att ·F S Finally, the weighted features are subjected to the convolutional layer. T Reconstruct, restore to the same shape as the input X'=Conv(F T ) By independently weighting space and time, the most important spatiotemporal information in the input data can be effectively captured.
4. A long-time series rPPG signal heart rate variability estimation system according to claim 3, characterized in that: The decoder module is responsible for recovering high-resolution spatiotemporal signals from the low-dimensional features generated by the encoder and finally predicting the rPPG signal. Each upsampling layer consists of a deconvolution layer (ConvTranspose3d), a batch normalization layer (Batch Normalization), and an ELU activation function. First, the decoder upsamples the features through the deconvolution layer to gradually restore its spatial resolution. After each upsampling, the network will concatenate the features of the current decoder layer with the corresponding high-resolution features in the encoder. The output after upsampling and feature fusion will be adjusted in shape through adaptive pooling to adapt to the shape of the rPPG signal generated by the last convolution block; Next, the data is input into the LSTM module. LSTM captures the long-term dependencies in the input sequence through its long short-term memory units, thereby modeling the temporal information in the video sequence.
5. A long-time series rPPG signal heart rate variability estimation system according to claim 4, characterized in that: The negative Pearson correlation coefficient (L ρ ) is used as the main loss function. In order to further reduce the amplitude error between the predicted rPPG signal and the true PPG signal, (L ρ ) is combined with the mean absolute error (MAE, L1) to form a combined loss function: Loss=α*L ρ +(1-a)*L1 Among them, N represents the length of the input, x and y represent the predicted rPPG signal and the true PPG signal, respectively, and the optimal value of α is 0.
7.
6. A long-time series rPPG signal heart rate variability estimation system according to claim 5, characterized in that: After obtaining the complete rPPG signal, the evaluation module first extracts the RR interval through peak detection. Based on the RR interval, the time domain index of HRV can be further calculated. In terms of frequency domain index, the main focus is on low frequency (LF) and high frequency (HF) components, which reflect the regulatory effects of the sympathetic and parasympathetic nerves on the heart respectively. In addition, the LF / HF ratio was used to represent the balance between the sympathetic and parasympathetic nerves, and the LF and HF were standardized and converted into the sum. The frequency domain indices were evaluated by the root mean square error (RMSE) and Pearson correlation coefficient (r) to comprehensively analyze the individual's cardiac health status and autonomic nervous system function.
7. A long-time series rPPG signal heart rate variability estimation system according to claim 6, characterized in that: MeanRR (mean RR interval) is the average value of consecutive heartbeat intervals, reflecting the overall level of heart rate. The normal range is 800-1200ms. The calculation formula is: Where is the i-th RR interval, N is the total number of RR intervals; SDNN (overall standard deviation) is the standard deviation of the RR interval, reflecting the overall variability of heart rate. The normal range is 102-180ms. The calculation formula is: RMSSD (mean square of difference) is the root mean square of the difference between adjacent RR intervals, reflecting the fast-changing components in HRV. The normal range is 15-39ms. The calculation formula is: SDSD (Successive Difference of Successive Differences) refers to the standard deviation of the difference between two adjacent RR intervals, which can reflect the effect of parasympathetic nerve activity on heart rate. The normal range is 20-50ms. The calculation formula is: PNN50 (Percentage of Successive RR Intervals Differing by More Than 50ms) indicates the ratio of the difference between adjacent RR intervals exceeding 50 milliseconds within a period of time, which reflects the degree of fluctuation between heartbeats, and the normal range is 10%-30%; The following evaluation indicators are used to evaluate the system: MAE (mean absolute error) is a commonly used indicator to measure the average absolute difference between the predicted value and the actual value. It can reflect the average absolute difference between the predicted value and the actual value. MAE is defined as: Where N is the number of data points, y i is the actual value, is the predicted value. RMSE (root mean square error) is another commonly used indicator to measure the difference between the predicted value and the actual value. It can reflect the standard deviation of the prediction error. A larger error will have a greater impact on RMSE. RMSE is defined as: The Pearson Correlation Coefficient (r) is an indicator to measure the linear correlation between the predicted value and the actual value. It reflects the strength and direction of the relationship between the two variables. The value range is between 1 and -1. The closer to 1, the stronger the positive correlation, and the closer to -1, the stronger the negative correlation. The Pearson correlation coefficient is defined as: Evaluation is carried out through the above indicators, so as to make targeted adjustments.
Citation Information
Cited By
Parkinson's disease assessment method and system based on face video multi-modal physiological feature fusion
CN120809239A