Remote photoplethysmography method and system based on long and short term space-time convolution network
Through multi-scale modeling and dynamic feature fusion of long-term and short-term spatiotemporal convolution networks, the signal accuracy problem of remote photoplethysmography technology under light and motion interference is solved, and high-precision physiological parameter monitoring is achieved.
Patent Information
- Application Number
- CN202510448847.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-10
AI Technical Summary
When facing challenges such as light interference, motion artifacts and reduced signal amplitude for dark-skinned people, existing remote photoplethysmography (rPPG) technology has problems such as decreased signal-to-noise ratio, increased phase error in motion scenes and imbalance in multi-parameter accuracy, making it difficult to achieve high-precision physiological parameter monitoring.
A method based on long and short-term spatiotemporal convolution network is adopted, combining 3D convolution with 1D expansion convolution, and a spatiotemporal graph convolution and self-attention mechanism are introduced. Through multi-scale spatiotemporal modeling and dynamic feature fusion, anti-interference capabilities are enhanced, and computational efficiency is improved through end-to-end optimization and lightweight design.
It effectively improves the accuracy and robustness of remote photoplethysmographic signals, can achieve high-quality physiological parameter extraction in complex environments, and reduces the impact of motion and light interference.
Smart Images

Figure CN120299077A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the cross - field of biomedical signal processing and computer vision, and particularly relates to a non - contact physiological signal detection method and system integrating spatio - temporal convolutional network and long - short - term memory network. Background Art
[0002] Photoplethysmography (PPG) technology monitors physiological parameters by detecting the absorption changes of blood for light of specific wavelengths. Traditional contact sensors are likely to cause skin discomfort, and the contact pressure fluctuation will increase the heart rate error by 15% - 20% (Smith et al., "Contact Pressure Effects on PPG Signals", IEEE TBME, 2017), so there are application limitations in groups such as burn patients. Remote PPG (rPPG) uses a camera to capture the subtle skin light changes for non - contact monitoring, but faces three major challenges: 1) The illumination interference causes the signal - to - noise ratio to drop by more than 40% (when the fluctuation is > 200 lux, Wang et al., "Illumination - Robust rPPG via Chromaticity Separation", CVPR, 2021.); 2) Micro - movements of the head (> 3 mm) result in motion artifacts, and the motion state error reaches 5 - 8 bpm; 3) The signal amplitude of people with dark skin tones is reduced by about 60% (McDuff et al., "The Impact of Skin Type on rPPG", Nature NPJ Digital Medicine, 2020.).
[0003] Among the existing deep - learning methods, the CNN model (such as PhysNet) lacks the ability of temporal modeling and is difficult to capture heart rate variability; the RNN solution has insufficient suppression of motion artifacts. Although the ST - rPPG network (Chen et al., "ST - rPPG: Spatial - Temporal Network for Remote Physiological Measurement", ICCV, 2021.) combines 3D - CNN and LSTM, there are still problems such as noise transmission caused by shallow spatio - temporal coupling and the fixed - weight fusion mechanism not adapting to dynamic scenarios. Its phase error in the motion scenario is 2.3 times higher than that in the static scenario.
[0004] The current technical bottlenecks are concentrated in: ① Spatio - temporal feature mismatch - The cascade processing ignores the spatio - temporal physical correlation of microvascular movement; ② Weak dynamic adaptability - It is difficult to cope with the coordinated changes of illumination - motion - physiological state; ③ Imbalance of multi - parameter accuracy - The blood oxygen detection error of the existing single - camera system exceeds ± 3%, exceeding the clinical standard of ± 2%. These defects restrict the reliable application of rPPG technology in real scenarios. Summary of the Invention
[0005] In view of the above technical problems, the present invention provides a method for extracting remote photoplethysmogram (rPPG) signals based on a long short-term spatio-temporal convolutional network. This method uses multi-scale spatio-temporal modeling, combines 3D convolution (local spatio-temporal features) with 1D dilated convolution (long-range periodic patterns), and introduces spatio-temporal graph convolution (GCN) to model spatial topological relationships, achieving dual capture of microscopic transient and macroscopic periodic features; dynamic feature fusion, generating spatio-temporal weights through a self-attention mechanism, and combining a gated residual structure (such as a Sigmoid-gated β t ) to adaptively fuse long- and short-term features and enhance anti-interference ability; lightweight and efficient design, using depthwise separable convolution, channel shuffle, etc. to compress the number of parameters, and combining parallel GPU acceleration and spatio-temporal pooling strategies to improve computational efficiency; end-to-end joint optimization, jointly using time-domain (MSE loss) and frequency-domain (KL divergence) supervision, and integrating a probability prediction framework (TCN encoder-decoder) to ensure that the signal has both waveform accuracy and physiological interpretability. It can effectively improve the perception of rPPG signal fluctuations in videos and enhance the accuracy and robustness of rPPG signal extraction.
[0006] The technical solution provided by the present invention includes the following.
[0007] The present application provides an rPPG signal processing method based on a long short-term spatio-temporal convolutional network, including: determining preprocessed data according to the acquired face video; determining short spatio-temporal convolutional features according to the preprocessed data; determining long spatio-temporal convolutional features according to the preprocessed data; determining dynamic spatio-temporal fusion features according to the short and long spatio-temporal convolutional features; and determining rPPG signals according to the dynamic spatio-temporal fusion features.
[0008] In some embodiments, determining preprocessed data according to the acquired face video; determining the ROI region of the target face video according to the target face video; determining the face ROI motion points according to the face ROI region; determining the face ROI segmentation and noise reduction information according to the face ROI motion points; determining the YUV color conversion space according to the face ROI segmentation and noise reduction information; and determining the time series according to the YUV color conversion space.
[0009] In some embodiments, determining short spatio-temporal convolutional features according to the preprocessed data; determining local spatio-temporal information in the video according to the time series; determining spatio-temporal non-linear activation information according to the local spatio-temporal information; determining spatio-temporal pooling information according to the spatio-temporal non-linear activation information; determining normalization information according to the spatio-temporal pooling information; determining residual connection information according to the normalization information; and determining short spatio-temporal convolutional features according to the residual connection information;
[0010] In some embodiments, long-term spatio-temporal convolution features are determined based on the preprocessed data; dilated temporal convolution in the video is determined according to the time series; temporal non-linear activation information is determined based on the dilated temporal convolution; temporal pooling information is determined according to the temporal non-linear activation information; normalization information is determined according to the temporal pooling information; residual connection information is determined according to the normalization information; temporal attention mechanism information is determined according to the residual connection information; and long-term spatio-temporal feature information is determined based on the above information.
[0011] In some embodiments, dynamic spatio-temporal fusion features are determined based on the long and short spatio-temporal convolution features; feature alignment information is determined according to the short spatio-temporal convolution features and the long-term spatio-temporal feature information; joint feature representation is determined according to the feature alignment information; attention weight information is determined according to the joint feature; and gated fusion information and residual output information are determined according to the attention weight information.
[0012] In some embodiments, an rPPG signal is determined based on the dynamic spatio-temporal fusion features. A 1D signal is determined according to the dynamic spatio-temporal fusion features; and a high-quality rPPG signal is determined by performing band-pass filtering on the 1D signal.
[0013] Advantages of the present invention: In this application, a target face video is obtained, and an effective region of the target face video is determined according to the target face video; long and short spatio-temporal convolution features of the target face video are determined according to the effective region of the target face video; dynamic spatio-temporal fusion features are determined according to the long and short spatio-temporal convolution features; a high-quality rPPG signal is determined according to the determined spatio-temporal fusion features; and by using the long and short spatio-temporal convolution method, environmental interferences such as moving light can be effectively avoided, and the accuracy and robustness of rPPG signal extraction are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The scope of the present disclosure can be better understood by reading the following detailed description of exemplary embodiments in conjunction with the accompanying drawings. The accompanying drawings include:
[0015] Figure 1 An overall flowchart for extracting an rPPG signal based on a long short-term spatio-temporal convolutional network provided by an embodiment of the present invention;
[0016] Figure 2 A flowchart for extracting short-term spatio-temporal information provided by an embodiment of the present invention;
[0017] Figure 3 A flowchart for extracting long-term spatio-temporal information provided by an embodiment of the present invention;
[0018] Figure 4 A structural block diagram for long and short spatio-temporal feature fusion provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0019] Implementing high-precision remote photoplethysmography (rPPG) signal detection under complex environmental interference poses severe technical challenges. When existing deep learning models process video data, it is difficult to effectively separate noise components unrelated to physiological signals (such as sudden changes in ambient light, facial motion artifacts, etc.), and there is a significant coupling effect between the weak light intensity changes related to blood flow on the facial skin surface (usually less than 0.1% brightness fluctuation) and background noise, resulting in a serious deterioration of the signal extraction signal-to-noise ratio. Therefore, as Figure 1 shown, this application proposes an rPPG signal extraction method based on a long short-term spatio-temporal convolutional network, which can be deployed on various electronic device platforms, including but not limited to: cloud servers, mobile terminals (smartphones / tablets), embedded computers, and dedicated neural network accelerators, etc. The extraction method of the rPPG signal includes:
[0020] In some embodiments, preprocessing data is determined according to the acquired face video; the ROI region of the target face video is determined according to the target face video; the face ROI face motion points are determined according to the face ROI region; the face ROI segmentation noise reduction information is determined according to the face ROI motion points; the YUV color conversion space is determined according to the face ROI segmentation noise reduction information; the time series is determined according to the YUV color conversion space.
[0021] Step S1: Acquire the target face video, and determine preprocessing data according to the acquired face video. The specific steps include:
[0022] Step S11, determine the ROI region of the target face video according to the target face video, and use the MTCNN algorithm for face monitoring. First, use the P-Net to quickly generate candidate face regions, then use the R-Net to filter out false detection candidate boxes, and finally use the O-Net to further refine the position.
[0023] Step S12, determine the face ROI face motion points according to the face ROI region. First, calculate the grayscale images of the current frame and the next frame, and characterize the pixel changes through the image gradient (x-direction derivative I x , y-direction derivative I y ) and the time derivative I t ; establish the optical flow equation, I x ·u + I y ·v + I t = 0; solve the displacement vector through the Lucas-Kanade algorithm; perform pixel aggregation, and average the RGB values of all pixels in the ROI by channel to generate a time series signal (R mean, G mean, B mean of each frame).
[0024] Step S13. Determine the face ROI segmentation noise reduction information based on the face ROI motion points, perform Gaussian blur or median filtering on the ROI region to reduce camera noise; use a skin mask to exclude non-skin pixels (such as the background and hair).
[0025] Step S14. Determine the YUV color conversion space based on the face ROI segmentation noise reduction information, and use the CHROM algorithm to convert the RGB signal into a YUV color space that is more sensitive to blood flow. Step S15. Determine the time series based on the YUV color conversion space, and extract a spatiotemporal cube from the preprocessed video, with the shape of: where T is the number of time steps (number of frames), H×W is the spatial resolution, and C is the number of color channels.
[0026] In some embodiments, determine the short spatiotemporal convolutional features based on the preprocessing data; determine the local spatiotemporal information in the video based on the time series; determine the spatiotemporal non-linear activation information based on the local spatiotemporal information; determine the spatiotemporal pooling information based on the spatiotemporal non-linear activation information; determine the normalization information based on the spatiotemporal pooling information; determine the residual connection information based on the normalization information; determine the short spatiotemporal convolutional features based on the residual connection information;
[0027] Step S2. Determine the short spatiotemporal convolutional features of the target face video based on the effective region of the target face video;
[0028] Step S21. Determine the local spatiotemporal information in the video based on the time series, and set the 3D convolution kernel parameters as where, k t : the size of the convolution kernel in the time dimension (such as 3 frames), k h , k w : the size of the convolution kernel in the spatial dimension (such as 3×3), D s : the output feature dimension (number of channels). For time step t, input the local time window The convolution output is The output spatial dimensions H′, W′ are determined by the convolution stride and padding. Typical parameters: k t =3, k h =k w =3, stride (1,1,1). Step S22. Determine the spatiotemporal non-linear activation information based on the local spatiotemporal information. The purpose of introducing non-linear activation is to enhance the model's expression ability. The activation function
[0029] Step S23: Determine the spatio-temporal pooling information based on the spatio-temporal non-linear activation information. Spatio-temporal pooling can reduce the spatial resolution, enhance the translational invariance, and reduce the computational amount. In the present invention, max pooling is performed along the spatial dimension, and the formula Pooling kernel size: 2×2, stride 2, and the spatial size is halved.
[0030] Step S24: Determine the normalization information based on the spatio-temporal pooling information, and normalize each feature channel. The formula Step S25: Determine the residual connection information based on the normalization information, and alleviate the problem of gradient disappearance through the residual connection, and retain the shallow detail information. If the input and output dimensions match, fuse the features through the skip connection: If the dimensions do not match, the number of channels needs to be adjusted through a 1×1×1 convolution.
[0031] Step S26: Determine the short spatio-temporal convolution features based on the residual connection information, compress the multi-dimensional features into a time series for subsequent fusion. Global average pooling (GAP): Aggregate information along the spatial dimension H″×W″, Time alignment: Arrange the features of each time step in sequence to form a short-term feature sequence:
[0032] In some embodiments, determine the long spatio-temporal convolution features based on the preprocessed data; determine the dilated temporal convolution in the video according to the time series; determine the temporal non-linear activation information according to the dilated temporal convolution; determine the temporal pooling information according to the temporal non-linear activation information; determine the normalization information according to the temporal pooling information; determine the residual connection information according to the normalization information; determine the temporal attention mechanism information according to the residual connection information; and determine the long-term spatio-temporal feature information according to the above information.
[0033] Step S3: Determine the long spatio-temporal convolution features based on the preprocessed data.
[0034] Step S31: Determine the dilated temporal convolution in the video according to the time series. The dilated convolution increases the temporal receptive field and captures long-term periodic patterns (such as heart rhythms). Apply a one-dimensional dilated convolution in the temporal dimension and ignore the spatial dimension (the spatial information has been extracted by the short-term branch). Mathematical formula: Let the parameters of the dilated convolution be where: k t : Convolution kernel size (such as 3), D l: Output feature dimension, r: dilation rate (e.g., r = 4). For time point t, the actual time window covered by the convolution is: t start = t - r×(k t - 1) / 2, t end = t + r×(k t - 1) / 2. The output feature is calculated as
[0035] Step S32: Determine the time non-linear activation information according to the dilated-time convolution. The non-linear activation function enhances the non-linear expression ability. The activation function
[0036] Step S33: Determine the time pooling information according to the time non-linear activation information. The pooling process can reduce the time resolution and enhance the robustness to low-frequency periodic signals. Perform max pooling along the time dimension:
[0037] Step S34: Determine the normalization information according to the time pooling information. The normalization process can stabilize the training process and accelerate convergence. Step S35: Determine the residual connection information according to the normalization information. The residual connection is used to alleviate the vanishing gradient and retain the long-term trend of the original time series. If the input dimensions match, fuse the starting input and the convolution result through a skip connection:
[0038] If the input dimensions do not match, use a 1×1×1 convolution to adjust the number of channels.
[0039] Step S36: Determine the time attention mechanism information according to the residual connection information. Adaptively focus on the key time steps (such as the heart rate peak moment), and introduce the time attention weight β t ∈[0,1], where is the short-term branch feature, σ is the Sigmoid function, W a , b a are learnable parameters, and the weighted output is: Step S37: Determine the long spatio-temporal feature information according to the above information. The long-term spatio-temporal feature formula can be expressed as:
[0040] In some embodiments, dynamic spatio-temporal fusion features are determined based on the long and short spatio-temporal convolution features; feature alignment information is determined based on the short spatio-temporal convolution features and long-term spatio-temporal feature information; joint feature representations are determined based on the feature alignment information; attention weight information is determined based on the joint features; and gated fusion information and residual output information are determined based on the attention weight information.
[0041] Step 4: Determine dynamic spatio-temporal fusion features from the long and short spatio-temporal convolution features;
[0042] Step 41: Determine feature alignment information based on the short spatio-temporal convolution features and long-term spatio-temporal feature information, and map the long and short features to the same dimension D through linear mapping:
[0043] Step 42: Determine joint feature representations based on the feature alignment information and use concatenation for joint representation.
[0044] Step 43: Determine attention weight information based on the joint features. Through the self-attention mechanism, dynamically learn the fusion weights of short-term features at each time step to generate Query, Key, and Value.
[0045] Project the joint feature X into the query, key, and value spaces:
[0046]
[0047]
[0048] Calculate the attention scores, measure the correlation between time steps through dot product, and scale to prevent gradient explosion. Finally, perform normalized weight calculation, apply Softmax to each row to obtain the normalized attention weight matrix: where α i,j represents the degree of attention of time step i to j.
[0049] Step 44: Determine gated fusion information and residual output information based on the attention weight information. Aggregate according to the context features, weight the values using the attention weights to generate context-aware global features. Weighted fusion of long and short-term features, apply the gating weight β t ∈[0,1], dynamically allocate long and short-term features: β t =σ(W g [L′ t ; S′ t +b g ), where σ is the Sigmoid function, and the final fused feature is In some embodiments, the rPPG signal is determined according to the dynamic spatio-temporal fusion feature. A 1D signal is determined according to the dynamic spatio-temporal fusion feature; a high-quality rPPG signal is determined by performing band-pass filtering on the 1D signal.
[0050] Step 5: Determine the rPPG signal according to the dynamic spatio-temporal fusion feature.
[0051] Step 51: Determine a 1D signal according to the dynamic spatio-temporal fusion feature, and decode the fused feature into a 1D signal through a fully connected layer: Connection layer
[0052] Step 52: Determine a high-quality rPPG signal by performing band-pass filtering on the 1D signal. Perform band-pass filtering (0.5 - 4 Hz, corresponding to 30 - 240 BPM) on the generated raw rPPG signal to filter out high-frequency noise (such as motion artifacts) and baseline drift, and obtain a high-quality rPPG signal, rPPG filtered = Butterworth_Bandpass(rPPG raw )
Claims
1. A remote photoplethysmogram (rPPG) signal processing method based on a long short-term spatio-temporal convolutional network, characterized in that It includes the following steps: Step S1: Obtain a target face video and determine preprocessing data according to the face video; Step S2: Extract short-term spatio-temporal convolutional features according to the preprocessing data, and the short-term spatio-temporal convolutional features capture local spatio-temporal patterns through a 3D convolutional network; Step S3: Extract long-term spatio-temporal convolutional features according to the preprocessing data, and the long-term spatio-temporal convolutional features capture long-range periodic patterns through a 1D dilated convolutional network; Step S4: Perform dynamic spatio-temporal fusion on the short-term spatio-temporal convolutional features and the long-term spatio-temporal convolutional features to generate fused features; Step S5: Determine the rPPG signal according to the fused features.
2. The method according to claim 1, characterized in that, Determining the preprocessing data in step S1 includes: Extracting the face ROI region based on the MTCNN algorithm; Calculating the moving points of the ROI region through the optical flow equation; Performing noise reduction processing and YUV color space conversion on the ROI region to generate a time series signal.
3. The method according to claim 2, characterized in that, The noise reduction processing includes: Performing Gaussian blur or median filtering on the ROI region; Excluding the interference of non-skin pixels based on the skin color mask.
4. The method according to claim 1, characterized in that Specifically, extracting the short-term spatio-temporal convolutional features in step S2 includes: Performing 3D convolution operation on the time series to capture local spatio-temporal information; Extracting high-dimensional features through spatio-temporal non-linear activation functions and spatio-temporal pooling operations; Adopting residual connection to fuse shallow and deep features to generate a short-term spatio-temporal feature sequence.
5. The method according to claim 4, wherein The kernel parameters of the 3D convolution operation are: the convolution kernel size K in the time dimension t = 3, and the convolution kernel size k in the spatial dimension h × k w = 3 × 3, and the stride is (1, 1, 1).
6. The method according to claim 1, wherein Specifically, extracting the long-term spatio-temporal convolutional features in step S3 includes: performing 1D dilated convolution operation on the time series to expand the time receptive field, with the dilation rate r≥2; weighting the features of key time steps through a time attention mechanism; combining residual connection and normalization operations to generate a long-term spatio-temporal feature sequence.
7. The method according to claim 6, wherein The time attention mechanism generates a weight β through the Sigmoid function t , satisfying: β t = σ(W a [L t ; S t ) + b a , where L t is the long-term feature, S t is the short-term feature, W a and b a are learnable parameters.
8. The method according to claim 1, characterized in that, The dynamic spatio-temporal fusion in step S4 includes: aligning short spatio-temporal features and long spatio-temporal features to the same dimension; calculating a spatio-temporal weight matrix through a self-attention mechanism; and adaptively fusing long- and short-term features using a gated residual structure. The final fused feature is: where β t is the gating weight, and o t is the global feature aggregated by self-attention.
9. The method according to claim 1, wherein Determining the rPPG signal in step S5 includes: Decoding the fused features into a 1D raw signal through a fully connected layer; Performing band-pass filtering (0.5 - 4Hz) on the raw signal to filter out high-frequency noise and baseline drift and generate a high-quality rPPG signal.
10. A remote photoplethysmogram (rPPG) signal processing system based on a long short-term spatio-temporal convolutional network, characterized in that, It includes: A preprocessing module for performing step S1 described in any one of claims 1 - 9; A short-term spatio-temporal feature extraction module for performing step S2 described in any one of claims 1 - 9; A long-term spatio-temporal feature extraction module for performing step S3 described in any one of claims 1 - 9; A dynamic fusion module for performing step S4 described in any one of claims 1 - 9; A signal generation module for performing step S5 described in any one of claims 1 - 9.
Citation Information
Patent Citations
Construction method and detection method of short-time rPPG signal detection model
CN113920387A
Remote plethysmography signal detection model construction method and device, remote plethysmography signal detection method and device and application
CN114628020A
Non-contact heart rate calculation method based on deep spatial-temporal characteristics
CN116013499A
Motion interference robust non-contact heartbeat physiological signal measurement method and system
CN117743832A
Method and device for extracting rPPG signal, medium, equipment and model training method
CN119049106A
Cited By
Lightweight identity authentication method based on remote photoelectric volume pulse wave signals
CN121744292A
Vital sign monitoring method and system based on multi-modal perception and space-time restoration
CN121971061A
Vital sign monitoring method and system based on multi-modal perception and space-time repair
CN121971061B
RPPG physiological index estimation method and system based on multi-scale time sequence Mama architecture
CN122163177A