RPPG signal extraction method based on 3D convolutional neural network (3D CNN)

By combining 3D CNN and attention mechanism in rPPG signal extraction, robust adaptation to rPPG signals is achieved when dealing with complex scenarios, solving the problem of signal-to-noise ratio drop in traditional methods under dynamic interference, and significantly improving the accuracy and robustness of signal extraction.

CN120198944APending Publication Date: 2025-06-24BEIJING QINGFENG QIHANG TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510256400.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The traditional rPPG method performs poorly when dealing with complex scenarios, especially when facing dynamic interference such as lighting changes and motion interference, the signal-to-noise ratio drops significantly, resulting in a large error in heart rate estimation.

Method used

The combination of 3D convolutional neural network (3D CNN) and attention mechanism is used to achieve robust adaptation to complex scenarios. By obtaining the target face video, differential features on the continuous time axis are extracted to form spatial and fusion features, and enhance the spatial and fusion features through the encoder and decoder modules, and finally extract the rPPG signal.

Benefits of technology

The time series dimension information is added through the 3D CNN network, which can effectively improve the perception of rPPG signal fluctuations in video, improve the accuracy and robustness of rPPG signal extraction, and reduce the heart rate estimation error to less than 1BPM.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198944A_ABST
    Figure CN120198944A_ABST
Patent Text Reader

Abstract

The invention discloses a remote photoplethysmography (rPPG) signal extraction method based on a 3D convolutional neural network (3D CNN), and belongs to the technical field of image processing and physiological signal detection. In order to solve the problem that space-time information utilization is insufficient in a complex scene (illumination change and motion interference) in a traditional method, space-time feature fusion is carried out through 3D CNN, and feature expression is enhanced in combination with an encoder-decoder structure and an attention mechanism. The method comprises the following steps: extracting continuous time difference features of a face video, generating space-time fusion features through a 3D CNN, enhancing the space-time fusion features through a codec, and outputting rPPG signals. A self-supervised training method is innovatively proposed, and the generalization ability of the model is improved by adopting data enhancement, time period alignment and positive and negative sample comparison. Experiments show that according to the method, the heart rate estimation error is reduced to be within 1 BPM, the signal-to-noise ratio is increased to 1.8 dB, and the motion robustness is superior to that of a traditional algorithm. The method is suitable for non-contact heart rate monitoring, respiratory rate detection and other scenes, has the characteristics of high precision and strong anti-interference, and provides a reliable technical scheme for intelligent health monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing and physiological signal detection, and particularly to a method for extracting remote photoplethysmogram (rPPG) signals based on a 3D convolutional neural network (3D CNN), which is used to non-contact extract physiological signals such as heart rate and respiratory rate from video sequences. Background Art

[0002] Remote photoplethysmogram (rPPG) is a technology for extracting physiological signals by analyzing minute color changes on the skin surface in a video. Traditional rPPG methods usually rely on manually designed feature extraction and signal processing techniques, and these methods perform poorly when dealing with complex scenarios (such as illumination changes, motion interference, etc.). In recent years, deep learning techniques, especially convolutional neural networks (CNNs), have made remarkable progress in the field of image processing. However, traditional 2D CNNs have limitations in processing time series data and cannot fully utilize the temporal information in videos. Traditional rPPG methods (such as the remote plethysmography technique based on ambient light proposed by Verkruysse et al.) rely on static scene assumptions and manually designed spectral analysis (Verkruysse, W., Svaasand, L. O., & Nelson, J. S. (2008). Remote plethysmographic imaging using ambient light. Optics Express, 16(26), 21434-21445.). However, when dealing with dynamic interferences (such as facial micro-movements, sudden illumination changes), due to the lack of spatio-temporal joint modeling ability, the signal-to-noise ratio of the signal drops significantly (typical heart rate error ≥ 3.2 BPM). In contrast, the present invention realizes robust adaptation to complex scenarios through the combination of 3D CNN and attention mechanism. After testing, the heart rate estimation error using 3D CNN is reduced to within 1 BPM, with obvious advantages. Summary of the Invention

[0003] In view of the above technical problems, the present invention provides a method for extracting remote photoplethysmography (rPPG) signals based on a 3D convolutional neural network (3DCNN). By introducing the spatio-temporal joint modeling method of 3D CNN, the present invention is trained on the publicly available dataset UBFC-rPPG and uses the synchronous heart rate labels of this dataset. This application obtains a target face video, and determines the continuous difference features of the face video on the continuous time axis according to the face video; determines the continuous spatio-temporal fusion features of the target face video according to the continuous face video difference features; determines the encoder features according to the continuous spatio-temporal fusion features; determines the decoder features from the encoder features; determines the enhanced spatio-temporal fusion features from the decoder features; and extracts the rPPG signal from the enhanced spatio-temporal fusion features and the encoder. By adding the time series dimension information through the 3D CNN network, the perception of the rPPG signal fluctuations in the video can be effectively improved, and the accuracy and robustness of the rPPG signal extraction can be enhanced.

[0004]

[0005] To solve the above technical problems, the technical solution provided by the present invention includes three aspects.

[0006] In the first aspect, this application provides a method for extracting rPPG signals, including: obtaining a target face video, and determining the continuous difference features on the continuous time axis in the target face video according to the target face video; determining the spatio-temporal fusion features of the target face video according to the continuous difference features; determining the encoder features according to the continuous spatio-temporal fusion features; determining the decoder features from the encoder features, and outputting the enhanced spatio-temporal fusion features; and extracting the rPPG signal from the enhanced spatio-temporal fusion features and the encoder.

[0007] In some embodiments, determining the continuous difference features on the continuous time axis in the target face video according to the target face video includes: generating a plurality of continuous temporal displacement images according to the effective region of the target face video; and determining the difference features of the target face video according to the plurality of continuous temporal displacement images.

[0008] In some embodiments, determining the spatio-temporal fusion features of the target face video according to the continuous difference features includes: determining the spatial features according to the continuous difference features; determining the continuous features according to the continuous difference features; and performing splicing or weighted fusion on the spatial features and the time difference features to form the spatio-temporal fusion features.

[0009] In some embodiments, determining the encoder features based on the continuous spatio-temporal fusion features includes: obtaining the encoder features with temporal features by passing the continuous spatio-temporal fusion features through multiple 3D encoder modules; fusing the continuous temporal features obtained during the encoding process with other dimensional features to obtain encoded enhanced features;

[0010] In some embodiments, determining the decoder features based on the encoder features includes: obtaining the decoder features with temporal features by passing the encoder enhanced feature information through multiple decoder modules; fusing the continuous temporal features obtained during the decoding process with other dimensional features to obtain enhanced decoding features; determining the enhanced spatio-temporal fusion features based on the enhanced decoding features;

[0011] In some embodiments, extracting the rPPG signal from the new spatio-temporal fusion features and the encoder includes: determining the rPPG signal by passing the enhanced spatio-temporal fusion features through the encoder;

[0012] Masked Autoencoder

[0013] In a second aspect, a training method for a self-supervised neural network model is proposed in the present application. The self-supervised neural network model is used to perform any of the steps in the first aspect. The method includes: obtaining a training video to be processed; performing cropping and alignment processing on the training video to obtain a training set; inputting the training set into the neural network model to obtain the output rPPG signal of the neural network model; determining the self-similarity loss of the self-supervised neural network model based on the rPPG signal and the label of the training set; adjusting the self-supervised neural network model according to the self-similarity loss.

[0014] In some embodiments, performing cropping processing on the training video to obtain a training set includes: determining facial coordinates based on the training video; determining the face center point of each frame image based on the facial coordinates; determining the face segmentation region of each frame image based on the face center point; performing time-axis alignment based on the face segmentation region; performing facial region cropping and alignment on the training video based on the face segmentation region to form a training set.

[0015] In some embodiments, determining the self-similarity loss of the self-supervised neural network model based on the rPPG signal and the label of the training set includes: performing data augmentation on the rPPG signal; performing time-period alignment on the rPPG signal; constructing positive and negative samples for the rPPG; determining the self-similarity loss based on the signal training.

[0016] In some embodiments, adjusting the self-supervised neural network model according to the self-similarity loss includes: integrating loss components according to the self-similarity loss; modifying the data augmentation strategy according to the self-similarity loss.

[0017] In a third aspect, the present invention proposes a method for extracting rPPG signals, including: a first module configured to obtain a target face video and configured to extract pixel-level color change features from the face video through an optical flow compensation algorithm, specifically including: performing a temporal difference operation ΔI t} on the continuous frame sequence {I t =I t+1 -I t , and performing motion correction based on the Horn-Schunck optical flow field; a second module, according to the first module, configured to perform spatio-temporal feature fusion through a 3D CNN network, including: using 3D convolution operations to extract spatial features, F t space =Conv3D(F diff [t-Δt:t+Δt,:,:,:]); extracting continuous difference features The spatio-temporal feature fusion is F fused ={F1 fused ,F2 fused ,...,F T fused}}. A third module, according to the second module, performs a masking operation on the spatio-temporal fusion features, with the formula X masked =M⊙X; performs an encoding operation on the spatio-temporal fusion features, with the formula Finally, the encoded enhanced feature Z multi-head =Concat(Z1,Z2,...,Z h )W O ; A fourth module, according to the second module, performs a masking operation on the spatio-temporal fusion features, with the formula X masked =M⊙X; performs an encoding operation on the spatio-temporal fusion features, with the formula Performs a linear projection on the enhanced decoded feature, with the formula Enhanced spatio-temporal block Obtains the enhanced spatio-temporal fusion feature; A fifth module, according to the second module, extracts the rPPG signal from the enhanced spatio-temporal fusion feature and the encoder, including calculating the average value of the three RGB channels, and using The formula extracts the rPPG signal.

[0018] Advantages of the present invention: In this application, a target face video is obtained, and the differential features of the target face video are determined based on the target face video; the continuous spatio-temporal fusion features of the target face video are determined according to the differential features of the continuous face video; encoder features are determined according to the continuous spatio-temporal fusion features; decoder features are determined from the encoder features; enhanced spatio-temporal fusion features are determined from the decoder features; and rPPG signals are extracted from the enhanced spatio-temporal fusion features and the encoder. Through the enhanced spatio-temporal fusion features, the ability to perceive the fluctuations of rPPG signals in the video can be effectively improved, and the accuracy and robustness of rPPG signal extraction can be enhanced. Description of the Drawings

[0019] The scope of the present disclosure can be better understood by reading the detailed description of the following exemplary embodiments in conjunction with the accompanying drawings. The accompanying drawings included are:

[0020] Figure 1 It is an overall flowchart of a method for extracting rPPG signals based on a 3D CNN network provided by an embodiment of the present invention;

[0021] Figure 2 It is an overall flowchart of a method for training a 3D CNN network provided by an embodiment of the present invention;

[0022] Figure 3 It is a structural block diagram of a spatio-temporal feature fusion module provided by an embodiment of the present invention. Detailed Embodiments

[0023] In order to make the objectives, technical solutions, and advantages of this application clearer, the following will further describe this application in detail in conjunction with the accompanying drawings. The described embodiments should not be regarded as limitations of this application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.

[0024] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0025] If similar descriptions such as "first / second / third" appear in the application documents, the following description shall be added. In the following description, the terms "first / second / third" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that, where permitted, "first / second / third" can be interchanged in a specific order or sequence so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0026] Implement technical solution 1:

[0027] Extracting high-precision rPPG signals in complex environments poses significant challenges. Existing deep learning frameworks have difficulty effectively separating noise interference unrelated to rPPG signals when processing video data. In addition, due to the extremely weak amplitude of color changes in the facial area, accurately extracting rPPG signals from video sequences is quite technically difficult.

[0028] Regarding the problems existing in the prior art, such as Figure 1 As shown, the present application provides a method for extracting rPPG signals. The method is applied to an electronic device, and the electronic device can be a server, a mobile terminal, a computer, a cloud platform, a neural network model, etc. The functions realized by the device data processing in the embodiments of the present application can be implemented by a processor of the electronic device calling program code, where the program code can be stored in a computer storage medium. The method for extracting rPPG signals includes:

[0029] Step S1: Obtain a target face video and determine the differential features of the target face video according to the target face video.

[0030] In some embodiments, step S1 obtains a target face video and determines the continuous differential features on the continuous time axis in the target face video. Due to reasons such as movement and light, various noises in the video sequence are relatively strong, and the facial color change is very weak. The rPPG signal is a series of periodic signals, and three-dimensional features can better improve the extraction efficiency and accuracy of the rPPG signal and can effectively remove noises such as light motion artifacts.

[0031] The specific steps include:

[0032] Step S11: Extract a continuous frame sequence according to the target face video.

[0033] Step S12: Divide the frame sequence into several spatial blocks to form spatio-temporal blocks, which contain differential features such as the spatial information and implementation information of the face video. Use a 3D CNN convolutional layer to extract spatio-temporal blocks, and the extraction formula is:

[0034] P i,j,k = V[i×t:(i + 1),j×h:(j + 1)×h,k×w:(k + 1)×w,:]

[0035] Suppose a video V has a size of 32×224×224×3 (32 frames, each frame is 224x224 pixels, 3 color channels), and the size of the spatio-temporal block is selected as 4×16×16×3 (4 frames, each block is 16×16 pixels, 3 color channels). Time dimension division Height division: Width division: Total number of blocks: N = 8×14×14 = 1568, and each spatio-temporal block P i,j,k can be extracted through the above formula.

[0036] The following is an example of the 3D CNN structure for rPPG signal extraction:

[0037]

[0038] In some embodiments, determine the spatio-temporal fusion feature of the target face video according to the continuous differential feature; determine the spatial feature according to the continuous differential feature; determine the continuous feature according to the continuous differential feature; perform splicing or weighted fusion according to the spatial feature and the time differential feature to form the spatio-temporal fusion feature;

[0039] Step S2: Determine the spatio-temporal fusion feature of the target face video according to the continuous differential feature, including:

[0040] Step S21: Determine the spatio-temporal fusion feature of the target face video according to the continuous differential feature, and the formula is The extracted spatial feature can be expressed as F space ={F1 space ,F2 space ,...,F T space}

[0041] Step S22: Determine the continuous feature according to the continuous differential feature, and use the formula ΔF t time = Encoder time (V t+1 - V t ) for extraction, and the extracted continuous differential feature can be expressed as

[0042] Step S23: Perform splicing and fusion based on spatial features and temporal difference features to form spatio-temporal fusion features. The formula is F t fused = Concat(F t space , ΔF t time ). The extracted spatio-temporal fusion features can be expressed as F fused = {F1 fused , F2 fused ,..., F T fused}.

[0043] In some embodiments, determining encoder features based on the continuous spatio-temporal fusion features includes: obtaining encoder features with temporal features by passing the continuous spatio-temporal fusion features through multiple 3D encoder modules; fusing the continuous temporal features obtained during the encoding process with other dimensional features to obtain encoded enhanced features; including:

[0044] Step S3: Determine encoder features according to the continuous spatio-temporal fusion features.

[0045] Step S31: Obtain encoder features with temporal features by passing the continuous spatio-temporal fusion features through multiple 3D encoder modules. Use the formula X masked = M⊙X to perform a masking operation on the continuous spatio-temporal features.

[0046] Step S32: Obtain encoder features with temporal features by passing the continuous spatio-temporal fusion features through multiple 3D encoder modules. Use the formula to perform an encoding operation on the continuous spatio-temporal features, X input = X masked + E pos .

[0047] Step S33: Fuse the continuous temporal features obtained during the encoding process with other dimensional features to obtain encoded enhanced features. Use to calculate the attention weights. Use Z multi-head = Concat(Z1, Z2,..., Z h )W O to obtain the encoded enhanced features.

[0048] In some embodiments, determining decoder features based on the encoder features includes: obtaining decoder features with temporal features by passing the encoder enhanced feature information through multiple decoder modules; fusing the continuous temporal features obtained during the decoding process with other dimensional features to obtain enhanced decoding features; including:

[0049] Step S34: Obtain the decoder features with temporal features from the encoder enhanced feature information by passing it through multiple decoder modules. Use to mark the encoder output features N masked = rN is the masked block; Add positional encoding to retain the positional information of the spatio-temporal block, X decoder =

[0050] Concat(F encoded , M token ) + E pos to obtain the decoder features;

[0051] Step S35: Fuse the continuous temporal features obtained during the decoding process with other dimensional features to obtain enhanced decoding features. Use to calculate the attention weights. Use Z multi-head = Concat(Z1, Z2,..., Z h )W O to obtain the enhanced decoder features.

[0052] In some embodiments, determining the enhanced spatio-temporal fusion feature based on the decoder features includes: passing the spatio-temporal fusion feature through an encoder to determine the encoded enhanced information; passing the encoded enhanced information through a decoder to determine the enhanced decoding feature; determining the enhanced spatio-temporal fusion feature through the enhanced decoding feature; including:

[0053] Step S4: Determine the enhanced spatio-temporal fusion feature based on the decoder features.

[0054] Step S41: Pass the spatio-temporal fusion feature through an encoder to determine the encoded enhanced information. Step S42, pass the encoded enhanced information through a decoder to determine the enhanced decoding feature. Step S43, determine the enhanced spatio-temporal fusion feature through the enhanced decoding feature. Linearly project the enhanced decoding feature. The formula to obtain the enhanced spatio-temporal fusion feature. is the enhanced spatio-temporal block. The loss function represents the enhanced spatio-temporal block and the difference between the original spatio-temporal block X.

[0055] In some embodiments, extracting the rPPG signal from the enhanced spatio-temporal fusion feature and the encoder includes: passing the enhanced spatio-temporal fusion feature through an encoder to determine the rPPG signal, including:

[0056] Step S5: Extract the rPPG signal from the enhanced spatio-temporal fusion feature and the encoder.

[0057] Step S51: Use the formula to calculate the average value of the R channel for each frame of video data.

[0058] Step S52: Use to calculate the average value of the R channel for each frame of video data.

[0059] Step S53: Use to calculate the average value of the R channel for each frame of video data.

[0060] Step S54: Use the formula to extract the rPPG signal.

[0061] During the human blood circulation process, as the heart beats, the blood volume in blood vessels changes periodically. When the blood is full, the absorption and scattering characteristics of light will change. The green channel G t signal is relatively sensitive to this blood volume change. By normalizing the green channel signal with the average signal of the RGB channels the rPPG signal can be obtained, which can enhance the signal related to the blood volume change and reduce the interference caused by factors such as light intensity change and skin color change. By performing time-domain or frequency-domain analysis on the rPPG signal. In time-domain analysis, the heart beat will cause periodic fluctuations of the signal, and the heart rate can be calculated by measuring the time interval between adjacent peaks. In frequency-domain analysis, the frequency component corresponding to the heart rate will appear as an obvious peak in the signal spectrum, so that the heart rate variability information can be extracted.

[0062] Embodiment 2: The present invention also discloses a training method for a self-supervised neural network model, and this training method can execute any of the steps described in the first aspect. This training method is one of the implementation methods of the first aspect, and other methods can also be adopted for training. The training method includes:

[0063] Step 6: Obtain the training video data to be processed.

[0064] Step 7: Preprocess the training video data.

[0065] Step S71, use the CNN network to extract the human face, determine the important point information of the face, the center of the left eye: (x left_eye , y left_eye ), the center of the right eye: (x right_eye , y right_eye ), the nose: (x nose , y nose ), and the coordinates of the bounding box [x1, y1, x2, y2].

[0066] Step S72, determine the center point of the human face according to the important point information of the face.

[0067] Step S73, perform face segmentation regions based on the center point of the face

[0068] Step S74, perform data alignment according to the face segmentation regions

[0069] Step S75, similarity transformation, map the current face key points to the reference coordinate system by rotation, translation and scaling. The transformation matrix has the formula:

[0070]

[0071] a = s cosθ (a combination of the scaling factor s and the rotation angle θ), b = s sinθ, t x , t y : translation amount. Step S76, align the face, apply an affine transformation to each frame of the image, and map the original face region to the standard coordinate system: The formula is Step S77, loop Steps S6 - S7 to establish a training data set

[0072] In some embodiments, determining the self - similarity loss of the self - supervised neural network model according to the rPPG signal and the labels of the publicly available UBFC - rPPG data set includes: performing data augmentation on the rPPG signal; performing time - period alignment on the rPPG signal; constructing positive and negative samples for the rPPG; and determining the self - similarity loss according to the training of the signal. Step 8: Determine the self - similarity loss of the self - supervised neural network model according to the rPPG signal and the labels of the training set

[0073] Perform data augmentation on the rPPG signal, and the specific steps are as follows

[0074] Step S81: Perform two random augmentations (such as adding noise and time cropping) on the same rPPG signal x to obtain x (1) and x (2) .

[0075] Step S82: The model generates feature vectors f (1) = Enc(x (1) ), f (2) =

[0076] Enc(x (2) ). The loss formula: Minimize the cosine similarity formula between features (cosine similarity)

[0077] Perform time - period alignment on the rPPG signal, and the specific steps are as follows

[0078] Step S83: Label-guided Period Segmentation: Calculate the period length HR (unit: BPM) according to the heart rate label y (seconds), and segment the signal into multiple period segments {x t}.

[0079] Step S84: Time Period Alignment, Force the alignment of features at the same phase position (e.g., the phase alignment of the i-th period and the j-th period).

[0080] Loss function where k represents the same phase position (such as the k-th time point within a period).

[0081] Construct positive and negative samples for the rPPG; the specific steps are as follows:

[0082] Step S85: Apply two different random masks to the same input data x, generating two masked views x (1) and x (2) .

[0083] Step S86: Feature Extraction: Extract the latent features of both through the encoder Enc:

[0084] f (1) = Enc(x (1) ), f (2) = Enc(x (2) )

[0085] Positive sample pair: (f (1) , f (2) ) comes from the same data and should have a high similarity.

[0086] Step S87: Negative sample construction strategy, using the masked versions of other samples in the same batch as negative samples.

[0087] Masked negative sample: Apply a mask to different data x ′ to generate x ′(1) , and its feature f ′(1) is used as a negative sample.

[0088] Train according to the signal to determine the self-similarity loss, and the specific steps are as follows:

[0089] Step S88: Loss function design, MAE reconstruction loss: The supervised model recovers the masked content:

[0090]

[0091] Step S89: Contrastive loss: Maximize the similarity of positive samples and minimize the similarity of negative samples:

[0092]

[0093] where f (k) includes positive samples and negative samples, and τ is the temperature coefficient.

[0094] Total loss in step S810: Jointly optimize the reconstruction and contrast objectives:

[0095] α is the weight hyperparameter.

[0096] In some embodiments, rPPG signals are typically used for non-contact heart rate monitoring, which requires high robustness and accuracy of the model. It is necessary to adjust the self-supervised neural network model according to the self-similarity loss.

[0097] Self-loss function formula

[0098] Self-similarity loss (such as contrast loss, feature consistency loss).

[0099] Task loss (such as reconstruction loss of MAE, mean square error of heart rate).

[0100] Step S9: Adjust the self-supervised neural network model according to the self-similarity loss

[0101] Step S91: Dynamic weight adjustment: Dynamically adjust α according to the training stage

[0102] Step S92: Feature consistency enhancement: Force feature similarity for different augmented views of the same data.

[0103] Perform signal enhancement on rPPG:

[0104] Step S93: Temporal enhancement, including the following steps

[0105] Step S94: Random Crop: Preserve the integrity of the physiological cycle.

[0106] Step S95: Time Warping: Scale the signal within a reasonable heart rate range (such as 40 - 180 BPM) according to the requirements of "ECG Monitoring Parameter Setting Range".

[0107] Step S96: Additive noise: Add Gaussian noise or motion artifacts to simulate noise. Noise formula, 0.3 < σ < 0.6.

[0108] Step S97: Frequency domain enhancement Step S98: Band-pass filtering: Preserve the heart rate related frequency band (0.7 - 4 Hz).

[0109] Step S99: Phase perturbation: Randomly shift the signal phase.

[0110] The loss function fusion strategy aims to jointly optimize the self - similarity loss (contrast learning) and the task - related loss (such as reconstruction, classification).

[0111] Formula design:

[0112]

[0113] * Self - similarity loss (such as contrast loss, feature consistency loss).

[0114] * Task loss (such as the reconstruction loss of MAE, the mean square error of heart rate).

[0115] * α, β, γ: Weight coefficients, which need to be tuned through experiments.

[0116] Key adjustment points:

[0117] 1. Dynamic weight adjustment: Dynamically adjust α according to the training stage (such as focusing on self - supervision in the early stage and on the task in the later stage).

[0118] 2. Feature consistency enhancement: Force feature similarity for different augmented views of the same data. Example 3: Based on the foregoing embodiments, an embodiment of the present invention provides a method for extracting rPPG signals. Each module included in this method and each unit included in each module can be implemented by a software program; this software program can run on hardware devices such as personal computers, servers, mobile phones, and embedded computers. As shown in Figure x, a method for extracting rPPG signals includes: a first execution module, a second execution module, a third execution module, a fourth execution module, a fifth module, and a sixth module. Xxxx (describe the logic)

[0119] It should be noted that the division of modules in the embodiments of this application is schematic, merely a logical function division, and there may be other division methods in actual implementation.

Claims

1. A rPPG signal extraction method based on 3D convolutional neural network, characterized in that , including the following steps: Step S1: Obtain a target face video, and extract difference features on a continuous time axis in the video through a time domain difference operation and an optical flow compensation algorithm; Step S2: input the difference features into a 3D convolutional neural network, and generate spatiotemporal fusion features through a spatial feature extraction layer and a time series modeling layer; Step S3: inputting the spatiotemporal fusion features into the encoder module for dimensionality reduction encoding to obtain encoder features; Step S4: input the encoder features into the decoder module for feature reconstruction to obtain decoder features; Step S5: Fusing the encoder features and the decoder features to generate enhanced spatiotemporal fusion features; Step S6: extracting the rPPG signal from the enhanced spatiotemporal fusion features, and removing noise through bandpass filtering and blind source separation.

2. The method according to claim 1, characterized in that , the step S1 in which the difference features are extracted includes: For a continuous frame sequence {I t }Perform time domain difference operation ΔI t =I t+1 -I t , and based on the Horn-Schunck optical flow field, motion artifact compensation is performed to generate a pixel-level color change feature map 3. The method according to claim 1, characterized in that , the step S2 of generating spatiotemporal fusion features includes: Extract spatial features F through 3D convolution kernel (kernel_size = 3×3×3) spatial , and a bidirectional LSTM network is used to model time dependencies. Concatenate bidirectional hidden states to obtain spatiotemporal fusion features 4. The method according to claim 1, characterized in that , the encoder module in step S3 is a stacked Transformer layer, including: Add position encoding P to the spatiotemporal fusion feature pos =SinusoidalEncoding(T, D), and capture the global contextual relationship through the multi-head self-attention mechanism, and output low-dimensional encoding features 5. The method according to claim 1, characterized in that , the step S5 of generating enhanced spatiotemporal fusion features includes: Align encoder features F via deformable convolution enc With the decoder feature F dec The spatial and temporal dimensions of the attention weight matrix W a Perform gated fusion: F enhanced =Sigmoid(W a ·F enc )☉F dec +F enc。 6. The method according to claim 1, characterized in that , extracting the rPPG signal in step S6 includes: Perform PCA dimensionality reduction on the enhanced spatiotemporal fusion features along the channel dimension to extract the principal component signal S raw , and output the standardized rPPG waveform through adaptive bandpass filter (passband 0.7-4Hz) and independent component analysis (ICA) 7. A training method for a self-supervised neural network model, used to implement any of the methods described in claims 1-6, characterized in that ,include: Step T1: Perform face detection, key point alignment and spatiotemporal cropping on the training video to generate a standardized training set; Step T2: randomly mask the spatiotemporal blocks of the input video through a masked autoencoder and reconstruct the masked content; Step T3: Based on reconstruction loss and contrast loss Calculate self-similar loss and dynamically adjust model parameters; The contrast loss is achieved by maximizing the feature similarity of the same video enhancement view:

8. The method according to claim 1, characterized in that , the encoder module in step S3 is a stacked Transformer layer, including: Add position encoding P to the spatiotemporal fusion feature pos =SinusoidalEncoding(T, D), and capture the global contextual relationship through the multi-head self-attention mechanism, and output low-dimensional encoding features 9. The method according to claim 1, characterized in that , the step S5 of generating enhanced spatiotemporal fusion features includes: Align encoder features F via deformable convolution enc With the decoder feature F dec The spatial and temporal dimensions of the attention weight matrix W a Perform gated fusion: F enhanced =Sigmoid(W a ·F enc )☉F dec +F enc。 10. The method according to claim 1, characterized in that , extracting the rPPG signal in step S6 includes: Perform PCA dimensionality reduction on the enhanced spatiotemporal fusion features along the channel dimension to extract the principal component signal S raw , and output the standardized rPPG waveform through adaptive bandpass filter (passband 0.7-4Hz) and independent component analysis (ICA)

Citation Information

Cited By

  • Video segmentation method and system

    CN120913134A