Face video heart rate detection method based on separable convolution
By constructing a face video heart rate detection method based on separable convolution, the space-time differential multi-head attention mechanism using nearest neighbor features and time-differential convolution is used to solve the problem of limited detection performance of the remote photoelectric capacitive pulse wave method under complex conditions, and efficient heart rate detection and robustness improvement are achieved.
Patent Information
- Application Number
- CN202510283678.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-15
AI Technical Summary
The existing remote photoelectric volume pulse wave method has limited heart rate detection performance under factors such as lighting conditions, skin color and facial angle in video, and the model parameters are complex and the calculation amount is large.
The heart rate detection method of face video based on separable convolution is adopted, including convolution-gated linear units of nearest neighbor features, spatiotemporal differential multi-head attention mechanism of time-differential convolution, and spatiotemporal feedforward network of separable convolution, heart rate data is obtained by refining local spatiotemporal features and suppressing noise.
Improves the accuracy and robustness of heart rate detection, reduces computational complexity, and is suitable for intelligent health monitoring and telemedicine fields.
Smart Images

Figure CN120318148A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent detection, and specifically relates to a face video heart rate detection method based on separable convolution. Background Art
[0002] Physiological signals such as heart rate can usually evaluate a person's health condition and predict the risk of potential diseases. Heart rate measurement can be performed by a heart rate meter. Although the measurement is accurate, contact measurement devices have some disadvantages. For example, long-term use can cause discomfort, and factors such as skin temperature, color, and patient movement can also affect the accuracy of its measurement. Therefore, researchers have begun to explore non-contact remote heart rate detection methods. Currently, the remote photoplethysmography (rPPG) method does not require wearing a contact device. Although it avoids the various disadvantages of the contact device, it is also suitable for long-term continuous monitoring.
[0003] The remote photoplethysmography method is mainly divided into the conventional rPPG method and the rPPG method based on deep learning. With the development of deep learning, the remote heart rate estimation method based on deep learning has become the mainstream. In the prior art, to improve the accuracy of rPPG estimated heart rate, researchers proposed a heart rate estimation model based on a multi-head self-attention mechanism, designed a spatio-temporal representation conversion method to suppress the noise introduced by non-interested regions. Researchers proposed a heart rate estimation method based on a 3D residual attention network, which fuses a 3D convolutional attention mechanism and spatio-temporal convolution, facilitating the model to extract channel and spatial feature information and focus on regions rich in physiological signals. The disadvantage is that this method does not consider the influence of factors such as light under complex conditions on heart rate detection. To reduce the influence of light, researchers constructed a self-supervised machine learning network for robust heart rate measurement, which was trained without using any labels and achieved ideal results. The disadvantage is that the model parameters are relatively complex and the computational load is large. To solve the problem of being limited by a large amount of data and improve the ability to cope with noise, researchers proposed a linear self-supervised reconstruction model with self-similar prior, designed a noise-insensitive strategy to reduce the interference of movement and illumination. Generally speaking, the heart rate detection method based on remote video has achieved good performance advantages, but factors such as light conditions, skin color, and the face angle in the video limit the performance of the remote photoplethysmography method. Summary of the Invention
[0004] The purpose of the present invention is to overcome the above-mentioned disadvantages and propose a face video heart rate detection method based on separable convolution with good non-contact heart rate detection performance.
[0005] A face video heart rate detection method based on separable convolution according to the present invention, wherein: the specific steps of the method include:
[0006] Step 1: Data collection: Collect face video data and perform preprocessing;
[0007] Step 2: Model construction: Construct a face video heart rate detection model based on separable convolution, including a convolutional gated linear unit based on nearest neighbor features, a spatio-temporal differential multi-head attention mechanism based on temporal difference convolution, and a spatio-temporal feed-forward network based on separable convolution; The convolutional gated linear unit based on nearest neighbor features inputs face video data to obtain fine-grained features of adjacent frames; The spatio-temporal differential multi-head attention mechanism based on temporal difference convolution is used to refine the anti-interference ability of local spatio-temporal representations; The spatio-temporal feed-forward network based on separable convolution processes the output of the spatio-temporal differential multi-head attention mechanism to obtain heart rate data;
[0008] The convolutional gated linear unit based on nearest neighbor features consists of two linear projection parts for generating gating signals and linear transformations for transmitting information, and a separable convolution is added to the activation function of the part for generating gating signals. The calculation formula is:
[0009] t1 = L(f)
[0010] t = t1·R(D C (L(f)))
[0011] x = f + t
[0012] where f(f1,…,f n ) represents the input of the convolutional gated linear unit, t and t1 represent intermediate quantities in the operation process, L represents a linear operation, R represents an activation function, D C represents depthwise separable convolution, x(x1,…,x n ) represents the output matrix of the convolutional gated linear unit, and n represents the number of video frames;
[0013] The spatio-temporal differential multi-head attention mechanism based on temporal difference convolution combines convolution operations, temporal difference calculations, and multi-head attention mechanisms to guide global attention to refine local spatio-temporal features. Its calculation process is as follows: Assume that the input matrix of the spatio-temporal differential multi-head attention mechanism is x(x1,…,x n ), which is mapped into Q, K, and V matrices through temporal difference convolution. The Q matrix and the K matrix are multiplied to obtain an intermediate attention matrix S. Then, the matrix S and the matrix V are multiplied mathematically for cross-operation to obtain an output matrix M, which is expressed as:
[0014] Q = B(T C (x))
[0015] K = B(T C (x))
[0016] V = TC (x)
[0017]
[0018] Among them, Q, K, and V respectively represent the Q matrix, K matrix, and V matrix after linear mapping, B represents the batch normalization operation, and T C represents the temporal difference convolution, Softmax represents the maximum batch normalization operation, d represents the random packet loss operation, S represents the intermediate attention matrix after multiplying the Q matrix and the K matrix, and M represents the output matrix of the spatio-temporal difference multi-head attention mechanism;
[0019] The spatio-temporal feed-forward network based on separable convolution consists of two linear transformation layers, which are used to refine local consistency and suppress some noise features, so as to obtain relative position clues. The calculation process is as follows: Assume that the input of this spatio-temporal feed-forward network is the output matrix M of the spatio-temporal difference multi-head attention mechanism. First, the M matrix passes through a 3D convolution with a convolution kernel of 1, batch normalization, and an activation function to obtain the matrix F1. Then, after passing through a depthwise separable convolution with a convolution kernel of 1, batch normalization, and an activation function, the matrix S T is obtained. Finally, after passing through a 3D convolution with a convolution kernel of 3 and batch normalization, the output matrix Z of this spatio-temporal feed-forward network is obtained. The calculation formula is:
[0020] F1 = ELU(B(C(M)))
[0021] S T = ELU(B(D C (F1)))
[0022] Z = B(C3(S T ))
[0023] Among them, C represents the 3D convolution with a convolution kernel of 1, ELU represents the exponential linear unit, B represents the 3D batch normalization, C3 represents the 3D convolution with a convolution kernel of 1, and F1, S T represent the intermediate output matrices, and Z represents the output matrix;
[0024] Step 3: Train and test the model: Use the collected face video data to train and test the face video heart rate detection model based on separable convolution;
[0025] Step 4: Input the face video into the trained face video heart rate detection model based on separable convolution, output the heart rate data, and perform face video heart rate detection.
[0026] In the above-mentioned face video heart rate detection method based on separable convolution, where: in Step 2, the D C represents the depthwise separable convolution with a convolution kernel of 3×3.
[0027] The above-mentioned face video heart rate detection method based on separable convolution, wherein: in step three, the collected face video data is used to train and test the face video heart rate detection model based on separable convolution, and the loss function in the training process is:
[0028] The loss function loss1 is a loss function that calculates the mean of the sum of squares of the differences between the predicted values and the target values, and its calculation formula is:
[0029]
[0030] where y i represents the predicted value of the sample, y represents the target value of the current sample, and n represents the number of samples;
[0031] The mean absolute error loss function loss2 calculates the average error between the predicted value and the target value, and its calculation formula is:
[0032]
[0033] The cross-entropy loss function loss3 calculates the distance metric between the predicted value and the true label value, and its calculation formula is:
[0034]
[0035] where p i represents the probability predicted by the model;
[0036] The calculation formula of the total loss function loss is:
[0037] loss = loss1 + loss2 + loss3
[0038] Compared with the prior art, the present invention has obvious beneficial effects. As can be seen from the above solutions, the present invention is based on a convolutional gated linear unit with nearest neighbor features, which consists of two linear projection parts for generating gating signals and for linear transformation of information transmission, and a separable convolution is added to the activation function of the part for generating gating signals to enhance the local modeling ability and prompt the model to obtain nearest neighbor image features. The spatio-temporal differential multi-head attention mechanism based on temporal difference convolution combines convolution operations, temporal difference calculations, and multi-head attention mechanisms to guide global attention to refine local spatio-temporal features and enhance the long spatio-temporal perception and interaction of facial modeling. The spatio-temporal feed-forward network based on separable convolution consists of two linear transformation layers, which are used to refine local consistency and suppress some noise features, so as to obtain relative position clues. In short, the present invention can capture weak feature information of the face, aggregate features, and convert the features into PPG signal information for heart rate detection.
[0039] The beneficial effects of the present invention will be further described below through specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a flowchart of the present invention. SPECIFIC EMBODIMENTS
[0041] The specific embodiments, features and effects of a face video heart rate detection method based on separable convolution proposed according to the present invention will be described in detail below in conjunction with the accompanying drawings and preferred embodiments.
[0042] See Figure 1 , a face video heart rate detection method based on separable convolution of the present invention, wherein: the specific steps of the method include:
[0043] Step 1: Data collection: Collect face video data and perform preprocessing;
[0044] Step 2: Model construction: As Figure 1 shown, construct a face video heart rate detection model based on separable convolution, including a convolutional gated linear unit based on nearest neighbor features, a spatio-temporal differential multi-head attention mechanism based on temporal difference convolution, and a spatio-temporal feed-forward network based on separable convolution; the convolutional gated linear unit based on nearest neighbor features inputs face video data to obtain fine-grained features of adjacent frames; the spatio-temporal differential multi-head attention mechanism based on temporal difference convolution is used to refine the anti-interference ability of local spatio-temporal representation; the spatio-temporal feed-forward network based on separable convolution processes the output of the spatio-temporal differential multi-head attention mechanism to obtain heart rate data;
[0045] The convolutional gated linear unit based on nearest neighbor features consists of two linear projection parts for generating a gating signal and a linear transformation for transmitting information, and a separable convolution is added to the activation function of the part for generating the gating signal, and the calculation formula is:
[0046] t1 = L(f)
[0047] t = t1 · R(D C (L(f)))
[0048] x = f + t
[0049] where, f(f1,..., f n ) represents the input of the convolutional gated linear unit, t and t1 represent intermediate quantities in the operation process, L represents a linear operation, R represents an activation function, D C represents a depthwise separable convolution with a convolution kernel of 3×3, x(x1,..., x n ) represents the output matrix of the convolutional gated linear unit, and n represents the number of video frames;
[0050] The spatio-temporal differential multi-head attention mechanism based on temporal-difference convolution combines convolution operation, temporal-difference calculation, and multi-head attention mechanism to guide global attention to refine local spatio-temporal features. Its calculation process is as follows: Assume that the input matrix of the spatio-temporal differential multi-head attention mechanism is x(x1,…,x n ), which is mapped into Q, K, and V matrices through temporal-difference convolution. The Q matrix and the K matrix are multiplied, and the intermediate attention matrix S is obtained. Then, the matrix S and the matrix V are multiplied and cross-operated mathematically to obtain the output matrix M, which is expressed as:
[0051] Q = B(T C (x))
[0052] K = B(T C (x))
[0053] V = T C (x)
[0054]
[0055] where Q, K, and V represent the Q matrix, K matrix, and V matrix after linear mapping respectively, B represents batch normalization operation, T C represents temporal-difference convolution, Softmax represents maximum batch normalization operation, d represents random packet loss operation, S represents the intermediate attention matrix after multiplying the Q matrix and the K matrix, and M represents the output matrix of the spatio-temporal differential multi-head attention mechanism;
[0056] The spatio-temporal feed-forward network based on separable convolution consists of two linear transformation layers, which are used to refine local consistency and suppress some noise features, so as to obtain relative position clues. Its calculation process is as follows: Assume that the input of this spatio-temporal feed-forward network is the output matrix M of the spatio-temporal differential multi-head attention mechanism. First, the M matrix passes through a three-dimensional convolution with a convolution kernel of 1, batch normalization, and activation function, and then the matrix F1 is obtained. Then, after passing through a depthwise separable convolution with a convolution kernel of 1, batch normalization, and activation function, the matrix S T is obtained to enrich the carried feature information. Finally, after passing through a three-dimensional convolution with a convolution kernel of 3 and batch normalization, the output matrix Z of this spatio-temporal feed-forward network is obtained. The calculation formula is:
[0057] F1 = ELU(B(C(M)))
[0058] S T = ELU(B(D C (F1)))
[0059] Z = B(C3(S T ))
[0060] Among them, C represents a three-dimensional convolution with a convolution kernel of 1, ELU represents an exponential linear unit, B represents three-dimensional batch normalization, C3 represents a three-dimensional convolution with a convolution kernel of 1, F1, S T represents the intermediate output matrix, and Z represents the output matrix;
[0061] Step 3: Train and test the model: Use the collected face video data to train and test the face video heart rate detection model based on separable convolution. The loss function during the training process is:
[0062] The loss function loss1 is a loss function that calculates the mean of the sum of the squares of the differences between the predicted values and the target values. Its calculation formula is:
[0063]
[0064] Among them, y i represents the predicted value of the sample, y represents the target value of the current sample, and n represents the number of samples;
[0065] The mean absolute error loss function loss2 calculates the average error between the predicted value and the target value. Its calculation formula is:
[0066]
[0067] The cross-entropy loss function loss3 calculates the distance metric between the predicted value and the true label value. Its calculation formula is:
[0068]
[0069] Among them, p i represents the probability predicted by the model;
[0070] The calculation formula for the total loss function loss is:
[0071] loss = loss1 + loss2 + loss3
[0072] Step 4: Input the face video into the trained face video heart rate detection model based on separable convolution, output the heart rate data, and perform face video heart rate detection.
[0073] Performance analysis
[0074] To ensure the objectivity and effectiveness of the experimental results, the experiments of the present invention are all implemented based on the PyTorch framework and trained and tested on an NVIDIA GTX v100 GPU with 32GB of video memory. rPPG-Toolbox is an open-source platform that integrates heart rate detection algorithms such as PhysFormer, DeepPhys, and PhysNet. The proposed FVSC-HR algorithm model of the present invention is implemented on the rPPG-Toolbox platform. The AdamW optimizer with an initial learning rate of 1e-4 is used, and the momentum is set to 0.9. A fixed-step decay strategy with a step size of 15 is used to adjust the learning rate. The total loss function is used as the loss function for network training. The number of epochs is set to 50, the batch size is set to 8, and the MSE loss function is used as the regression loss.
[0075] The summary of the training and testing results of the model of the present invention and the comparison models on the UBFC-rPPG dataset is shown in Table 1. From the experimental data in Table 1, it can be seen that from the perspective of the MAE evaluation index, the MAE evaluation index of the model POS is 4.0, and the MAE evaluation index of the model LGI is 15.8. The MAE of the STSC model is 2.15, and the MAE of the EEMD+LA-SSA model is 1.37, which is the best value achieved by the models in the MAE evaluation index. From the perspective of the RMSE evaluation index, the RMSE of the model POS is 7.6, and the RMSE of the model LGI is 28.5. The RMSE of the STSC model is 3.82, and the RMSE of the EEMD+LA-SSA model is 3.61, which is the best value achieved by the models in the RMSE evaluation index. The RMSE of the proposed FVSC-HR model is 10.8, exceeding the LGI model. From the perspective of the MAPE evaluation index, the MAPE of the model POS is 3.9, and the MAPE of the model LGI is 14.6. The MAPE of the STSC model is 3.15, which is the best value achieved by the models in the MAPE evaluation index. The MAPE of the proposed FVSC-HR model is 3.8, ranking second among all comparison models. This shows that the model of the present invention has obvious advantages on the UBFC-rPPG dataset.
[0076] In addition, to reasonably evaluate the effectiveness of the proposed model, the model of the present invention and the comparative model were trained on the UBFC-rPPG dataset and tested on the UBFC-phys dataset. The results are summarized in Table 2. As can be seen from Table 2, from the perspective of the MAE evaluation index, the MAE of the DeepPhys model is 30.1, the MAE of PhysFormer is 13.5, the MAE of PhysNet is 21.1, the MAE of TS-CAN is 12.0, the MAE of TransPhys is 5.1, and the MAE of the proposed FVSC-HR model is 7.3, ranking second among the compared models. From the perspective of the RMSE evaluation index, the RMSE of the DeepPhys model is 30.1, the RMSE of PhysFormer is 15.9, the RMSE of PhysNet is 26.3, the RMSE of TS-CAN is 12.4, the RMSE of TransPhys is 10.2, and the RMSE of the proposed FVSC-HR model is 7.4, which is the lowest among all the compared models and the performance is the best. From the perspective of the MAPE evaluation index, the MAPE of the DeepPhys model is 26.8, the MAPE of PhysFormer is 12.5, the MAPE of PhysNet is 23.8, the MAPE of TS-CAN is 13.7, and the MAPE of the proposed FVSC-HR model is 7.2, which is the lowest among all the compared models and the performance is the best.
[0077] In summary, benefiting from the fine-grained feature ability of the convolutional gated linear unit and the ability of the temporal difference multi-head attention mechanism to obtain video frame context information, the model proposed in the present invention has achieved highly competitive performance compared with the representative algorithm models. The research results can be applied to fields such as intelligent health monitoring and telemedicine, providing a solution for remote heart rate estimation.
[0078] Table 1 Performance statistical results of different models on the UBFC-rPPG dataset
[0079]
[0080] Table 2 Performance statistical results of different models on the UBFC-phys dataset
[0081]
[0082] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A face video heart rate detection method based on separable convolution, characterized in that: The specific steps of this method include: Step 1: Data collection: Collect face video data and perform preprocessing; Step 2: Model construction: Construct a face video heart rate detection model based on separable convolution, including a convolutional gated linear unit based on nearest neighbor features, a spatio-temporal differential multi-head attention mechanism based on temporal difference convolution, and a spatio-temporal feed-forward network based on separable convolution; The convolutional gated linear unit based on nearest neighbor features inputs face video data to obtain fine-grained features of adjacent frames; The spatio-temporal differential multi-head attention mechanism based on temporal difference convolution is used to refine the anti-interference ability of local spatio-temporal representations; The spatio-temporal feed-forward network based on separable convolution processes the output of the spatio-temporal differential multi-head attention mechanism to obtain heart rate data; The convolutional gated linear unit based on nearest neighbor features consists of two linear projection parts for generating gating signals and for linear transformation of information transmission, and a separable convolution is added to the activation function of the part for generating gating signals. The calculation formula is: t1 = L(f) t = t1·R(D C (L(f)) x = f + t Among them, f(f1, …, f n ) represents the input of the convolutional gated linear unit, t and t1 represent intermediate quantities in the operation process, L represents a linear operation, R represents an activation function, D C represents depthwise separable convolution, x(x1, …, x n ) represents the output matrix of the convolutional gated linear unit, and n represents the number of video frames; The spatio-temporal differential multi-head attention mechanism based on temporal difference convolution combines convolution operation, temporal difference calculation, and multi-head attention mechanism to guide global attention to refine local spatio-temporal features. Its calculation process is as follows: Assume that the input matrix of the spatio-temporal differential multi-head attention mechanism is x(x1,…,x n ), which is mapped into Q, K, and V matrices through temporal difference convolution. The Q matrix and the K matrix are multiplied to obtain the intermediate attention matrix S. Then, the matrix S and the matrix V are multiplied and cross-operated mathematically to obtain the output matrix M, which is expressed as: Q = B(T C (x)) K = B(T C (x)) V=T C (x) Among them, Q, K, and V respectively represent the Q matrix, K matrix, and V matrix after linear mapping, B represents the batch normalization operation, T C represents the temporal difference convolution, Softmax represents the maximum batch normalization operation, d represents the random packet loss operation, S represents the intermediate attention matrix after multiplying the Q matrix and the K matrix, and M represents the output matrix of the spatio-temporal difference multi-head attention mechanism; The spatio-temporal feed-forward network based on separable convolution consists of two linear transformation layers, which are used to refine local consistency and suppress partial noise features, so as to obtain relative position clues. The calculation process is as follows: assuming that the input of the spatio-temporal feed-forward network is the output matrix M of the spatio-temporal difference multi-head attention mechanism. First, the M matrix passes through a 3D convolution with a convolution kernel of 1, batch normalization, and an activation function to obtain matrix F1. Then, after passing through a depthwise separable convolution with a convolution kernel of 1, batch normalization, and an activation function, matrix S is obtained. T , finally, after passing through a 3D convolution with a convolution kernel of 3 and batch normalization, the output matrix Z of the spatio-temporal feed-forward network is obtained. The calculation formula is as follows: F1 = ELU(B(C(M))) S T = ELU(B(D C (F1))) Z = B(C3(S T )) Among them, C represents a three-dimensional convolution with a convolution kernel of 1, ELU represents an exponential linear unit, B represents three-dimensional batch normalization, C3 represents a three-dimensional convolution with a convolution kernel of 1, F1, S T represents the intermediate output matrix, and Z represents the output matrix; Step 3: Train and test the model: Use the collected face video data to train and test the face video heart rate detection model based on separable convolution; Step 4: Input the face video into the trained face video heart rate detection model based on separable convolution, output the heart rate data, and perform face video heart rate detection.
2. The face video heart rate detection method based on separable convolution according to claim 1, wherein: In step two, the D C represents a depthwise separable convolution with a 3×3 convolution kernel.
3. A face video heart rate detection method based on separable convolution according to claim 1 or 2, characterized in that: In Step 3, use the collected face video data to train and test the face video heart rate detection model based on separable convolution. The loss function in the training process is: The loss function loss1 is a loss function that calculates the mean of the sum of squares of the differences between the predicted value and the target value. Its calculation formula is: Among them, y i represents the predicted value of the sample, y represents the target value of the current sample, and n represents the number of samples; The mean absolute error loss function loss2 calculates the average error between the predicted value and the target value. Its calculation formula is: The cross-entropy loss function loss3 calculates the distance metric between the predicted value and the true label value. Its calculation formula is: where p i represents the probability predicted by the model; The calculation formula for the total loss function loss is: loss = loss1 + loss2 + loss3.
Citation Information
Cited By
Indoor environment data prediction method and system based on LSTM-iTransform
CN120744481A
Non-contact atrial fibrillation recognition and multi-physiological index monitoring system and method based on deep learning
CN121400797A