A Heart Rate Detection Method Based on Joint Attention and Multi-Scale Fusion

By constructing a multi-scale video pyramid and joint attention mechanism, the problem of incomplete scale and time information in the existing heart rate detection methods is solved, and a higher-precision heart rate detection is achieved.

CN114694061BActive Publication Date: 2025-07-29ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210256849.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-16
Publication Date
2025-07-29
Estimated Expiration
2042-03-16

AI Technical Summary

Technical Problem

The existing contactless heart rate detection methods are limited by a single-scale region of interest, lacking the integrity of the scale and the continuity of time information, resulting in low detection accuracy.

Method used

A heart rate detection method based on joint attention and multi-scale fusion is adopted, and a multi-scale video pyramid is constructed to extract and fusion multi-scale feature, and an inter-layer attention mechanism and rPPG signal extraction network are used to extract more complete heart rate signals.

Benefits of technology

It effectively improves the accuracy of heart rate detection, overcomes the limitations of single-scale and attention modules, and achieves higher detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114694061B_ABST
    Figure CN114694061B_ABST
Patent Text Reader

Abstract

A heart rate detection method based on joint attention and multi-scale fusion constructs a multi-scale image sequence to ensure the integrity of the scale, then extracts signal features from multiple scales, performs multi-scale feature fusion through an inter-layer attention mechanism, and finally sends the fused features into the rPPG signal extraction network. The designed rPPG signal extraction network consists of channel-time joint attention (CTJA) and spatio-temporal joint attention (STJA) to extract signals in terms of time, space, and channels. The present invention can effectively improve the accuracy of heart rate detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of video image processing, computer vision, and signal processing. Specifically, it relates to a heart rate detection method based on face video. Background Art

[0002] Heart rate is one of the important physiological indicators of humans and is often used as an important parameter to measure the physiological or emotional state of the human body. Currently, the common heart rate detection methods are mostly contact-based and require skin contact with the subject. This makes the contact-based measurement method restricted by the use site, and for groups with inconvenient skin contact, the contact-based measurement method is not applicable. With the development of computational vision technology, video image processing technology has been applied to the field of heart rate detection, and non-contact video heart rate detection has shown obvious advantages. It can achieve real-time heart rate monitoring only with a camera, which is more convenient and fast. The limitations of existing non-contact heart rate measurement methods include: 1) restricted by the single-scale region of interest (ROI), the network mostly uses single-scale images as input, lacking scale integrity; 2) restricted by the attention module, the attention module focuses on feature extraction from the spatial perspective and has weak focus on the temporal perspective, lacking the continuity and integrity of temporal information. Summary of the Invention

[0003] To overcome the deficiencies of the prior art, in view of the limitations of the prior art, the present invention proposes a heart rate detection method based on joint attention and multi-scale fusion that can effectively improve the accuracy of heart rate detection.

[0004] To solve the above technical problems, the present invention adopts the following technical solutions:

[0005] A heart rate detection method based on joint attention and multi-scale fusion, the method comprising the following steps:

[0006] Step 1, constructing a multi-scale video pyramid

[0007] Obtaining image sequences of different scales by performing Gaussian filtering and downsampling operations on the cropped face video;

[0008] Step 2, multi-scale feature extraction

[0009] Inputting the obtained multi-scale video sequences into a feature extraction network respectively for feature extraction operations;

[0010] Step 3, multi-scale feature fusion

[0011] Performing feature fusion on the feature maps obtained at each scale through an inter-layer attention mechanism to obtain a fused feature map;

[0012] Step 4, rPPG signal extraction

[0013] Input the fused feature map into the rPPG signal extraction network for signal extraction operations to obtain the final rPPG signal.

[0014] Furthermore, in step 1, the Gaussian pyramid can be regarded as a low-pass filter to preliminarily filter out motion interference and quantization noise. Different-scale image sequences are obtained through Gaussian filtering and downsampling operations. Using multi-scale image analysis can make up for the shortcoming of incomplete information in single-scale image analysis and meet the requirement of extracting more complete information at different scales. For the rPPG pulse extraction task, in single-scale analysis, the pulse signal, noise, and motion artifacts are combined together, increasing the difficulty of signal separation. Project the combined observations into the scale space. Assume that at a certain scale, the pulse signal dominates, and at other scales, the motion artifacts dominate. The preliminary separation of the signal is achieved, reducing the difficulty of signal separation.

[0015] Even further, in step 2, a lightweight feature extraction network is designed to extract features at different scales. This network consists of 5 layers, namely 3 convolutional layers, 1 adaptive average pooling layer, and 1 spatio-temporal joint attention (STJA) layer. Using 2D STJA can enhance spatio-temporal feature extraction and extract as many features as possible at different scales.

[0016] Still further, in step 3, by averaging the feature maps of each layer, according to the differences in feature extraction at different levels of the pyramid, the layer-wise weights are obtained using the softmax activation function, and the multi-layer feature maps are fused according to the weight values. This method uses the correlation of layer-wise features to improve the effect of feature fusion, expressed as:

[0017] α = softmax(ODC(GAP(β)))

[0018]

[0019] where β represents the multi-scale feature maps generated in step 2, α represents a set of layer-wise connection weights generated using the softmax activation function, GAP represents global average pooling, ODC represents one-dimensional temporal convolution with a convolutional kernel size of [1, 1], n represents the number of layers of the image pyramid, and F represents the fused feature map.

[0020] In step 4, pulse signals are extracted from the fused feature maps obtained in step 3. Using the characteristics of rPPG signals, a joint attention mechanism is adopted to make full use of the characteristics of rPPG signals in the spatial, channel, and time dimensions. It includes channel-time joint attention (CTJA) and spatial-time joint attention (STJA). By combining CTJA and STJA, an rPPG extraction network is constructed to make the extraction of rPPG signals more sufficient and complete. The first part of the rPPG extraction network is a two-dimensional convolutional layer that processes the input feature maps in the spatial dimension. The second part is a stack composed of four sub-networks for joint-dimensional feature learning. The sub-network consists of a max-pooling layer, a 3D convolutional layer, a CTJA layer, a 3D convolutional layer, and an STJA layer. In the four sub-networks, the correlations in the spatial, time, and channel dimensions are fully utilized and extracted. Finally, a global average pooling layer and a one-dimensional convolutional layer are used to obtain an output with the same size as the pulse signal.

[0021] Preferably, the processes of CTJA and STJA are as follows:

[0022] For channel-time joint attention (CTJA), the features on the feature map are compressed into the channel and time dimensions through an average operation. Dilated convolutions with dilation rates of 1, 2, and 4 are used to obtain richer feature information, and then depthwise separable convolution is used to integrate the information obtained by the dilated convolution. Finally, the channel-time joint weight is obtained through a sigmoid activation function. After expanding the obtained weight to the same size as the input feature map, it is multiplied corresponding to the input feature map, which is expressed as:

[0023] A weight = σ(DSC(S))

[0024]

[0025] where S represents the concatenation of the feature information obtained by dilated convolutions with 3 different dilation rates, DSC represents depthwise separable convolution, including two parts: depth convolution (DW) and point convolution (PW), σ represents the sigmoid activation function; A weight represents the generated channel-time joint weight, represents the element-wise multiplication.;

[0026] For spatial-time joint attention (STJA), the input feature map is respectively passed through global average and channel average to obtain D t and D thw and D t and D thwPerform convolution and sigmoid activation operations on two branches respectively to learn the weights in time and space. Then multiply the two weight values and replicate them in the channel dimension to obtain a weight matrix of the same size as the input. Finally, apply the weight matrix to the input feature map. The process is as follows:

[0027] Learn the weights in time and space on two branches respectively:

[0028] T weight = σ(conv(D t ))

[0029] B weight = σ(conv(D thw ))

[0030] Where conv represents the convolution operation and σ represents the sigmoid activation function;

[0031] Obtain the spatio-temporal joint weight and apply it to the input feature:

[0032]

[0033]

[0034] Where REP represents replication in the channel dimension, represents element-wise multiplication;

[0035] For a given face video, segment the long video into specific lengths and input it into the proposed algorithm network to obtain the final pulse signal output.

[0036] The beneficial effects of the present invention are as follows: effectively improve the heart rate detection accuracy. Brief Description of the Drawings

[0037] Figure 1 is a flowchart of a heart rate detection method based on joint attention and multi-scale fusion.

[0038] Specific Implementation Process

[0039] The present invention will be further described in detail below.

[0040] Refer to Figure 1 , a heart rate detection method based on joint attention and multi-scale fusion, the method includes the following steps:

[0041] Step 1, construct a multi-scale video pyramid

[0042] The Gaussian pyramid can be regarded as a low-pass filter, which preliminarily filters out motion interference and quantization noise. An image sequence of different scales is obtained through Gaussian filtering and downsampling operations. Using multi-scale image analysis can make up for the shortcoming of incomplete information in single-scale image analysis and meet the requirement of extracting more complete information at different scales. For the rPPG pulse extraction task, in single-scale analysis, the pulse signal, noise, and motion artifacts are combined together, increasing the difficulty of signal separation. Projecting the combined observations into the scale space, assuming that at a certain scale, the pulse signal dominates, and at other scales, the motion artifacts dominate, the preliminary separation of the signal is achieved, reducing the difficulty of signal separation.

[0043] Step 2, Multi-scale feature extraction

[0044] A lightweight feature extraction network is designed to extract features of different scales. The network consists of 5 layers, namely 3 convolutional layers, 1 adaptive average pooling layer, and 1 spatio-temporal joint attention STJA layer. Using 2D STJA can enhance spatio-temporal feature extraction and extract as many features as possible at different scales.

[0045] Step 3, Multi-scale feature fusion

[0046] By averaging the feature maps of each layer, according to the differences in feature extraction at different levels of the pyramid, the sofmax activation function is used to obtain the inter-layer weights, and the multi-layer feature maps are fused according to the weight values. This method uses the correlation of inter-layer features to improve the effect of feature fusion, expressed as:

[0047] α = softmax(ODC(GAP(β)))

[0048]

[0049] Among them, β represents the multi-scale feature maps generated in Step 2, α represents a set of inter-layer connection weights generated using the softmax activation function, GAP represents global average pooling, ODC represents one-dimensional temporal convolution with a convolutional kernel size of [1, 1], n represents the number of layers of the image pyramid, and F represents the fused feature map;

[0050] Step 4, rPPG signal extraction

[0051] Extract the pulse signal from the fused feature map. Using the characteristics of the rPPG signal, a joint attention mechanism is adopted to fully utilize the characteristics of the rPPG signal in the spatial, channel, and time dimensions, including channel-time joint attention (CTJA) and spatial-time joint attention (STJA). Combine CTJA and STJA to construct an rPPG extraction network, making the extraction of rPPG signals more sufficient and complete. The first part of the rPPG extraction network is a two-dimensional convolutional layer that processes the input feature map in the spatial dimension. The second part is a stack composed of four sub-networks for joint-dimensional feature learning. The sub-network consists of a max-pooling layer, a 3D convolutional layer, a CTJA layer, a 3D convolutional layer, and an STJA layer. In the four sub-networks, the correlations in the spatial, time, and channel dimensions are fully utilized and extracted. Finally, a global average pooling layer and a one-dimensional convolutional layer are used to obtain an output with the same size as the pulse signal.

[0052] The processes of CTJA and STJA are as follows:

[0053] For channel-time joint attention (CTJA), the features on the feature map are compressed to the channel and time dimensions through an average operation. Dilated convolutions with dilation rates of 1, 2, and 4 are used to obtain richer feature information, and then depthwise separable convolution is used to integrate the information obtained by the dilated convolution. Finally, the channel-time joint weight is obtained through a sigmoid activation function. After expanding the obtained weight to the same size as the input feature map, it is multiplied by the corresponding input feature map, which is expressed as:

[0054] A weight = σ(DSC(S))

[0055]

[0056] where S represents the concatenation of the feature information obtained by dilated convolutions with 3 different dilation rates, DSC represents depthwise separable convolution, including two parts: depth convolution (DW) and point convolution (PW), and σ represents the sigmoid activation function; A weight represents the generated channel-time joint weight, represents the multiplication of elements;

[0057] For spatial-time joint attention (STJA), the input feature map is respectively passed through global average and channel average to obtain D t and D thw , and D t and D thw are respectively subjected to convolution and sigmoid activation operations on two branches to learn the weights in the time and space dimensions. Then, the two weight values are multiplied and then replicated in the channel dimension to obtain a weight matrix of the same size as the input. Finally, the weight matrix is applied to the input feature map. The process is as follows:

[0058] Learn the weights in time and space on two branches respectively:

[0059] T weight = σ(conv(D t ))

[0060] B weight = σ(conv(D thw ))

[0061] where conv represents the convolution operation and σ represents the sigmoid activation function;

[0062] Obtain the spatio-temporal joint weight and apply it to the input features:

[0063]

[0064]

[0065] where REP represents replication in the channel dimension, represents element-wise multiplication.

[0066] For a given face video, this embodiment segments the long video into specific lengths and inputs it into the proposed algorithm network to obtain the final pulse signal output.

Claims

1. A heart rate detection method based on joint attention and multi-scale fusion, characterized in that, The method includes the following steps: Step 1, construct a multi-scale video pyramid Obtain image sequences of different scales from the cropped face video through Gaussian filtering and downsampling operations; Step 2, multi-scale feature extraction Input the obtained multi-scale video sequences into the feature extraction network respectively for feature extraction operations; Step 3, multi-scale feature fusion Fuse the feature maps obtained at each scale through an inter-layer attention mechanism to obtain the fused feature maps; Step 4, rPPG signal extraction Input the fused feature maps into the rPPG signal extraction network for signal extraction operations to obtain the final rPPG signal; In step 4, extract the pulse signal from the fused feature maps obtained in step 3, and utilize the characteristics of the rPPG signal. The joint attention mechanism is adopted to make full use of the characteristics of the rPPG signal in the spatial, channel, and time dimensions; It includes channel-time joint attention CTJA and spatial-time joint attention STJA. Combine CTJA and STJA to construct the rPPG extraction network, making the rPPG signal extraction more sufficient and complete; The first part of the rPPG extraction network is a two-dimensional convolutional layer, which processes the input feature maps in the spatial dimension; The second part is a stack composed of four sub-networks for joint-dimensional feature learning. The sub-network consists of a max-pooling layer, a 3D convolutional layer, a CTJA layer, a 3D convolutional layer, and an STJA layer. In the four sub-networks, the correlations of the spatial, time, and channel dimensions are fully utilized and extracted; Finally, use the global average pooling layer and a one-dimensional convolutional layer to obtain an output with the same size as the pulse signal; The processes of CTJA and STJA are as follows: Channel-time joint attention CTJA compresses the features on the feature maps to the channel and time dimensions through an average operation, uses dilated convolutions with dilation rates of 1, 2, and 4 respectively to obtain richer feature information, then uses depthwise separable convolutions to integrate the information obtained by the dilated convolutions, and finally obtains the channel-time joint weights through the sigmoid activation function. After expanding the obtained weights to the same size as the input feature maps, multiply them corresponding to the input feature maps, expressed as: A weight = σ(DSC(S)) Among them, S represents the concatenation of feature information obtained by dilation convolutions with three different dilation rates, DSC represents depthwise separable convolution, including two parts: depthwise convolution (DW) and pointwise convolution (PW), σ represents the sigmoid activation function; A weight represents the generated channel-time joint weight, represents the element-wise multiplication; Spatial-temporal joint attention STJA obtains D by performing global average and channel average on the input feature map respectively t and D thw , and performs convolution and sigmoid activation operations on D t and D thw on two branches respectively to learn the weights in time and space. Then, the two weight values are multiplied and replicated in the channel dimension to obtain a weight matrix of the same size as the input. Finally, the weight matrix is applied to the input feature map. The process is as follows: Learn the weights in time and space on two branches respectively: T weight = σ(conv(D t )) B weight =σ(conv(D thw )) Among them, conv represents the convolution operation, and σ represents the sigmoid activation function; Obtain the spatial-time joint weights and apply them to the input features: Among them, REP represents replication in the re-channel dimension, represents the multiplication of elements; For a given face video, segment the long video into a specific length and input it into the proposed algorithm network to obtain the final pulse signal output.

2. The heart rate detection method based on joint attention and multi-scale fusion according to claim 1, wherein In the said step 1, the Gaussian pyramid can be regarded as a low-pass filter, which preliminarily filters out motion interference and quantization noise. An image sequence of different scales is obtained through Gaussian filtering and downsampling operations. The use of multi-scale image analysis can make up for the shortcoming of incomplete information in single-scale image analysis and meet the requirement of extracting more complete information at different scales. For the rPPG pulse extraction task, in single-scale analysis, the pulse signal, noise, and motion artifacts are combined together, increasing the difficulty of signal separation. By projecting the combined observations into the scale space, it is assumed that at a certain scale, the pulse signal dominates, and at other scales, the motion artifacts dominate. The preliminary separation of the signal is achieved, and the difficulty of signal separation is reduced.

3. A heart rate detection method based on joint attention and multi-scale fusion according to claim 1 or 2, characterized in that In the said step 2, a lightweight feature extraction network is designed to extract features of different scales. This network consists of 5 layers, namely 3 convolutional layers, 1 adaptive average pooling layer, and 1 spatio-temporal joint attention (STJA) layer. The use of 2D STJA can enhance spatio-temporal feature extraction and extract as many features as possible at different scales.

4. A heart rate detection method based on joint attention and multi-scale fusion according to claim 1 or 2, characterized in that, In the said step 3, by averaging the feature maps of each layer, according to the differences in feature extraction at different levels of the pyramid, the layer weights are obtained using the softmax activation function, and the multi-layer feature maps are fused according to the weight values. This method uses the correlation between inter-layer features to improve the effect of feature fusion, which is expressed as: α = softmax(ODC(GAP(β))) where β represents the multi-scale feature maps generated in step 2, α represents a set of inter-layer connection weights generated using the softmax activation function, GAP represents global average pooling, ODC represents one-dimensional temporal convolution with a convolutional kernel size of [1, 1], n represents the number of layers of the image pyramid, and F represents the fused feature map.

Citation Information

Patent Citations

  • Person Re-Identification Method Combining Random Batch Mask and Multi-Scale Representation Learning

    JP6830707B1

  • Detection method using fusion network based on attention mechanism, and terminal device

    US11222217B1