Non-contact heart rate detection method, system and device based on visual Transform and multi-scale feature aggregation and medium
Through the aggregation method of visual Transformer and multi-scale feature, the problem that signals are susceptible to environmental interference in contactless heart rate detection is solved, and high-precision and stable heart rate detection are achieved, which is suitable for remote monitoring in complex environments.
Patent Information
- Application Number
- CN202510279650.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-11
AI Technical Summary
In the existing non-contact heart rate detection technology, rPPG signals are susceptible to ambient light changes, facial movements and background noise, resulting in low signal accuracy. Deep learning methods have limitations in processing long-term series data and capturing periodic signals.
Using a method based on visual Transformer and multi-scale feature aggregation, the heart rate value is extracted and predicted through feature enhancement module, multi-scale mask feature aggregation module and Transformer timing modeling module, combined with sparse attention mechanism and loss function optimization.
It significantly improves the extraction accuracy and stability of rPPG signals, can perform high-precision heart rate detection in complex environments, is suitable for remote heart rate monitoring, and improves the robustness and generalization performance of detection.
Smart Images

Figure CN120298311A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of physiological signal detection and image processing, and particularly to a non-contact heart rate detection method, system, device and medium based on vision Transformer and multi-scale feature aggregation. Background Art
[0002] Heart rate, as a key physiological indicator to measure human health, plays an important role in many fields such as medical monitoring, affective computing, and anti-counterfeiting detection. Traditional detection means mainly rely on contact sensors, such as finger clip pulse oximeters or chest strap devices. However, these methods may cause discomfort or limit the user's freedom of movement when in use, and are not convenient for remote or continuous monitoring. With the progress of photoplethysmography (PPG) and remote photoplethysmography (rPPG) technologies, non-contact heart rate detection technology based on facial videos has gradually become the focus of research.
[0003] In order to accurately calculate the heart rate from the rPPG signal, appropriate algorithms and models are also needed for feature extraction and analysis. Currently, there are mainly two common methods. The first category is traditional signal processing methods, which extract color change signals from facial videos through algorithms such as filtering and Fourier transform to calculate the heart rate. However, this method is relatively sensitive to environmental conditions and is difficult to cope with complex lighting and motion interferences, thus affecting the accuracy of signal extraction. The second category of methods is deep learning technologies based on convolutional neural networks, which can automatically extract features to improve the detection accuracy. However, deep learning methods still have limitations in processing long-time series data and capturing periodic signals. In addition, the rPPG signal is easily affected by environmental light changes, facial movements, and background noise, which poses significant challenges to the accurate extraction and stability of the signal. Moreover, the existing models have limited capabilities in feature extraction and noise suppression, which may lead to misjudgments and omissions in the detection results. Summary of the Invention
[0004] In order to overcome the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a non-contact heart rate detection method, system, device and medium based on vision Transformer and multi-scale feature aggregation, which solves the problems that in the existing non-contact heart rate detection technology, the rPPG signal is easily affected by environmental light changes, facial movements, and background noise, resulting in low signal accuracy, and the deep learning method has limitations in processing long-time series data and capturing periodic signals.
[0005] The present invention is realized through the following technical solutions:
[0006] A non-contact heart rate detection method based on vision Transformer and multi-scale feature aggregation includes the following steps:
[0007] S1. Obtain the facial visible light video, unify the video frame rate through frame rate resampling, perform face key point localization and region division on each frame of image, and obtain the facial visible light video sequence;
[0008] S2. Use time offset operation on the facial visible light video sequence obtained in step S1 to obtain a differential frame sequence, perform channel fusion on the facial video sequence and the differential frames, and downsample the spatial dimension to obtain a low-resolution feature tensor;
[0009] S3. Construct a non-contact heart rate extraction model, input the low-resolution feature tensor into the non-contact heart rate extraction model, and output the heart rate value of the target; the non-contact heart rate extraction model includes a feature enhancement module, a multi-scale mask feature aggregation module, a Transformer temporal modeling module, and an rPPG predictor;
[0010] Among them, in step S3, input the low-resolution feature tensor into the feature enhancement module, extract local details and global context information through parallel dilated convolution and multi-branch convolution respectively to obtain an enhanced feature map; input the enhanced feature map into the sparse attention mechanism, and screen and weight the channels and pixel positions most relevant to rPPG within the spatio-temporal range through lightweight and grouped self-attention to obtain a sparse attention output feature map;
[0011] Input the sparse attention output feature map into the multi-scale mask feature aggregation module, divide it into multiple parallel branches, set different dilation rate convolutions in each branch to extract local and global skin textures and illumination changes at different scales; fuse the outputs of each branch through splicing and convolution to obtain a feature representation after multi-scale mask aggregation; segment and fuse the linear position encoding of the feature representation after multi-scale mask aggregation to obtain a time series sequence that meets the periodic time modeling requirements, and output a feature sequence fused with position encoding;
[0012] Input the feature sequence fused with position encoding into the Transformer temporal modeling module, perform global feature extraction through the multi-head self-attention mechanism, and obtain high-dimensional features with both spatio-temporal context and periodic information;
[0013] Input the high-dimensional features into the rPPG predictor, and output the heart rate value of the target.
[0014] Further, the feature enhancement module in step S3 consists of a linear fusion branch. After the multi-branch convolutions remap the input channels through 1×1 convolutions respectively, they combine 3×3 convolution kernels with different sizes and dilation rates to alternately extract local details and global context information in the horizontal and vertical directions. Subsequently, the outputs of the parallel atrous convolutions are concatenated in the channel dimension and mapped to the specified number of channels through 1×1 convolutions to obtain the fused features. After the low-resolution feature tensor output in step S2 and the fused features are weighted and added together, the enhanced feature map is output through the ReLU activation function.
[0015] Further, in step S3, the sparse attention output feature map is a high-order feature. First, the high-order feature is dimensionally reduced through a convolution operation and then interpolated to the same spatial size as the low-order feature. Then, the two and the additional spatial mask are divided into blocks by channel and fed into four parallel branches respectively. Each branch contains normalization and convolutions with different dilation rates. The outputs of each branch are fused with spatial weight and spatial mask guidance, giving higher weights to the effective skin regions, while suppressing the interference regions that may affect the rPPG signal extraction and the shadow regions caused by pose changes, to obtain the feature representation after multi-scale mask aggregation.
[0016] Further, in step S3, the Transformer temporal modeling module first reduces the resolution in the time dimension through a temporal downsampling layer to extract long-term dependencies. Subsequently, it inputs multiple self-attention layers, uses the grouped self-attention mechanism to screen the most relevant time windows, strengthens the periodic features, and enhances the model's ability to identify heart rate signals. Finally, the resolution in the time dimension is restored through a temporal upsampling layer to ensure the complete transmission of the feature temporal information.
[0017] Further, the grouped self-attention mechanism adopts a local time block processing method, divides fixed-length blocks in the time dimension, performs local self-attention calculations within each time block, and adopts the Top-K screening strategy to only retain the key frames that are the most relevant within the local temporal blocks.
[0018] Further, in step S3, the rPPG predictor consists of multiple three-dimensional convolutional layers and a linear mapping layer. First, the three-dimensional convolutional layers are used for feature extraction, the low-dimensional temporal features are restored through upsampling operations to maintain the temporal integrity of the heart rate information, and the linear mapping layer is used for feature decoding and conversion to map the extracted high-dimensional features to a preliminary rPPG waveform. The fast Fourier transform and power spectral density analysis are performed on the rPPG waveform, the main peak frequency is calculated, and the heart rate value of the target is obtained by inferring the heart beat period based on the located main peak frequency.
[0019] Further, the loss function of the contactless heart rate extraction model is:
[0020] Based on the predicted rPPG signal and the pre - acquired reference PPG signal, calculate the negative Pearson correlation coefficient between the predicted rPPG signal and the reference PPG signal, measure the time - series correlation of the predicted rPPG signal, and optimize the fitting ability of the predicted rPPG signal to the reference PPG signal to construct the time - domain loss function \(L\). time :
[0021]
[0022] where \(T\) is the signal length, \(s(t)\), are the reference PPG signal and the predicted rPPG signal respectively;
[0023] In the frequency domain, by calculating the power spectral density of the predicted rPPG signal and the reference PPG signal, and using the cross - entropy loss function to measure the frequency matching degree between the two, to ensure that the main frequency components of the model output signal are consistent with the true heart rate:
[0024]
[0025] where \(maxIdx\) represents the peak frequency index of obtaining the power spectral density, \(PSD(s)\) represents the power spectral density of the true signal \(s\), represents the predicted signal 's power spectral density;
[0026] Adopt the optimal time - shifted mean - square error loss, allowing for optimal time alignment within a given time window \([-\delta\) t , \(\delta\) t to minimize the error between the predicted rPPG signal and the reference PPG signal:
[0027]
[0028] where \(s(t)\) and are the reference PPG signal and the predicted rPPG signal respectively, \(\tau\) is the time offset, ensuring that the model can adapt to small time offsets and improve the alignment accuracy;
[0029] Combine the time - domain loss, frequency - domain loss and optimal time - shifted mean - square error loss to construct the final total loss function:
[0030] Loss all =\(\lambda_1\cdot L\) time +\(\lambda_2\cdot L\) freq +\(\lambda_3\cdot L\) MSE
[0031] where \(\lambda_1\), \(\lambda_2\), \(\lambda_3\) are hyper - parameter weights used to balance the contributions of each loss term in the optimization process.
[0032] A non-contact heart rate detection system based on Vision Transformer and multi-scale feature aggregation, comprising:
[0033] Facial visible light video acquisition module: acquires facial visible light video, unifies the video frame rate through frame rate resampling, performs face key point localization and region division on each frame of image, and obtains a facial visible light video sequence;
[0034] Low-resolution feature tensor acquisition module: uses time offset operation according to the facial visible light video sequence to obtain a differential frame sequence, performs channel fusion on the facial video sequence and the differential frames, and downsamples the spatial dimension to obtain a low-resolution feature tensor;
[0035] Non-contact heart rate extraction model construction module: constructs a non-contact heart rate extraction model, inputs the low-resolution feature tensor into the non-contact heart rate extraction model, and outputs the heart rate value of the target; the non-contact heart rate extraction model includes a feature enhancement module, a multi-scale mask feature aggregation module, a Transformer temporal modeling module, and an rPPG predictor;
[0036] Among them, the low-resolution feature tensor is input into the feature enhancement module, and local details and global context information are respectively extracted through parallel dilated convolution and multi-branch convolution to obtain an enhanced feature map; the enhanced feature map is input into the sparse attention mechanism, and the channels and pixel positions most relevant to rPPG are screened and weighted within the spatio-temporal range through lightweight and grouped self-attention to obtain a sparse attention output feature map;
[0037] The sparse attention output feature map is input into the multi-scale mask feature aggregation module, divided into multiple parallel branches, different dilation rate convolutions are set in each branch to extract local and global skin textures and illumination changes at different scales; the outputs of each branch are fused through splicing and convolution to obtain a multi-scale mask aggregated feature representation; the multi-scale mask aggregated feature representation is segmented and fused with linear position encoding to obtain a time series that meets the periodic time modeling requirements, and a feature sequence fused with position encoding is output;
[0038] The feature sequence fused with position encoding is input into the Transformer temporal modeling module, and global feature extraction is performed through the multi-head self-attention mechanism to obtain high-dimensional features with both spatio-temporal context and periodic information;
[0039] The high-dimensional features are input into the rPPG predictor, and the heart rate value of the target is output.
[0040] A non-contact heart rate detection device based on Vision Transformer and multi-scale feature aggregation, comprising:
[0041] Memory: Store a computer program of a non-contact heart rate detection method based on Vision Transformer and multi-scale feature aggregation, which is a computer-readable device;
[0042] Processor: When executing the computer program, it implements a non-contact heart rate detection method based on Vision Transformer and multi-scale feature aggregation.
[0043] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement a non-contact heart rate detection method based on Vision Transformer and multi-scale feature aggregation.
[0044] Compared with the prior art, the beneficial effects of the present invention are:
[0045] 1. Through the feature enhancement module, multi-scale mask feature aggregation module and Transformer temporal modeling module, the present invention significantly improves the accuracy and efficiency of rPPG signal extraction; the feature enhancement module forms multi-scale feature extraction through multiple branches, expands the receptive field and enhances weak signals; the multi-scale mask feature aggregation module guides the model to focus on the effective area through spatial masking, improves signal accuracy and efficiency, avoids noise interference, and improves signal stability; the Vision Transformer backbone network combines periodic time hierarchical encoding to efficiently extract periodic features in the rPPG signal. Finally, the rPPG predictor accurately predicts the signal, which can effectively improve the accuracy, robustness and generalization performance of heart rate detection, and is especially suitable for remote heart rate detection in complex environments. At the same time, the non-contact heart rate detection method based on the Transformer temporal modeling module helps to improve the stability and detection efficiency of signal extraction by capturing long-range spatio-temporal dependence relationships and periodic features, so as to meet the actual needs of high-precision and physiological monitoring.
[0046] 2. The present invention uses the pixel-level light intensity change in the facial video signal to infer the heart rate. First, the input visible light video is frame rate standardized, and face key point detection is performed. Based on the key point coordinates, the facial area is divided, and the skin area with high confidence is extracted to reduce environmental noise interference; then the time offset difference method is used to calculate the inter-frame light intensity change, generate a difference image sequence, and fuse the original frames to form a feature tensor; finally, the multi-scale feature extraction module performs hierarchical encoding on local and global information to further enhance the signal expression ability related to the heart rate.
[0047] 3. In the feature screening stage of the present invention, a sparse attention mechanism is adopted. Through an adaptive channel weighting strategy, the region most relevant to the rPPG signal is retained, and irrelevant background interference is suppressed. The screened features are input into a periodic time Transformer, and its self-attention mechanism is used to model the temporal dependence of the signal to enhance the stability and phase consistency of the rPPG signal. Finally, through a time-frequency domain joint optimization method, with the support of fast Fourier transform and power spectrum analysis, the strongest signal component is extracted and the heart rate value is calculated to achieve high-precision non-contact heart rate measurement.
[0048] 4. Through the joint optimization of the loss function, the present invention can synchronously constrain the predicted signal in the time domain and frequency domain, ensure the temporal stability, spectral consistency of the rPPG signal and the accuracy of heart rate estimation, and improve the robustness and generalization ability of the model. Description of the Drawings
[0049] Figure 1 It is a flowchart of the method provided by an embodiment of the present invention.
[0050] Figure 2 It is a schematic diagram of the network structure of the non-contact heart rate extraction model provided by an embodiment of the present invention.
[0051] Figure 3 It is a schematic diagram of the network structure of the feature enhancement module provided by an embodiment of the present invention.
[0052] Figure 4 It is a schematic diagram of the visualization interface of the heart rate detection result provided by an embodiment of the present invention. Detailed Embodiments
[0053] The following further describes the present invention in detail with specific embodiments, which are explanations rather than limitations of the present invention.
[0054] As Figure 1 shown, a non-contact heart rate detection method based on vision Transformer and multi-scale feature aggregation, a non-contact heart rate extraction model based on Transformer, and the non-contact heart rate extraction model includes a feature enhancement module, a multi-scale mask feature aggregation module, a Transformer temporal modeling module, and an rPPG predictor. The method includes:
[0055] First, the video captured by the original camera is obtained. Due to the reasons of the acquisition device, the frame rates of different videos in the same dataset may also be different. In order to eliminate the influence of uneven rPPG data frame rates on the model, the present invention unifies the video frame rate to 30fps through frame rate resampling. When the input video frame rate is 25fps or other non-fixed frame rates, frame interpolation is used to make the frame rate match the target frame rate, ensuring the stability and consistency of cross-dataset learning and improving the generalization ability of the model. The original video frame sequence is: F = {f1, f2,..., f N}, the original video frame rate is fps orig , and the target frame rate is fps target = 30. After the adjusted video frame sequence is F' = {f'1, f'2,..., f' M}. After resampling the video frame rate to 30fps, face key point localization and region division are then performed on each frame image.
[0056] Using face detection and key point detection algorithms, the coordinates of 108 facial key points are obtained, and the entire face area is divided into 6 main regions according to the key point distribution, including the forehead region, the left and right cheek regions, the nose wing region, the mandible region, and the temporal region. Sub-regions with obvious dynamic interference or low signal-to-noise ratio are excluded, and only regions with high confidence are retained for subsequent rPPG signal extraction.
[0057] After determining the key facial regions, the present invention further calculates the differential frames using time offset operations to refine the weak color changes caused by human blood pulsation. Specifically, for adjacent frames {X t , X t+1} or more distant frames {X t , X t+k}, the differential frames D i = X t+i - X t are respectively obtained, thereby obtaining the implicit heart rate dynamic information. After channel fusion of the original frame and the corresponding differential frame, downsampling is performed on the spatial dimension, and a set of feature tensors with lower resolution but rich in dynamic information is output.
[0058] After the low-resolution feature tensor is input into the feature enhancement module, the model will capture local and global skin optical features through a multi-branch convolution structure. Specifically, each branch first uses standard convolution for channel adjustment, and then combines standard convolution and dilated convolution at different levels to extract rich and complementary feature information under different receptive fields. Through this parallel multi-branch design, on the one hand, the attention to weak rPPG signals can be amplified, enhancing their saliency and distinguishability; on the other hand, noise interference such as ambient light changes or local movements can also be suppressed. After the above enhancement, the present invention inputs the enhanced feature map into the sparse attention mechanism, and filters and weights the feature channels or pixel positions most relevant to rPPG in the spatio-temporal range through lightweight or grouped self-attention, so as to effectively suppress irrelevant background interference and further improve the attention to blood pulsation signals and modeling accuracy.
[0059] After completing the sparse attention processing, the present invention inputs the sparse attention output feature map into the multi-scale mask feature aggregation module. This module first generates a spatial mask according to the effective area of the human face skin, and divides the features into several parallel branches according to the channel or spatial dimension; each branch uses convolution kernels with different dilation rates to capture texture information at different scales such as local details and global light changes. Subsequently, spatial weight or mask-guided fusion is performed on the outputs of each branch, assigning higher weights to the effective skin areas, while suppressing interference areas that may affect the extraction of rPPG signals, including eyes, mouth, eyebrows, beard, highlight areas (such as forehead reflections), and shadow areas caused by pose changes, so as to obtain the feature representation after multi-scale mask aggregation, further enhancing the attention to rPPG signals and improving robustness.
[0060] The feature representation after multi-scale mask aggregation is segmented and linearly position-encoded, and converted into a time series sequence that meets the requirements of periodic time modeling; the feature sequence integrated with position encoding is input into the Transformer time series modeling module, and high-dimensional feature vectors with spatio-temporal context and periodic information are obtained under the global and fine-grained modeling of multi-head self-attention;
[0061] The above high-dimensional feature vectors are input into the upsampling or decoding structure of the rPPG predictor, the features are resized and linearly transformed, and a preliminary rPPG waveform is output; post-processing is performed on the obtained rPPG waveform, including calculating the main peak frequency based on fast Fourier transform (FFT) or power spectrum analysis, and finally obtaining the target heart rate value.
[0062] Preferably, as Figure 3As shown, the feature enhancement module consists of multiple parallel branches and a linear fusion branch, which is used to enhance the input features in different receptive fields and channel mapping dimensions: First, each parallel branch remaps the input channels through 1×1 convolution respectively, and then combines 3×3 convolution kernels with different sizes or dilation rates to alternately extract local and global semantic information in the horizontal and vertical directions; Subsequently, the outputs of the parallel branches are concatenated in the channel dimension and mapped to the specified number of channels through 1×1 convolution; At the same time, there is a path for the input features, which is weighted and added to the fused features and then output the final enhanced features through the ReLU activation function.
[0063] The core idea of the sparse attention mechanism described in the present invention is that after calculating the pre-attention scores, only the spatio-temporal regions most relevant to the rPPG signal are retained, and then the Top-K strategy is used to screen out several positions with the highest similarity to perform fine-grained attention operations, so as to significantly reduce the computational amount of irrelevant backgrounds while ensuring the full capture of blood pulsation information. The sparse attention mechanism first performs a coarse-grained matching on the query matrix Q r and the key matrix K r to calculate the pre-attention score matrix A r , then selects the K column indices with the highest scores for each row of the matrix A r to form the index set TopK(i), and calculates the attention weight α i,j within this limited range. Finally, the corresponding value matrix V j is aggregated to obtain the output Z i , and its calculation formula is as follows:
[0064]
[0065] Among them, Q r is the query matrix, representing the query vector of the input feature map; K r is the key matrix, representing the key vector of the input feature map; V j is the value matrix (Value Matrix), storing the value information of the feature map; A r is the pre-attention score matrix, calculating the correlation between the query vector and the key vector; TopK(i) represents the K key indices with the highest correlation with the query vector Q[i]; α i,j is the attention weight, measuring the importance of the query vector Q[i] to the key vector K[j]; β is the scaling factor, adjusting the attention distribution; Z i is the final aggregated output vector, serving as the input for subsequent feature processing.
[0066] Compared with the traditional global attention, the implementation of the present invention can effectively suppress background noise and save a large amount of computational overhead, improving the ability to capture subtle physiological signals in videos and the inference efficiency of the model.
[0067] The described multi-scale mask feature aggregation module realizes the fusion and mask filtering of multi-scale skin texture and illumination changes by registering and grouping and aggregating high-order and low-order features, highlighting the effective regions related to the rPPG signal. Specifically, first, the input high-order features are dimension-reduced through a convolution operation and then interpolated to the same spatial size as the low-order features; then, the two and the additional mask are divided into blocks by channel and fed into four parallel branches respectively. Each branch contains normalization and convolutions with different dilation rates to obtain local details and global context; finally, the outputs of each branch are fused through concatenation and convolution to form the final mask aggregation feature. With the above multi-branch design and dilated convolution, the model can be effectively guided to focus on high-confidence skin regions and suppress irrelevant background noise, thereby enhancing the extraction and robustness of the rPPG signal at different scales.
[0068] The Transformer temporal modeling module proposed in the present invention is used to capture the time dependence and quasi-periodic characteristics of the rPPG signal. Specifically, taking the feature map output by the multi-scale mask feature aggregation module as the input, first, the resolution of the time dimension is reduced through a temporal downsampling layer to extract long-term dependencies; subsequently, it is input into multiple self-attention layers, and the grouped self-attention is used to screen the most relevant time windows to strengthen the periodic characteristics and enhance the model's ability to identify the heart rate signal; then, the resolution of the time dimension is restored through a temporal upsampling layer to ensure the complete transmission of the feature temporal information. In the self-attention mechanism, the present invention adopts a local time block processing method, that is, a fixed-length block is divided in the time dimension, and local self-attention calculation is performed within each time block. To further optimize the calculation efficiency, the model adopts a Top-K screening strategy, only retaining the most relevant key frames within the local temporal block, thereby reducing the global computational complexity and enhancing the focusing ability on key rPPG features.
[0069] Finally, the high-dimensional features after Transformer temporal modeling are input into the lightweight rPPG predictor, which is composed of multiple three-dimensional convolutions and linear mapping layers, and is responsible for mapping the high-dimensional time features output by the Transformer to the rPPG waveform.
[0070] Preferably, the process of defining the loss function of the non-contact heart rate extraction model includes the following steps:
[0071] Based on the predicted rPPG signal and the pre-collected reference PPG signal, calculate the negative Pearson correlation coefficient between the predicted rPPG signal and the reference PPG signal, measure the time series correlation of the predicted rPPG signal, and optimize the fitting ability of the predicted rPPG signal to the reference PPG signal to construct the time-domain loss function L time :
[0072]
[0073] Among them, T is the signal length, s(t), are the reference PPG signal and the predicted rPPG signal respectively;
[0074] In the frequency domain, by calculating the power spectral density of the predicted rPPG signal and the reference PPG signal, and using the cross-entropy loss function to measure the frequency matching degree between the two, to ensure that the main frequency components of the model output signal are consistent with the true heart rate:
[0075]
[0076] Among them, maxIdx represents the peak frequency index of obtaining the power spectral density, PSD(s) represents the power spectral density of the true signal s, represents the predicted signal of the power spectral density;
[0077] Adopt the optimal time-shifted mean square error loss, allowing optimal time alignment within a given time window [-δ t , δ t to minimize the error between the predicted rPPG signal and the reference PPG signal:
[0078]
[0079] Among them, s(t) and are the reference PPG signal and the predicted rPPG signal respectively, τ is the time offset, ensuring that the model can adapt to small time offsets and improve the alignment accuracy;
[0080] Combine the time-domain loss, frequency-domain loss and optimal time-shifted mean square error loss to construct the final total loss function:
[0081] Loss all = λ1·L time + λ2·L freq + λ3·L MSE
[0082] Among them, λ1, λ2, λ3 are hyperparameter weights, used to balance the contributions of each loss term in the optimization process.
[0083] Through the joint optimization of the above loss functions, the present invention can synchronously constrain the predicted signal in the time domain and frequency domain, ensure the timing stability, spectral consistency of the rPPG signal and the accuracy of heart rate estimation, and improve the robustness and generalization ability of the model.
[0084] Such as Figure 2As shown, the non-contact heart rate extraction model includes a feature enhancement module, a multi-scale mask feature aggregation module, a Transformer time series modeling module, and an rPPG predictor, specifically:
[0085] The feature enhancement module is used to receive the channel attention spatio-temporal block and extract skin optical features with different receptive fields through a multi-branch parallel convolution structure. In this module, first, a standard convolution operation is performed on the input pixel-level skin area, and then dilated convolutions with different dilation rates are combined to enhance long-distance dependence features.
[0086] The sparse attention mechanism module is used to receive the high-dimensional spatio-temporal feature map output by the feature enhancement module and filter the most relevant regions or channels in the time dimension and space dimension. In this module, first, the attention score of the input data is calculated through a local area weighting mechanism, and then the Top-K filtering strategy is adopted to retain only the pixel regions with the highest correlation with the rPPG signal to reduce the interference of the background region and improve the spatial contrast of the signal. In addition, this module also adopts a channel adaptive weighting method to enable the model to dynamically adjust the attention to different facial regions.
[0087] The multi-scale mask feature aggregation module is used to perform branch processing on the output features of the sparse attention mechanism module to extract skin texture and illumination change information at different scales. This module uses four groups of parallel convolutions, and each branch corresponds to a different dilation rate to adapt to the physiological signal changes at different resolutions. At the same time, aiming at the uneven illumination and skin reflection characteristics of the facial region, a spatial weight mask is adopted to adaptively adjust the attention distribution according to the pixel-level signal-to-noise ratio, and only the high-confidence regions are retained.
[0088] The Transformer time series modeling module is used to receive the time series feature sequence output by the multi-scale mask feature aggregation module and capture the quasi-periodic pattern and phase information of the rPPG signal using the self-attention mechanism within different time windows. In addition, this module adopts a hierarchical time block division method to perform feature extraction at different time scales to enhance the adaptability to long and short period changes. The lightweight rPPG predictor module is used to receive the time series features output by the periodic time Transformer module and construct an efficient signal regression model through one-dimensional convolution and linear transformation.
[0089] The spectrum analysis and heart rate calculation module is used to receive the time series signal output by the rPPG predictor and extract the main heart rate components through fast Fourier transform and power spectrum analysis. First, the predicted signal is band-pass filtered to remove low-frequency drift and high-frequency noise; then the power spectral density curve is calculated, and the maximum peak detection method is used to lock the main frequency component, and finally it is converted to the number of heartbeats per minute to output the final heart rate value.
[0090] AsFigure 4 As shown, it presents the interface of a non-contact heart rate detection system implemented based on the present invention. The system can extract rPPG signals from facial videos in a non-contact and natural state and predict the heart rate in real time. In the figure, the face of the subject is automatically detected and marked with a red bounding box to accurately lock the region of interest (ROI), ensuring the stability and accuracy of signal extraction. The system captures the video stream through a computer camera, then performs face detection and ROI selection, extracts weak blood flow change signals from areas such as the forehead and cheeks. These signals are then subjected to feature extraction and input into a Transformer model based on deep learning to learn the blood flow change patterns in the time series and predict the rPPG waveform. Finally, the system uses fast Fourier transform and power spectral density analysis to calculate the main heart rate components and displays the predicted heart rate value in real time in the upper left corner of the interface.
[0091] A non-contact heart rate detection system based on vision Transformer and multi-scale feature aggregation, comprising:
[0092] Facial visible light video acquisition module: Acquire facial visible light videos, unify the video frame rate through frame rate resampling, perform face key point localization and region division on each frame image to obtain a facial visible light video sequence;
[0093] Low-resolution feature tensor acquisition module: Use time offset operation according to the facial visible light video sequence to obtain a differential frame sequence, perform channel fusion on the facial video sequence and the differential frames, and downsample the spatial dimension to obtain a low-resolution feature tensor;
[0094] Non-contact heart rate extraction model construction module: Construct a non-contact heart rate extraction model, input the low-resolution feature tensor into the non-contact heart rate extraction model, and output the target heart rate value; the non-contact heart rate extraction model includes a feature enhancement module, a multi-scale mask feature aggregation module, a Transformer time series modeling module, and an rPPG predictor;
[0095] Among them, input the low-resolution feature tensor into the feature enhancement module, extract local details and global context information respectively through parallel dilated convolution and multi-branch convolution to obtain an enhanced feature map; input the enhanced feature map into the sparse attention mechanism, and screen and weight the channels and pixel positions most relevant to rPPG within the spatio-temporal range through lightweight and grouped self-attention to obtain a sparse attention output feature map;
[0096] Input the sparse attention output feature map into the multi-scale mask feature aggregation module, which is divided into multiple parallel branches. Set dilated convolutions with different dilation rates in each branch to extract local and global skin textures and illumination changes at different scales. Fuse the outputs of each branch through concatenation and convolution to obtain the feature representation after multi-scale mask aggregation. Segment and fuse the linear positional encoding of the feature representation after multi-scale mask aggregation to obtain a time series that meets the requirements of periodic time modeling, and output the feature sequence fused with positional encoding.
[0097] Input the feature sequence fused with positional encoding into the Transformer time series modeling module, and perform global feature extraction through the multi-head self-attention mechanism to obtain high-dimensional features with both spatio-temporal context and periodic information.
[0098] Input the high-dimensional features into the rPPG predictor to output the target heart rate value.
[0099] A non-contact heart rate detection device based on vision Transformer and multi-scale feature aggregation, comprising:
[0100] A memory: storing a computer program for a non-contact heart rate detection method based on vision Transformer and multi-scale feature aggregation, which is a computer-readable device;
[0101] A processor: used to implement a non-contact heart rate detection method based on vision Transformer and multi-scale feature aggregation when executing the computer program.
[0102] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement a non-contact heart rate detection method based on vision Transformer and multi-scale feature aggregation.
Claims
1. A non-contact heart rate detection method based on Vision Transformer and multi-scale feature aggregation, characterized in that, It includes the following steps: S1. Obtain a facial visible light video, unify the video frame rate through frame rate resampling, perform face key point localization and region division on each frame of the image, and obtain a facial visible light video sequence; S2. Use time offset operation on the facial visible light video sequence obtained in step S1 to obtain a differential frame sequence, perform channel fusion on the facial video sequence and the differential frames, and downsample the spatial dimension to obtain a low-resolution feature tensor; S3. Build a non-contact heart rate extraction model, input the low-resolution feature tensor into the non-contact heart rate extraction model, and output the heart rate value of the target; the non-contact heart rate extraction model includes a feature enhancement module, a multi-scale mask feature aggregation module, a Transformer temporal modeling module, and an rPPG predictor; Among them, in step S3, the low-resolution feature tensor is input into the feature enhancement module, and local details and global context information are respectively extracted through parallel dilated convolution and multi-branch convolution to obtain an enhanced feature map; the enhanced feature map is input into the sparse attention mechanism, and the channels and pixel positions most relevant to rPPG are screened and weighted within the spatio-temporal range through lightweight and grouped self-attention to obtain a sparse attention output feature map; The sparse attention output feature map is input into the multi-scale mask feature aggregation module, which is divided into multiple parallel branches. Different dilation rate convolutions are set in each branch to extract local and global skin textures and illumination changes at different scales; the outputs of each branch are fused through splicing and convolution to obtain a multi-scale mask aggregated feature representation; the multi-scale mask aggregated feature representation is segmented and fused with linear position encoding to obtain a temporal sequence that meets the requirements of periodic time modeling, and a feature sequence with fused position encoding is output; The feature sequence with fused position encoding is input into the Transformer temporal modeling module, and global feature extraction is performed through the multi-head self-attention mechanism to obtain high-dimensional features with both spatio-temporal context and periodic information; The high-dimensional features are input into the rPPG predictor, and the heart rate value of the target is output.
2. The non-contact heart rate detection method based on vision Transformer and multi-scale feature aggregation according to claim 1, characterized in that In step S3, the feature enhancement module consists of a linear fusion branch. After the multi-branch convolution remaps the input channels through 1×1 convolution respectively, combined with 3×3 convolution kernels of different sizes and dilation rates, local details and global context information are alternately extracted in the horizontal and vertical directions; Subsequently, the outputs of the parallel dilated convolutions are concatenated in the channel dimension and mapped to the specified number of channels through 1×1 convolution to obtain a fused feature; The low-resolution feature tensor output in step S2 and the fused feature are weighted and added, and then the enhanced feature map is output through the ReLU activation function.
3. A non-contact heart rate detection method based on Vision Transformer and multi-scale feature aggregation according to claim 1, characterized in that, In step S3, the sparse attention output feature map is a high-order feature. First, the high-order feature is dimensionally reduced through a convolution operation and then interpolated to the same spatial size as the low-order feature. Then, the two and the additional spatial mask are divided into blocks by channel and fed into four parallel branches respectively. Each branch contains normalization and convolutions with different dilation rates. The outputs of each branch are fused by spatial weight and spatial mask guidance, with higher weights assigned to the effective skin regions, while suppressing the interference regions that may affect the extraction of the rPPG signal and the shadow regions caused by pose changes, resulting in a feature representation after multi-scale mask aggregation.
4. A non-contact heart rate detection method based on Vision Transformer and multi-scale feature aggregation according to claim 1, characterized in that, In step S3, the Transformer temporal modeling module first reduces the resolution in the time dimension through a temporal downsampling layer to extract long-term dependencies. Subsequently, it is input into multiple self-attention layers. The grouped self-attention mechanism is used to screen the most relevant time windows, strengthen the periodic features, and enhance the model's ability to identify the heart rate signal. Finally, the resolution in the time dimension is restored through a temporal upsampling layer to ensure the complete transmission of the feature temporal information.
5. A non-contact heart rate detection method based on vision Transformer and multi-scale feature aggregation according to claim 4, characterized in that, The grouped self-attention mechanism adopts a local time block processing method, divides fixed-length blocks in the time dimension, performs local self-attention calculations within each time block, and adopts a Top-K screening strategy to retain only the key frames that are the most relevant within the local temporal blocks.
6. The non-contact heart rate detection method based on Vision Transformer and multi-scale feature aggregation according to claim 1, characterized in that, In step S3, the rPPG predictor consists of multiple three-dimensional convolutional layers and a linear mapping layer. First, the three-dimensional convolutional layers are used for feature extraction, and the low-dimensional temporal features are restored through an upsampling operation to maintain the temporal integrity of the heart rate information. The linear mapping layer is used for feature decoding and conversion, mapping the extracted high-dimensional features to a preliminary rPPG waveform. The rPPG waveform is subjected to a fast Fourier transform and power spectral density analysis to calculate the main peak frequency, and the heart rate value of the target is obtained by inferring the heart beat period based on the location of the main peak frequency.
7. A non-contact heart rate detection method based on Vision Transformer and multi-scale feature aggregation according to claim 1, characterized in that The loss function of the contactless heart rate extraction model is: Based on the predicted rPPG signal and the pre-acquired reference PPG signal, calculate the negative Pearson correlation coefficient between the predicted rPPG signal and the reference PPG signal, measure the time series correlation of the predicted rPPG signal, and optimize the fitting ability of the predicted rPPG signal to the reference PPG signal to construct the time-domain loss function L time : where T is the signal length, s(t), are the reference PPG signal and the predicted rPPG signal, respectively; In the frequency domain, by calculating the power spectral densities of the predicted rPPG signal and the reference PPG signal and using the cross-entropy loss function to measure the frequency matching degree between the two, it is ensured that the main frequency components of the model output signal are consistent with the true heart rate: where maxIdx represents the peak frequency index for obtaining the power spectral density, and PSD(s) represents the power spectral density of the true signal s. represents the predicted signal of the power spectral density; Using the optimal time-shifted mean squared error loss, allows for optimal time alignment within a given time window [-δ t , δ t to minimize the error between the predicted rPPG signal and the reference PPG signal: where s(t) and are the reference PPG signal and the predicted rPPG signal respectively, and τ is the time offset to ensure that the model can adapt to small time offsets and improve the alignment accuracy; Combining the time-domain loss, the frequency-domain loss, and the optimal time shift mean square error loss, the final total loss function is constructed: Loss all = λ1·L time + λ2·L freq + λ3·L MSE where λ1, λ2, and λ3 are hyperparameter weights used to balance the contributions of each loss term during the optimization process.
8. A non-contact heart rate detection system based on Vision Transformer and multi-scale feature aggregation, characterized in that, Including: Facial visible light video acquisition module: Acquire the facial visible light video, unify the video frame rate through frame rate resampling, perform face key point localization and region division on each frame image, and obtain the facial visible light video sequence; Low-resolution feature tensor acquisition module: Using the time shift operation according to the facial visible light video sequence, obtain the differential frame sequence, perform channel fusion on the facial video sequence and the differential frames, and downsample the spatial dimension to obtain the low-resolution feature tensor; Non-contact heart rate extraction model construction module: Construct a non-contact heart rate extraction model, input the low-resolution feature tensor into the non-contact heart rate extraction model, and output the heart rate value of the target; the non-contact heart rate extraction model includes a feature enhancement module, a multi-scale mask feature aggregation module, a Transformer temporal modeling module, and an rPPG predictor; Among them, input the low-resolution feature tensor into the feature enhancement module, and extract local details and global context information through parallel dilated convolution and multi-branch convolution respectively to obtain an enhanced feature map; input the enhanced feature map into the sparse attention mechanism, and screen and weight the channels and pixel positions most relevant to rPPG within the spatio-temporal range through lightweight and grouped self-attention to obtain a sparse attention output feature map; Input the sparse attention output feature map into the multi-scale mask feature aggregation module, which is divided into multiple parallel branches; set dilated convolutions with different dilation rates in each branch to extract skin textures and illumination changes at different local and global scales; fuse the outputs of each branch through splicing and convolution to obtain a multi-scale mask aggregated feature representation; segment and fuse the linear position encoding of the multi-scale mask aggregated feature representation to obtain a time series that meets the requirements of periodic time modeling, and output a feature sequence with fused position encoding; Input the feature sequence with fused position encoding into the Transformer temporal modeling module, and perform global feature extraction through the multi-head self-attention mechanism to obtain high-dimensional features with both spatio-temporal context and periodic information; Input the high-dimensional features into the rPPG predictor, and output the heart rate value of the target.
9. A non-contact heart rate detection device based on Vision Transformer and multi-scale feature aggregation, characterized in that, Including: Memory: Store the computer program of a non-contact heart rate detection method based on visual Transformer and multi-scale feature aggregation according to any one of claims 1-7, which is a computer-readable device; Processor: Used to implement a non-contact heart rate detection method based on visual Transformer and multi-scale feature aggregation according to any one of claims 1-7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it can implement a non-contact heart rate detection method based on visual Transformer and multi-scale feature aggregation according to any one of claims 1-7.
Citation Information
Cited By
Global-local perception violation behavior identification method and system based on multiple modes
CN121170712A
Non-contact heart rate measurement method and system based on space-time enhancement network
CN121370107A
Non-contact body surface temperature prediction method and device based on C-BTANeXt network
CN121902100A
Non-contact real-time heart rate measurement method and system and medium
CN122376066A
A motion-artifact-resistant video heart rate estimation model, method, device and medium
CN122510940A