Lightweight real-time heart rate monitoring model and heart rate monitoring method

Through the deep learning model of lightweight dual-branch architecture, the existing camera heart rate monitoring technology has solved the problems of large computing volume, poor real-time performance and cross-scene adaptability, and efficient and real-time heart rate monitoring on mobile devices is achieved, adapting to complex scenarios and improving monitoring accuracy.

CN120339728AActive Publication Date: 2025-07-18JILIN UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510807876.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-18
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

The existing camera-based heart rate monitoring technology has problems such as huge computing volume, difficulty in real-time operation, poor ability to generalize across scenes, and serious facial motion interference, which cannot meet the real-time heart rate monitoring needs on mobile devices.

Method used

The deep learning model adopts a lightweight dual-branch architecture, including spatiotemporal feature branches and motion compensation branches, is processed in parallel, combined with multi-task loss function, reduces computational complexity through separable 3D convolution and deformable convolution, enhances model adaptability, and improves heart rate monitoring accuracy using a second-order Butterworth low-pass filter and peak detection algorithm.

Benefits of technology

It realizes efficient and real-time heart rate monitoring on mobile devices, can run above 30FPS, adapt to different skin tones, lighting conditions and facial movements, improves monitoring accuracy and stability, and reduces computing resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339728A_ABST
    Figure CN120339728A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of crossing of computer vision and biomedical sensing, and particularly provides a lightweight real-time heart rate monitoring model and a heart rate monitoring method. The model comprises a video preprocessing module, a lightweight double-branch network module and a heart rate calculation module, wherein the video preprocessing module is used for carrying out standardized preprocessing on an input video; the lightweight double-branch network module is used for carrying out spatio-temporal feature and motion compensation branch parallel processing on a video image; the heart rate calculation module is used for receiving heart rate prediction related features output by the lightweight double-branch network module and carrying out heart rate calculation; according to the method, a multi-task loss function module considering multiple loss cascade optimization is adopted, training data is used for training the model, lightweight real-time heart rate monitoring can be achieved through the model after training is completed, and the heart rate can be predicted and obtained according to real-time images collected by a camera. Based on the rPPG technology, high-precision real-time heart rate estimation is achieved through the deep learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - technical field of computer vision and biomedical sensing, and particularly relates to a lightweight real - time heart rate monitoring model and a heart rate monitoring method. Background Art

[0002] Traditional heart rate monitoring mainly relies on contact sensors, such as electrocardiogram (ECG) devices and pulse oximeters. Although these devices can provide relatively accurate measurement results in clinical settings, they have obvious drawbacks. First, they are invasive, and long - term wearing can cause discomfort to the skin. For example, after wearing ECG electrode patches continuously for several hours, the skin may show symptoms such as allergies and redness, affecting the user experience and the continuity of monitoring. Second, the usage scenarios are limited. During exercise, due to factors such as body movement and sweating, the signals are easily interfered with, resulting in inaccurate measurement results; in an environment with frequent light changes, such as when moving from an indoor strong - light environment to an outdoor natural - light environment, the measurement accuracy of the pulse oximeter will be severely affected. Third, the cost of professional medical equipment is relatively high, making it difficult to be widely popularized in ordinary families, schools, gyms and other places, which limits the widespread application of heart rate monitoring technology.

[0003] The remote photoplethysmography (rPPG) technology based on cameras has received much attention in recent years. It extracts heart rate signals by analyzing the changes in skin - reflected light, opening up a new path for non - contact heart rate monitoring. However, this technology currently faces many challenges. On the one hand, facial micro - movements (such as talking, chewing, and facial expression changes) and environmental light fluctuations (such as indoor light flickering and outdoor sunlight intensity changes) can easily cause the extracted heart rate signals to be distorted. Facial micro - movements may interfere with the changes in skin - reflected light, making it difficult to accurately isolate the signals related to heart rate; environmental light fluctuations will change the light intensity collected by the camera, affecting the signal quality. On the other hand, traditional rPPG algorithms have poor generalization. Methods such as independent component analysis (ICA) and CHROM rely too much on manually designed features and cannot adaptively adjust when facing complex and variable real - world scenarios with different skin colors, lighting conditions, and facial postures, resulting in a significant decrease in monitoring accuracy. Moreover, some complex deep - learning - based models (such as 3D - CNN) show high accuracy on specific datasets, but have huge computational requirements and are difficult to run in real - time on resource - constrained devices such as mobile phones, unable to meet the user's need for instant monitoring.

[0004] Currently, the existing camera - based rPPG technologies mainly include signal - processing methods, supervised - learning methods, and related open - source tools.

[0005] In terms of signal processing methods, the POS algorithm is representative. Based on the principle of physiological optics, the algorithm constructs a projection plane orthogonal to the skin tone, projects the spatially normalized and averaged pixel values onto the plane, thereby restoring the photoplethysmogram (PPG) waveform and extracting the heart rate signal. However, in actual applications, when the face moves violently, such as turning the head quickly or changing facial expressions drastically, the POS algorithm is difficult to accurately track the changes in the light reflected from the skin, resulting in large errors in heart rate monitoring.

[0006] In terms of supervised learning methods, PhysNet uses a 3D convolutional network to model the spatiotemporal features of videos. Through multiple layers of 3D convolutional layers and pooling layers, the heart rate-related features in the video are automatically learned. However, the model has a parameter volume of up to 10M, and a lot of computing resources and time are required when predicting heart rate. Taking common mobile devices as an example, when running the PhysNet model for heart rate monitoring, it may take hundreds of milliseconds or even longer to process a frame of video, which is far from meeting the requirements of real-time monitoring (real-time monitoring usually requires a frame rate of 30FPS or above, that is, the processing time per frame is within 33 milliseconds).

[0007] In terms of related open source tools, rPPG-Toolbox integrates 6 supervised and unsupervised algorithms, providing a convenient platform for the research and application of rPPG technology. It supports the preprocessing, model training and evaluation of a variety of public data sets, promoting the development of the rPPG field. However, the tool lacks lightweight models optimized specifically for mobile terminals. When running on mobile devices, it cannot fully utilize the performance of the device and it is difficult to achieve efficient real-time heart rate monitoring.

[0008] Computational efficiency is a key factor limiting the application of existing rPPG technology. Deep learning models such as PhysNet have huge parameters and complex model structures. When running on mobile devices, serious delays will occur due to the limited computing power and memory of the device. This not only fails to meet the needs of real-time heart rate monitoring, but may also lead to data loss or inaccurate monitoring results, greatly limiting its application in practical scenarios.

[0009] At the feature extraction level, traditional methods rely heavily on manually designed color space transformation and feature engineering, which makes it difficult for the algorithm to adapt to complex and changing real-world scenarios. Differences in skin color among different races and changes in lighting conditions in different environments may cause the manually designed features to fail, thus affecting the accuracy of heart rate monitoring.

[0010] The cross-scenario generalization ability is also a shortcoming of existing rPPG technologies. When existing rPPG models are trained, they often rely on specific datasets, which have certain limitations in terms of skin color, lighting, facial movements, etc. When the model is applied to scenarios that are significantly different from the training set, its performance will drop significantly. For example, for a model trained mainly on indoor white-face video data, when performing heart rate monitoring on outdoor-acquired Asian-face videos, the error increases significantly.

[0011] Therefore, there is an urgent need to design a lightweight end-to-end heart rate detection method to overcome the limitations of traditional technologies in terms of computational efficiency, feature extraction, and cross-scenario generalization applications. Summary of the Invention

[0012] In view of this, the present invention aims to provide a lightweight real-time heart rate monitoring model and a heart rate monitoring method. By adopting remote photoplethysmography (rPPG), facial videos are collected with an ordinary camera, and a deep learning model is used to achieve high-precision and real-time heart rate estimation. A lightweight double-branch architecture is adopted, with the spatio-temporal feature and motion compensation branches processed in parallel. While ensuring the accuracy of heart rate monitoring, the computational complexity of the model is significantly reduced, and the real-time heart rate monitoring frame rate ≥ 30 FPS, which can meet the user's need for instant access to heart rate data.

[0013] To achieve the above object, the technical solution of the present invention is realized as follows: On the one hand, the present invention provides a lightweight real-time heart rate monitoring model, including: A video preprocessing module, which is used to preprocess the input video, detect and align the face in the input video, and perform spatio-temporal partitioning on the input video to obtain standardized video blocks; A lightweight double-branch network module, which includes a spatio-temporal feature branch and a motion compensation branch with the input being video blocks, and a feature fusion unit. Among them, the spatio-temporal feature branch includes multiple layers of separable 3D convolutions. Each layer of separable 3D convolution is used to perform depth convolution on each channel of the input data respectively, and then the results of the depth convolution are combined in the channel dimension through spatial convolution to obtain the spatio-temporal features in the video block; the motion compensation branch is used to dynamically capture the motion information of the face, obtain the motion offset field to perform motion compensation on the features in the video block; the feature fusion unit is used to perform feature fusion on the outputs of the spatio-temporal feature branch and the motion compensation branch to obtain relevant features for heart rate prediction; A heart rate calculation module, which calculates the heart rate based on the heart rate prediction-related features output by the lightweight double-branch network module and outputs the predicted heart rate; A multi-task loss function module, which includes a heart rate loss, a noise suppression loss, and a regularization loss, and optimizes the model based on these three losses.

[0014] Preferably, the video preprocessing module uses the MTCNN algorithm to perform face detection and alignment on the input video, and crops the face region into a region of interest with a fixed resolution.

[0015] Preferably, after obtaining the region of interest, the video preprocessing module divides the input video into a sequence of non-overlapping T frames in a spatio-temporal block manner, and divides each frame image into an N×N grid to extract local spatio-temporal features.

[0016] Preferably, after extracting the local spatio-temporal features of the input video, it further includes performing dynamic normalization processing on the RGB channels of the input video. The calculation formula for dynamic normalization processing is: ; where, represents the normalized pixel value of the t-th frame in channel c, represents any channel in the RGB channels, represents the pixel value of the t-th frame in channel c, represents the pixel mean of the t-th frame in channel c, represents the pixel standard deviation of the t-th frame in channel c, is a constant used to prevent division by zero errors during the calculation process.

[0017] Preferably, the spatio-temporal feature branch includes 4 layers of separable 3D convolution. The kernel size of each layer of separable 3D convolution is 3×3×3, and the stride is 2×2×2. The operation of separable 3D convolution is: ; where, represents the -th layer of separable 3D convolution, represents the output result of the -th layer of separable 3D convolution, represents the separable 3D convolution operation.

[0018] Preferably, the motion compensation branch includes at least two layers of deformable convolution. The deformable convolution adaptively adjusts the sampling position of the convolution kernel according to the video block to obtain face motion information. The operation of each layer of deformable convolution is: ; where, represents the output result of the deformable convolution, represents the position on the output feature map, represents the fixed position of the convolution kernel, that is, in the case of no offset, the -th sampling point of the convolution kernel should be located at the coordinate on the input feature map, represents the learned relative to the position at ​ The offset of represents the weight of the convolutional kernel, represents the total number of sampling points in the convolutional kernel, represents the input feature map. The feature map input to the first deformable convolution is the normalized video block X, and the feature map input to the second deformable convolution is the output result of the first deformable convolution.

[0019] Preferably, the motion compensation branch uses optical flow estimation to obtain the motion offset field, thereby achieving motion compensation.

[0020] Preferably, the feature fusion unit splices the final output result of the spatio-temporal feature branch with the final output result of the motion compensation branch

[0021] and performs dimensionality reduction processing through 1×1 convolution to obtain and output the features related to heart rate prediction.

[0022] Preferably, the heart rate calculation module obtains the preliminary heart rate prediction value based on the heart rate prediction-related features output by the lightweight double-branch network module, filters the preliminary heart rate prediction value through a second-order Butterworth low-pass filter, performs peak detection on the filtered signal, and calculates the heart rate based on the time interval between adjacent peaks to obtain the predictable heart rate that can be read and used. The expression of the heart rate loss is: ; where represents the true heart rate of the th training sample, represents the predicted heart rate of the th training sample, represents the number of samples for model training, represents the average value of the true heart rate, represents the average value of the predicted heart rate; The expression of the noise suppression loss is: ; where represents the energy of the predicted heart rate at the frequency ; The regularization loss uses the L2 regularization term, and the expression of the regularization loss is: ; where represents the regularization coefficient, represents the set of parameters of the model, Represents any parameter in the model.

[0023] Another aspect of the present invention provides a lightweight real-time heart rate monitoring method, which uses a lightweight real-time heart rate monitoring model to perform heart rate detection and obtain a predicted heart rate.

[0024] Compared with the prior art, the invention can achieve the following beneficial effects: Based on rPPG technology, the present invention designs and implements a lightweight end-to-end neural network model that can predict the heart rate of people in videos with high precision and in real time. In the feature extraction process, a lightweight dual-branch architecture is adopted to design the spatiotemporal feature extraction and motion compensation branches in parallel. This architecture innovatively separates the spatiotemporal feature extraction and motion compensation, so that the model can learn different aspects of information more attentively. The spatiotemporal feature branch uses separable 3D convolution to efficiently extract the spatiotemporal features of the video, greatly reducing the computational complexity of the model and the amount of calculation, so that the model can run efficiently on mobile devices, realize real-time heart rate monitoring, and meet the user's demand for instant acquisition of heart rate data. Compared with traditional complex models, such as PhysNet, which is difficult to deploy locally on the mobile terminal due to the amount of calculation, the present invention greatly improves the computational efficiency while ensuring the accuracy of heart rate monitoring, reduces the time to process each frame of video, and can achieve a smooth real-time monitoring effect, and achieve a real-time heart rate monitoring frame rate of ≥30FPS.

[0025] The present invention also adds a motion compensation branch to the model. The motion compensation branch adopts an adaptive motion compensation mechanism and uses technologies such as deformable convolution to dynamically capture facial motion information, effectively reducing the interference of facial motion on heart rate monitoring and providing support for the accurate extraction of heart rate signals. Compared with traditional POS algorithms, when facing intense facial movements, deformable convolution breaks through the limitations of the fixed sampling position of the traditional convolution kernel, and adaptively adjusts the sampling position according to the image content, so that the model can better adapt to changes in facial movement and can more accurately track changes in skin reflected light, thereby more stably and accurately extracting heart rate signals and improving the accuracy and stability of heart rate monitoring results.

[0026] The present invention designs a multi-task loss function, introduces a noise suppression loss to suppress the noise in non-heart rate frequency bands, takes heart rate estimation and noise suppression as jointly optimized tasks, and makes full use of the correlation between the two. During the training process, the heart rate loss guides the model to learn accurate heart rate features, the noise suppression loss constrains the model to reduce the noise in non-heart rate frequency bands, and the regularization loss prevents the model from overfitting, improving the generalization ability of the model in different datasets and complex scenarios. The model of the present invention can better adapt to changes in different skin colors, lighting conditions, and facial movement situations, and can stably and accurately monitor the heart rate in various actual application scenarios, overcoming the problem that the performance of existing models significantly decreases in cross-scenario applications.

[0027] In the process of calculating the heart rate, the present invention processes the preliminary heart rate prediction value output by the lightweight double-branch network through a second-order Butterworth low-pass filter combined with a peak detection algorithm, effectively removing the noise interference and further improving the accuracy and stability of the finally output heart rate.

[0028] The present invention greatly reduces the model calculation amount. The number of model parameters is designed to be close to 0.8M, and the peak memory consumption during the model operation is expected to be ≤50MB, with low resource requirements for mobile devices, enabling it to be deployed on mobile devices for local operation. Compared with some existing models with large numbers of parameters and high memory occupancy, when running on mobile devices, it will not have a great impact on other functions of the device, ensuring the smooth operation of the device, broadening the application scope of the technology, and being particularly suitable for resource-constrained mobile devices and embedded systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings: Figure 1 is the overall architecture modular flowchart of the lightweight real-time heart rate monitoring model provided by the embodiment of the present invention; Figure 2 is the working flowchart of the video preprocessing module provided by the embodiment of the present invention; Figure 3 is the architecture diagram of the lightweight double-branch network module provided by the embodiment of the present invention; Figure 4 is the working flowchart of the heart rate calculation module provided by the embodiment of the present invention; Figure 5 is the working flowchart of the multi-task loss function module provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation to the present invention. Similar elements in different embodiments are labeled with related similar element numbers. In the following embodiments, many details are described to enable a better understanding of the present invention. However, those skilled in the art can easily recognize that some of the features can be omitted in different situations, or can be replaced by other elements, materials, and methods. In some cases, some operations related to the present invention are not shown or described in the specification, which is to avoid the core part of the present invention being overwhelmed by excessive description. For those skilled in the art, it is not necessary to describe these related operations in detail, and they can fully understand the related operations based on the description in the specification and the general technical knowledge in the field.

[0031] It should be noted that, without conflict, the embodiments and features in the embodiments of the present invention can be combined with each other to form various embodiments. At the same time, the steps or actions in the method description can also be adjusted in the order that is obvious to those skilled in the art. Therefore, the various sequences in the specification and drawings are only for clearly describing a certain embodiment and do not mean that they are the necessary sequences, unless it is stated that a certain sequence must be followed.

[0032] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.

[0033] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations.

[0034] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0035] Please refer to Figure 1 , in an embodiment of the present invention, a lightweight real-time heart rate monitoring model is provided to solve the problems of poor timeliness, serious noise interference, large computational complexity, and inability to be deployed on mobile devices with general computing capabilities in traditional heart rate prediction models. The model architecture mainly includes: a video preprocessing module, a lightweight dual-branch network module, a heart rate calculation module, and a multi-task loss function module.

[0036] Please refer to Figure 2, the video preprocessing module is used to perform standardized preprocessing on the input video. Specifically, the model is deployed on a smartphone (a typical device with general computing capabilities), and a 30Hz video is captured through the smartphone camera and used as the input video for the heart rate monitoring model. After the 30Hz video captured by the camera is input into the video preprocessing module, since the resolution of the original input video is affected by specific devices, cameras, and mobile phone camera parameters, etc., the resolution of the input video may be diverse. To facilitate subsequent processing, it needs to be converted to a unified resolution. Assume that the resolution of the original input video is the common 1280×720, and it is converted to a preset unified resolution. After the resolution is unified, the MTCNN (Multi-task Cascaded Convolutional Networks) algorithm is used to perform face detection and alignment on each frame image of the input video. The MTCNN algorithm can detect the face region in each frame image of the input video through a cascaded convolutional neural network, quickly and accurately locate multiple key points of the face, such as the positions of eyes, nose, mouth, etc. Based on these key points, the face region is cropped into a ROI (Region of Interest) region with a fixed resolution. The fixed resolution can usually be selected as 128×128, 256×256, etc. During the cropping process, usually centered on a certain face key point, the key parts of the face are completely included in the cropping region for cropping, focusing on the face and removing the interference of background information, providing purer data for subsequent heart rate signal extraction. Through the above processing, the diverse resolutions in the original input video can be uniformly processed, and the face region of interest can be accurately cropped. MTCNN can detect and align the cropped face images, providing high-quality data for subsequent video processing tasks.

[0037] After face detection and alignment, continue to divide the cropped image into blocks. Specifically, the cropped and aligned video is divided into multiple sequences of T frame images, where T = 180. That is, it is divided into an input sequence every 180 frames (6 seconds). In addition, during the process of dividing the image sequences, overlapping division or non-overlapping division can be adopted. Non-overlapping division means that each image sequence has no overlapping time. The first sequence: frame 1 to frame 180; the second sequence: frame 181 to frame 360, and so on. This division method can ensure that each sequence is independent in the time dimension, facilitating subsequent analysis of the dynamic changes of the heart rate signal. Overlapping division means that each image sequence has overlapping time. In the embodiment of the present invention, an overlapping rate of 50% is set for adjacent image sequences. This design can ensure data continuity while making full use of the information in the video, avoiding information loss caused by block division and making it difficult for the model to learn the complete change trend of the heart rate signal.

[0038] After obtaining the image sequence by partitioning, for each sequence of 180 frames, each frame image is further divided into an N×N grid, where N = 8. Each frame image is divided into 64 grid regions, achieving spatio-temporal partitioning of the video image. In the time dimension, 180 consecutive frame images can capture the dynamic changes of the heart rate signal; in the space dimension, dividing each frame image into an 8×8 grid can obtain the subtle feature changes in different regions of the face. For example, different grid regions may correspond to different parts of the face, such as the forehead, cheeks, etc., and the correlations between the light reflection changes in these regions and the heart rate are different. Through partitioning, these features can be analyzed more carefully.

[0039] After spatio-temporal partitioning, dynamic normalization processing is performed on the RGB channels of each frame image in the video. For any t-th frame image, calculate the sum of the pixel values of the R, G, and B channels respectively, and then divide it by the number of pixels to obtain the mean value of any channel c of the t-th frame image. . Then, based on the sum of the squares of the differences between the pixel value of any channel c of each pixel in the t-th frame image and the corresponding mean value , and then divide it by the total number of pixels to obtain the pixel variance of channel c of the t-th frame. , and further take the square root to obtain the standard deviation , thus realizing dynamic normalization processing and obtaining the normalized video blocks. The calculation formula for dynamic normalization processing is: ; where, represents the normalized pixel value of channel c of the t-th frame, represents any channel in the RGB channels, represents the pixel value of channel c of the t-th frame, represents the pixel mean value of channel c of the t-th frame, represents the pixel standard deviation of channel c of the t-th frame, is a very small constant used to prevent division by zero errors during the calculation process.

[0040] After the above dynamic normalization processing, the RGB channel data of different videos can have a unified scale, which can eliminate the influence of factors such as illumination intensity differences on the data, improve the stability and convergence speed of model training. At the same time, it can also ensure the robustness during the application of the model.

[0041] Please refer to Figure 3 , the normalized video blocks output by the video preprocessing module Input lightweight double-branch network module to extract spatio-temporal features and perform motion compensation on video blocks. Here, \(R\) represents the set of numbers, \(H\) represents the height of each image frame of the video block, \(W\) represents the width of each image frame of the video block, \(T = 180\), \(H = 128\), \(W = 128\), and 3 represents the three RGB channels. Traditional designs usually only use 3D convolution for feature extraction, lacking the motion compensation process, resulting in serious interference of model-predicted heart rate by noise such as facial movements. Or the spatio-temporal feature extraction and motion compensation are combined for processing, and the traditional 3D convolution is an integrated calculation method, with a large amount of calculation and low calculation efficiency, which cannot be applied to devices with small calculation amounts and has relatively high requirements for the deployed devices. Therefore, the embodiment of the present invention designs the spatio-temporal feature extraction and motion compensation branches in parallel, adopting a lightweight double-branch structure. Specifically, the video block is simultaneously input into the parallel spatio-temporal feature branch and motion compensation branch. Among them, the spatio-temporal feature branch decomposes the traditional 3D convolution into two steps: spatial convolution and depth convolution. In terms of the branch structure, the spatio-temporal feature branch includes 4 progressive separable 3D convolutions, with the kernel size of each layer being \(3\times3\times3\) and the stride being \(2\times2\times2\). This design can decompose the traditional 3D convolution into two steps: spatial convolution and depth convolution, so as to significantly reduce the number of model parameters and the amount of calculation without losing too much performance.

[0042] For each separable 3D convolution, its operation is as follows: ; Among them, represents the th separable 3D convolution, represents the output result of the th separable 3D convolution, , represents the separable 3D convolution operation.

[0043] For any separable 3D convolution, first perform depth convolution on the input image data (i.e., the output of the previous separable 3D convolution, and the input of the first separable 3D convolution is the video block ). Perform a \(3\times3\times3\) convolution operation on each channel respectively to extract the local spatio-temporal features of each channel. Then combine the results of the depth convolution in the channel dimension to achieve cross-channel information fusion. This decomposition method can not only significantly reduce the amount of calculation and the number of parameters, but also effectively extract the spatio-temporal features of the video block. After 4 layers of separable 3D convolution processing, according to the input video block output the convolution feature map (the final output result of the spatio-temporal feature branch), and .

[0044] The motion compensation branch includes two layers of deformable convolution (Deformable Conv). By introducing learnable offsets, it can adaptively and dynamically adjust the sampling positions of the convolution kernel according to the content of each frame of the image, thus being able to better adapt to the changes in the image content. In contrast, traditional motion compensation uses fixed sampling positions (usually regular grid points), and its compensation ability for complex dynamic changes is limited.

[0045] During actual calculation, by sampling the input feature map at position and multiplying it with the weight and then summing them up, the output feature map is obtained at the value at position . This operation mechanism enables the convolution kernel to dynamically adjust the sampling position according to the facial motion situation and better capture the motion features. The specific operation for each layer of deformable convolution is as follows: ; Among them, represents the output result of the deformable convolution, represents the position on the output feature map, represents the fixed position of the convolution kernel, that is, in the case of no offset, the th sampling point of the convolution kernel should be located at the coordinates on the input feature map, represents the learned offset relative to at position , represents the weight of the convolution kernel, represents the total number of sampling points in the convolution kernel, represents the input feature map. The input feature map of the first layer of deformable convolution is the normalized video block , and the input feature map of the second layer of deformable convolution is the output result of the first layer of deformable convolution.

[0046] After two layers of deformable convolution, a motion offset field can be output, where 2 represents the offset in the x and y directions. The input video block is adjusted according to the motion offset field to obtain the motion-compensated feature (the final output result of the motion compensation branch).

[0047] After the spatio-temporal feature branch and the motion compensation branch perform feature extraction and motion compensation, the corresponding output results are concatenated with the input feature fusion unit. The concatenation operation means concatenating with Merge by channel dimension to form a high - dimensional feature map. The concatenated high - dimensional feature map usually has a high number of channels, which increases the complexity of subsequent calculations. To reduce the feature dimension and fuse information, 1×1 convolution can be used for dimensionality reduction. The role of 1×1 convolution is to adjust the number of channels without changing the spatial size of the feature map, realizing feature fusion and dimensionality reduction, enabling the model to more effectively utilize these features for subsequent heart rate calculation. Map the concatenated high - dimensional features to a low - dimensional space and output the features related to heart rate prediction.

[0048] Please refer to Figure 4 , input the features related to heart rate prediction output by the lightweight double - branch network module into the heart rate calculation module. The heart rate calculation module calculates the preliminary heart rate prediction value based on the spatio - temporal features of the video and the features related to motion compensation information learned by the lightweight double - branch network module. However, this preliminary heart rate prediction value still has noise and cannot be directly read and output. Therefore, a second - order Butterworth low - pass filter is further used to filter the preliminary heart rate prediction value. By setting an appropriate cut - off frequency, high - frequency noise is filtered out, making the prediction signal smoother and reducing the interference of noise on heart rate calculation. In practical applications, according to the frequency characteristics of the heart rate signal, selecting an appropriate cut - off frequency range can better retain the characteristics of the heart rate signal while removing most other interfering frequency components. According to actual human physiological parameters, the cut - off frequency range is usually selected as 0.75 - 2.5Hz. Therefore, the transfer function of the second - order Butterworth low - pass filter can be expressed as: ; where, are the poles of the second - order Butterworth low - pass filter, is the complex frequency variable, is the order of the denominator polynomial, indicating that there are poles.

[0049] After filtering, peak detection is performed on the filtered preliminary heart rate prediction signal. In the filtered signal, local peak points are found, and these peak points correspond to the fluctuations of the heart rate signal. The heart rate can be calculated according to the time interval between adjacent peaks. The calculation formula for heart rate is: ; where, represents the heart rate (beats per minute), is the average time interval (seconds) between adjacent peaks. After the above - mentioned processing, the preliminary heart rate prediction value can be converted into a practically readable and usable predicted heart rate. Thus, the predicted heart rate of the person in the input video can be obtained and output.

[0050] In addition, there is also a multi - task loss function module for model training and optimization. For details, please refer toFigure 5 , in the embodiments of the present invention, the performance of the model is improved by jointly optimizing heart rate estimation (main task) and frequency domain noise suppression (auxiliary task). The total loss function of the model includes heart rate loss, noise suppression loss, and regularization loss. The model is optimized by feedback based on these three losses. This multi-task learning strategy can, while optimizing the main task (heart rate estimation), utilize the auxiliary task (noise suppression) to provide additional supervision information, and prevent the model from overfitting through regularization, thereby improving the overall performance and generalization ability of the model.

[0051] Among them, the heart rate loss uses the negative Pearson correlation coefficient to measure the correlation between the final predicted heart rate and the true heart rate. The final predicted heart rate is the predicted heart rate value finally output by the model, and the true heart rate is the actual heart rate of the person in the training samples during the training process. When the model is trained and applied, there is no actual heart rate input, so the loss function will not be used to optimize the model. However, after the model is trained, additional training samples and test samples can still be used to verify and optimize the model. Heart rate loss The expression is: ; Among them, represents the true heart rate of the th training sample, represents the predicted heart rate of the th training sample, represents the number of samples for model training, represents the average value of the true heart rate, represents the average value of the predicted heart rate. The Pearson correlation coefficient is used to measure the linear correlation between two variables, and its value range is between -1 and +1. In the embodiments of the present invention, the negative Pearson correlation coefficient is used as the loss function, aiming to make the correlation between the predicted heart rate and the true heart rate as high as possible, that is, the smaller the loss function value, the closer the heart rate predicted by the model is to the true value.

[0052] The noise suppression loss is used to constrain the energy of the predicted signal outside the heart rate frequency band (0.7 - 2.5 Hz). The noise suppression loss The expression is: ; Among them, represents the energy of the predicted heart rate at frequency .

[0053] Through the noise suppression loss the energy of the predicted heart rate outside the heart rate frequency band can be constrained, non-heart rate related noise signals can be suppressed, and the quality of the heart rate signal can be improved. The noise suppression loss During the training process, it can reflect the noise filtering effect of the evaluation model, ensuring that when the model predicts the heart rate, it mainly focuses on the signals within the heart rate frequency band and ignores the noise interference in other frequency bands.

[0054] The regularization loss uses the L2 regularization term to prevent overfitting in model training. The regularization loss has the following expression: ; where represents the regularization coefficient, represents the set of model parameters, represents any parameter in the model. L2 regularization penalizes through the sum of squares of model parameters to prevent model overfitting. During the training process, large parameter values may cause the model to be too complex and prone to overfitting. Through L2 regularization, the model parameters can be kept within a reasonable range, thereby improving the generalization ability of the model.

[0055] Summarize the above three loss calculations into the total loss calculation module of the model, and feedback the total loss to the model, then the model optimization training can be achieved. When the total loss meets the preset conditions, it can be confirmed that the model training is completed. The total loss can adopt the weighted sum operation of the above three losses, and the weight coefficients of the three losses can be adaptively adjusted according to actual requirements. By jointly optimizing the heart rate loss, noise suppression loss, and regularization loss, the model can improve the accuracy of heart rate estimation while suppressing noise interference and preventing overfitting, thus significantly improving the performance and generalization ability of the model.

[0056] As an alternative embodiment, the model proposed in the embodiments of the present invention can be deployed on a mobile device processor, including but not limited to mobile phone terminals, PC terminals, cloud platforms, and chips with performance equivalent to that of current mainstream mobile processors, etc. In actual deployment, it is necessary to adapt and optimize the model according to the architectural characteristics of the deployment platform, such as using its specific instruction set to accelerate the calculation process and improve the inference efficiency.

[0057] As an alternative embodiment, for the motion compensation branch, optical flow estimation (such as RAFT) can be used to replace deformable convolution for motion compensation. Optical flow estimation aims to calculate the motion vectors of pixels between adjacent frames in a video, and its core principle is based on the assumptions of brightness constancy, small motion, and spatial consistency. Taking only the optical flow estimation based on variational method as an example, its basic calculation formula is: ; where is the energy function for optical flow calculation, is the image intensity, is the optical flow vector, representing the motion displacement of pixels in the x and y directions, represents the image intensity, that is, the pixel value of the image at the position and time The pixel value at is the regularization parameter, which is used to balance the weights of the data term and the smooth term.

[0058] Optical flow estimation can capture large-scale motion more precisely. By calculating the motion vector of each pixel, it can describe the changes in facial motion in more detail. However, since optical flow estimation requires complex calculations for each pixel to solve the optical flow equation, its computational complexity is higher compared to deformable convolution. Therefore, the disadvantage of optical flow estimation is that the computational amount will increase and the computational complexity is higher. In practical applications, if the accuracy requirement for motion compensation is extremely high and the computing resources of the device are relatively sufficient, optical flow estimation can be considered to replace deformable convolution.

[0059] As an alternative embodiment, for the design of the loss function during the training process, adversarial training (GAN) can be used to replace the original frequency-domain loss calculation. By introducing a generator and a discriminator to evaluate and feedback on the training situation. The goal of the generator is to output a prediction signal that is as close as possible to the true heart rate signal, while the discriminator is used to distinguish the spectrum of the prediction signal output by the generator from the true heart rate signal. Specifically, the generator attempts to minimize , and the discriminator attempts to maximize . The objective function of adversarial training can be expressed as: ; ; where is the loss function of the generator, representing the objective that the generator attempts to minimize, is the loss function of the discriminator, representing the objective that the discriminator attempts to maximize, is the true heart rate signal, that is, the true heart rate signal sampled from the training dataset, is the input noise of the generator, usually a random vector, is the distribution of the true heart rate signal data, is the distribution of the noise, is the output of the generator, representing the heart rate signal generated by the input noise , is the judgment of the discriminator on the true heart rate signal , representing the probability that the discriminator believes is the true signal, represents the judgment of the discriminator on the true signal in the loss function of the generator, represents the judgment of the discriminator on the generated signal in the loss function of the generator Judgment Indicates the discriminator's judgment on the real signal in the discriminator's loss function Judgment Indicates the discriminator's judgment on the generated signal in the discriminator's loss function Judgment.

[0060] Through adversarial training, the generator generates realistic heart rate signals, and the discriminator distinguishes between the generated signals and the real signals, which can enhance the model's learning ability of the spectral characteristics of heart rate signals and enable the generator to generate prediction results closer to the real heart rate signals. However, adversarial training is relatively complex, the training process requires careful adjustment of hyperparameters, and problems such as mode collapse are likely to occur, that is, the generator can only generate a limited number of types of samples, resulting in unstable training.

[0061] As an alternative embodiment, the model quantization can use floating-point numbers (such as FP32) or integers (such as INT8). Among them, converting the FP32 weights of the model to INT8 reduces the space occupied by each weight value from 32 bits to 8 bits, which can further compress the model volume, reduce the model storage requirements and computational amount. The principle of model quantization is to reduce the representation precision of weights and activation values, and reduce the storage and computational requirements of the model while keeping the model accuracy loss small. During the conversion process, the model needs to be calibrated. By running the model on a small amount of calibration data, the distribution of activation values is statistically analyzed to determine the quantization parameters to ensure that the performance loss of the quantized model is within an acceptable range. The compression of the model volume can not only reduce the storage overhead, but also speed up the model loading speed and improve the inference efficiency.

[0062] As an alternative embodiment, during the model deployment process, the model can be deployed to a specific hardware platform (such as NVIDIA Jetson series development boards) through a dedicated hardware acceleration framework (such as TensorRT). The hardware acceleration framework can optimize the model, such as layer fusion, quantization and other operations, to improve the inference speed of the model on specific hardware. During the deployment process, it is necessary to optimize the configuration of the model according to the characteristics of the hardware, such as the number of GPU cores, memory bandwidth, etc., to give full play to the hardware performance. At the same time, it is also necessary to consider the data transfer efficiency between the model and the hardware, reduce the data transfer time, and further improve the overall inference speed.

[0063] The effectiveness of the present invention has not been verified. When the optimized model is run on mobile devices such as smartphones, the average latency can be controlled within 40 ms, thus meeting the requirement of real-time processing at 30 FPS and enabling users to obtain heart rate monitoring results almost instantaneously. The number of model parameters is designed to be about 0.8M, and the peak memory consumption during operation is expected to be ≤ 50MB to ensure that when running on mobile devices, it will not have a significant impact on other functions of the device and guarantee the smooth operation of the device.

[0064] After the above lightweight real-time heart rate monitoring model is trained, the trained model can be used to detect the heart rate of the person in the video and output the corresponding predicted heart rate.

[0065] In summary, the above description is only a preferred embodiment of this specification and is not intended to limit the protection scope of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this specification shall be included within the protection scope of this specification.

[0066] The system, device, module or unit described in the above one or more embodiments can be specifically implemented by a computer chip or entity, or by a product with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0067] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent in such process, method, commodity or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the said element.

[0068] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.

[0069] The above description has been made of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A lightweight real-time heart rate monitoring model, characterized in that, Including: A video preprocessing module, which is used to preprocess the input video, detect and align the faces in the input video, and perform spatio-temporal partitioning on the input video to obtain standardized video blocks; A lightweight dual-branch network module, which includes a spatio-temporal feature branch and a motion compensation branch with the video blocks as inputs, and a feature fusion unit. Among them, the spatio-temporal feature branch includes multiple layers of separable 3D convolutions. Each layer of separable 3D convolution is used to perform depth convolution on each channel of the input data respectively, and then combine the results of the depth convolution through spatial convolution in the channel dimension to obtain the spatio-temporal features in the video blocks; the motion compensation branch is used to dynamically capture the motion information of the face, obtain a motion offset field to perform motion compensation on the features in the video blocks; the feature fusion unit is used to perform feature fusion on the outputs of the spatio-temporal feature branch and the motion compensation branch to obtain relevant features for heart rate prediction; A heart rate calculation module, which calculates the heart rate based on the heart rate prediction relevant features output by the lightweight dual-branch network module and outputs the predicted heart rate; A multi-task loss function module, which includes a heart rate loss, a noise suppression loss, and a regularization loss, and optimizes the model according to these three losses.

2. The lightweight real-time heart rate monitoring model according to claim 1, wherein The video preprocessing module uses the MTCNN algorithm to detect and align the faces in the input video, and crops the face region into a region of interest with a fixed resolution.

3. The lightweight real-time heart rate monitoring model according to claim 2, wherein After obtaining the region of interest, the video preprocessing module divides the input video into non-overlapping sequences of T frames in a spatio-temporal partitioning manner, and divides each frame image into an N×N grid to extract local spatio-temporal features.

4. The lightweight real-time heart rate monitoring model according to claim 1, wherein After extracting the local spatio-temporal features in the input video, it further includes performing dynamic normalization processing on the RGB channels of the input video. The calculation formula for the dynamic normalization processing is: ; Among them, represents the normalized pixel value of the t-th frame in channel c, represents any channel in the RGB channels, represents the pixel value of the t-th frame in channel c, represents the pixel mean value of the t-th frame in channel c, represents the pixel standard deviation of the t-th frame in channel c, is a constant used to prevent division-by-zero errors during the calculation process.

5. The lightweight real-time heart rate monitoring model according to claim 1, wherein The spatio-temporal feature branch includes 4 layers of separable 3D convolutions. The kernel size of each layer of separable 3D convolution is 3×3×3, and the stride is 2×2×2. The operation of the separable 3D convolution is: ; Among them, represents the layer of separable 3D convolution, represents the output result of the layer of separable 3D convolution, represents the separable 3D convolution operation.

6. The lightweight real-time heart rate monitoring model according to claim 1, wherein The motion compensation branch includes at least two layers of deformable convolutions. The deformable convolutions adaptively adjust the sampling positions of the convolution kernels according to the video blocks to obtain the face motion information. The operation of each layer of deformable convolution is: ; Among them, represents the output result of the deformable convolution, represents the position on the output feature map, represents the fixed position of the convolution kernel, that is, in the case of no offset, the th sampling point of the convolution kernel should be located at the coordinates on the input feature map, represents the offset learned at position relative to , represents the weight of the convolution kernel, represents the total number of sampling points in the convolution kernel, represents the input feature map. The input feature map of the first-layer deformable convolution is the normalized video block X, and the input feature map of the second-layer deformable convolution is the output result of the first-layer deformable convolution.

7. The lightweight real-time heart rate monitoring model according to claim 1, wherein The motion compensation branch uses optical flow estimation to obtain a motion offset field, and further realizes motion compensation.

8. The lightweight real-time heart rate monitoring model according to claim 1, characterized in that, The feature fusion unit combines the final output result of the spatio-temporal feature branch with the final output result of the motion compensation branch to perform splicing and dimensionality reduction through 1×1 convolution, and obtain and output features related to heart rate prediction.

9. The lightweight real-time heart rate monitoring model according to claim 1, wherein The heart rate calculation module obtains a preliminary heart rate prediction value based on the heart rate prediction relevant features output by the lightweight dual-branch network module, filters the preliminary heart rate prediction value through a second-order Butterworth low-pass filter, and performs peak detection on the filtered signal, and calculates the heart rate according to the time interval between adjacent peaks to obtain a readable and usable predicted heart rate.

10. The lightweight real-time heart rate monitoring model according to claim 1, wherein The heart rate loss uses the negative Pearson correlation coefficient to measure the correlation between the final predicted heart rate and the true heart rate. The heart rate loss has the following expression: ; Among them, represents the true heart rate of the th training sample, represents the predicted heart rate of the th training sample, represents the number of samples for model training, represents the average value of the true heart rate, represents the average value of the predicted heart rate; The noise suppression loss has the following expression: ; Among them, represents the energy of the predicted heart rate at the frequency ; The regularization loss adopts an L2 regularization term, and the regularization loss has the following expression: ; Among them, represents the regularization coefficient, represents the set of parameters of the model, represents any parameter in the model.

11. A heart rate monitoring method, characterized in that, Using the lightweight real-time heart rate monitoring model according to any one of claims 1 to 10 to perform heart rate detection and obtain the predicted heart rate.

Citation Information

Patent Citations

  • Non-contact heart rate measurement method based on space-time attention network and input optimization

    CN113343821A

  • Lightweight video super-resolution reconstruction method based on hybrid space-time convolution

    CN117830095A

  • Method for multi-motion flow deep convolutional network model for video prediction

    WO2020037965A1

  • Video encoding and decoding method and apparatus, device, system and storage medium

    WO2023206420A1