Contactless heart rate measurement method, system, apparatus, and storage medium

By extracting feature maps in RGB and YCbCr spaces and combining them with encoder and transformer networks, the problem of background noise interference was solved, and high-accuracy detection of non-contact heart rate measurement was achieved.

CN117058569BActive Publication Date: 2025-11-11SOUTH CHINA NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310826559.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-06
Publication Date
2025-11-11
Estimated Expiration
2043-07-06

AI Technical Summary

Technical Problem

Existing non-contact heart rate measurement methods cannot effectively reduce the interference of background noise on rPPG signal extraction, resulting in low accuracy of heart rate detection results.

Method used

By extracting feature maps from RGB and YCbCr spaces respectively, and combining an RGB encoder and a visual transformer network, the rPPG feature information is extracted using the characteristics of RGB and YCbCr spaces respectively. The target rPPG feature information is then weighted and synthesized on each channel, and finally the heart rate is extracted through Fourier transform and bandpass filter.

Benefits of technology

It effectively reduces redundant information interference between video images, suppresses the influence of noise, and improves the accuracy of non-contact heart rate detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117058569B_ABST
    Figure CN117058569B_ABST
Patent Text Reader

Abstract

This invention discloses a non-contact heart rate measurement method, system, device, and storage medium, applicable to the field of heart rate monitoring technology. The invention extracts RGB and YCbCr spatial sample data from video data within a dataset, thereby obtaining the overall RGB and YCbCr spatial feature maps. Based on these maps, corresponding rPPG feature information is extracted, and the target rPPG feature information is obtained. This leverages the characteristics of RGB space to ensure temporal and spatial information extraction redundancy, reducing interference from significant redundancy in video images. Simultaneously, the characteristics of YCbCr space effectively suppress noise, thus improving the accuracy of non-contact heart rate measurement and enhancing the overall accuracy of the results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of heart rate monitoring technology, and in particular to a non-contact heart rate measurement method, system, device, and storage medium. Background Technology

[0002] Among related technologies, heart rate is one of the most important physiological indicators, playing a crucial role in assessing heart health. Traditional heart rate measurement requires the subject to wear the measuring device for extended periods, which is inconvenient for real-time monitoring of the user's health and causes considerable discomfort. Furthermore, typical measuring devices are bulky and expensive, resulting in high measurement costs. Therefore, non-contact heart rate monitoring based on facial video offers significant advantages. This monitoring method requires only a consumer-grade camera to extract heart rate from changes in optical information on the skin surface, effectively overcoming the drawbacks of contact-based heart rate measurement and offering portability and non-invasiveness. However, while existing non-contact heart rate measurement methods employ deep learning models for heart rate detection, this method cannot effectively reduce background noise interference in rPPG (Remote Photoplethysmography) signal extraction, thus failing to improve the accuracy of heart rate detection results. Summary of the Invention

[0003] The present invention aims to at least solve one of the technical problems existing in the prior art. To this end, the present invention proposes a non-contact heart rate measurement method, system, device, and storage medium, which can effectively improve the accuracy of non-contact heart rate detection results.

[0004] On one hand, embodiments of the present invention provide a non-contact heart rate measurement method, comprising the following steps:

[0005] Obtain a dataset of the target object, which includes facial videos and physiological signal labels;

[0006] The dataset is subjected to a first preprocessing step to obtain RGB space sample data;

[0007] The RGB space sample data is subjected to a second preprocessing to obtain YCbCr space sample data;

[0008] The feature data of the RGB space sample data are extracted to form a total RGB space feature map;

[0009] The feature data of the YCbCr spatial sample data are extracted to form the overall feature map of YCbCr space;

[0010] After inputting the total feature map of the RGB space into the RGB encoder and RGB decoder, the rPPG feature information of the RGB space is obtained as the first rPPG feature information;

[0011] The total feature map of the YCbCr space is input into the visual transformer network to obtain the rPPG feature information of the YCbCr space as the second rPPG feature information;

[0012] The first rPPG feature information is added to the second rPPG feature information on each channel to obtain the target rPPG feature information;

[0013] The measured heart rate of the target object is extracted from the target rPPG feature information.

[0014] In some embodiments, the first preprocessing of the dataset to obtain RGB space sample data includes:

[0015] Each frame of the video in the dataset is scaled to a preset resolution to obtain the video data to be processed.

[0016] The video data to be processed is divided into video segments using a preset sliding window as RGB space sample data.

[0017] In some embodiments, the second preprocessing of the RGB space sample data to obtain YCbCr space sample data includes:

[0018] Face tracking is performed on each frame of the RGB space sample data, and the faces are cropped.

[0019] Select several regions of interest in each cropped frame of the face image;

[0020] The average pixel count is obtained by averaging all pixels in each region of interest across the R, G, and B channels.

[0021] The average pixel value is converted to the YCbCr space, and the converted pixels of each frame are combined in the Y, Cb, and Cr channels to obtain YCbCr space sample data.

[0022] In some embodiments, extracting feature data from the RGB space sample data to form a total RGB space feature map includes:

[0023] The RGB space sample data is input into the first two-dimensional neural network to obtain the first two-dimensional feature map;

[0024] The RGB space sample data is input into the first three-dimensional neural network to obtain the first three-dimensional feature map;

[0025] Add the first two-dimensional feature map and the first three-dimensional feature map to obtain a preliminary RGB spatial data sample feature map;

[0026] The initial RGB spatial data sample feature map is input into the residual attention module to obtain the total RGB spatial feature map.

[0027] In some embodiments, the residual attention module includes a first convolutional layer, a second convolutional layer, an attention module, and a residual module. The attention module includes a channel attention module and a spatial attention module. The step of inputting the preliminary RGB spatial data sample feature map into the residual attention module to obtain the total RGB spatial feature map includes:

[0028] The initial RGB spatial data sample feature map is processed sequentially through the first convolutional layer and the second convolutional layer to obtain the first input feature map;

[0029] The first feature map to be input is input into the channel attention module to obtain the channel attention feature map;

[0030] The channel attention feature map is multiplied with the first input feature map to obtain the second input feature map;

[0031] The second feature map to be input is input into the spatial attention module to obtain the spatial attention feature map;

[0032] The spatial attention feature map is multiplied with the second input feature map to obtain the third input feature map;

[0033] The third feature map to be input and the preliminary RGB space data sample feature map are input into the residual module and added together to obtain the total RGB space feature map.

[0034] In some embodiments, extracting feature data from the YCbCr spatial sample data to form a total YCbCr spatial feature map includes:

[0035] The YCbCr spatial sample data is input into a second two-dimensional neural network to obtain a second two-dimensional feature map;

[0036] The YCbCr spatial sample data is input into the second three-dimensional neural network to obtain the second three-dimensional feature map;

[0037] The second two-dimensional feature map and the third three-dimensional feature map are added together to obtain the first YCbCr spatial data sample feature map;

[0038] Input the first YCbCr spatial data sample feature map into the residual attention module, and the second YCbCr spatial data sample feature map;

[0039] The second YCbCr spatial data sample feature map is input into the adaptive average pooling layer to obtain the total YCbCr spatial feature map.

[0040] In some embodiments, extracting the measured heart rate of the target object from the target rPPG feature information includes:

[0041] The target rPPG feature information is input into a bandpass filter to obtain the rPPG feature information to be converted within a preset frequency range;

[0042] Perform a Fourier transform on the rPPG feature information to be converted to obtain the power spectral density;

[0043] Determine the peak frequency corresponding to the peak value from the power spectral density;

[0044] The peak frequency is multiplied by a preset value to obtain the measured heart rate of the target object.

[0045] On the other hand, embodiments of the present invention provide a non-contact heart rate measurement system, comprising:

[0046] The first module is used to acquire a dataset of the target object, which includes facial videos and physiological signal labels;

[0047] The second module is used to perform a first preprocessing on the dataset to obtain RGB space sample data;

[0048] The third module is used to perform a second preprocessing on the RGB space sample data to obtain YCbCr space sample data;

[0049] The fourth module is used to extract feature data from the RGB space sample data to form a total RGB space feature map;

[0050] The fifth module is used to extract feature data from the YCbCr spatial sample data to form the overall YCbCr spatial feature map;

[0051] The sixth module is used to input the total feature map of the RGB space into the RGB encoder and RGB decoder to obtain the rPPG feature information of the RGB space as the first rPPG feature information;

[0052] The seventh module is used to input the total feature map of the YCbCr space into the visual transformer network to obtain the rPPG feature information of the YCbCr space as the second rPPG feature information.

[0053] The eighth module is used to add the second rPPG feature information to the first rPPG feature information on each channel to obtain the target rPPG feature information;

[0054] The ninth module is used to extract the measured heart rate of the target object from the target rPPG feature information.

[0055] On the other hand, embodiments of the present invention provide a non-contact heart rate measurement device, comprising:

[0056] At least one memory for storing programs;

[0057] At least one processor is used to load the program to execute the non-contact heart rate measurement method.

[0058] On the other hand, embodiments of the present invention provide a computer storage medium storing a computer-executable program, which, when executed by a processor, is used to implement the non-contact heart rate measurement method.

[0059] This invention provides a non-contact heart rate measurement method, which has the following beneficial effects:

[0060] This embodiment extracts RGB and YCbCr space sample data from the video data within the dataset, thereby obtaining the RGB and YCbCr space total feature maps respectively. Based on these RGB and YCbCr space feature maps, corresponding rPPG feature information is extracted. The target rPPG feature information is then obtained based on this rPPG feature information. This approach leverages the characteristics of the RGB space to ensure temporal and spatial information extraction redundancy, reducing interference from significant redundancy in the video images. Simultaneously, the characteristics of the YCbCr space effectively suppress noise, thus improving the accuracy of non-contact heart rate measurement and enhancing the overall accuracy of the non-contact heart rate detection results.

[0061] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0062] The present invention will be further described below with reference to the accompanying drawings and embodiments, wherein:

[0063] Figure 1 This is a flowchart of a non-contact heart rate measurement method according to an embodiment of the present invention;

[0064] Figure 2 This is a flowchart illustrating the extraction of the total feature map in RGB space according to an embodiment of the present invention;

[0065] Figure 3 This is a schematic diagram of the structure of a residual attention module according to an embodiment of the present invention;

[0066] Figure 4 This is a schematic diagram of the structure of an attention module according to an embodiment of the present invention;

[0067] Figure 5 This is a schematic diagram of the structure of a channel attention module according to an embodiment of the present invention;

[0068] Figure 6 This is a schematic diagram of the structure of a spatial attention module according to an embodiment of the present invention;

[0069] Figure 7 This is a flowchart illustrating the extraction process of the total spatial feature map of YCbCr according to an embodiment of the present invention;

[0070] Figure 8 This is a flowchart illustrating the extraction of first rPPG feature information according to an embodiment of the present invention.

[0071] Figure 9 This is a schematic diagram of the structure of an RGB encoder according to an embodiment of the present invention;

[0072] Figure 10 This is a schematic diagram of the structure of an RGB decoder according to an embodiment of the present invention;

[0073] Figure 11 This is a flowchart illustrating the extraction of second rPPG feature information according to an embodiment of the present invention.

[0074] Figure 12 This is a schematic diagram of the structure of a visual transformer network according to an embodiment of the present invention;

[0075] Figure 13 This is a flowchart of a heart rate extraction process according to an embodiment of the present invention. Detailed Implementation

[0076] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0077] In the description of this invention, it should be understood that the orientation descriptions, such as up, down, front, back, left, right, etc., are based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting this invention.

[0078] In the description of this invention, "several" means one or more, "multiple" means two or more, "greater than," "less than," and "exceeding" are understood to exclude the stated number, while "above," "below," and "within" are understood to include the stated number. The use of "first" and "second" in the description is merely for distinguishing technical features and should not be construed as indicating or implying relative importance, or implicitly indicating the number of indicated technical features, or implicitly indicating the order of the indicated technical features.

[0079] In the description of this invention, unless otherwise explicitly defined, terms such as "setting," "installing," and "connecting" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this invention in conjunction with the specific content of the technical solution.

[0080] In the description of this invention, the terms "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0081] Before describing specific embodiments, the terms used in this application are explained as follows:

[0082] RGB space: also known as RGB color space, is based on three primary colors: R (Red), G (Green), and B (Blue). Different levels of superposition are used to produce a rich and wide range of colors, so it is commonly known as the three primary color mode.

[0083] YCbCr color space: also called YCbCr color space. In the YCbCr color space, Y is the luminance channel, Cb is the blue component, and Cr is the red component.

[0084] rPPG, short for Remote Photoplethysmography, measures subtle changes in skin brightness by using reflected ambient light. These subtle changes in skin brightness are caused by blood flow due to heartbeats. rPPG typically produces a signal similar to BVP (Body Velocity Spectrometry), which can be used to predict information such as heart rate and respiratory rate.

[0085] Among related technologies, heart rate is one of the most important physiological indicators, playing a crucial role in assessing heart health. Traditional heart rate measurement requires the subject to wear the device for extended periods, which is inconvenient for real-time monitoring of the user's health and causes considerable discomfort. Furthermore, typical measuring devices are bulky, expensive, and costly. Therefore, non-contact heart rate monitoring based on facial video offers significant advantages. This monitoring method requires only a consumer-grade camera to extract heart rate from changes in optical information on the skin's surface. It effectively overcomes the drawbacks of contact-based heart rate measurement, offering portability and non-invasive characteristics.

[0086] Early non-contact heart rate measurement methods involved the following steps: detecting faces in videos and extracting Regions of Interest (ROIs); extracting physiological signals, such as independent component analysis based on blind source separation algorithms and orthogonal skin frontal analysis based on optical reflection models; processing the physiological signals, such as detrending and bandpass filtering; and estimating the heart rate by performing Fourier transform on the extracted physiological signals to the frequency domain, calculating the frequency of the peak, or performing peak detection on the physiological signals, ultimately obtaining the heart rate signal.

[0087] The creation of large public datasets and the increased computing power of GPUs have provided opportunities for deep learning to perform rPPG heart rate detection tasks. Convolutional Neural Networks (CNNs) are among the most commonly used models; for example, the HR-CNN framework for heart rate measurement shows a significant improvement in detection performance compared to traditional models. Furthermore, attention mechanisms are widely used in deep learning, such as the adaptive "Squeeze-and-Excitation (SE)" attention mechanism and spatial and channel attention constructed using a "bottom-up top-down" technique. By employing attention mechanisms, deep learning models can mimic the human brain's tendency to prioritize regions of interest when processing massive amounts of data. One example is an rPPG extraction network utilizing spatial attention. This network consists of two modules: a Motion Model and an Appearance Model. Guiding the learning of the Appearance Model through the Motion Model helps improve the accuracy of rPPG signal extraction. The proposed SynRhythm and RhythmNet models generate spatiotemporal graph representations of facial regions of interest from multiple regions, and then use these spatiotemporal graph representations to predict heart rate. Compared with traditional methods, these deep learning models have significantly improved the accuracy of heart rate detection, but they are still subject to some limitations. For example, they can handle head movements and changes in illumination intensity, ignore the influence of redundant information in long video sequences, cannot fully utilize the contextual information in the video sequence, and cannot effectively reduce the interference of background noise on rPPG signal extraction, thus failing to improve the accuracy of heart rate detection results.

[0088] Reference Figure 1 This invention provides a non-contact heart rate measurement method, which can be applied to the processing end, server, or cloud of a non-contact heart rate measurement platform.

[0089] In application, the method of this embodiment includes, but is not limited to, the following steps:

[0090] Step S110: Obtain the dataset containing face videos and physiological signal labels corresponding to the target object; perform the first preprocessing on the dataset to obtain RGB space sample data;

[0091] In this embodiment, a dataset containing face videos and physiological signal labels can be obtained. It is understood that if the current processing directly detects the target object, the dataset is the dataset corresponding to that target object; if the current processing trains a heart rate detection model, the dataset can be the PURE, UBFC-rPPG, or VIPL-HR datasets. After obtaining the corresponding dataset, each frame of the video within the dataset can be scaled to a preset resolution to obtain the video data to be processed. Then, the video data to be processed is divided into video segments using a preset sliding window as RGB space sample data. For example, the physiological signals in the dataset are resampled using the video frame rate as the sampling rate, so that the length of the resampled physiological signals is equal to the number of frames in each face video. Each frame of the video is aligned with the physiological signals, and then each frame in the aligned video is scaled to 180*180 pixels. For the RGB space, we only need to divide each face video into separate video segments using a sliding window of x seconds and a window length of T frames, generating v1, v2...vn video segments and saving them. These separate video segments are then used as RGB space sample data.

[0092] Step S120: Perform a second preprocessing on the RGB space sample data to obtain YCbCr space sample data;

[0093] In this embodiment, face tracking can be performed on each frame of the RGB space sample data, and the face can be cropped. Then, several regions of interest can be selected in each cropped face image, and all pixels in each region of interest can be averaged in the R, G, and B channels to obtain the average pixel. Then, the average pixel is converted to the YCbCr space, and the pixels obtained from the conversion in each frame are combined in the Y, Cb, and Cr channels to obtain the YCbCr space sample data.

[0094] Understandably, this embodiment requires importing a cascaded face tracker from the cv2 library in Python to perform face tracking on each frame of the RGB video clip in the YCbCr space, and cropping the faces. If no face is detected in a frame, that frame is replaced with a black background. Then, for each cropped frame, m regions of interest (ROIs) are selected, and all pixels in each ROI are averaged across the RGB three channels to obtain the average pixel value. This average pixel value is then converted to the YCbCr space, and each frame is combined across the Y, Cb, and Cr channels to obtain the final YCbCr data sample. Each data sample is then subjected to min-max normalization and scaled to the range [0,1]. The size of each data sample is T*3*m (T is the window length, 3 is the three channels of Y, Cb, and Cr, and m is the number of ROIs). The number of data samples in both the YCbCr space and the RGB space is n.

[0095] In this embodiment, the expression for converting a video clip from RGB space to YCbCr space is shown in formula (1):

[0096] Formula (1).

[0097] For the preprocessing of the physiological signal labels contained in the dataset in this application embodiment, the physiological signals in the dataset can be resampled at the video frame rate as the sampling rate, so that the length of the sampled physiological signal is equal to the number of frames in each face video, so that each frame of the video is aligned with the physiological signal. Then, the label signal is also divided into segments p1, p2, ..., pn using a sliding window with a step size of x seconds and a window length of T, and each segment of the label signal is normalized. Specifically, the sampling min-max normalization method is shown in formula (2):

[0098] Formula (2)

[0099] Where x represents any tag signal, This represents the tag signal after normalization.

[0100] Step S130: Extract feature data from RGB space sample data to form a total feature map in RGB space;

[0101] In the embodiments of this application, such as Figure 2As shown, RGB space sample data can be input into a first two-dimensional neural network to obtain a first two-dimensional feature map; simultaneously, RGB space sample data can be input into a first three-dimensional neural network to obtain a first three-dimensional feature map; then, the first two-dimensional feature map and the first three-dimensional feature map are added together to obtain a preliminary RGB space data sample feature map. This preliminary RGB space data sample feature map is then input into a residual attention module to obtain a total RGB space feature map. Specifically, as... Figure 3 As shown, the residual attention module includes a first convolutional layer, a second convolutional layer, an attention module, and a residual module. Figure 4 As shown, the attention module includes a channel attention module and a spatial attention module. Based on Figure 3 and Figure 4 The network structure diagram shows that the process of inputting the initial RGB space data sample feature map into the residual attention module to obtain the total RGB space feature map may include, but is not limited to, the following steps:

[0102] After the initial RGB spatial data sample feature map is processed sequentially through the first convolutional layer and the second convolutional layer, the first input feature map is obtained.

[0103] The first feature map to be input is input into the channel attention module to obtain the channel attention feature map;

[0104] The channel attention feature map is multiplied with the first input feature map to obtain the second input feature map;

[0105] The second feature map to be input is input into the spatial attention module to obtain the spatial attention feature map;

[0106] The spatial attention feature map is multiplied with the second input feature map to obtain the third input feature map;

[0107] The third feature map to be input and the initial RGB space data sample feature map are input into the residual module for addition to obtain the total RGB space feature map.

[0108] Understandably, the input data samples in RGB space have a shape of [B, C, T, W, H], where B represents the batch size; C represents the number of image channels (RGB space has 3 channels, so C is 3); T represents the number of video frames per sample, based on preprocessing settings; W and H represent the width and height, scaled from 180*180 to 90*90, with a data shape of [B, 3, T, 90, 90]. For a given RGB space data sample, it is first simultaneously input into a 3D convolutional neural network with two convolutional blocks (the first 3D neural network) and a 2D convolutional neural network with two convolutional blocks (the first 2D neural network). In the first 3D neural network, the first convolutional block contains a 3D convolutional layer with a kernel of [3,3,3], 3 input channels, 16 output channels, a stride of [1,1,1], and padding of [1,1,1], plus a BatchNorm3D layer with 16 input and output channels, and a ReLU activation layer. The output feature map shape is [B,16,T,90,90]. The second convolutional block contains a 3D convolutional layer with a kernel of [3,3,3], 16 input and 16 output channels, a stride of [1,1,1], and padding of [1,1,1], plus a BatchNorm3D layer with 16 input and output channels, and a ReLU activation layer. Finally, the feature map size obtained from the first 3D convolutional network is [B,16,T,90,90]. When input into the first two-dimensional neural network, the first two dimensions of the data are merged, resulting in a shape of [B*T,C,W,H], which is [B*T,3,90,190]. In the first two-dimensional neural network, the first convolutional block contains a two-dimensional convolutional layer with a kernel of [3,3], 3 input channels, 16 output channels, a stride of [1,1], and padding of [1,1]. This is followed by a BatchNorm2D layer with 16 input and output channels and a ReLU activation layer. The output feature map has a shape of [B*T,16,90,90]. The second convolutional block contains a two-dimensional convolutional layer with a kernel of [3,3], 16 input channels, 16 output channels, a stride of [1,1], and padding of [1,1]. This is followed by a BatchNorm2d layer with 16 input and output channels and a ReLU activation layer. The final feature map obtained from the 2D-CNN network is [B*T,16,90,90], which is then reshaped to [B,16,T,90,90]. The first two-dimensional feature map obtained from the first two-dimensional neural network and the first three-dimensional feature map obtained from the first three-dimensional neural network are added to obtain the preliminary RGB spatial data sample feature map, with the shape [B,16,T,90,90]. Then, the preliminary RGB spatial data sample feature map is input into... Figure 3The residual CBAM (attention mechanism) module shown first modifies the feature map shape to [B*T, 16, 90, 90], then inputs it into the first convolutional layer, which consists of 2D convolution (16 input and output channels, kernel [3, 3], stride [1, 1], padding [1, 1]), BatchioNorm2D (16 input and output channels), and the ReLU activation function. The resulting shape is [B*T, 16, 90, 90]. This is then input into the second convolutional layer, which consists of 2D convolution (16 input and output channels, kernel [3, 3], stride [1, 1], padding [1, 1]) and BatchioNorm2D (16 input and output channels). The output shape is [B*T, 16, 90, 90], representing the first input feature map. Finally, the first input feature map is input into the CBAM module (attention mechanism module). Figure 4 As shown, it consists of a channel attention module and a spatial attention module. The first feature map to be input obtained in the previous step has a shape of [B*T, 16, 90, 90]. Each feature map F has a shape of [16, 90, 90], and there are a total of B*T feature maps. Each first feature map F to be input is sequentially passed through the channel attention module ( Figure 5 (as shown) and spatial attention module ( Figure 6 As shown, the shape of the first input feature map F is modified to [90, 90, 16]. It is then subjected to global max pooling and global average pooling on the height and width, respectively, resulting in two 1×1×16 feature maps. These are then fed into a two-layer neural network (MLP). The first layer has 16 neurons / r (r is the reduction rate) and uses ReLU activation. The second layer has 16 neurons, and the two layers are shared. The MLP output features are then summed element-wise, followed by sigmoid activation to generate the final channel attention feature map. (1×1×16). Finally, the channel attention feature map... Perform element-wise multiplication with the first input feature map F to generate the second input feature map needed by the subsequent spatial attention module. (90×90×16).

[0109] The specific formula is shown in formula (3):

[0110] Formula (3)

[0111] In formula (3), AvgPool represents average pooling, MaxPool represents max pooling, and F represents the first input feature map. This represents the channel attention feature map. MLP represents a two-layer neural network. This represents the sigmoid activation function.

[0112] In this embodiment, the spatial attention mechanism process is as follows: For the aforementioned second input feature map... Two (90×90×16) feature maps are obtained by performing global average pooling and global max pooling along the channel dimension. Then, a 7*7 convolution kernel is used for convolution, followed by sigmoid activation to obtain the spatial attention feature map. (90×90×1), finally compared with the second feature map to be input. Performing element-wise multiplication yields the optimized third input feature map. (90×90×16).

[0113] The specific formula is shown in formula (4):

[0114] Formula (4)

[0115] Formula (4), where F represents the first input feature map, AvgPool represents average pooling, and MaxPool represents max pooling. This indicates a convolutional layer with a kernel size of 7x7. Representing the sigmoid activation function, we obtain a feature map with the shape [B*T,16,90,i]. This map is then modified to [B,T,16,90,90]. Through the residual structure, we add the third input feature map obtained in the previous step and the initial RGB space data sample feature map initially input to the residual CBAM module to obtain the final RGB space total feature map with the shape [B,T,16,90,90]. This completes the feature extraction of the RGB space data samples.

[0116] Step S140: Extract feature data from YCbCr spatial sample data to form the overall feature map of YCbCr space;

[0117] In the embodiments of this application, such as Figure 7 As shown, YCbCr spatial sample data can be input into a second two-dimensional neural network to obtain a second two-dimensional feature map; simultaneously, YCbCr spatial sample data can be input into a second three-dimensional neural network to obtain a second three-dimensional feature map; then, the second two-dimensional feature map and the third-dimensional feature map are added together to obtain a first YCbCr spatial data sample feature map; then, the first YCbCr spatial data sample feature map is input into a residual attention module, and the second YCbCr spatial data sample feature map is input into an adaptive average pooling layer to obtain a total YCbCr spatial feature map.

[0118] It is understandable that for a given YCbCr spatial sample data, its data shape is [B,C,T,m], that is, [B,3,T,m]. m is generally a number that can be square rooted, such as 9, 16, 25, etc. In this embodiment, m is chosen as 25. Therefore, the data shape can be changed to [B,3,T,i,i], where m=i*i, that is, [B,3,T,5,5]. YCbCr spatial sample data are simultaneously input into a second two-dimensional neural network and a second three-dimensional neural network. Before being input into the second two-dimensional neural network, the first two dimensions of the data are merged into [B*T,3,5,5]. The second two-dimensional neural network contains two convolutional blocks. The first convolutional block contains a two-dimensional convolutional layer with 3 input channels, 16 output channels, a kernel of [3,3], a stride of [1,1], and padding of [1,1], a BatchNorm2D layer with 16 input and output channels, and a ReLU activation layer with an output shape of [B*T,16,i,i]. The second convolutional block contains a two-dimensional convolutional layer with 16 input channels, 16 output channels, a kernel of [3,3], a stride of [1,1], and padding of [1,1], a BatchNorm2D layer with 16 input and output channels, and a ReLU activation layer with an output shape of [B*T,16,5,5], which is a 2D-CNN feature map. Before being input into the second 3D neural network, the data samples are modified to [B,3,T,5,5] by matrix transpose. The second 3D neural network contains two convolutional blocks. The first convolutional block contains a 3D convolutional layer with 3 input channels and 16 output channels, a kernel of [3,3,3], a stride of [1,1,1], and padding of [1,1,1], a BatchNorm3D layer with 16 input and output channels, and a ReLU activation layer with an output shape of [B,T,16,5,5]. The second convolutional block contains a 3D convolutional layer with 16 input channels and 16 output channels, a kernel of [3,3,3], a stride of [1,1,1], and padding of [1,1,1], a BatchNorm3D layer with 16 input and output channels, and a ReLU activation layer with an output shape of [B,16,T,i,i], which is a 3D-CNN feature map. The shape of the second 2D feature map is modified to [B,16,T,5,5], and then added to the 3D-CNN feature map to obtain the preliminary feature map of the YCbCr data sample, with the shape [B,16,T,5,5]. The second 3D feature map is then input into the residual CBAM (attention mechanism) module, as follows: Figure 3As shown, the feature map shape is first modified to [B*T,16,5,5], and then input into the first layer, which consists of 2D convolution (16 input and output channels, kernel [3,3], stride [1,1], padding [1,1]), BatchoNorm2D (16 input and output channels), and ReLU activation function. The resulting shape is [B*T,16,5,5]. Then, it is input into the second layer, which consists of 2D convolution (16 input and output channels, kernel [3,3], stride [1,1], padding [1,1]) and BatchoNorm2D (16 input and output channels). The output shape is [B*T,16,5,5]. Finally, it is input into the CBAM module (attention mechanism module), as shown in the diagram. Figure 4 As shown, it consists of a channel attention module ( Figure 5 (as shown) and spatial attention module ( Figure 6 The input consists of the feature maps obtained in the previous step, with a shape of [B*T, 16, 5, 5]. Each feature map F has a shape of [16, 5, 5], and there are a total of B*T feature maps. Each feature map F is sequentially processed through channel attention and spatial attention mechanisms to modify its shape to [5, 5, 16]. It then undergoes global max pooling and global average pooling on the height and width sides, respectively, resulting in two 1×1×16 feature maps. These are then fed into a two-layer neural network (MLP). The first layer has 16 neurons / r (r is the reduction rate) and uses ReLU activation. The second layer has 16 neurons, and this two-layer neural network is shared. The MLP output features are then summed element-wise, followed by sigmoid activation to generate the final channel attention features. (1×1×16). Finally, Perform element-wise multiplication with the input feature map F to generate the input features needed by the subsequent spatial attention module. (5×5×16).

[0119] The specific formula is shown in formula (5):

[0120] Formula (5)

[0121] Formula (5), where AvgPool represents average pooling, MaxPool represents max pooling, F represents the input feature map, and M represents the channel attention map. MLP represents a two-layer neural network. This represents the sigmoid activation function.

[0122] In this embodiment, the spatial attention mechanism process is as follows: [The process involves processing the aforementioned feature maps.] (5×5×16) Global average pooling and global max pooling are performed along the channel dimension to obtain two (i×i×16) feature maps. Then, a 7*7 convolution kernel is used for convolution, followed by sigmoid activation to obtain a two-dimensional spatial attention map. (5×5×1), finally compared with the input feature map Element-wise multiplication is performed to obtain the optimized feature map. (5×5×16).

[0123] The specific formula is shown in formula (6):

[0124] Formula (6)

[0125] In formula (6), F represents the input feature map, AvgPool represents average pooling, and MaxPool represents max pooling. This indicates a convolutional layer with a kernel size of 7x7. The sigmoid activation function is used to obtain a feature map with shape [B*T,16,5,5]. This map is then modified to [B,T,16,5,5]. Using a residual structure, the feature map obtained in the previous step is added to the feature map initially input into the residual CBAM module to obtain a further feature map with shape [B,T,16,5,5]. Finally, this feature map is input into an adaptive average pooling layer. The adaptive average pooling layer outputs the final YCbCr total feature map with shape [B,T,16,4,4], thus completing the feature extraction of the YCbCr spatial data samples.

[0126] Step S150: After inputting the total feature map of RGB space into the RGB encoder and RGB decoder, the rPPG feature information of RGB space is obtained as the first rPPG feature information;

[0127] In the embodiments of this application, such as Figure 8 As shown, the shape of the total RGB spatial feature map is [B,T,16,90,90]. Before being input into the RGB encoder, the shape of the RGB spatial feature map is modified to [B,16,T,90,90] through matrix transpose. The RGB encoder (e.g.) Figure 9The first layer is a 3D average pooling layer with a kernel size of [2,2,2], a stride of [2,2,2], and zero padding. The output shape is [B,16,T / 2,45,45]. This is then fed into the second layer, a 3D convolutional block containing 3D convolution, batch processing (BatchNorm3d), and the ELU activation function. The 3D convolution has 16 input channels and 32 output channels, with a kernel size of [3,3,3], a stride of [1,1,1], and zero padding. This is then passed through a batch pooling layer. The nNorm3d layer has 32 input and output channels. After passing through the ELU activation function, the output shape is [B, 32, T / 2, 45, 45]. The result from the previous step is then fed into the third layer, which is a 3D average pooling layer with a kernel size of [2, 2, 2], a stride of [2, 2, 2], and zero padding, resulting in an output shape of [B, 32, T / 4, 22, 22]. This output is then fed into the fourth layer of the RGB encoder, which is also a 3D convolutional block with 32 input channels and 64 output channels. The convolution kernel size is [3,3,3], stride is [1,1,1], and padding is [1,1,1]. BatchNorm3d has 64 input and output channels. Finally, the activation function ELU is applied, and the output shape is [B,64,T / 4,11,11]. This is then fed into the fifth layer of the RGB encoder, which is a 3D average pooling layer with a kernel size of [1,2,2], stride of [1,2,2], and padding of 0, resulting in an output shape of [B,64,T / 4,5,5]. This is then fed into the R... The sixth and seventh layers of the GB encoder are both 3D convolutional blocks. The sixth layer's 3D convolution has 64 input and output channels, a kernel size of [3,3,3], a stride of [1,1,1], and padding of [1,1,1]. Then comes BatchNorm3d, with 64 input and output channels, followed by the ELU activation function. The structure and input parameters of the sixth and seventh layers are identical, resulting in an RGB encoder output with a shape of [B,64,T / 4,5,5], completing the RGB space data sample encoding. The encoded result of the RGB encoder is then input into the RGB decoder (e.g., ...). Figure 10The first layer of the RGB decoder is a 3D transposed convolutional block, containing a 3D transposed convolution, BatchNorm3d, and the activation function ELU. The 3D transposed convolution has 64 input and output channels, a kernel of [4,1,1], a stride of [2,1,1], and padding of [1,0,0]. The shape of the result after the 3D transposed convolution is [B,64,T / 2,5,5]. BatchNorm3d has 64 input and output channels, and then after passing through the activation function ELU, the output shape is [B,64,T / 2,5,5]. Then, the input is fed into the second layer, which is also a 3D transposed convolutional block. The input and output channels of the 3D transposed convolution are both 64, the kernel is [4,1,1], the stride is [2,1,1], and the padding is [1,0,0]. The shape of the result obtained after the 3D transposed convolution is [B,64,T,5,5]. The input and output channels of BatchNorm3d are both 64. Then, after passing through the activation function ELU, the final output shape is [B,64,T,5,5]. The rPPG feature information in RGB space is obtained as the first rPPG feature information.

[0128] Step S160: Input the total feature map of YCbCr space into the visual transformer network to obtain the rPPG feature information of YCbCr space as the second rPPG feature information;

[0129] In the embodiments of this application, such as Figure 11 As shown, the shape of the total feature map in the YCbCr space is [B,T,16,4,4]. It is input into the Vision Transformer (ViT, a visual transformer network; the network model is as follows...) Figure 12 In the ViT model configuration of this invention, the input, output size, and process of each part are as follows:

[0130] 1. YCbCr total feature map: size is (B, T, 16, 4, 4), B is Batchsize, T is the number of frames, 16 is the number of channels, and (4, 4) is the width and height.

[0131] 2. Patch Division: This invention sets patch_size=2. Based on patch_size=2, the input image is divided into 2x2 blocks (patches). For each batch of 256 images, T / 2 2x2 patches will be obtained. Therefore, the output size after patch division is (B, T, T / 2, 2, 2, 16).

[0132] 3. Embedding Layer: This layer converts each patch into an embedding vector. The tensor input to the embedding layer has a shape of (B, T, T / 2, 2, 2, 16). The embedding layer converts each 2x2 patch into an embedding vector of size num_hiddens, which is chosen to be 128 in this implementation. Therefore, the output size of the embedding layer is (B, T, T / 2, 128).

[0133] 4. Positional Embedding: Positional embeddings encode the positional information of each patch by adding positional embedding vectors. The shape of the positional embedding vector is (1, T / 2, num_hiddens=128), where 1 represents a batch. The positional embedding vector is added to the output of the embedding layer to obtain the final input vector with the shape (B, T, T / 2, num_hiddens=128).

[0134] 5. Transformer Encoder: The Transformer Encoder encodes the input vector. It consists of multiple Transformer encoder layers, each containing a multi-head self-attention mechanism and a feedforward neural network. The multi-head attention mechanism captures dependencies between global and local features by computing attention weights at each position in the input sequence. The feedforward neural network layer is used for feature mapping and non-linear transformations. After each Transformer encoder block, a block-level dropout operation is applied to enhance the model's robustness. In the given model configuration, eight encoder layers are used. The output shape of the Transformer encoder is the same as the input, i.e., (B, T, T / 2, num_hiddens=128). The principle and process of the multi-head attention mechanism: The multi-head attention mechanism allows the model to simultaneously focus on information from different positions in the input sequence to capture dependencies between global and local features. In ViT, the multi-head attention mechanism is applied to the self-attention layer of each Transformer encoder block. The self-attention layer divides the input sequence into multiple heads (num_heads=8, 8 in this invention), and each head learns a set of attention weights. The attention weights for each head are calculated using three input vectors: a query vector (Q), a key vector (K), and a value vector (V). The query vector determines the location of interest, while the key and value vectors represent features in the sequence. By calculating the similarity between the query vector and the key vector, the attention weights of each location relative to other locations can be obtained. The multi-head attention mechanism weights each head's attention weights with their corresponding value vectors to obtain the final output features. In this way, the multi-head attention mechanism can simultaneously capture features at different locations and scales, and extract relationships between global and local features.

[0135] 6. Global Average Pooling: Performs global average pooling on the output of the Transformer encoder, averaging the feature maps across each channel. This reduces the feature map of each image from shape (T, num_hiddens) to shape (T, 1).

[0136] 7. Classifier: A fully connected layer maps the globally average pooled features to a vector of size num_classes for classification. In the given model configuration, num_classes is 25, so the classifier's output shape is (B, T, 25). In this embodiment, the output size of the classifier is modified to (B, 1, T, 5, 5), obtaining rPPG feature information in the YCbCr space, which is then added to each channel of the rPPG feature information in the RGB space of the RGB encoder-decoder output.

[0137] Step S170: Add the second rPPG feature information to each channel of the first rPPG feature information to obtain the target rPPG feature information;

[0138] In this embodiment, the output of the RGB encoder-decoder is added to each channel with the output of YCbCr-VIT, resulting in a shape of [B,64,T,5,5]. This is then input into the final encoder-decoder. The first part of the final encoder consists of a 3D average pooling layer followed by a 3D convolutional block. The kernel size of the 3D average pooling layer is [2,2,2], the stride is [2,2,2], and the padding is 0. The resulting data after passing through this layer has a shape of [B,64,T / 2,2,2]. The 3D convolutional block is activated by a 3D convolutional kernel (64 input and output channels, kernel size [3,3,3], stride [1,1,1], padding [1,1,1]) and a batch processing (BatchNorm3d, 64 input and output channels) kernel. The first part of the result, composed of the ELU function, has a shape of [B,64,T / 2,2,2]. The second part consists of a 3D average pooling layer plus a 3D convolutional block. The kernel size of the 3D average pooling layer is [2,2,2], the stride is [2,2,2], and the padding is 0. The data obtained after passing through this layer has a shape of [B,64,T / 4,1,1]. The 3D convolutional block consists of a 3D convolution (both input and output channels are 64, the kernel size is [3,3,3], the stride is [1,1,1], and the padding is [1,1,1]) and a batch processing (BatchNorm3d, both input and output channels are 64) kernel activation function ELU. The second part of the result has a shape of [B,64,T / 4,1,1], which is the final output of the final encoder. The output of the final encoder is then decoded, and the result from the previous step is input into the final decoder. The final decoder consists of two 3D transposed convolutional blocks. The first transposed convolutional block consists of a 3D transposed convolution (64 input and output channels, kernel size [4,1,1], stride [2,1,1], padding [1,0,0]), BatchNorm3d (64 input and output channels), and the activation function ELU. The result from the previous step is input into the first 3D transposed convolutional block and the shape of the result is [B,64,T / 2,1,1]. The result is then input into the second 3D transposed convolutional block, which consists of a 3D transposed convolution (64 input and output channels, kernel size [4,1,1], stride [2,1,1], padding [1,0,0]), BatchNorm3d (64 input and output channels), and the activation function ReLU. The output decoding result is [B,64,T,1,1].The decoding result is input into the final rPPG signal extraction layer to obtain the desired rPPG signal. The rPPG signal extraction layer consists of a 3D convolution (kernel size [1,1,1], 64 input channels, 1 output channel, stride [1,1,1], padding with 0) and the ReLU activation function. The decoding result passes through this extraction layer to obtain the final rPPG signal, which serves as the target rPPG feature information. The shape of the target rPPG feature information is initially [B,1,T,1,1], which is ultimately modified to [B,T], where B is the batch size and T is the time length.

[0139] Step S180: Extract the measured heart rate of the target object from the target rPPG feature information.

[0140] In the embodiments of this application, such as Figure 13 As shown, after obtaining the target rPPG feature information, the target rPPG feature information is input into a bandpass filter to obtain the rPPG feature information to be converted within a preset frequency range; then, a Fourier transform is performed on the rPPG feature information to be converted to obtain the power spectral density; the peak frequency corresponding to the peak value is determined from the power spectral density; the peak frequency is multiplied by a preset value to obtain the measured heart rate of the target object. Specifically, after obtaining the target rPPG feature information, since the human heart rate range is generally between 40 and 250 beats per minute, which translates to a frequency range of 0.667 Hz to 4.167 Hz, the preset frequency range of the required rPPG signal is 0.667 Hz to 4.167 Hz. Using a bandpass filter, the target rPPG feature information can be filtered through frequencies [0.667 Hz, 4.167 Hz]. Then, the power spectral density (PSD) of the target rPPG feature information is obtained through Fourier transform. The peak frequency corresponding to the peak value is found from the power spectral density. Multiplying this frequency f by 60 gives the heart rate, thus completing the extraction of the heart rate from the rPPG signal.

[0141] In this embodiment, if the heart rate extraction model is being trained, it is considered that the label signal is mostly derived from the PPG signal of the fingertip pulse oximeter and the face rPPG, which should have a certain time delay. That is, the true value label may not completely correspond to the true label of the face rPPG. Directly using the mean squared error loss function will lead to network non-convergence. Therefore, the loss function of the network model in this embodiment includes calculating the linear correlation between the output value and the true value. The linear correlation can reduce the absolute value dependence between the network output and the true label, and instead use the degree of linearity to measure it. The formula for calculating the Pearson correlation coefficient is shown in formula (7):

[0142] Formula (7)

[0143] Furthermore, in order to obtain high signal-to-noise ratio (SNR) rPPG signals, which can greatly improve the accuracy and reliability of video heart rate measurement, it is necessary for the neural network to output high-quality rPPG signals. Referring to the cross-entropy formula and SNR formula in image classification tasks, the SNR loss as shown in formula (9) is proposed:

[0144]

[0145]

[0146] Formula (8)

[0147] In formula (8), The subscript frequency (true heart rate frequency) represents the maximum value of the PPG signal power spectral density obtained from a pulse oximeter. The subscript frequency represents the maximum value of the rPPG signal power spectral density output by the network (heart rate prediction frequency).

[0148] Therefore, the total loss function is shown in equation (9):

[0149] Formula (9)

[0150] In formula (9), It is L2 regularization. , , It's a hyperparameter.

[0151] Repeat the above process, use the Adam optimizer to optimize the model parameters, set the learning rate to 0.00005, and obtain the optimal rPPG signal extraction model. Use this model to extract rPPG signals from face videos.

[0152] In summary, in existing technologies, feeding video frames one by one into the network can easily lead to the neglect of temporal correlation information. While continuous video segments are input into a 3D convolutional network, although the number of frames is dense, the content changes relatively slowly, and the video content contains a significant amount of background information irrelevant to the signal extraction task. Furthermore, the choice of color space has a crucial impact on rPPG signal extraction and heart rate recovery; a good color space selection can effectively suppress noise and improve the signal-to-noise ratio of the target signal.

[0153] Therefore, this embodiment combines feature extraction in RGB and YCbCr spaces. Feature extraction in the RGB space using 2D and 3D CNNs ensures sufficient information extraction in both time and space, reducing interference from redundant information in video images that could affect rPPG signal extraction. The YCbCr space is generally used for skin segmentation; therefore, 2D / 3D-CNN feature extraction in the YCbCr space effectively suppresses noise and better captures facial information conducive to rPPG signal recovery. Based on extensive previous research and empirical rules, adding an average pooling layer provides better robustness than using individual pixels. The combined use of RGB and YCbCr spaces achieves sufficient extraction and utilization of temporal and spatial information and suppresses the effects of noise (such as illumination and motion), resulting in better robustness to measurements under different conditions.

[0154] Furthermore, ViT (Vison Transformer) is used to initially extract rPPG signals from the feature map in YCbCr space. The self-attention mechanism in the ViT model allows the model to model long-distance dependencies between pixels globally. This enables ViT to capture the global correlation between each location in the YCbCr feature map, ensuring more comprehensive temporal and spatial information and more thorough initial extraction of rPPG signals. The rPPG information obtained from YCbCr is combined with the rPPG information extracted through an RGB encoder-decoder structure in RGB space and input into the final encoder-decoder to recover the final rPPG signal. This structure allows the model to learn higher-level rPPG signal features. The decoder can better recover the rPPG signal through these features and can perform anomaly detection by comparing the differences between the input data and the reconstructed data, reducing the interference of outliers, lowering the noise interference of the rPPG signal, and making the rPPG signal cleaner and closer to the label signal.

[0155] Meanwhile, by adding a signal-to-noise ratio loss function to the loss function, the rPPG signal generated by the model can not only have a stronger linear relationship with the label data in the time domain, but also be as close as possible to the frequency domain information of the label signal in the frequency domain.

[0156] This invention provides a non-contact heart rate measurement system, comprising:

[0157] The first module is used to acquire the dataset of the target object, which contains facial videos and physiological signal labels;

[0158] The second module is used to perform the first preprocessing of the dataset to obtain RGB space sample data;

[0159] The third module is used to perform a second preprocessing on the RGB space sample data to obtain YCbCr space sample data.

[0160] The fourth module is used to extract feature data from RGB space sample data to form a total feature map in RGB space;

[0161] The fifth module is used to extract feature data from YCbCr spatial sample data to form the overall feature map of YCbCr space;

[0162] The sixth module is used to input the total feature map of the RGB space into the RGB encoder and RGB decoder, and obtain the rPPG feature information of the RGB space as the first rPPG feature information.

[0163] The seventh module is used to input the total feature map of YCbCr space into the visual transformer network to obtain the rPPG feature information of YCbCr space as the second rPPG feature information.

[0164] The eighth module is used to add the second rPPG feature information to the first rPPG feature information in each channel to obtain the target rPPG feature information;

[0165] The ninth module is used to extract the target object's measured heart rate from the target rPPG feature information.

[0166] The content of the method embodiments of the present invention is applicable to the system embodiments. The specific functions implemented in the system embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above methods.

[0167] This invention provides a non-contact heart rate measurement device, comprising:

[0168] At least one memory for storing programs;

[0169] At least one processor is used to load the program for execution. Figure 1 The non-contact heart rate measurement method shown.

[0170] The content of the method embodiments of the present invention is applicable to the device embodiments. The specific functions implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above methods.

[0171] This invention provides a computer storage medium storing a computer-executable program, which, when executed by a processor, is used to implement... Figure 1 The non-contact heart rate measurement method shown.

[0172] The content of the method embodiments of the present invention is applicable to the storage medium embodiments. The specific functions implemented by the storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above methods.

[0173] Furthermore, embodiments of the present invention also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform... Figure 1 The non-contact heart rate measurement method shown.

[0174] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made within the scope of knowledge possessed by those skilled in the art without departing from the spirit of the present invention. Furthermore, the embodiments of the present invention and the features thereof can be combined with each other unless otherwise specified.

Claims

1. A non-contact heart rate measurement method, characterized in that, Includes the following steps: Obtain a dataset of the target object, which includes facial videos and physiological signal labels; The dataset is subjected to a first preprocessing step to obtain RGB space sample data; The RGB space sample data is subjected to a second preprocessing to obtain YCbCr space sample data; The feature data of the RGB space sample data are extracted to form a total RGB space feature map; The feature data of the YCbCr spatial sample data are extracted to form the overall feature map of YCbCr space; After inputting the total feature map of the RGB space into the RGB encoder and RGB decoder, the rPPG feature information of the RGB space is obtained as the first rPPG feature information; The total feature map of the YCbCr space is input into the visual transformer network to obtain the rPPG feature information of the YCbCr space as the second rPPG feature information; The first rPPG feature information is added to the second rPPG feature information on each channel to obtain the target rPPG feature information; The measured heart rate of the target object is extracted from the target rPPG feature information.

2. The non-contact heart rate measurement method according to claim 1, characterized in that, The first preprocessing of the dataset to obtain RGB space sample data includes: Each frame of the video in the dataset is scaled to a preset resolution to obtain the video data to be processed. The video data to be processed is divided into video segments using a preset sliding window as RGB space sample data.

3. The non-contact heart rate measurement method according to claim 1, characterized in that, The second preprocessing of the RGB space sample data to obtain YCbCr space sample data includes: Face tracking is performed on each frame of the RGB space sample data, and the faces are cropped. Select several regions of interest in each cropped frame of the face image; The average pixel count is obtained by averaging all pixels in each region of interest across the R, G, and B channels. The average pixel value is converted to the YCbCr space, and the converted pixels of each frame are combined in the Y, Cb, and Cr channels to obtain YCbCr space sample data.

4. The non-contact heart rate measurement method according to claim 1, characterized in that, The feature data extracted from the RGB space sample data are used to form the overall RGB space feature map, including: The RGB space sample data is input into the first two-dimensional neural network to obtain the first two-dimensional feature map; The RGB space sample data is input into the first three-dimensional neural network to obtain the first three-dimensional feature map; Add the first two-dimensional feature map and the first three-dimensional feature map to obtain a preliminary RGB spatial data sample feature map; The initial RGB spatial data sample feature map is input into the residual attention module to obtain the total RGB spatial feature map.

5. A non-contact heart rate measurement method according to claim 4, characterized in that, The residual attention module includes a first convolutional layer, a second convolutional layer, an attention module, and a residual module. The attention module includes a channel attention module and a spatial attention module. The process of inputting the preliminary RGB spatial data sample feature map into the residual attention module to obtain the total RGB spatial feature map includes: The initial RGB spatial data sample feature map is processed sequentially through the first convolutional layer and the second convolutional layer to obtain the first input feature map; The first feature map to be input is input into the channel attention module to obtain the channel attention feature map; The channel attention feature map is multiplied with the first input feature map to obtain the second input feature map; The second feature map to be input is input into the spatial attention module to obtain the spatial attention feature map; The spatial attention feature map is multiplied with the second input feature map to obtain the third input feature map; The third feature map to be input and the preliminary RGB space data sample feature map are input into the residual module and added together to obtain the total RGB space feature map.

6. The non-contact heart rate measurement method according to claim 1, characterized in that, The feature data extracted from the YCbCr spatial sample data are used to form the overall YCbCr spatial feature map, including: The YCbCr spatial sample data is input into a second two-dimensional neural network to obtain a second two-dimensional feature map; The YCbCr spatial sample data is input into the second three-dimensional neural network to obtain the second three-dimensional feature map; The second two-dimensional feature map and the second three-dimensional feature map are added together to obtain the first YCbCr spatial data sample feature map; Input the first YCbCr spatial data sample feature map into the residual attention module, and the second YCbCr spatial data sample feature map; The second YCbCr spatial data sample feature map is input into the adaptive average pooling layer to obtain the total YCbCr spatial feature map.

7. The non-contact heart rate measurement method according to claim 1, characterized in that, The step of extracting the measured heart rate of the target object from the target rPPG feature information includes: The target rPPG feature information is input into a bandpass filter to obtain the rPPG feature information to be converted within a preset frequency range; Perform a Fourier transform on the rPPG feature information to be converted to obtain the power spectral density; Determine the peak frequency corresponding to the peak value from the power spectral density; The peak frequency is multiplied by a preset value to obtain the measured heart rate of the target object.

8. A non-contact heart rate measurement system, characterized in that, include: The first module is used to acquire a dataset of the target object, which includes facial videos and physiological signal labels; The second module is used to perform a first preprocessing on the dataset to obtain RGB space sample data; The third module is used to perform a second preprocessing on the RGB space sample data to obtain YCbCr space sample data; The fourth module is used to extract feature data from the RGB space sample data to form a total RGB space feature map; The fifth module is used to extract feature data from the YCbCr spatial sample data to form the overall YCbCr spatial feature map; The sixth module is used to input the total feature map of the RGB space into the RGB encoder and RGB decoder to obtain the rPPG feature information of the RGB space as the first rPPG feature information; The seventh module is used to input the total feature map of the YCbCr space into the visual transformer network to obtain the rPPG feature information of the YCbCr space as the second rPPG feature information. The eighth module is used to add the second rPPG feature information to the first rPPG feature information on each channel to obtain the target rPPG feature information; The ninth module is used to extract the measured heart rate of the target object from the target rPPG feature information.

9. A non-contact heart rate measuring device, characterized in that, include: At least one memory for storing programs; At least one processor is configured to load the program to perform the non-contact heart rate measurement method as described in any one of claims 1-7.

10. A computer storage medium, characterized in that, It contains a computer-executable program, which, when executed by a processor, is used to implement the non-contact heart rate measurement method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Non-contact fatigue detection method and system

    CN113420624A

  • Non-contact mental stress detection method and system

    CN115553777A