Driver heart rate monitoring method and system and vehicle

The method improves heart rate monitoring for drivers by using RGB and YCbCr color spaces with transformer architectures to analyze facial features, addressing the limitations of contact and visual methods, enhancing accuracy and reliability.

CN120318801APending Publication Date: 2025-07-15KAIRUI AUTOMOBILE TECHNOLOGY (ANHUI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510409457.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Traditional heart rate measurement equipment is large in size, expensive and does not have real-time performance. The vision-based driver's status monitoring method is low in reliability, making it difficult to monitor the driver's psychological and physiological status in real time and accurately.

Method used

The driver's heart rate monitoring method is adopted with multi-color spatial analysis combined with Transformer architecture, and the driver's heart rate signal is output through RGB and YCbCr image feature extraction, and the driver's heart rate signal is output using the multi-layer perceptron layer. The first Transformer architecture captures high-frequency information in the skin area, and the second Transformer architecture filters background noise to improve the accuracy of heart rate prediction.

Benefits of technology

It realizes contactless, real-time and accurate monitoring of the driver's heart rate, reduces light and motion interference, and improves the reliability and safety of driver status monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318801A_ABST
    Figure CN120318801A_ABST
Patent Text Reader

Abstract

The invention provides a method for monitoring the heart rate of a driver, and the method specifically comprises the following steps: collecting an RGB image of the face of the driver, converting the RGB image from an RGB space to a YCbCr space, and obtaining a YCbCr image; respectively extracting an RGB feature map and a YCbCr feature map of the RGB image and the YCbCr image; inputting a skin area in the RGB feature map into a first Transform architecture at a frequency of an RGB image acquisition frequency, and inputting the YCbCr feature map containing the background into a second Transform architecture at a frequency lower than the RGB image acquisition frequency; the features output by the first Transform architecture and the second Transform architecture are spliced and then input into the multi-layer perceptron layer, and the multi-layer perceptron layer outputs a heart rate signal of the driver. The heart rate of the driver is monitored in a non-contact mode, the method is used for judging the driving state of the driver, and active safety measures can be collected when the driving state is not good, so that the safety in the driving process is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of driving state monitoring, and provides a driver heart rate monitoring method, system and vehicle. Background Art

[0002] As the driving time increases, the heart rate of the driver usually decreases, while the heart rate variability (HRV) increases. This indicates that the sympathetic nerve activity of the driver weakens and the vagus nerve activity strengthens, reflecting the exacerbation of driving fatigue. When the driver feels fatigued, their reaction speed and attention will significantly decrease, thus increasing the driving risk.

[0003] Heart rate is an important physiological information for evaluating the driving state. Traditional heart rate measurement requires wearing a device on the driver for a long time. The driving state monitoring method based on physiological monitoring is not conducive to real-time monitoring, and the measurement device is usually large in size, expensive, and has a high measurement cost. There are problems such as complex and difficult to use, lack of real-time performance, and the need to install a monitoring device on the driver.

[0004] To solve the inconvenience of contact heart rate measurement, a non-contact heart rate measurement method - a vision-based driver state monitoring method has emerged, which extracts relevant facial features through visual images. However, most of them are limited to simple facial feature extraction, and the driver's poor state is a combined manifestation of psychological and physiological factors, which makes the reliability of the vision-based driver state monitoring method relatively low. Summary of the Invention

[0005] In view of this, the present application provides a driver heart rate monitoring method, aiming to improve the above problems.

[0006] Specifically, it includes the following technical solutions:

[0007] On the one hand, an embodiment of the present application provides a driver heart rate monitoring method, and the method is as follows:

[0008] (1) Collect the RGB image of the driver's face, convert the RGB image from the RGB space to the YCbCr space to obtain a YCbCr image;

[0009] (2) Extract the RGB feature map and YCbCr feature map of the RGB image and YCbCr image respectively;

[0010] (3) Input the skin area in the RGB feature map into the first Transformer architecture at the frequency of the RGB image acquisition frequency, and input the YCbCr feature map containing the background into the second Transformer architecture at a frequency lower than the RGB image acquisition frequency;

[0011] (4) After concatenating the features output by the first Transformer architecture and the second Transformer architecture, input them into a multi-layer perceptron layer, and the multi-layer perceptron layer outputs the heart rate signal of the driver.

[0012] In some embodiments of the present invention, the extraction process of the RGB feature map is specifically as follows:

[0013] Input the RGB image into the first three-dimensional convolutional neural network to obtain a three-dimensional RGB feature map T3. Input the data of two channels in the RGB image into the first two-dimensional convolutional neural network to obtain a two-dimensional RGB feature map T2. Expand the two-dimensional RGB feature map T2 into a three-dimensional RGB feature map T*3. Add the expanded three-dimensional RGB feature map T*3 to the three-dimensional RGB feature map T3 to obtain an initial RGB feature map. Input the initial RGB feature map into the residual attention module and output the RGB feature map.

[0014] In some embodiments of the present invention, the extraction process of the YCbCr feature map is specifically as follows:

[0015] Input the YCbCr image into the second three-dimensional convolutional neural network to obtain a three-dimensional YCbCr feature map Y3. Input the data of two channels in the YCbCr image into the second two-dimensional convolutional neural network to obtain a two-dimensional YCbCr feature map Y2. Expand the two-dimensional YCbCr feature map Y2 into a three-dimensional YCbCr feature map Y*3. Add the expanded three-dimensional YCbCr feature map Y*3 to the three-dimensional YCbCr feature map Y3 to obtain an initial YCbCr feature map. Input the initial YCbCr feature map into the residual attention module and the adaptive average pooling layer in sequence to obtain the YCbCr feature map.

[0016] In some embodiments of the present invention, the first Transformer architecture is stacked by a number of Transformer modules.

[0017] In some embodiments of the present invention, the second Transformer architecture is stacked by a number of Transformer modules. Taking the center of the face detection box as the center point, the area covered by twice the size of the face detection box is used as the YCbCr feature map including the background.

[0018] In some embodiments of the present invention, the YCbCr feature map including the background area is input into the second Transformer architecture at 1 / 4 of the RGB image acquisition frame rate.

[0019] In some embodiments of the present invention, the data of the R channel and the G channel or the data of the R channel and the B channel are input into the first two-dimensional convolutional neural network.

[0020] In some embodiments of the present invention, the data of the Cr channel and the Y channel or the data of the Cr channel and the Cb channel are input into the second two-dimensional convolutional neural network.

[0021] On the other hand, an embodiment of the present application provides a driver heart rate monitoring system, which includes:

[0022] A camera, an image conversion unit, and a heart rate monitoring model connected to the camera and the image conversion unit, where the heart rate monitoring model includes:

[0023] A first two-dimensional neural convolutional network and a first three-dimensional convolutional network connected to the camera, a first fusion unit connected to the first two-dimensional convolutional network and the first three-dimensional convolutional network, a first residual attention module connected to the first fusion unit, a first image processing unit connected to the first residual attention module, and a first Transformer architecture connected to the first image processing module;

[0024] A second two-dimensional convolutional network and a second three-dimensional convolutional network connected to the image conversion unit, a second fusion unit connected to the second two-dimensional convolutional network and the second three-dimensional convolutional network, a second residual attention module and an adaptive average pooling layer sequentially connected to the second fusion unit, a second image processing unit connected to the adaptive average pooling layer, and a second Transformer architecture connected to the second image processing unit; A third fusion unit connected to the first Transformer architecture and the second Transformer architecture, and a multi-layer perceptron layer connected to the third fusion unit.

[0025] The camera is used to collect the RGB image of the driver's face and input it into the image conversion unit. The image conversion unit converts the input RGB image from the RGB space to the YCbCr space to obtain the YCbCr image, and inputs the RGB image and the YCbCr image into the heart rate monitoring model, and the heart rate monitoring model outputs the heart rate signal of the driver.

[0026] On the other hand, an embodiment of the present application provides a vehicle, and the above driver heart rate monitoring system is integrated on the vehicle.

[0027] The present invention analyzes the skin from multiple color spaces to obtain more characteristic information related to heart rate. There is a lack of information in a single color space, which can be compensated by the information in another color space to improve the accuracy of heart rate prediction. In addition, the first Transformer architecture captures the subtle changes in skin color on a short time scale and is sensitive to high-frequency heart rate signals, which helps to track the changes in heartbeats. The second Transformer architecture can capture the long-term trends and low-frequency information in the image sequence, filter out some noise interferences, and helps to stably extract the heart rate signal, reducing the interferences caused by factors such as light, breathing, and slight movements. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0029] Figure 1 It is a flowchart of the driver heart rate monitoring method provided by the embodiment of the present invention;

[0030] Figure 2 It is a schematic structural diagram of the driver heart rate monitoring system provided by the embodiment of the present invention;

[0031] Through the above drawings, the clear embodiments of the present application have been shown, and more detailed descriptions will be given later. These drawings and the text descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application. YCbCr image

[0033] Unless otherwise defined, all technical terms used in the embodiments of the present application have the same meaning as commonly understood by those of ordinary skill in the art.

[0034] Figure 1 It is a flowchart of the driver heart rate monitoring method provided by the embodiment of the present invention, and the method is as follows:

[0035] (1) Collect the RGB image of the driver's face, convert the RGB image from the RGB space to the YCbCr space, and obtain the YCbCr image;

[0036] (2) Extract the RGB feature map and YCbCr feature map of the RGB image and YCbCr image respectively;

[0037] In the embodiment of the present invention, the extraction process of the RGB feature map is specifically as follows:

[0038] Directly input the RGB image into the first three-dimensional convolutional neural network to obtain the three-dimensional RGB feature map T3. Input the data of two channels in the RGB image into the first two-dimensional convolutional neural network. In the present invention, it is selected to input the data of the R channel and the G channel or the data of the R channel and the B channel into the first two-dimensional convolutional neural network to obtain the two-dimensional RGB feature map T2. Expand the two-dimensional RGB feature map T2 into a three-dimensional RGB feature map T*3. Add the expanded three-dimensional RGB feature map T*3 to the three-dimensional RGB feature map T3 to obtain the initial RGB feature map. Input the initial RGB feature map into the residual attention module to output the RGB feature map. The residual attention module includes channel attention and spatial attention. Through the residual attention module, important information in the RGB space can be better focused, and the feature representation ability can be enhanced.

[0039] In the embodiment of the present invention, the extraction process of the YCbCr feature map is specifically as follows:

[0040] Input the YCbCr image into the second three-dimensional convolutional neural network to obtain the three-dimensional YCbCr feature map Y3. Input the data of two channels in the YCbCr image into the second two-dimensional convolutional neural network. In the present invention, it is selected to input the data of the Cr channel and the Y channel or the data of the Cr channel and the Cb channel into the second two-dimensional convolutional neural network to obtain the two-dimensional YCbCr feature map Y2. Expand the two-dimensional YCbCr feature map Y2 into a three-dimensional YCbCr feature map Y*3. Add the expanded three-dimensional YCbCr feature map Y*3 to the three-dimensional YCbCr feature map Y3 to obtain the initial YCbCr feature map. Input the initial YCbCr feature map into the residual attention module and the adaptive average pooling layer in sequence to obtain the YCbCr feature map. This process utilizes the advantages of the YCbCr space in skin segmentation and noise suppression to enhance the feature extraction of skin regions and noise information.

[0041] (3) Input the skin region in the RGB feature map into the first Transformer architecture at the acquisition frequency of the RGB image, and input the YCbCr feature map with background into the second Transformer architecture at a frequency lower than the acquisition frequency of the RGB image;

[0042] In an embodiment of the present invention, the first Transformer architecture is stacked by a number of Transformer modules, which extracts the skin area in the RGB feature map and inputs it into the first Transformer architecture at the RGB image acquisition frame rate. Each Transformer module includes a multi-head self-attention mechanism and a feed-forward neural network, which calculates the attention from the spatial dimension and the time dimension and quickly processes local and detailed information, such as the detailed changes within the facial skin area, time features, etc., and can better extract time features and the rPPG information (remote photoplethysmography) contained in the skin area. Its input is the local information of the facial skin area (including the cheek area and the forehead area), which is directly obtained from the RGB feature map, and the input frame rate is consistent with the original sampling rate to ensure high-frequency capture of detailed information.

[0043] The second Transformer architecture is stacked by a number of Transformer modules, and inputs the YCbCr feature map containing the background area into the second Transformer architecture at 1 / 4 of the RGB image acquisition frame rate. It is obtained by downsampling the global information, based on the face detection box, with a wider range and containing more background and global information. However, through a lower frame rate and a larger receptive field, it can better focus on spatial information such as background noise, head movement, and facial action units and eliminate the influence brought by spatial changes. The input is the global information with the center of the face detection box as the center point and twice the size of the face detection box, and is sampled to twice the size of the local feature map, and calculates the attention from the spatial dimension and the time dimension.

[0044] (4) After splicing the features output by the first Transformer architecture and the second Transformer architecture, input them into the multi-layer perceptron layer, and the multi-layer perceptron layer outputs the heart rate signal of the driver.

[0045] The present invention analyzes the skin from multiple color spaces to obtain more feature information related to the heart rate. There is a lack of information in a single color space, which can be compensated by the information in another color space to improve the accuracy of heart rate prediction. In addition, the first Transformer architecture captures the subtle changes in skin color on a short time scale and is sensitive to high-frequency heart rate signals, which helps to track heart rate changes. The second Transformer architecture can capture the long-term trends and low-frequency information in the image sequence, can filter some noise interference, and helps to stably extract the heart rate signal, and can reduce the interference brought by factors such as light, breathing, and slight movement.

[0046] Figure 2Schematic diagram of the driver heart rate monitoring system provided by the embodiments of the present invention. For the sake of convenience of description, only the parts related to the embodiments of the present invention are shown. The system includes:

[0047] A camera, an image conversion unit, and a heart rate monitoring model connected to the camera and the image conversion unit. The heart rate monitoring model includes:

[0048] A first two-dimensional neural convolutional network and a first three-dimensional convolutional network connected to the camera, a first fusion unit connected to the first two-dimensional convolutional network and the first three-dimensional convolutional network, a first residual attention module connected to the first fusion unit, a first image processing unit connected to the first residual attention module, and a first Transformer architecture connected to the first image processing module;

[0049] A second two-dimensional convolutional network and a second three-dimensional convolutional network connected to the image conversion unit, a second fusion unit connected to the second two-dimensional convolutional network and the second three-dimensional convolutional network, a second residual attention module and an adaptive average pooling layer connected in sequence to the second fusion unit, a second image processing unit connected to the adaptive average pooling layer, and a second Transformer architecture connected to the second image processing unit; A third fusion unit connected to the first Transformer architecture and the second Transformer architecture, and a multi-layer perceptron layer connected to the third fusion unit.

[0050] The camera is used to collect the RGB image of the driver's face and input it into the image conversion unit. The image conversion unit converts the input RGB image from the RGB space to the YCbCr space to obtain the YCbCr image.

[0051] The RGB image is directly input into the first three-dimensional convolutional network to obtain a three-dimensional RGB feature map T3. The data of two channels in the RGB image are input into the first two-dimensional convolutional network to obtain a two-dimensional RGB feature map T2. The first fusion unit expands the two-dimensional RGB feature map T2 into a three-dimensional RGB feature map T*3, adds the three-dimensional RGB feature map T*3 to the three-dimensional RGB feature map T3 to obtain an initial RGB feature map, inputs the initial RGB feature map into the first residual attention module, outputs an RGB feature map, and the first image processing unit extracts the skin area in the RGB feature map and inputs it into the first Transformer architecture at the RGB image acquisition frame rate.

[0052] The YCbCr image is input into the second three-dimensional convolutional neural network to obtain a three-dimensional YCbCr feature map Y3. The two channel data of the YCbCr image are input into the second two-dimensional convolutional neural network to obtain a two-dimensional YCbCr feature map Y2. The second fusion unit expands the two-dimensional YCbCr feature map Y2 into a three-dimensional YCbCr feature map Y*3. The three-dimensional YCbCr feature map Y*3 is added to the three-dimensional YCbCr feature map Y3 to obtain an initial YCbCr feature map. The initial YCbCr feature map is sequentially input into the second residual attention module and the adaptive average pooling layer to obtain a YCbCr feature map. The second image processing unit processes it into a YCbCr feature map containing the background, and inputs it into the second Transformer architecture at 1 / 4 of the RGB image acquisition frame rate. The third fusion unit splices the features output by the first Transformer architecture and the second Transformer architecture, and inputs them into the multi-layer perceptron layer, which outputs the driver's heart rate signal.

[0053] The heart rate monitoring model needs to be trained with a large number of samples before use. The sample data includes facial images of the driver under different heart rate signals. The samples are divided into training samples and test samples. The heart rate monitoring model is trained based on the training samples, and tested based on the test samples. After detecting that the recognition accuracy of the heart rate monitoring model reaches the set standard, the training of the heart rate monitoring model is completed, and the trained heart rate monitoring model is used for the driver's heart rate signal collection, that is, the collected RGB image and YCbCr image of the driver's face are input into the heart rate monitoring model, and the heart rate monitoring model outputs the heart rate signal corresponding to the driver's facial features.

[0054] An embodiment of the present invention also provides a vehicle, which is integrated with the above-mentioned driver heart rate monitoring system, which monitors the driver's heart rate in a non-contact manner and is used to determine the driver's driving state. When the driving state is not good, active safety measures can be taken to improve safety during driving.

[0055] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the present application disclosed herein. The present application is intended to cover any variations, uses or adaptations of the present application, which follow the general principles of the present application and include common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only.

[0056] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A driver heart rate monitoring method, characterized in that, The method is as follows: (1) Collect the RGB image of the driver's face, convert the RGB image from the RGB space to the YCbCr space to obtain the YCbCr image; (2) Extract the RGB feature map and YCbCr feature map of the RGB image and YCbCr image respectively; (3) Input the skin region in the RGB feature map into the first Transformer architecture at the frequency of the RGB image acquisition frequency, and input the YCbCr feature map containing the background into the second Transformer architecture at a frequency lower than the RGB image acquisition frequency; (4) After splicing the features output by the first Transformer architecture and the second Transformer architecture, input them into the multi-layer perceptron layer, and the multi-layer perceptron layer outputs the heart rate signal of the driver.

2. The driver heart rate monitoring method according to claim 1, wherein The extraction process of the RGB feature map is as follows: Input the RGB image into the first three-dimensional convolutional neural network to obtain the three-dimensional RGB feature map T3. Input the data of two channels in the RGB image into the first two-dimensional convolutional neural network to obtain the two-dimensional RGB feature map T2. Expand the two-dimensional RGB feature map T2 into the three-dimensional RGB feature map T*3. Add the three-dimensional RGB feature map T*3 to the three-dimensional RGB feature map T3 to obtain the initial RGB feature map. Input the initial RGB feature map into the residual attention module to output the RGB feature map.

3. The driver heart rate monitoring method according to claim 1, wherein The extraction process of the YCbCr feature map is as follows: Input the YCbCr image into the second three-dimensional convolutional neural network to obtain the three-dimensional YCbCr feature map Y3. Input the data of two channels in the YCbCr image into the second two-dimensional convolutional neural network to obtain the two-dimensional YCbCr feature map Y2. Expand the two-dimensional YCbCr feature map Y2 into the three-dimensional YCbCr feature map Y*3. Add the three-dimensional YCbCr feature map Y*3 to the three-dimensional YCbCr feature map Y3 to obtain the initial YCbCr feature map. Input the initial YCbCr feature map into the residual attention module and the adaptive average pooling layer in sequence to obtain the YCbCr feature map.

4. The driver heart rate monitoring method according to claim 1, wherein The first Transformer architecture is stacked by a number of Transformer modules.

5. The driver heart rate monitoring method according to claim 1, wherein, The second Transformer architecture is stacked by a number of Transformer modules. Taking the center of the face detection box as the center point, the area covered by twice the size of the face detection box is used as the YCbCr feature map containing the background.

6. The driver heart rate monitoring method according to claim 1, characterized in that The YCbCr feature map containing the background region is input into the second Transformer architecture at 1 / 4 of the RGB image acquisition frame rate.

7. The driver heart rate monitoring method according to claim 2, wherein Input the data of the R channel and the G channel, or the data of the R channel and the B channel into the first two-dimensional convolutional neural network.

8. The driver heart rate monitoring method according to claim 3, wherein, Input the data of the Cr channel and the Y channel, or the data of the Cr channel and the Cb channel into the second two-dimensional convolutional neural network.

9. A driver heart rate monitoring system, characterized in that, The system includes: A camera, an image conversion unit, and a heart rate monitoring model connected to the camera and the image conversion unit. The heart rate monitoring model includes: Connected to the first two-dimensional neural convolutional network and the first three-dimensional convolutional neural network connected to the camera, connected to the first two-dimensional convolutional neural network and the first three-dimensional convolutional neural network is the first fusion unit, connected to the first fusion unit is the first residual attention module, connected to the first residual attention module is the first image processing unit, and connected to the first image processing module is the first Transformer architecture; Connected to the image conversion unit are the second two-dimensional convolutional neural network and the second three-dimensional convolutional neural network, connected to the second two-dimensional convolutional neural network and the second three-dimensional convolutional neural network is the second fusion unit, sequentially connected to the second fusion unit are the second residual attention module and the adaptive average pooling layer, connected to the adaptive average pooling layer is the second image processing unit, and connected to the second image processing unit is the second Transformer architecture; connected to the first Transformer architecture and the second Transformer architecture is the third fusion unit, and connected to the third fusion unit is the multi-layer perceptron layer. The camera is used to collect the RGB image of the driver's face and send it to the image conversion unit. The image conversion unit converts the input RGB image from the RGB space to the YCbCr space to obtain the YCbCr image, and inputs the RGB image and the YCbCr image into the heart rate monitoring model, and the heart rate monitoring model outputs the heart rate signal of the driver.

10. A vehicle, characterized in that, The vehicle is integrated with the driver heart rate monitoring system as described in claim 9.