Driver heart rate monitoring method, storage medium, program product, electronic equipment and vehicle
By using optical and infrared cameras to acquire images in a coordinated manner, and combining a dual-branch neural network model and motion compensation algorithm, the problem of accuracy in heart rate monitoring under complex environments was solved, and efficient heart rate detection was achieved in vehicle environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-12
- Publication Date
- 2026-04-10
AI Technical Summary
In existing technologies, driver heart rate monitoring is susceptible to noise interference in complex driving environments such as drastic changes in external light, low illumination at night, or strong backlight, leading to failure of heart rate estimation or a decrease in accuracy.
The system uses optical and infrared cameras to acquire driver images in collaboration, and combines a pre-trained dual-branch neural network model with motion compensation and noise removal algorithms to output the driver's heart rate value.
The accuracy and robustness of heart rate detection have been improved in complex driving environments, meeting automotive-grade health monitoring requirements.
Smart Images

Figure CN121817834A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent vehicles, and in particular to a driver heart rate monitoring method, a storage medium, a program product, an electronic device and a vehicle. BACKGROUND
[0002] The driver in an unsafe driving state is an important reason for causing traffic accidents. Abnormal heart rate is one of the unsafe states that the driver often appears in the driving process.
[0003] In the related art, driver heart rate monitoring relies on a single optical camera to collect visible light images, and extracts weak skin blood flow change signals based on remote photoplethysmography (RPPG). However, in complex driving environments such as severe external light changes, low light at night, or strong backlights, visible light images are easily disturbed by noise, resulting in a sharp decline in the signal-to-noise ratio of the RPPG signal, and heart rate estimation failure or accuracy decline. SUMMARY
[0004] The present application aims to at least solve one of the technical problems existing in the prior art. To this end, the present application provides a driver heart rate monitoring method, a storage medium, a program product, an electronic device and a vehicle. In the present application, the images of the driver are collected based on an optical camera and an infrared camera; the images of the driver are input into a pre-trained neural network model, and the heart rate value of the driver is output. Compared with the related art, the present application uses optical and near-infrared dual-mode cameras to cooperatively collect images, which is not limited by complex driving scenes such as night and backlight, and can improve the accuracy of heart rate detection, thereby solving the problem that the related art cannot accurately estimate the heart rate of the driver in complex environments.
[0005] To achieve the above-mentioned purpose, the application adopts the following technical solutions:
[0006] In a first aspect, the application provides a driver heart rate monitoring method, which comprises: collecting images of a driver based on an optical camera and an infrared camera; inputting the images of the driver into a pre-trained neural network model to output the heart rate value of the driver.
[0007] In some embodiments of the present application, before the images of the driver are input into the pre-trained neural network model, the method further comprises: performing face detection on the images of the driver to obtain images containing only the face of the driver; and performing motion compensation on the images of the face of the driver to obtain denoised face images.
[0008] In some embodiments of the present application, the motion compensation on the face image of the driver to obtain a de-noised face image comprises: spatially averaging the pixel intensity of a face region of interest in the face image of the driver to obtain a multi-dimensional RPPG signal matrix P containing noise; determining a multi-dimensional noise space matrix Q according to a multi-dimensional horizontal motion signal matrix H and a multi-dimensional vertical motion signal matrix V and a multi-dimensional background illumination signal matrix B in the face image of the driver; and calculating a de-noised face image feature matrix based on the multi-dimensional RPPG signal matrix P and the multi-dimensional noise space matrix Q.
[0009] In some embodiments of the present application, the calculation of the de-noised face image feature matrix based on the signal matrix P and the noise space matrix Q comprises: orthogonally projecting the signal matrix P to the noise space matrix Q to obtain a projection matrix of the signal matrix P; and determining a de-noised face image feature matrix according to the difference between the signal matrix P and the projection matrix.
[0010] In some embodiments of the present application, the method further comprises: performing target operation processing on the face image of the driver in each frame to obtain a multi-dimensional background illumination signal matrix B varying with time; wherein the target operation processing comprises: selecting a plurality of background regions not containing face regions from the face image of the driver, dividing each background region into n*n pixel blocks, determining a background illumination signal according to the intensity value of the pixel blocks, and n is greater than or equal to 2.
[0011] In some embodiments of the present application, the method further comprises: determining a multi-dimensional horizontal motion signal matrix H and a vertical motion signal matrix V varying with time.
[0012] In some embodiments of the present application, the method further comprises: performing target operation processing on the face image of the driver in each frame to obtain a multi-dimensional background illumination signal matrix B varying with time; wherein the target operation processing comprises: selecting a plurality of background regions not containing face regions from the face image of the driver, dividing each background region into n*n pixel blocks, determining a background illumination signal according to the intensity value of the pixel blocks, and n is greater than or equal to 2; and determining a multi-dimensional horizontal motion signal matrix H and a vertical motion signal matrix V varying with time according to the multi-dimensional RPPG signal matrix P containing noise.
[0013] In some embodiments of the present application, the face detection on the image of the driver to obtain an image containing only the face of the driver comprises: performing grayscale processing on the image of the driver to obtain a grayscale image; identifying the face in the grayscale image using a classifier and cropping the grayscale image according to a detection frame to output an image containing only the face of the driver.
[0014] In some embodiments of this application, the neural network model includes at least a first sub-network and a second sub-network. The first sub-network is used to perform difference calculation on the images of the previous and next frames to obtain the pixel changes between the previous and next frames, and the second sub-network is used to enhance the image features of the current frame.
[0015] In some embodiments of this application, the first sub-network and the second sub-network are respectively composed of a convolutional layer with a kernel size of 3*3 and b convolutional layers with a kernel size of 2*2; wherein: a is greater than or equal to 4, b is greater than or equal to 2; and / or, the dual-branch neural network model further includes a prediction head, which includes a flattening layer and a linear layer. The flattening layer is used to flatten the three-dimensional tensor into a one-dimensional tensor for calculation, and the linear layer is used to project and map the one-dimensional tensor for prediction.
[0016] Secondly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the computer to perform the method described in the first aspect.
[0017] Thirdly, this application provides a computer program product that stores instructions which, when executed by a computer, cause the computer to perform the method described in the first aspect.
[0018] Fourthly, this application provides an electronic device, comprising: a memory having a computer program stored thereon; and a processor for executing the computer program in the memory to implement the method as described in the first aspect.
[0019] Fifthly, this application provides a vehicle comprising: an electronic device as described in the fourth aspect; or, a processor configured to perform the method as described in the first aspect.
[0020] The advantages and control methods of the vehicle and electronic equipment compared to the prior art are the same, and will not be elaborated here.
[0021] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] To gain a more complete understanding of this application and its beneficial effects, the following description will be provided in conjunction with the accompanying drawings, wherein the same reference numerals in the following description denote the same parts.
[0024] Figure 1 This is a flowchart illustrating a driver's heart rate monitoring method according to an embodiment of the present invention;
[0025] Figure 2 This is a schematic diagram of the installation layout of an optical camera and an infrared camera according to an embodiment of the present invention;
[0026] Figure 3 This is a schematic diagram of the face detection process provided in an embodiment of the present invention;
[0027] Figure 4 This is a schematic diagram of the process for facial image motion compensation according to an embodiment of the present invention;
[0028] Figure 5 This is a schematic diagram of the structure of a dual-branch neural network model provided in an embodiment of the present invention;
[0029] Figure 6 This is an overall flowchart of a heart rate monitoring method provided according to an embodiment of the present invention;
[0030] Figure 7 This is a flowchart of the heart rate detection model board deployment according to an embodiment of the present invention;
[0031] Figure 8 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of the present invention. Detailed Implementation
[0032] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the protection scope of this application.
[0033] In related technologies, driver heart rate monitoring often relies on a single optical camera to acquire visible light images and extracts weak skin blood flow change signals based on remote photoplethysmography (RPPG). However, in complex driving environments such as drastic changes in external light, low illumination at night, or strong backlight, visible light images are easily affected by noise, causing a sharp drop in the signal-to-noise ratio of the RPPG signal, resulting in failure of heart rate estimation or a decrease in accuracy.
[0034] To address this, the present invention proposes a driver heart rate monitoring method, storage medium, program product, electronic device, and vehicle. This solution uses an optical camera and an infrared camera to acquire images of the driver; the driver's images are then input into a pre-trained neural network model, which outputs the driver's heart rate value. Compared to related technologies, this solution utilizes a dual-modal camera (optical and near-infrared) to collaboratively acquire images, overcoming limitations imposed by complex driving scenarios such as nighttime and backlighting, thereby improving the accuracy of heart rate detection.
[0035] The present invention will now be described in further detail with reference to the embodiments.
[0036] like Figure 1 The diagram shown is a flowchart illustrating a driver heart rate monitoring method according to an embodiment of the present invention. The method includes:
[0037] 101. Images of the driver are acquired using optical and infrared cameras.
[0038] 102. Input the driver's image into a pre-trained neural network model and output the driver's heart rate value.
[0039] The aforementioned neural network model can be a two-branch neural network model, which includes at least a first sub-network and a second sub-network. The first sub-network is used to perform difference calculations on the images of the preceding and following frames to obtain the pixel changes between the preceding and following frames, and the second sub-network is used to enhance the image features of the current frame.
[0040] In the above embodiments, by simultaneously acquiring optical and near-infrared dual-modal images, the system can stably acquire the driver's facial physiological information under full illumination conditions: the optical camera provides visible light RGB images, preserving details of skin texture, color, and microvascular distribution; the near-infrared camera can still output grayscale images with high signal-to-noise ratio in low-light or high-contrast scenarios such as nighttime, tunnel entry / exit, and strong backlighting. It is sensitive to the absorption characteristics of hemoglobin and can effectively penetrate the superficial skin layer, enhancing the visibility of microcirculation pulsation. The dual-branch neural network model decouples physiological signals from motion noise by separating the modeling of inter-frame dynamic changes from the local texture of the current frame: the first sub-network takes two consecutive frames as input, calculates the difference map pixel by pixel, focuses on the weak light intensity modulation (i.e., RPPG signal) caused by capillary pulsation, and suppresses inter-frame displacement artifacts caused by head shaking and vehicle vibration; the second sub-network performs multi-scale nonlinear enhancement on the current frame image, expands the receptive field through dilated convolution, and strengthens the high-frequency micro-motion response of areas with rich blood supply such as the cheekbone area, nasal alar area, and forehead. The outputs of the two branches are spliced together and then input into the prediction head to achieve end-to-end heart rate regression. The model significantly improves the monitoring accuracy in complex driving environments while keeping the inference latency below 15ms / frame.
[0041] In some embodiments, the optical camera and near-infrared camera in step 100 can be integrated into the same binocular camera module, installed in the vehicle rearview mirror base, with a field of view covering the central area of the driver's face, a sampling frequency of 30fps, and a uniform image resolution of 1280×720. The image acquisition module is directly connected to the vehicle's AI computing module via a MIPI CSI-2 interface, achieving zero-copy data transmission with an end-to-end latency of less than 50ms. The dual-modal images undergo spatial alignment during the preprocessing stage, employing sub-pixel-level feature point matching to ensure pixel-level synchronization and avoid signal misalignment caused by lens parallax.
[0042] like Figure 2 The diagram shows the installation layout of the optical camera and infrared camera provided in an embodiment of the present invention. As a preferred implementation, the optical camera and infrared camera in step 100 above can be installed above the rearview mirror and at the A-pillar of the driver's seat, respectively, with their orientation adjusted so that the camera's field of view includes the driver's seat. The installation of the two cameras needs to take into account the possible changes in the driver's face position during actual driving. During actual driving, when the driver turns their head left and right to observe the rearview mirror, the angle of rotation towards the passenger seat is usually larger. Therefore, the optical camera, which captures more RPPG signals, needs to be installed in the center of the front compartment to maximize the acquisition of the driver's face signal. The infrared camera is installed near the A-pillar of the driver's seat to prevent noise from affecting the detection effect.
[0043] During dataset collection and actual deployment, it is necessary to ensure that the positions, angles, and layouts of the two cameras are consistent to avoid the impact of different angles on face exposure. Inter-frame artifacts are usually caused by vehicle body shaking or driver head rotation during driving, and their presence can easily lead to the loss of RPPG signals in the face, thus causing the model to deteriorate. To eliminate the impact of inter-frame artifacts, the original image and the reference image are first registered using consistency-sensitive hashing. The hash function family used is as follows: .
[0044] in, For image blocks, To be from the interval Real numbers selected evenly from the middle Walsh-Hadamard kernel functions, , Typically, 8 is chosen. After selecting a hash function family that meets the criteria during the indexing phase, [the system] creates... Hash function table Each entry is composed of hash codes formed by concatenating the entries using the function above. ,Right now Then, all image patches in the image are mapped and stored in this hash table. In the process of registering the two consecutive input images, a set of registered images is obtained, with exposure parameters aligned with the original image and content aligned with the reference image. The registered image sequence is then further fused to obtain the final image.
[0045] In some embodiments, the training dataset for the dual-branch neural network model comes from over 20,000 real driver videos collected from the cockpits of multiple vehicle models. These videos cover 20 extreme scenarios, including different skin tones, drivers wearing glasses or hats, low-light conditions at night, strong glare, head tilting, and side profiles. Each dataset simultaneously includes an electrocardiogram (ECG) as the gold standard label. The model employs a combination of cross-entropy loss and MAE optimization, achieving a heart rate estimation error of less than ±4 bpm on the validation set, meeting automotive-grade health monitoring requirements.
[0046] In some embodiments, prior to step 102 above, the method further includes: 102a, performing face detection on the driver's image to obtain an image containing only the driver's face; 102b, performing motion compensation on the driver's face image to obtain a denoised face image.
[0047] In the above embodiments, the pre-processing module effectively eliminates interference signals from non-facial areas such as the steering wheel, dashboard, and A-pillar, and removes rigid displacement artifacts caused by slight head movements, nodding, or tilting of the driver's head, ensuring that subsequent RPPG signal extraction only reflects real physiological changes. This preprocessing significantly reduces input noise to the neural network, improving model convergence efficiency and generalization ability.
[0048] For example, step 102a above specifically includes the following: 102a1, performing grayscale processing on the driver's image to obtain a grayscale image; 102a2, using a classifier to identify faces in the grayscale image and cropping the grayscale image according to the detection box, outputting an image containing only the driver's face.
[0049] In the above embodiments, grayscale processing can employ a weighted average method: I_gray = 0.299×I_red+0.587×I_green+0.114×I_blue (optical channel) or directly take the intensity value of the infrared channel. Face recognition can be achieved using a Haar feature cascade classifier finely tuned with a cockpit dataset. For example, positive samples can be face images of 20 poses, and negative samples can be easily confused structures. Optical and infrared images are input into independent cascade classifiers, and cross-modal matching is performed after the output detection boxes: only when the center distance between the two modal detection boxes is less than 10 pixels and the overlap area is greater than 85% is it retained as a valid detection result. The cropped image size can be uniformly set to 128×128 pixels as input for motion compensation and neural networks, ensuring consistent input scale and avoiding feature extraction deviations caused by changes in face size, thus providing high-quality input for subsequent motion compensation.
[0050] like Figure 3 The diagram illustrates the face detection process provided in this embodiment of the invention. In the actual acquisition process of vehicle-mounted cameras, the original image often contains complex visual information, capturing not only the driver's facial features but also the faces of fellow passengers, other pedestrians on the road, the external environment, and a large amount of irrelevant content. For the driver's remote heart rate detection model, only the driver's facial details can provide effective information; the remaining parts not only fail to contribute effective features to the model but also incur additional computational costs and may even interfere with model training. Therefore, a Haar feature cascade classifier is used for face detection. The original image is cropped based on the detection bounding box to obtain an image containing only the driver's face. For details, please refer to... Figure 3 The content shown.
[0051] In some embodiments, step 102b specifically includes the following: 102b1, spatially averaging the pixel intensity of the region of interest in the driver's face image to obtain a noise-containing RPPG signal matrix P; 102b2, determining the noise spatial matrix Q based on the time-varying multidimensional horizontal motion signal matrix H and vertical motion signal matrix V and the background illumination signal matrix B in the driver's face image; 102b3, calculating the denoised face image feature matrix based on the signal matrix P and the noise spatial matrix Q.
[0052] In the above embodiments, the region of interest (ROI) on the face is composed of pixels from areas such as the cheeks, forehead, and bridge of the nose. The spatially averaged P matrix generated from this matrix is a mixed signal containing physiological signals and various types of noise. By independently modeling horizontal motion H, vertical motion V, and background illumination B, a three-dimensional noise basis Q is constructed, whose column space represents linearly predictable interference components. This method overcomes the limitations of traditional single motion compensation and, for the first time, achieves joint modeling and orthogonal separation of motion and illumination noise in RPPG signal processing, significantly improving denoising accuracy.
[0053] It should be noted that there is no strict order between steps 102b1 and 102b2. Step 102b1 can be executed first and then step 102b2, or vice versa, or they can be executed simultaneously. There is no restriction on this.
[0054] In some embodiments, step 102b3 above specifically includes the following: 102b31, orthogonally projecting the multidimensional RPPG signal matrix P onto the multidimensional noise space matrix Q to obtain the projection matrix of the multidimensional signal matrix P; 102b32, determining the denoised face image feature matrix based on the difference between the multidimensional signal matrix P and the projection matrix.
[0055] In the above embodiment, the projection matrix P_proj = Q·(QT·Q){-1}·Q^T·P is calculated using the generalized least squares method, which preserves the noise components in P that are collinear with H, V, and B. The denoised feature matrix P_denoised = P - P_proj represents the components of P in the Q orthogonal complement space, mainly preserving the periodic physiological signals. This mathematical modeling method is implemented on an embedded GPU through SVD decomposition optimization, with a single-frame processing latency of less than 5ms, meeting real-time requirements. Furthermore, it does not rely on prior physiological models and possesses good generalization ability.
[0056] In some embodiments, step 102b2 specifically includes the following: concatenating the multidimensional horizontal motion signal matrix H, the multidimensional vertical motion signal matrix V, and the multidimensional background illumination signal matrix B to obtain a multidimensional noise space matrix Q. For example, if H, V, and B are all T×5 matrices, then the concatenated Q is a T×15 noise matrix Q = .
[0057] In some embodiments, the above method further includes: 102b4, performing target operation processing on the driver's face image in each frame to obtain a time-varying multidimensional background illumination signal matrix B; wherein, the target operation processing includes: selecting multiple background regions that do not contain facial regions from the driver's face image, dividing each background region into n*n pixel blocks, determining the background illumination signal based on the intensity value of the pixel blocks, where n is greater than or equal to 2; and / or, 102b5, determining a time-varying multidimensional horizontal motion signal matrix H and a vertical motion signal matrix V.
[0058] In the above embodiment, step 102b4 may specifically include the following: When modeling background lighting, at least three non-facial regions (such as the inner side of the A-pillar, the roof lining, and the edge of the dashboard) are selected in each frame image. Each region is divided into 4×4 pixel blocks (n=4). The average intensity of the pixels in each block is calculated to form a background lighting vector. This vector is then filtered by median filtering and low-pass filtering to eliminate transient interference, and finally concatenated into a B matrix. This method avoids the physiological signal contamination caused by using facial regions to estimate lighting.
[0059] In some embodiments, step 102b5 may specifically include the following: calculating the average displacement vector within a ring-shaped region of a certain pixel width at the outer edge of the ROI using optical flow to obtain Ht and Vt; or utilizing the correlation between the low-frequency fluctuation characteristics of Pt (<0.3Hz) and known motion patterns to back-calculate the estimated values of Ht and Vt through linear regression. The two methods serve as redundant checks for each other. When optical flow fails (e.g., strong glare causing texture loss), the system automatically switches to regression estimation mode to ensure continuous and stable output of the motion signal. By constructing motion signal matrices H and V as supplementary dimensions to Q, the completeness of noise modeling is enhanced, especially improving robustness in low-texture regions.
[0060] like Figure 4 The diagram illustrates the process of facial image motion compensation according to the present invention. For example, RPPG signals can be extracted from videos by spatially averaging pixel groups to reduce noise in the image. The original RPPG signal is obtained from the video frame by spatially averaging the pixel intensities of 48 regions of interest (ROIs) on the driver's face. In this method, 68 facial landmarks are first detected, and then the detected landmarks are interpolated and extrapolated to obtain a total of 145 points, including the forehead region, to define the 48 ROIs. The original RPPG signal is obtained from the average intensity. It is a one-dimensional time series signal, where It is a length of The video frame index within the time window of the frame. The facial signals from each region are stacked into an RPPG matrix. Size is During the cardiac cycle, blood flow exhibits roughly the same temporal variation pattern across all facial regions, therefore the underlying RPPG signals in each facial region should be low-rank. However, different facial regions possess different levels of noise, which is high-dimensional. An orthogonal projection method is used to transform the noisy RPPG signals... Projected onto the noise subspace Up, and then from Subtracting this projection signal from the middle is equivalent to projecting the noisy RPPG signal onto the orthogonal complement space of the noise subspace.
[0061] This method uses two time-varying 5D signals (a 5D horizontal motion signal and a 5D horizontal motion signal). and 5D vertical motion signals This is used to summarize facial movements, thus approximating motion noise. Horizontal motion signal The extraction is achieved through the following steps:
[0062] A1. Horizontal motion measurement of small areas: For each of the 48 small facial areas, the horizontal position of the area in the current frame is obtained by averaging the horizontal coordinates of the four corner points of the area in each frame; the difference in horizontal position between adjacent frames is the instantaneous horizontal motion amplitude of the small area.
[0063] A2. Median Region Dimension Reduction: The horizontal motion signals of the 48 small regions are grouped according to their respective 5 median regions; the median of the horizontal motion signals of all small regions within each median region is taken to obtain the horizontal motion value of that median region.
[0064] A3. Formation Matrix: Repeat the above operation for each frame within a 10-second sliding time window (30fps, 300 frames in total), ultimately forming a matrix with dimension [missing information]. Horizontal motion signal matrix Each column corresponds to a horizontal motion time series of a median region.
[0065] Vertical motion signal The extraction logic is completely symmetrical to the horizontal motion signal and is obtained by the same steps described above.
[0066] These two signals measure the horizontal / vertical motion of each of the 48 facial regions by spatially averaging the positions of the four corners of each region in each frame. Then, by calculating the median of the motion signals of all sub-regions belonging to each large region, the 48-dimensional data is reduced to 5 dimensions (one dimension per large region). The signal sequence of all timestamps within the time window is then used to construct a data structure of size [size missing]. matrix and .
[0067] In addition, a background illumination signal that varies over time is also required. By selecting five regions in the background that do not contain facial areas, these five background regions are divided into... For each pixel block, the intensity values within each small region are spatially averaged, and then the median of these averages is taken as the background illumination signal for each background region. This process is repeated for all five background regions in each frame, resulting in a sequence of pixels of size [size missing]. Background lighting matrix .
[0068] The above three noise signals are spliced together to obtain a value of [value]. noise matrix The noisy RPPG signal matrix Orthogonal projection onto the noise subspace Up, then from the RPPG signal matrix Subtracting the projection from the middle yields the RPPG signal after orthogonal projection denoising. : .
[0069] In some embodiments, the first and second sub-networks in the above-described dual-branch neural network model are respectively composed of convolutional layers a with a kernel size of 33 and b with a kernel size of 22; wherein: a is greater than or equal to 4, and b is greater than or equal to 2; and / or, the dual-branch neural network model further includes a prediction head, which includes a flattening layer and a linear layer. The flattening layer is used to flatten the three-dimensional tensor into a one-dimensional tensor for computation, and the linear layer is used to project and map the one-dimensional tensor for prediction. In the above embodiments, through a carefully designed dual-branch convolutional structure and a lightweight prediction head, the model meets automotive-grade computing power and power consumption constraints while maintaining high accuracy.
[0070] like Figure 5 The diagram shows the structure of the dual-branch neural network model provided in this embodiment of the invention. The first sub-network in the dual-branch architecture can be a Motion sub-network. The Motion sub-network uses the pixel difference between two frames of the driver's face image as the change in the RPPG signal and is composed of multiple convolutional networks. For example, it can consist of 4 convolutional layers with a kernel size of and 2 convolutional layers with a kernel size of . Since heart rate detection is highly dependent on pixel changes between consecutive frames, the model's feature extraction should not be too deep; therefore, only two levels of downsampling are used for feature extraction. The second sub-network in the dual-branch architecture can be an Appearance sub-network. It also uses a two-level downsampling feature extraction method, feeding the downsampled features of the t-th frame image back into the Motion sub-network as an auxiliary network to extract the spatial features of the image, and then fusing them into the temporal features. This includes one... Convolution and a Normalization layer: Normalize the data to a standard normal distribution, where: Indicates features, The mean, Standard deviation, These are the normalized features. 1x1 convolutions are used to align Appearance and Motion network features, while Normalization removes scale information, allowing spatial features to be better integrated with temporal features. The prediction head in the two-branch neural network model contains one flattening layer and two linear layers. The flattening layer flattens the three-dimensional tensor into a one-dimensional tensor for subsequent linear layer computation, and the two linear layers project and map the one-dimensional tensor for prediction.
[0071] like Figure 6 The diagram shown is an overall flowchart of a heart rate monitoring method provided by an embodiment of the present invention. Specifically, it includes the following steps:
[0072] S101: Enables simultaneous image acquisition by optical and infrared cameras.
[0073] Image acquisition is performed along with inter-frame artifact removal. On the hardware side, a combination of a high-resolution optical camera and a near-infrared camera is used for image acquisition, fusing multispectral information to improve all-weather adaptability.
[0074] S102: Perform driver face detection on the image.
[0075] S103: Perform motion compensation on the face image.
[0076] S104: Motion subnetwork in a dual-branch architecture.
[0077] S105: Appearance subnetwork in a dual-branch architecture.
[0078] S106: Predictive head.
[0079] The specific implementation process of steps S101-S106 above is described in the corresponding sections above and will not be repeated here.
[0080] like Figure 7 This is a flowchart of the deployment process of the heart rate detection model board of the present invention. The specific steps are as follows:
[0081] S201: Dataset Acquisition. Before model training, driver facial data needs to be acquired using two types of cameras. Ground truth values are recorded by wearing a fingernail-shaped heart rate monitor while driving, and optical and near-infrared cameras are installed in front of the cockpit to capture driver video. The acquired dataset includes various driving scenarios and time periods to improve the generalization ability of the trained model. It is important to note that the camera layout during data acquisition should be consistent with the actual model deployment to avoid differences in exposure distribution due to different orientations and positions. The frame rate is set to 30fps during data acquisition, with each 30-second segment containing 900 frames. At least 150 video segments are required for the model to achieve good performance and convergence.
[0082] S202: Data Preprocessing. Before model training, the acquired video clips need to be preprocessed, including reading video frames, adjusting color channels, resizing, and normalizing. The data preprocessing differs between the training and inference phases. During training, image frames need to be aligned with ground truth values, while during inference, 128 frames need to be acquired as input for each inference iteration.
[0083] S203: Model Training. Based on the dichroic reflectance model, the reflectance of each skin pixel in the recorded image sequence is defined as a time-varying function in the RGB channels:
[0084] In the above formula, specular reflection Defined as: .
[0085] in, Let be the unit color vector of the light source spectrum. This is the static part of the specular reflection. This represents the dynamic part of specular reflection, caused by motion.
[0086] Diffuse reflection Defined as: .
[0087] in, A unit color vector representing skin tissue. Indicates static reflection intensity. This indicates the relative pulse intensity of the RGB channels. This indicates the pulse signal.
[0088] Specular reflection and diffuse reflection Each part of a reflection contains a static component that does not change over time and a dynamic component that changes over time. Combining these two static components, the reflection can be represented as: It can be divided into inherent light intensity and the light intensity that changes over time , represented as:
[0089] Substituting the two equations into the original model, we obtain the skin reflection model: .
[0090] in, Represents the unit color vector of skin reflection. Indicates the intensity of reflection. Let be the unit color vector of the light source spectrum. This refers to the dynamic component induced by human movement in specular reflection. This indicates the relative pulse intensity of the RGB channels. Indicates pulse signal, This represents the quantization noise of the camera sensor.
[0091] The training of the neural network model based on RPPG signals is primarily aimed at learning parameters that can fit a skin reflex model. The model training uses the Adam optimizer, and the parameter updates are as follows: .
[0092] in, These are parameters to be optimized. It's the learning rate. It's a very small value added to prevent division by zero. and These are the estimates of the first and second moments of the gradient, respectively.
[0093] The heart rate detection result is the predicted heart rate amplitude. To maximize trend similarity and minimize peak position error, the NegPearson loss function is used to guide model training. .
[0094] in, The length of the signal. For the predicted RPPG signal, It is a true value.
[0095] S204: Model Conversion. To meet the requirements of edge computing device frameworks during model deployment, the trained PyTorch model file (pth) needs to be converted into an Open Neural Network Exchange (ONNX) model. Additionally, static graph optimization of the dynamic computation graph used during training can pre-optimize the execution path, reduce computational redundancy, and shorten inference time. Pseudo-inputs are randomly constructed based on the input dimensions of the pth model, while specifying the dimensions and instruction set version of the dynamic inputs.
[0096] S205: Pre- and Post-processing Rewriting. Most pre- and post-processing used during model training employs torch.tensor for data processing. To match edge computing devices, the data processing before model inference needs to be rewritten, replacing the tensor data type.
[0097] This invention provides a vision-based remote driver heart rate detection method. It uses a camera to capture driver images as input and trains a neural network model to detect the driver's heart rate. The model employs a dual-branch structure to fuse multi-scale and motion features, exhibiting robustness under conditions such as bumps, driver head movements, and uneven lighting. The key difference compared to existing technologies lies in:
[0098] 1. A combination of high-resolution optical and near-infrared cameras is used to acquire RGB and infrared images for model training. Compared to near-infrared cameras, optical cameras are more susceptible to changes in ambient light. Furthermore, due to the weakening of ambient light in the narrow-band near-infrared (especially 940 nm) band, combining two cameras avoids scorching light changes during driving that could overwhelm the weak RPPG signal, thus improving the model's robustness. Additionally, the combination of optical and infrared cameras offers advantages over using only a single infrared camera: because skin has lower reflectivity in the near-infrared band compared to visible light, a single infrared camera would result in a lower signal-to-noise ratio, making it more difficult for the model to learn RPPG features; furthermore, the resolution and contrast of infrared images are generally lower than high-quality visible light images, potentially affecting the accuracy of motion compensation algorithms based on facial feature point tracking. Therefore, this invention employs a combination of optical and infrared cameras to capture images.
[0099] 2. A motion compensation algorithm is employed to mitigate the impact of driver head movements. After obtaining the RPPG signal intensity by spatial averaging of the facial region of interest, noisy facial landmarks are discarded. Simultaneously, orthographic projection is used to project high-dimensional noise into a noise subspace. This method effectively suppresses high-noise scenes caused by rigid head movements during driving actions such as braking and accelerator pedal presses, as well as non-rigid facial movements induced by conscious actions like talking and eating. This significantly improves the model's robustness to interference in complex driving scenarios.
[0100] 3. A dual-branch network architecture is adopted, comprising a Motion sub-network and an Appearance sub-network. The Motion sub-network calculates the pixel changes between two consecutive frames by difference, and then performs feature extraction through multiple convolutional layers. The Appearance sub-network is an auxiliary network that mainly enhances the features of the current frame, feeding the features into the Motion sub-network at different stages of feature extraction for feature fusion with the Motion features. This architecture is beneficial for extracting rich detailed features and improving model accuracy.
[0101] In summary, this invention aims to improve the robustness and accuracy of remote driver heart rate detection technology under complex lighting conditions and head movements. By using a combination of optical and near-infrared cameras as sensors, and training a neural network model with the acquired images, coupled with a motion compensation algorithm and a dual-branch structure design, it can effectively cope with various complex environmental factors in driving scenarios, providing strong support for the intelligentization of in-vehicle health.
[0102] like Figure 8 The above is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. The electronic device 700 includes a processor 701 with one or more processing cores, a memory 702 with one or more computer-readable storage media, and a computer program stored in the memory 702 and executable on the processor. The processor 701 and the memory 702 are electrically connected.
[0103] The processor 701 is the control center of the electronic device 700. It connects various parts of the electronic device 700 via various interfaces and lines. By running or loading software programs and / or units stored in the memory 702, and by calling data stored in the memory 702, it executes various functions and processes data of the electronic device 700, thereby providing overall monitoring of the electronic device 700. The processor 701 can be a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a Network Processor (NP), etc., and can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application.
[0104] In this embodiment of the application, the processor 701 in the electronic device 700 loads the computer program corresponding to the process of one or more applications into the memory 702 according to the method or steps of the above embodiment, and the processor 701 runs the applications stored in the memory 702 to execute the above method.
[0105] According to an embodiment of the present invention, an electronic device acquires images of the driver using an optical camera and an infrared camera by performing the above-described method; the driver's images are input into a pre-trained neural network model, and the driver's heart rate value is output. Compared with related technologies, this solution utilizes a dual-modal camera (optical and near-infrared) to acquire images collaboratively, which is not limited by complex driving scenarios such as nighttime and backlighting, and can improve the accuracy of heart rate detection, thereby solving the problem of inaccurate estimation of the driver's heart rate in complex environments that arises with related technologies.
[0106] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, enables the computer to implement the vehicle control method described above. For example, the computer-readable storage medium may be the aforementioned memory including program instructions, which may be executed by a processor of an electronic device to implement or execute the methods, steps, and logic diagrams disclosed in the embodiments of this application.
[0107] This invention also provides a computer program product storing instructions that, when executed by a computer, cause the computer to implement the vehicle control method described above. For example, when executed by a computer, the instructions implement or execute the methods, steps, and logic diagrams disclosed in the embodiments of this application.
[0108] Embodiments of the present invention also provide a vehicle comprising the electronic equipment described above, or a processor, the processor being used to execute the methods described above. The vehicle may be a gasoline-powered vehicle, a plug-in hybrid electric vehicle, or a new energy vehicle, etc., and this specification does not specifically limit it.
[0109] According to embodiments of the present invention, a vehicle executes the above-described method via an electronic device, control system, or controller to acquire images of the driver using both an optical camera and an infrared camera; the driver's images are then input into a pre-trained neural network model, which outputs the driver's heart rate value. Compared to related technologies, this solution utilizes a dual-modal camera system (optical and near-infrared) to acquire images, which is not limited by complex driving scenarios such as nighttime or backlighting, thus improving the accuracy of heart rate detection and solving the problem of inaccurate estimation of the driver's heart rate in complex environments encountered by related technologies.
[0110] The above-described embodiments are only used to illustrate the technical solutions of applying the above methods to vehicles, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the method can also be used in motor vehicles, trains, and ships, etc., without causing the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0111] In one embodiment, the vehicle can be configured for fully or partially autonomous driving. For example, the vehicle can control itself while in autonomous driving mode, and can determine the current state of the vehicle and its surrounding environment through human intervention, determine the possible behaviors of at least one other vehicle in the surrounding environment, and determine the confidence level corresponding to the probability of that other vehicle performing a possible behavior, and control the vehicle based on the determined information. When the vehicle is in autonomous driving mode, it can be configured to operate without human interaction.
[0112] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0113] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," "optional example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0114] The embodiments, implementation methods, and related technical features of this application can be combined and substituted for each other without conflict.
[0115] The above are merely preferred embodiments of this application and are not intended to limit this application in any way. Although the descriptions of each embodiment in this application have different focuses, and the parts not described in detail in a certain embodiment can be referred to the relevant embodiments of other embodiments, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of this application without departing from the content of the technical solution of this application shall still fall within the scope of the technical solution of this application.
Claims
1. A method for monitoring a driver's heart rate, characterized in that, The method includes: Images of the driver are acquired using optical and infrared cameras; The driver's image is input into a pre-trained neural network model, which outputs the driver's heart rate value.
2. The method according to claim 1, characterized in that, Before inputting the driver's image into the pre-trained neural network model, the method further includes: Face detection is performed on the driver's image to obtain an image containing only the driver's face; Motion compensation is applied to the image of the driver's face to obtain a denoised face image.
3. The method according to claim 2, characterized in that, The process of performing motion compensation on the driver's facial image to obtain a denoised facial image includes: Spatial averaging of pixel intensities in the region of interest on the driver's face image yields a noisy, multidimensional RPPG signal matrix P. The multidimensional noise space matrix Q is determined based on the time-varying multidimensional horizontal motion signal matrix H, multidimensional vertical motion signal matrix V, and multidimensional background illumination signal matrix B in the driver's face image. The denoised face image feature matrix is calculated based on the multidimensional RPPG signal matrix P and the multidimensional noise space matrix Q.
4. The method according to claim 3, characterized in that, The step of calculating the denoised face image feature matrix based on the multidimensional RPPG signal matrix P and the multidimensional noise space matrix Q includes: The projection matrix of the multidimensional signal matrix P is obtained by orthogonally projecting the multidimensional RPPG signal matrix P onto the multidimensional noise space matrix Q. The denoised face image feature matrix is determined based on the difference between the multidimensional signal matrix P and the projection matrix.
5. The method according to claim 3, characterized in that, The method further includes: The driver's face image in each frame is subjected to target operation processing to obtain a multidimensional background illumination signal matrix B that varies with time; wherein, the target operation processing includes: selecting multiple background regions that do not contain facial regions from the driver's face image, dividing each background region into n*n pixel blocks, and determining the background illumination signal based on the intensity value of the pixel blocks, wherein n is greater than or equal to 2. And / or, determine the time-varying multidimensional horizontal motion signal matrix H and vertical motion signal matrix V.
6. The method according to claim 2, characterized in that, The step of performing face detection on the driver's image to obtain an image containing only the driver's face includes: The image of the driver is processed to obtain a grayscale image; A classifier is used to identify faces in the grayscale image and the grayscale image is cropped according to the detection box to output an image containing only the driver's face.
7. The method according to claim 1, characterized in that, The neural network model includes at least a first sub-network and a second sub-network. The first sub-network is used to perform difference calculation on the images of the previous and next frames to obtain the pixel changes between the previous and next frames, and the second sub-network is used to enhance the image features of the current frame.
8. The method according to claim 7, characterized in that, The first sub-network and the second sub-network are respectively composed of a convolutional layer with a kernel size of 3*3 and b convolutional layers with a kernel size of 2*2; wherein: a is greater than or equal to 4, b is greater than or equal to 2; and / or, the dual-branch neural network model further includes a prediction head, which includes a flattening layer and a linear layer. The flattening layer is used to flatten the three-dimensional tensor into a one-dimensional tensor for calculation, and the linear layer is used to project and map the one-dimensional tensor for prediction.
9. An electronic device, characterized in that, include: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the method of any one of claims 1-7.
10. A vehicle, characterized in that, include: The electronic device according to claim 9; Alternatively, a processor, said processor being configured to perform the method according to any one of claims 1-7.