Sample generation method for non-contact heart rate monitoring and heart rate monitoring method and system
Through the sample generation method of contactless heart rate monitoring, random stepping waveforms and local brightness changes are used to generate expanded training samples, which solves the problem of insufficient sample size in the prior art and improves the performance of the heart rate monitoring model under complex lighting conditions.
Patent Information
- Application Number
- CN202510453401.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The existing heart rate monitoring technology is insufficient in brightness changes scenarios, resulting in weak robustness and generalization ability of the heart rate monitoring model under complex lighting conditions.
By designing a sample generation method for contactless heart rate monitoring, the training samples are adjusted brightness by using random stepping waveforms and local brightness changes, and multiple expanded training samples with different brightness levels are generated to enhance the model's ability to resist ambient light changes.
It significantly improves the robustness and generalization ability of the heart rate monitoring model under complex lighting conditions during the actual driving of the vehicle, ensuring the continuous reliability of heart rate monitoring.
Smart Images

Figure CN119992258A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of physiological state monitoring, and in particular relates to a sample generation method for non-contact heart rate monitoring, a heart rate monitoring method and a system. Background Art
[0002] Heart rate is a reliable indicator for analyzing health status and a forward-looking clinical diagnostic tool. As one of the four major vital signs of the human body, the stability of heart rate directly reflects the quality of heart function.
[0003] The development of driver heart rate monitoring technology provides a new way to optimize the personalized driving experience. By analyzing heart rate data, the car can dynamically adjust the air conditioning, seats, lights, audio and other configurations in the car to relieve driving fatigue and adjust the driver's mood. In addition, by collecting the driver's heart rate data in real time, it can effectively identify potential risks such as fatigue driving, excessive tension or sudden health conditions, so as to issue early warnings or take the initiative to intervene in time, and it can also be linked with the telemedicine system to win precious time for emergency treatment of sudden health problems. The demand for driver heart rate monitoring reflects the deep integration of modern intelligent transportation technology in the fields of safety and health management. Heart rate monitoring technology can not only improve driving safety and the efficiency of driver health monitoring, but also promote the transformation of smart cars from traditional means of transportation to more humane and health-oriented travel partners through data sharing and intelligent linkage.
[0004] Electrocardiogram (ECG) is a gold standard for analyzing heart rate data. Electrodes attached to the skin are commonly used in clinical practice. The electrodes are connected to the ECG recorder through wires to obtain ECG signals. Another commonly used method is photoplethysmography (PPG), which is a method of detecting changes in blood volume in blood vessels under human skin tissue by photoelectric means to obtain pulse waves. This is an optical non-invasive technology. A light beam of a certain wavelength is emitted to the skin surface by a specific light source transmitter connected to the wrist, finger, etc. Since the blood volume in the artery fluctuates periodically with the relaxation and contraction of the heart, the contraction and expansion of blood vessels during each heartbeat will affect the transmission or reflection of light. Therefore, when the light passes through the skin tissue and then reflects to the photoelectric receiver, the intensity of the light will be attenuated to a certain extent. By converting the optical signal into an electrical signal and extracting the AC signal therein, the characteristics of blood flow can be reflected, and then the pulse wave signal can be obtained. Since the PPG signal can reflect many physiological information related to the cardiovascular system, and has the advantages of low cost, non-invasiveness, portability, and no cross infection, it has been widely used. Currently, many commercial wearable devices on the market, such as sports bracelets and smart watches, use this detection method.
[0005] In recent years, due to the impact of the epidemic, people's demand for telemedicine has continued to increase, and some image photoplethysmography (iPPG) based on video to obtain heart rate has emerged. iPPG is a method that can obtain PPG signals from captured videos. It has the advantages of non-contact measurement, easy implementation, and machine operation. In addition, remote PPG (rPPG) technology has also attracted much attention. This method can obtain PPG signals using only an ordinary digital camera without the need for contact probes and dedicated light sources.
[0006] However, the above existing methods all have certain limitations in the application process. Specifically, when using the traditional PPG method, since the contact sensor needs to maintain contact with the skin of the person being tested, it may cause discomfort to the driver and affect the normal driving process. When using the iPPG method to detect heart rate, it is necessary to place a standard white board as a background, and use the color change of the white board in the video to correct the skin color. The method of placing a white board does not conform to the actual driving environment, and cannot be used for measurements at a long distance. It is difficult to meet the monitoring needs in complex driving processes and is only suitable for research in laboratory environments. The existing rPPG method has significant advantages over other methods, but when changes in ambient light affect the signal quality, the system is prone to output erroneous or inaccurate heart rate data, and its ability to resist changes in ambient light is weak. Therefore, there is an urgent need for a method that can improve the adaptability of raw data in brightness change scenarios and improve the confidence of heart rate monitoring results. Summary of the invention
[0007] In view of this, the present invention aims to provide a sample generation method for non-contact heart rate monitoring that can significantly increase the training samples of the heart rate monitoring model, and on this basis, a non-contact heart rate monitoring method and system that can resist changes in ambient light are designed, which effectively solves the problem of insufficient sample number of original data in brightness change scenarios, and improves the robustness and generalization ability of the heart rate monitoring model under complex lighting conditions during actual vehicle driving.
[0008] To achieve the above object, the technical solution created by the present invention is implemented as follows: The first aspect of the invention provides a sample generation method for non-contact heart rate monitoring, comprising: for each frame of training samples for heart rate monitoring, setting a time step and a brightness change interval within each time step, randomly sampling a brightness change value within the brightness change interval at each time node, superimposing the brightness change value with the global brightness value of the training sample at the previous time node, and generating a random step waveform of the training sample; The training samples are partitioned and located, and the local brightness change of each block is expressed as: ; in, Indicates the location The local brightness value of the block at time t, represents the brightness amplitude, represents the brightness change frequency, Indicates phase shift; The brightness amplitude of each block , brightness change frequency and phase shift Random assignment; The brightness of each block of each frame of training samples is adjusted according to the random step waveform and the local brightness change, and multiple expanded training samples with different brightness are generated for each frame of training samples.
[0009] Preferably, the random step waveform of the training sample is: ; in, represents the global brightness value of the training sample at time t, represents the global brightness value of the training sample at time t-1, Indicates the amount of brightness change within each time step.
[0010] Preferably, the brightness change value is obtained by random sampling in the uniform distribution of the brightness change interval. .
[0011] Preferably, the block is a ROI block of a training sample.
[0012] Preferably, the random step waveform and the local brightness change are fused to generate a dynamic brightness change as follows: ; in, represents the dynamic brightness change of the training sample, Represents the original brightness of the training sample.
[0013] Preferably, the method further includes: superimposing random noise during the brightness adjustment of each block, fusing the random step waveform, the local brightness change and the random noise, and generating a dynamic brightness change as follows: ; in, represents the dynamic brightness change of the training sample, represents the original brightness of the training sample, represents random noise.
[0014] The second aspect of the present invention provides a non-contact heart rate monitoring method, comprising: Collect facial video images of the driver while driving the vehicle; For each static image frame in the facial video image, the facial ROI features are selected, the ROI area is obtained by locating the face detector, and the ROI area is divided into multiple ROI blocks; Detect the light intensity, occlusion, and motion artifacts of the ROI block of each static image frame. When the light intensity, occlusion, and motion artifacts meet the set conditions, extract the static image frame as the sample to be tested; The trained heart rate detection model is used to monitor the driver's heart rate in real time. The training sample set of the heart rate detection model is: using the sample generation method of non-contact heart rate monitoring, an expanded training sample is generated for each training sample, and the training sample and the expanded training sample are combined into a training sample set.
[0015] The third aspect of the present invention provides a non-contact heart rate monitoring system, comprising: A video acquisition module for acquiring facial video images of the driver during driving of the vehicle; A positioning and partitioning module is used to select the ROI features of the face for each static image frame in the facial video image, obtain the ROI area through face detector positioning, and divide the ROI area into multiple ROI blocks; The signal quality analysis and processing module is used to detect the light intensity, occlusion, and motion artifacts in the ROI area of each frame of static image. When the light intensity, occlusion, and motion artifacts meet the set conditions, the static image is extracted as a sample to be tested; and for each sample to be tested, the signal quality analysis and processing module performs global average pooling on all ROI blocks, compresses the pixel values of each ROI block into a time series, and constructs the sample to be tested into a time-space mapping matrix based on the time series; A heart rate monitoring model for real-time monitoring of a driver's heart rate completes model training using expanded training samples and training samples generated by a sample generation method for non-contact heart rate monitoring. After the heart rate monitoring model inputs the spatiotemporal mapping matrix of the sample to be tested, it calculates and obtains the corresponding predicted heart rate value.
[0016] Preferably, it also includes a confidence detection module, which performs a physiological range test and a time consistency test on the predicted heart rate value. When the predicted heart rate value is consistent with the heart rate range of a healthy adult and the average heart rate fluctuation is less than the set threshold, the predicted heart rate value calculated by the heart rate monitoring model is output; otherwise, the signal quality analysis and processing module is fed back to the signal quality analysis and processing module to reprocess the test samples corresponding to the predicted heart rate values that do not meet the output.
[0017] Preferably, the video acquisition module includes a visible light camera; Or, the video acquisition module includes a visible light camera and an infrared camera, and the priority of the visible light camera and the infrared camera is determined according to the ratio of the number of samples to be tested extracted by the signal quality analysis and processing module to the number of static images, and the determination results include: using the visible light camera to collect the facial video image of the driver during the driving of the vehicle, using the infrared camera to collect the facial video image of the driver during the driving of the vehicle, and using the visible light camera and the infrared camera to collect the facial video image of the driver during the driving of the vehicle at the same time; When the visible light camera and the infrared camera are used to simultaneously collect facial video images of the driver during driving, the features of the RGB image collected by the visible light camera and the IR image collected by the infrared camera are extracted through a dual-branch convolutional neural network respectively, and the attention mechanism is used to fuse them, and the contribution ratio of the RGB and IR modalities is dynamically adjusted to form a fused image, and the fused image is used as the sample to be tested.
[0018] Compared with the prior art, the invention can achieve the following beneficial effects: The present invention innovatively designs a fluctuating brightness method, which superimposes brightness fluctuations on the original samples and quickly generates brightness changes through mathematical functions. It can easily expand the data set without relying on additional shooting equipment or complex scene settings. It not only effectively solves the problem of insufficient real data in practical applications, but also avoids the occurrence of overfitting in the model training process, significantly improves the generalization ability of the model, and makes up for the defect of the lack of complex lighting sample images in the original samples, and enhances the training model's ability to resist ambient light changes. In different driving environments and conditions, the heart rate monitoring model can maintain stable performance, providing drivers with continuous and reliable health monitoring services.
[0019] The present invention also makes an innovative design to the traditional PPG method. The traditional PPG method requires the use of electrode patches or light source probes to contact the human body to obtain heart rate signals, while the present invention can accurately extract heart rate through video signals without any physical contact. This not only avoids the discomfort caused by the driver wearing the device, ensures the naturalness and comfort of the driving process, but is also particularly suitable for remote monitoring of the driver's health status, providing a strong guarantee for driving safety. Compared with the iPPG method, the present invention is more competitive in terms of equipment cost and ease of application. The iPPG method requires specific light sources or dedicated hardware to obtain high-quality signals, which increases the complexity and cost of the equipment. The present invention only requires common smart devices such as ordinary cameras on the dashboard to collect video images of the driver's face, without the need for additional special hardware, to achieve accurate measurement of heart rate. It has the advantages of low cost and low threshold, and is easy to be widely used in various types of vehicles, which promotes the popularization of driving health monitoring technology.
[0020] In terms of signal quality assurance, the present invention adopts a multi-dimensional detection method to evaluate and optimize the signal from multiple angles such as brightness mean and standard deviation, key point texture analysis, Laplace transform, etc., and monitors the influence of light intensity, occlusion and motion artifacts on signal quality in real time, effectively optimizes and processes unqualified signals, and ensures the high accuracy and robustness of the measurement results through comprehensive and fine signal quality control. In addition, the present invention proposes for the first time to use dual cameras for image acquisition. According to the extraction ratio of the samples to be tested fed back by the signal quality analysis and processing module, the visible light camera and the infrared camera can be flexibly switched, and the infrared camera is used to make up for the shortcomings of the visible light camera in low light environments, which greatly improves the image acquisition capability of the video acquisition module in low light environments such as night and tunnels, as well as other complex environments where the visible light camera cannot accurately capture facial images. Moreover, through the fusion of the extracted features of the RGB image and the IR image, the fused image can accurately reflect the physiological characteristics of the driver.
[0021] Compared with the existing iPPG and rPPG methods, the present invention places more emphasis on adaptability to complex lighting changes in real driving environments. It can stably provide reliable monitoring data regardless of strong light, weak light or dynamic lighting conditions, greatly improving the practicality and reliability of in-vehicle applications.
[0022] The present invention innovatively adopts the space-time mapping matrix method to effectively integrate and model the spatial and temporal information in the video signal. This signal processing method can capture richer and more comprehensive heart rate characteristics, deeply explore the subtle changes in heart rate signals in different regions and time points, and further improve the accuracy and credibility of heart rate monitoring results. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The drawings constituting part of the present invention are used to provide a further understanding of the present invention. The exemplary embodiments and descriptions of the present invention are used to explain the present invention and do not constitute an improper limitation on the present invention. In the drawings: Figure 1 is a structural framework diagram of a non-contact heart rate monitoring system provided according to an embodiment of the present invention; Figure 2 is a schematic diagram of the in-vehicle position design of a video acquisition module provided by an embodiment of the present invention; Figure 3 Schematic diagram of a dual-camera video acquisition method provided according to an embodiment of the present invention. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical scheme and advantages of the invention clearer, the invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the invention and do not constitute a limitation to the invention. Similar components in different embodiments use associated similar component numbers. In the following embodiments, many detailed descriptions are to enable the invention to be better understood. However, those skilled in the art can easily recognize that some of the features can be omitted in different situations, or can be replaced by other components, materials, and methods. In some cases, some operations related to the invention are not shown or described in the specification, in order to avoid the core part of the invention being overwhelmed by too much description, and for those skilled in the art, it is not necessary to describe these related operations in detail, and they can fully understand the related operations according to the description in the specification and the general technical knowledge in the art.
[0025] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to form various implementation methods. At the same time, the steps or actions in the method description can also be interchanged or adjusted in a manner that is obvious to those skilled in the art. Therefore, the various sequences in the specification and the drawings are only for the purpose of clearly describing a certain embodiment and are not meant to be a necessary sequence, unless otherwise specified that a certain sequence must be followed.
[0026] In the description of the invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the drawings, and are only for the convenience of describing the invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first", "second", etc. may explicitly or implicitly include one or more of the features. In the description of the invention, unless otherwise specified, the meaning of "multiple" is two or more.
[0027] In the description of the invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installation", "connection" and "connection" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the invention can be understood according to specific circumstances.
[0028] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments.
[0029] Embodiment 1: In embodiment 1 of the present invention, a sample generation method for contactless heart rate monitoring is provided, which is mainly used to solve the problem that the number of training samples of traditional heart rate monitoring models in intelligent driving applications is insufficient and cannot cope with complex lighting changes. By expanding the amount of training sample data, the ability of the trained model to resist ambient light changes is improved. The sample generation method specifically includes the following steps: First, in the training process of the non-contact heart rate monitoring model, the public dataset UBFC-RPPG Dataset is usually used. The dataset contains synchronously collected facial videos and heart rate data, which provides rich learning resources for the model. By using this dataset to train the heart rate monitoring model, the heart rate monitoring model can learn effective features and patterns to predict heart rate values from facial videos. However, the amount of data in the public dataset UBFC-RPPG Dataset is limited, especially the data samples in complex light environments are limited, so when using it for model training, the model's ability to resist changes in ambient light is relatively weak. Therefore, the embodiment of the present invention mainly expands the training samples used for heart rate monitoring in each frame of the public dataset UBFC-RPPG Dataset, and supplements the data samples of complex ambient light changes.
[0030] For each frame of training samples for heart rate monitoring, a random step waveform is first generated for the brightness of the entire picture to simulate the random fluctuations in the brightness of the entire picture, reflecting the changes in the light and shadow of the driver's face when the vehicle enters the tunnel or exits the shadow area. Specifically, the expression of the random step waveform of the training sample is established: ; in, represents the global brightness value of the training sample at time t, represents the global brightness value of the training sample at time t-1, Indicates the amount of brightness change within each time step.
[0031] It should be noted that t-1 and t represent two time nodes, and the time step between any two time nodes is set manually. The brightness change interval within each time step is set according to the set time step, and the interval is expressed as , Indicates the step range of brightness change, which is used to control the brightness fluctuation amplitude.
[0032] In the process of generating random step waveforms, first of all, the global brightness value needs to be Initialize, that is: .
[0033] For each time node t, from the brightness change interval The brightness change value is obtained by random sampling from the uniform distribution of , the brightness change value The global brightness value of the training sample at the previous time node By superimposing, you can get the global brightness value at the current moment. By repeatedly iterating the above process, a random step waveform with random brightness changes and a change amplitude limited by the brightness change amplitude can be obtained. Subsequently, the random step waveform is superimposed with the original training sample to obtain a set of expanded samples with constantly changing global brightness.
[0034] In order to ensure that the subsequent expansion of training samples conforms to the actual scene, it is further necessary to simulate the local illumination changes of the driver's face image to reflect the dynamic characteristics of local light and shadow on the driver's face. Specifically, in Example 1 of the present invention, each training sample needs to be partitioned and positioned, and the training sample is divided into multiple blocks, and the block position is positioned. The positioning coordinates can agree to select the coordinates of a certain pixel point in the block, and the coordinates of the central pixel point in the block are usually selected as the position representation of the entire block. As for the size of the block, it can be determined according to the amount of sample data that needs to be expanded. Since the heart rate is mainly predicted based on the driver's facial features during the driver's heart rate monitoring process, the cheek area is usually selected as the region of interest for heart rate monitoring, that is, the ROI area. Therefore, in the process of adjusting the brightness of the local area, it is also possible to adjust and expand only the brightness of the ROI area. Here, the ROI area refers to the entire driver's cheek part, and the specific ROI area extraction can be achieved through existing image processing methods. In order to achieve accurate adjustment of the illumination on the driver's cheek, it is also necessary to partition and locate the ROI area. Assume that the ROI area is divided into n rectangular ROI blocks. For each ROI block, the local dynamic waveform of its local brightness change is expressed as: ; in, Indicates the location The local brightness value of the ROI block at time t, represents the brightness amplitude, represents the brightness change frequency, Represents the phase offset, i.e. the initial phase of the brightness curve.
[0035] The block brightness amplitude in the above formula , brightness change frequency and phase shift is a variable, so the brightness amplitude can be , brightness change frequency and phase shift Random assignment is performed to adjust the local brightness change of each block. The local brightness change of the ROI block of each training sample is expressed as , by assigning different brightness amplitudes , brightness change frequency and phase shift , the randomness and diversity of the brightness change in the ROI area can be expanded. In order to meet the light changes in actual driving scenes, the local brightness change adjustment of the ROI block can also be set with a certain rule. For example, the ROI area is adjusted from the first ROI block in the upper left to the last ROI block in the lower right, and the brightness fluctuation continues to increase, simulating local light and dark changes.
[0036] After setting the global brightness change as a random step waveform and the local brightness change as a local dynamic waveform, the brightness of each ROI block of each frame of training samples can be adjusted according to the random step waveform and the local brightness change. The brightness of each ROI block of the training sample is adjusted according to the time step to simulate the changes in driver's facial illumination caused by tunnels, sunlight, tree shadows, etc. during actual driving. The training samples are adjusted after the random step waveform and the local brightness change are fused to generate the expanded training samples after brightness adjustment. The dynamic brightness change of the training samples is: ; in, represents the dynamic brightness change of the training sample, that is, the brightness of the expanded training sample, Represents the original brightness of the training sample.
[0037] In order to further simulate the slight jitter of the illumination or the random change of the ambient illumination and enhance the diversity and authenticity of the data, it is also necessary to superimpose random noise. In the embodiment of the present invention, the random noise is set as follows: For any time t, the brightness noise value is , from the normal distribution is randomly generated from . Among them, the variance According to the light interference intensity setting, after generating the noise matrix, the randomly generated brightness noise value It is superimposed on each ROI block in a certain proportion to simulate the slight jitter of the light or the random change of the ambient light, thereby further enhancing the diversity and authenticity of the data. Specifically, the random step waveform, local brightness change and random noise are fused to generate dynamic brightness changes as follows: ; in, represents the dynamic brightness change of the training sample, that is, the brightness of the expanded training sample after superimposing random noise. represents the original brightness of the training sample, represents random noise.
[0038] The above method is a general method for each training sample. For each frame of training samples initially obtained, the brightness can be adjusted frame by frame according to the above method. Each frame of training samples can generate multiple expanded training samples with dynamic brightness changes. It should be noted that the label of the expanded training sample is consistent with the data label of the corresponding training sample, that is, the heart rate monitoring data of the expanded training sample is equal to the heart rate monitoring data of the original training sample.
[0039] As an optional embodiment, the expanded training samples may be generated only for part of the training samples.
[0040] Finally, the training samples and the expanded training samples are combined into a training data set, which can be used to train the heart rate monitoring model. Due to the effective expansion of the training data, the trained heart rate monitoring model can greatly improve the ability to resist ambient light changes during the heart rate detection process.
[0041] Embodiment 2: Based on the above-mentioned sample generation method for non-contact heart rate monitoring, the present invention further proposes a non-contact heart rate monitoring method, comprising: The vehicle's video acquisition device is used to collect the driver's facial video images in real time during the driving process. The video acquisition device can generally use a high-resolution, high-frame-rate camera, specifically one or more of a visible light camera, an infrared camera, a multi-spectral camera, etc. And set an appropriate video resolution and frame rate, such as 1080p resolution and 30 frames per second frame rate, to ensure image clarity and continuity.
[0042] For each static image in the facial video image, an open source face detector with high accuracy and good real-time performance is used to perform face detection on the static image to accurately locate the facial area. The facial ROI features for the driver's heart rate detection are selected, and the ROI area is obtained by locating the open source face detector. In the embodiment of the present invention, the driver's cheeks are selected as ROI features based on the different distribution characteristics of the heart rate signal on the face. These areas have rich blood flow and more obvious heart rate signals.
[0043] After obtaining the ROI area, the located ROI area is further divided into multiple small ROI blocks, such as evenly dividing the cheek area into multiple rectangular ROI blocks. The size of each ROI block can be adjusted according to actual needs and computing resources, such as dividing the ROI area into n ROI blocks, each ROI block includes m pixels.
[0044] After partition positioning, the light intensity, occlusion, and motion artifacts of the ROI block of each static image frame are detected. Specifically, the brightness mean and standard deviation of all ROI blocks in the ROI area are calculated, and compared with the preset brightness threshold range to determine whether the light is too strong or too weak. If the brightness mean is within the preset brightness threshold range, the light intensity of the static image frame is judged to be qualified and heart rate monitoring can be performed; if the brightness mean exceeds this brightness threshold range, the light is considered unqualified.
[0045] It is judged whether there is occlusion affecting heart rate monitoring in each static image frame, and whether there is occlusion is determined by analyzing the texture consistency of the area near the key points of the ROI area. For example, the small area image around the key points such as the left and right cheekbones and the chin edge points is detected. If the brightness of the area near the key points changes suddenly or the expected texture features are missing, it is determined to be occluded; if the texture features of the small area image around the key points are basically consistent, it is considered to be unoccluded.
[0046] It is determined whether there are motion artifacts in each static image that affect heart rate monitoring. The clarity of the ROI block is detected using Laplace transform, and whether there are motion artifacts is determined based on the set clarity threshold. For example, the variance of the Laplace transform result is calculated. If the variance is lower than the threshold, the image is judged to be blurred and there are motion artifacts; otherwise, there are no motion artifacts.
[0047] Only when the light intensity, occlusion, and motion artifacts meet the set conditions, that is, the light intensity meets the set threshold, there is no occlusion, and there are no motion artifacts, the static image frame is extracted as the sample to be tested to ensure the reliable quality of the data input into the heart rate detection model.
[0048] After the samples to be tested are extracted, the trained heart rate detection model is used to monitor the driver's heart rate in real time. The training sample set in the model training process is expanded and generated by the method in Example 1. The sample generation method for non-contact heart rate monitoring in Example 1 is used to generate an expanded training sample for each training sample. The training sample and the expanded training sample are combined to form a training sample set.
[0049] The sample to be tested is input into the trained heart rate detection model, and the driver's predicted heart rate value can be obtained through the model output.
[0050] Example 3: Based on the sample generation method of Example 1 and the heart rate monitoring method of Example 2, please refer to Figure 1 Embodiment 3 of the present invention also proposes a non-contact heart rate monitoring system, which mainly includes a signal input and processing part and a heart rate detector (heart rate prediction and output), wherein the signal input and processing part includes a video acquisition module, a positioning and partitioning module, and a signal quality analysis and processing module.
[0051] See also Figure 2 The main functional component of the video acquisition module is a camera installed on the dashboard, usually a visible light camera is used to obtain the driver's facial video image.
[0052] The main function of the positioning and partitioning module is to locate the facial ROI features through the open source face detector. During the heart rate monitoring process, the ROI features usually select the cheek position. Therefore, the cheek ROI area can be located through the open source face detector. The ROI area positioning method is: each static image in the facial video image is detected and divided frame by frame to obtain a stable ROI area positioning. After the ROI area is located, it is adjusted to a rectangle, and the rectangular ROI area is divided to obtain n rectangular ROI blocks.
[0053] The main functions of the signal quality analysis and processing module include detecting the light intensity, occlusion, and motion artifacts on the driver's face, and converting the samples to be tested that meet the environmental conditions into a time-space mapping matrix for input into the heart rate detection module for processing.
[0054] Lighting conditions have a significant impact on the quality of video signals. Extreme lighting conditions can affect the accuracy of heart rate monitoring. Therefore, it is necessary to detect whether the light is too dark or too bright. The embodiment of the present invention uses the calculation of the brightness mean and standard deviation of each frame of video image to determine the light intensity level. The brightness threshold range is specifically set. If the calculated brightness mean exceeds this range, the light is judged to be unqualified. The specific calculation method of the brightness mean is: ; in, is the mean brightness, express The brightness of the block, H represents the number of ROI blocks contained in the x direction of each static image frame, and W represents the number of ROI blocks contained in the y direction of each static image frame. is the total number of ROI blocks in each static image frame.
[0055] Facial occlusion will lead to inaccurate ROI area extraction results, which will interfere with the heart rate estimation results. Therefore, occlusion detection is a key step to ensure the quality of the input signal. In the embodiment of the present invention, detection based on the area near the key points is adopted to detect the key points related to the cheek area such as the left and right cheekbone areas and the chin edge points, and the position and distribution of these points are used to infer the integrity of the cheek area. The specific method is to locate the key points near the left and right cheeks, extract small rectangular area images around the key points, and analyze the texture consistency of these areas. If the brightness of the area near the key points changes suddenly, or the expected texture features are missing, it is determined to be cheek occlusion, and it is difficult to obtain an accurate predicted heart rate value through the static image of this frame; if the texture of the small rectangular area image extracted around the key points is basically the same, it is considered that the face of the static image of this frame is not occluded. The specific implementation method of the above-mentioned facial occlusion judgment can be implemented by an existing image processing method or feature comparison model.
[0056] Motion artifacts are the phenomenon that images are blurred due to the rapid movement of the driver, which can be evaluated by detecting the clarity of the image. The embodiment of the present invention uses Laplace transform to detect the clarity of the image, sets the clarity threshold according to the use requirements, and determines that the image is blurred if the variance of the transform result is lower than the threshold. The specific calculation method is: The static image in the original facial video image is converted from RGB to grayscale. The grayscale conversion formula based on the ITU-R BT.601 standard is: ; in, Indicates the pixel grayscale value converted from RGB to grayscale.
[0057] Apply Laplace transform to the grayscale image obtained after conversion and calculate the second-order gradient of each pixel : ; in, Denotes Laplace transform.
[0058] Calculate the pixel value variance of the Laplace transform result , The calculation formula is: ; in, represents the mean of the Laplace transform results, , Represents the number of pixels in the x direction of the grayscale image, Represents the number of pixels in the x direction of the grayscale image, Represents the total number of pixels of the grayscale image.
[0059] Determine pixel value variance Is it less than the set clarity threshold? If the pixel value variance If the value is less than the clarity threshold, the image is judged to be blurred and there are motion artifacts; otherwise, the image is judged to be clear and there are no motion artifacts.
[0060] When the three tests of light intensity, occlusion and motion artifact are all qualified, it is considered that the environmental condition of the static image frame is qualified, and the static image frame is extracted as the sample to be tested.
[0061] After obtaining the test samples with qualified environmental conditions, the signal quality analysis and processing module compresses the pixel data into a time series by performing global average pooling on all ROI blocks in the test samples, so as to generate a spatiotemporal mapping diagram for each face video sequence and obtain a space-time representation matrix, i.e., a spatiotemporal mapping matrix. This process converts complex image data into concise time series data, which is convenient for the subsequent heart rate detection model to perform feature extraction and heart rate prediction.
[0062] For T frames of samples to be tested, each frame of samples has three color channels: R, G, and B, and each channel is a two-dimensional image. Indicates the position of one of the three channels R, G, and B in the sample to be tested in the tth frame Each ROI block obtained by the positioning and partitioning module is averaged and the pixel values of all ROI blocks are averaged to obtain a time series.
[0063] For any of the three channels R, G, and B, the average pixel value of each ROI block at time t It is expressed as: ; Among them, i represents the index of the i-th ROI block, Represents the number of pixels in the i-th ROI block. Through average pooling, the pixel dynamic information of each ROI block of the T-frame sample to be tested is simplified to a time series , the total length of the time series is T.
[0064] Perform maximum and minimum normalization on each time series and transform the time series Narrowed to the range of [0,225], the maximum and minimum normalization calculation formula is:
[0065] in, Represents the maximum and minimum normalized time series values, represents the maximum pixel value of the i-th ROI block at all time points, Represents the minimum pixel value of the i-th ROI block at all time points.
[0066] By normalizing all signal values to a uniform range, the difference in signal value amplitudes between different ROIs or different videos can be eliminated to a certain extent.
[0067] Furthermore, the T-frame samples to be tested are converted into R, G, and B three-channel spatiotemporal mapping. The constructed spatiotemporal mapping is represented by a data matrix of n×T for each channel, where each row corresponds to the time series of an ROI block, and each column represents the value of all ROI blocks at the same time point. The output spatiotemporal map is represented by the ROI time series of the three channels of R, G, and B being spliced into an n×T×3 three-dimensional data structure, where n is the number of ROIs, T is the number of time frames, and 3 represents the three channels of R, G, and B. For each RGB channel, The normalized time series of ROIs are arranged in rows to construct T. The matrices of the three channels R, G, and B are stacked to obtain a spatiotemporal mapping matrix of size n×T×3. This matrix is used as a compressed and normalized representation of the original sample to be tested and input into the heart rate detector. The heart rate monitoring model in the heart rate detector extracts the features of the spatial and temporal dimensions from it, and then predicts the heart rate value.
[0068] The heart rate monitoring model used in the embodiment of the present invention is based on the original heart rate regression model, which is obtained by training the training sample set obtained by expanding the sample generation method of the non-contact heart rate monitoring in Example 1. For the training of the heart rate monitoring model, the embodiment of the present invention uses the mean square error loss function for optimization, thereby optimizing the model training effect. The loss function of the model is: ; Where T is the time frame number of the heart rate signal, is the predicted heart rate value measured by the model at time t, is the real heart rate signal of the driver at time t. When L is stable and reaches the set value, it indicates that the model training is completed.
[0069] The n×T×3 spatiotemporal mapping matrix of the sample to be tested is used as the input of the heart rate monitoring model. A 3D convolutional network is used to perform convolution operations in both the time and space dimensions to extract the joint representation of dynamic and static features. The extracted features are downsampled through the pooling layer to reduce the computational complexity while retaining the key time-space information. The extracted time-space features are input into the Transformer time series modeling module to capture the dynamic dependencies of the time dimension and model the changes between frames. The features of each frame are mapped and the corresponding predicted heart rate value is output to estimate the heart rate of each frame of the sample to be tested.
[0070] After the predicted heart rate value is estimated by the heart rate monitoring model, it is also necessary to test the predicted heart rate value through the confidence detection module to determine whether the predicted heart rate value is credible. In the embodiment of the present invention, the confidence detection module tests the predicted heart rate value through two aspects: physiological range test and time consistency test. Specifically, the physiological range test is used to check whether the predicted heart rate value is within a reasonable physiological range. According to the heart rate range of healthy adults at rest and during strenuous exercise, the embodiment of the present invention stipulates that the reasonable physiological range is When the predicted heart rate value is within this interval, it is considered that the predicted heart rate value is in line with the physiological range and has a high confidence level. Otherwise, it is considered that the predicted heart rate value has a low confidence level.
[0071] The time consistency test is to evaluate whether the predicted heart rate value changes smoothly by calculating the average heart rate in the sliding window. The calculation formula for the average heart rate in the sliding window is: ; Where N represents the size of the sliding window.
[0072] If the average heart rate fluctuation in the sliding window exceeds the set threshold, the confidence of the predicted heart rate value in this window is considered to be low; if the average heart rate fluctuation in the sliding window does not exceed the set threshold, the confidence of the predicted heart rate value in this window is considered to be high.
[0073] When both the physiological range test and the time consistency test obtain high confidence, the predicted heart rate value is considered to be close to the driver's actual heart rate, and the predicted heart rate value is output through the output module. In addition, the reliability detection module can use a dynamic physiological range based on an individualized model for testing, or introduce more complex statistical test methods to improve the personalized adaptation of the driver, which can enhance the signal robustness, but the dependence on historical data is enhanced, and the detection judgment is more complicated.
[0074] As a preferred embodiment, under extreme conditions such as insufficient lighting at night, in tunnels or shadow areas, it is difficult to obtain a clear driver's facial image using only a single visible light camera according to the traditional design, that is, there is a problem with the input to be detected, which further leads to a low reliability of the predicted heart rate value. To solve this problem, the video acquisition module is equipped with an additional infrared camera on the basis of the traditional single visible light camera. Figure 3 The dual-camera collaborative working mode shown.
[0075] When using the dual-lens collaborative working mode, it is necessary to determine whether it is in a low-light scene based on the analysis results of the signal quality analysis and processing module, and then determine the priority of the visible light camera and the infrared camera. Specifically, the visible light camera is an RGB camera with a higher resolution and frame rate, which is mainly used to collect high-quality color images of the driver under normal lighting conditions; the infrared camera has near-infrared or far-infrared imaging technology, supports low-light conditions, and is used to clearly capture facial contours and feature points in low-light or even no-light environments.
[0076] Therefore, the ratio of the number of samples to be tested to the number of static images extracted by the signal quality analysis and processing module can reflect the lighting conditions over a period of time, and the priority of the visible light camera and the infrared camera can be determined based on the lighting conditions. In a scene with sufficient lighting, the visible light camera can obtain clear facial video images, so only the visible light camera is used to collect the driver's facial video images during driving; in a scene with insufficient lighting, the visible light camera cannot obtain clear facial video images, so the visible light camera and the infrared camera are used to collect the driver's facial video images during driving. The two modes are dynamically switched according to the feedback from the signal quality analysis and processing module, and the data priority of the visible light and infrared cameras is dynamically switched.
[0077] When the visible light camera and the infrared camera simultaneously collect the driver's facial video images, the data collected by the two cameras are synchronized to avoid information lag or loss. A dynamic weight allocation algorithm is used to ensure synchronization and delay-free camera switching. Then, the RGB images collected by the visible light camera and the IR images collected by the infrared camera are aligned, enhanced, denoised, and segmented. The attention mechanism feature-level fusion method is used to extract features from the RGB images collected by the visible light camera and the IR images collected by the infrared camera through a dual-branch convolutional neural network, and the attention mechanism is used for fusion. The contribution ratio of the RGB and IR modalities is dynamically adjusted under complex lighting conditions, and finally a unified sample data is formed.
[0078] As a preferred embodiment, other cameras may also be added. For example, a multi-spectral camera device may be used to collect facial videos in multiple wavelength ranges such as visible light and near-infrared light, and the PPG signal may be extracted by separating and analyzing signals of different wavelengths. Multi-spectral data provides richer information, can enhance the robustness of signal extraction, and can also work under partial occlusion or complex light conditions, but multi-spectral cameras are expensive, data processing is complex, and have high requirements for computing resources and technical barriers.
[0079] As a preferred embodiment, the heart rate can also be derived by using ultrasonic millimeter wave radar technology to detect tiny movements such as skin vibrations on the surface of the human body or internal chest movements. This method can also achieve non-contact measurement and does not rely on optical equipment. It has strong tolerance to occlusion, and the millimeter wave radar can penetrate clothing to monitor the driver's heart rate. However, the equipment is not simple and portable, and the tiny movement signal is easily interfered by breathing or environmental noise, requiring a complex signal separation algorithm.
[0080] In short, the above description is only a preferred embodiment of this specification and is not intended to limit the protection scope of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included in the protection scope of this specification.
[0081] The systems, devices, modules or units described in one or more of the above embodiments may be implemented by a computer chip or entity, or by a product having a certain function. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0082] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0083] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0084] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A sample generation method for non-contact heart rate monitoring, characterized in that: include: For each frame of training samples for heart rate monitoring, a time step and a brightness change interval within each time step are set, a brightness change value is randomly sampled within the brightness change interval at each time node, and the brightness change value is superimposed with the global brightness value of the training sample at the previous time node to generate a random step waveform of the training sample; The training samples are partitioned and located, and the local brightness change of each block is expressed as: ; in, Indicates the location The local brightness value of the block at time t, represents the brightness amplitude, represents the brightness change frequency, Indicates phase shift; The brightness amplitude of each block , brightness change frequency and phase shift Random assignment; The brightness of each block of each frame of training samples is adjusted according to the random step waveform and the local brightness change, and multiple expanded training samples with different brightness are generated for each frame of training samples.
2. The sample generation method for non-contact heart rate monitoring according to claim 1, characterized in that: The random step waveform of the training sample is: ; in, represents the global brightness value of the training sample at time t, represents the global brightness value of the training sample at time t-1, Indicates the amount of brightness change within each time step.
3. The sample generation method for non-contact heart rate monitoring according to claim 2, characterized in that: Randomly sample the uniform distribution of the brightness change range to obtain the brightness change value .
4. The sample generation method for non-contact heart rate monitoring according to claim 1, characterized in that: The block is the ROI block of the training sample.
5. The sample generation method for non-contact heart rate monitoring according to claim 1, characterized in that: The random step waveform and local brightness change are combined to generate dynamic brightness changes: ; in, represents the dynamic brightness change of the training sample, Represents the original brightness of the training sample.
6. The sample generation method for non-contact heart rate monitoring according to claim 1, characterized in that: Also includes: In the process of brightness adjustment for each block, random noise is superimposed, and the random step waveform, local brightness change and random noise are fused to generate dynamic brightness change: ; in, represents the dynamic brightness change of the training sample, represents the original brightness of the training sample, represents random noise.
7. A non-contact heart rate monitoring method, characterized in that: include: Collect facial video images of the driver while driving the vehicle; For each static image frame in the facial video image, the facial ROI features are selected, the ROI area is obtained by locating the face detector, and the ROI area is divided into multiple ROI blocks; Detect the light intensity, occlusion, and motion artifacts of the ROI block of each static image frame. When the light intensity, occlusion, and motion artifacts meet the set conditions, extract the static image frame as the sample to be tested; The trained heart rate detection model is used to monitor the driver's heart rate in real time. The training sample set of the heart rate detection model is: using the sample generation method for non-contact heart rate monitoring as described in any one of claims 1 to 6, an expanded training sample is generated for each training sample, and the training sample and the expanded training sample are combined into a training sample set.
8. A non-contact heart rate monitoring system, characterized in that: include: A video acquisition module for acquiring facial video images of the driver during driving of the vehicle; A positioning and partitioning module is used to select the ROI features of the face for each static image frame in the facial video image, obtain the ROI area through face detector positioning, and divide the ROI area into multiple ROI blocks; The signal quality analysis and processing module is used to detect the light intensity, occlusion, and motion artifacts in the ROI area of each frame of static image. When the light intensity, occlusion, and motion artifacts meet the set conditions, the static image is extracted as a sample to be tested; and for each sample to be tested, the signal quality analysis and processing module performs global average pooling on all ROI blocks, compresses the pixel values of each ROI block into a time series, and constructs the sample to be tested into a time-space mapping matrix based on the time series; A heart rate monitoring model for real-time monitoring of a driver's heart rate, wherein the heart rate monitoring model completes model training using an expanded training sample and a training sample generated by the sample generation method for non-contact heart rate monitoring as described in any one of claims 1 to 6; after the heart rate monitoring model inputs the space-time mapping matrix of the sample to be tested, the corresponding predicted heart rate value is calculated.
9. The non-contact heart rate monitoring system according to claim 8, characterized in that: It also includes a confidence detection module, which performs a physiological range test and a time consistency test on the predicted heart rate value. When the predicted heart rate value is consistent with the heart rate range of a healthy adult and the average heart rate fluctuation is less than a set threshold, the predicted heart rate value calculated by the heart rate monitoring model is output; otherwise, the signal quality analysis and processing module is fed back to the signal quality analysis and processing module to reprocess the test samples corresponding to the predicted heart rate values that do not meet the output.
10. The non-contact heart rate monitoring system according to claim 8, characterized in that: The video acquisition module includes a visible light camera; Or, the video acquisition module includes a visible light camera and an infrared camera, and the priority of the visible light camera and the infrared camera is determined according to the ratio of the number of samples to be tested extracted by the signal quality analysis and processing module to the number of static images, and the determination results include: using the visible light camera to collect the facial video image of the driver during the driving of the vehicle, using the infrared camera to collect the facial video image of the driver during the driving of the vehicle, and using the visible light camera and the infrared camera to collect the facial video image of the driver during the driving of the vehicle at the same time; When the visible light camera and the infrared camera are used to simultaneously collect facial video images of the driver during driving, the features of the RGB image collected by the visible light camera and the IR image collected by the infrared camera are extracted through a dual-branch convolutional neural network respectively, and the attention mechanism is used to fuse them, and the contribution ratio of the RGB and IR modalities is dynamically adjusted to form a fused image, and the fused image is used as the sample to be tested.
Citation Information
Patent Citations
Large-scale total element training sample set generation method
CN115223014A
Soil image enhancement method based on brightness migration and local information fusion
CN116363003A
Image automatic exposure correction and enhancement method and device
CN117793538A
Non-invasive driver driving fatigue state identification method and system
CN118051810A
Multi-modal fatigue driving detection method based on thermal imaging
CN119741744A