Sample generation method, heart rate monitoring method and system for non-contact heart rate monitoring

By generating random brightness waveform expansion training samples and dual camera acquisition technology, the robustness of heart rate monitoring in complex lighting environments is solved, contactless and low-cost heart rate monitoring is achieved, and monitoring accuracy and safety during driving is improved.

CN119992258BActive Publication Date: 2025-07-04JILIN UNIVERSITY

Patent Information

Application Number
CN202510453401.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-04
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The existing heart rate monitoring methods have weak resistance to ambient light changes in complex lighting environments, making it difficult to meet the monitoring needs during actual driving. In addition, traditional contact methods affect driving comfort, and contactless methods have high equipment complexity and cost.

Method used

By generating random brightness waveform adjustment training samples, expanding the data set, combining dual-camera acquisition technology and signal quality analysis, monitoring and processing light changes in real time, and using the spatiotemporal mapping matrix method to extract heart rate characteristics.

Benefits of technology

It improves the robustness and accuracy of the heart rate monitoring model under complex lighting conditions, ensures stable monitoring during driving, reduces equipment complexity and cost, and improves the reliability of driving safety and health monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992258B_ABST
    Figure CN119992258B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of physiological state monitoring, and specifically provides a method for generating samples for non-contact heart rate monitoring, a heart rate monitoring method and a system. For the original samples, the overall brightness of the existing face image is dynamically adjusted through a randomly generated brightness waveform to simulate the light environment in which the driver's face brightens and darkens during actual driving, and the generated image is used as an augmented data set for training a monitoring model for detecting the driver's heart rate. After the model training is completed, the ROI region of the static image collected by the video acquisition module is located and partitioned through the positioning and partitioning module, and the static image is subjected to signal detection and processing through the signal quality analysis and processing module, and the obtained sample to be measured is input into the heart rate monitoring model to estimate the driver's heart rate value. The present invention effectively expands the training data set through fluctuating brightness, and greatly improves the anti-environmental light change ability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of physiological state monitoring, and particularly relates to a method for generating samples for non-contact heart rate monitoring, a heart rate monitoring method and a system. Background Art

[0002] Heart rate is a reliable indicator for analyzing health conditions and a prospective clinical diagnostic tool. As one of the four vital signs of the human body, the stability of heart rate directly reflects the health of the heart function.

[0003] The development of driver heart rate monitoring technology provides a new way to optimize the personalized driving experience. By analyzing heart rate data, the car can dynamically adjust the configurations such as in-vehicle air conditioner, seat, lights, audio, etc., to relieve driving fatigue and regulate the driver's mood. In addition, by real-time collecting the driver's heart rate data, potential risks such as fatigue driving, excessive tension or sudden health conditions can be effectively identified, so as to issue warnings in time or take active intervention, and can also be linked with the remote medical system to win valuable time for the emergency treatment of sudden health problems. The demand for driver heart rate monitoring reflects the deep integration of modern intelligent transportation technology in the fields of safety and health management. Heart rate monitoring technology can not only improve the efficiency of driving safety and driver health monitoring, but also promote the transformation of intelligent vehicles from traditional transportation tools to more user-friendly and health-oriented travel partners through data sharing and intelligent linkage.

[0004] Electrocardiogram (ECG) is a gold standard for analyzing heart rate data. Currently, electrodes attached to the skin are commonly used clinically, and the electrodes are connected to an ECG recorder through wires to obtain electrocardiogram signals. Another commonly used method is photoplethysmography (PPG), that is, a method of detecting the change in blood volume in blood vessels under human skin tissue by photoelectric means to obtain a pulse wave. This is an optical non-invasive technology. A specific light source emitter connected to positions such as the wrist and finger emits a light beam with a certain wavelength to the skin surface. Since the blood volume in the artery fluctuates periodically with the diastolic and systolic of the heart, the contraction and expansion of the blood vessel during each heartbeat will affect the transmission or reflection of light. Therefore, when the light passes through the skin tissue and then is reflected to the photoelectric receiver, the intensity of the light will have a certain attenuation. Converting the light signal into an electrical signal and then extracting the alternating current signal therein can reflect the characteristics of blood flow, and then obtain the pulse wave signal. Since the PPG signal can reflect many human physiological information related to the cardiovascular system and has the advantages of low cost, non-invasiveness, portability, no cross-infection, etc., it has been widely used. Currently, many commercial wearable devices such as sports bracelets and smart watches on the market utilize this detection method.

[0005] In recent years, due to the impact of the epidemic, people's demand for telemedicine has continued to increase, and some image photoplethysmography (iPPG) based on video to obtain heart rate has emerged. iPPG is a method that can obtain PPG signals from captured videos. It has the advantages of non-contact measurement, easy implementation, and machine operation. In addition, remote PPG (rPPG) technology has also attracted much attention. This method can obtain PPG signals using only an ordinary digital camera without the need for contact probes and dedicated light sources.

[0006] However, the above existing methods all have certain limitations in the application process. Specifically, when using the traditional PPG method, since the contact sensor needs to maintain contact with the skin of the person being tested, it may cause discomfort to the driver and affect the normal driving process. When using the iPPG method to detect heart rate, it is necessary to place a standard white board as a background, and use the color change of the white board in the video to correct the skin color. The method of placing a white board does not conform to the actual driving environment, and cannot be used for measurements at a long distance. It is difficult to meet the monitoring needs in complex driving processes and is only suitable for research in laboratory environments. The existing rPPG method has significant advantages over other methods, but when changes in ambient light affect the signal quality, the system is prone to output erroneous or inaccurate heart rate data, and its ability to resist changes in ambient light is weak. Therefore, there is an urgent need for a method that can improve the adaptability of raw data in brightness change scenarios and improve the confidence of heart rate monitoring results. Summary of the invention

[0007] In view of this, the present invention aims to provide a sample generation method for non-contact heart rate monitoring that can significantly increase the training samples of the heart rate monitoring model, and on this basis, a non-contact heart rate monitoring method and system that can resist changes in ambient light are designed, which effectively solves the problem of insufficient sample number of original data in brightness change scenarios, and improves the robustness and generalization ability of the heart rate monitoring model under complex lighting conditions during actual vehicle driving.

[0008] To achieve the above object, the technical solution created by the present invention is implemented as follows:

[0009] The first aspect of the invention provides a sample generation method for non-contact heart rate monitoring, comprising: for each frame of training samples for heart rate monitoring, setting a time step and a brightness change interval within each time step, randomly sampling a brightness change value within the brightness change interval at each time node, superimposing the brightness change value with the global brightness value of the training sample at the previous time node, and generating a random step waveform of the training sample;

[0010] The training samples are partitioned and located, and the local brightness change of each block is expressed as:

[0011] ;

[0012] Among them, represents the local luminance value of the block at position at time t, represents the luminance amplitude, represents the luminance change frequency, represents the phase shift;

[0013] is the luminance amplitude of each block , the luminance change frequency and the phase shift are randomly assigned;

[0014] According to the random step waveform and the local luminance change, the luminance of each block of each frame of training samples is adjusted, and multiple augmented training samples with different luminances are generated for each frame of training samples.

[0015] Preferably, the random step waveform of the training samples is:

[0016] ;

[0017] Among them, represents the global luminance value of the training sample at time t, represents the global luminance value of the training sample at time t - 1, represents the luminance change amount within each time step.

[0018] Preferably, the luminance change value is randomly sampled from the uniform distribution in the luminance change interval

[0019] Preferably, the block is the ROI block of the training sample.

[0020] Preferably, the random step waveform and the local luminance change are fused to generate a dynamic luminance change as:

[0021] ;

[0022] Among them, represents the dynamic luminance change of the training sample, represents the original luminance of the training sample.

[0023] Preferably, it further includes: adding random noise during the luminance adjustment process of each block, and fusing the random step waveform, the local luminance change and the random noise to generate a dynamic luminance change as:

[0024] ;

[0025] Among them, represents the dynamic luminance change of the training sample, represents the original brightness of the training sample, represents random noise.

[0026] The second aspect of the present invention provides a non-contact heart rate monitoring method, including:

[0027] Collect the facial video images of the driver during vehicle driving;

[0028] For each static image in the facial video image, select the ROI features of the face, locate the ROI area through a face detector, and divide the ROI area into multiple ROI blocks;

[0029] Detect the light intensity, occlusion, and motion artifacts of each ROI block of the static image. When the light intensity, occlusion, and motion artifacts all meet the set conditions, extract this static image as a sample to be measured;

[0030] Use the completed heart rate detection model to monitor the driver's heart rate in real time. The training sample set of the heart rate detection model is: use the sample generation method of non-contact heart rate monitoring to generate the augmented training samples of each training sample, and merge the training samples and the augmented training samples into a training sample set.

[0031] The third aspect of the present invention provides a non-contact heart rate monitoring system, including:

[0032] A video acquisition module for collecting facial video images of the driver during vehicle driving;

[0033] A positioning and partitioning module, which is used for each static image in the facial video image, select the ROI features of the face, locate the ROI area through a face detector, and divide the ROI area into multiple ROI blocks;

[0034] A signal quality analysis and processing module, which is used to detect the light intensity, occlusion, and motion artifacts of each static image ROI area. When the light intensity, occlusion, and motion artifacts all meet the set conditions, extract this static image as a sample to be measured; and, for each sample to be measured, the signal quality analysis and processing module performs global average pooling on all its ROI blocks, compresses the pixel values of each ROI block into a time series, and constructs the sample to be measured into a spatio-temporal mapping matrix according to the time series;

[0035] A heart rate monitoring model for real-time monitoring of the driver's heart rate. The heart rate monitoring model completes model training using the augmented training samples and training samples generated by the sample generation method of non-contact heart rate monitoring; after inputting the spatio-temporal mapping matrix of the sample to be measured, the heart rate monitoring model calculates the corresponding predicted heart rate value.

[0036] Preferably, it further includes a confidence detection module. The confidence detection module conducts physiological range test and time consistency test on the predicted heart rate value. When the predicted heart rate value conforms to the heart rate range of healthy adults and the average heart rate fluctuation is less than the set threshold, the predicted heart rate value calculated by the heart rate monitoring model is output. Otherwise, it feeds back to the signal quality analysis and processing module, and the signal quality analysis and processing module reprocesses the test sample corresponding to the predicted heart rate value that does not meet the output requirement.

[0037] Preferably, the video acquisition module includes a visible light camera;

[0038] Alternatively, the video acquisition module includes a visible light camera and an infrared camera. According to the ratio of the number of test samples extracted by the signal quality analysis and processing module to the number of static images, the priorities of the visible light camera and the infrared camera are judged. The judgment results include: using the visible light camera to collect the facial video image of the driver during vehicle driving, using the infrared camera to collect the facial video image of the driver during vehicle driving, and using the visible light camera and the infrared camera to collect the facial video image of the driver during vehicle driving at the same time;

[0039] When using the visible light camera and the infrared camera to collect the facial video image of the driver during vehicle driving at the same time, the features of the RGB image collected by the visible light camera and the IR image collected by the infrared camera are respectively extracted through a dual-branch convolutional neural network, and an attention mechanism is used for fusion to dynamically adjust the contribution ratio of the RGB and IR modalities to form a fused image, and the fused image is used as the test sample.

[0040] Compared with the prior art, the present invention can achieve the following beneficial effects:

[0041] The present invention innovatively designs a fluctuating brightness method, superimposes brightness fluctuations on the original samples, and quickly generates brightness changes through a mathematical function. Without relying on additional shooting equipment or complex scene settings, the dataset can be easily expanded. It not only effectively solves the problem of insufficient real data in practical applications, but also avoids the occurrence of overfitting phenomena during model training, significantly improves the generalization ability of the model, and makes up for the defect of the lack of complex illumination sample images in the original samples, enhancing the anti-environmental light change ability of the training model. In different driving environments and conditions, the heart rate monitoring model can maintain stable performance and provide continuous and reliable health monitoring services for drivers.

[0042] The present invention also makes an innovative design to the traditional PPG method. The traditional PPG method requires the use of electrode patches or light source probes to contact the human body to obtain heart rate signals, while the present invention can accurately extract heart rate through video signals without any physical contact. This not only avoids the discomfort caused by the driver wearing the device, ensures the naturalness and comfort of the driving process, but is also particularly suitable for remote monitoring of the driver's health status, providing a strong guarantee for driving safety. Compared with the iPPG method, the present invention is more competitive in terms of equipment cost and ease of application. The iPPG method requires specific light sources or dedicated hardware to obtain high-quality signals, which increases the complexity and cost of the equipment. The present invention only requires common smart devices such as ordinary cameras on the dashboard to collect video images of the driver's face, without the need for additional special hardware, to achieve accurate measurement of heart rate. It has the advantages of low cost and low threshold, and is easy to be widely used in various types of vehicles, which promotes the popularization of driving health monitoring technology.

[0043] In terms of signal quality assurance, the present invention adopts a multi-dimensional detection method to evaluate and optimize the signal from multiple angles such as brightness mean and standard deviation, key point texture analysis, Laplace transform, etc., and monitors the influence of light intensity, occlusion and motion artifacts on signal quality in real time, effectively optimizes and processes unqualified signals, and ensures the high accuracy and robustness of the measurement results through comprehensive and fine signal quality control. In addition, the present invention proposes for the first time to use dual cameras for image acquisition. According to the extraction ratio of the samples to be tested fed back by the signal quality analysis and processing module, the visible light camera and the infrared camera can be flexibly switched, and the infrared camera is used to make up for the shortcomings of the visible light camera in low light environments, which greatly improves the image acquisition capability of the video acquisition module in low light environments such as night and tunnels, as well as other complex environments where the visible light camera cannot accurately capture facial images. Moreover, through the fusion of the extracted features of the RGB image and the IR image, the fused image can accurately reflect the physiological characteristics of the driver.

[0044] Compared with the existing iPPG and rPPG methods, the present invention places more emphasis on adaptability to complex lighting changes in real driving environments. It can stably provide reliable monitoring data regardless of strong light, weak light or dynamic lighting conditions, greatly improving the practicality and reliability of in-vehicle applications.

[0045] The present invention innovatively adopts the space-time mapping matrix method to effectively integrate and model the spatial and temporal information in the video signal. This signal processing method can capture richer and more comprehensive heart rate characteristics, deeply explore the subtle changes in heart rate signals in different regions and time points, and further improve the accuracy and credibility of heart rate monitoring results. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The accompanying drawings, which form a part of the present invention, are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not unduly limit the present invention. In the drawings:

[0047] Figure 1 is a structural framework diagram of a non-contact heart rate monitoring system provided according to an embodiment of the present invention;

[0048] Figure 2 is a schematic diagram of the in-vehicle position design of a video acquisition module provided according to an embodiment of the present invention;

[0049] Figure 3 is a schematic diagram of a dual-camera video acquisition method provided according to an embodiment of the present invention. Detailed implementation manners

[0050] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation to the present invention. Similar elements in different embodiments are labeled with related similar element numbers. In the following embodiments, many details are described to enable a better understanding of the present invention. However, those skilled in the art can easily recognize that some of the features can be omitted in different situations, or can be replaced by other elements, materials, or methods. In some cases, some operations related to the present invention are not shown or described in the specification to avoid the core part of the present invention being overwhelmed by excessive description. For those skilled in the art, it is not necessary to describe these related operations in detail, and they can fully understand the related operations based on the description in the specification and the general technical knowledge in the art.

[0051] It should be noted that, without conflict, the embodiments and features in the embodiments of the present invention can be combined with each other to form various embodiments. At the same time, the steps or actions in the method description can also be adjusted in the order that can be obvious to those skilled in the art. Therefore, the various sequences in the specification and drawings are only for clearly describing a certain embodiment and do not mean that they are the necessary sequences, unless it is stated that a certain sequence must be followed.

[0052] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "plural" is two or more.

[0053] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "mounted", "connected", "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific circumstances.

[0054] The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0055] Embodiment 1: In Embodiment 1 of the present invention, a method for generating samples for non-contact heart rate monitoring is provided, which is mainly used to solve the problem that the traditional heart rate monitoring model training sample size is insufficient in intelligent driving applications and cannot cope with complex light changes. By expanding the data volume of the training samples, the anti-environmental light change ability of the trained model is improved. The sample generation method specifically includes the following steps:

[0056] First, during the training process of the non-contact heart rate monitoring model, the publicly available UBFC-RPPG Dataset is usually adopted. This dataset contains synchronously collected facial videos and heart rate data, providing rich learning resources for the model. By using this dataset to train the heart rate monitoring model, the model can learn effective features and patterns for predicting heart rate values from facial videos. However, the data volume of the UBFC-RPPG Dataset is limited, especially the data samples in complex light environments are limited. Therefore, when using it for model training, the model's ability to resist ambient light changes is weak. Thus, in the embodiments of the present invention, the training samples for heart rate monitoring in each frame of the UBFC-RPPG Dataset are mainly augmented to supplement the data samples with complex ambient light changes.

[0057] For each training sample used for heart rate monitoring, first, a random step waveform is generated for the brightness of the entire picture to simulate the random fluctuations of the brightness of the entire picture, reflecting the light changes of the driver's face when the vehicle enters a tunnel, exits the shadow area, etc. Specifically, the expression for establishing the random step waveform of the training sample is:

[0058] ;

[0059] Among them, represents the global brightness value of the training sample at time t, represents the global brightness value of the training sample at time t - 1, represents the amount of brightness change within each time step.

[0060] It should be noted that t - 1 and t represent two consecutive time nodes, and the time step between any two time nodes is set artificially. And according to the set time step, the brightness change interval within each time step is set, and this interval is represented as , represents the step range of brightness change, which is used to control the fluctuation amplitude of brightness.

[0061] During the generation process of the random step waveform, first, the global brightness value needs to be initialized, that is: .

[0062] For each time node t, a brightness change value is randomly sampled from the uniform distribution of the brightness change interval , and this brightness change value is superimposed on the global brightness value of the training sample at the previous time node to obtain the global brightness value , by continuously repeating and iterating the above process, a random stepped waveform with randomly varying brightness values and the variation amplitude limited by the brightness change amplitude can be obtained. Subsequently, by superimposing this random stepped waveform on the original training samples, a set of augmented samples with continuously changing global brightness can be obtained.

[0063] To ensure that the subsequent augmented training samples conform to the actual scenario, it is further necessary to simulate the local illumination changes of the driver's face image to reflect the dynamic characteristics of the local light and shadow on the driver's face. Specifically, in Embodiment 1 of the present invention, each training sample needs to be partitioned and located. The training sample is divided into multiple blocks, and the positions of the blocks are located. The positioning coordinates can be uniformly selected as the coordinates of a certain pixel point within the block. Usually, the coordinates of the central pixel point within the block are selected as the position representation of the entire block. As for the size of the block, it can be determined according to the amount of sample data that needs to be expanded actually. Since during the driver's heart rate monitoring process, the heart rate is mainly predicted based on the driver's facial features, the cheek area is usually selected as the region of interest (ROI) for heart rate monitoring. Therefore, during the process of adjusting the brightness of the local area, the brightness of only the ROI region can also be adjusted and augmented. Here, the ROI region refers to the entire cheek part of the driver. Specifically, the extraction of the ROI region can be achieved through existing image processing methods. To achieve precise adjustment of the illumination on the driver's cheek, it is also necessary to partition and locate the ROI region. Assuming that the ROI region is divided into n rectangular ROI blocks, for each ROI block, the local dynamic waveform of its local brightness change is expressed as:

[0064] ;

[0065] where, represents the local brightness value of the ROI block at position at time t, represents the brightness amplitude, represents the brightness change frequency, represents the phase shift, that is, the initial phase of the brightness curve.

[0066] The block brightness amplitude , brightness change frequency and phase shift in the above formula are variables. Therefore, by randomly assigning values to the brightness amplitude , brightness change frequency and phase shift , the local brightness change of each block can be adjusted. The local brightness change of the ROI blocks of each training sample is expressed as , by assigning different brightness amplitudes , brightness change frequencies and phase shifts , it is possible to expand the randomness and diversity of the brightness change in the ROI region. To conform to the light changes in the actual driving scenario, the local brightness change adjustment of the ROI block can also be set with a certain rule. For example, the brightness adjustment of the ROI region starts from the first ROI block in the upper left and continues until the last ROI block in the lower right, and the brightness fluctuation continuously increases to simulate the local bright and dark changes.

[0067] After setting the global brightness change as a random step waveform and the local brightness change as a local dynamic waveform, the brightness of each ROI block in each frame of the training sample can be adjusted according to the random step waveform and the local brightness change. The brightness of each ROI block in the training sample is adjusted according to the time step to simulate the illumination change of the driver's face caused by tunnels, sunlight, tree shadows, etc. during the actual driving process. After fusing the random step waveform and the local brightness change, the training sample is adjusted to generate an expanded training sample with adjusted brightness. The dynamic brightness change of the training sample is:

[0068] ;

[0069] Among them, represents the dynamic brightness change of the training sample, that is, the brightness of the expanded training sample, represents the original brightness of the training sample.

[0070] To further simulate the slight jitter of light or the random change of ambient light and enhance the diversity and authenticity of the data, it is also necessary to superimpose random noise. In the embodiment of the present invention, the setting method of the random noise is:

[0071] For the brightness noise value at any time t, it is randomly generated from the normal distribution . Among them, the variance is set according to the light interference intensity. After generating the noise matrix, the randomly generated brightness noise value is superimposed on each ROI block according to a certain ratio to simulate the slight jitter of light or the random change of ambient light, thereby further enhancing the diversity and authenticity of the data. Specifically, the random step waveform, the local brightness change, and the random noise are fused to generate a dynamic brightness change of:

[0072] ;

[0073] Among them, represents the dynamic brightness change of the training sample, that is, the brightness of the expanded training sample after superimposing the random noise, represents the original brightness of the training sample, represents the random noise.

[0074] The above method is a general method for each training sample. For each frame of training sample initially obtained, the brightness can be adjusted frame by frame according to the above method, and multiple augmented training samples with dynamic brightness changes can be generated for each frame of training sample. It should be noted that the label of the augmented training sample is consistent with the data label of its corresponding training sample, that is, the heart rate monitoring data of the augmented training sample is equal to the heart rate monitoring data of the original training sample.

[0075] As an optional embodiment, the generation of augmented training samples can be performed only for some training samples.

[0076] Finally, the training samples and the augmented training samples are combined to form a training data set, and the heart rate monitoring model can be trained using this training data set. Due to the effective augmentation of the training data, the trained heart rate monitoring model can significantly improve its ability to resist ambient light changes during heart rate detection.

[0077] Embodiment 2: On the basis of the above method for generating samples for non-contact heart rate monitoring, the present invention further proposes a non-contact heart rate monitoring method, including:

[0078] Using the video acquisition device of the vehicle to collect the facial video images of the driver during the vehicle driving process in real time. The video acquisition device generally can use a high-resolution and high-frame-rate camera, and specifically can adopt one or more of a visible light camera, an infrared camera, a multi-spectral camera, etc. And set appropriate video resolution and frame rate, such as 1080p resolution and 30 frames per second frame rate, to ensure the clarity and coherence of the images.

[0079] For each static image in the facial video image, an open-source face detector with high accuracy and good real-time performance is used to detect the face in the static image and accurately locate the facial area. Select the facial ROI features for driver heart rate detection, and obtain the ROI area through the open-source face detector. In the embodiment of the present invention, according to the different distribution characteristics of the heart rate signal on the face, the driver's cheeks are selected as the ROI features, and the blood flow in these areas is rich and the heart rate signal is relatively obvious.

[0080] After obtaining the ROI area, the located ROI area is further divided into multiple small ROI blocks, such as evenly dividing the cheek area into multiple rectangular ROI blocks, and the size of each ROI block can be adjusted according to actual needs and computing resources, such as dividing the ROI area into n ROI blocks, and each ROI block includes m pixels.

[0081] After partition positioning, the light intensity, occlusion, and motion artifacts of the ROI block of each frame of static image are detected. Specifically, the brightness mean and standard deviation of all ROI blocks in the ROI region are calculated and compared with the preset brightness threshold range to determine whether the light is too strong or too weak. If the brightness mean is within the preset brightness threshold range, it is determined that the light intensity of the frame of static image is qualified and heart rate monitoring can be performed; if the brightness mean exceeds this brightness threshold range, the light is considered unqualified.

[0082] It is determined whether there is occlusion affecting heart rate monitoring in each frame of static image. By analyzing the texture consistency of the area near the key points in the ROI region, it is judged whether there is occlusion. For example, small area images around key points such as the left and right cheekbone regions and the chin edge points are detected. If there is a sudden change in brightness or a lack of expected texture features in the area near the key points, it is determined as occlusion; if the texture features of the small area images around the key points are basically consistent, it is considered unoccluded.

[0083] It is determined whether there are motion artifacts affecting heart rate monitoring in each frame of static image. The Laplace transform is used to detect the sharpness of the ROI block, and it is judged whether there are motion artifacts according to the set sharpness threshold. For example, the variance of the Laplace transform result is calculated. If the variance is lower than the threshold, it is determined that the image is blurred and there are motion artifacts; otherwise, there are no motion artifacts.

[0084] Only when the light intensity, occlusion, and motion artifacts all meet the set conditions, that is, the light intensity meets the set threshold, there is no occlusion, and there are no motion artifacts, will the frame of static image be extracted as a sample to be tested to ensure the reliability of the data quality input into the heart rate detection model.

[0085] After the samples to be tested are extracted, the heart rate of the driver is monitored in real time using the trained heart rate detection model. The training sample set in the training process of this model is generated by expanding through the method in Embodiment 1. The expanded training samples of each training sample are generated using the sample generation method of non-contact heart rate monitoring in Embodiment 1, and the training samples and the expanded training samples are combined to form the training sample set.

[0086] Input the sample to be tested into the trained heart rate detection model, and the predicted heart rate value of the driver can be obtained through the output of the model.

[0087] Embodiment 3: Based on the sample generation method of Embodiment 1 and the heart rate monitoring method of Embodiment 2, please refer to Figure 1 , Embodiment 3 of the present invention also proposes a non-contact heart rate monitoring system, which mainly includes a signal input and processing part and a heart rate detector (heart rate prediction and output). Among them, the signal input and processing part includes a video acquisition module, a positioning and partitioning module, and a signal quality analysis and processing module.

[0088] Please refer toFigure 2 , the main functional device of the video acquisition module is a camera installed on the instrument panel. Usually, a visible light camera is used to obtain the facial video image of the driver.

[0089] The main function of the positioning and partitioning module is to locate the facial ROI features through an open-source face detector. During the heart rate monitoring process, the ROI features are usually selected at the cheek position. Therefore, the cheek ROI area can be located through the open-source face detector. The positioning method of the ROI area is as follows: each frame of static image in the facial video image is detected and divided frame by frame to obtain a stable ROI area positioning. After the ROI area is located, it is adjusted to a rectangle, and this rectangular ROI area is divided to obtain n rectangular ROI blocks.

[0090] The main functions included in the signal quality analysis and processing module include detecting the light intensity, occlusion, and motion artifacts of the driver's face, and converting the test samples with qualified environmental conditions into a spatio-temporal mapping matrix for easy input into the heart rate detection module for processing.

[0091] The light condition has a significant impact on the quality of the video signal. Extreme lighting conditions will affect the accuracy of heart rate monitoring. Therefore, it is necessary to detect whether the light is too dark or too bright. In the embodiment of the present invention, the brightness mean value and standard deviation of each frame of video image are calculated to judge the light intensity level. Specifically, the brightness threshold range is set. If the calculated brightness mean value exceeds this range, it is determined that the light is unqualified. The specific calculation method of the brightness mean value is:

[0092] ;

[0093] Among them, is the brightness mean value, represents the brightness of the block, H represents the number of ROI blocks included in the x direction of each frame of static image, W represents the number of ROI blocks included in the y direction of each frame of static image, is the total number of ROI blocks in each frame of static image.

[0094] Facial occlusion can lead to inaccurate extraction results of the ROI area, which will in turn interfere with the heart rate estimation results. Therefore, occlusion detection is a key step to ensure the quality of the input signal. In the embodiments of the present invention, detection based on the area near key points is adopted to detect key points related to the cheek area such as the left and right zygomatic regions and the chin edge points, and the integrity of the cheek area is inferred using the positions and distributions of these points. The specific method is to locate the key points near the left and right cheeks, extract small rectangular area images around the key points, and analyze the texture consistency of these areas. If there is a sudden change in the brightness of the area near the key point, or the expected texture features are missing, it is determined as cheek occlusion, and it is difficult to obtain an accurate predicted heart rate value through this frame of static image; if the textures of the small rectangular area images extracted around the key points are basically consistent, it is considered that the face in this frame of static image is unoccluded. The specific implementation method of the above facial occlusion judgment can be realized by existing image processing methods or feature comparison models.

[0095] Motion artifacts are phenomena where the image is blurred due to the rapid movement of the driver, and can be evaluated by detecting the image sharpness. In the embodiments of the present invention, the Laplace transform is used to detect the sharpness of the image, and a sharpness threshold is set according to the usage requirements. If the variance of the transformation result is lower than the threshold, the image is determined to be blurred. The specific calculation method is as follows:

[0096] Convert the static image in the original facial video image from RGB to grayscale. The grayscale conversion formula based on the ITU-R BT.601 standard is:

[0097] ;

[0098] where represents the pixel grayscale value converted from RGB to grayscale.

[0099] Apply the Laplace transform to the obtained grayscale image and calculate the second-order gradient of each pixel point :

[0100] ;

[0101] where represents performing the Laplace transform.

[0102] Calculate the variance of the pixel values of the Laplace transform result , and the calculation formula is:

[0103] ;

[0104] where represents the mean value of the Laplace transform result, , represents the number of pixels in the x direction of the grayscale image, represents the number of pixels in the x - direction of the grayscale image, represents the total number of pixels of the grayscale image.

[0105] Judge the variance of pixel values whether it is less than the set clarity threshold. If the variance of pixel values is less than the clarity threshold, it is determined that the image is blurred and there are motion artifacts; otherwise, it is determined that the image is clear and there are no motion artifacts.

[0106] When the above three tests of light intensity, occlusion, and motion artifacts are all qualified, it is considered that the environmental conditions of the static image of this frame are qualified, and the static image of this frame is extracted as the sample to be tested.

[0107] After obtaining the sample to be tested with qualified environmental conditions, the signal quality analysis and processing module performs global average pooling on all ROI blocks in the sample to be tested, compresses the pixel data into a time series, so as to generate a spatio - temporal mapping diagram for each face video sequence, and obtains a spatio - temporal representation matrix, that is, the spatio - temporal mapping matrix. This process converts complex image data into concise time - series data, which is convenient for the subsequent heart rate detection model to perform feature extraction and heart rate prediction.

[0108] For the T samples to be tested, each sample has three color channels of R, G, and B, and each channel is a two - dimensional image. Use to represent the pixel value at the position of one of the R, G, and B channels in the t - th sample to be tested. And perform average pooling on each ROI block obtained by the positioning and partitioning module, and average the pixel values of all ROI blocks to obtain a time series.

[0109] For any one of the R, G, and B channels, the average pixel value of each ROI block at time t is expressed as:

[0110] ;

[0111] where i represents the index of the i - th ROI block, represents the number of pixels of the i - th ROI block. Through average pooling, the pixel dynamic information of each ROI block of the T samples to be tested is simplified into a time series , and the total length of the time series is T.

[0112] Perform min - max normalization on each time series, and scale the time series to the range of [0, 225]. The calculation formula for min - max normalization is:

[0113]

[0114] Among them, represents the time series value after maximum-minimum normalization, represents the maximum pixel value of the i-th ROI block at all time points, represents the minimum pixel value of the i-th ROI block at all time points.

[0115] Through the above normalization, all signal values are mapped to a unified range, which can, to a certain extent, eliminate the signal value amplitude differences between different ROIs or different videos.

[0116] Furthermore, the T-frame sample to be measured is converted into a spatio-temporal mapping of R, G, and B channels. The constructed spatio-temporal mapping is represented as a data matrix of n×T for each channel, where each row corresponds to the time series of an ROI block, and each column represents the values of all ROI blocks at the same time point. The output spatio-temporal map is represented as a three-dimensional data structure of n×T×3 in which the ROI time series of R, G, and B channels are concatenated. Among them, n is the number of ROIs, T is the number of time frames, and 3 represents the three channels of R, G, and B. For each RGB channel, the normalized time series of the ROI are arranged by row to construct a matrix of T. The matrices of the R, G, and B channels are stacked to obtain a spatio-temporal mapping matrix of size n×T×3. This matrix, as the compressed and normalized representation of the original sample to be measured, is input into the heart rate detector, and the heart rate monitoring model in the heart rate detector extracts features in the spatial and temporal dimensions from it, and then predicts the heart rate value.

[0117] The heart rate monitoring model adopted in the embodiment of the present invention is based on the original heart rate regression model and is trained by the training sample set expanded by the sample generation method of non-contact heart rate monitoring in Embodiment 1. For the training of the heart rate monitoring model, the mean square error loss function is adopted in the embodiment of the present invention to optimize the training effect of the model. The loss function of the model is:

[0118] ;

[0119] Among them, T is the number of time frames of the heart rate signal, is the predicted heart rate value measured by the model at time t, is the true heart rate signal of the driver at time t. When L is stable and reaches the set value, it indicates that the model training is completed.

[0120] The sample to be measured is used as the input of the heart rate monitoring model with an n×T×3 spatio-temporal mapping matrix. A 3D convolutional network is used to perform convolutional operations in both the time and space dimensions simultaneously, so as to extract the joint representation of dynamic and static features. The extracted features are downsampled through a pooling layer, which reduces the computational complexity while retaining the key time-space information. The extracted time-space features are input into the Transformer time series modeling module to capture the dynamic dependencies in the time dimension and model the changes between frames. For each frame of features, a mapping is performed to output the corresponding predicted heart rate value, realizing the heart rate estimation for each frame of the sample to be measured.

[0121] After obtaining the predicted heart rate value estimated by the heart rate monitoring model, it is also necessary to check the predicted heart rate value through a confidence detection module to determine whether the predicted heart rate value is credible. In the embodiment of the present invention, the confidence detection module checks the predicted heart rate value through two aspects: physiological range check and time consistency check. Specifically, the physiological range check is used to check whether the predicted heart rate value is within a reasonable physiological range. According to the heart rate ranges of healthy adults at rest and during strenuous exercise, the embodiment of the present invention stipulates that the reasonable physiological range is , when the obtained predicted heart rate value is within this interval, it is considered that the predicted heart rate value conforms to the physiological range and has a high confidence. Otherwise, it is considered that the predicted heart rate value has a low confidence.

[0122] The time consistency check is to calculate the average heart rate within a sliding window to evaluate whether the predicted heart rate value changes smoothly. The calculation formula for the average heart rate within the sliding window is:

[0123] ;

[0124] where N represents the size of the sliding window.

[0125] If the fluctuation of the average heart rate within the sliding window exceeds the set threshold, it is considered that the predicted heart rate value within this window has a low confidence; if the fluctuation of the average heart rate within the sliding window does not exceed the set threshold, it is considered that the predicted heart rate value within this window has a high confidence.

[0126] When both the above physiological range check and time consistency check obtain high confidence, it is considered that the predicted heart rate value is close to the actual heart rate of the driver, and the predicted heart rate value is output through the output module. In addition, the confidence detection module can use the dynamic physiological range based on an individualized model for the check, or introduce more complex statistical test methods to improve the personalized adaptation to the driver, which can enhance the signal robustness, but increases the dependence on historical data and makes the detection and judgment more complex.

[0127] As a preferred embodiment, in extreme conditions such as at night, in tunnels, or in shadow areas where there is insufficient light, it is difficult to obtain a clear facial image of the driver by only using a single visible light camera according to the traditional design, that is, there are problems with the input to be detected, which further leads to a low credibility of the predicted heart rate value. To solve this problem, on the basis of the traditional single visible light camera, the video acquisition module additionally adds an infrared camera and adopts a dual-lens collaborative working mode as shown in Figure 3 .

[0128] When adopting the dual-lens collaborative working mode, it is necessary to judge whether it is in a scene with insufficient light according to the analysis result of the signal quality analysis and processing module on the light, and then determine the priority of the visible light camera and the infrared camera. Specifically, the visible light camera is an RGB camera with high resolution and frame rate, which is mainly used to collect high-quality color images of the driver under normal light conditions; the infrared camera has near-infrared or far-infrared imaging technology and supports low-light conditions, and is used to clearly capture facial contours and feature points in low-light or even lightless environments.

[0129] Therefore, according to the ratio of the number of samples to be measured extracted by the signal quality analysis and processing module to the number of static images, the light condition within a period of time can be reflected, and the priority of the visible light camera and the infrared camera can be judged according to the light condition. In a scene with sufficient light, the visible light camera can obtain clear facial video images, so only the visible light camera is used to collect the facial video images of the driver during vehicle driving; in a scene with insufficient light, the visible light camera cannot obtain clear facial video images, so the visible light camera and the infrared camera are used to collect the facial video images of the driver during vehicle driving at the same time. Dynamically switch between the two modes according to the feedback of the signal quality analysis and processing module, and dynamically switch the data priority of the visible light and infrared cameras.

[0130] When the visible light camera and the infrared camera collect the facial video images of the driver at the same time, the data collected by the two cameras are synchronized to avoid information lag or loss. The dynamic weight distribution algorithm is adopted to ensure the synchronization and no delay of the camera switching. Then, a series of processes such as alignment processing, image enhancement, denoising, and region segmentation are performed on the RGB image collected by the visible light camera and the IR image collected by the infrared camera. The attention mechanism feature-level fusion method is adopted to extract features from the RGB image collected by the visible light camera and the IR image collected by the infrared camera through a two-branch convolutional neural network respectively, and the attention mechanism is used for fusion to dynamically adjust the contribution ratio of the RGB and IR modalities under complex light conditions, and finally form unified sample data.

[0131] As a preferred embodiment, other cameras can also be added. For example, multispectral imaging devices can be used to capture facial videos in multiple wavelength ranges such as visible light and near-infrared light, and PPG signals can be extracted by separating and analyzing signals of different wavelengths. Multispectral data provides richer information, which can enhance the robustness of signal extraction and also enable operation under partial occlusion or complex lighting conditions. However, multispectral cameras are costly, data processing is complex, and they have high requirements for computing resources and technical thresholds.

[0132] As a preferred embodiment, ultrasonic millimeter-wave radar technology is used to detect minute movements such as vibrations on the surface of the human body or movements of the internal chest cavity, and the heart rate can also be deduced. This method can also achieve non-contact measurement and does not rely on optical devices, has strong tolerance to occlusion, and millimeter-wave radar can penetrate clothing to monitor the heart rate of the driver. However, the device is not simple and portable enough, and minute movement signals are easily interfered by breathing or environmental noise, and complex signal separation algorithms need to be adopted.

[0133] In summary, the above are only the preferred embodiments of this specification and are not used to limit the protection scope of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this specification shall be included within the protection scope of this specification.

[0134] The systems, devices, modules, or units illustrated by the above one or more embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0135] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitations, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, commodity or device comprising the said element.

[0136] Each embodiment in this specification is described in a progressive manner. For the identical or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the relevant parts in the method embodiments for the relevant content.

[0137] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A method for generating samples for non-contact heart rate monitoring, characterized in that, Including: For each training sample for heart rate monitoring in each frame, set a time step and the brightness change interval within each time step. At each time node, randomly sample a brightness change value within the brightness change interval, and superimpose this brightness change value on the global brightness value of the training sample at the previous time node to generate a random step waveform of the training sample; Partition and locate the training sample, and represent the local brightness change of each block as: ; Among them, represents the local luminance value of the block at position at time t, represents the luminance amplitude, represents the luminance change frequency, represents the phase shift; Randomly assign values to the brightness amplitude, brightness change frequency, and phase shift for each block; Adjust the brightness of each block of each frame of the training sample according to the random step waveform and the local brightness change, and generate multiple augmented training samples with different brightnesses for each frame of the training sample.

2. The method for generating a sample for non-contact heart rate monitoring according to claim 1, wherein The random step waveform of the training sample is: ; Among them, represents the global brightness value of the training sample at time t, represents the global brightness value of the training sample at time t-1, represents the amount of brightness change within each time step.

3. The method for generating a sample for non-contact heart rate monitoring according to claim 2, wherein Randomly sample a brightness change value from a uniform distribution within the brightness change range .

4. The method for generating a sample for non-contact heart rate monitoring according to claim 1, characterized in that, The block is the ROI block of the training sample.

5. The method for generating a sample for non-contact heart rate monitoring according to claim 1, wherein Fuse the random step waveform and the local brightness change to generate a dynamic brightness change as: ; Among them, represents the dynamic brightness change of the training sample, represents the original brightness of the training sample, represents the global brightness value of the training sample at time t; this dynamic brightness change is the brightness of the augmented training sample.

6. The method for generating a sample for non-contact heart rate monitoring according to claim 1, wherein Also including: During the brightness adjustment process of each block, superimpose random noise, and fuse the random step waveform, the local brightness change, and the random noise to generate a dynamic brightness change as: ; Among them, represents the dynamic brightness change of the training sample, represents the original brightness of the training sample, represents random noise, represents the global brightness value of the training sample at time t; this dynamic brightness change is the brightness of the augmented training sample after superimposing random noise.

7. A non-contact heart rate monitoring method, characterized in that, Including: Collect the facial video images of the driver during vehicle driving; For each static image in the facial video images, select the ROI features of the face, obtain the ROI area through face detector positioning, and divide the ROI area into multiple ROI blocks; Detect the light intensity, occlusion, and motion artifacts of each frame of the static image ROI block. When the light intensity, occlusion, and motion artifacts all meet the set conditions, extract this frame of the static image as the sample to be tested; Use the trained heart rate detection model to monitor the driver's heart rate in real time. The training sample set of the heart rate detection model is: Use the sample generation method for non-contact heart rate monitoring described in any one of claims 1 to 6 to generate the augmented training samples of each training sample, and merge the training samples and the augmented training samples into the training sample set.

8. A non-contact heart rate monitoring system, characterized in that, Including: A video acquisition module for collecting the facial video images of the driver during vehicle driving; A positioning and partitioning module, which is used for each static image in the facial video images, select the ROI features of the face, obtain the ROI area through face detector positioning, and divide the ROI area into multiple ROI blocks; A signal quality analysis and processing module, which is used to detect the light intensity, occlusion, and motion artifacts of each frame of the static image ROI area. When the light intensity, occlusion, and motion artifacts all meet the set conditions, extract this frame of the static image as the sample to be tested; and for each sample to be tested, the signal quality analysis and processing module performs global average pooling on all its ROI blocks, compresses the pixel values of each ROI block into a time series, and constructs the sample to be tested into a spatio-temporal mapping matrix according to the time series; A heart rate monitoring model for real-time monitoring of the driver's heart rate. The heart rate monitoring model completes model training using the augmented training samples and training samples generated by the sample generation method for non-contact heart rate monitoring described in any one of claims 1 to 6; after inputting the spatio-temporal mapping matrix of the sample to be tested, the heart rate monitoring model calculates the corresponding predicted heart rate value.

9. The non-contact heart rate monitoring system according to claim 8, wherein, It further includes a confidence detection module, which conducts physiological range inspection and time consistency inspection on the predicted heart rate value. When the predicted heart rate value conforms to the heart rate range of healthy adults and the average heart rate fluctuation is less than the set threshold, the predicted heart rate value calculated by the heart rate monitoring model is output. Otherwise, it feeds back to the signal quality analysis and processing module, and the signal quality analysis and processing module reprocesses the test sample corresponding to the predicted heart rate value that does not meet the output requirement.

10. The non-contact heart rate monitoring system according to claim 8, characterized in that, The video acquisition module includes a visible light camera; Alternatively, the video acquisition module includes a visible light camera and an infrared camera. According to the ratio of the number of test samples extracted by the signal quality analysis and processing module to the number of static images, the priorities of the visible light camera and the infrared camera are judged. The judgment results include: using the visible light camera to acquire the facial video image of the driver during vehicle driving, using the infrared camera to acquire the facial video image of the driver during vehicle driving, and using the visible light camera and the infrared camera to simultaneously acquire the facial video image of the driver during vehicle driving; When using the visible light camera and the infrared camera to simultaneously acquire the facial video image of the driver during vehicle driving, features are respectively extracted from the RGB image acquired by the visible light camera and the IR image acquired by the infrared camera through a dual-branch convolutional neural network, and an attention mechanism is used for fusion to dynamically adjust the contribution ratio of the RGB and IR modalities to form a fused image, and the fused image is used as the test sample.

Citation Information

Patent Citations

  • Large-scale total element training sample set generation method

    CN115223014A

  • Soil image enhancement method based on brightness migration and local information fusion

    CN116363003A

Cited By

  • Driver heart monitoring system based on non-contact electrocardio and visual pulse waves

    CN121621995A