An rppg anti-shake method and device and a storage medium
By extracting facial and respiratory region offsets and global optical flow features from video images and combining them with vehicle motion features, artifacts are identified and eliminated, improving the accuracy and stability of in-vehicle breathing rate detection and solving the detection accuracy problem caused by camera shake.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHERY COMMERCIAL VEHICLE (ANHUI) CO LTD
- Filing Date
- 2026-07-02
- Publication Date
- 2026-07-31
AI Technical Summary
During vehicle operation, factors such as camera shake and vehicle vibration can cause image synchronization displacement confusion, affecting the accuracy of non-contact breath detection.
By extracting image motion features such as overall facial offset, respiratory region offset, and global optical flow from video images, and combining them with vehicle motion features, we can identify and remove respiratory signal artifacts caused by vehicle motion and personnel movements, retaining only the true respiratory signals.
It improves the accuracy, stability and reliability of respiratory rate detection in complex dynamic environments inside the vehicle, and reduces the interference of vehicle movement and personnel movements on respiratory signals.
Smart Images

Figure CN122493436A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of vital sign monitoring technology, and more specifically, this invention relates to an rPPG anti-shake method, device and storage medium. Background Technology
[0002] In the fields of intelligent cockpit health monitoring, driver status perception, and in-vehicle vital sign monitoring, non-contact respiratory rate detection is gradually becoming an important component of in-vehicle perception systems due to its advantages such as not requiring the wearing of sensors, not requiring active cooperation, and continuous online monitoring.
[0003] rPPG enables non-contact monitoring of vital signs such as heart rate and respiration using ordinary cameras. This technology acquires multi-channel respiratory signals through a single camera and extracts respiratory signals directly from the video ROI. It mainly relies on local brightness changes, edge changes, or optical flow changes in the nostrils, lips, or neck area for respiratory estimation. However, during vehicle operation, factors such as camera shake, vehicle vibration, door closing impact, and road bumps can cause synchronous displacement of the entire image or a large area. These movements are easily confused with respiratory micro-movements in the temporal and spatial domains, resulting in obvious artifacts in the respiratory signal and seriously affecting the accuracy of respiratory detection. Summary of the Invention
[0004] In view of this, this application provides an rPPG anti-jitter method, which aims to improve at least one of the above-mentioned problems.
[0005] Specifically, the following technical solutions are included:
[0006] On the one hand, embodiments of this application provide an rPPG anti-jitter method, the method being as follows:
[0007] (1) Extract current image motion features from video images containing the target person;
[0008] (2) Extract the current vehicle motion features from the vehicle's motion signals;
[0009] (3) Select a breathing signal that reflects real breathing by combining the current vehicle motion characteristics and the current image motion characteristics of the target person;
[0010] (4) Extract the current respiratory rate of the target person from the selected respiratory signal.
[0011] In some embodiments of the present invention, image motion features include the current overall facial offset, the breathing region offset, and the global optical flow of the current image.
[0012] In some embodiments of the present invention, the current vehicle motion features and image motion features determine the jitter source that causes the change in the breathing signal. The jitter source is the movement of the vehicle, the action of the target person, or the breathing change of the target person. If the jitter source is the breathing change of the target person, the current breathing signal is taken as a valid breathing signal; otherwise, the current breathing signal is considered an invalid breathing signal.
[0013] In some embodiments of the present invention, the process for determining the jitter source is as follows:
[0014] When there is a significant change in the vehicle's motion characteristics, and the global optical flow, breathing region offset, and overall facial offset of the current image frame increase synchronously relative to the previous image frame, the jitter source is identified as the vehicle's motion.
[0015] When the vehicle's motion characteristics are relatively stable, and only the offset of the breathing area and the overall facial offset of the current image frame increase synchronously relative to the previous image frame, the source of the shaking is identified as the movement of the target person.
[0016] When the vehicle's motion characteristics are relatively stable, and the global optical flow and overall facial offset of the current image frame relative to the previous image frame are close to zero, and the low-frequency signal of the breathing area shows periodic, slight changes in the video image, then the source of the jitter is identified as the breathing changes of the target person.
[0017] In some embodiments of the present invention, the breathing area includes at least one of the nostril area, lip area, and throat area.
[0018] In some embodiments of the present invention, the raw respiratory signals extracted from the nostril region, lip region, and throat region are fused to form a first respiratory signal. The first respiratory signal is used for extracting the respiratory rate, wherein the first respiratory signal is represented as follows:
[0019] ;
[0020] in, This represents the first respiratory signal at the current time t. This represents the raw respiratory signal extracted from the nostril region at the current time t. This represents the raw respiratory signal extracted from the lip region at the current time t. This represents the raw respiratory signal extracted from the throat region at the current time t. , as well as These are the weighting coefficients.
[0021] In some embodiments of the present invention, the current image motion features and vehicle motion features are used to calculate the stability score of the first respiratory signal, and the respiratory rate is extracted from the first respiratory signal with a high stability score.
[0022] In some embodiments of the present invention, vehicle motion characteristics include: longitudinal dynamic acceleration and angular velocity of the vehicle.
[0023] In some embodiments of the present invention, the stability score of the first respiratory signal is represented as follows:
[0024] ;
[0025] ;
[0026] in, The jitter score represents the first respiratory signal at time t; The stability score of the first respiratory signal at time t; , , These represent the overall facial offset at time t. Breathing area offset Global optical flow The overall facial offset was calculated separately. Breathing area offset Global optical flow Normalization is performed to obtain a uniform overall facial offset. Homogenized respiratory region offset Uniform global optical flow ; , Let these represent the longitudinal dynamic acceleration and angular velocity of the vehicle at time t, respectively. Regarding the longitudinal dynamic acceleration... angular velocity Normalization is performed to obtain a uniform longitudinal dynamic acceleration. Homogenized angular velocity ; , , , as well as These are the weighting coefficients.
[0027] On the other hand, embodiments of this application provide an rPPG anti-jitter device, the device comprising:
[0028] The image feature extraction unit is used to extract current image motion features from video images containing target personnel;
[0029] The motion feature extraction unit is used to extract the current vehicle motion features from the vehicle's motion signal;
[0030] The breathing signal filtering unit is used to select breathing signals that reflect real breathing based on the current vehicle motion characteristics and image motion characteristics.
[0031] The respiratory detection unit is used to extract the current respiratory rate of the target person from the selected respiratory signal.
[0032] In some embodiments of the present invention, image motion features include the current overall facial offset, the breathing region offset, and the global optical flow of the current image.
[0033] In some embodiments of the present invention, the motion feature extraction unit extracts vehicle motion features including longitudinal dynamic acceleration and angular velocity from the vehicle's three-axis acceleration and three-axis angular velocity.
[0034] In some embodiments of the present invention, the breathing signal filtering unit determines the jitter source that causes the change in breathing signal by combining the current vehicle motion characteristics and image motion characteristics. The jitter source is the movement of the vehicle, the action of the target person, or the breathing change of the target person. If the jitter source is the breathing change of the target person, the current breathing signal is regarded as a valid breathing signal; otherwise, the current breathing signal is considered as an invalid breathing signal.
[0035] In some embodiments of the present invention, when there is a significant change in vehicle motion characteristics, and the global optical flow, breathing region offset, and overall facial offset of the current image frame increase synchronously relative to the previous image frame, the breathing signal filtering unit identifies the jitter source as vehicle motion; when the vehicle motion characteristics are relatively stable, and only the breathing region offset and overall facial offset of the current image frame increase synchronously relative to the previous image frame, the breathing signal filtering unit identifies the jitter source as the target person's movement; when the vehicle motion characteristics are relatively stable, and the global optical flow and overall facial offset of the current image frame are close to zero relative to the previous image frame, and the low-frequency signal of the breathing region changes periodically and slightly in the video image, the breathing signal filtering unit identifies the jitter source as the target person's breathing changes.
[0036] On the other hand, embodiments of this application provide a storage medium storing a computer program, which is executed by a processor to implement the rPPG anti-jitter method as described above.
[0037] This invention extracts image motion features such as overall facial offset, respiratory region offset, and global optical flow of the current image from video images. It also combines the motion features output by the vehicle's IMU to identify respiratory signals that reflect real breathing. The respiratory rate is extracted from the respiratory signals that reflect real breathing, which greatly reduces the interference of vehicle movement and target personnel's actions on the respiratory signals, thereby improving the accuracy, stability, and reliability of respiratory rate detection in complex dynamic environments inside the vehicle. Attached Figure Description
[0038] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 A flowchart of the rPPG anti-jitter method provided in an embodiment of the present invention;
[0040] Figure 2 This is a schematic diagram of the structure of the rPPG anti-jitter device provided in an embodiment of the present invention;
[0041] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0042] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. Unless otherwise defined, all technical terms used in the embodiments of this application have the same meaning as commonly understood by those skilled in the art.
[0043] This invention extracts image motion features such as overall facial offset, respiratory region offset, and global optical flow of the current image from video images. It also combines the motion features output by the vehicle's IMU to identify the breathing signal that reflects real breathing. The breathing signal is then scored for stability, and the breathing rate is extracted from the breathing signals with high stability scores. This greatly reduces the interference of vehicle movement and target personnel's actions on the breathing signal, thereby improving the accuracy, stability, and reliability of breathing rate detection in complex dynamic environments inside the vehicle. Figure 1 A flowchart of the rPPG anti-jitter method provided in this embodiment of the invention is shown below:
[0044] (1) Extract current image motion features from video images containing the target person;
[0045] In this embodiment of the invention, the image motion features include the current overall facial offset, the breathing region offset, and the global optical flow of the current image. The method for extracting the above image motion features is as follows:
[0046] (11) Install a camera inside the vehicle to capture video images of the target person in the target area, extract the facial region of the target person from the current image frame, and detect the overall facial offset of the current image frame relative to the previous image frame. Select the areas involved in the overall facial offset from relatively stable regions such as the corners of the eyes, bridge of the nose, tip of the nose, chin, and facial contours. The key point of detection is that the aforementioned cameras can be in-vehicle DMS cameras, OMS cameras, cockpit cameras, or ordinary RGB cameras;
[0047] (12) Extract the breathing region of the target person from the current image frame and detect the breathing region offset of the current image frame relative to the previous image frame. In this invention, the nostril region, lip region, and throat region are taken as breathing regions. The offset of the nostril region, lip region, and throat region between adjacent image frames is calculated respectively. That is, the average of the above offsets is taken as the breathing region offset. Alternatively, the nostril region, lip region, and throat region are taken as a breathing region and the offset of the breathing region between adjacent image frames is detected. .
[0048] (13) Calculate the global optical flow of the background region and the face region between the current image frame and the previous image frame, respectively. The image consists of a facial region and a background region outside the facial region.
[0049] (2) Extract the current vehicle motion features from the vehicle's motion signals;
[0050] In this embodiment of the invention, based on the inertial measurement unit (IMU) installed on the vehicle, if the inertial measurement unit (IMU) is integrated on the vehicle's camera, then there is no need to separately deploy an inertial measurement unit (IMU) on the vehicle. The inertial measurement unit (IMU) is used to collect the vehicle's three-axis acceleration and three-axis angular velocity. Since the acquisition frequency of the inertial measurement unit (IMU) and the camera is different, the time of the video frame is used as the reference time. The IMU data with the smallest time difference from the reference time is found, and the IMU data with the smallest time difference is used as the IMU data for the reference time of the corresponding video frame. The vehicle's three-axis acceleration and three-axis angular velocity are used as the vehicle's motion signals. Motion features including the vehicle's longitudinal dynamic acceleration and angular velocity are extracted from the vehicle's three-axis acceleration and three-axis angular velocity. The motion features are used to reflect whether the vehicle has physical vibration. The specific formula for calculating the vehicle's longitudinal dynamic acceleration is as follows:
[0051] ;
[0052] in, This represents the longitudinal (z-axis) dynamic acceleration at the current time t. This represents the z-axis acceleration collected by the inertial measurement unit (IMU) at time t. It represents the acceleration due to gravity.
[0053] The formula for calculating the angular velocity of a vehicle is as follows:
[0054] ;
[0055] in, , , These represent the angular velocities along the x, y, and z axes acquired by the inertial measurement unit (IMU) at time t, respectively.
[0056] When a vehicle experiences idling vibration, road bumps, sudden braking, turning, door closing impact, or camera structure vibration, the vehicle's motion characteristics will change accordingly. These physical motion characteristics can be used to help determine whether the motion in the video is caused by the shaking of an external vehicle.
[0057] (3) Select a breathing signal that reflects real breathing by combining the current vehicle motion characteristics and the current image motion characteristics of the target person;
[0058] In this embodiment of the invention, since the movement of the vehicle, the action of the target person, and the changes in the breathing of the target person may all cause changes in the breathing signal, the present invention aims to monitor the changes in the breathing signal caused by the actual breathing of the target person, identify and eliminate false changes in the breathing signal caused by the movement of the vehicle and the movement of the target person. Based on this, the present invention combines the current vehicle movement characteristics and the current image movement characteristics of the target person to determine the jitter source that causes the change in the breathing signal. The jitter source is the movement of the vehicle, the action of the target person, or the change in the breathing of the target person. If the jitter source is the change in the breathing of the target person, the current breathing signal is taken as a valid breathing signal; otherwise, the current breathing signal is considered an invalid breathing signal.
[0059] In this embodiment of the invention, the method for determining the source of jitter is as follows:
[0060] When there are significant changes in the vehicle's longitudinal dynamic acceleration and angular velocity, and the global optical flow, breathing region offset, and overall facial offset of the target person increase synchronously in the current image frame relative to the previous image frame, the change in the current breathing signal is determined to be caused by vehicle vibration, and the current breathing signal is considered an invalid breathing signal.
[0061] When the vehicle's longitudinal dynamic acceleration and angular velocity are relatively stable, and only the offset of the breathing area and the overall offset of the target person's face in the current image frame relative to the previous image frame increase synchronously, it is determined that the change in the current breathing signal is caused by the target person's actions, such as turning the head, lowering the head, speaking, adjusting the sitting posture, etc., and the current breathing signal is determined to be an invalid breathing signal.
[0062] When the vehicle's longitudinal dynamic acceleration and angular velocity are relatively stable, and the global optical flow of the current image frame relative to the previous image frame and the overall facial offset of the target person are close to zero, while the low-frequency signal of the breathing area shows periodic slight changes in the video image, and the frequency band of the breathing signal is located in the low-frequency band of 0.5Hz-2.2Hz, then the change in the current breathing signal is considered to truly reflect the breathing changes of the target person, and the corresponding breathing signal is recognized as a valid breathing signal.
[0063] When the change in the vehicle's current longitudinal dynamic acceleration compared to the previous moment is small, and the change in the vehicle's current angular velocity compared to the previous moment is also small, the vehicle's longitudinal dynamic acceleration and angular velocity are considered to be relatively stable at present.
[0064] (4) Extract the current respiratory rate of the target person from the selected respiratory signal.
[0065] In this embodiment of the invention, multiple raw breathing signals are extracted from multiple breathing regions of the current image frame, and these signals are fused to form a first breathing signal. The breathing rate is extracted from the first breathing signal. The breathing regions include at least one of the nostril region, lip region, and throat region. The nostril region is used to extract nasal airflow or nasal wing micro-movement information; the lip region is used to extract mouth breathing or lip opening and closing information; and the throat region is used to extract the minute undulations of the neck caused by breathing. The corresponding first breathing signal is represented as follows:
[0066] ;
[0067] in, This represents the first respiratory signal at the current time t. This represents the raw respiratory signal extracted from the nostril region at the current time t. This represents the raw respiratory signal extracted from the lip region at the current time t. This represents the raw respiratory signal extracted from the throat region at the current time t. , as well as These are weighting coefficients, which can be dynamic or static. Weights are assigned based on the current clarity, stability, and interference level of each region. For example, when the nostril region is clear, the weight of the original respiratory signal in that region is increased; when the occupant breathes through their mouth, the weight of the original respiratory signal in the lip region is increased; when the nostril or lip regions are obstructed (partially or completely), the weight of the original respiratory signal in the throat region is increased. By jointly analyzing the original respiratory signals from multiple respiratory regions, even if the original respiratory signal from one region is unavailable, respiratory monitoring can continue using signals from other respiratory regions.
[0068] To improve the accuracy of respiratory rate and intensity detection, the quality of the first respiratory signal is scored, and its stability is assessed using a stability score. A higher stability score indicates better stability and higher quality of the first respiratory signal. Respiratory rate and intensity are extracted from the first respiratory signal with a high stability score. The specific formula for calculating the stability score of the first respiratory signal is as follows:
[0069] ;
[0070] ;
[0071] in, The jitter score represents the first respiratory signal at time t; The stability score of the first respiratory signal at time t; , , These represent the overall facial offset at time t. Breathing area offset Global optical flow The overall facial offset was calculated separately. Breathing area offset Global optical flow Normalization is performed to obtain a uniform overall facial offset. Homogenized respiratory region offset Uniform global optical flow ; , Let these represent the longitudinal dynamic acceleration and angular velocity of the vehicle at time t, respectively. Regarding the longitudinal dynamic acceleration... angular velocity Normalization is performed to obtain a uniform longitudinal dynamic acceleration. Homogenized angular velocity ; , , , as well as These are the weighting coefficients.
[0072] After forming a stability score for the current respiratory signal, the quality of the first respiratory signal is rated based on the stability score, and divided into multiple quality levels. This invention uses four quality levels as an example, namely the first quality level to the fourth quality level, where the fourth quality level is the highest quality level and corresponds to the highest stability score; in the stability score... At this time, the current quality level is fourth, and the respiratory rate and respiratory intensity are extracted from the corresponding first respiratory signal; At this time, the current quality level is third, and the respiratory rate and respiratory intensity are extracted from the corresponding first respiratory signal, but the confidence level of the current first respiratory signal is reduced; At this time, the current quality level is second, and motion compensation, filtering suppression, or significant reduction of the confidence level of the first respiratory signal are applied to the respiratory signal; If the current quality level is first, the respiratory signal is discarded, and respiratory rate and intensity are not extracted.
[0073] Figure 2 This is a schematic diagram of the rPPG anti-jitter device provided in an embodiment of the present invention. For ease of explanation, only the parts related to the embodiment of the present invention are shown. The device includes:
[0074] The system includes an image feature extraction unit, a motion feature extraction unit, a respiratory signal screening unit, and a respiratory monitoring unit. The image feature extraction unit extracts current image motion features from video images containing the target person and sends them to the respiratory signal screening unit. The motion feature extraction unit extracts current vehicle motion features from vehicle motion signals and sends them to the respiratory signal screening unit. The respiratory signal screening unit selects respiratory signals reflecting actual breathing based on the current vehicle motion features and the target person's current image motion features. The respiratory monitoring unit extracts the target person's current respiratory rate from the selected respiratory signals.
[0075] In this embodiment of the invention, the image motion features include the current overall facial offset, the respiratory region offset, and the global optical flow of the current image. Therefore, the image feature extraction unit includes a facial offset monitoring module, a respiratory region offset detection module, and a global optical flow monitoring module. A camera is installed inside the vehicle to capture video images of the target person within the target area. The camera can be an in-vehicle DMS camera, an OMS camera, a cabin camera, or a regular RGB camera. The captured video of the target person is sent to the facial offset monitoring module, the respiratory region offset detection module, and the global optical flow monitoring module. The facial offset monitoring module extracts the facial region of the target person from the current image frame and detects the overall facial offset of the current image frame relative to the previous image frame. Select the areas involved in the overall facial offset from relatively stable regions such as the corners of the eyes, bridge of the nose, tip of the nose, chin, and facial contours. Key points of detection: The breathing region offset detection module extracts the breathing region of the target person from the current image frame and detects the breathing region offset of the current image frame relative to the previous image frame. This invention uses the nostril region, lip region, and throat region as the breathing region, and calculates the offset of the nostril region, lip region, and throat region between adjacent image frames, using the average of these offsets as the breathing region offset. Alternatively, it can use the nostril region, lip region, and throat region as a single breathing region and detect the offset of this breathing region between adjacent image frames. The global optical flow monitoring module calculates the global optical flow of the background region and the facial region between the current image frame and the previous image frame, respectively. The image consists of the facial area and the background area outside the facial area;
[0076] In this embodiment of the invention, based on the inertial measurement unit (IMU) installed on the vehicle, if the inertial measurement unit (IMU) is integrated on the vehicle's camera, there is no need to install an additional inertial measurement unit (IMU) on the vehicle. The inertial measurement unit (IMU) is used to collect the vehicle's three-axis acceleration and three-axis angular velocity. Since the acquisition frequency of the inertial measurement unit (IMU) is different from that of the camera, the IMU data with the smallest time difference from the reference time is found using the time of the video frame as the reference time. The IMU data with the smallest time difference is compared with the IMU data of the corresponding video frame at the reference time. The vehicle's three-axis acceleration and three-axis angular velocity are then sent as motion signals to the motion feature extraction unit. The motion feature extraction unit extracts motion features, including the vehicle's longitudinal dynamic acceleration and angular velocity, from the vehicle's three-axis acceleration and three-axis angular velocity. The motion features are used to reflect whether the vehicle has physical vibration. The specific formula for calculating the vehicle's longitudinal dynamic acceleration is as follows:
[0077] ;
[0078] in, This represents the longitudinal (z-axis) dynamic acceleration at the current time t. This represents the z-axis acceleration collected by the inertial measurement unit (IMU) at time t. It represents the acceleration due to gravity.
[0079] The formula for calculating the angular velocity of a vehicle is as follows:
[0080] ;
[0081] in, , , These represent the angular velocities along the x, y, and z axes acquired by the inertial measurement unit (IMU) at time t, respectively.
[0082] When a vehicle experiences idling vibration, road bumps, sudden braking, turning, door closing impact, or camera structure vibration, the vehicle's motion characteristics will change accordingly. These physical motion characteristics can be used to help determine whether the motion in the video is caused by the shaking of an external vehicle.
[0083] In this embodiment of the invention, since the movement of the vehicle, the action of the target person, and the breathing changes of the target person may all cause changes in the breathing signal, the breathing signal filtering unit aims to monitor the changes in the breathing signal caused by the actual breathing of the target person, identify and eliminate false changes in the breathing signal caused by the movement of the vehicle and the movement of the target person. Based on this, the breathing signal filtering unit combines the current vehicle movement characteristics and the current image movement characteristics of the target person to determine the jitter source that causes the change in the breathing signal. The jitter source is the movement of the vehicle, the action of the target person, or the breathing changes of the target person. If the jitter source is the breathing changes of the target person, the current breathing signal is regarded as a valid breathing signal; otherwise, the current breathing signal is considered as an invalid breathing signal.
[0084] In this embodiment of the invention, when there are significant changes in the vehicle's longitudinal dynamic acceleration and angular velocity, and the global optical flow, breathing region offset, and overall facial offset of the target person increase synchronously in the current image frame relative to the previous image frame, the change in the current breathing signal is determined to be caused by vehicle vibration, and the breathing signal filtering unit considers the current breathing signal as an invalid breathing signal; when the vehicle's longitudinal dynamic acceleration and angular velocity are relatively stable, and only the breathing region offset and overall facial offset of the target person increase synchronously in the current image frame relative to the previous image frame, the change in the current breathing signal is determined to be caused by the target person's movement, such as... Actions such as turning the head, lowering the head, speaking, and adjusting the sitting posture of the target person are considered invalid breathing signals by the breathing signal screening unit. When the longitudinal dynamic acceleration and angular velocity of the vehicle are relatively stable, and the global optical flow of the current image frame relative to the previous image frame and the overall facial offset of the target person are close to zero, while the low-frequency signal of the breathing area changes periodically and slightly in the video image, and the frequency band of the breathing signal is in the low-frequency band of 0.5Hz-2.2Hz, then the change in the current breathing signal is considered to truly reflect the breathing changes of the target person, and the breathing signal screening unit identifies the corresponding breathing signal as a valid breathing signal.
[0085] When the change in the vehicle's current longitudinal dynamic acceleration compared to the previous moment is small, and the change in the vehicle's current angular velocity compared to the previous moment is also small, the vehicle's longitudinal dynamic acceleration and angular velocity are considered to be relatively stable at present.
[0086] In this embodiment of the invention, the respiratory monitoring unit includes a respiratory signal extraction module, a fusion module, and a monitoring module. The respiratory signal extraction module extracts multiple raw respiratory signals from multiple respiratory regions in the current image frame. The fusion module fuses the extracted raw respiratory signals to form a first respiratory signal. The monitoring module extracts the respiratory rate from the first respiratory signal. The respiratory regions include at least one of a nostril region, a lip region, and a throat region. The nostril region is used to extract nasal airflow or nasal wing micro-movement information; the lip region is used to extract mouth breathing or lip opening and closing information; and the throat region is used to extract the minute undulations of the neck caused by breathing. The first respiratory signal in the fusion module is represented as follows:
[0087] ;
[0088] in, This represents the first respiratory signal at the current time t. This represents the raw respiratory signal extracted from the nostril region at the current time t. This represents the raw respiratory signal extracted from the lip region at the current time t. This represents the raw respiratory signal extracted from the throat region at the current time t. , as well as These are weighting coefficients, which can be dynamic or static. Weights are assigned based on the current clarity, stability, and interference level of each region. For example, when the nostril region is clear, the weight of the original respiratory signal in that region is increased; when the occupant breathes through their mouth, the weight of the original respiratory signal in the lip region is increased; when the nostril or lip regions are obstructed (partially or completely), the weight of the original respiratory signal in the throat region is increased. By jointly analyzing the original respiratory signals from multiple respiratory regions, even if the original respiratory signal from one region is unavailable, respiratory monitoring can continue using signals from other respiratory regions.
[0089] To improve the accuracy of respiratory rate and intensity detection, the quality of the first respiratory signal is scored, and its stability is assessed using a stability score. A higher stability score indicates better stability and higher quality of the first respiratory signal. The respiratory monitoring unit also includes a scoring module for calculating the stability score of the first respiratory signal and extracting the respiratory rate and intensity from first respiratory signals with high stability scores. The stability score of the first respiratory signal is expressed as follows:
[0090] ;
[0091] ;
[0092] in, The jitter score represents the first respiratory signal at time t; The stability score of the first respiratory signal at time t; , , These represent the overall facial offset at time t. Breathing area offset Global optical flow The overall facial offset was calculated separately. Breathing area offset Global optical flow Normalization is performed to obtain a uniform overall facial offset. Homogenized respiratory region offset Uniform global optical flow ; , Let these represent the longitudinal dynamic acceleration and angular velocity of the vehicle at time t, respectively. Regarding the longitudinal dynamic acceleration... angular velocity Normalization is performed to obtain a uniform longitudinal dynamic acceleration. Homogenized angular velocity ; , , , as well as These are the weighting coefficients.
[0093] The respiratory monitoring unit also includes a processing module, which rates and grades the quality of the first respiratory signal based on a stability score, dividing the respiratory signal quality into multiple quality levels. This invention uses four quality levels as an example, from the first to the fourth quality level, where the fourth quality level is the highest and corresponds to the highest stability score. At this time, the current quality level is fourth, and the respiratory rate and respiratory intensity are extracted from the corresponding first respiratory signal; At this time, the current quality level is third, and the respiratory rate and respiratory intensity are extracted from the corresponding first respiratory signal, but the confidence level of the current first respiratory signal is reduced; At this time, the current quality level is second, and motion compensation, filtering suppression, or significant reduction of the confidence level of the first respiratory signal are applied to the respiratory signal; If the current quality level is first, the respiratory signal is discarded, and respiratory rate and intensity are not extracted.
[0094] It should be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0095] This device extracts image motion features such as overall facial offset, respiratory region offset, and global optical flow from video images. Combined with motion features output from the vehicle's IMU, it identifies respiratory signals reflecting actual breathing. The device performs a stability score on these signals and extracts the respiratory rate from those with high stability scores. This significantly reduces interference from vehicle movement and the actions of the target person, thereby improving the accuracy, stability, and reliability of respiratory rate detection in complex dynamic environments within the vehicle. Furthermore, it can be implemented using existing in-vehicle cameras and onboard IMUs, eliminating the need for occupants to wear contact sensors such as chest straps, nose clips, or finger clips, and also eliminating the need for additional multi-camera arrays or complex medical equipment. It is well-suited for deployment in vehicle cockpit domain controllers, DMS / OMS controllers, smart rearview mirrors, vehicle infotainment systems, in-vehicle camera modules, or mobile terminals, offering significant engineering feasibility and cost advantages.
[0096] One embodiment of this application provides a terminal device including a processor and a memory. The processor may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). The processor may also include a main processor and a coprocessor. The main processor is used to process data in the wake-up state, also known as a central processing unit (CPU); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, the processor may also include an AI processor, which is used to handle computational operations related to machine learning. The memory may include one or more computer-readable storage media, which may be non-transitory. The memory may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, a non-transitory computer-readable storage medium in the memory is used to store a computer program configured to be executed by one or more processors to implement the rPPG anti-jitter method described above.
[0097] In some embodiments, the terminal device may also optionally include: a peripheral device interface and at least one peripheral device. The processor, memory, and peripheral device interface can be connected via a bus or signal lines. Each peripheral device can be connected to the peripheral device interface via a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of: radio frequency circuitry, a display screen, audio circuitry, and a power supply. Those skilled in the art will understand that the above structure does not constitute a limitation on the terminal device, and may include more or fewer components than illustrated, or combine certain components, or employ different component arrangements.
[0098] In an exemplary embodiment, a computer-readable storage medium is also provided, wherein a computer program is stored in the storage medium, and the computer program, when executed by a processor, implements the above-described rPPG anti-jitter method. Optionally, the computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), solid-state drives (SSDs), or optical discs, etc. The random access memory may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM).
[0099] In an exemplary embodiment, a computer program product is also provided, the computer program product including a computer program stored in a computer-readable storage medium. A processor of a terminal device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, causing the terminal device to perform the above-described rPPG anti-jitter method.
[0100] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the application disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only.
[0101] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. An rPPG anti-jitter method, characterized in that, The method is as follows: (1) Extract current image motion features from video images containing the target person; (2) Extract the current vehicle motion features from the vehicle's motion signals; (3) Select a breathing signal that reflects real breathing by combining the current vehicle motion characteristics and the current image motion characteristics of the target person; (4) Extract the current respiratory rate of the target person from the selected respiratory signal.
2. The rPPG anti-jitter method as described in claim 1, characterized in that, Image motion features include the current overall facial offset, the respiratory region offset, and the global optical flow of the current image.
3. The rPPG anti-jitter method as described in claim 1, characterized in that, The current vehicle motion characteristics and image motion characteristics determine the jitter source that causes the change in the breathing signal. The jitter source is the movement of the vehicle, the action of the target person, or the breathing change of the target person. If the jitter source is the breathing change of the target person, the current breathing signal is considered a valid breathing signal; otherwise, the current breathing signal is considered an invalid breathing signal.
4. The rPPG anti-jitter method as described in claim 3, characterized in that, The process for determining the jitter source is as follows: When there is a significant change in the vehicle's motion characteristics, and the global optical flow, breathing region offset, and overall facial offset of the current image frame increase synchronously relative to the previous image frame, the jitter source is identified as the vehicle's motion. When the vehicle's motion characteristics are relatively stable, and only the offset of the breathing area and the overall facial offset of the current image frame increase synchronously relative to the previous image frame, the source of the shaking is identified as the movement of the target person. When the vehicle's motion characteristics are relatively stable, and the global optical flow and overall facial offset of the current image frame relative to the previous image frame are close to zero, and the low-frequency signal of the breathing area shows periodic, slight changes in the video image, then the source of the jitter is identified as the breathing changes of the target person.
5. The rPPG anti-jitter method as described in claim 1, characterized in that, The breathing area includes at least one of the nostril area, lip area, and throat area.
6. The rPPG anti-jitter method as described in claim 5, characterized in that, The raw respiratory signals extracted from the nostril region, lip region, and throat region are fused to form a first respiratory signal, which is used for respiratory rate extraction. The first respiratory signal is represented as follows: ; in, This represents the first respiratory signal at the current time t. This represents the raw respiratory signal extracted from the nostril region at the current time t. This represents the raw respiratory signal extracted from the lip region at the current time t. This represents the raw respiratory signal extracted from the throat region at the current time t. , as well as These are the weighting coefficients.
7. The rPPG anti-jitter method as described in claim 6, characterized in that, The current image motion features and vehicle motion features are used to calculate the stability score of the first respiratory signal, and the respiratory rate is extracted from the first respiratory signal with a high stability score.
8. The rPPG anti-jitter method as described in claim 7, characterized in that, Vehicle motion characteristics include: longitudinal dynamic acceleration and angular velocity.
9. The rPPG anti-jitter method as described in claim 8, characterized in that, The stability score of the first respiratory signal is expressed as follows: ; ; in, The jitter score represents the first respiratory signal at time t; The stability score of the first respiratory signal at time t; , , These represent the overall facial offset at time t. Breathing area offset Global optical flow The overall facial offset was calculated separately. Breathing area offset Global optical flow Normalization is performed to obtain a uniform overall facial offset. Homogenized respiratory region offset Uniform global optical flow ; , Let these represent the longitudinal dynamic acceleration and angular velocity of the vehicle at time t, respectively. Regarding the longitudinal dynamic acceleration... angular velocity Normalization is performed to obtain a uniform longitudinal dynamic acceleration. Homogenized angular velocity ; , , , as well as These are the weighting coefficients.
10. An rPPG anti-jitter device, characterized in that, The device includes: The image feature extraction unit is used to extract current image motion features from video images containing target personnel; The motion feature extraction unit is used to extract the current vehicle motion features from the vehicle's motion signal; The breathing signal filtering unit is used to select breathing signals that reflect real breathing based on the current vehicle motion characteristics and image motion characteristics. The respiratory detection unit is used to extract the current respiratory rate of the target person from the selected respiratory signal.
11. The rPPG anti-jitter device as described in claim 10, characterized in that, Image motion features include the current overall facial offset, the respiratory region offset, and the global optical flow of the current image.
12. The rPPG anti-jitter device as described in claim 10, characterized in that, The motion feature extraction unit extracts vehicle motion features, including longitudinal dynamic acceleration and angular velocity, from the vehicle's three-axis acceleration and three-axis angular velocity.
13. The rPPG anti-jitter device as described in claim 10, characterized in that, The breathing signal screening unit combines the current vehicle motion characteristics and image motion characteristics to determine the jitter source that causes the change in breathing signal. The jitter source is the movement of the vehicle, the action of the target person, or the breathing change of the target person. If the jitter source is the breathing change of the target person, the current breathing signal is considered a valid breathing signal; otherwise, the current breathing signal is considered an invalid breathing signal.
14. The rPPG anti-jitter device as described in claim 13, characterized in that, When there is a significant change in vehicle motion characteristics, and the global optical flow, breathing region offset, and overall facial offset of the current image frame increase synchronously relative to the previous image frame, the breathing signal filtering unit identifies the jitter source as vehicle motion. When the vehicle motion characteristics are relatively stable, and only the breathing region offset and overall facial offset of the current image frame increase synchronously relative to the previous image frame, the breathing signal filtering unit identifies the jitter source as the target person's movement. When the vehicle motion characteristics are relatively stable, and the global optical flow and overall facial offset of the current image frame are close to zero relative to the previous image frame, and the low-frequency signal of the breathing region changes periodically and slightly in the video image, the breathing signal filtering unit identifies the jitter source as the target person's breathing changes.
15. A storage medium, characterized in that, The storage medium stores a computer program that is executed by a processor to implement the rPPG anti-jitter method as described in any one of claims 1 to 9.