Remote heart rate measurement method, device and equipment based on adaptive ROI selection
By employing an adaptive ROI selection method to segment facial image regions, calculate reflection angles and signal-to-noise ratios, and construct a signal-to-noise ratio model, the problem of ROI tracking loss caused by facial motion is solved, thereby improving the accuracy and robustness of remote heart rate measurement.
Patent Information
- Application Number
- CN202310335761.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-28
- Publication Date
- 2026-05-15
- Estimated Expiration
- 2043-03-28
AI Technical Summary
Existing technologies lack robustness of the region of interest (ROI) during remote heart rate measurement when the subject's face is moving, leading to inaccurate heart rate measurements.
An adaptive ROI selection method is adopted. By continuously acquiring facial images, segmenting them into multiple regions, calculating the reflection angle and signal-to-noise ratio, constructing a signal-to-noise ratio estimation model, and selecting the most suitable ROI region for signal synthesis, the accuracy of heart rate measurement is improved.
Under conditions of facial movement and changes in illumination, it significantly improved the signal-to-noise ratio (SNR) by 0.3152 dB and reduced the heart rate error PET6 by less than 0.5651%, thereby improving the robustness and accuracy of heart rate measurement.
Smart Images

Figure CN116327155B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of heart rate measurement technology, and in particular to a remote heart rate measurement method, apparatus, and device based on adaptive ROI selection. Background Technology
[0002] With the rise of infectious disease pandemics and increasing mortality rates from heart disease, public awareness of heart health has grown. Intra-periodic pulsed light (iPPG) technology is a contactless and non-invasive technique that measures and monitors heart rate without the sensor touching the skin. Long-term, accurate, and timely monitoring, combined with medical diagnosis, can significantly reduce the mortality rate from heart disease. iPPG works by detecting subtle changes in blood flow intensity across the skin during each cardiac cycle, caused by varying blood volume over time. Currently, many remote measurement technologies, such as radar and optical depth measurement, are used to measure heart rate (HR), blood pressure, respiratory rate (RR), and heart rate variability (HRV). Its non-contact nature makes it particularly suitable for drivers, newborns, and patients with skin conditions, thus attracting considerable research interest.
[0003] One of the key steps in iPPG heart rate measurement is the selection of the Region of Interest (ROI). The iPPG signal is generated by averaging images of the selected ROI over a period of time; this is the most commonly used method by researchers. The image averaging of a frame within an ROI will give the average G-channel value of all G-channel values within the ROI. Therefore, the already subtle variations in the G-channel values within the ROI are also averaged; this makes the iPPG signal less strong in facial videos. Numerous methods for ROI selection have been proposed over the past decade.
[0004] Since Poh et al. initially proposed modifying ROI selection from selecting the entire face to selecting ROIs representing 60% of the face's width and height, cheeks, forehead, and the area below the eyes have been subsequently proposed as ROIs. A more accurate ROI selection method based on facial segmentation was then proposed. Lam et al. also proposed an ROI selection criterion using facial landmarks, employing random patches of landmarks from the area below the eyes and within the lips. These methods select these ROIs in the first frame and do not change with video variations. Therefore, these methods are not suitable for head rotation and unstable lighting conditions.
[0005] To overcome unstable illumination, weighted selection of Regions of Interest (ROIs) has been proposed. Kumar et al. proposed selecting ROIs from the entire face using 20×20 pixel blocks and using weights to estimate the optimal combination of ROIs. Ewa M. Nowara et al. segmented the face into many regions using facial landmarks, evaluated the signal quality of each region to select ROIs, and selected five background regions to eliminate interference from illumination variations. This method has proven effective and is frequently used for ROI selection.
[0006] However, weighted ROI selection methods typically require choosing a time window to calculate the selection result and measure the optimal ROI within that time period. On the other hand, the length of this time window varies from 5 seconds to tens of seconds, and this method can fail because the subject's facial area is not always captured by the camera as the head rotates. When the subject's face moves, target area tracking can be lost, making it impossible to select the optimal ROI and resulting in failed heart rate measurements.
[0007] Therefore, improving the robustness of the ROI during remote heart rate measurement when the subject's face is moving is a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0008] The main technical problem to be solved by this invention is to improve the robustness of the ROI during remote heart rate measurement when the subject's face is moving, thereby improving the accuracy of remote heart rate measurement. In order to solve this technical problem, this invention provides a remote heart rate measurement method, device and equipment based on adaptive ROI selection, which combines a new ROI selection method to measure heart rate when the head is rotating.
[0009] According to a first aspect of the present invention, the present invention provides a remote heart rate measurement method based on adaptive ROI selection, comprising the following steps:
[0010] Continuously acquire multiple frames of facial images of the subject with their head at rest;
[0011] Each frame of facial image is segmented into multiple regions, and the reflection angle of each region is calculated;
[0012] Calculate the signal-to-noise ratio of each region in each frame of the facial image;
[0013] A signal-to-noise ratio estimation model is constructed by using the reflection angle, G-channel value, and signal-to-noise ratio of each region in each frame of facial image.
[0014] ROI is selected for each frame of facial image based on the signal-to-noise ratio estimation model;
[0015] The iPPG signal is obtained by synthesizing the ROI of multiple facial images.
[0016] The subject's heart rate was calculated based on the iPPG signal.
[0017] Furthermore, before the step of segmenting each frame of facial image into multiple regions and calculating the reflection angle of each region, the method further includes:
[0018] MediaPipeFaceMesh was applied to each acquired facial image to detect 468 3D landmarks on the subject's face;
[0019] Using the 3D landmarks of the face as reference points for calculating the surface normals and reflection angles of facial pixels, the set of possible surface normals at 3D landmark j is defined as follows:
[0020]
[0021] In the formula, This represents the vector from landmark j to the neighboring landmark i. Let m represent the vector from landmark j to another neighboring landmark k, where m is a positive integer representing the set P of possible surface normal vectors used to estimate landmark j. j The number of adjacent landmarks;
[0022] Using covariance analysis, from the surface normal vector set P j Select the most representative surface normal at landmark j.
[0023] Calculate the reflection angle θ of landmark j based on the surface normal. j :
[0024]
[0025] in, It is the surface normal of landmark j. It is the unit vector of the camera axis.
[0026] Furthermore, the step of segmenting each frame of facial image into multiple regions and calculating the reflection angle of each region includes,
[0027] The subject's facial image was segmented in three steps, resulting in N quadrilateral ROIs;
[0028] The reflection angle θ of the segmented region m Provided in the following manner:
[0029]
[0030] Where m is the index of the N segmented regions, θ m,n It is the normal angle of the four vertices of the nth region, and the reflection angle of each region is represented by gray.
[0031] Furthermore, the step of segmenting the subject's facial image into three steps includes:
[0032] Remove landmarks from the eyes, eyebrows, and mouth;
[0033] Remove landmarks caused by facial expression movements, including around the corners of the eyes, lips, and nose;
[0034] Connect the remaining landmarks in a straight line.
[0035] Furthermore, the step of constructing a signal-to-noise ratio estimation model using the reflection angle, G-channel value, and signal-to-noise ratio of each region in each frame of the facial image includes:
[0036] For each captured facial image frame, the average G channel value G of N regions is recorded. N Compared with the average reflection angle value θ N Then, record the number of frames t∈{1,…,T*FPS} of the current facial image, and store the recorded data in an N×t×2 matrix, where T is the number of seconds of video capture and FPS is the frame rate captured by the camera;
[0037] The average G channel value, average reflection angle value, and signal-to-noise ratio within the preset time window are recorded as a set of data.
[0038] For the average G channel value G N The two influencing factors, the average reflection angle, were tested for central tendency separately to remove outliers from the data.
[0039] The data after removing outliers is fitted with a nonlinear curve to obtain a signal-to-noise ratio estimation model with the reflection angle and G channel value as inputs and the signal-to-noise ratio as output.
[0040] Furthermore, the step of selecting the ROI for each frame of facial image based on the signal-to-noise ratio estimation model includes:
[0041] The minimum number of frames required to calculate the human heart rate is determined based on the camera frequency and the human heart rate.
[0042] The ROI is selected when the number of consecutive occurrences of the region with the highest signal-to-noise ratio estimated by the signal-to-noise ratio model is greater than or equal to the minimum number of frames.
[0043] Preferably, the number of segmented regions N = 91.
[0044] According to a second aspect of the present invention, the present invention provides a remote heart rate measurement device based on adaptive ROI selection, comprising the following modules:
[0045] The acquisition module is used to continuously acquire multiple frames of facial images of the subject with their head at rest.
[0046] The segmentation module is used to segment each frame of facial image into multiple regions;
[0047] The calculation module is used to calculate the reflection angle of each area;
[0048] The calculation module is also used to calculate the signal-to-noise ratio of each region in each frame of facial image;
[0049] The module is used to build a signal-to-noise ratio estimation model using the reflection angle, G channel value, and signal-to-noise ratio of each region in each frame of facial image.
[0050] The selection module is used to select the ROI for each frame of facial image based on the signal-to-noise ratio estimation model.
[0051] The synthesis module is used to synthesize the ROIs of multiple facial images to obtain iPPG signals;
[0052] The calculation module is also used to calculate the subject's heart rate based on the iPPG signal.
[0053] According to a third aspect of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the remote heart rate measurement method.
[0054] According to another aspect of the invention, the invention also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the remote heart rate measurement method.
[0055] The technical solution provided by this invention has the following beneficial effects:
[0056] This invention proposes a remote heart rate measurement method based on adaptive ROI selection. This method segments each frame of a facial image into multiple regions. By calculating the reflection angle, G-channel value, and signal-to-noise ratio (SNR) of each region, an SNR estimation model is constructed. Based on the SNR estimation model, the most suitable ROI for iPPG heart rate extraction in a frame can be identified, and a synthesized signal is used to estimate the heart rate. This method primarily addresses the issue of target region tracking loss due to facial motion, providing a new approach for dynamic recognition of facial ROI regions in iPPG technology. Experimental results show that, considering facial motion and changes in lighting conditions, using this method before and after processing improves the SNR by 0.3152 dB and the PET6 (absolute error heart rate less than or equal to 6) by 0.5651%. Attached Figure Description
[0057] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:
[0058] Figure 1 This is a flowchart illustrating the overall process of a remote heart rate measurement method based on adaptive ROI selection according to the present invention.
[0059] Figure 2 This is a schematic diagram illustrating the adaptive ROI selection process of the present invention;
[0060] Figure 3 This invention provides a complete pipeline for generating angle maps from images, where (a) is the calculation of surface reflection angle, (b) is the SNR estimation method, (c) is the establishment of the SNR estimation model, and (d) is the selection of ROI robustness.
[0061] Figure 4 This is a schematic diagram of the data in this invention;
[0062] Figure 5 This is a schematic diagram of the window sliding mechanism of the present invention;
[0063] Figure 6 For the present invention U t (f) Measurement signal template (a) and the spectrum of the IPPG signal (b);
[0064] Figure 7 This is a box plot of the G channel value and signal-to-noise ratio of the present invention. The solid line represents the median value, which is used to observe the overall trend of the results.
[0065] Figure 8 The ROI changes frequently across several consecutive frames;
[0066] Figure 9 A schematic diagram illustrating the robustness selection method for ROI;
[0067] Figure 10 This is a schematic diagram of the structure of a remote heart rate measurement device based on adaptive ROI selection according to the present invention;
[0068] Figure 11 This is a schematic diagram of the structure of an electronic device according to the present invention;
[0069] Figure 12 This is a comparison chart of SNR and PET6 results before and after using the ROI selection method of this invention. Detailed Implementation
[0070] To provide a clearer understanding of the technical features, objectives, and effects of the present invention, specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0071] Please refer to Figure 1This invention provides a remote heart rate measurement method based on adaptive ROI selection, the method comprising the following steps:
[0072] S1: Continuously acquire multiple frames of facial images of the subject with the head at rest;
[0073] S2: Divide each frame of facial image into multiple regions and calculate the reflection angle of each region;
[0074] S3: Calculate the signal-to-noise ratio of each region in each frame of the facial image;
[0075] S4: Construct a signal-to-noise ratio estimation model using the reflection angle, G channel value, and signal-to-noise ratio of each region in each frame of facial image;
[0076] S5: Select the ROI for each frame of facial image based on the signal-to-noise ratio estimation model;
[0077] S6: Synthesize the ROIs of multiple facial images to obtain the iPPG signal;
[0078] S7: Calculate the subject's heart rate based on the iPPG signal.
[0079] The adaptive ROI selection process is as follows: Figure 2 As shown, it mainly includes four steps:
[0080] 1) Segment the face into 91 regions and calculate the reflection angle for each region. 2) Calculate the signal-to-noise ratio (SNR) for each region. 3) Construct an SNR estimation model using the reflection angle and G-channel value, along with the calculated SNR. 4) Robustly select ROIs based on the SNR estimation model.
[0081] Next, the key points of the implementation process of this invention will be described in detail with reference to the specific accompanying drawings:
[0082] (1) Facial segmentation and reflection angle calculation:
[0083] like Figure 3 As shown, the overall pipeline for generating angle maps from images is illustrated. MediaPipeFaceMesh is applied to each frame to detect 468 3D landmarks on the subject's face, such as... Figure 3 As shown in (a), 3D facial landmarks serve as reference points for subsequent calculation of the surface normal vectors and reflection angles of facial pixels. Figure 2 Let P be the set of possible surface normal vectors at landmark j. j ,Right now
[0084]
[0085] In the formula, This represents the vector from landmark j to the neighboring landmark i. Let m represent the vector from landmark j to another neighboring landmark k, where m is a positive integer representing the set P of possible surface normal vectors used to estimate landmark j. j The number of adjacent landmarks;
[0086] Then, the analysis of covariance (ANCOVA) method was used to extract the surface normal vector set P. j Select the most representative surface normal at landmark j. The reflection angle θ of landmark j. j It is estimated to be:
[0087]
[0088] in, It is the surface normal of landmark j. It is the unit vector of the camera axis. The normal vector of the estimated coordinate point is as follows: Figure 3 As shown in (b).
[0089] Next, the subject's face was segmented in three steps: 1) Landmarks around the eyes, eyebrows, and mouth were removed, as these landmarks were useless for BVP signals. 2) Landmarks caused by facial expression movements, including the corners of the eyes, lips, and around the nose, were removed. 3) The remaining landmarks were connected with straight lines. Therefore, the subject's face was segmented into 91 quadrilateral Regions of Interest (ROIs), as shown below. Figure 3 As shown in (c). Finally, the reflection angle of the segmented region can be given in the following way:
[0090]
[0091] Where m is the index of the 91 segmented regions, θ m,n These are the normal angles of the four vertices of the nth region. The reflection angle of each region is represented by grayscale, such as... Figure 3 As shown in (d).
[0092] Within the 91 defined regions, the signal-to-noise ratio (SNR) is used to measure the impact of different reflection angles and illumination intensities on signal quality. Previous research has demonstrated that among the RGB channels of a color camera, the G channel contains heart rate information most strongly correlated with the BVP signal. Therefore, the G channel signal strength is used here to measure the SNR.
[0093] For each captured image frame, the average G channel value G of N=91 regions is recorded. N Compared with the average reflection angle value θ N Where N∈{1,…,91}, and the frame number t∈{1,…,T*FPS} of the current image is recorded. Therefore, the data information can be stored in an N×t×2 matrix, where T is the number of seconds of video capture and FPS is the frame rate captured by the camera.
[0094] (2) Signal-to-noise ratio estimation:
[0095] A brief time-frequency analysis was performed on the acquired G-channel signal. To ensure the effectiveness of this research method in the short term, such as... Figure 5 As shown, a sliding analysis was performed using a 10-second window. The spectrum of the pulse signal acquired within 10 seconds was normalized to obtain the estimated heart rate signal. The average measurement result within 10 seconds of contact with the pulse oximeter was then used as the true heart rate signal. Since the subject's face was required to remain still, the average reflection angle within 10 seconds could be used to represent the reflection angle of the signal during that time period.
[0096] The signal-to-noise ratio (SNR) evaluation method of DeHaan et al. was used to assess the availability and quality of rPPG signals. Frequency domain evaluation was performed on a time-domain dataset within a 10-second filter. This method calculates the energy and residual portions of the power spectrum around the harmonics; the ratio of these two is then expressed as the SNR in dB. The formulas for signal, noise, and SNR calculation are as follows:
[0097]
[0098]
[0099]
[0100] in The spectrum of the iPPG signal is given by f, where f is the frequency in beats per minute (bpm). The required signal band is within the normal human pulse frequency range, from 30 bpm to 240 bpm.
[0101] like Figure 6 As shown, the dashed line represents U. t (f) Measurement signal template, the solid line represents the spectrum of the iPPG signal. The signal-to-noise ratio (SNR) is calculated using a measurement signal template within the estimated heart rate signal spectrum, centered on the signal peak with a range of 2.5 frequency units above and below it, and centered on the contact sensor pulse rate with a range of 5 frequency units above and below it, to account for heart rate variability. The SNR is measured by the energy ratio of the components inside and outside the template.
[0102] (3) Establishment of signal-to-noise ratio estimation model
[0103] Data was recorded within a 10-second time window. Since the face was stationary and the light source constant, the reflection angle and the mean G-channel value of a fixed facial area did not change significantly. Therefore, the average G-channel value, average reflection angle, and signal-to-noise ratio within 10 seconds were recorded as a set of data. To establish the mathematical relationship between them, box plots were created for these two influencing factors to perform a central tendency test separately, removing outliers to achieve preliminary data screening and reduce the interference of outliers on the data model. Figure 7 The data results for the G channel and SNR are shown in the figure. "+" in the figure indicates outliers. By observing the trend of the median value, the processed data is fitted with a nonlinear curve. Similarly, the data for reflection angle and SNR are processed, ultimately obtaining the reflection angle and G channel value as inputs, and the signal-to-noise ratio estimation model as the output.
[0104] (4) ROI robustness selection
[0105] The obtained signal-to-noise ratio (SNR) estimation model is directly applied to the collected simulated driving dataset. In each frame, a region with the highest SNR is estimated. The original iPPG signal is obtained by synthesizing the signals from the best region frame by frame. However, when synthesizing signals from selected regions, issues such as… Figure 8 As shown, selecting different regions as ROIs in several consecutive frames of images can cause the iPPG signal synthesized from these frames to become invalid, resulting in a decrease in the signal-to-noise ratio. This is because the values collected due to changes in the ROI region are changes in light intensity caused by non-physiological phenomena, and the expected physiological information cannot be obtained from them. Therefore, it is necessary to increase the robustness of the selection of ROI regions.
[0106] Here, the highest normal heart rate is 4Hz, with a period of 0.25 seconds. When using a camera with a frequency of 30Hz, at least 7.5 frames are needed to reconstruct and calculate the human heart rate. Figure 9 As shown, a threshold F is set here to limit the transformation of ROI regions. First, the best region IDs within 1 second are stored. When the best regions are consecutively the same and the number is greater than the threshold F (F≥8), the ROI transformation is performed.
[0107] The following describes a remote heart rate measurement device based on adaptive ROI selection provided by the present invention. The remote heart rate measurement device described below can be referred to in correspondence with the remote heart rate measurement method described above.
[0108] like Figure 10 As shown, a remote heart rate measurement device based on adaptive ROI selection includes the following modules:
[0109] The acquisition module 001 is used to continuously acquire multiple frames of facial images of the subject with the head at rest.
[0110] The segmentation module 002 is used to segment each frame of facial image into multiple regions;
[0111] Calculation module 003 is used to calculate the reflection angle of each area;
[0112] The calculation module 003 is also used to calculate the signal-to-noise ratio of each region in each frame of facial image;
[0113] Module 004 is used to construct a signal-to-noise ratio estimation model using the reflection angle, G channel value, and signal-to-noise ratio of each region in each frame of facial image;
[0114] Module 005 is used to select the ROI for each frame of facial image based on the signal-to-noise ratio estimation model.
[0115] The synthesis module 006 is used to synthesize the ROI of multiple frames of facial images to obtain the iPPG signal;
[0116] The calculation module 003 is also used to calculate the subject's heart rate based on the iPPG signal.
[0117] Based on, but not limited to, the above-described apparatus, the computing module 003 is further configured to:
[0118] MediaPipeFaceMesh was applied to each acquired facial image to detect 468 3D landmarks on the subject's face;
[0119] Using the 3D landmarks of the face as reference points for calculating the surface normals and reflection angles of facial pixels, the set of possible surface normals at 3D landmark j is defined as follows:
[0120]
[0121] In the formula, This represents the vector from landmark j to the neighboring landmark i. Let m represent the vector from landmark j to another neighboring landmark k, where m is a positive integer representing the set P of possible surface normal vectors used to estimate landmark j. j The number of adjacent landmarks;
[0122] Using covariance analysis, from the surface normal vector set P j Select the most representative surface normal at landmark j.
[0123] Calculate the reflection angle θ of landmark j based on the surface normal. j :
[0124]
[0125] in, It is the surface normal of landmark j. It is the unit vector of the camera axis.
[0126] Based on, but not limited to, the above-described apparatus, the segmentation module 002 is specifically used for:
[0127] The subject's facial image was segmented in three steps: landmarks on the eyes, eyebrows, and mouth were removed; landmarks caused by facial expression movements, including the corners of the eyes, lips, and around the nose, were removed; and the remaining landmarks were connected by straight lines, ultimately segmenting the image into N=91 quadrilateral ROIs.
[0128] The reflection angle θ of the segmented region m Provided in the following manner:
[0129]
[0130] Where m is the index of the N segmented regions, θ m,n It is the normal angle of the four vertices of the nth region, and the reflection angle of each region is represented by gray.
[0131] Based on, but not limited to, the above-described apparatus, the construction module 004 is specifically used for:
[0132] For each captured facial image frame, the average G channel value G of N regions is recorded. N Compared with the average reflection angle value θ N Then, record the number of frames t∈{1,…,T*FPS} of the current facial image, and store the recorded data in an N×t×2 matrix, where T is the number of seconds of video capture and FPS is the frame rate captured by the camera;
[0133] The average G channel value, average reflection angle value, and signal-to-noise ratio within the preset time window are recorded as a set of data.
[0134] The two influencing factors, the average G-channel value and the average reflection angle value, were tested for central tendency separately to remove outliers from the data.
[0135] The data after removing outliers is fitted with a nonlinear curve to obtain a signal-to-noise ratio estimation model with the reflection angle and G channel value as inputs and the signal-to-noise ratio as output.
[0136] Based on, but not limited to, the above-described apparatus, the selection module 005 is specifically used for:
[0137] The minimum number of frames required to calculate the human heart rate is determined based on the camera frequency and the human heart rate.
[0138] The ROI is selected when the number of consecutive occurrences of the region with the highest signal-to-noise ratio estimated by the signal-to-noise ratio model is greater than or equal to the minimum number of frames.
[0139] like Figure 11 The diagram illustrates the physical structure of an electronic device, which may include a processor 610, a communication interface 620, a memory 630, and a communication bus 640. The processor 610, communication interface 620, and memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute the steps of a remote heart rate measurement method, specifically including: continuously acquiring multiple frames of facial images of the subject with their head at rest; segmenting each frame of the facial image into multiple regions and calculating the reflection angle of each region; calculating the signal-to-noise ratio (SNR) of each region in each frame of the facial image; constructing an SNR estimation model using the reflection angle, G-channel value, and SNR of each region in each frame of the facial image; selecting a Region of Interest (ROI) for each frame of the facial image based on the SNR estimation model; synthesizing the ROIs of multiple frames of the facial image to obtain an iPPG signal; and calculating the subject's heart rate based on the iPPG signal.
[0140] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0141] In another aspect, embodiments of the present invention also provide a storage medium storing a computer program, which, when executed by a processor, implements the steps of the aforementioned remote heart rate measurement method, specifically including: continuously acquiring multiple frames of facial images of the subject with their head at rest; segmenting each frame of facial image into multiple regions and calculating the reflection angle of each region; calculating the signal-to-noise ratio (SNR) of each region in each frame of facial image; constructing an SNR estimation model using the reflection angle, G-channel value, and SNR of each region in each frame of facial image; selecting a Region of Interest (ROI) for each frame of facial image based on the SNR estimation model; synthesizing the ROIs of multiple frames of facial images to obtain an iPPG signal; and calculating the subject's heart rate based on the iPPG signal.
[0142] This invention, through the implementation of the aforementioned remote heart rate measurement method, device, and equipment based on adaptive ROI selection, along with the corresponding storage medium, can identify the most suitable region for iPPG heart rate extraction within a single image frame as the ROI. Multiple image frames are then synthesized to obtain an iPPG signal. Fourier transform is applied to the iPPG signal for frequency domain analysis, and the frequency corresponding to the point with the largest amplitude is selected as the heart rate. This method primarily addresses the issue of target region tracking loss due to facial movement, providing a new approach for the dynamic recognition of facial ROI regions using iPPG technology. Experimental results are as follows... Figure 12 As shown, Figure 12 The results show that, taking into account facial movement and changes in lighting conditions, using the method of the present invention before and after processing can improve the signal-to-noise ratio (SNR) by 0.3152 dB and the PET6 (absolute value of error heart rate less than or equal to 6) by 0.5651%.
[0143] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0144] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. In the unit claims listing several devices, several of these devices may be embodied by the same hardware item. The use of the terms first, second, and third, etc., does not indicate any order and can be interpreted as identifiers.
[0145] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A remote heart rate measurement method based on adaptive ROI selection, characterized in that, Includes the following steps: Continuously acquire multiple frames of facial images of the subject with their head at rest; Each frame of facial image is segmented into multiple regions, and the reflection angle of each region is calculated; Calculate the signal-to-noise ratio of each region in each frame of the facial image; A signal-to-noise ratio estimation model is constructed by using the reflection angle, G-channel value, and signal-to-noise ratio of each region in each frame of facial image. ROI is selected for each frame of facial image based on the signal-to-noise ratio estimation model; The iPPG signal is obtained by synthesizing the ROI of multiple facial images. The subject's heart rate was calculated based on the iPPG signal; Before the step of segmenting each frame of facial image into multiple regions and calculating the reflection angle of each region, the method further includes: MediaPipe FaceMesh was applied to each acquired facial image to detect 468 3D landmarks on the subject's face; Using 3D facial landmarks as reference points for calculating the surface normal vectors and reflection angles of facial pixels, the 3D landmarks... j The set of possible surface normal vectors is defined as follows: In the formula, Indicates from the landmark j To nearby landmarks i vector, Indicates from the landmark j To another nearby landmark k vector, m It is a positive integer representing the value used to estimate the landmark. j Set of possible surface normal vectors The number of adjacent landmarks; Using covariance analysis to analyze the set of surface normal vectors Select landmark j The most representative surface normal ; Calculate the landmark based on the surface normal. j Reflection angle : in, It is a landmark j Surface normal, It is the unit vector of the camera axis; The step of segmenting each frame of facial image into multiple regions and calculating the reflection angle of each region includes, The subject's facial image was segmented in three steps, resulting in... N The ROI portion is a quadrilateral; Reflection angle of the segmented region Provided in the following manner: in, m yes N The sequence number of each segmented region It is the first n The normal angles of the four vertices of each region, and the reflection angle of each region is represented by grayscale.
2. The remote heart rate measurement method according to claim 1, characterized in that, The step of segmenting the subject's facial image into three steps includes: Remove landmarks from the eyes, eyebrows, and mouth; Remove landmarks caused by facial expression movements, including the corners of the eyes, lips, and around the nose; Connect the remaining landmarks in a straight line.
3. The remote heart rate measurement method according to claim 1, characterized in that, The step of constructing a signal-to-noise ratio estimation model using the reflection angle, G-channel value, and signal-to-noise ratio of each region in each frame of facial image includes: For each captured facial image frame, the average G-channel value of N regions is recorded. Compared with the average reflection angle value Then record the frame number t∈{1,…,T*FPS} of the current facial image, and store the recorded data in a... In the matrix, T is the number of seconds of video capture, and FPS is the frame rate captured by the camera; The average G channel value, average reflection angle value, and signal-to-noise ratio within the preset time window are recorded as a set of data. The two influencing factors, the average G-channel value and the average reflection angle value, were tested for central tendency separately to remove outliers from the data. The data after removing outliers is fitted with a nonlinear curve to obtain a signal-to-noise ratio estimation model with the reflection angle and G channel value as inputs and the signal-to-noise ratio as output.
4. The remote heart rate measurement method according to claim 1, characterized in that, The step of selecting the Region of Interest (ROI) for each frame of facial image based on the signal-to-noise ratio estimation model includes: The minimum number of frames required to calculate the human heart rate is determined based on the camera frequency and the human heart rate. The ROI is selected when the number of consecutive occurrences of the region with the highest signal-to-noise ratio estimated by the signal-to-noise ratio model is greater than or equal to the minimum number of frames.
5. The remote heart rate measurement method according to claim 1 or 3, characterized in that, The number of regions N=91.
6. A remote heart rate measurement device based on adaptive ROI selection, characterized in that, The steps for implementing the remote heart rate measurement method as described in any one of claims 1-5 include the following modules: The acquisition module is used to continuously acquire multiple frames of facial images of the subject with their head at rest. The segmentation module is used to segment each frame of facial image into multiple regions; The calculation module is used to calculate the reflection angle of each area; The calculation module is also used to calculate the signal-to-noise ratio of each region in each frame of facial image; The module is used to build a signal-to-noise ratio estimation model using the reflection angle, G channel value, and signal-to-noise ratio of each region in each frame of facial image. The selection module is used to select the ROI for each frame of facial image based on the signal-to-noise ratio estimation model. The synthesis module is used to synthesize the ROIs of multiple facial images to obtain iPPG signals; The calculation module is also used to calculate the subject's heart rate based on the iPPG signal.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the remote heart rate measurement method as described in any one of claims 1-5.
8. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the remote heart rate measurement method as described in any one of claims 1-5.