Video heart rate detection method based on target region set determination and multi-signal fusion

By identifying a set of target regions that are robust to uneven lighting and motion, and employing a multi-signal fusion method, the accuracy and robustness issues of video heart rate detection under complex lighting and motion scenarios were resolved, achieving higher detection accuracy and reliability.

CN118298486BActive Publication Date: 2026-05-19HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HEFEI UNIV OF TECH
Filing Date
2024-04-02
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing video heart rate detection methods lack accuracy and robustness in complex lighting and motion scenarios, especially due to the effects of uneven lighting and motion artifacts.

Method used

By identifying a set of target regions that are robust to uneven lighting and motion, a multi-signal fusion method is employed, including face detection, feature point tracking, chromatic aberration signal processing, and image pyramid operation, to extract high-quality heart rate signals.

Benefits of technology

It improves the accuracy and robustness of heart rate detection in complex lighting and motion scenarios, reduces noise interference, and enhances detection accuracy and reliability under real-world conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118298486B_ABST
    Figure CN118298486B_ABST
Patent Text Reader

Abstract

The application discloses a video heart rate detection method based on target region set determination and multi-signal fusion, which comprises the following steps: 1, adaptive region generation and target region determination against uneven illumination; 2, multi-scale target region generation and determination; 3, target region generation and determination against disturbance; 4, target region set construction against uneven illumination and motion robustness based on an image pyramid; and 5, signal fusion and heart rate estimation based on the target region set. In one aspect, the application obtains a target region set against uneven illumination and motion robustness from video image frames, so that a pulse signal with good quality can be obtained. In another aspect, the application fuses signals obtained from the target region set based on Gaussian prior and uniform prior, further improves the accuracy of heart rate detection, and further expands the application scene of non-contact and robust heart rate detection technology in more life scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of heart rate detection technology, specifically relating to a video heart rate detection method based on target region set determination and multi-signal fusion. Background Technology

[0002] Heart rate (HR) is one of the most important vital signs in the human body, serving as a crucial indicator of health and mental well-being. Therefore, monitoring heart rate is of great significance. Currently, commonly used clinical methods for heart rate monitoring typically require professional operation and involve viscous sensors, electrodes, leads, wires, and chest straps. These methods have limitations, such as causing discomfort or allergies to the subject due to sensor or electrode placement, and are not suitable for long-term monitoring. Furthermore, they may not be usable in certain specific scenarios, such as non-invasive monitoring, sensitive populations, burn patients, and monitoring in intensive care units (ICUs). Non-contact heart rate monitoring methods can effectively address these challenges.

[0003] Existing non-contact heart rate detection methods include those based on the Doppler effect, thermal imaging, and imaging photo-plethysmography (IPPG). Among these, heart rate detection based on imaging photo-plethysmography, also known as video-based heart rate detection, leverages the difference in light absorption between blood and other skin tissues. With each heartbeat, the blood volume within the vessels undergoes periodic changes, resulting in periodic variations in light absorption and reflection, macroscopically reflected in changes in skin color. This variation can be captured by capturing video of facial skin. Signal processing and analysis then enable heart rate detection. Its advantages—requiring only a camera, no physical contact or specialized operation, low cost, convenience, and the ability to detect multiple individuals—have garnered widespread attention. Initially, video-based heart rate measurements were performed under well-lit, stationary, and controlled conditions, yielding good results. However, the heart rate signal caused by the pulse is very weak and easily affected by changes in lighting and motion artifacts, thus greatly limiting its application scenarios.

[0004] Current research on video-based heart rate detection generally focuses on only one aspect: lighting or motion. However, in everyday life scenarios, motion and lighting rarely exist independently. Therefore, this invention proposes a video heart rate detection method resistant to uneven lighting and motion artifacts, supplementing research on accurate and rapid heart rate detection under real-world conditions involving complex lighting and motion. Because different parts of the skin absorb light differently, different regions of interest (ROIs) on the face, such as the cheeks, chin, and forehead, have varying degrees of sensitivity to light absorption. Therefore, spatial averaging of the entire face is inaccurate, and a large skin area leads to more noise in the extracted signal. Summary of the Invention

[0005] The present invention addresses the shortcomings of the prior art by providing a video heart rate detection method based on target region set determination and multi-signal fusion. This method aims to reduce the impact of uneven lighting and motion artifacts on heart rate detection, thereby improving the accuracy and robustness of heart rate detection in complex lighting and motion scenarios.

[0006] The present invention adopts the following solution to solve the technical problem:

[0007] The video heart rate detection method based on target region set determination and multi-signal fusion of the present invention is characterized by the following steps:

[0008] Step 1: Obtain the target region I1 for resistance to uneven illumination from the region of interest of the facial skin in the subject's T-frame video image;

[0009] Step 2: Based on the face detection bounding boxes of T frames of video images, generate J bounding boxes with different aspect ratios for T frames of video images, and select the bounding box with the highest signal-to-noise ratio as the multi-scale target region M1;

[0010] Step 3: Based on the target region M1, generate L rectangular boxes for the T frames of video images. According to the mean square error of the color difference signal of the T frames of video images in the L rectangular boxes, select the bounding box containing the signal with the largest mean square error as the target region for anti-disturbance, that is, the motion-robust target region M2.

[0011] Step 4: Based on image pyramid operations, combine the target region I1 which is resistant to uneven illumination and the target region M2 which is robust to motion to determine the set of target regions that are both resistant to uneven illumination and robust to motion. Among them, TargetROI f This represents the target region of the f-th layer image;

[0012] Step 5: Based on the target region set The signals of the target regions in each layer of the image are extracted layer by layer and fused to obtain the dominant frequency f of the fused signal. max This allows us to estimate the heart rate.

[0013] The video heart rate detection method based on target region set determination and multi-signal fusion described in this invention is characterized in that step 1 includes:

[0014] Step 1.1: Use a face detector to locate the face in the first frame of the video image and obtain the face rectangle detection box B0 = [x0, y0, w0, h0] in the first frame of the video image, where (x0, y0) are the coordinates of the upper left point of the face rectangle detection box B0, and w0 and h0 are the width and height of the face rectangle detection box B0, respectively.

[0015] Step 1.2: Based on the face rectangle detection box B0, use the face feature point detection algorithm to determine the coordinate information (x1) of N facial feature points in the first frame of the video image. (1) ,y1 (1) ),(x2 (1) ,y2 (1) ),(x n (1) ,y n (1) ),...,(x N (1) ,y N (1) ), where (x n (1) ,y n (1) ) represents the coordinate information of the nth facial feature point in the first frame of the video image, and is derived from the N facial feature points (x1) in the first frame of the video image. (1) ,y1 (1) ),(x2 (1) ,y2 (1) ),(x n (1) ,y n (1) ),...,(x N (1) ,y N (1) In the first frame of the video image, select M facial feature points to form the region of interest (ROI) for the facial skin; M≤N; n∈[1,N];

[0016] Step 1.3: Based on the region of interest (ROI1) of the facial skin in the first frame of the video image, use a feature point tracking algorithm to track M facial feature points in the second to Tth frames of the video image, where the region of interest (ROI1) of the facial skin in the i-th frame of the video image is... iThe coordinate information of M facial feature points is denoted as (x1) (i) ,y1 (i) ),(x2 (i) ,y2 (i) ),...,(x m (i) ,y m (i) ),...,(x M (i) ,y M (i) ), where (x m (i) ,y m (i) ) represents the region of interest (ROI) of facial skin in the i-th frame of the video image. i The coordinate information of the m-th facial feature point, 2≤i≤T;

[0017] Step 1.4: Identify the Region of Interest (ROI) for the facial skin in the i-th frame of the video image. i The pixel space is converted from RGB color space to LAB color space, and the region of interest (ROI) of facial skin in the i-th frame of the video image is calculated. i The L-channel luminance value of each pixel is used to calculate the maximum L-channel luminance value. max The Euclidean distance of the ROI i Divide into Q-block sub-regions;

[0018] Step 1.5: Calculate the average pixel value of each sub-region in the T-frame video image frame by frame, and calculate the signal-to-noise ratio of each sub-region based on the average pixel value of each sub-region as a quality index. Based on the quality index of each sub-region, select V best sub-regions from the Q sub-regions as the target region I1.

[0019] Step 2 includes:

[0020] Step 2.1: Calculate the center point of the face rectangle detection box B0 Where, x c0 y c0 C represents the center point of the rectangular detection box B0 for the face. c0 The horizontal and vertical coordinates;

[0021] With C c0 Generate J bounding boxes with different aspect ratios centered on the boundary. Where j = 1, 2, ..., J, J is the total number of bounding boxes, B j Let B represent the j-th bounding box, and B j =[x j ,y j ,w j,h j ], w j h j Represents the j-th bounding box B j The width and height are obtained from equations (1) and (2), x j y j Represents the j-th bounding box B j The coordinates of the upper left corner point are obtained from equations (3) and (4);

[0022] w j =σ w ·w0 (1)

[0023] h j =σ h ·w j (2)

[0024] In equations (1) and (2), σ w , σ h These represent coefficients for width and height, respectively.

[0025]

[0026]

[0027] Step 2.2: Calculate the j-th bounding box B frame by frame. j The pixel mean time series is used to obtain the j-th bounding box B. j The RGB channel time series of a T-frame video image, including: the j-th bounding box B j Time series R channel of T-frame video images [j] =[R1 [j] R2 [j] ,...,R i [j] ,...,R T [j] ], the j-th bounding box B j Time series G channels of T-frame video images [j] =[G1 [j] G2 [j] ,...,G i [j] ,...,G T [j] ] and the j-th bounding box B j Time series B channel of T-frame video image [j] =[B1 [j] B2 [j] ,...,B i [j] ,...,B T[j] ], where R i [j] Represents the j-th bounding box B j The average pixel value of the R channel of the i-th frame of the video image, G i [j] Represents the j-th bounding box B j The average pixel value of the G channel of the i-th frame of the video image, B i [j] Represents the j-th bounding box B j The average pixel value of the B channel in the i-th frame of the video image;

[0028] Step 2.3: Use equation (5) to preprocess the RGB channel time series of the T-frame video image into the j-th bounding box B. j single-channel color difference signal p of a medium-T frame video image [j] Thus, J bounding boxes are obtained. Color difference signal of a medium-T frame video image [P] [1] ,P [2] ,...,P [j] ,...,P [J] ];

[0029]

[0030] In equation (5), R' [j] G' [j] B' [j] These are the j-th bounding box B j Transpose of the R-channel time series, G-channel time series, and B-channel time series of a mid-T frame video image; p1 [j] This indicates the use of the j-th bounding box B j The chroma signal defined by the projection of the G-channel and B-channel time series of a mid-T frame video image, p2 [j] This indicates the use of the j-th bounding box B j The chrominance signal defined by the projection of the R, G, and B channel time series of a mid-T frame video image, α j This represents the tuning parameters for separating the illumination and the pulse signal, and σ(·) represents the standard deviation of a given signal;

[0031] Step 2.4: Calculate the j-th bounding box B using equation (6). j Single-channel color difference signal P in a medium-T frame video image [j] Signal-to-noise ratio (SNR) j Thus, the signal-to-noise ratio set is obtained.

[0032]

[0033] In equation (6), U fs Represents a gate function;

[0034] Step 2.5: Record the maximum value in SNR as SNR. max1 and from Select SNR max1 The corresponding bounding box B max1 As the multi-scale target region M1, B max1 The width and height are denoted as w. m1 h m1 B max1 The center point is denoted as C. P0 =[x p0 ,y p0 ]; where x p0 y p0 Indicates center point C P0 The horizontal and vertical coordinates.

[0035] Step 3 includes:

[0036] Step 3.1: In B max1 center point C P0 Determine L neighboring points around the perimeter, and use equation (7) to obtain the l-th neighboring point C. pl Thus, we obtain C pl Centered on a point with width and height w m1 h m1 The l-th rectangle BP l This results in L rectangular frames.

[0037] C pl =C p0 +p(l)·s (7)

[0038] In equation (7), p(l) represents the l-th neighboring point C. pl Relative center point C P0 The position s represents any neighboring point relative to the center point C. P0 The distance;

[0039] Step 3.2: Calculate the frame-by-frame BP of the l-th rectangle. l The pixel mean time series is used to obtain the l-th bounding box BP. l The RGB channel time series of a T-frame video image, including: the l-th bounding box BP l Time series of the R channel of a medium-T frame video image The l-th rectangle BP l Time series of G channels of medium T-frame video images The l-th rectangle BP lTime series of B channel of T-frame video image in, BP represents the l-th rectangle. l The average pixel value of the R channel of the i-th frame of the video image. BP represents the l-th rectangle. l The average pixel value of the G channel of the i-th frame of the video image. BP represents the l-th rectangle. l The average pixel value of the B channel in the i-th frame of the video image;

[0040] Step 3.3: Use equation (8) to obtain the l-th rectangular frame BP l Single-channel color difference signal of a medium-T frame video image This results in L rectangular frames. Color difference signal of a medium-T frame video image

[0041]

[0042] In equation (8), These represent the l-th rectangle BP. l Transpose of the R-channel time series, G-channel time series, and B-channel time series of a T-frame video image; This indicates the use of the l-th rectangle BP l The chroma signal defined by the projection of the G-channel and B-channel time series of a mid-T frame video image. This indicates the use of the l-th rectangle BP l The chroma signal defined by the projection of the time series of the R, G, and B channels of a mid-T frame video image. Indicates the tuning parameters, and

[0043] Step 3.4: Calculate the l-th rectangle BP using equation (9). l Single-channel color difference signal of a medium-T frame video image Root mean square error (MSE) l Thus, the root mean square error set is obtained.

[0044]

[0045] In equation (9), β represents the summation function, and α l Let represent the l-th mean square error function, and W is signal length, P ref [l] Indicates the reference pulse signal, and

[0046] Step 3.5: Record the maximum value in MSE as MSE. max2 and from Select MSE max2 The corresponding rectangle BP max2 M2 is the target area for anti-disturbance; BP max2 The width and length are denoted as w. m2 h m2 , BP max2 The center point is denoted as C. m2 =(x m2 ,y m2 ), where x m2 y m2 Indicates center point C m2 The horizontal and vertical coordinates.

[0047] Step 4 includes:

[0048] Step 4.1: Using C m2 =(x m2 ,y m2 Using a given center, generate a set of 2F+1 rectangular frames representing F layers with the same center point but different scaling scales. Where F represents the number of layers to be enlarged or reduced, and D... f Let D represent the enlarged or reduced frame of the f-th layer rectangle, and D f =[x m2 ,y m2 ,w f ,h f ], where w f h f Represents the rectangle D at the f-th level. f The width and height are obtained from equations (10) and (11):

[0049] w f =w m2 ×K f (10)

[0050] h f =h m2 ×K f (11)

[0051] In equations (10) and (11), -F≤f≤F, and K represents the scale factor that controls the ratio of the size of the target region M2;

[0052] Step 4.2: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require the full context.] By taking the intersection of each of the target regions I1 (resistance to uneven lighting) with the target region I1, we can obtain the set of target regions for both uneven lighting and motion. Among them, TargetROI f This represents the target region of the image at layer f.

[0053] Step 5 includes:

[0054] Step 5.1: Calculate the target region (TargetROI) of the f-th layer image frame by frame. f The pixel mean time series is used to obtain the RGB channel time series of T frames of video images, including: TargetROI f Time series of the R channel of a medium-T frame video image TargetROI f Time series of G channels of medium T-frame video images TargetROI f Time series of B channel of T-frame video image in, Represents the target region TargetROI of the f-th layer image. f The average pixel value of the R channel of the i-th frame of the video image. Represents the target region TargetROI of the f-th layer image. f The average pixel value of the G channel of the i-th frame of the video image. Represents the target region (TargetROI) of the f-th layer image. f The average pixel value of the B channel in the i-th frame of the video image;

[0055] Step 5.2: Use equation (12) to obtain the target region TargetROI of the f-th layer image. f Single-channel color difference signal of the i-th frame of video image Thus, the target region TargetROI of the f-th layer image is obtained. f Single-channel color difference signal of a medium-T frame video image This leads to the target region set. Color difference signal of a medium-T frame video image

[0056]

[0057] In equation (12), These represent the target regions (TargetROI) of the f-th layer image, respectively. f Transpose the R-channel time series, G-channel time series, and B-channel time series of a T-frame video image; Indicates the use of TargetROI f The chroma signal defined by the projection of the G-channel and B-channel time series of a mid-T frame video image. Indicates the use of TargetROIf The chrominance signal defined by the time-series projection of the R, G, and B channels of a mid-T frame video image. Indicates the tuning parameters, and

[0058] Step 5.3: Obtain the target region (TargetROI) of the f-th layer image using equations (13) and (14). f Single-channel color difference signal of a medium-T frame video image Average fusion weight λ f Gaussian fusion weight λ f (μ,σ):

[0059]

[0060]

[0061] In equation (14), μ and σ represent the mean and standard deviation of all values ​​of f = [-F, F].

[0062] Step 5.4: Calculate the fused signal

[0063] Step 5.5: Heart rate estimation;

[0064] Perform spectral analysis on the fused signal Pulse to obtain the dominant frequency f of the fused signal. max Thus, the estimated heart rate value is time × f. max Where time represents the duration of T frames of video image.

[0065] The present invention provides an electronic device, including a memory and a processor, wherein the memory is used to store a program that supports the processor in executing the video heart rate detection method, and the processor is configured to execute the program stored in the memory.

[0066] The present invention discloses a computer-readable storage medium on which a computer program is stored, wherein the computer program, when executed by a processor, performs the steps of the video heart rate detection method.

[0067] Compared with existing technologies, the beneficial effects of this invention are reflected in:

[0068] 1. The target region set obtained by this invention, which is robust to uneven lighting and motion, can acquire better quality signals compared to the overall spatial averaging of a single ROI. Especially when uneven lighting exists, the brightness information of different facial regions varies, and the signal quality obtained is also significantly different. Therefore, determining the target region set is necessary and important. At the same time, the target region set can be combined with more heart rate information, complementing each other. In addition, compared with single signal methods, by fusing multi-scale signals based on prior knowledge to extract potential common heart rate components from multiple sets of signals, it is more robust to noise interference in video-based heart rate detection in reality.

[0069] 2. This invention segments facial skin regions based on facial illumination intensity, adaptively classifies pixels within the facial region based on facial illumination intensity information, obtains multiple Regions of Interest (ROIs) with different brightness distributions on the face, and filters out target regions robust to uneven illumination based on signal quality evaluation indicators. Compared with previous region rule segmentation, this clustering segmentation is more flexible in its operation on pixels, and can concentrate pixels with good illumination information into one region instead of scattering them into different sub-ROIs. Finally, region selection is based on signal-to-noise ratio, which can effectively obtain regions robust to illumination and effectively improve the accuracy of heart rate detection under uneven illumination.

[0070] 3. This invention adapts to the faces of different subjects by generating ROIs with different aspect ratios, maximizing the removal of background noise and improving extraction accuracy. Secondly, a bounding box perturbation strategy is proposed to generate bounding boxes around selected locations to reduce tracking errors caused by head translation and rotation. Finally, for the fusion signal from multi-scale ROIs, Gaussian and uniform priors are used to assign weights to ROIs of different scales to fuse the signals selected from the multi-scale bounding boxes. The pulse signals extracted from the multi-scale bounding boxes are complementary, and fusion improves measurement accuracy, thereby effectively improving heart rate detection accuracy in motion scenarios. Attached Figure Description

[0071] Figure 1 This is a flowchart of the method of the present invention;

[0072] Figure 2 Flowchart for determining the target region I1 for resisting uneven illumination in this invention;

[0073] Figure 3 This is a flowchart illustrating the multi-scale target region M1 determination process of the present invention.

[0074] Figure 4 Flowchart for determining the target region M2 for disturbance resistance in this invention;

[0075] Figure 5 Flowchart for determining the target region set for the non-uniform illumination and motion robustness of this invention;

[0076] Figure 6 This is a flowchart of the signal fusion process based on a set of target regions according to the present invention. Detailed Implementation

[0077] In this example, a video heart rate detection method resistant to uneven lighting and motion artifacts is presented. Specifically, it is a video heart rate detection method based on target region set determination and multi-signal fusion. It focuses on the selection and processing of ROIs, choosing a more reliable set of ROI target regions to acquire higher-quality and more robust signals. Combined with a fusion strategy, it achieves robust video heart rate detection under uneven lighting and motion scenarios. This method is of great significance for the accurate and rapid detection of heart rate in real-world conditions involving complex lighting and motion, and its widespread application. Specifically, for example... Figure 1 As shown, the method is performed according to the following steps:

[0078] Step 1: As Figure 2 As shown, the target region I1 for resisting uneven illumination was obtained from the T-frame video images of the subject;

[0079] Step 1.1: Use a face detector to locate the face in the first frame of the video image. In this example, the Viola-Jones face detector is used to obtain the face rectangle detection box B0 = [x0, y0, w0, h0] in the first frame of the video image, where (x0, y0) are the coordinates of the upper left point of the face rectangle detection box B0, and w0 and h0 are the width and height of the face rectangle detection box B0, respectively.

[0080] Step 1.2: Based on the face rectangle detection box B0, use the face feature point detection algorithm to determine the coordinate information (x1) of N facial feature points in the first frame of the video image. (1) ,y1 (1) ),(x2 (1) ,y2 (1) ),(x n (1) ,y n (1) ),...,(x N (1) ,y N (1) ), where (x n (1) ,y n (1) (x1) represents the coordinate information of the nth facial feature point in the first frame of the video image. In this example, the Landmark68 face feature point detection algorithm is used, N=68, and the coordinates of the nth facial feature point (x1) in the first frame of the video image are determined. (1) ,y1 (1) ),(x2(1) ,y2 (1) ),(x n (1) ,y n (1) ),...,(x N (1) ,y N (1) Select M facial feature points to form the region of interest (ROI) of the facial skin in the first frame of the video image; M≤N; n∈[1,N].

[0081] Step 1.3: Based on the Region of Interest (ROI) 1 of the facial skin in the first frame of the video image, a feature point tracking algorithm is used to track M facial feature points in the second to Tth frames of the video image. In this example, the Kanade-Lucas-Tomasi algorithm is used for feature point tracking along the video frames. The Region of Interest (ROI) 1 of the facial skin in the i-th frame of the video image is... i The coordinate information of M facial feature points is denoted as (x1) (i) ,y1 (i) ),(x2 (i) ,y2 (i) ),...,(x m (i) ,y m (i) ),...,(x M (i) ,y M (i) ), where (x m (i) ,y m (i) ) represents the region of interest (ROI) of facial skin in the i-th frame of the video image. i The coordinate information of the m-th facial feature point, 2≤i≤T.

[0082] Step 1.4: Identify the Region of Interest (ROI) for the facial skin in the i-th frame of the video image. i The pixel space is converted from RGB color space to LAB color space, and the region of interest (ROI) of facial skin in the i-th frame of the video image is calculated. i The L-channel luminance value of each pixel is used to calculate the maximum L-channel luminance value. max The Euclidean distance of the ROI i The area is divided into Q sub-regions; in this example, Q is empirically set to 10.

[0083] Step 1.5: Calculate the average pixel value of each sub-region in the T-frame video image frame by frame, and calculate the signal-to-noise ratio of each sub-region based on the average pixel value of each sub-region as a quality index. Based on the quality index of each sub-region, select V best sub-regions from the Q sub-regions as the target region I1.

[0084] Step 2: As Figure 3 As shown, the multi-scale target region M1 is determined;

[0085] Step 2.1: Calculate the center point of the face rectangle detection box B0 Where, x c0 y c0 C represents the center point of the rectangular detection box B0 for the face. c0 The horizontal and vertical coordinates;

[0086] With C c0 Generate J bounding boxes with different aspect ratios centered on the boundary. Where j = 1, 2, ..., J, J is the total number of bounding boxes. In this example, 5 width coefficients and 5 height coefficients are combined, and J = 25 is chosen. B j Let B represent the j-th bounding box, and B j =[x j ,y j ,w j ,h j ], w j h j Represents the j-th bounding box B j The width and height are obtained from equations (1) and (2), x j y j Represents the j-th bounding box B j The coordinates of the upper left corner point are obtained from equations (3) and (4);

[0087] w j =σ w ·w0 (1)

[0088] h j =σ h ·w j (2)

[0089] In equations (1) and (2), σ w , σ h These represent coefficients for width and height, respectively; in this example, σ is set to... w =[0.7,0.75,0.80,0.85,0.90],σ h =[1.1,1.2,1.3,1.4,1.5].

[0090]

[0091]

[0092] Step 2.2: Calculate the j-th bounding box B frame by frame. j The pixel mean time series is used to obtain the j-th bounding box B. j The RGB channel time series of a T-frame video image, including: the j-th bounding box B j Time series R channel of T-frame video images [j] =[R1 [j] R2 [j] ,...,R i [j] ,...,R T [j] ], the j-th bounding box B j Time series G channels of T-frame video images [j] =[G1 [j] G2 [j] ,...,G i [j] ,...,G T [j] ] and the j-th bounding box B j Time series B channel of T-frame video image [j] =[B1 [j] B2 [j] ,...,B i [j] ,...,B T [j] ], where R i [j] Represents the j-th bounding box B j The average pixel value of the R channel of the i-th frame of the video image, G i [j] Represents the j-th bounding box B j The average pixel value of the G channel of the i-th frame of the video image, B i [j] Represents the j-th bounding box B j The average pixel value of the B channel of the i-th frame of the video image.

[0093] Step 2.3: The purpose of pulse signal extraction is to extract pulse signals from the input video. Various pulse signal extraction methods have been proposed. In this invention, the planar orthogonal skin algorithm is chosen because of its motion robustness and ease of implementation. The RGB channel time series of the T-frame video image is preprocessed into the j-th bounding box B using equation (5). j single-channel color difference signal p of a medium-T frame video image [j] Thus, J bounding boxes are obtained. Color difference signal of a medium-T frame video image [P] [1] ,P [2] ,...,P [j] ,...,P [J] ];

[0094]

[0095] In equation (5), R' [j] G' [j] B' [j] These are the j-th bounding box B j Transpose of the R-channel time series, G-channel time series, and B-channel time series of a mid-T frame video image; p1 [j] This indicates the use of the j-th bounding box B j The chroma signal defined by the projection of the G-channel and B-channel time series of a mid-T frame video image, p2 [j] This indicates the use of the j-th bounding box B j The chrominance signal defined by the projection of the R, G, and B channel time series of a mid-T frame video image, α j This represents the tuning parameters for separating the illumination and the pulse signal, and σ(·) represents the standard deviation of a given signal.

[0096] Step 2.4: Calculate the j-th bounding box B using equation (6). j Single-channel color difference signal P in a medium-T frame video image [j] Signal-to-noise ratio (SNR) j Thus, the signal-to-noise ratio set is obtained.

[0097]

[0098] In equation (6), U fs This represents the gate function; in this example, signals with frequencies in the range of [0.8, 3.5] are considered signals, while other signals outside this range are considered noise.

[0099] Step 2.5: Record the maximum value in SNR as SNR. max1 and from Select SNR max1 The corresponding bounding box B max1 As the multi-scale target region M1, B max1 The width and height are denoted as w. m1 h m1 B max The center point is denoted as C. P0 =[x p0 ,y p0 ]; where x p0y p0 Indicates center point C P0 The horizontal and vertical coordinates.

[0100] Step 3: As Figure 4 As shown, the target region M2 for disturbance resistance is determined. During tracking, the bounding box may deviate from the target due to rapid movement, rigid / non-rigid rotation, translation, etc. This can significantly negatively impact the accuracy of HR heart rate detection. A distance C is selected. P0 A new rectangular region is generated from the s neighboring pixels of the center point to reduce the impact of motion disturbance.

[0101] Step 3.1: In B max center point C P0 Determine L neighboring points around the perimeter, and use equation (7) to obtain the l-th neighboring point C. pl Thus, we obtain C pl Centered on a point with width and height w m1 h m1 The l-th rectangle BP l This results in L rectangular frames.

[0102] C pl =C p0 +p(l)·s (7)

[0103] In equation (7), p(l) represents the l-th neighboring point C. pl Relative center point C P0 The position s represents any neighboring point relative to the center point C. P0 The distance; in this example, L is set to 8 and s = 7.

[0104] Step 3.2: Calculate the frame-by-frame BP of the l-th rectangle. l The pixel mean time series is used to obtain the l-th bounding box BP. l The RGB channel time series of a T-frame video image, including: the l-th bounding box BP l Time series of the R channel of a medium-T frame video image The l-th rectangle BP l Time series of G channels of medium T-frame video images The l-th rectangle BP l Time series of B channel of T-frame video image in, BP represents the l-th rectangle. l The average pixel value of the R channel of the i-th frame of the video image. BP represents the l-th rectangle. l The average pixel value of the G channel of the i-th frame of the video image. BP represents the l-th rectangle. l The average pixel value of the B channel of the i-th frame of the video image.

[0105] Step 3.3: In this example, the planar orthogonal skin algorithm is chosen to extract the pulse signal. The l-th rectangular frame BP is obtained using equation (8). l Single-channel color difference signal of a medium-T frame video image This results in L rectangular frames. Color difference signal of a medium-T frame video image

[0106]

[0107] In equation (8), These represent the l-th rectangle BP. l Transpose of the R-channel time series, G-channel time series, and B-channel time series of a T-frame video image; This indicates the use of the l-th rectangle BP l The chroma signal defined by the projection of the G-channel and B-channel time series of a mid-T frame video image. This indicates the use of the l-th rectangle BP l The chroma signal defined by the projection of the time series of the R, G, and B channels of a mid-T frame video image. Indicates the tuning parameters, and

[0108] Step 3.4: Since the contents of the rectangular regions obtained after determining the multi-scale bounding boxes are basically the same, the SNR is no longer sufficient to distinguish the differences between them. Therefore, a maximum weight selection scheme based on mean squared error (MSE) is proposed. The BP of the l-th rectangular box is calculated using equation (9). l Single-channel color difference signal of a medium-T frame video image Root mean square error (MSE) l Thus, the root mean square error set is obtained.

[0109]

[0110] In equation (9), β represents the summation function, and α l Let α represent the l-th mean square error function. In this example, α l Specifically, this represents the mean square error between the extracted pulse signal and the reference signal. W is The signal length, in this example, is selected as W = 150 frames of video image, P ref [l] Indicates the reference pulse signal, and In this example, the average fusion is defined as the reference pulse signal P. ref [l] The estimate.

[0111] Step 3.5: Record the maximum value in MSE as MSE. max2 and from Select MSE max2 The corresponding rectangle BP max2 M2 is the target area for anti-disturbance; BP max2 The width and length are denoted as w. m2 h m2 , BP max2 The center point is denoted as C. m2 =(x m2 ,y m2 ), where x m2 y m2 Indicates center point C m2 The horizontal and vertical coordinates.

[0112] Step 4: As Figure 5 As shown, the target region set is determined;

[0113] Step 4.1: The scaling factor used when constructing a video pyramid is generally fixed. This invention modifies this limitation by introducing a scaling factor, allowing the algorithm to handle different face shapes more flexibly. Based on the image pyramid principle, using C... m2 =(x m2 ,y m2 Using a given center, generate a set of 2F+1 rectangular frames representing F layers with the same center point but different scaling scales. Where F represents the number of layers to be enlarged or reduced, and D... f Let D represent the enlarged or reduced frame of the f-th layer rectangle, and D f =[x m2 ,y m2 ,w f ,h f ], where w f h f Represents the rectangle D at the f-th level. f The width and height are obtained from equations (10) and (11):

[0114] w f =w m2 ×K f (10)

[0115] h f =h m2 ×K f (11)

[0116] In equations (10) and (11), -F ≤ f ≤ F, and K represents the scale factor controlling the ratio of the size of the target region M2; all scaled bounding boxes have the same center as the original bounding boxes. f > 0, the bounding box shrinks, and f < 0, the bounding box expands. To prevent the image size from being too small, it is recommended that K be less than and close to 1. In this embodiment, K is empirically set to 0.9.

[0117] Step 4.2: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require the full context.] By taking the intersection of each of the target regions I1 (resistance to uneven lighting) with the target region I1, we can obtain the set of target regions for both uneven lighting and motion. Among them, TargetROI f This represents the target region of the image at layer f.

[0118] Step 5: As Figure 6 As shown, based on the target region set Signal fusion and heart rate estimation;

[0119] Step 5.1: Calculate the target region (TargetROI) of the f-th layer image frame by frame. f The pixel mean time series is used to obtain the RGB channel time series of T frames of video images, including: TargetROI f Time series of the R channel of a medium-T frame video image TargetROI f Time series of G channels of medium T-frame video images TargetROI f Time series of B channel of T-frame video image in, Represents the target region (TargetROI) of the f-th layer image. f The average pixel value of the R channel of the i-th frame of the video image. Represents the target region (TargetROI) of the f-th layer image. f The average pixel value of the G channel of the i-th frame of the video image. Represents the target region (TargetROI) of the f-th layer image. f The average pixel value of the B channel of the i-th frame of the video image.

[0120] Step 5.2: Use equation (12) to obtain the target region TargetROI of the f-th layer image. f Single-channel color difference signal of the i-th frame of video image Thus, the target region TargetROI of the f-th layer image is obtained. f Single-channel color difference signal of a medium-T frame video image This leads to the target region set. Color difference signal of a medium-T frame video image

[0121]

[0122] In equation (12), These represent the target regions (TargetROI) of the f-th layer image, respectively. f Transpose the R-channel time series, G-channel time series, and B-channel time series of a T-frame video image; Indicates the use of TargetROI f The chroma signal defined by the projection of the G-channel and B-channel time series of a mid-T frame video image. Indicates the use of TargetROI f The chrominance signal defined by the time-series projection of the R, G, and B channels of a mid-T frame video image. Indicates the tuning parameters, and

[0123] Step 5.3: In this example, F is set to 2. Empirically, levels with better signal quality should have higher weights. However, we cannot determine the signal quality because we do not know the true value. We assign weights based on prior knowledge. The target region TargetROI of the f-th layer image is obtained through equations (13) and (14), respectively. f Single-channel color difference signal of a medium-T frame video image Average fusion weight λ f Gaussian fusion weight λ f (μ,σ):

[0124]

[0125]

[0126] In equation (14), μ and σ represent the mean and standard deviation of all values ​​of f = [-F, F].

[0127] Step 5.4: Calculate the fused signal

[0128] Step 5.5: Heart rate estimation;

[0129] Perform spectral analysis on the fused signal Pulse to obtain the dominant frequency f of the fused signal. max Thus, the estimated heart rate value is time × f. max Where time represents the duration of T frames of video image.

[0130] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the above-described method, and the processor is configured to execute the program stored in the memory.

[0131] In this embodiment, a computer-readable storage medium stores a computer program, which is executed by a processor to perform the steps of the above method.

[0132] To verify the robustness of the proposed heart rate detection method to uneven lighting and motion, experiments were conducted on the publicly available datasets PURE, UBFC-RPPG, and COHFACE. The heart rate was evaluated using the mean absolute error (MAE), root mean square error (RMSE), and Pierre correlation coefficient r.

[0133] Table 1. Comparison of experimental results using the PURE dataset.

[0134]

[0135] Table 1 presents the experimental results of the proposed method and the comparison method on the PURE dataset. As can be seen from Table 1, the proposed method achieves the best results on the PURE dataset, where HR... mae It is 2.45 bpm, HR rmse It achieves a performance of 1.57 bpm and an r of 0.92, even outperforming deep learning methods like PhysNet.

[0136] Table 2 Comparison of HR measurement results under different motion scenarios on the PURE dataset.

[0137]

[0138] Table 2 (Continued) Comparison of HR measurement results under different motion scenarios on the PURE dataset

[0139]

[0140] Table 2 details the experimental results of the proposed method and comparative methods in different scenarios (stable, speaking, translation, head turning) on ​​the PURE dataset. As shown in Table 2, the proposed method achieves the best results compared to traditional methods in stable and translational scenarios. In static scenarios, it even outperforms the deep learning method of PhysNet. However, in speaking and rotation scenarios, the proposed method is slightly inferior to model-based signal extraction methods (CHROM, POS). This may be due to the adaptive region generation and target region determination modules in the proposed method, where rotation and facial expression distortion during speaking cause changes in facial brightness distribution, leading to changes in brightness-based region segmentation and thus affecting detection performance.

[0141] Table 3 Comparison Experiment Results of UBFC-RPPG Dataset

[0142]

[0143]

[0144] Table 3 presents the experimental results of the proposed method and several other methods on the UBFC database. As can be seen from Table 3, the proposed method achieves the best results across all metrics compared to other methods, with HR... mae Reduced to 1.54 bpm, HR rmse The efficiency was reduced to 2.62 bpm, and the r was increased to 0.96. Furthermore, performance comparisons using motion model-based (CHROM, POS) methods show that the POS-based method of this invention outperforms the CHROM-based method, and even surpasses the cross-validation method based on deep learning's PulseGAN, effectively demonstrating the effectiveness of the non-uniform illumination and motion-robust region identified in this invention.

[0145] Table 4. Experimental results of the COHFACE dataset comparison method

[0146]

[0147] Table 4 presents the experimental results of the proposed method versus other comparative methods on the COHFACE dataset. As can be seen from Table 4, the proposed method based on target region set determination and multi-signal fusion outperforms the model-based signal extraction method (CHROM, POS) on the COHFACE dataset. The optimal HR of the proposed method combined with CHROM is [missing information]. mae It can reach 5.85 bpm, HR rmseIt achieves a performance of 7.32 bpm and an r of 0.65, and its overall performance under both natural and laboratory lighting conditions is superior to that of deep learning methods such as HR-CNN and Two-stream CNN, demonstrating certain advantages.

[0148] In summary, the experimental results on the three datasets demonstrate the robustness of the proposed method for video heart rate detection under non-uniform lighting and motion conditions. The video heart rate detection method based on target region set determination and multi-signal fusion proposed in this invention improves the accuracy and detection rate of video heart rate detection in non-uniform lighting and motion scenarios, exhibiting good robustness.

Claims

1. A video heart rate detection method based on target region set determination and multi-signal fusion, characterized in that, The procedure is as follows: Step 1: Obtain the target region I1 for resistance to uneven illumination from the region of interest of the facial skin in the subject's T-frame video image; Step 2: Based on the face detection bounding boxes of T frames of video images, generate J bounding boxes with different aspect ratios for T frames of video images, and select the bounding box with the highest signal-to-noise ratio as the multi-scale target region M1; Step 3: Based on the target region M1, generate L rectangular boxes for the T frames of video images. According to the mean square error of the color difference signal of the T frames of video images in the L rectangular boxes, select the bounding box containing the signal with the largest mean square error as the target region for anti-disturbance, that is, the motion-robust target region M2. Step 4: Based on image pyramid operations, combine the target region I1 which is resistant to uneven illumination and the target region M2 which is robust to motion to determine the set of target regions that are both resistant to uneven illumination and robust to motion. ,in, This represents the target region of the f-th layer image; Step 5: Based on the target region set The signals of the target regions in each layer of the image are extracted layer by layer and fused to obtain the main frequency of the fused signal. This allows for the estimation of heart rate. Step 5.1: Calculate the target region of the f-th layer image frame by frame. The pixel mean time series is used to obtain the RGB channel time series of T frames of video images, including: Time series of the R channel of a medium-T frame video image , Time series of G channels of medium T-frame video images , Time series of B channel of T-frame video image ,in, Represents the target region of the f-th layer image. The average pixel value of the R channel of the i-th frame of the video image. Represents the target region of the f-th layer image. The average pixel value of the G channel of the i-th frame of the video image. Represents the target region of the f-th layer image. The average pixel value of the B channel in the i-th frame of the video image; Step 5.2: Obtain the target region of the f-th layer image using equation (12). Single-channel color difference signal of the i-th frame of video image Thus, the target region of the f-th layer image is obtained. Single-channel color difference signal of a medium-T frame video image This leads to the set of target regions. Color difference signal of a medium-T frame video image : (12) In equation (12), , , These represent the target regions of the f-th layer image, respectively. Transpose the R-channel time series, G-channel time series, and B-channel time series of a T-frame video image; Indicates the use of The chroma signal defined by the projection of the G-channel and B-channel time series of a mid-T frame video image. Indicates the use of The chrominance signal defined by the time-series projection of the R, G, and B channels of a mid-T frame video image. Indicates the tuning parameters, and ; Step 5.3: Obtain the target region of the f-th layer image using equations (13) and (14) respectively. Single-channel color difference signal of a medium-T frame video image Average fusion weight Gaussian fusion weights : (13) (14) In equation (14), μ, express The mean and standard deviation of all values; Step 5.4: Calculate the fused signal ; Step 5.5: Heart rate estimation; For the fused signal Perform spectrum analysis to obtain the dominant frequency of the fused signal. Thus, the estimated heart rate value is obtained. , where time represents the duration of T frames of video image.

2. The video heart rate detection method based on target region set determination and multi-signal fusion according to claim 1, characterized in that, Step 1 includes: Step 1.1: Use a face detector to locate faces in the first frame of the video image, and obtain the face bounding boxes in the first frame of the video image. ,in, face rectangular detection box The coordinates of the top left point, , Rectangular detection boxes for faces Width and height; Step 1.2: Based on the face rectangle detection box The coordinates of N facial feature points in the first frame of the video image are determined using a facial feature point detection algorithm. ,in, This represents the coordinate information of the nth facial feature point in the first frame of the video image, and is derived from the N facial feature points in the first frame of the video image. Select M facial feature points to form the region of interest (ROI) of the facial skin in the first frame of the video image. ; ; ; Step 1.3: Based on the region of interest (ROI) of facial skin in the first frame of the video image. The feature point tracking algorithm is used to track M facial feature points in the video images from frame 2 to frame T, where the region of interest for the facial skin in the i-th video image is... The coordinate information of M facial feature points is denoted as ,in, This represents the region of interest (ROI) for facial skin in the i-th frame of the video image. The coordinate information of the m-th facial feature point. ; Step 1.4: Identify the region of interest (ROI) for the facial skin in the i-th frame of the video image. The pixel space is converted from RGB color space to LAB color space, and the region of interest for facial skin in the i-th frame of the video image is calculated. The L-channel luminance value of each pixel is used to calculate the maximum L-channel luminance value. max The Euclidean distance will Divide into Q-block sub-regions; Step 1.5: Calculate the average pixel value of each sub-region in the T-frame video image frame by frame, and calculate the signal-to-noise ratio of each sub-region based on the average pixel value of each sub-region as a quality index. Based on the quality index of each sub-region, select V best sub-regions from the Q sub-regions as the target region I1.

3. The video heart rate detection method based on target region set determination and multi-signal fusion according to claim 2, characterized in that, Step 2 includes: Step 2.1: Calculate the face rectangle detection box center point ,in, , Represents a rectangular detection box for faces center point The horizontal and vertical coordinates; by Generate J bounding boxes with different aspect ratios centered on the boundary. Where j = 1, 2, ..., J, and J is the total number of bounding boxes. Let j be the bounding box, and , , Represents the j-th bounding box The width and height are obtained from equations (1) and (2). , Represents the j-th bounding box The coordinates of the upper left corner point are obtained from equations (3) and (4); (1) (2) In equations (1) and (2), , These represent coefficients for width and height, respectively. (3) (4) Step 2.2: Calculate the j-th bounding box frame by frame. The pixel mean time series is used to obtain the j-th bounding box. The RGB channel time series of a T-frame video image, including: the j-th bounding box. Time series of the R channel of a medium-T frame video image The j-th bounding box Time series of G channels of medium T-frame video images and the j-th bounding box Time series of B channel of T-frame video image ,in, Represents the j-th bounding box The average pixel value of the R channel of the i-th frame of the video image. Represents the j-th bounding box The average pixel value of the G channel of the i-th frame of the video image. Represents the j-th bounding box The average pixel value of the B channel in the i-th frame of the video image; Step 2.3: Use equation (5) to preprocess the RGB channel time series of the T-frame video image into the j-th bounding box. Single-channel color difference signal of a medium-T frame video image Thus, J bounding boxes are obtained. Color difference signal of a medium-T frame video image ; (5) In equation (5), , , These are the j-th bounding boxes. Transpose of the R-channel time series, G-channel time series, and B-channel time series of a T-frame video image; This indicates the use of the j-th bounding box. The chroma signal defined by the projection of the G-channel and B-channel time series of a mid-T frame video image. This indicates the use of the j-th bounding box. The chroma signal defined by the projection of the time series of the R, G, and B channels of a mid-T frame video image. This represents the tuning parameters for separating the illumination and the pulse signal, and ; This represents the standard deviation of a given signal; Step 2.4: Calculate the j-th bounding box using equation (6) Single-channel color difference signal in a medium-T frame video image signal-to-noise ratio Thus, the signal-to-noise ratio set is obtained. : (6) In equation (6), Represents a gate function; Step 2.5: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require The maximum value in is denoted as and from Select The corresponding bounding box As a multi-scale target region M1, The width and height are denoted as follows: , ,Will The center point is denoted as ;in, , Indicates the center point The horizontal and vertical coordinates.

4. The video heart rate detection method based on target region set determination and multi-signal fusion according to claim 3, characterized in that, Step 3 includes: Step 3.1: In center point Determine L neighboring points around the perimeter, and use equation (7) to obtain the l-th neighboring point. Thus, to obtain With the center point as the center, the width and height are... , The l-th rectangle This results in L rectangular frames. : (7) In equation (7), Represents the l-th nearest neighbor. relative center point The position 's' represents any neighboring point relative to the center point. The distance; Step 3.2: Calculate the l-th bounding box frame by frame. The pixel mean time series is used to obtain the l-th rectangle. The RGB channel time series of a T-frame video image, including: the l-th bounding box Time series of the R channel of a medium-T frame video image The l-th rectangle Time series of G channels of medium T-frame video images The l-th rectangle Time series of B channel of T-frame video image ,in, Represents the l-th rectangle The average pixel value of the R channel of the i-th frame of the video image. Represents the l-th rectangle The average pixel value of the G channel of the i-th frame of the video image. Represents the l-th rectangle The average pixel value of the B channel in the i-th frame of the video image; Step 3.3: Use equation (8) to obtain the l-th rectangle. Single-channel color difference signal of a medium-T frame video image This results in L rectangular frames. Color difference signal of a medium-T frame video image : (8) In equation (8), , , These represent the l-th rectangle. Transpose of the R-channel time series, G-channel time series, and B-channel time series of a T-frame video image; This indicates the use of the l-th rectangle. The chroma signal defined by the projection of the G-channel and B-channel time series of a mid-T frame video image. This indicates the use of the l-th rectangle. The chroma signal defined by the projection of the time series of the R, G, and B channels of a mid-T frame video image. Indicates the tuning parameters, and ; Step 3.4: Calculate the l-th rectangle using equation (9) Single-channel color difference signal of a medium-T frame video image root mean square error Thus, the root mean square error set is obtained. : (9) In equation (9), Denotes the summation function, and , Let represent the l-th mean square error function, and W is The signal length, Indicates the reference pulse signal, and ; Step 3.5: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require the full context.] The maximum value in is denoted as and from Select The corresponding rectangle M2 is designated as the target area for disturbance mitigation; The width and length are denoted as follows: , ,Will The center point is denoted as ,in, , Indicates the center point The horizontal and vertical coordinates.

5. The video heart rate detection method based on target region set determination and multi-signal fusion according to claim 4, characterized in that, Step 4 includes: Step 4.1: with Centered on the same point, the total number of F layers generated with different scaling ratios is: collection of rectangles Where F represents the number of layers to be enlarged or reduced. This represents the enlarged or reduced f-th layer rectangle, and ,in, , Represents the rectangle of the f-th layer The width and height are obtained from equations (10) and (11): (10) (11) In equations (10) and (11), K represents the scale factor that controls the ratio of the size of the target region M2; Step 4.2: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require the full context.] By taking the intersection of each of the target regions I1 (resistance to uneven lighting) with the target region I1, we can obtain the set of target regions for both uneven lighting and motion. ,in, This represents the target region of the image at layer f.

6. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports the processor in executing any of the video heart rate detection methods of claims 1-5, and the processor is configured to execute the program stored in the memory.

7. A computer-readable storage medium storing a computer program thereon, characterized in that, The computer program, when run by a processor, performs the steps of the video heart rate detection method according to any one of claims 1-5.