Non-contact heart rate estimation method and system based on multi-modal signal fusion
Through multimodal signal fusion and dynamic parameter adjustment, combined with image sharpening and multi-color spatial signal processing, the problems of noise interference and environmental changes in contactless heart rate estimation are solved, and high-precision, stability and real-time heart rate monitoring is achieved.
Patent Information
- Application Number
- CN202510668302.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-23
AI Technical Summary
The prior art faces the challenges of noise interference, motion artifacts and lighting changes in contactless heart rate estimation, resulting in limited accuracy of heart rate estimation, and traditional algorithms are difficult to dynamically adapt to complex environments, while deep learning methods sacrifice real-time and computing efficiency.
A method based on multimodal signal fusion is adopted, combining image sharpening, multi-color spatial signal fusion, background brightness compensation, color change enhancement, adaptive normalization and bandpass filtering technology, as well as time-frequency domain joint estimation and historical data smoothing strategies, to achieve accurate extraction and stability improvement of heart rate signals.
It improves the accuracy and stability of color change information under complex lighting conditions, realizes effective purification and accurate analysis of heart rate signals, ensures the stability and reliability of heart rate estimation results, and can provide continuous and accurate heart rate monitoring during exercise or environmental changes.
Smart Images

Figure CN120183005A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a non-contact heart rate estimation method and system based on multi-modal signal fusion. Background Art
[0002] In spatial domain enhancement, directly operating on pixels is the most intuitive way. Through point processing methods such as contrast stretching or gamma correction, the gray distribution can be non-linearly adjusted. For example, logarithmic transformation can highlight dark details, while histogram equalization improves the global contrast by redistributing pixel gray values. However, global equalization may lead to local overexposure or noise amplification, so adaptive histogram equalization (CLAHE) is proposed to avoid distortion by restricting the contrast change in local areas. In addition, spatial filtering techniques such as Gaussian smoothing and median filtering can effectively suppress noise but may blur edges; while sharpening filters (such as Laplacian operator) highlight textures by enhancing high-frequency components, but need to be carefully processed to avoid noise interference.
[0003] Frequency domain enhancement is based on the Fourier transform to convert the image into the frequency domain and achieve effects by adjusting different frequency components. Low-pass filters retain low-frequency information to smooth the image and are often used for denoising; high-pass filters enhance high-frequency edge information, but noise control needs attention; band-pass or band-stop filtering is for specific frequency ranges, such as eliminating periodic noise. The advantage of this method is that it can separate the structural and random components in the image, but the computational cost of frequency domain conversion is high, and the error between the frequency domain and the spatial domain conversion needs to be processed; for color images, color space conversion is a key step in enhancement. Independently adjusting each channel in the RGB space may lead to color distortion, so the image is often converted to the HSV or HSL space, and only the brightness (V or L component) is adjusted while keeping the hue and saturation unchanged. The Retinex theory further simulates the adaptability of the human eye to light and improves the problem of uneven illumination by separating the illumination and reflection components.
[0004] Remote photoplethysmography (rPPG), as a non-contact physiological signal detection technology, estimates the heart rate by analyzing the weak fluctuations of the facial skin color with the change of blood volume in the video. Its core principle is similar to that of traditional contact photoplethysmography (PPG), but it does not need to be in direct contact with the skin, so it has broad application prospects in the fields of health monitoring, emotion recognition, etc. However, the existing technologies face significant challenges of noise interference, motion artifacts and illumination changes in practical applications, resulting in limited heart rate estimation accuracy.
[0005] Traditional rPPG algorithms mainly rely on signal processing techniques. For example, the GREEN algorithm extracts the pulse wave through the green channel signal because it is related to the absorption characteristics of hemoglobin; the POS algorithm uses linear projection to separate the noise components in the skin reflection signal; the CHROM algorithm suppresses motion interference by constructing chromaticity signals. However, these methods generally rely on fixed-bandwidth filtering and a single color space (such as RGB or the green channel), making it difficult to dynamically adapt to complex environments. In addition, traditional algorithms are highly sensitive to skin color differences. Individuals with dark skin are vulnerable to ambient light interference due to the small signal amplitude, resulting in insufficient robustness. The application of image sharpening technology in remote photoplethysmography (rPPG) mainly aims to highlight the weak physiological signals in the video. Its core principle is to use spatial domain convolution operations to enhance high-frequency details, thereby highlighting the periodic changes in the skin surface color caused by cardiac pulsation. These changes stem from the periodic fluctuations in blood volume. When the heart contracts, the subcutaneous capillaries dilate, and the absorption of light of a specific wavelength by the blood increases, leading to a decrease in the intensity of the reflected light. The opposite occurs during heart relaxation. This optical phenomenon appears as tiny periodic fluctuations in pixel values in the video, but its amplitude is extremely small and is easily masked by ambient light interference, motion artifacts, and sensor noise. The sharpening operation realizes signal enhancement through the design of a specific convolution kernel. A typical implementation uses the Laplacian kernel, which is essentially a high-pass filter that amplifies the high-frequency components in the image by calculating the differences between local pixels. For example, a 3×3 convolution kernel enhances the local contrast by assigning a higher positive weight to the central pixel and negative weights to the surrounding pixels. This operation is mathematically equivalent to performing a second-order differential on the image, which can significantly enhance high-frequency details such as edges and textures, and these details often contain weak color changes related to blood flow. Smaller convolution kernels focus on the microscopic pixel neighborhood and can capture the tiny color changes caused by blood flow more precisely. The fluctuations in skin reflected light caused by cardiac pulsation usually only affect local pixel groups. Large-size convolution kernels are prone to introducing signal mixing in irrelevant regions and blurring the effective physiological features. Smaller weights reduce the non-linear distortion of the original signal, ensuring that the fundamental frequency and harmonic components of the heart rate signal are closer to the true physiological characteristics. In addition, the local pixel offsets caused by small head movements are significantly amplified by large-weight kernels, while the gentle enhancement of small-weight kernels can reduce the impact of such interference.
[0006] In recent years, deep learning-based methods have gradually become a research hotspot. For example, ST-ASENet enhances facial feature extraction capabilities through spatiotemporal attention mechanisms, and rPPG-MAE uses masked autoencoders to mine the self-similarity of physiological signals. This type of method automatically learns spatiotemporal features through neural networks, and exhibits higher accuracy under complex backgrounds and low-light conditions. However, deep learning models rely heavily on large-scale labeled data training and have high computing resource requirements, making them difficult to deploy in real-time monitoring devices. In addition, the model has limited generalization capabilities for extreme scenarios such as sudden changes in lighting and intense exercise, and lacks interpretability, which limits its practical application in the field of medical health. For example, although the Transformer-based fusion framework can improve the ability to model long-distance features, it has high computing latency and cannot meet real-time requirements.
[0007] The core problem of existing technologies is that traditional algorithms are difficult to adapt to dynamic environmental changes due to fixed parameter design, and deep learning methods improve accuracy but sacrifice real-time and computational efficiency. In addition, most solutions do not fully integrate multimodal signals or dynamically adjust parameters, resulting in insufficient stability in signal extraction. For example, a single color space is easily affected by the wavelength of light, and a fixed bandwidth filter cannot track dynamic changes in heart rate, especially during exercise or when the physiological state fluctuates, where errors accumulate significantly. Therefore, an optimization solution that takes into account accuracy, robustness, and real-time performance is highly needed to improve the performance of heart rate estimation in complex environments through multimodal signal fusion and dynamic parameter adjustment mechanisms. Summary of the invention
[0008] In order to solve the above problems, the present invention proposes a non-contact heart rate estimation method and system based on multimodal signal fusion, which realizes more accurate and stable heart rate detection of facial videos by combining image sharpening, multi-color space signal fusion, background brightness compensation, color change enhancement, adaptive normalization and bandpass filtering technology, as well as time-frequency domain joint estimation and historical data smoothing strategy.
[0009] The specific plan is as follows:
[0010] On the one hand, a non-contact heart rate estimation method based on multimodal signal fusion includes:
[0011] S1, obtain sample face video;
[0012] S2, using image sharpening technology to highlight the slight color changes in the sample face video and obtain a sharpened face video image;
[0013] S3, locates the region of interest based on facial feature points in the sharpened face video;
[0014] S4. Extract the signals in the HSV and Lab color spaces from the region of interest, fuse the signals in these two color spaces, and obtain a mixed signal containing the information of both color spaces;
[0015] S5. Adopt a background brightness compensation strategy for the mixed signal to suppress the ambient light interference and obtain an original brightness compensation signal;
[0016] S6. Perform median filtering on the original brightness compensation signal to obtain the brightness compensation signal after median filtering; subtract the brightness compensation signal after median filtering from the original brightness compensation signal to obtain a color transformation signal;
[0017] S7. Use the maximum and minimum values of the color transformation signal as the range for signal normalization to normalize the color transformation signal and obtain a normalized signal;
[0018] S8. Adjust the bandwidth of the original band-pass filter based on the historical heart rate, and perform band-pass filtering on the normalized signal through the adjusted band-pass filter to obtain an rPPG signal;
[0019] S9. Perform time-domain analysis and frequency-domain analysis on the rPPG signal, and weight the heart rate values estimated by time-domain analysis and the heart rate values estimated by frequency-domain analysis to obtain a heart rate estimation result;
[0020] S10. Correct the heart rate estimation result through a historical data smoothing strategy to obtain a final heart rate estimation value.
[0021] Further, the S4 specifically includes:
[0022] Extract color signals from the HSV color space and the Lab color space to obtain color signal components H, S, V, L, a, and b. Weight the color signal components H, S, and V to obtain a mixed signal in the HSV color space; weight the color signal components L, a, and b to obtain a mixed signal in the Lab color space.
[0023] Further, the S5 specifically includes: Extract the brightness signal outside the region of interest as a reference benchmark signal, subtract the reference benchmark signal from the mixed signals in the HSV color space and the Lab color space to obtain brightness compensation, and perform a weighted addition operation on the mixed signals in the HSV color space and the Lab color space after brightness compensation to obtain an original brightness compensation signal.
[0024] Further, in S8, the adjustment of the bandwidth of the original band-pass filter based on the historical heart rate specifically includes:
[0025] Obtain the bandwidth range of the preset original band-pass filter, obtain the average heart rate within a specified period based on the original band-pass filter, convert the average heart rate into a heart rate frequency, adaptively adjust the bandwidth range of the original band-pass filter based on the heart rate frequency, and obtain the adjusted band-pass filter. The calculation formula is as follows:
[0026] ;
[0027] ;
[0028] Among them, represents the band-pass range on the left side of the reference frequency of the adjusted band-pass filter; represents the band-pass range on the right side of the reference frequency of the adjusted band-pass filter; represents the heart rate frequency corresponding to the average heart rate in the specified period; and represent the range of the adjusted band-pass filter, which is used to filter out the part with frequency components outside the range of , .
[0029] Furthermore, the S9 specifically includes: using FFT to extract the frequency-domain information in the rPPG signal, calculating the frequency component with the largest amplitude in the frequency spectrum based on the frequency-domain information to obtain the heart rate value estimated by frequency-domain analysis; calculating the heart rate value estimated by time-domain analysis based on the peak interval in the rPPG signal; combining the heart rate value estimated by time-domain analysis and the heart rate value estimated by frequency-domain analysis through weighted averaging to obtain the heart rate estimation result. The calculation formula of the heart rate estimation result is as follows:
[0030] ;
[0031] Among them, represents the heart rate value estimated by frequency-domain analysis; represents the heart rate value estimated by time-domain analysis; represents the heart rate estimation result; is the weighting coefficient, and the value range is (0, 1), which is used to control the weights of frequency-domain and time-domain information in heart rate estimation.
[0032] Furthermore, the S10 specifically includes: calculating the average heart rate of the previous five times based on the heart rate values of the last five times. The calculation formula is as follows:
[0033] ;
[0034] Among them, , , , and are the heart rate values of the previous five times respectively; Indicates the average heart rate of the first five times;
[0035] Take the heart rate estimation result and the average of the first five times Calculate the difference. If the difference exceeds the preset threshold , that is:
[0036] ;
[0037] It indicates that there is an abnormal fluctuation between the current heart rate estimation result and the historical estimation result. Perform weighted fusion on the current heart rate estimation result and the average of the first five times to eliminate the influence of the abnormal fluctuation. The weighted fusion formula is as follows:
[0038] ;
[0039] Among them, represents the weight of the current heart rate estimation result, and , which is used to ensure that the current heart rate estimation result occupies a greater proportion during fusion; represents the final heart rate estimation value.
[0040] On the other hand, a non-contact heart rate estimation system based on multi-modal signal fusion includes:
[0041] A face video acquisition module for acquiring a sample face video;
[0042] A sharpening module for highlighting the weak color changes in the sample face video through image sharpening technology to obtain a sharpened face video image;
[0043] A region of interest (ROI) localization module for localizing the region of interest based on the facial feature points in the sharpened face video;
[0044] A mixed signal acquisition module for extracting signals in the HSV and Lab color spaces from the region of interest, fusing the signals in these two color spaces to obtain a mixed signal containing information from both color spaces;
[0045] An original brightness compensation signal acquisition module for suppressing ambient light interference on the mixed signal using a background brightness compensation strategy to obtain an original brightness compensation signal;
[0046] A color transformation signal acquisition module for performing median filtering on the original brightness compensation signal to obtain a median-filtered brightness compensation signal; subtracting the median-filtered brightness compensation signal from the original brightness compensation signal to obtain a color transformation signal;
[0047] A normalization signal acquisition module for normalizing the color transformation signal using the maximum and minimum values of the color transformation signal as the range for signal normalization to obtain a normalized signal;
[0048] The rPPG signal acquisition module is used to adjust the bandwidth of the original band-pass filter based on the historical heart rate, perform band-pass filtering on the normalized signal through the adjusted band-pass filter, and obtain the rPPG signal;
[0049] The heart rate estimation module is used to perform time-domain analysis and frequency-domain analysis on the rPPG signal, and weight the heart rate values estimated by the time-domain analysis and the frequency-domain analysis to obtain the heart rate estimation result;
[0050] The final heart rate estimation value acquisition module corrects the heart rate estimation result through the historical data smoothing strategy to obtain the final heart rate estimation value.
[0051] The present invention adopts the above technical solutions and has the following beneficial effects:
[0052] (1) By using image sharpening technology and extracting and fusing signals from the HSV and Lab color spaces, the present invention highlights the weak color changes in the face video, and adopts a background brightness compensation strategy to suppress environmental light interference, thereby improving the accuracy and stability of color change information under complex lighting conditions;
[0053] (2) By subtracting the signal after median filtering from the signal without median filtering to highlight the color change information, combining the normalization processing based on the maximum and minimum values of the signal and the band-pass filtering technology that dynamically adjusts the bandwidth according to the historical heart rate, the present invention realizes the effective purification and accurate analysis of the heart rate signal;
[0054] (3) By correcting the estimated heart rate value through the historical data smoothing strategy, the present invention effectively reduces the influence of accidental errors, ensures the stability and reliability of the finally output heart rate estimation result, and enables continuous and accurate heart rate monitoring even in the presence of interference factors such as movement or environmental changes. Description of the Drawings
[0055] Figure 1 It is a flowchart of a non-contact heart rate estimation method based on multi-modal signal fusion according to an embodiment of the present invention;
[0056] Figure 2 It is a schematic diagram of the module for detecting 468 key points of the face according to an embodiment of the present invention;
[0057] Figure 3 It is a system diagram of a non-contact heart rate estimation based on multi-modal signal fusion according to an embodiment of the present invention. Detailed Embodiments
[0058] The present invention will be further described in detail below in conjunction with the embodiments and the drawings, but the embodiments of the present invention are not limited thereto.
[0059] AsFigure 1 As shown in Figure 1 , the non-contact heart rate estimation method based on multi-modal signal fusion of the present invention includes:
[0060] S1. Obtain a sample face video.
[0061] Specifically, in this embodiment, facial images are extracted frame by frame from the input video, and the entire video is divided into frames of corresponding lengths. The entire video image can be regarded as consisting of N frames, and subsequent processing is also performed based on each frame image.
[0062] S2. Highlight the weak color changes in the sample face video through image sharpening technology to obtain a sharpened face video image.
[0063] Specifically, in this embodiment, each frame image is sharpened through image sharpening technology to highlight the weak color changes in the facial area. During the image sharpening process, the convolution kernel exists in the form of a matrix, and it performs convolution operations with each pixel in the image and its surrounding neighborhood to change the values of the pixels in the image and achieve the sharpening effect. The formula for the convolution operation is as follows:
[0064] ;
[0065] where I(x,y) is the pixel value of the input image, K(i,j) is the weight in the convolution kernel, and I’(x,y) is the pixel value of the output image after convolution. The sharpening adopts a 3×3 convolution kernel design with a higher weight in the center and negative weights around it, using the convolution operation to highlight the high-frequency components of the image and making the skin color fluctuations caused by blood volume changes more significant.
[0066] S3. Locate the region of interest based on the facial feature points in the sharpened face video.
[0067] Specifically, as shown in Figure 2 Figure 2As shown, in this embodiment, the Mediapipe Face Mesh module is used to detect 468 3D key points of the face in real time. By setting two groups of key points to correspond to the cheek regions respectively, the algorithm can accurately locate the cheek regions. The reason for choosing the cheek regions is that they are relatively flat, less affected by head movement, and have dense subcutaneous capillaries and strong blood flow signals. Two independent ROI regions (left cheek and right cheek) are delimited in this way to avoid background noise interference; Face Mesh is a part of the Mediapipe framework, aiming to provide high-precision face key point positioning for each face image and support 2D and 3D face analysis. The input of this model is an image or video stream, and the output is the coordinates of 468 face key points. Each key point has a position in 3D space, and the depth information of this face region can be extracted. Face Mesh detects the key points in the face image through a deep learning model, accurately locates the face feature regions such as eyes, eyebrows, nose, mouth, and face contour. The algorithm selects some feature points on the left and right cheeks respectively, uses these feature points as the boundaries of the ROI regions, and takes the regions surrounded by these feature points as the ROI regions required by the algorithm, and then extracts the signals from these two ROI regions.
[0068] S4. Extract the signals in the HSV and Lab color spaces from the region of interest, and fuse the signals in these two color spaces to obtain a mixed signal containing the information of the two color spaces.
[0069] Specifically, the S4 specifically includes:
[0070] Extract the color signals in the HSV color space and the Lab color space to obtain the color signal components H, S, V, L, a, and b. Weight the color signal components H, S, and V to obtain the mixed signal in the HSV color space; weight the color signal components L, a, and b to obtain the mixed signal in the Lab color space.
[0071] Specifically, extract the HSV color space and the Lab color space for color signals from the two ROI regions respectively, extract each different color component in these two color spaces one by one and then perform weighted summation to obtain the mixed signal. The mixed signal contains the color change information of these two color spaces, that is, contains more features, which helps to extract a signal closer to the real pulse signal.
[0072] Specifically, in this embodiment, the HSV color space and the Lab color space are selected for color signal extraction. The color signals in these two color spaces are extracted one by one and then weighted and added to obtain the mixed signals of their respective color spaces. The mixed signals contain the color change information of these two color spaces. If the mixed signals of the two color spaces are weighted and fused, that is, more features are included, which helps to extract a signal closer to the true pulse signal. Different from the commonly used RGB color space, the HSV color space is closer to the way humans perceive colors. It describes colors through three components, namely hue, saturation, and value. These three components are the extraction targets in the HSV color space. Since the default color space of OpenCV is BGR, each frame of the image needs to be converted to the HSV color space before extracting these three components. Given a pixel point in the BGR color space, B, G, and R are the pixel values of the blue, green, and red channels respectively. Before conversion, it is first necessary to normalize the BGR values, and the values of each channel of BGR are normalized to the range of [0, 1]. This step is for the convenience of subsequent calculations:
[0073] ;
[0074] Then, for each pixel, calculate the maximum value Cmax and the minimum value Cmin:
[0075] ;
[0076] Next, calculate the color difference (i.e., the difference between the maximum value and the minimum value):
[0077] ;
[0078] After that, the calculation of the three components can be carried out. The hue H is calculated based on the color difference of the pixel. There are different formulas for calculating the hue, specifically depending on the color channel where the maximum value Cmax is located. :
[0079] ;
[0080] :
[0081] ;
[0082] :
[0083] ;
[0084] Saturation describes the purity of the color. The calculation formula for saturation is:
[0085] ;
[0086] If , it means that the pixel is gray (no color). At this time . The lightness is the maximum value , that is, the brightness value of the pixel:
[0087] ;
[0088] The Lab color space (CIE 1976 (L*, a*, b*) color space) is a color space defined by the International Commission on Illumination (CIE) in 1976 to provide a device-independent color representation and to simulate the human eye's perception of color as closely as possible. The Lab color space also consists of three components, where L* represents the lightness of the color, ranging from 0 to 100, 0 representing black and 100 representing white, and it is the component used to describe the light and dark degree of the color. a* (red-green axis) represents the change of the color between green and red, negative values representing green and positive values representing red, with a range of approximately -128 to 127. b* (blue-yellow axis) represents the change of the color between blue and yellow, negative values representing blue and positive values representing yellow, with a range of approximately -128 to 127.
[0089] To convert an image from the BGR color space to the Lab color space, it is necessary to first convert from the BGR color space to the XYZ color space, and then from the XYZ color space to the Lab color space. The converted X, Y, Z are the components of the XYZ color space. Next, each pixel in the XYZ color space is converted to the Lab color space. First, the XYZ values need to be normalized. Usually, the D65 white point is used as a reference:
[0090] ;
[0091] Normalize each pixel using the reference value:
[0092] ;
[0093] Next, the L*, a*, and b* values can be calculated using the following formulas:
[0094] First, calculate L* (lightness):
[0095] ;
[0096] Among them, is a piecewise function:
[0097] ;
[0098] Next, calculate the values of a* and b* respectively:
[0099] ;
[0100] ;
[0101] The different components of the two color spaces are weighted and added together according to certain weights to obtain the mixed signals of their respective color spaces.
[0102] S5. Adopt a background brightness compensation strategy for the mixed signal to suppress ambient light interference and obtain the original brightness compensation signal.
[0103] Specifically, the S5 specifically includes: extracting the brightness signal outside the region of interest as the reference benchmark signal, subtracting the reference benchmark signal from the mixed signals of the HSV color space and the Lab color space to obtain brightness compensation, and performing a weighted addition operation on the mixed signals of the HSV color space and the Lab color space after brightness compensation to obtain the original brightness compensation signal.
[0104] Specifically, to suppress the influence of ambient light, the algorithm extracts the background brightness outside the ROI region (V of HSV and L of Lab), subtracts it from the mixed signals of the corresponding color spaces to achieve brightness compensation. Finally, the compensated signals of the two color spaces are fused according to weights to generate a mixed signal containing multi-dimensional features, enhancing the robustness to light changes.
[0105] S6. Perform median filtering on the original brightness compensation signal to obtain the brightness compensation signal after median filtering; subtract the brightness compensation signal after median filtering from the original brightness compensation signal to obtain the color transformation signal.
[0106] Specifically, remove the baseline drift through median filtering. Perform median filtering on the mixed signal (window length 300 frames), and subtract the filtering result from the original signal to obtain the detrended signal. This step can highlight the periodic characteristics of blood volume changes and suppress low-frequency noise at the same time.
[0107] S7. Use the maximum and minimum values of the color transformation signal as the range for signal normalization to normalize the color transformation signal and obtain the normalized signal.
[0108] Specifically, the mixed signal needs to be normalized to eliminate the influence of individual skin color differences. The traditional method uses global normalization, but the signal dynamic range of individuals with dark skin may be compressed. Therefore, the algorithm proposes local adaptive normalization: the signal is divided into sliding windows of 300 frames, the maximum and minimum values are independently calculated for each window, and the signal is linearly mapped to the range of [0,1] accordingly. This method ensures that the signal changes of individuals with different skin colors can be fully developed within their own dynamic ranges.
[0109] S8. Adjust the bandwidth of the original band-pass filter based on the historical heart rate, and perform band-pass filtering on the normalized signal through the adjusted band-pass filter to obtain the rPPG signal.
[0110] Specifically, in S8, the adjustment of the bandwidth of the original band-pass filter based on the historical heart rate specifically includes:
[0111] Obtain the preset bandwidth range of the original band-pass filter, obtain the average heart rate within a specified period based on the original band-pass filter, convert the average heart rate into a heart rate frequency, and adaptively adjust the bandwidth range of the original band-pass filter based on the heart rate frequency to obtain the adjusted band-pass filter. The calculation formula is as follows:
[0112] ;
[0113] ;
[0114] Among them, represents the band-pass range on the left side of the reference frequency of the adjusted band-pass filter; represents the band-pass range on the right side of the reference frequency of the adjusted band-pass filter; represents the heart rate frequency corresponding to the average heart rate within the specified period; and represent the range of the adjusted band-pass filter, which is used to filter out the parts with frequency components outside the range of , .
[0115] Specifically, in this embodiment, the bandwidth of the original band - pass filter is set to [0.7, 3] Hz corresponding to the normal heart rate range of [42, 180]. After the signal is normalized, an adaptive band - pass filter is applied to the signal. Traditional rPPG algorithms usually use a band - pass filter with a fixed bandwidth (such as 0.7, 3 Hz) to filter out non - heart - rate - related noise. However, in dynamic scenarios (such as exercise or sudden changes in lighting), a fixed bandwidth is difficult to adapt to the changing range of the heart rate and the frequency band shift of interference signals. Therefore, the present invention proposes an adaptive band - pass filtering strategy based on historical heart rate estimates. The core of this strategy lies in dynamically adjusting the upper and lower limit frequencies of the filter to balance the conflict between noise suppression and effective signal retention. The application of the adaptive band - pass filter in heart rate signal extraction not only effectively avoids sudden noise interference and signal mutations caused by head movement or facial expressions, but also ensures that the heart rate signal remains stable within the normal heart rate range, thus significantly improving the accuracy and stability of signal extraction. When the heart rate is in a stable range, the bandwidth of the band - pass filter will automatically shrink, more precisely focusing on the characteristic frequency band of the heart rate signal. This dynamic adjustment mechanism enables the filter to flexibly adapt to different motion states and physiological changes, making the extracted heart rate signal smoother and avoiding the loss of effective information due to over - filtering or too narrow a bandwidth. Through this adaptive band - pass filter method, not only is the interference from external factors such as the environment, facial expressions, and head movement effectively reduced, but also the accurate extraction of heart rate signals under different physiological states is ensured, making the extraction of heart rate signals more in line with the actual physiological change law rather than following a fixed filtering standard.
[0116] S9. Conduct time - domain analysis and frequency - domain analysis on the rPPG signal, and weight the heart rate values estimated by time - domain analysis and frequency - domain analysis to obtain the heart rate estimation result.
[0117] Specifically, the S9 specifically includes: using FFT to extract the frequency - domain information in the rPPG signal, calculating the frequency component with the largest amplitude in the spectrum based on the frequency - domain information to obtain the heart rate value estimated by frequency - domain analysis; calculating the heart rate value estimated by time - domain analysis based on the peak intervals in the rPPG signal; combining the heart rate value estimated by time - domain analysis and the heart rate value estimated by frequency - domain analysis through weighted averaging to obtain the heart rate estimation result. The formula for the heart rate estimation result is as follows:
[0118] ;
[0119] Where represents the heart rate value estimated by frequency - domain analysis; represents the heart rate value estimated by time - domain analysis; represents the heart rate estimation result; is the weighting coefficient, and its value range is (0, 1), which is used to control the weights of frequency - domain and time - domain information in heart rate estimation.
[0120] Specifically, traditional heart rate estimation methods usually perform a fast Fourier transform (FFT) on the signal, then find the frequency component with the largest amplitude in the frequency spectrum, and multiply it by 60 to convert it to the number of heartbeats per minute (bpm). The core assumption of this method is that the spectral peak of the signal corresponds to the main component of the heart rate, so the heart rate can be estimated through the main frequency in the frequency spectrum. In the time domain, the fluctuations of the heart rate signal are periodic, and each heartbeat generates a peak. The interval between two adjacent peaks in the heart rate signal reflects the time length of a cycle, and this interval has a direct relationship with the heart rate. The heart rate can be directly estimated by measuring the peak interval in the signal. This method does not rely on frequency domain information but directly uses the time domain characteristics of the signal to estimate the heart rate. Although this method is simple and intuitive, it may be affected by noise and motion artifacts, resulting in inaccurate peak identification and thus affecting the accuracy of heart rate estimation. To overcome the limitations of relying solely on time domain or frequency domain information, the present invention proposes a weighted method that combines the two. First, the FFT is used to extract the frequency domain information in the signal, and the frequency component with the largest amplitude in the frequency spectrum is calculated to obtain a preliminary heart rate estimate (in bpm). Then, the peak interval in the signal is further analyzed to calculate the heart rate estimate in the time domain. After obtaining these two estimation results, they are combined by weighted averaging to obtain a more accurate heart rate estimate.
[0121] Specifically, since relying solely on frequency domain or time domain analysis has limitations: frequency domain methods (such as FFT main peak detection) are vulnerable to periodic noise interference, while time domain methods (such as peak interval calculation) are sensitive to instantaneous noise (such as facial micro-movements). Therefore, the advantages of both are combined to achieve more robust heart rate estimation. In frequency domain analysis, the algorithm performs a fast Fourier transform (FFT) on the filtered signal and extracts the frequency component with the largest amplitude in the power spectrum , and converts it to a heart rate value:
[0122] ;
[0123] In time domain analysis, the algorithm detects the peak-to-peak interval (PPI) of the signal, eliminates outliers (such as PPI < 0.3 s or PPI > 1.5 s), and calculates the median interval to obtain the heart rate in the time domain:
[0124] ;
[0125] To balance the contributions of both, the algorithm introduces a dynamic weight coefficient α. When the amplitude of the main peak in the frequency domain is significantly higher than that of the secondary peak (e.g., the amplitude of the main peak > 2 times the amplitude of the secondary peak), α is set to a larger value to give priority to trusting the frequency domain result. When the stability of the time domain peak interval is high (e.g., the standard deviation of PPI < 0.1 s), α is set to a smaller value to enhance the time domain contribution. The final estimated heart rate after weighting can be expressed by the following formula:
[0126] .
[0127] S10, correct the heart rate estimation result through the historical data smoothing strategy to obtain the final heart rate estimation value.
[0128] Specifically, the S10 specifically includes: calculating the average heart rate of the previous five times based on the heart rate values of the last five times, and the calculation formula is as follows:
[0129] ;
[0130] Among them, , , , and are the heart rate values of the previous five times respectively; represents the average heart rate of the previous five times;
[0131] Subtract the previous five - time average from the heart rate estimation result to calculate the difference. If the difference exceeds the preset threshold , that is:
[0132] ;
[0133] It indicates that there is an abnormal fluctuation between the current heart rate estimation result and the historical estimation result. Perform weighted fusion on the current heart rate estimation result and the previous five - time average to eliminate the influence of the abnormal fluctuation. The weighted fusion formula is as follows:
[0134] ;
[0135] Among them, represents the weight of the current heart rate estimation result, and , which is used to ensure that the current heart rate estimation result occupies a larger proportion during fusion; represents the final heart rate estimation value.
[0136] Specifically, the average heart rate within a specified period is mainly used as the reference frequency of the new band - pass filter to adjust the band - pass range, paying more attention to the current heart rate. Generally, the average heart rate of the previous 3 times is taken. Here, the average heart rate is mainly used for historical data smoothing, so the average heart rate of the previous 5 times is taken, which is more biased towards historical heart rate data.
[0137] Specifically, under normal circumstances, the heart rate value does not fluctuate violently in a short period of time. Therefore, the recent several heart rate estimation results can be referred to "smooth" the current heart rate estimation value, avoiding unnecessary influence of sudden abnormal signal fluctuations on the final result and reducing the estimation error caused by instantaneous noise, motion artifacts or other sudden interferences. Based on this idea, the heart rate estimation smoothing strategy refers to introducing the mean value of the previous five heart rate estimation results as a reference and analyzing the difference between the current estimated heart rate and the mean value of the previous five times. If the gap between the current heart rate value and the mean value of the previous five heart rates exceeds the preset threshold, the two are weighted averaged to correct the current heart rate estimation value and achieve a more stable heart rate estimation. The specific implementation can be divided into three steps. Outlier detection: Calculate the mean value of the previous five heart rate estimation values. If the difference between the current heart rate estimation value and the mean value of the previous five heart rate estimations exceeds the threshold, it is determined as an outlier (assuming this may be caused by noise or other sudden interferences). If it is less than the threshold, it is determined as a normal value and no smoothing is required; Weighted correction: Smooth the outlier. Usually, the value can be adjusted according to the reliability of the signal. When the current heart rate signal is more stable and accurate, it will tend to a larger value. When the threshold is continuously exceeded, it is considered that the current heart rate signal is interfered, and it will be appropriately reduced to rely more on historical data. The final heart rate estimation value after S10 processing and the PPG estimation waveform after S8 filtering will be displayed on the final page.
[0138] As Figure 3 shown, this embodiment also discloses a non-contact heart rate estimation system based on multi-modal signal fusion, including:
[0139] A face video acquisition module 31, configured to acquire a sample face video;
[0140] A sharpening module 32, configured to highlight weak color changes in the sample face video through image sharpening technology to obtain a sharpened face video image;
[0141] A region of interest positioning module 33, configured to locate a region of interest based on facial feature points in the sharpened face video;
[0142] A mixed signal acquisition module 34, configured to extract signals in two color spaces, HSV and Lab, from the region of interest, fuse the signals in these two color spaces, and obtain a mixed signal containing information in two color spaces;
[0143] An original brightness compensation signal acquisition module 35, configured to suppress ambient light interference for the mixed signal using a background brightness compensation strategy to obtain an original brightness compensation signal;
[0144] A color transformation signal acquisition module 36 is configured to perform median filtering on the original luminance compensation signal to obtain a luminance compensation signal after median filtering; subtract the luminance compensation signal after median filtering from the original luminance compensation signal to obtain a color transformation signal;
[0145] A normalization signal acquisition module 37 is configured to normalize the color transformation signal by using the maximum and minimum values of the color transformation signal as the range for signal normalization to obtain a normalized signal;
[0146] An rPPG signal acquisition module 38 is configured to adjust the bandwidth of the original band-pass filter based on the historical heart rate, and perform band-pass filtering on the normalized signal through the adjusted band-pass filter to obtain an rPPG signal;
[0147] A heart rate estimation module 39 is configured to perform time-domain analysis and frequency-domain analysis on the rPPG signal, and weight the heart rate value estimated by time-domain analysis and the heart rate value estimated by frequency-domain analysis to obtain a heart rate estimation result;
[0148] A final heart rate estimation value acquisition module 310 is configured to correct the heart rate estimation result through a historical data smoothing strategy to obtain a final heart rate estimation value.
[0149] The specific implementation of the non-contact heart rate estimation system based on multi-modal signal fusion is the same as that of the non-contact heart rate estimation method based on multi-modal signal fusion, and this embodiment will not be repeated.
[0150] Although the present invention is specifically shown and described in conjunction with the preferred embodiments, those skilled in the art should understand that various changes can be made to the present invention in terms of form and details without departing from the spirit and scope of the present invention defined by the appended claims, and all of them fall within the protection scope of the present invention.
Claims
1. A non-contact heart rate estimation method based on multi-modal signal fusion, characterized in that, Including: S1, obtaining a sample face video; S2, highlighting the weak color changes in the sample face video through image sharpening technology to obtain a sharpened face video image; S3, locating the region of interest based on the facial feature points in the sharpened face video; S4, extracting the signals in the HSV and Lab color spaces from the region of interest, fusing the signals in these two color spaces to obtain a mixed signal containing the information of the two color spaces; S5, suppressing the ambient light interference on the mixed signal by using a background brightness compensation strategy to obtain an original brightness compensation signal; S6, performing median filtering on the original brightness compensation signal to obtain a median-filtered brightness compensation signal; subtracting the median-filtered brightness compensation signal from the original brightness compensation signal to obtain a color transformation signal; S7, normalizing the color transformation signal by using the maximum and minimum values of the color transformation signal as the range for signal normalization to obtain a normalized signal; S8, adjusting the bandwidth of the original band-pass filter based on the historical heart rate, and performing band-pass filtering on the normalized signal through the adjusted band-pass filter to obtain an rPPG signal; S9, performing time-domain analysis and frequency-domain analysis on the rPPG signal, and weighting the heart rate values estimated by time-domain analysis and the heart rate values estimated by frequency-domain analysis to obtain a heart rate estimation result; S10, correcting the heart rate estimation result through a historical data smoothing strategy to obtain a final heart rate estimation value.
2. The non-contact heart rate estimation method based on multi-modal signal fusion according to claim 1, characterized in that, In S4, it specifically includes: Extracting color signals from the HSV color space and the Lab color space to obtain color signal components H, color signal component S, color signal component V, color signal component L, color signal component a, and color signal component b, weighting the color signal components H, color signal component S, and color signal component V to obtain a mixed signal in the HSV color space; weighting the color signal components L, color signal component a, and color signal component b to obtain a mixed signal in the Lab color space.
3. The non-contact heart rate estimation method based on multi-modal signal fusion according to claim 1, characterized in that, In S5, it specifically includes: extracting the brightness signal outside the region of interest as a reference benchmark signal, subtracting the reference benchmark signal from the mixed signal in the HSV color space and the mixed signal in the Lab color space to obtain brightness compensation, and performing a weighted summation operation on the brightness-compensated mixed signal in the HSV color space and the mixed signal in the Lab color space to obtain an original brightness compensation signal.
4. The non-contact heart rate estimation method based on multi-modal signal fusion according to claim 1, characterized in that, In S8, the adjusting the bandwidth of the original band-pass filter based on the historical heart rate specifically includes: Obtaining the bandwidth range of the preset original band-pass filter, obtaining the average heart rate within a specified period based on the original band-pass filter, converting the average heart rate into a heart rate frequency, and adaptively adjusting the bandwidth range of the original band-pass filter based on the heart rate frequency to obtain an adjusted band-pass filter. The calculation formula is as follows: ; ; Among them, represents the band - pass range to the left of the reference frequency of the adjusted band - pass filter; represents the band - pass range to the right of the reference frequency of the adjusted band - pass filter; represents the heart - rate frequency corresponding to the average heart - rate during the specified period; and represent the range of the adjusted band - pass filter, which is used to filter out the parts with frequency components outside the range of , .
5. The non-contact heart rate estimation method based on multi-modal signal fusion according to claim 1, characterized in that, The S9 specifically includes: using FFT to extract the frequency-domain information in the rPPG signal, calculating the frequency component with the largest amplitude in the spectrum based on the frequency-domain information to obtain the heart rate value estimated by frequency-domain analysis; calculating the heart rate value estimated by time-domain analysis based on the peak intervals in the rPPG signal; combining the heart rate value estimated by time-domain analysis and the heart rate value estimated by frequency-domain analysis through weighted averaging to obtain the heart rate estimation result; the formula for the heart rate estimation result is as follows: ; Among them, represents the heart rate value estimated by frequency domain analysis; represents the heart rate value estimated by time domain analysis; represents the heart rate estimation result; is a weighting coefficient, and its value range is (0, 1), which is used to control the weights of frequency domain and time domain information in heart rate estimation.
6. The non-contact heart rate estimation method based on multi-modal signal fusion according to claim 1, characterized in that, The S10 specifically includes: calculating the average of the previous five heart rate values based on the recent five heart rate values, and the formula is as follows: ; Among them, , , , and are the heart rate values of the first five times respectively; represents the average heart rate of the first five times. The heart rate estimation result and the average of the previous five times are used to calculate the difference. If the difference exceeds the preset threshold , that is: ; It indicates that there is abnormal fluctuation between the current heart rate estimation result and the historical estimation result. The current heart rate estimation result and the average of the previous five values are weighted and fused to eliminate the influence of abnormal fluctuation. The weighted fusion formula is as follows: ; Among them, represents the weight of the current heart rate estimation result, and is used to ensure that the current heart rate estimation result occupies a greater proportion during fusion; represents the final heart rate estimated value.
7. A non-contact heart rate estimation system based on multi-modal signal fusion, characterized in that Includes: A face video acquisition module, used to acquire a sample face video; A sharpening module, used to highlight the weak color changes in the sample face video through image sharpening technology to obtain a sharpened face video image; A region of interest (ROI) positioning module, used to locate the region of interest based on the facial feature points in the sharpened face video; A mixed signal acquisition module, used to extract the signals in the HSV and Lab color spaces from the region of interest, and fuse the signals in these two color spaces to obtain a mixed signal containing the information of the two color spaces; An original brightness compensation signal acquisition module, used to suppress the ambient light interference on the mixed signal by using the background brightness compensation strategy to obtain the original brightness compensation signal; A color transformation signal acquisition module, used to perform median filtering on the original brightness compensation signal to obtain the brightness compensation signal after median filtering; subtracting the brightness compensation signal after median filtering from the original brightness compensation signal to obtain the color transformation signal; A normalized signal acquisition module, used to normalize the color transformation signal by using the maximum and minimum values of the color transformation signal as the range for signal normalization to obtain the normalized signal; An rPPG signal acquisition module, used to adjust the bandwidth of the original band-pass filter based on the historical heart rate, and perform band-pass filtering on the normalized signal through the adjusted band-pass filter to obtain the rPPG signal; A heart rate estimation module, used to perform time-domain analysis and frequency-domain analysis on the rPPG signal, and weight the heart rate value estimated by time-domain analysis and the heart rate value estimated by frequency-domain analysis to obtain the heart rate estimation result; A final heart rate estimation value acquisition module, used to correct the heart rate estimation result through the historical data smoothing strategy to obtain the final heart rate estimation value.
Citation Information
Patent Citations
Non-contact real-time heart rate detection method based on RGB-NIR camera
CN114220152A
Non-contact mental stress detection method and system
CN115553777A
Non-contact heart rate extraction method and device
CN116994292A
Non-contact heart rate detection method based on three-color channel signal fusion technology
CN119770014A
Non-contact fatigue detection system and method based on rppg
US20240081705A1
Cited By
Monitoring video enhancement method for farm
CN120634901A
Medical post-sending system for search and rescue type aircraft and data processing method
CN121015151A