Adaptive reconstruction method of non-contact pulse wave signal

By employing multi-ROI partitioning and adaptive signal decomposition algorithms, the problems of interference such as facial occlusion and shaking in non-contact pulse wave signal reconstruction were solved, enabling high-quality signal extraction and accurate physiological parameter measurement under extreme conditions.

CN116502062BActive Publication Date: 2026-02-10SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310477273.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-28
Publication Date
2026-02-10
Estimated Expiration
2043-04-28

AI Technical Summary

Technical Problem

Existing non-contact pulse wave signal reconstruction methods are easily affected by factors such as facial obstruction, head movement, and facial expressions, resulting in inaccurate signal quality.

Method used

A multi-ROI partitioning and adaptive signal decomposition algorithm, including variational mode extraction (VME) and multi-set canonical correlation analysis (MCCA), was used to screen high-quality pulse wave signals by combining time-domain and frequency-domain indices.

Benefits of technology

Under extreme conditions such as motion and changes in light, it can stably and accurately extract physiological parameters, improving the robustness and accuracy of signal reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116502062B_ABST
    Figure CN116502062B_ABST
Patent Text Reader

Abstract

The application relates to a kind of adaptive reconstruction methods of non-contact pulse wave signal, comprising the following steps: S10, from human face image, divide M region of interest ROI, calculate the pixel average value of R, G, B three channels of ROI, and it is converted to Cg' channel according to formula Cg'=G-0.75xR-0.25xB, obtain M initial rPPG signal;S20, eliminate baseline drift of initial rPPG signal;S30, to the initial rPPG signal after eliminating baseline drift, decomposition is carried out, obtains M groups of signal components;S40, the M groups of signal components are carried out multi-set canonical correlation analysis, and a plurality of groups of canonical variables are calculated, each group of canonical variables includes M canonical components, the sum of M canonical variables in each group of canonical variables is calculated, and the sum of at least two groups of canonical variables is retained as a pulse wave signal to be selected;S60, based on the index set, select a pulse wave signal from the pulse wave signal to be selected as the pulse wave signal obtained by reconstruction.The application can stably and accurately extract physiological parameters under the condition of greater interference.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of physiological information measurement, and particularly provides a self-adaptive reconstruction method of non-contact pulse wave signal. BACKGROUND

[0002] The pulse wave signal is an important physiological signal, which can be used for measuring physiological parameters such as heart rate, heart rate variability, blood pressure and blood oxygen saturation. The traditional method acquires the pulse wave through a pressure sensor or a finger clip pulse oximeter. These methods need to contact the human skin and need special equipment, and have problems such as poor comfort, unhygienic, limited application scenarios (such as measurement of burn patients and patients with skin diseases), etc. Remote photoplethysmography (rPPG) is a technology that uses optical principles to non-contact detect the blood volume changes in blood vessels. The basic principle is that the heart beats will cause the blood volume in the blood vessels to change periodically, so the absorption of light by capillaries will also change periodically, and in turn the color of the facial skin will change slightly. This change can be captured by an ordinary camera (such as a webcam) that remotely captures facial videos. Computer vision, signal processing and other technologies are used to reconstruct the hidden pulse wave signal in the video, and then the corresponding physiological parameters are calculated and analyzed. The rPPG technology has simple principles and low equipment cost, and can non-contact and continuously measure human physiological parameters in daily life, effectively overcoming the shortcomings of contact physiological parameter measurement methods, and has broad application prospects in remote medical care, special patient monitoring and other fields.

[0003] Although the rPPG technology has many advantages compared to the traditional contact method, a large amount of interference introduced by factors such as motion and environmental light reduces the quality of the pulse wave signal obtained by this technology, which seriously restricts the development and popularization of the rPPG technology. The traditional rPPG-based method selects a single facial region of interest (ROI) to calculate the pixel mean value, and uses the G channel signal in the RGB video to reconstruct the pulse wave. The reconstructed signal is easily disturbed by factors such as facial occlusion, head shaking and facial expressions. SUMMARY

[0004] The present application aims to improve the problem that the reconstructed signal is easily disturbed by factors such as facial occlusion, head shaking and facial expressions in the prior art, and provides a self-adaptive reconstruction method of non-contact pulse wave signal.

[0005] In order to achieve the above-mentioned application purpose, the embodiments of the present application provide the following technical solutions:

[0006] The self-adaptive reconstruction method of non-contact pulse wave signal comprises the following steps:

[0007] S10, detecting a face position from the obtained face image, and dividing M regions of interest (ROIs) at forehead, cheek and nose bridge positions, calculating the average pixel values of R, G and B channels of each ROI in each frame of face image, converting the average pixel values of R, G and B channels to Cg' channel according to the formula Cg'=G-0.75xR-0.25xB to obtain M initial rPPG signals, wherein R, G and B represent the average pixel values of the three channels respectively, and M is an integer greater than 1;

[0008] S20, removing baseline drift of the initial rPPG signal;

[0009] S30, decomposing the initial rPPG signal after removing baseline drift to obtain M groups of signal components, and obtaining one group of signal components for each ROI;

[0010] S40, performing multiple-set canonical correlation analysis on the M groups of signal components to calculate multiple groups of canonical variables, each group of canonical variables including M canonical components, summing the M canonical variables in each group of canonical variables, and retaining the sum of the first two groups of canonical variables as a candidate pulse wave signal;

[0011] S60, based on the set index, screening one pulse wave signal from the candidate pulse wave signals as the reconstructed pulse wave signal.

[0012] In the optimized scheme, the ROIs are 11, of which 4 ROIs are divided at the forehead, 3 ROIs are divided at each of the two cheeks, and 1 ROI is divided at the nose bridge. In this embodiment, the number of ROIs divided at the forehead, nose bridge, left cheek and right cheek is 4, 1, 3 and 3 respectively. The multiple ROIs can make full use of the signals of better quality in the unobstructed area, and the multiple ROIs are more likely to extract complete and continuous rPPG signals.

[0013] In the optimized scheme, when removing baseline drift, if the frequency of the main peak with the largest amplitude value is less than 1 Hz and the ratio of the frequency of the second largest peak to the amplitude value of the main peak is less than 0.67, the parameter a is taken as 2500, and if the frequency of the main peak is greater than 1 Hz, the parameter a is taken as 1000. In the processing of removing rPPG signal baseline drift by using VME algorithm, only one estimated center frequency needs to be set to extract the signal component of interest, which can effectively avoid modal aliasing, has high calculation efficiency, and has better adaptability. Moreover, according to the above manner to determine the parameter a, the removal effect of baseline drift can be further improved.

[0014] In the optimized scheme, in step S30, the initial rPPG signal after removing baseline drift is decomposed to obtain multiple groups of signal components, including:

[0015] The frequency range for variational mode extraction (VME) decomposition is set to [0.7, 4 Hz], τ = 0.001, α = 5000, and the center frequency is... The dominant frequency of the rPPG signal within the range of [0.7, 4Hz];

[0016] The VME algorithm is used to decompose the initial rPPG signal after baseline drift removal to obtain the extracted desired mode and residual components.

[0017] The remaining components after each VME decomposition are used as new signals to be decomposed and VME decomposition is performed again. The VME decomposition process is repeated until the signal energy reaches the convergence criterion and the iteration stops. The expected mode obtained from each decomposition is a signal component.

[0018] The above scheme is used to decompose rPPG signals. The entire decomposition process is adaptive based on the signal energy and signal quality, which can achieve accurate and efficient signal decomposition.

[0019] In the optimized scheme, in step S30, the convergence criterion is: Where R is an index characterizing the significance of the amplitude value of the main peak, R∈(0,1]; P i The amplitude value of the signal within the frequency band [0.7, 4Hz] is sorted from largest to smallest, with index i representing the amplitude value. When i < 5 and there exists... When N=i, N=i; otherwise N=5; E r E represents the energy of the residual signal within the frequency band [0.7, 4Hz] after the signal is decomposed. s This represents the total energy of the initial decomposed signal within the frequency band [0.7, 4Hz]. In this scheme, the convergence condition allows for fewer mode extractions when signal quality is good to avoid over-decomposition, while ensuring sufficient decomposition when signal quality is poor, thus improving the algorithm's efficiency and accuracy. The convergence criterion for decomposition is determined based on the quality and energy of the rPPG signal. The signal decomposition process is adaptive, efficient, and accurate, and it also avoids the shortcomings of traditional algorithms, such as mode aliasing and reliance on empirical parameters.

[0020] In the optimized scheme, step S50 is included before step S60, where the search range is set to [5000, 15000], the step size is 5000, α is updated, and steps S30 and S40 are repeated until the search is completed. Besides the typical variables in different orders in the MCCA results affecting pulse signal extraction, the decomposition effect of the PVME algorithm also affects the accuracy of the MCCA results. This scheme uses a cyclic iterative approach to search for the optimal value of α, which allows the PVME algorithm to achieve better signal decomposition results.

[0021] In the optimized scheme, in step S60, the indicators include waveform factor, peak factor, and heart rate standard deviation, wherein the waveform factor is... Peak factor is It is the root mean square value of the signal. It is the rectified average value of the signal, where n represents the signal length and x is the average value of the signal. i This represents the value of the signal at point i; v peak Indicates the peak value of the signal;

[0022] Heart rate is calculated by sliding windows of different sizes across the selected pulse wave signal, and the standard deviation of all obtained heart rate values ​​is calculated. This standard deviation is the heart rate standard deviation D. hr ;

[0023] F form -1.11、1 / C crest and D hr Normalize each to [0,1], and then add the three normalized indices together to obtain a comprehensive index. Search for the signal corresponding to the minimum value of the comprehensive index from the candidate pulse wave signals as the reconstructed pulse wave signal.

[0024] If we directly select a pulse signal that falls within the normal human heart rate range and use the amplitude of the main peak to determine the pulse signal, noise may be obtained, as noise can also exist within the normal heart rate range. In this scheme, through F... form -1.11、1 / C crest and D hr By comprehensively selecting a signal from the candidate pulse wave signals as the reconstructed signal, the influence of noise can be eliminated as much as possible, that is, the signal quality can be further improved, making the reconstructed signal closer to the real pulse wave signal.

[0025] In the optimized scheme, step S60 is followed by step S70, where the signal selected in step S60 is used as the input signal, and the dominant frequency of this signal is used as the center frequency of the signal to be extracted by the VME. VME signal extraction is then performed, and the resulting signal is the final reconstructed pulse wave signal. This scheme can further eliminate noise that may exist in the MCCA reconstructed signal, and further improve the quality of the time-domain waveform of the final reconstructed signal. In other words, this scheme can solve the problem of noise that may be present in the MCCA reconstructed signal.

[0026] Compared with existing technologies, the algorithm provided by this invention has the following advantages:

[0027] Dividing the data into multiple ROIs helps to extract the complete rPPG signal related to the heartbeat as much as possible under extreme conditions such as facial occlusion and large head movements; converting RGB data to the Cg′ color channel helps to reduce motion interference. Therefore, this invention can effectively improve the problem of inaccurate reconstructed signals caused by interference such as facial occlusion and head movements, and improve the accuracy of the reconstructed signals.

[0028] This invention proposes a novel decomposition algorithm (PVME) for rPPG signals. The entire decomposition process is adaptive based on the signal energy and signal quality, enabling accurate and efficient signal decomposition.

[0029] Using a multi-set canonical correlation analysis algorithm, under extreme conditions of low light and strong motion interference, information related to heartbeat is fully extracted. Combined with signal screening indicators, a high-quality pulse wave signal is reconstructed. These steps enable the algorithm provided by this invention to exhibit good robustness to interference caused by motion and changes in lighting.

[0030] Therefore, the method of this invention for non-contact pulse wave signal reconstruction can stably and accurately extract physiological parameters under conditions of significant interference, thus improving the practicality of rPPG technology. Attached Figure Description

[0031] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 The flowchart illustrates the adaptive reconstruction method for non-contact pulse wave signals provided in this embodiment of the invention.

[0033] Figure 2 This is a schematic diagram of ROI region division in the embodiment.

[0034] Figure 3a The waveform of the initial rPPG signal with baseline drift is shown. Figure 3b The waveform of the initial rPPG signal after baseline drift removal.

[0035] Figure 4a For a given initial rPPG signal after baseline drift removal, Figure 4b , Figure 4c , Figure 4d , Figure 4e , Figure 4f The waveforms of the signal components obtained by decomposing the initial signal are shown below.

[0036] Figure 5 This is a comparison diagram of the pulse wave reconstructed by the method described in the embodiments of the present invention and the actual pulse wave. Detailed Implementation

[0037] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0038] Please see Figure 1 The adaptive reconstruction method for non-contact pulse wave signals provided in this embodiment includes the following steps:

[0039] S10: The camera captures face video to obtain consecutive frames of face images. The face position is detected from the face images, and 11 regions of interest (ROIs) are divided in three locations: forehead, cheek, and bridge of nose. The average pixel value of the R, G, and B channels of each ROI in each frame of the face image is calculated. The data of the R, G, and B channels are converted to the Cg′ channel to obtain 11 initial rPPG signals.

[0040] Facial landmark detection, also known as facial landmark localization or face alignment, is the process of locating facial features based on the specific position of a face in an image obtained through face detection. These facial feature locations are discrete points with well-defined semantics, hence the name facial landmarks. Typically, facial landmarks are defined in the cheek, mouth, eyes, nose, and eyebrow regions of a face. Connecting these facial landmarks allows us to depict the geometric features of the face.

[0041] In this step, the face position is first detected and the coordinates of 468 key points (based on the open-source MediaPipe Face Mesh implementation) are obtained. To avoid background interference and maintain the accuracy of ROI segmentation, ROIs are segmented based on facial features using key point coordinates, so that ROIs can be dynamically adjusted according to facial key points. Dynamic adjustment includes: (1) When head rotation or facial tilt causes the relative position of the face in the image to change, the ROI will also automatically locate to the new position along with the key points. (2) When the face is deformed due to speaking or facial expressions, the ROI will also change along with the changes in key points.

[0042] To minimize interference from factors such as significant head movement and facial occlusion when defining Regions of Interest (ROIs), 11 ROIs were selected, covering the forehead, cheeks, and bridge of the nose. Specifically, 4 ROIs were defined on the forehead, 3 on each cheek, and 1 on the bridge of the nose. Figure 2 As shown in the figure (not all keypoints are displayed), the numbers represent the keypoint numbers, and ROI1…ROI11 represent the 11 ROIs. The four ROIs for the forehead are determined by keypoints numbered 67, 333, 297, 103, 109, 333, 109, 333, 337, and 109, 332, 333, 337, respectively. The three ROIs for each cheek are determined by keypoints numbered 116, 142, 206, 50, 203, 207, 36, 118, 187, 340, 349, 376, 411, 329, 411, 423, and 348, 411, respectively. The one ROI for the bridge of the nose is determined by keypoints numbered 114 and 355.

[0043] In this embodiment, the number of ROIs divided at the forehead, bridge of the nose, left cheek, and right cheek are 4, 1, 3, and 3 respectively. Figure 2 As shown, the Regions of Interest (ROIs) were selected in areas rich in capillaries on the face, and multiple ROIs of different locations and sizes were chosen. This has the following advantages: 1) When the face is partially occluded (e.g., by hair or hands), the occluded area may not contain physiological information or may be subject to strong interference. Multiple ROIs can fully utilize the better signal quality from the unoccluded areas; 2) Significant tilting and shaking of the face may cause some ROIs to be out of the video frame, resulting in missing data for those ROIs. Under these conditions, multiple ROIs are more likely to extract complete and continuous rPPG signals; 3) Due to interference from factors such as light source angle, illumination changes, and facial movement, ROIs of different locations and sizes on the face are subject to varying degrees of interference, resulting in different signal quality. Despite this, these ROIs can be considered to contain pulse information from the same source but with different intensities, while the noise within them is considered more random and uncorrelated. This lays the foundation for subsequent use of signal decomposition and blind source separation techniques to fully extract pulse signals under strong interference conditions.

[0044] In this step, the data from the R, G, and B channels are converted to the Cg′ channel according to the formula Cg′=G-0.75×R-0.25×B, where R, G, and B represent the pixel average values ​​of the three channels, respectively. Through extensive testing and analysis, the inventors discovered that, using the above formula for data conversion, when the G channel exhibits large amplitude fluctuations or unclear heart rhythms, Cg′ can better suppress such amplitude abrupt changes and clarify indistinct heart rhythms.

[0045] Traditional methods typically select only one Region of Interest (ROI) to calculate the average pixel value and reconstruct the pulse wave using the G channel signal from the RGB video image. This reconstructed signal is easily affected by factors such as facial occlusion, head movement, and facial expressions. In this step, dividing the face into multiple Regions of Interest (ROIs) avoids the problem of poor data quality or even unusable data in some ROIs due to facial occlusion or significant head movement. Furthermore, the Cg′ color channel is more resistant to motion interference than the G channel in the RGB color space, providing a clearer heartbeat rhythm.

[0046] S20, the baseline drift of the initial rPPG signal is removed using the variational mode extraction (VME) algorithm.

[0047] When removing baseline drift, based on the characteristic that the larger the VME algorithm parameter α is, the smaller the extracted mode bandwidth, the parameter α for VME to remove baseline drift is determined according to the main peak (the main peak refers to the peak with the largest amplitude value) in the frequency domain [0.7, 4Hz] of the initial rPPG signal. If the frequency of the main peak is less than 1Hz and the ratio of the second largest peak (the peak with the second largest amplitude value) to it is less than 0.67 (an optimal value obtained through experiments on four datasets (three public datasets UBFC_rPPG, PURE, and MMSE-HR, and one self-built dataset; this ratio reflects the significance of the main peak in the frequency domain), a larger α (α = 2500) is used to avoid VME excessively removing heartbeat-related signal components; if the main peak is greater than 1Hz, a smaller α (α = 1000) is used to reduce low-frequency noise interference. Larger α can be adjusted appropriately between [2000, 3000], and smaller α can be adjusted appropriately between [800, 1200]. However, in this step, using α = 2500 or α = 1000 can improve the baseline drift removal effect.

[0048] Baseline drift in rPPG signals can be caused by factors such as human respiration, facial movements, and changes in ambient light. It is characterized by low frequency, non-periodicity, and large amplitude variations. Baseline drift severely degrades the processing and analysis of subsequent signals in the time or frequency domains. Traditional methods for eliminating baseline drift (such as EMD, wavelet analysis, and polynomial fitting) have drawbacks such as the need for pre-selection of basis functions and the risk of mode aliasing. This invention fully considers the advantages of the VME algorithm in extracting specific modes and the characteristic that rPPG signal baseline drift is a narrowband signal within a specific frequency band. By applying it to the processing of rPPG signal baseline drift removal, only an estimated center frequency needs to be set to extract the signal component of interest. This effectively avoids mode aliasing, has high computational efficiency, and exhibits better adaptability.

[0049] VME transforms the signal decomposition process into solving variational problems within a variational framework. Assume the signal f(t) is decomposed into desired modes u. d (t) and residual component f r The specific process of the algorithm (t) can be summarized as follows:

[0050] 1. For the desired mode u d (t), its analytic signal is obtained by Hilbert transform, and then compacted at the estimated center frequency. Its variational expression is:

[0051]

[0052] In the formula, δ represents the Dirac distribution; * denotes convolution.

[0053] 2. In order to minimize the residual signal f r (t) and the desired mode u d The spectrum of (t) overlaps, and the following penalty function is considered:

[0054]

[0055] In the formula, β(t) is the impulse response of the filter used.

[0056] 3. Constraint condition f(t)=u d (t)+f r (t) guarantees that f(t) can be completely reconstructed, therefore the problem of finding the required mode can be expressed as a constraint minimization problem:

[0057]

[0058] subject to:u d (t)+f r (t)=f(t) (3)

[0059] In the formula, α is a parameter used to balance J1 and J2.

[0060] VME transforms the constrained variational problem into an unconstrained variational problem, and uses the alternating direction method of multipliers (ADMM) to find the saddle point of the Lagrangian function, i.e. the optimal solution of equation (3), to obtain the desired mode.

[0061] VME extracts the corresponding modes by setting the desired modal center frequency, and it only needs to set an estimated value for the desired center frequency to extract the desired signal components. The interference components causing baseline drift in the aforementioned initial rPPG signal are located in a specific low-frequency range, such as respiratory signals (0.1-0.4 Hz). Therefore, by setting only an estimated value (e.g., 0.1 Hz), the baseline drift of the initial rPPG signal can be eliminated using VME.

[0062] Based on the characteristics of the VME algorithm and the distribution of heart rate, this invention removes the baseline drift of the initial rPPG signal in two cases:

[0063] 1. The main peak of the rPPG signal in the frequency domain is less than 1Hz, and the ratio of the amplitude of the second largest peak to the amplitude of the main peak is less than 0.67. In order to avoid excessive removal of heartbeat-related signal components by VME, a larger α (α = 2500) is used.

[0064] 2. If the main peak of the rPPG signal in the frequency domain is greater than 1Hz, a smaller α (α = 1000) should be used to reduce low-frequency noise interference.

[0065] After determining the VME parameter α in both cases, the estimated frequency of its extracted mode is set to 0.1Hz, and the baseline drift can be removed by subtracting the extracted low-frequency component from the initial rPPG signal.

[0066] like Figure 3a and Figure 3b As shown, Figure 3a This is the original rPPG signal image. Figure 3b This image shows the result of using VME to remove baseline drift from the initial rPPG signal. It can be seen that the rPPG signal with baseline drift was strongly interfered with by a low-frequency component, resulting in significant amplitude fluctuations. After removing the baseline drift, a certain heartbeat rhythm can be observed in the rPPG signal.

[0067] S30, decompose the initial rPPG signal after removing baseline drift, and extract the modal components based on the main peak of the rPPG signal in the frequency band [0.7, 4Hz]. When the signal energy reaches the convergence criterion... The iteration stops when the time is right; this algorithm is called the PVME (Pulse Variational Mode Extraction) algorithm. After the initial rPPG signal of each ROI is decomposed, multiple signal components are obtained as a set of signal components. Thus, 11 sets of signal components can be obtained from 11 ROIs.

[0068] This step employs the PVME algorithm for signal decomposition to improve computational efficiency and accuracy. Specifically, it includes the following steps:

[0069] 1. Initialize parameters n←0, set the frequency range of VME decomposition to [0.7, 4Hz], τ=0.001, α=5000, center frequency The dominant frequency of the rPPG signal is within the range of [0.7, 4Hz].

[0070] 2. When n = n + 1, execute the loop.

[0071] 3. Update all values ​​with ω≥0 (Expected center frequency of the current cycle):

[0072]

[0073] 4. Update ω d (Signal component currently being extracted in the loop):

[0074]

[0075] 5. Update all Lagrange multipliers with ω≥0 according to equation (6):

[0076]

[0077] 6. Repeat steps 2 to 5 until the conditions of formula (7) are met, thus completing one desired mode extraction, i.e., completing one decomposition, and extracting the desired mode u. d (t) represents the result of each signal decomposition.

[0078]

[0079] 7. Assuming the input signal to be decomposed is f(t), and the desired mode extracted in the first step is u1(t), the remaining components of the signal are expressed as f′(t) = f(t) - u1(t). Then, using f′(t) as the new signal to be decomposed, steps 1 to 6 are repeated until the energy of the signal in the frequency band [0.7, 4Hz] satisfies the decomposition convergence condition set by formula (8):

[0080]

[0081] Where R is an index characterizing the significance of the amplitude value of the main peak, R∈(0,1]; P0 represents the amplitude value of the main peak of the signal within the frequency band [0.7,4Hz], P i This represents the peak value with index i, sorted from largest to smallest amplitude. When i < 5 and there exists... When N=i, N=i; otherwise N=5; E r E represents the energy of the residual signal within the frequency band [0.7, 4Hz] after the signal is decomposed. s This represents the total energy of the initially decomposed signal within the frequency band [0.7, 4Hz].

[0082] In step 7, the remaining energy of the rPPG signal is decomposed. The convergence condition set by formula (8) ensures that fewer mode extractions are performed when the signal quality is good to avoid over-decomposition. Conversely, when the signal quality is poor, the main peak within [0.7, 4Hz] is not obvious. This convergence condition ensures that the signal is fully decomposed, thus improving the efficiency and accuracy of the algorithm. Specifically, formula (8) and its constraints in step 7 ensure that the more obvious the main peak within the heartbeat frequency band, the smaller R is, and the more peaks with amplitudes similar to the main peak, i.e., the greater the noise interference, the larger R is. This convergence condition ensures that the rPPG signal is decomposed fewer times when the quality is good; for example, when the signal quality is extremely good, R approaches 0. R approaches 1, therefore VME will only perform one extraction, that is, extract the heartbeat signal corresponding to the main peak. Conversely, when the signal quality is extremely poor, R approaches 1 (approximately 0.97). 3 The coefficient before is determined to be -7.5. When R = 0.97, At approximately 0.001, the signal will be fully decomposed. This invention determines the convergence criterion for decomposition based on the quality and energy of the rPPG signal, resulting in a signal decomposition process that is adaptive, efficient, and accurate. Furthermore, it avoids the shortcomings of traditional algorithms, such as mode aliasing and reliance on empirical parameters.

[0083] Taking the result of the first ROI signal decomposition as an example, such as Figure 4a - Figure 4f As shown. Among them. Figure 4a This represents the unresolved rPPG signal. Figure 4b , Figure 4c , Figure 4d , Figure 4e , Figure 4f The corresponding IMF1 to IMF5 represent the five signal components obtained from the decomposition. It should be noted that the number of signals obtained after the above seven processing steps is not fixed, but depends on the quality of the rPPG signal. Figure 4b - Figure 4f This is merely an example to demonstrate the results of decomposing a portion of the rPPG signal after processing in this step.

[0084] S40, this invention divides the face into 11 ROIs. After decomposing the signals of each ROI based on PVME, 11 sets of data with homologous pulse wave information are obtained. These 11 sets of data are input into the MCCA algorithm to calculate multiple sets of canonical variables. The canonical variables of each set are summed, and the sum of the first two sets of canonical variables is retained as the candidate pulse wave signal.

[0085] When the face is simultaneously obscured, tilted significantly, subjected to changes in lighting, or affected by facial expressions during natural conversation, the rPPG signal will be strongly interfered with, manifesting as a large amount of low-frequency and high-frequency noise within the heartbeat frequency range. If frequency domain analysis is performed directly on the components decomposed by PVME, using the IMF (Intracranial Multiple Component) with its main peak within the heartbeat frequency range as the heartbeat signal, the noise components can be easily filtered out. To avoid the shortcomings of directly filtering the IMF as the pulse signal, step S30 of this invention utilizes PVME to decompose each sub-ROI of the face into a multi-dimensional dataset. These datasets contain common underlying heartbeat information because each heartbeat almost simultaneously affects the blood volume of all ROIs of the face.

[0086] Based on the above analysis, the problem of extracting heartbeat information from multiple multidimensional datasets can be viewed as a blind source separation problem. MCCA (multiset canonical correlation analysis) is a generalization of CCA (canonical correlation analysis) to multiple variable sets. It first transforms the original high-dimensional data into a low-dimensional subspace through multiple transformations, and then extracts relevant information from the transformed data. It is a joint blind source separation technique. A brief introduction to MCCA is as follows:

[0087] Given M datasets, let X m ∈R V×N Let V represent the m-th dataset, where V is the number of channels and N is the number of samples. Assuming each dataset is a linear mixture of multiple potentially independent sources, then:

[0088] X m =A m S m ,m=1,2,...,M (9)

[0089] Where A m ∈R V×H Let S represent the mixing matrix, H represent the number of potential sources, and S represent the number of potential sources. m ∈R H×N Unknown sources can be represented as: The superscript T indicates the transpose operation.

[0090] JBSS assumes that multiple datasets share a common latent source, and its goal is to estimate the mixture matrix A. m The inverse matrix W m To find the source component S m The formula for estimating one of the source vectors can be expressed as:

[0091]

[0092] Among them, w m ∈R V×1 It is the unmixing vector, which makes y = ... m Maximize the correlation; y m It is also known as a canonical variable (CV).

[0093] w m This can be obtained through an appropriate objective function; the objective function chosen in this paper is "SSQCOR". Finally, all obtained w m The unmixing matrix W will be formed m ∈R V×H We can obtain:

[0094]

[0095] Among them, Y m ∈R H×N It is an unknown source S m An approximate estimate of all Y m The composition Y∈R H×N×M .

[0096] In Y, each group of CVs has M CVs, and there are H groups in total. In this paper, M is the number of ROIs used, and H is the data dimension of each ROI. As defined by formula (10), each group of M CVs has maximized correlation. When an ROI is strongly disturbed, its CVs may have poor representation ability of the source signal. Therefore, in order to avoid the shortcomings of using a single CV and further reduce noise interference, the sum of all CVs in each group is directly used as the candidate signal for the pulse wave.

[0097] After summing each set of CVs, H pulse wave candidate signals will be reconstructed.

[0098] Of the H pulse wave candidate signals, some may be true pulse wave signals, while others mainly contain noise. Therefore, it is necessary to determine the signal that is closest to the actual pulse. Experiments showed that the true pulse signal is usually present in the first of the reconstructed H signals. This is consistent with the assumption that pulse signals are more correlated across multiple ROIs, while noise tends to vary more across different ROIs. However, in some cases, the true pulse signal appears as the second signal in the reconstructed signal, possibly because most ROIs are subjected to strong homogeneous noise. Based on this, the method of this invention retains the first two of the aforementioned H pulse wave candidate signals for further analysis. In other embodiments, more candidate signals can be retained, but this would increase the data processing workload.

[0099] S60, based on the set time-domain and frequency-domain indicators, selects a pulse wave signal from the candidate pulse wave signals as the reconstructed pulse wave signal.

[0100] Besides the impact of different CV sorting in the MCCA results on pulse signal extraction, the decomposition effect of the aforementioned PVME algorithm also affects the accuracy of the MCCA results. If PVME effectively separates noise and pulse signals in the rPPG signal, then the pulse signal is more correlated and the noise is more distinct during MCCA processing, making it easier for MCCA to recover the correct pulse signal. However, the PVME parameter α affects its performance. The smaller α is, the larger the bandwidth of each extracted mode, which may not effectively separate noise and pulse signals; conversely, the larger α is, the smaller the bandwidth of the extracted mode, which may lead to over-decomposition of the signal and distortion. Therefore, this invention uses a cyclic iterative approach to search for the optimal value of α to achieve better signal decomposition results for the PVME algorithm. The specific search steps are as follows:

[0101] 1) Initialize parameter α = 5000 (i.e., α = 5000 in step S30 during the first loop); the search range is [5000, 15000], and the step size is 5000;

[0102] 2) The PVME algorithm is used to decompose the rPPG signal of each ROI after removing baseline drift, resulting in multiple sets of multidimensional datasets. These datasets are input into MCCA to reconstruct H pulse wave candidate signals, and the first two signals are retained for further analysis. Step 2) here is actually the aforementioned steps S30 and S40.

[0103] 3) Based on the search criteria in 1), update α and repeat step 2) until the search is complete. Finally, six retained pulse wave candidate signals will be obtained. This step corresponds to... Figure 1 Step S50 in the process.

[0104] If the pulse signal is directly selected based on the amplitude of the main peak within the normal human heart rate range, noise may be obtained, as noise can also exist within the normal heart rate range. After fully analyzing the characteristics of noise and pulse signals, this invention proposes an index that integrates the frequency and time domain characteristics of rPPG signals.

[0105] First, pulse signals are more periodic than noise signals and more closely resemble the characteristics of a sine wave. Therefore, the form factor is defined as an indicator of the signal's time domain as follows:

[0106]

[0107] in, It is the root mean square value of the signal. It is the average rectified value of the signal, where n represents the signal length and x represents the signal value. i This represents the value of the signal at point i.

[0108] The waveform factor of a sine wave is approximately 1.11. Based on the periodic characteristics of pulse signals, it is believed that the closer the waveform factor of the aforementioned candidate pulse wave signals is to 1.11, the more likely it is to be the desired pulse signal to be extracted. Furthermore, the waveform factor reflects the degree of signal distortion; therefore, using the waveform factor as a time-domain indicator for screening and reconstructing pulse wave signals is reasonable.

[0109] Secondly, the less noise contained in the reconstructed pulse signal, the more distinct and obvious its peak value will be in the frequency domain. The peak factor characterizes the prominence of the signal peak; a larger peak factor indicates a more pronounced peak. Therefore, the peak factor is used as a frequency domain indicator reflecting the noise content of the reconstructed signal, and its definition is as follows:

[0110]

[0111] Among them, v peak Indicates the peak value of the signal, v rms This represents the root mean square value of the signal.

[0112] Finally, a normal person's heart rate is unlikely to undergo drastic changes in a short period, i.e., rapid and repeated increases and decreases over a short time. Therefore, it can be assumed that for a good quality pulse signal, using a sliding window exceeding one heartbeat cycle to calculate the heart rate across the entire data segment should result in a small standard deviation. Similarly, using sliding windows of different sizes to calculate the heart rate on the same pulse signal should still result in a relatively small standard deviation. Therefore, by using sliding windows of different sizes to calculate the heart rate on the reconstructed pulse signal and calculating the standard deviation of all heart rate values, this standard deviation is used as the third indicator for screening true pulse signals. Furthermore, the heart rate is calculated based on the main peak of the signal frequency domain within each window, thus obtaining an indicator that combines time-domain and frequency-domain analysis of the signal. The calculation steps of this indicator are briefly summarized as follows:

[0113] 1) Input the signal to be analyzed s (the pulse wave signal to be selected); set the loop range of the sliding window to [20, N], and the step size to 2, where N is the length of the signal.

[0114] 2) Update the sliding window size W win The sliding window moves across the signal 's' with a step size of 1, and the heart rate is calculated using the dominant peak value in the frequency domain within each window. The sliding calculation is performed over the entire signal to obtain... Heart rate values.

[0115] 3) Repeat step 2) until the loop is completed within [20, N], obtaining a set of heart rate values ​​h calculated using sliding windows of different sizes. Calculate the standard deviation of all heart rate values ​​in h to obtain the third index D for screening pulse signals. hr .

[0116] This invention employs a strategy that integrates the above three indicators to screen the signal from six candidate pulse signals that best represents heartbeat information. First, F... form -1.11、1 / C crest and D hr Normalized to [0,1], C crest Taking the reciprocal ensures that the smaller the three indicators are, the more likely it is to be a pulse signal. Then, the three normalized indicators are directly added together to obtain the comprehensive indicator I. Finally, the signal corresponding to the minimum value of I among the six signals is searched as the final extracted pulse signal, and the α corresponding to this signal is the optimal α obtained through optimization.

[0117] The following analysis uses a specific signal as an example:

[0118] Still with Figure 3b Taking the signal shown as an example, the rPPG signal after baseline drift removal will yield 6 candidate pulse wave signals, and the F corresponding to the 6 candidate pulse wave signals will be... form -1.11、1 / C crest and D hr The results, normalized to [0,1] respectively, and the comprehensive index I obtained by summing the three indices are shown in the table below:

[0119]

[0120] As can be seen from the table, the minimum value of the comprehensive index is 1, which corresponds to the signal 3, with a heart rate of 1.15Hz or 69bpm, which is exactly the same as the actual pulse signal of 1.15Hz.

[0121] S70, using the signal selected in step S60 as the input signal, the main frequency of the signal is used as the center frequency of the signal to be extracted by VME, VME signal extraction is performed, and the resulting signal is the pulse wave signal finally reconstructed.

[0122] Because the selected signal is reconstructed through MCCA, it may contain some noise, and the waveform quality in the time domain may not be very good. Therefore, after obtaining the correct pulse wave signal, in order to further improve the quality of the signal's time domain waveform, the dominant frequency of this signal (the signal selected in step S60) is again used as the center frequency of the signal that the VME expects to extract, and the heartbeat signal is accurately extracted using the VME. At this time, what needs to be extracted is a narrowband signal with a known frequency. Therefore, it is only necessary to set the VME parameter to a large value, that is, to extract a signal with a narrow bandwidth. α is set to 200000 in this step, and the extracted signal is the final reconstructed pulse wave signal.

[0123] exist Figure 3a In the process, the rPPG signal without baseline removal has extremely poor quality, exhibiting strong low-frequency interference and a large amount of high-frequency noise. After processing through steps S20-S70 of this invention, the final result is as follows: Figure 5 The results show that the pulse wave signal reconstructed by rPPG is in excellent agreement with the true value, which lays the foundation for improving the accuracy of non-contact pulse wave signal estimation of physiological parameters.

[0124] Based on the reconstructed pulse wave signal, physiological parameters such as heart rate and heart rate variability can be calculated. For example, by performing a Fast Fourier Transform on the reconstructed pulse wave signal, the maximum peak value in the frequency domain is the frequency f corresponding to the heartbeat. hr Then, the heart rate can be calculated using the following formula:

[0125] HR = f hr ×60 (9)

[0126] HR stands for heart rate.

[0127] To further verify the authenticity and reliability of the pulse wave signal reconstructed by the method described above in this invention, experiments were conducted on a self-built dataset (hereinafter referred to as DATA_30 in Tables 1 and 2). The dataset contains 30 test subjects, divided into five lighting variations and motion environments: low illumination, head shaking, high illumination, unbalanced illumination (uniform brightness on the left and right sides of the face), and normal illumination. In addition, experiments were also conducted on three publicly available datasets. The results of the method provided in this invention are shown in Table 1, and Table 2 shows the results of the traditional method (GREEN method, which uses a bandpass filter to filter the green channel signal to obtain the pulse wave signal. Here, the GREEN method averages the data from the green channels of all ROIs as the original rPPG signal to be filtered). The bolded data indicates the optimal result.

[0128] Table 1. Experimental results of the method provided by this invention on multiple datasets.

[0129]

[0130]

[0131] Among them, mean absolute error (MAE), root mean square error (RMSE), and Pearson correlation coefficient (PCC) are commonly used evaluation indicators in the industry. In the formula, R i For the i-th heart rate estimate, G i Let be the i-th true heart rate value, N be the data length for calculating heart rate, cov(R,G) be the covariance between the estimated and true heart rate values, and σ be the true heart rate value. R and σ G These represent the standard deviations of the estimated and actual heart rate values, respectively.

[0132] As shown in Table 1, the method provided by this invention exhibits small errors on three publicly available datasets (UBFC_rPPG, PURE, and MMSE-HR) and a self-built dataset (DATA_30), with a minimum MAE of only 0.848 bpm and a maximum of only 1.64 bpm. Even in environments with significant interference such as low illumination, head movement, and unbalanced illumination, the method still achieves low MAEs of 1.533 bpm, 1.1 bpm, and 1.333 bpm, respectively. The minimum RMSE is 2.258 bpm, and the maximum is 3.759 bpm, thus demonstrating the good robustness and stability of the method. Furthermore, the highest PCC is 0.994, and the lowest is 0.962, further proving the consistency and effectiveness of the method in different scenarios.

[0133] As shown in Table 2, compared with the traditional method of directly reconstructing the pulse wave using the green channel signal, the method provided by this invention, as shown in Table 1, reduces the minimum MAE by 4.796 bpm and the maximum MAE by 12.26 bpm; the minimum RMSE is reduced by 9.591 bpm and the maximum RMSE is reduced by 17.742 bpm; the minimum PCC of this invention is 0.981 higher than that of the traditional method, and the maximum PCC is 0.286 higher. Therefore, the advantages of the method provided by this invention compared to the traditional method have been verified.

[0134] Table 2 Experimental results of traditional methods on multiple datasets.

[0135] Dataset MAE / bpm RMSE / bpm PCC UBFC_rPPG 7.708 20.532 0.558 PURE 5.644 17.86 0.679 MMSE-HR 7.145 14.448 0.582 DATA_30 low light 12.2 19.147 0.222 DATA_30 head motion 9.6 15.04 0.566 DATA_30 high light 12.943 20.759 -0.019 DATA_30 unbalanced light 13.9 21.501 0.125 DATA_30 normal light 6.4 11.849 0.708

[0136] This invention proposes a novel method for non-contact pulse wave signal reconstruction using rPPG. This method improves the algorithm's performance in several aspects, including multi-channel ROI data, color space conversion, baseline drift removal, signal decomposition, MCCA analysis, and time-domain and frequency-domain indices. Experimental results on the self-built dataset DATA_30 under five different environments demonstrate that this method is robust to changes in illumination and motion interference. Even under extreme interference conditions, the method provided by this invention can still reconstruct a high-quality signal with good consistency with the real pulse wave. Furthermore, results from three publicly available datasets verify the good generalization ability of this method, laying a solid foundation for the promotion and application of rPPG technology in remote measurement of human physiological parameters.

[0137] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. An adaptive reconstruction method for non-contact pulse wave signals, characterized in that, Includes the following steps: S10: Detect the face location from the captured face image, and divide the face into M Regions of Interest (ROIs) at the forehead, cheeks, and bridge of the nose. Calculate the average pixel values ​​of the R, G, and B channels for each ROI in each frame of the face image, and then calculate the average pixel values ​​of the R, G, and B channels according to the formula... Switch to The channel yields M initial rPPG signals, where R, G, and B represent the pixel average values ​​of the three channels, and M is an integer greater than 1. S20, removes baseline drift of the initial rPPG signal; S30, decompose the initial rPPG signal after removing baseline drift to obtain M groups of signal components, with each ROI corresponding to a group of signal components; S40, perform multi-set canonical correlation analysis on the M groups of signal components to calculate multiple groups of canonical variables. Each group of canonical variables includes M canonical components. Sum the M canonical variables in each group of canonical variables and retain the sum of at least the first two groups of canonical variables as the candidate pulse wave signal. S60, based on the set index, selects a pulse wave signal from the candidate pulse wave signals as the reconstructed pulse wave signal; In step S30, the initial rPPG signal after baseline drift removal is decomposed to obtain multiple sets of signal components, including: The frequency range for variational mode extraction of VME decomposition is set to [0.7, 4 Hz]. , Center frequency The dominant frequency of the rPPG signal within the range of [0.7, 4 Hz]; The VME algorithm is used to decompose the initial rPPG signal after baseline drift removal to obtain the extracted desired mode and residual components. The remaining components after each VME decomposition are used as new signals to be decomposed and VME decomposition is performed again. The VME decomposition process is repeated until the signal energy reaches the convergence criterion and the iteration stops. The desired mode obtained from each decomposition is a signal component. In step S60, the indicators include waveform factor, peak factor, and heart rate standard deviation, wherein the waveform factor is... Peak factor is , It is the root mean square value of the signal. It is the rectified average value of the signal, where n represents the signal length. This represents the value of the signal at point i. Indicates the peak value of the signal; Heart rate is calculated by sliding windows of different sizes across the selected pulse wave signal, and the standard deviation of all obtained heart rate values ​​is calculated. This standard deviation is the heart rate standard deviation D. hr ; Will , and D hr Normalize each of the three indices to [0, 1], and then add the three normalized indices together to obtain a comprehensive index. Search for the signal corresponding to the minimum value of the comprehensive index from the candidate pulse wave signals as the reconstructed pulse wave signal.

2. The adaptive reconstruction method for non-contact pulse wave signals according to claim 1, characterized in that, There are 11 ROIs, including 4 ROIs on the forehead, 3 ROIs on each side of the cheeks, and 1 ROI on the bridge of the nose.

3. The adaptive reconstruction method for non-contact pulse wave signals according to claim 1, characterized in that, In step S20, the variational mode extraction (VME) algorithm is used to remove the baseline drift of the initial rPPG signal.

4. The adaptive reconstruction method for non-contact pulse wave signals according to claim 3, characterized in that, When eliminating baseline drift, if the frequency of the main peak with the largest amplitude is less than 1 Hz and the ratio of the frequency of the second largest amplitude peak to the amplitude of the main peak is less than 0.67, then the parameters are taken as follows: If the frequency of the main peak is greater than 1 Hz, take the parameter. .

5. The adaptive reconstruction method for non-contact pulse wave signals according to claim 1, characterized in that, In step S30, the convergence criterion is: Where R is an indicator representing the significance of the amplitude value of the main peak. ;P i The amplitude value of the signal within the frequency band [0.7, 4 Hz], sorted from largest to smallest, with index i. And exist When N=i, N=i; otherwise N=5; E r E represents the energy of the residual signal within the frequency band [0.7, 4 Hz] after the signal is decomposed. s This represents the total energy of the initially decomposed signal within the frequency band [0.7, 4 Hz].

6. The adaptive reconstruction method for non-contact pulse wave signals according to any one of claims 1-4, characterized in that, In step S40 , , , As typical variables, It is the unmixing vector. Obtained through the objective function SSQCOR, all obtained Composition of the unmixing matrix , Let H represent the mixing matrix, and H represent the number of potential sources. The unknown source is V, the number of channels is N, the number of samples is H, the number of potential sources is H, and the superscript T indicates the transpose operation.

7. The adaptive reconstruction method for non-contact pulse wave signals according to claim 1, characterized in that, Before step S60, there is also step S50, which sets the search range to [5000, 15000], the step size to 5000, and updates... Repeat steps S30 and S40 until the search is complete.

8. The adaptive reconstruction method for non-contact pulse wave signals according to claim 1, characterized in that, Step S60 is followed by step S70, in which the signal selected in step S60 is used as the input signal, and the main frequency of the signal is used as the center frequency of the signal to be extracted by VME. VME signal extraction is performed, and the resulting signal is the pulse wave signal that is finally reconstructed.

Citation Information

Patent Citations

  • Non-contact video heart rate detection method based on multivariate empirical mode decomposition and joint blind source separation

    CN110269600A

  • Pulse wave extraction method and device, electronic equipment and storage medium

    CN113712526A