Video pulse wave extraction method and system based on automatic noise recognition
By combining empirical mode decomposition and adaptive filtering techniques with multi-dimensional physiological feature recognition of motion artifacts, the problem of spectral overlap in video pulse signal extraction was solved, achieving high-quality pulse signal extraction in complex motion scenarios.
Patent Information
- Application Number
- CN202511408204.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2025-11-21
AI Technical Summary
现有视频脉搏信号提取技术在处理运动伪影与生理信号频谱重叠问题上存在困难,传统频域滤波方法无法有效分离运动伪影和脉搏信号,尤其在复杂运动场景下效果不佳。
An automatic noise identification method is adopted, which identifies motion artifacts by combining multiple physiological characteristics such as empirical mode decomposition, peak-to-peak interval variation coefficient, amplitude stability index and spectral purity. Noise is eliminated by using an adaptive filter and the signal is restored by an autoregressive model, thus achieving accurate separation of spectral overlapping signals and accurate extraction of pulse signals.
Effective signal separation was achieved under spectral overlap, ensuring accurate extraction and physiological integrity of pulse signals, adapting to complex motion scenarios, and improving signal quality.
Smart Images

Figure CN120997884A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of physiological signal detection, and more specifically, to a video pulse wave extraction method and system based on automatic noise recognition; Background Technology
[0002] With the rapid development of computer vision technology and biomedical engineering, video-based non-contact physiological parameter detection has become a research hotspot. This technology extracts pulse signals by analyzing subtle color changes on the human skin surface caused by blood pulsation, providing a convenient solution for applications such as remote health monitoring, driver fatigue detection, and newborn monitoring. Compared to traditional contact measurement devices, video pulse signal extraction technology has significant advantages such as non-invasiveness, low cost, and ease of deployment.
[0003] However, video pulse signal extraction faces severe technical challenges in practical applications. The most critical problem is the significant overlap in the frequency spectrum between motion artifacts and the pulse signal. The human pulse signal is primarily distributed within the 0.8-2.5 Hz frequency band, while motion artifacts generated by daily activities, such as slight head movements, facial expressions, and breathing, often have their spectral components concentrated in the 0.5-3 Hz range, highly overlapping with the pulse signal frequency band. This spectral overlap renders traditional frequency domain processing methods such as bandpass filtering and wavelet denoising ineffective, failing to effectively separate motion artifacts from the true pulse signal. Furthermore, the amplitude of motion artifacts is typically several orders of magnitude stronger than the pulse signal, further exacerbating the difficulty of signal separation.
[0004] Existing technologies mainly employ methods such as frequency domain filtering or independent component analysis to handle motion interference. For example, Chinese patent application CN202210856743.2 discloses a pulse wave extraction method based on adaptive filtering, but this method assumes that noise and signal are separable in the frequency domain, resulting in poor performance in handling spectral overlap. US patent US10,485,425B2 proposes a motion artifact removal method based on independent component analysis, but its performance significantly degrades when the signal components do not satisfy the statistical independence assumption.
[0005] In summary, existing video pulse signal extraction technologies have fundamental shortcomings in handling the problem of overlapping spectra of motion artifacts and physiological signals. Traditional frequency domain separation methods can no longer meet the application requirements in complex motion scenarios. There is an urgent need to develop new signal decomposition and noise identification strategies to achieve effective signal separation under spectral overlap conditions. Summary of the Invention
[0006] To address the problem that motion artifacts and physiological signal spectrum overlap during video pulse signal extraction, which causes traditional frequency domain filtering methods to fail in signal separation, this application provides a video pulse wave extraction method and system based on automatic noise identification. This method can achieve effective signal separation and accurate pulse signal extraction even when motion artifacts and pulse signal spectrums severely overlap.
[0007] One aspect of this application provides a video pulse wave extraction method based on automatic noise recognition, comprising: S1, acquiring a video signal containing a face, performing face detection on the video signal to obtain a facial region, dividing the facial region into multiple sub-regions, and extracting the original pulse signal of each sub-region; S2, performing bandpass filtering on the original pulse signals of the multiple sub-regions, calculating the signal-to-noise ratio of each sub-region in the pulse frequency band, and performing weighted fusion of the original pulse signals of each sub-region based on the signal-to-noise ratio to obtain a fused signal that suppresses local motion interference; S3, performing ensemble empirical mode decomposition on the fused signal to decompose the fused signal into multiple intrinsic mode function components; S4, further decomposing each intrinsic mode function component into multiple intrinsic mode function components. The system performs noise identification based on physiological signal characteristics. It identifies motion-induced irregular components using the peak-to-peak interval variation coefficient, identifies motion artifact variation characteristics using the amplitude stability index, and distinguishes between broadband noise and narrowband pulse signals using spectral purity. Based on the peak-to-peak interval variation coefficient, amplitude stability index, and spectral purity, it determines the noise component dominated by motion artifacts. In step S5, the identified noise components are reconstructed into a reference noise signal. An adaptive filter is then used, with the reference noise signal as input, to adaptively filter the fused signal, resulting in a filtered fused signal. In step S6, distortion detection is performed on the filtered fused signal, and signal restoration processing is applied to the detected distortion areas to obtain the extracted pulse signal.
[0008] Further, S2, bandpass filtering is performed on the original pulse signals of multiple sub-regions, the signal-to-noise ratio (SNR) of each sub-region within the pulse frequency band is calculated, and the original pulse signals of each sub-region are weighted and fused based on the SNR to obtain a fused signal that suppresses local motion interference, including: k = 1 to N; i = 1 to N; where S merge (t) represents the fused signal at time t, S k (t) represents the bandpass filtered pulse signal of the k-th sub-region at time t, where k is the sub-region index and N represents the total number of sub-regions; SNR i This represents the signal-to-noise ratio of the i-th sub-region, where i is the sub-region index, with a value ranging from 1 to N; in, This represents the signal power of the k-th sub-region within the 0.8-2.5Hz pulse frequency band. This represents the noise power of the k-th sub-region within the 0.1-0.5Hz low-frequency noise band;
[0009] Furthermore, S3 performs ensemble empirical mode decomposition on the fused signal, decomposing the fused signal into multiple intrinsic mode function components, including: j = 1 to J; where j represents the index of the intrinsic mode function component, ranging from 1 to J, and J represents the total number of intrinsic mode function components obtained from the decomposition; IMF j (t) represents the value of the j-th eigenmode function component at time t; r J (t) represents the value of the residual component after decomposition at time t;
[0010] Furthermore, S4, based on the peak-to-peak interval variation coefficient, amplitude stability index, and spectral purity, determines the noise component dominated by motion artifacts, including: for the j-th intrinsic mode function component (IMF) j (t) Perform peak detection, setting the minimum interval between adjacent peaks to 0.5×f s , where f s This represents the sampling frequency; if the number of detected peaks is less than 3, the corresponding component is determined to be a noise component; when the number of peaks is greater than or equal to 3, the peak-to-peak interval variation coefficient (CV) is calculated respectively. np Amplitude stability index MDR j and spectral purity SPR j According to the peak-to-peak interval coefficient of variation (CV) np Amplitude stability index MDR j and spectral purity SPR j The noise component score is calculated through weighted fusion. j Score j =w1×CV np +w2×MDR j +w3×SPR j Where w1, w2, and w3 are weighting coefficients;
[0011] When the noise component score j When the value exceeds the threshold, the corresponding component is identified as a noise component.
[0012] In particular, although motion artifacts and pulse signals overlap significantly in the spectrum (both concentrated in 0.5-3Hz), they exhibit distinctly different physiological characteristics in the time domain: the real pulse signal originates from the periodic pumping activity of the heart and has high quasi-periodicity (small coefficient of variation in peak-to-peak intervals), amplitude stability (smooth amplitude changes between adjacent sampling points), and spectral concentration (energy concentrated in a narrow band); while motion artifacts are generated by random body movements and exhibit irregularity in the time domain, drastic amplitude changes, and a broadband spectral distribution.
[0013] This application, based on the intrinsic mode function components after empirical mode decomposition, uses three complementary physiological characteristic indicators (peak-to-peak interval coefficient of variation, CV) to... np Amplitude stability index MDR j Spectral purity SPR j A multi-dimensional identification system is constructed. The key advantage of this method is that even if motion artifacts and pulse signals have completely overlapping spectra on a certain IMF component, they can still be distinguished through quasi-periodicity in the time domain and amplitude stability. By calculating a comprehensive score through weighted fusion of three indicators, accurate identification of noise components in complex aliased signals is achieved, fundamentally solving the signal separation failure problem caused by spectral overlap.
[0014] Furthermore, the coefficient of variation (CV) of the peak-to-peak interval is calculated. np The formula is as follows: Where σ×ΔT represents the peak-to-peak interval sequence ΔT of the j-th intrinsic mode function component. j The standard deviation; μ×ΔT represents the peak-to-peak interval sequence ΔT of the j-th intrinsic mode function component. j The mean; peak-to-peak interval sequence ΔT j It is calculated by the difference between adjacent peak time points;
[0015] Furthermore, the amplitude stability index MDR is calculated. j The formula is as follows: t = 1 to T-1; where, IMF j (t+1) represents the amplitude value of the j-th intrinsic mode function component at the (t+1)-th sampling time; IMF j (t) represents the amplitude value of the j-th intrinsic mode function component at the t-th sampling time; T represents the total number of sampling points of the j-th intrinsic mode function component; t represents the sampling time index;
[0016] Furthermore, the spectral purity SPR was calculated. j The formula is as follows: Among them, P j (f) represents the power spectral density function of the j-th intrinsic mode function component; f represents the frequency in Hz; f max This represents the maximum analysis frequency, which is half the sampling frequency. This indicates the power within the pulse signal frequency band of 0.8Hz to 2.5Hz; Indicates the range from 0 to f max Total power within the frequency range;
[0017] Furthermore, in step S5, the identified noise components are reconstructed into a reference noise signal. An adaptive filter, using the reference noise signal as input, is then applied to the fused signal to obtain the filtered fused signal, which includes: the adaptive filter e(t): e(t) = S merge (t)-w T (t)×N ref (t), where e(t) represents the output signal of the adaptive filter at time t, i.e., the filtered and fused signal; S merge (t) represents the fused signal at time t; w T (t) represents the transpose of the adaptive filter weight vector at time t; N ref (t) represents the reference noise signal at time t; the weight update formula for the adaptive filter e(t) is: Where w(t+1) represents the filter weight vector at time t+1; w(t) represents the filter weight vector at time t; ||N ref (t)|| 2 Represents the reference noise signal N ref The square of the second norm of (t), i.e., the energy of the reference noise signal; 10 -6 To prevent division by zero, a regularization constant is used; Among them, J noise This represents the set of intrinsic mode function component indices that are determined to be noise components based on the noise comprehensive score in step S4; IMF j (t) represents the value of the j-th intrinsic mode function component at time t;
[0018] Specifically, traditional adaptive filters require a separate noise reference channel, but a clean motion noise reference cannot be obtained in video pulse signal extraction. This application utilizes the noise-dominated IMF component identified in step S4 to reconstruct a reference noise signal N that is highly correlated with motion artifacts in the original signal. ref (t). Although this reference signal is not pure noise, it contains the main temporal characteristics and dynamic change patterns of motion artifacts.
[0019] Even if motion artifacts and pulse signals completely overlap in the spectrum, they still exhibit different correlations in the time-domain waveform. An adaptive filter dynamically adjusts the weight vector w(t) by minimizing the energy of the output signal e(t), such that the reference noise signal N... ref (t) After weighting, it can match and cancel the fused signal S to the greatest extent. merge The motion artifact component in (t). Since the pulse signal has low correlation with the reference noise signal, the pulse signal is preserved while the noise is eliminated.
[0020] This application does not rely on the separability of noise and signal in the frequency domain, but rather utilizes their correlation differences in the time domain to achieve separation. The reference noise signal obtained through IMF reconstruction can adaptively track the time-varying characteristics of motion artifacts, achieving precise elimination of spectral overlap noise, which is impossible for traditional fixed-band filters.
[0021] Further, in step S6, distortion detection is performed on the filtered fused signal, and signal restoration processing is performed on the detected distorted regions to obtain the extracted pulse signal, including: detecting signals that satisfy |e(t)-median(e)|>3×σ e The amplitude jump point, where σ e The standard deviation of the filtered fused signal e(t) is represented; median(e) represents the median of the filtered fused signal e(t) over the entire time series; detected abrupt changes are marked as distortion points;
[0022] An autoregressive model is used to restore the signal at the distortion points: in, The value of the recovered signal at time t is represented; P represents the order of the autoregressive model, which takes the value of 6; a p This represents the p-th autoregressive coefficient, where p ranges from 1 to P; the autoregressive coefficient a p The Yule-Walker equation is used to solve the signal, which is estimated using P normal signal points before the distortion point; s(tp) represents the value of the filtered and fused signal at time tp.
[0023] This also includes: quality assessment of the extracted pulse signal; calculation of the time-domain consistency feature F1. Calculate the frequency domain purity feature F2: Where H represents the spectral entropy of the extracted pulse signal in the 0.8-2.5Hz frequency band; H max H represents the theoretical maximum spectral entropy. max =log2(N), where N is the number of spectrum points;
[0024] Calculate the relative signal-to-noise ratio feature F3: Among them, SNR output The output signal-to-noise ratio (SNR) of the extracted pulse signal is obtained by calculating the ratio of the power in the 0.8-2.5Hz frequency band to the power in the 0.1-0.5Hz frequency band; input This represents the input signal-to-noise ratio of the fused signal in step S2;
[0025] The overall confidence level is calculated using a weighted average:
[0026] In particular, when dealing with motion artifacts with severely overlapping spectra, even after noise identification in S4 and adaptive filtering in S5, overcompensation or undercompensation may still occur at certain times, leading to local signal distortion (manifested as amplitude abrupt changes). Although these distortion points have a short time span, they can seriously affect the accuracy of subsequent applications such as heart rate variability analysis.
[0027] This application utilizes the physiological characteristics of pulse signals—quasi-periodicity and smooth continuity. It employs an autoregressive model... A time-series prediction model is established using the normal signal segment preceding the distortion point to achieve accurate restoration of the distortion point. The unique advantage of this method lies in the autoregressive coefficient α. p Adaptive estimation using the Yule-Walker equation can capture individual-specific pulse waveform characteristics and current physiological state, rather than simple interpolation or smoothing.
[0028] This application establishes a distortion detection criterion based on the median and standard deviation: |e(t)-median(e)|>3×σ e This approach can accurately locate abnormalities caused by residual motion artifacts or over-filtering, while avoiding misinterpreting normal physiological variations (such as respiratory sinus arrhythmia) as distortions. This two-step detection-recovery strategy ensures that the physiological integrity of the pulse signal is maintained to the greatest extent while eliminating motion artifacts, laying the foundation for accurate extraction of physiological parameters.
[0029] Another aspect of this application provides a video pulse signal extraction system, comprising: a video acquisition and preprocessing module, which acquires a video signal containing a face, performs face detection on the video signal to obtain a facial region, divides the facial region into multiple sub-regions, and extracts the original pulse signal of each sub-region; a signal fusion module, which performs bandpass filtering on the original pulse signals of the multiple sub-regions, calculates the signal-to-noise ratio of each sub-region in the pulse frequency band, and performs weighted fusion on the original pulse signals of each sub-region based on the signal-to-noise ratio to obtain a fused signal that suppresses local motion interference; a signal decomposition module, which performs ensemble empirical mode decomposition on the fused signal to decompose the fused signal into multiple intrinsic mode function components; and a noise recognition module, which identifies each intrinsic mode function component. The system performs noise identification based on physiological signal characteristics, identifies motion-induced irregular components through the peak-to-peak interval variation coefficient, identifies motion artifact variation characteristics through the amplitude stability index, and distinguishes broadband noise from narrowband pulse signals through spectral purity. The noise component dominated by motion artifacts is determined based on the peak-to-peak interval variation coefficient, amplitude stability index, and spectral purity. An adaptive filtering module reconstructs the identified noise components into a reference noise signal, and uses an adaptive filter with the reference noise signal as input to perform adaptive filtering on the fused signal, obtaining the filtered fused signal. A signal restoration module performs distortion detection on the filtered fused signal, and performs signal restoration processing on the detected distortion areas to obtain the extracted pulse signal.
[0030] Compared to existing technologies, the advantages of this application are:
[0031] This application initially suppresses local motion interference through multi-sub-region signal-to-noise ratio weighted fusion; it utilizes ensemble empirical mode decomposition to decompose aliased signals with overlapping spectra into multiple intrinsic mode function components in the time domain, achieving adaptive separation of motion artifacts and pulse signals at different components; it identifies noise components based on multi-dimensional physiological signal characteristics such as peak-to-peak interval variation coefficient, amplitude stability index, and spectral purity, rather than solely relying on spectral features, thereby accurately identifying motion artifact components overlapping with the pulse signal spectrum; it reconstructs the identified noise components into a reference signal for adaptive filtering, achieving precise denoising; and it further improves signal quality through distortion detection and signal restoration. This application overcomes the limitations of traditional frequency domain filtering in handling spectral overlap problems, enabling stable extraction of high-quality pulse signals in complex motion scenarios. Attached Figure Description
[0032] Figure 1 This is a schematic diagram of the process of an embodiment of this application;
[0033] Figure 2 This is a flowchart of the automatic noise component identification algorithm in an embodiment of this application;
[0034] Figure 3This is a diagram illustrating the effect of the denoising algorithm in an embodiment of this application. Detailed Implementation
[0035] The present application will now be described in detail with reference to the accompanying drawings and specific embodiments;
[0036] like Figure 1 As shown in this embodiment, a video pulse wave extraction method based on adaptive noise recognition and distortion restoration is provided, including the following steps: acquiring a video signal containing a face, obtaining and dividing the face detection region from the video signal; combining the signals of multiple sub-regions by maximum signal-to-noise ratio and then decomposing them by EEMD; extracting reference noise through an automatic noise component identification algorithm, using an adaptive filter for deep denoising, and repairing signal distortion based on an autoregressive model; evaluating the quality of the extracted pulse signal based on the time-frequency characteristics of the pulse signal, and simultaneously outputting the pulse signal and the pulse signal confidence level.
[0037] Specifically, in this embodiment, a high-definition network camera is used to collect facial video data in an indoor environment. The camera configuration parameters are: resolution 640×480 pixels, frame rate 30fps, shooting distance 0.5-0.8 meters, and video duration 1 minute. The acquisition environment is ordinary office lighting with an intensity of 200-300 lux, and no direct sunlight on the face. Using the existing SeetaFace facial recognition tool, the entire face is divided into multiple regions, including the forehead, cheeks, and eyes. Key regions are extracted from the facial detection areas and further divided into several sub-regions. In this embodiment, the key regions include the cheek region and the chin region, excluding noisy areas such as the facial features and the forehead, which are easily obscured. Signals can be extracted from each of the divided regions using the sub-regions, and the heart rate is calculated using the facial video signal from the most stable region. In this embodiment, the key regions are divided into 42 sub-regions. Since each sub-region corresponds to a different part of the face, different sub-regions provide blood flow information with varying stability when facing different driving environments. Furthermore, the original pulse signal is dynamically extracted from the facial video signals in the 42 sub-regions using existing video photoplethysmography.
[0038] Furthermore, a bandpass filter (0.8-2.5Hz) is first used for filtering to remove high-frequency noise and baseline drift. Then, the signal-to-noise ratio (SNR) of each sub-region is calculated, and the signals from each sub-region are weighted and fused to obtain a high-quality fused signal.
[0039]
[0040] Signal-to-noise ratio The energy range is 0.8-2.5Hz. This represents the baseline noise energy.
[0041] Furthermore, ensemble empirical mode decomposition (EEMD) is performed on the fused signal. Specific parameters are set as follows: white noise amplitude is added at 0.2 times the signal standard deviation, and the ensemble averaging is performed 100 times. This decomposition yields 8 intrinsic mode function (IMF) components and 1 residual component. The EEMD algorithm effectively solves the mode aliasing problem and is adaptable to the non-stationary and nonlinear characteristics of physiological signals. By processing the fused signal through ensemble empirical mode decomposition (EEMD), the set of intrinsic mode functions (IMFs) is obtained:
[0042]
[0043] Furthermore, adopt Figure 2 The noise component automatic identification algorithm shown performs a triple criterion analysis on each IMF component: First, it calculates the peak-to-peak interval variation coefficient:
[0044]
[0045] Where ΔT is the peak-to-peak interval sequence of the j-th IMF component, obtained by difference after detecting the peak position using the find_peaks function. The minimum peak interval is set to 0.5 seconds (corresponding to 120 bpm) to ensure physiological rationality.
[0046] Next, calculate the amplitude stability index:
[0047]
[0048] This metric reflects the relative rate of change of signal amplitude; noise components typically exhibit higher instability.
[0049] Finally, the spectral purity was calculated:
[0050]
[0051] The overall noise score is calculated based on the above three indicators:
[0052] NS j =0.4·CV j +0.3·MDR j +0.3·(1-SPR j )
[0053] The weighting coefficients were optimized based on a large amount of experimental data. When NS j When the value is greater than 0.7, it is determined to be the dominant noise component.
[0054] Furthermore, the reconstructed reference noise signal is used as the input reference noise for the adaptive filter:
[0055]
[0056] J noise This is the index set for noise components.
[0057] Furthermore, based on an adaptive filter, the depth components of the signal and noise are extracted to obtain a relatively clean pulse signal.
[0058] Construct an adaptive filter for noise cancellation:
[0059] e(t) = S merge (t)-ω T (t)N ref (t)
[0060] The filter weight update formula is:
[0061]
[0062] Algorithm effect as follows Figure 3 As shown, this set of images illustrates the denoising and processing of IPPG (Integrated Photoplethysmography) signals. The first image (Original Noisy IPPG Signal) shows the original noisy IPPG signal. IPPG acquires pulse signals non-contactly through optical means (such as a camera), but it is susceptible to noise interference from ambient light, motion, etc. The fluctuations in the blue waveform in the image reflect this noisy characteristic. The second image (Extracted Reference Noise Signal) shows the extracted reference noise signal. For subsequent denoising, the noise components need to be separated first. The yellow waveform in the image represents the noise reference extracted from the original signal, which helps filter out interference in the original signal. The third image (Filtered Signal vs Ground Truth) compares the filtered signal with the true value (Ground Truth).
[0063] Solid green line: Pulse signal extracted after phase alignment and autoregression processing.
[0064] The solid black line represents the pulse signal obtained after adaptive filtering.
[0065] Red dashed line: Ground truth, the real pulse signal obtained from the fingertip (usually a more precise contact acquisition, such as a finger clip sensor), is used as a standard for evaluating the noise reduction effect.
[0066] The comparison shows that the signals processed by both filtering methods (green and black solid lines) are approaching the true signal (red dashed line), demonstrating the optimization effect of the denoising algorithm on the original noisy IPPG signal.
[0067] Furthermore, to address pulse signal distortion caused by strong interference, an autoregressive model is used to restore the distortion. In the distortion restoration step:
[0068] The test satisfies |e(t)-median(e)|>3σ e The amplitude jump point is used to reconstruct the signal using an autoregressive model:
[0069]
[0070] Where the coefficient {a p Solve using the Yule-Walker equation.
[0071] Finally, the quality of the extracted pulse signal is evaluated based on the time-frequency characteristics of the pulse signal, and the pulse signal and pulse signal confidence level are output simultaneously.
[0072] Three evaluation features are defined:
[0073] Temporal consistency characteristics:
[0074]
[0075] Frequency domain purity characteristics:
[0076]
[0077] Relative signal-to-noise ratio characteristics:
[0078]
[0079] The overall confidence level is calculated using a weighted average:
[0080]
[0081] The weights are ω1 = 0.4, ω2 = 0.3, and ω3 = 0.3.
[0082] The foregoing illustrative description of the invention and its embodiments is not restrictive and can be implemented in other specific forms without departing from the spirit or essential characteristics of the invention. The accompanying drawings are only one embodiment of the invention, and the actual structure is not limited thereto. Therefore, if a person skilled in the art is inspired by this description and designs a similar structure and embodiment without departing from the spirit of the invention, such design should fall within the protection scope of the invention. Furthermore, the use of the word "include" does not exclude other elements or steps, and a word preceding an element does not exclude the inclusion of multiple elements. Terms such as "first" and "second" are used to indicate names and do not indicate any specific order.
Claims
1. A video pulse wave extraction method based on automatic noise recognition, characterized in that, include: S1, acquire video signals containing faces, perform face detection on the video signals to obtain the facial region, divide the facial region into multiple sub-regions, and extract the original pulse signal of each sub-region respectively; S2, bandpass filtering is performed on the original pulse signals of multiple sub-regions, the signal-to-noise ratio of each sub-region in the pulse frequency band is calculated, and the original pulse signals of each sub-region are weighted and fused according to the signal-to-noise ratio to obtain a fused signal that suppresses local motion interference; S3, perform ensemble empirical mode decomposition on the fused signal, decomposing the fused signal into multiple intrinsic mode function components; S4. Noise identification based on physiological signal characteristics is performed on each intrinsic mode function component. Irregular components caused by motion are identified by peak-to-peak interval variation coefficient, and the variation characteristics of motion artifacts are identified by amplitude stability index. Broadband noise and narrowband pulse signals are distinguished by spectral purity. The noise component dominated by motion artifacts is determined based on peak-to-peak interval variation coefficient, amplitude stability index and spectral purity. S5, the identified noise components are reconstructed into a reference noise signal, and an adaptive filter is used to perform adaptive filtering on the fused signal with the reference noise signal as input to obtain the filtered fused signal; S6 performs distortion detection on the filtered fused signal, performs signal restoration processing on the detected distorted areas, and obtains the extracted pulse signal.
2. The video pulse wave extraction method based on automatic noise recognition according to claim 1, characterized in that: S2, bandpass filtering is performed on the original pulse signals of multiple sub-regions, the signal-to-noise ratio (SNR) of each sub-region within the pulse frequency band is calculated, and the original pulse signals of each sub-region are weighted and fused based on the SNR to obtain a fused signal that suppresses local motion interference, including: k = 1 to N; i = 1 to N; Among them, S merge (t) represents the fused signal at time t, S k (t) represents the bandpass filtered pulse signal of the k-th sub-region at time t, where k is the sub-region index and N represents the total number of sub-regions; SNR i This represents the signal-to-noise ratio (SNR) of the i-th sub-region, where i is the sub-region index, ranging from 1 to N; k This represents the signal-to-noise ratio of the k-th sub-region; in, This represents the signal power of the k-th sub-region within the 0.8-2.5Hz pulse frequency band. This represents the noise power of the k-th sub-region within the 0.1-0.5Hz low-frequency noise band.
3. The video pulse wave extraction method based on automatic noise recognition according to claim 1, characterized in that: S3, performs ensemble empirical mode decomposition on the fused signal, decomposing the fused signal into multiple intrinsic mode function components, including: j = 1 to J; Where j represents the index of the intrinsic mode function component, ranging from 1 to J, and J represents the total number of intrinsic mode function components obtained from the decomposition; IMF j (t) represents the value of the j-th eigenmode function component at time t; r J (t) represents the value of the residual component after decomposition at time t.
4. The video pulse wave extraction method based on automatic noise recognition according to claim 1, characterized in that: S4, based on the peak-to-peak interval variation coefficient, amplitude stability index, and spectral purity, determines the noise component dominated by motion artifacts, including: For the j-th intrinsic mode function component (IMF) j (t) Perform peak detection, setting the minimum interval between adjacent peaks to 0.5×f s , where f s Indicates the sampling frequency; If the number of detected peaks is less than 3, the corresponding component is determined to be a noise component. When the number of peaks is greater than or equal to 3, calculate the peak-to-peak interval variation coefficient (CV) respectively. np Amplitude stability index MDR j and spectral purity SPR j ; Based on the peak-to-peak interval variation coefficient CV np Amplitude stability index MDR j and spectral purity SPR j The noise component score is calculated through weighted fusion. j Score j =w1×CV np +w2×MDR j +w3×SPR j Where w1, w2, and w3 are weighting coefficients; When the noise component score j When the value exceeds the threshold, the corresponding component is identified as a noise component.
5. The video pulse wave extraction method based on automatic noise recognition according to claim 4, characterized in that: Calculate the coefficient of variation (CV) of the peak-to-peak interval np The formula is as follows: Where σ×ΔT represents the peak-to-peak interval sequence ΔT of the j-th intrinsic mode function component. j The standard deviation; μ×ΔT represents the peak-to-peak interval sequence ΔT of the j-th intrinsic mode function component. j The mean; peak-to-peak interval sequence ΔT j It is calculated by the difference between adjacent peak time points.
6. The video pulse wave extraction method based on automatic noise recognition according to claim 4, characterized in that: Calculate the amplitude stability index MDR j The formula is as follows: t = 1 to T-1; Among them, IMF j (t+1) represents the amplitude value of the j-th intrinsic mode function component at the (t+1)-th sampling time; IMF j (t) represents the amplitude value of the j-th intrinsic mode function component at the t-th sampling time; T represents the total number of sampling points of the j-th intrinsic mode function component; t represents the sampling time index.
7. The video pulse wave extraction method based on automatic noise recognition according to claim 4, characterized in that: Calculate spectral purity SPR j The formula is as follows: Among them, P j (f) represents the power spectral density function of the j-th intrinsic mode function component; f represents the frequency in Hz; f max This represents the maximum analysis frequency, which is half the sampling frequency. This indicates the power within the pulse signal frequency band of 0.8Hz to 2.5Hz; Indicates the range from 0 to f max Total power within the frequency range.
8. The video pulse wave extraction method based on automatic noise recognition according to any one of claims 5 to 7, characterized in that: S5, the identified noise components are reconstructed into a reference noise signal. An adaptive filter is then used, with the reference noise signal as input, to perform adaptive filtering on the fused signal, resulting in the filtered fused signal, including: Adaptive filter e(t): e(t) = S merge (t)-w T (t)×N ref (t), where e(t) represents the output signal of the adaptive filter at time t, i.e., the filtered and fused signal; S merge (t) represents the fused signal at time t; w T (t) represents the transpose of the adaptive filter weight vector at time t; N ref (t) represents the reference noise signal at time t; The formula for updating the weights of the adaptive filter e(t) is: Where w(t+1) represents the filter weight vector at time t+1; w(t) represents the filter weight vector at time t; ||N ref (t)|| 2 Represents the reference noise signal N ref The square of the second norm of (t), i.e., the energy of the reference noise signal; 10 -6 To prevent division by zero, a regularization constant is used; Among them, J noise This represents the set of intrinsic mode function component indices that are determined to be noise components based on the noise comprehensive score in step S4; IMF j (t) represents the value of the j-th intrinsic mode function component at time t.
9. The video pulse wave extraction method based on automatic noise recognition according to claim 8, characterized in that: S6, perform distortion detection on the filtered fused signal, and perform signal restoration processing on the detected distorted areas to obtain the extracted pulse signal, including: The test satisfies |e(t)-median(e)|>3×σ e The amplitude jump point, where σ e The standard deviation of the filtered and fused signal e(t) is represented; median(e) represents the median of the filtered and fused signal e(t) over the entire time series. The detected abrupt change points are marked as distortion points; An autoregressive model is used to restore the signal at the distortion points: in, The value of the recovered signal at time t is represented by P; P represents the order of the autoregressive model; a p This represents the p-th autoregressive coefficient, where p ranges from 1 to P; the autoregressive coefficient a p The Yule-Walker equation is used to solve the signal, which is estimated using P normal signal points before the distortion point; s(tp) represents the value of the filtered fused signal at time tp.
10. A system for extracting video pulse signals, characterized in that, include: The video acquisition and preprocessing module acquires video signals containing faces, performs face detection on the video signals to obtain the facial region, divides the facial region into multiple sub-regions, and extracts the original pulse signal of each sub-region. The signal fusion module performs bandpass filtering on the original pulse signals of multiple sub-regions, calculates the signal-to-noise ratio of each sub-region in the pulse frequency band, and performs weighted fusion on the original pulse signals of each sub-region based on the signal-to-noise ratio to obtain a fused signal that suppresses local motion interference. The signal decomposition module performs ensemble empirical mode decomposition on the fused signal, decomposing the fused signal into multiple intrinsic mode function components; The noise identification module performs noise identification based on physiological signal characteristics for each intrinsic mode function component. It identifies irregular components caused by motion through the peak-to-peak interval variation coefficient, identifies the variation characteristics of motion artifacts through the amplitude stability index, and distinguishes broadband noise from narrowband pulse signals through spectral purity. Based on the peak-to-peak interval variation coefficient, amplitude stability index, and spectral purity, it determines the noise component dominated by motion artifacts. The adaptive filtering module reconstructs the identified noise components into a reference noise signal, and uses an adaptive filter with the reference noise signal as input to perform adaptive filtering on the fused signal to obtain the filtered fused signal. The signal restoration module performs distortion detection on the filtered fused signal, performs signal restoration processing on the detected distorted areas, and obtains the extracted pulse signal.
Citation Information
Patent Citations
Preparation method of perovskite quantum dot / precious metal nanoparticle composite material
CN115197691A
Apparatus and methods for structured light scatteroscopy
US10485425B2
Cited By
Heart rate calculation method and device, intelligent wearable equipment and storage medium
CN121570151A