Blood pressure measurement device based on multi-frequency modal fusion

By using multi-frequency modal fusion technology, the accuracy of pulse signal detection in blood pressure monitoring devices under complex environments has been solved, enabling rapid and accurate calculation of heart rate and noise filtering, thereby improving the robustness and detection effect of the device.

CN114652282BActive Publication Date: 2026-02-03XIAN SINGULARITY FUSION INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210270275.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-18
Publication Date
2026-02-03
Estimated Expiration
2042-03-18

AI Technical Summary

Technical Problem

Existing blood pressure monitoring devices struggle to accurately estimate pulse signals when faced with complex noise signals such as changes in ambient light conditions and facial movement. Furthermore, traditional methods are susceptible to interference, leading to decreased detection accuracy.

Method used

A multi-frequency modal fusion method is adopted. The spatiotemporal feature mapping is extracted through the data preprocessing module, the multi-frequency modal signal separator is used to decompose it into initial multi-frequency band signals, the signal refinement network performs multi-scale feature extraction, the signal reconstruction network performs waveform reconstruction, and finally the multi-frequency modal signal fusion device determines the blood volume pulse wave signal, thereby reducing the overfitting of the training network.

Benefits of technology

It achieves accurate heart rate calculation with fewer parameters and in a shorter time, improving the robustness and detection accuracy of the device, effectively filtering out noise in complex environments, and improving the signal-to-noise ratio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114652282B_ABST
    Figure CN114652282B_ABST
Patent Text Reader

Abstract

The present application relates to a blood pressure measuring device based on multi-frequency modal fusion, comprising: a data preprocessing module for extracting a space-time feature map from a face video; a multi-frequency modal signal separator for decomposing the space-time feature map into an initial multi-frequency band signal, the initial multi-frequency band signal comprising n signal components; a signal refining network for performing multi-scale feature extraction on the n signal components respectively, thereby corresponding to generate n high-dimensional intrinsic modal signals; a signal reconstruction network for performing waveform reconstruction on the n high-dimensional intrinsic modal signals respectively, thereby obtaining n high-dimensional signals; and a multi-frequency modal signal fusion device for selecting effective components from the n high-dimensional signals for fusion, and determining the positions of wave crests and troughs to obtain a blood volume pulse wave signal. The present application is based on a lightweight short-time heart rate real-time monitoring network of multi-modal fusion, which can refine effective signal features with a small number of parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of blood pressure measurement device technology, and in particular to a blood pressure measurement device based on multi-frequency mode fusion. Background Technology

[0002] Blood volume pulse wave (BVP) is an important physiological signal with great potential in areas such as blood pressure monitoring, heart rate detection, arteriosclerosis detection, blood oxygen saturation detection, and alcohol consumption detection. Current blood pressure monitoring devices, which use contact photoplethysmography (cPPG), require skin contact with the subject, limiting their application scope and areas of expertise.

[0003] In recent years, an increasing number of methods have been designed into blood pressure measurement devices for remote estimation of heart rate or pulse signals. However, most devices can only estimate heart rate and not accurately estimate pulse signals. Some solutions improve detection accuracy by using more advanced acquisition devices that are difficult to implement in most cases. To reduce reliance on devices, methods designed into devices include remote optical volumetric spectroscopy (rPPG) and globocentesis (BBG) to enable heart rate monitoring based on traditional cameras. However, these methods are susceptible to complex noise signals caused by changes in ambient light conditions and facial movement. Summary of the Invention

[0004] The purpose of this invention is to use the concept of empirical mode decomposition and fusion to decompose the input signal into multiple intrinsic mode signals in the face of strong interference noise, and finally select the effective components for fusion to obtain the final measured blood volume pulse wave signal. At the same time, it also avoids the problem of overfitting of the training network to a certain extent, and provides a blood pressure measurement device based on multi-frequency mode fusion.

[0005] To achieve the above-mentioned objectives, the embodiments of the present invention provide the following technical solutions:

[0006] Blood pressure measurement devices based on multi-frequency modal fusion include:

[0007] The data preprocessing module is used to extract spatiotemporal feature maps from face videos;

[0008] A multi-frequency mode signal separator is used to decompose a spatiotemporal feature map into an initial multi-frequency band signal, wherein the initial multi-frequency band signal includes n signal components;

[0009] The signal refinement network is used to extract multi-scale features from n signal components, thereby generating n high-dimensional intrinsic mode signals.

[0010] A signal reconstruction network is used to reconstruct the waveforms of n high-dimensional intrinsic mode signals, thereby obtaining n high-dimensional signals.

[0011] A multi-frequency modal signal fusion unit is used to select effective components from n high-dimensional signals, fuse them, and then determine the positions of peaks and troughs to obtain the blood volume pulse wave signal.

[0012] Furthermore, the data preprocessing module includes a face video recognition engine and a spatiotemporal feature acquisition module, wherein,

[0013] The face video recognition engine is used to acquire face videos and perform landmark detection on the face video in frame t using MediaPipe. The face region is then segmented using the landmark to obtain a set of average pixel values ​​from M regions, forming the initial region features of frame t. t∈T, where T represents the total number of frames;

[0014] The spatiotemporal feature acquisition module is used to integrate the initial region features of the T-frame to obtain the initial signal features. And perform overall time-domain normalization on the initial signal features to obtain normalized signal features:

[0015]

[0016] And the normalized signal feature C t After adding white noise for enhancement, a spatiotemporal feature map is generated.

[0017] Furthermore, the signal refinement network includes four refinement modules with different scales but the same structure connected in sequence. Each refinement module includes a convolutional layer, a temporal multi-scale perceptron, an activation layer, and a batch normalization (BN) layer connected in sequence.

[0018] Each refining module at a different scale is used to progressively capture the time-domain characteristics of the signal component X, thereby generating the corresponding high-dimensional intrinsic mode signal.

[0019] Furthermore, the time-domain multi-scale perceptron includes a dimension transformation module and an extraction module, wherein,

[0020] The dimension transformation module is used to transform each input signal component. Perform compression transformation to form set up This reduces the network redundancy parameters of the multi-scale perceptron;

[0021] The extraction module includes a multi-scale filter for X U Adaptive extraction of refined features X R The refining feature X R Represented as:

[0022] Furthermore, the signal reconstruction network includes three reconstruction modules with different scales but the same structure that are connected in sequence. Each reconstruction module includes a spectrum-based self-attention mechanism layer, a convolutional layer, a batch normalization (BN) layer, an activation layer, and a pooling layer that are connected in sequence.

[0023] The spectrum-based self-attention mechanism layer is used to perform discrete cosine transform on the high-dimensional intrinsic mode signal to obtain the frequency characteristics of the high-dimensional intrinsic mode signal; then, based on the frequency characteristics, the high-dimensional intrinsic mode signal of the discrete cosine transform is spliced ​​into a one-dimensional feature signal, and self-attention weights are extracted; finally, the self-attention weights are aggregated with the high-dimensional intrinsic mode signal, and the one-dimensional feature signal is recombined to improve the signal-to-noise ratio.

[0024] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0025] (1) The present invention is based on a lightweight short-term real-time heart rate monitoring network with multimodal fusion, which can refine effective signal features with a small number of parameters;

[0026] (2) The present invention can accurately calculate heart rate values ​​using only 15 seconds of facial video signal;

[0027] (3) The present invention designs an oversampling training scheme for heart rate monitoring tasks, which helps the network to fully learn and fit the features of each heart rate zone with the assistance of a limited amount of training data.

[0028] (4) The time-domain multi-scale perceptron designed in this invention can further improve the robustness of the network. Attached Figure Description

[0029] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a block diagram of the blood pressure measurement device module of the present invention. Detailed Implementation

[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0032] It should be noted that similar reference numerals and letters in the following figures denote similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, the terms "first," "second," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance, or suggesting any such actual relationship or order between these entities or operations.

[0033] Example 1:

[0034] This invention is achieved through the following technical solutions, such as... Figure 1 As shown, the blood pressure measurement device based on multi-frequency mode fusion mainly consists of five parts: a data preprocessing module, a multi-frequency mode signal separator, a signal refinement network, a signal reconstruction network, and a multi-frequency mode signal fusion device. These five parts are integrated into one device.

[0035] After acquiring the face video, the data preprocessing module extracts spatiotemporal mapping features from it. The multi-frequency modal signal separator decomposes the spatiotemporal mapping features into initial multi-frequency band signals, which include n signal components. The signal refinement network performs multi-scale feature extraction on the n signal components respectively, generating n high-dimensional intrinsic modal signals. The signal reconstruction network performs waveform reconstruction on the n high-dimensional intrinsic modal signals respectively to obtain n high-dimensional signals. The multi-frequency modal signal fusion unit selects effective components from the n high-dimensional signals and fuses them to determine the positions of peaks and troughs, thereby obtaining the blood volume pulse wave signal.

[0036] As the first part, the data preprocessing module includes a face video recognition engine and a spatiotemporal feature extraction module. Facial information contains weak signals reflecting blood flow in blood vessels, but noise generated by changes in face position and facial expression can contaminate the effective signal. Therefore, in order to more robustly identify physiological signals, the first part first acquires face video and then obtains spatiotemporal feature mapping after preprocessing the face video.

[0037] The face video recognition engine is used to acquire face videos and uses MediaPipe to perform landmark detection on the face video in frame t. The face region is segmented using the landmark to obtain a set of average pixel values ​​of M regions, forming the initial region features of frame t. t∈T, where T represents the total number of frames.

[0038] The spatiotemporal feature acquisition module is used to integrate the initial region features of the T-frame to obtain the initial signal features. And perform overall time-domain normalization on the initial signal features to obtain normalized signal features:

[0039]

[0040] Normalization can effectively help networks accurately recover physiological signals when there is a lot of noise interference. In order to effectively simulate various potential noise signals caused by harsh ambient light conditions and motion such as facial expressions during the training phase, the normalized signal feature C is optimized. t White noise is added for enhancement, thereby generating a spatiotemporal feature map.

[0041] As the second part, the multi-frequency mode signal separator decomposes the spatiotemporal feature map into an initial multi-band signal, which includes n signal components. Traditionally, various time-domain and frequency-domain filters are used with blind source signal separation algorithms to effectively filter noise from facial reflected light signals and extract blood volume pulse wave signals. However, this "separation" only separates the signal of interest. Most blind source signal separation algorithms employ independent component analysis, decomposing the received mixed signal into several independent components, which serve as an approximate estimate of the source signal. However, in the task of extracting photoelectric blood volume pulse wave signals, the problem involves composite signals with unknown distributions and a large amount of noise, which greatly limits the effectiveness of blind source signal separation algorithms in independent component analysis. Many blind source separation algorithms use the frequency domain to filter noise, but noise exists in various complex forms. Simply using manually defined thresholds for signal extraction and fixed weights for fusion is insufficient to effectively handle strong noise interference.

[0042] Therefore, this scheme uses a multi-frequency mode signal separator to perform discrete cosine transform on the spatiotemporal feature map to obtain an initial multi-band signal, which includes n signal components {X1, ..., X...} n}

[0043] As the third part, the signal refinement network performs multi-scale feature extraction on each of the n signal components, generating n high-dimensional intrinsic mode signals. Please refer to [link to relevant documentation]. Figure 1The signal refinement network comprises four sequentially connected refinement modules, each with a different scale but identical internal structure. Each refinement module includes a convolutional layer, a temporal multi-scale perceptron, an activation layer, a batch normalization (BN) layer, and a pooling layer s, all connected sequentially. The refinement modules at different scales are used to progressively capture the temporal features of the signal component X, thereby refining the network to output the high-dimensional intrinsic mode signals corresponding to each signal component X.

[0044] The refining module consists of a time-domain multi-scale perceptron and other network layers. Its main purpose is to extract effective components from the waveform of the signal components, filter out noise signal features to refine the features, and finally, the signal refining network outputs a high-dimensional intrinsic mode signal. The design of the time-domain multi-scale perceptron in the refining module is based on the need to capture the time-domain features of the signal at multiple frequencies. Although traditional standard-scale filters can effectively extract neighborhood waveform features, the final waveform simulation requires a consistent understanding of the entire frequency range. Since the time-domain features of the signal components at various frequencies provide numerous clues that can be used to determine the waveform trend, a time-domain multi-scale perceptron was designed to capture these features.

[0045] The time-domain multi-scale perceptron includes a dimension transformation module and an extraction module. The dimension transformation module is used to process each input signal component. (W, H, and C represent the width, height, and number of channels of the feature, respectively) are compressed and transformed to form set up This reduces the network redundancy parameters of the multi-scale perceptron.

[0046] The extraction module includes a multi-scale filter for X U Adaptive extraction of refined features X R The refining feature X R Represented as: The basic principle of the extraction module is to design gates to control the flow of information from filters of different scales. These filters, carrying the signal features obtained by the multi-scale perceptron, will be passed to the next layer of network neurons.

[0047] In traditional computer vision tasks, 3×3 convolutional kernels are used. However, the task of this scheme is to extract temporal features. Therefore, in order to prevent neurons from extracting too many spatial features and interfering with the primary task, the size of the convolutional kernels designed in the temporal multi-scale perceptron is k×1, where k is different in filters of different scales.

[0048] As the fourth part, the signal reconstruction network reconstructs the waveforms of each of the n high-dimensional intrinsic mode signals to obtain the corresponding n high-dimensional signals. Please refer to [link to documentation]. Figure 1The signal reconstruction network consists of three sequentially connected reconstruction models. These three reconstruction modules have different scales but the same internal structure. Each reconstruction module includes a sequentially connected spectrum-based self-attention mechanism layer, a convolutional layer, a batch normalization (BN) layer, and an activation layer.

[0049] The spectrum-based self-attention mechanism layer filters features from a time-domain multi-scale perceptron. It assumes that within a 15-second waveform, the frequency domain features at different time points should be similar. However, if the waveform features at a certain time point show significant changes compared to other time points, it can be determined that a large noise signal exists in that time period. Through this mechanism, the self-attention weights of each filter can be obtained. These weights are then placed into the self-attention mechanism layer to globally group the features of each filter, further reducing pulse wave noise.

[0050] The spectrum-based self-attention mechanism layer can be divided into three processes: (1) After segmenting the high-dimensional intrinsic mode signal, perform discrete cosine transform to obtain the frequency characteristics of the high-dimensional intrinsic mode signal; (2) According to the frequency characteristics, use a 1×1 convolution kernel to splice the high-dimensional intrinsic mode signal of the discrete cosine transform into a one-dimensional feature signal, and extract the self-attention weights of each time period; (3) After aggregating the self-attention weights with the high-dimensional intrinsic mode signal, reorganize the one-dimensional feature signal to improve the signal-to-noise ratio.

[0051] As can be seen, the purpose of the signal reconstruction network is to reconstruct signal features. It uses a deconvolutional network that focuses on the time domain. At the same time, in order to prevent the loss of important information, the spectrum-based self-attention mechanism layer is fused with the features of the signal refinement network at the same resolution before signal reconstruction at each resolution. The main purpose of this design is to improve the SNR of the reconstructed signal and help the network to more accurately locate the position of each peak and trough of the signal.

[0052] As the fifth part, the multi-frequency modal signal fusion unit adaptively selects effective components from n high-dimensional signals for fusion to determine the positions of peaks and troughs, thereby obtaining the blood volume pulse wave signal.

[0053] Example 2:

[0054] Based on the technical solution of Embodiment 1, this embodiment conducts experiments and verifications. Three common physiological signals—heart rate (HR), heart rate variability (HRV), and respiratory rate (RF)—were tested on three publicly available datasets: ASPD2, UBFC-rPPG, and MMSE-HR.

[0055] (1) Introduction to datasets, metrics and tasks

[0056] ASPD2 contains 2166 RGB videos collected from 2166 individuals, all at a frame rate of 30fps. The dataset was used for training, and cross-dataset testing and ablation studies were conducted using MMSEHR. The metrics used included standard deviation (MAE), mean absolute error (MAE), root mean square error (RMSE), and Pearson correlation coefficient (r).

[0057] UBFC-rPPG is a challenging remote physiological measurement dataset with both sunlight and indoor lighting conditions. It contains 42 RGB videos, all captured at 30 frames per second using a Logitech Cq20 HD Pro webcam, along with corresponding ground truth BVP signals acquired using a CMS50E. UBFC-rPPG was used for heart rate and heart rate variability testing. For the heart rate estimation task, MAE, Std, RMSE, and r are reported. For the heart rate variability and respiratory rate estimation tasks, Std, RMSE, and r for low-frequency (LF), high-frequency (HF), and LF / HF ratios are reported.

[0058] MMSE-HR is a large-scale remote heart rate test dataset containing 102 videos from 40 individuals. The videos, which include various facial expressions and head movements, were recorded at 25fps. MMSE-HR has been used for cross-dataset validation and ablation studies, using Std, RMSE, and r as evaluation metrics.

[0059] (2) In-dataset testing

[0060] The mean heart rate (HR) estimation of UBFC-rPPG was evaluated, comparing traditional handcrafted methods (SAMC, POS, CHROM) with deep learning-based methods (SynRhythm, PulseGAN, Dual-GAN). The results of these commonly used methods were directly extracted from Dual-GAN. Table 1 shows the experimental results of traditional methods and our device (Proposed-15s, Proposed-30s).

[0061] Table 1

[0062] Method MAE↓ RMSE↓ MER↓ r↓ POS 8.35 10.00 9.85% 0.24 CHROM 8.20 9.92 9.17% 0.27 Green 6.01 7.87 6.48% 0.29 SynRhythm 5.59 6.82 5.5% 0.72 PulseGAN 1.19 2.10 1.24% 0.98 DualGAN 0.44 0.67 0.42% 0.99 Proposed-15s 1.53 2.04 1.57% 0.98 Proposed-30s 0.75 1.12 0.83% 0.99

[0063] Traditional methods process multiple segments of 30-second video data and then average the results. Our proposed device uses both a 30-second video and the first 15 seconds of video to generate a BVP (Body Volume Pulse Wave) for calculating heart rate. The results show that, in a single calculation, the proposed 30-second method achieves impressive results: RMSE 1.12 bpm, MAE 0.75 bpm, MER 0.83%, and r = 0.99. This surpasses all state-of-the-art traditional and deep learning methods, except for Dual-GAN, which require averaging multiple overlapping video segments. This indicates that the waveform quality of the generated blood volume pulse wave is higher, and a more accurate heart rate value can be obtained with only one calculation. On the other hand, the proposed 15-second method also achieves comparable results, second only to PulseGAN and Dual-GAN.

[0064] Since the heart rate result of this device is calculated from the pulse wave waveform, it is more difficult to calculate the accurate heart rate from a short-duration pulse wave, which further verifies the effectiveness of this device in reconstructing the BVP signal.

[0065] Heart rate variability (HRV) and respiratory rate (RF) estimations for UBFC-rPPG were evaluated. Following the Dual-GAN protocol, the first 30 participants were trained, and the remaining 12 were tested. For the tasks of HRV and RF estimation, several traditional methods (including POS, CHROM, Green, CVD, and DualGAN) were compared, with results obtained from CVD.

[0066] The estimation results of HRV and RF are shown in Table 2. Using low frequency (LF), high frequency (HF), LF / HF and respiratory rate (RF) as evaluation indicators, it can be seen that this device is superior to all traditional methods under all measurement criteria. This shows that this device can achieve accurate recovery of blood volume pulse wave shape in a network with fewer parameters.

[0067] Table 2

[0068]

[0069] (3) Cross-database testing

[0070] In addition to in-database testing on the VIPL-HR and UBFC-rPPG datasets, our device was trained on VIPL-HR and tested on MMSE-HR. The results of our device and other traditional methods are shown in Table 3. Li2014, CHROM, SAMC, RhythmNet, and CVD are all derived from CVD. Table 3 shows that our proposed 30-second video method achieves excellent results compared to other traditional methods, indicating its good versatility under unconstrained conditions. Furthermore, our proposed 15-second video method achieves a Std of 6.23 and an RMSE of 6.36, second only to CVD and significantly outperforming other traditional methods, demonstrating its ability to effectively calculate heart rate within a short timeframe.

[0071] Table 3

[0072] Method Std↓ RMSE↓ r↑ Li2014

[41] 20.02 19.95 0.38 CHROM

[36] 14.08 13.97 0.55 SAMC

[37] 12.24 11.37 0.71 RhythmNet

[24] 6.98 7.33 0.78 CVD

[27] 6.06 6.04 0.84 Proposed-15s 6.23 6.36 0.84 Proposed-30s 5.61 5.75 0.90

[0073] (4) Ablation Research

[0074] To verify the effectiveness of oversampling (OSS), the network was trained using both oversampling and normal sampling. The results show that when all training set samples are placed in one epoch for network training, the 30-second heart rate estimation results are worse than when using the oversampling strategy.

[0075] The MAE decreased from 3.36 bpm to 4.65 bpm, and the RMSE decreased from 9.9 bpm to 7.5 bpm, indicating that different heart rate intervals have similar temporal and spatial characteristics. Training the network with these intervals equally will help improve the network's generalization performance and noise reduction capabilities across the entire heart rate range.

[0076] The effectiveness of the Multi-Mode Signal Fusion Unit (MFMSF) was further evaluated, and experiments were conducted without the MFMSF. Table 4 shows that when the device lost the MFMSF, the results plummeted: MAE decreased from 3.36 bpm to 4.65 bpm, and Std and RMSE decreased by 4.36 bpm and 4.46 bpm, respectively. This once again verifies the importance of the proposed MFMSF for improving network computing power and robustness.

[0077] Table 4

[0078]

[0079] The effectiveness of the Temporal Multiscale Perceptron (TMSCB) was further evaluated. The results in Table 4 show that replacing the traditional residual convolutional module with the TMSCB further enhanced the network's robustness, reducing RMSE from 5.95 bpm to 5.75 bpm and Std from 5.95 bpm to 5.61 bpm. This verifies that the multiscale sensing mechanism of the TMSCB can significantly improve the noise filtering ability of physiological tasks.

[0080] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A blood pressure measurement device based on multi-frequency mode fusion, characterized in that: include: The data preprocessing module is used to extract spatiotemporal feature maps from face videos; A multi-frequency mode signal separator is used to decompose a spatiotemporal feature map into an initial multi-frequency band signal, wherein the initial multi-frequency band signal includes n signal components; The signal refinement network is used to extract multi-scale features from n signal components, thereby generating n high-dimensional intrinsic mode signals. The signal refinement network includes four sequentially connected refinement modules of different scales but identical structure. Each refinement module includes a sequentially connected convolutional layer, a temporal multi-scale perceptron, an activation layer, and a batch normalization (BN) layer. Each refinement module of different scales is used to progressively capture the temporal features of the signal component X, thereby generating the corresponding high-dimensional intrinsic mode signal. The time-domain multi-scale perceptron includes a dimension transformation module and an extraction module, wherein... The dimension transformation module is used to transform each input signal component. Perform compression transformation to form Where W, H, and C represent the width, height, and number of channels of the feature, respectively. This reduces the network redundancy parameters of the multi-scale perceptron. The extraction module includes a multi-scale filter for... Adaptive extraction of refined features The refining features Represented as: ; A signal reconstruction network is used to reconstruct the waveforms of n high-dimensional intrinsic mode signals, thereby obtaining n high-dimensional signals. A multi-frequency modal signal fusion unit is used to select effective components from n high-dimensional signals, fuse them, and then determine the positions of peaks and troughs to obtain the blood volume pulse wave signal.

2. The blood pressure measurement device based on multi-frequency mode fusion according to claim 1, characterized in that: The data preprocessing module includes a face video recognition engine and a spatiotemporal feature acquisition module, wherein... The face video recognition engine is used to acquire face videos and perform landmark detection on the face video in frame t using MediaPipe. The face region is then segmented using the landmark to obtain a set of average pixel values ​​from M regions, forming the initial region features of frame t. , , T represents a total of T frames; The spatiotemporal feature acquisition module is used to integrate the initial region features of the T-frame to obtain the initial signal features. And by performing overall time-domain normalization on the initial signal features, the normalized signal features are obtained: ; And the characteristics of the normalized signal After adding white noise for enhancement, a spatiotemporal feature map is generated.

3. The blood pressure measurement device based on multi-frequency mode fusion according to claim 2, characterized in that: The signal reconstruction network includes three reconstruction modules with different scales but the same structure that are connected in sequence. Each reconstruction module includes a spectrum-based self-attention mechanism layer, a convolutional layer, a batch normalization (BN) layer, an activation layer, and a pooling layer that are connected in sequence. The spectrum-based self-attention mechanism layer is used to perform discrete cosine transform on the high-dimensional intrinsic mode signal to obtain the frequency characteristics of the high-dimensional intrinsic mode signal; then, based on the frequency characteristics, the high-dimensional intrinsic mode signal of the discrete cosine transform is spliced ​​into a one-dimensional feature signal, and self-attention weights are extracted; finally, the self-attention weights are aggregated with the high-dimensional intrinsic mode signal, and the one-dimensional feature signal is recombined to improve the signal-to-noise ratio.

Citation Information

Patent Citations

  • Non-contact blood pressure measuring equipment based on face video

    CN113827208A