Non-contact blood pressure prediction method based on multi-modal fusion
Through multimodal fusion technology, the feature extraction and dynamic weighted fusion of visible light and infrared video are used to solve the stability problem of contactless blood pressure monitoring during light changes, and high-precision blood pressure prediction in low-light or light-free environments are achieved.
Patent Information
- Application Number
- CN202510581074.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-04
AI Technical Summary
The existing non-contact blood pressure monitoring technology is difficult to obtain stable and clear long-range photoelectric capacity pulse wave signals in low-light, backlight or luminous dramatic scenarios, resulting in an increase in blood pressure prediction error and imaging failure in light-free environments, limiting its scope of application.
The multimodal fusion method is adopted, combining visible light video and infrared video, and multi-scale and multi-channel depth features are extracted through the feature extraction module, and the dynamic weighted fusion mechanism of L1/L2 mixed norm and Softmax normalization is used to generate adaptive weights, fuse it into a unified feature vector, and finally rPPG signal fitting and reconstruction are performed through the blood pressure prediction module.
RPPG signal is robustly extracted when the light conditions are extremely changing, maintaining the system's availability in low-light or light-free environments, significantly improving the robustness and scope of contactless blood pressure monitoring, and solving the problem of light sensitivity.
Smart Images

Figure CN120241016A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of non-contact blood pressure monitoring, and relates to a non-contact blood pressure prediction method based on multimodal fusion. Background Art
[0002] In the field of non-contact blood pressure prediction, existing methods mostly rely on a single-modal video signal (usually visible light video) to extract the remote photoplethysmogram (rPPG) signal and then predict blood pressure. However, due to the high sensitivity of visible light imaging to skin texture, skin color differences, and the intensity and spectral distribution of ambient light, single-modal methods often struggle to obtain stable and clear rPPG waveforms in low-light, backlight, or rapidly changing light scenarios, leading to an increase in blood pressure prediction errors and difficulty in meeting the accuracy requirements of clinical and daily monitoring. In addition, in completely lightless or extremely low-light environments, visible light imaging almost fails, further limiting the applicable range of non-contact blood pressure monitoring technology. Summary of the Invention
[0003] In order to solve the above technical problems existing in the prior art, the present invention proposes a non-contact blood pressure prediction method based on multimodal fusion, and its specific technical solution is as follows: A non-contact blood pressure prediction method based on multimodal fusion: Preprocess the dynamically captured visible light video and infrared video sequences; Through the feature extraction module, respectively extract features from the preprocessed visible light and infrared light images; Through the feature fusion module, perform real-time weight optimization on the feature correlation between visible light and near-infrared images in the channel and spatial dimensions, and fuse them into a unified feature vector according to the complementary principle; Through the blood pressure prediction module, fit and reconstruct the rPPG signal from the fused features to obtain the blood pressure prediction value.
[0004] Further, the preprocessing includes spatio-temporal alignment and face region cropping.
[0005] Further, the feature extraction module is based on the swin-Transformer model and has 4 layers of MS-SwinBlock extraction blocks, and the number of MS-Swin Block extraction blocks in each layer is 2, 4, 4, and 2 respectively.
[0006] Furthermore, the MS-Swin Block extraction block includes a multi-scale CNN block, layer normalization, a window-based multi-head self-attention W-MSA layer, a shifted-window-based multi-head self-attention SW-MSA layer, and a multi-layer perceptron MLP; in each layer structure, the W-MSA layer is adopted in the first MS-Swin Block extraction block, while the SW-MSA layer is adopted in subsequent blocks, forming an alternating stacking pattern.
[0007] Furthermore, different dilation rates are configured in the multi-scale CNN block of the MS-Swin Block extraction block.
[0008] Furthermore, the feature fusion module uses a dynamic weighted fusion mechanism combining L1 / L2 mixed norm evaluation and Softmax normalization to generate adaptive weights for visible light image features and infrared light image features in the spatial-channel dimension, and fuses the visible light image features and infrared light image features into a unified feature vector according to the complementary principle.
[0009] Furthermore, the spatial-channel weights of the visible light image features are obtained through normalization, while the spatial-channel weights of the infrared light image features are obtained through complementary calculation.
[0010] Furthermore, the blood pressure prediction module includes an encoder and a multi-scale CNN structure containing three basic convolutional kernels of 1×1, 3×3, and 5×5. It gradually extracts the signal time-domain features through hierarchical stacked convolutional operations, dynamically weights the key time node information using the attention mechanism, and finally obtains the blood pressure prediction value through mapping by a fully connected layer.
[0011] Furthermore, the attention mechanism adopts a Top-K sparse attention mechanism.
[0012] Furthermore, no dilation rate is set in the multi-scale CNN structure of the blood pressure prediction module.
[0013] Beneficial effects: The present invention breaks through the limitations of traditional non-contact blood pressure monitoring technologies. It can not only robustly extract rPPG signals under extremely changing lighting conditions, but also maintain the continuous availability of the system relying on infrared imaging in low-light or no-light environments, thus significantly improving the robustness and application scope of non-contact blood pressure monitoring technologies. Its collaborative analysis of visible light and infrared light videos successfully solves the sensitivity problem of existing face-video-based methods to lighting conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is an overall process model diagram of a non-contact blood pressure prediction method based on multi-modal fusion in this embodiment; Figure 2It is the structure diagram of the MS-Swin Block extraction block of the feature extraction module in this embodiment; Figure 3 It is the structure diagram of the blood pressure prediction module in this embodiment. Specific implementation manners
[0015] In order to make the objectives, technical solutions and technical effects of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings of the specification and embodiments.
[0016] As Figure 1 shown, a non-contact blood pressure prediction method based on multi-modal fusion in this embodiment first performs preprocessing such as spatio-temporal alignment and face region cropping on dynamically captured visible light video and infrared video sequences to ensure the spatial consistency and temporal synchronization of bimodal data. Subsequently, through an independent feature extraction module, multi-scale and multi-channel deep feature representations are obtained from visible light images and infrared light images respectively. A dynamic weighted fusion mechanism combining L1 / L2 hybrid norm evaluation and Softmax normalization is used to generate adaptive weights for the two modalities in the spatial-channel dimension, and they are fused into a unified feature vector according to the complementary principle. Finally, based on the fused features, the rPPG signal is fitted and reconstructed, and the reconstructed pulse wave signal is input into the terminal blood pressure prediction network to achieve high-precision non-contact estimation of systolic blood pressure and diastolic blood pressure.
[0017] The feature extraction module integrates the sliding window mechanism of the swin-Transformer model with multi-scale CNN to effectively capture local and detailed features related to the rPPG signal, thereby realizing rPPG signal prediction based on video frames. The feature extraction process in this embodiment is divided into four layers, and each layer contains 2, 4, 4, and 2 extraction blocks respectively.
[0018] As Figure 2 shown, the main components of the extraction block include a multi-scale CNN block, layer normalization, a window-based multi-head self-attention (W-MSA) layer, a shifted window-based multi-head self-attention (SW-MSA) layer, and a multi-layer perceptron (MLP). In each hierarchical structure, the first block adopts the W-MSA layer, while the subsequent blocks adopt the SW-MSA layer, forming an alternating stacking pattern. In the multi-scale CNN, by configuring different dilation rates, such as r = 3, 5, 7, a wider range of context information is captured, making it suitable for global feature modeling of input data.
[0019] Traditional multimodal fusion methods typically use simple weighted average operations to generate static weighted graphs for data fusion. However, these methods often fail to effectively capture the complex interactions and cross-modal correlations between different modalities, resulting in insufficient utilization of complementary information in multimodal data. To address this limitation, the fusion module of the present invention significantly improves the robustness and accuracy of the fusion process through a dynamic weight evaluation and adjustment mechanism.
[0020] The core innovation of this fusion module lies in dynamically modeling the feature correlations between different modalities such as visible light and near-infrared, and performing real-time weight optimization in both channel and spatial dimensions. The specific process is as follows: First, RGB and NIR image data are processed to form feature representations.
[0021] Subsequently, a weighted combination of L1 norm and L2 norm is used to measure the absolute importance of each modality's feature vector, and it is combined with the Softmax function to generate dynamic weights in the spatial-channel dimension. Specifically, the spatial-channel weights of the RGB mode are obtained through normalization, while those of the NIR mode are obtained through complementary calculation. This dynamic weighting mechanism can adaptively adjust the fusion ratio according to the characteristics of the input data. For example, it can automatically enhance the contribution of the infrared mode in scenes with drastic lighting changes, or increase the weight of the RGB mode in scenes with rich texture details.
[0022] Since the rPPG signal is essentially one-dimensional time data, the decoder part is removed from the Transformer architecture, and a multi-scale CNN structure and encoder are used for feature extraction, combined with a simple output layer to achieve efficient rPPG signal processing. Among them, by introducing the Top-K sparse attention mechanism, the model's ability to model local features is enhanced, improving the computational efficiency and reducing complexity.
[0023] Specifically, as Figure 3 shown, in the blood pressure prediction module, in order to focus on the fine-grained changes and local feature expressions of the data, the dilation rate is not set, and three basic convolutional kernels of 1×1, 3×3, and 5×5 are used for feature capture. The time-domain features of the signal are gradually extracted through hierarchical stacked convolutional operations, and the attention mechanism is used to dynamically weight the information of key time nodes. Finally, the blood pressure prediction value is obtained through mapping by a fully connected layer. This architecture effectively avoids the computational redundancy problem of traditional Transformer models in long sequence processing, while ensuring the time modeling ability.
[0024] In this embodiment, the experimental results are compared with representative studies. As shown in the results of Table 1 below, the method of the present invention is significantly superior to the existing SOTA methods.
[0025] Table 1:
[0026] As described above, it is only the preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Although the implementation process of the present invention has been described in detail above, for those familiar with the art, they can still modify the technical solutions recorded in the foregoing examples or make equivalent replacements for some of the technical features. Any modifications, equivalent replacements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A non-contact blood pressure prediction method based on multi-modal fusion, characterized in that: Preprocess the dynamically captured visible light video and infrared video sequences; Through the feature extraction module, extract features from the preprocessed visible light and infrared light images respectively; Through the feature fusion module, optimize the feature correlation between visible light and near-infrared images in real time in the channel and spatial dimensions, and fuse them into a unified feature vector according to the complementary principle; Through the blood pressure prediction module, fit and reconstruct the rPPG signal from the fused features to obtain the blood pressure prediction value.
2. The non-contact blood pressure prediction method according to claim 1, wherein: The preprocessing includes spatio-temporal alignment and face region cropping.
3. The non-contact blood pressure prediction method according to claim 1, characterized in that: The feature extraction module is based on the swin-Transformer model and has 4 layers of MS-Swin Block extraction blocks, and the number of MS-Swin Block extraction blocks in each layer is 2, 4, 4, and 2 respectively.
4. The non-contact blood pressure prediction method according to claim 3, wherein: The MS-Swin Block extraction block includes a multi-scale CNN block, layer normalization, a window-based multi-head self-attention W-MSA layer, a shifted-window-based multi-head self-attention SW-MSA layer, and a multi-layer perceptron MLP; In each layer structure, the W-MSA layer is used in the first MS-Swin Block extraction block, and the SW-MSA layer is used in the subsequent blocks, forming an alternating stacking pattern.
5. The non-contact blood pressure prediction method according to claim 4, characterized in that: Configure different dilation rates in the multi-scale CNN block.
6. The non-contact blood pressure prediction method according to claim 1, characterized in that: The feature fusion module uses a dynamic weighted fusion mechanism combining L1 / L2 mixed norm evaluation and Softmax normalization to generate adaptive weights for visible light image features and infrared light image features in the spatial-channel dimension, and fuse the visible light image features and infrared light image features into a unified feature vector according to the complementary principle.
7. The non-contact blood pressure prediction method according to claim 6, wherein: The spatial-channel weights of the visible light image features are obtained by normalization, while the spatial-channel weights of the infrared light image features are obtained by complementary calculation.
8. The non-contact blood pressure prediction method according to claim 1, characterized in that: The blood pressure prediction module includes an encoder and a multi-scale CNN structure containing three basic convolutional kernels of 1×1, 3×3, and 5×5. It gradually extracts the signal time-domain features through hierarchical stacked convolutional operations, uses the attention mechanism to dynamically weight the key time node information, and finally obtains the blood pressure prediction value through a fully connected layer mapping.
9. The non-contact blood pressure prediction method according to claim 8, characterized in that: The attention mechanism adopts the Top-K sparse attention mechanism.
10. The non-contact blood pressure prediction method according to claim 8, characterized in that: The multi-scale CNN structure in the blood pressure prediction module does not set a dilation rate.
Citation Information
Cited By
Multi-scale wireless channel propagation loss prediction method based on finite element method
CN120785448A
Cardiovascular function related parameter non-contact estimation method and system fusing and applying RGB, multiband infrared and thermal infrared videos
CN122336646A