Multi-site rPPG fusion physiological signal monitoring method, device and storage medium
Through the multi-part rPPG fusion physiological signal monitoring method, the superconvolution neural network and asynchronous distillation technology are used to dynamically adjust the attention weight, solving the monitoring difficulties of traditional rPPG technology when the face and upper body are blocked, achieving high-precision physiological signal monitoring and robustness enhancement.
Patent Information
- Application Number
- CN202411775406.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2044-12-05
AI Technical Summary
Traditional rPPG technology relies on facial signals and cannot work when the patient's head and upper body are covered. Light changes affect signal accuracy. Single-part monitoring is insufficient and the overall physiological status cannot be fully monitored.
The multi-part rPPG fusion physiological signal monitoring method is adopted. By collecting multi-part skin videos, combining superconvolution neural networks and asynchronous distillation technology, attention weights are dynamically adjusted, rPPG pulse waveform features are extracted and fused, comprehensive signals are generated, and complex environments are adapted to.
When the face and upper body are blocked, high-precision multi-part physiological signal monitoring is achieved, enhancing the robustness and anti-interference ability of the model, ensuring the stable output of physiological signals.
Smart Images

Figure CN119600371B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer deep learning technology, and specifically relates to a multi-site rPPG fusion physiological signal monitoring method, device and storage medium. Background Art
[0002] Remote photoplethysmography (rPPG) accurately captures changes in blood volume by monitoring tiny fluctuations in light reflected from the skin, providing physiological indicators such as heart rate, respiratory rate, and blood oxygen saturation. However, traditional methods typically extract rPPG signals from the patient's face, which is easily affected by facial expressions, head movement, and lighting conditions, and therefore has limitations. Existing rPPG technology has the following drawbacks:
[0003] 1. Dependence on facial signals: When a patient enters a medical device such as a CT scan, the patient's head and upper body are covered. Since traditional rPPG technology relies on facial signals, it cannot work properly in such scenarios.
[0004] 2. Impact of light changes: The lighting environment in medical equipment is complex. Traditional rPPG technology is prone to noise under strong or weak light conditions, which affects the accuracy of the rPPG signal.
[0005] 3. Insufficient single-site monitoring: Traditional rPPG technology relies on the extraction of rPPG signals from a single site (such as the face) and cannot obtain physiological signals from multiple sites simultaneously, which also limits the comprehensive monitoring of the overall physiological state. Summary of the Invention
[0006] In order to solve the problems existing in the prior art, the present invention provides a multi-site rPPG fusion physiological signal monitoring method, device and storage medium.
[0007] In order to solve the above technical problems and achieve the above technical effects, the present invention is implemented through the following technical solutions:
[0008] A multi-site rPPG fusion physiological signal monitoring method, comprising:
[0009] Step 1) Collect skin videos of multiple parts of the current patient;
[0010] Step 2) Preprocessing the collected skin videos of multiple parts to eliminate environmental noise and interference;
[0011] Step 3) editing the physiological region prompt information (BioRegionPrompt) for indicating the source region of the collected multi-region skin video signals;
[0012] Step 4) The pre-processed multi-site skin video and the edited physiological site prompt information are used as input data and fed into a multi-site rPPG fusion physiological signal extraction model trained based on an improved hyperconvolutional neural network and asynchronous distillation technology for processing;
[0013] The multi-site rPPG fusion physiological signal extraction model adjusts the attention weight of skin videos of different sites based on the input physiological site prompt information, dynamically optimizes the feature extraction capability of specific areas, extracts the rPPG pulse waveform features of the skin video space of each site, and combines time series analysis to ensure the spatiotemporal integrity and stability of the rPPG pulse waveform features. Finally, the rPPG pulse waveform features extracted from each site are weighted and fused to generate a comprehensive rPPG waveform signal.
[0014] Step 5) The extracted integrated rPPG pulse waveform signal is post-processed, and the physiological signal values including heart rate and respiratory rate are calculated through frequency domain analysis methods (such as fast Fourier transform FFT), and a comprehensive report containing these physiological data is output.
[0015] Furthermore, the multi-site skin video is a skin video obtained by capturing multiple different parts of the patient (such as limbs or other skin areas) through a camera device including a near-infrared camera or a high-frame-rate RGB camera. The multi-site skin video can reflect subtle changes in light reflected by the skin of multiple different parts of the patient, so as to capture blood flow information.
[0016] Furthermore, the preprocessing focuses on the segmentation and enhancement of skin areas in multi-site skin videos to ensure the stability and accuracy of subsequent rPPG pulse waveform feature extraction.
[0017] The skin region segmentation process is to extract a mask including the face and limbs.
[0018] Skin area enhancement processing is to improve the signal quality of multiple skin parts videos by filtering and removing artifacts through the use of image processing technology; including using bandpass filters to remove high-frequency noise and low-frequency noise and retaining the frequency range related to the pulse; or using smoothing filtering and artifact removal technology to eliminate motion artifacts to improve signal reliability.
[0019] Furthermore, the physiological part prompt information is a physiological part code generated by editing the physiological part prompt word input by the user using one-hot encoding, or is a vector corresponding to each of the multiple part visual features converted by the CLIP model into the physiological part prompt word input by the user;
[0020] The physiological part prompt words are prompt words with part keywords (such as face, left arm, right arm, left hand, right hand, left leg, right leg, left foot, right foot, etc.) manually input by the user based on the characteristics of different body parts in the multi-part skin video collected;
[0021] Each of the physiological part codes or the vectors corresponds to a separate or combined physiological part, and is used to indicate the source parts of the multi-part skin video signals and to guide the attention of these parts by weighting.
[0022] Furthermore, the architecture of the multi-site rPPG fusion physiological signal extraction model includes an input layer, a feature extraction module, a physiological site prompt module, a feature weighting module and a dense layer.
[0023] The input layer is responsible for receiving the collected skin videos of multiple parts and performing preliminary standardization processing. The skin videos of multiple parts are skin videos obtained by shooting multiple different parts of the patient using a camera device; it is also responsible for receiving the physiological part prompt words input by the user.
[0024] The feature extraction module consists of multiple convolutional layers, a superconvolution module, a time translation module, and a dynamic attention layer. It is responsible for extracting the rPPG pulse waveform features of each part from the skin video of multiple parts, and obtaining a comprehensive rPPG waveform signal after fusion and weighting for subsequent physiological signal analysis.
[0025] The multi-layer convolutional layer (ConvLayer) is the core of the model, responsible for extracting spatial features from multiple skin videos. These spatial features reflect local changes in the skin's reflected light. Each convolutional layer uses multiple convolution kernels to scan multiple skin videos to extract low-level features such as edges and textures. By superimposing multiple layers of convolution, the model can gradually capture high-level features in multiple skin videos, including characteristic changes in pulse waveforms.
[0026] The hyper-conv module is responsible for dynamically adjusting the size and receptive field of the convolution kernel according to the body part source of the multi-part skin video, so as to improve the model's ability to capture the details of the multi-part skin video during the feature extraction process, ensuring that the model can efficiently extract the unique signal characteristics of each part of the skin video.
[0027] The time shift module (TSM) is responsible for injecting the time series of multi-site skin videos into the multi-layer convolutional layer, enabling the model to capture the temporal features in the multi-site skin videos. It is used to capture the dynamic fluctuations of the multi-site skin videos and perform time series analysis on the spatial features, thereby obtaining stable and accurate spatiotemporal features, namely rPPG pulse waveform features. The introduction of the time series ensures that the model has better processing capabilities for the time-dependent characteristics of multi-site skin videos, enabling the model to better capture the changing patterns of the pulse waveform at different time points.
[0028] The dynamic attention layer is responsible for automatically adjusting the model's attention to different areas in the multi-part skin video based on the characteristics of the multi-part skin video, thereby enhancing the model's attention to high-quality signal areas and reducing the weight of areas with greater noise.
[0029] The physiological part prompt module (BioRegionPromptModule) is responsible for converting the physiological part prompt words (such as left arm, right arm, left hand, right hand, left leg, right leg, left foot, right foot, etc.) input by the user into vectors that match the visual features of multiple parts through the CLIP model, or using one-hot encoding to edit the physiological part prompt words input by the user into physiological part codes, and passing the vectors or the physiological part codes as input to the super-convolution module to provide a basis for the super-convolution module to identify the body part source of the input multi-part skin video, and help the super-convolution module dynamically adjust the convolution kernel size and receptive field, thereby enhancing the model's attention to specific body parts.
[0030] The feature weighting module is responsible for fusing and weighting the rPPG pulse waveform features of various parts extracted by the feature extraction module according to the weighting mechanism through the weighted gate layer to ensure that the model can effectively fuse the rPPG pulse waveform features of multiple parts and adjust the weighted proportion according to the quality of the rPPG pulse waveform features.
[0031] The weighting mechanism weights the rPPG pulse waveform features of each part according to the quality and time series of the rPPG pulse waveform features of each part. That is, rPPG pulse waveform features that are relatively stable and have higher signal quality will occupy a larger proportion in the model output, while rPPG pulse waveform features with greater noise or instability will occupy a smaller proportion in the model output.
[0032] The dense layer is responsible for weighting and nonlinear mapping the feature vector output by the feature weighting module, integrating the multi-dimensional features into a compact representation, namely, the comprehensive rPPG waveform signal.
[0033] Furthermore, the training method of the multi-site rPPG fusion physiological signal extraction model is:
[0034] 1) Collect a certain number of skin videos of multiple sites in a clinical setting from a certain number of past patients, and simultaneously obtain the corresponding true values of PPG pulse waveform signals and true values of physiological parameters including heart rate, blood oxygen, and respiratory rate, to form an rPPG dataset that contains the temporally and spatially aligned skin videos of multiple sites, the true values of PPG pulse waveform signals, and the true values of physiological parameters;
[0035] 2) Using the constructed rPPG signal dataset, asynchronous distillation technology is employed in the distillation network to train the multi-site rPPG fusion physiological signal extraction model. The model training process adopts a staged guidance approach, with the goal of enhancing the model's ability to extract features from multi-site skin videos. During training, the multi-site skin videos are used as input, the true values of the PPG pulse waveform signals are used as the primary labels, and the true values of physiological parameters are used as supplementary labels. The model learns the mapping relationship between the features of the video frames of the multi-site skin videos and the corresponding true values of the PPG pulse waveforms. This enables the model to accurately extract the rPPG pulse waveform features of each site based on the input multi-site rPPG signals, and output a comprehensive rPPG waveform signal after fusion and weighting.
[0036] 3) Use a specific loss function to optimize the model parameters.
[0037] Furthermore, the asynchronous distillation technique decouples the optimization process of different body part information and dynamically adjusts the learning weights of each body part signal, enabling efficient extraction of signal features from each body part. Signals from all body parts are considered equally important, and the model adjusts the optimization process based on the physiological body part prompts, ensuring that signal features from all body parts are fully learned while also supporting the independent prediction and fusion of multi-body part signals. Facial signals, due to their strong reflective properties and large data volume, provide more guidance in model training. However, since facial signals are obscured in scenes such as CT equipment, the asynchronous distillation technique enhances the learning and optimization of signals from body parts other than the face, enabling the model to independently and accurately predict physiological signal features from these parts even when facial signals are missing. During the multi-body part signal fusion process, the model dynamically adjusts the weighting of each body part signal to ensure that the output physiological signal features from the fusion process remain consistent or synchronized with the facial signals in terms of timing and amplitude. Features of "non-target" parts are asynchronously optimized with weights from a different cycle, enabling the student model to maintain a certain understanding of the features of non-target parts without interfering with the primary target features, thereby achieving efficient separation and optimization of multi-body part information.
[0038] Furthermore, the loss function includes mean absolute error (MAE), root mean square error (RMSE) and signal-to-noise ratio (SNR); wherein,
[0039] The mean absolute error is used to evaluate the mean absolute difference between the predicted value and the true value of the rPPG signal;
[0040] The root mean square error is used to evaluate the root mean square of the square error between the predicted value and the true value of the rPPG signal;
[0041] The signal-to-noise ratio is used to evaluate the ratio of signal intensity to noise intensity in the rPPG signal;
[0042] In a CT equipment environment, the mean absolute error, the root mean square error, and the signal-to-noise ratio are used to calculate the error for signals of specific parts other than the face, and finally the total error is weighted to improve the generalization ability of the overall model and its ability to adapt to complex environments.
[0043] Furthermore, an adaptive algorithm based on illumination changes is used to post-process the integrated rPPG pulse waveform signal to correct the signal distortion caused by changes in light intensity. Specifically, according to the extracted brightness value, cross-attention is performed on the real and imaginary parts respectively, and finally an inverse fast Fourier transform is performed to convert it into a brightness-adapted rPPG waveform signal output.
[0044] A computer device comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus, and the memory is used to store at least one executable instruction, wherein the executable instruction causes the processor to perform operations corresponding to the above-mentioned multi-site rPPG fusion physiological signal monitoring method.
[0045] A computer-readable storage medium stores at least one executable instruction, which enables a processor to perform operations corresponding to the above-mentioned multi-site rPPG fusion physiological signal monitoring method.
[0046] The beneficial effects of the present invention are:
[0047] This invention proposes a physiological signal monitoring technology that integrates rPPG signals from multiple parts of the body. This technology is based on remote photoplethysmography (rPPG) signal detection and aims to achieve comprehensive physiological monitoring by collecting and fusing rPPG signals from multiple parts of the body, such as the limbs. It solves the problem of being unable to extract rPPG signals due to occlusion of the face and upper body, and complements real-time multi-part physiological signal monitoring.
[0048] The present invention adopts a multi-site rPPG fusion method. When the patient's face or upper body is partially covered (such as when the patient is in medical equipment such as CT), rPPG signals are extracted from exposed limbs (such as left and right arms and left and right legs, etc.), and the rPPG signals of multiple sites are fused to obtain comprehensive physiological signal data (such as heart rate, respiratory rate), ensuring high-precision physiological monitoring even when the patient's head or upper body is covered.
[0049] This paper combines hyperconvolution technology with asynchronous distillation technology, and then combines it with embedded physiological site prompts, so that the model has the ability to dynamically adapt to the rPPG signals of the limbs, thereby realizing feature extraction of rPPG signals in multiple parts, enhancing the model's robustness to complex signal environments, and effectively reducing computational costs.
[0050] The asynchronous distillation technology of this invention enables the model to effectively distinguish between target and non-target regions through feature learning, ensuring high-quality physiological signal output even in complex signal environments. This enhances the model's anti-interference ability and ensures stable physiological signal output, making it suitable for real-time physiological monitoring. Asynchronous distillation technology transfers features between the teacher and student models over different time periods to achieve more stable and accurate model training, further optimizing the model's learning performance and reducing its computational complexity.
[0051] Experimental results demonstrate that this technology can provide accurate physiological signal monitoring in multi-site rPPG signal detection. This technology is expected to provide important support for the widespread clinical application of physiological signal monitoring, while also improving the accuracy and efficiency of real-time monitoring.
[0052] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the invention and to implement it according to the contents of the description, the following preferred embodiments of the present invention are described in detail with reference to the accompanying drawings. The specific implementation methods of the present invention are given in detail by the following embodiments and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0054] Figure 1 This is a flowchart of the steps of the multi-site rPPG fusion physiological signal monitoring method of the present invention;
[0055] Figure 2 Schematic diagram of rPPG pulse waveform feature collection and fusion in the multi-site rPPG fusion physiological signal monitoring method of the present invention;
[0056] Figure 3 This is a training flow chart of the multi-site rPPG fusion physiological signal extraction model of the present invention under the distillation network framework;
[0057] Figure 4 Schematic diagram of the structure of the superconvolution module in the multi-site rPPG fusion physiological signal extraction model of the present invention;
[0058] Figure 5 Schematic diagram of the process of embedding physiological site prompt words into the super-convolution module in the multi-site rPPG fusion physiological signal extraction model of the present invention;
[0059] Figure 6 This is a sample diagram of a data set collected in an embodiment of the present invention;
[0060] Figure 7 Schematic diagram of the technical route of the multi-site rPPG fusion physiological signal extraction model of the present invention;
[0061] Figure 8 Schematic diagram of multi-site signal extraction weights in the distillation network framework of the present invention. DETAILED DESCRIPTION
[0062] The following will be described in detail with reference to the accompanying drawings to better understand the purpose, features and advantages of the invention. It should be understood that the embodiments shown in the accompanying drawings are not intended to limit the scope of the invention, but are only intended to illustrate the essential spirit of the technical solution of the invention.
[0063] In the following description, for the purpose of illustrating the various disclosed embodiments, certain specific details are set forth in order to provide a thorough understanding of the various disclosed embodiments. However, those skilled in the relevant art will recognize that the embodiments may be practiced without one or more of these specific details. In other cases, well-known devices, structures, and techniques associated with this application may not be shown or described in detail to avoid unnecessarily obscuring the description of the embodiments.
[0064] Unless the context requires otherwise, throughout the specification and claims, the word "comprise" and variations such as "include" and "have" should be construed in an open, inclusive sense, that is, should be interpreted to mean "including, but not limited to."
[0065] Reference throughout this specification to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of "in one embodiment" or "in an embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any manner in one or more embodiments.
[0066] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. It should be noted that the term "or" is generally employed in its sense including "and / or" unless the context clearly dictates otherwise.
[0067] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0068] To address the shortcomings of traditional physiological signal monitoring when the face and upper body are partially covered, as well as the accuracy and robustness of signal acquisition in complex environments, this paper proposes a comprehensive physiological monitoring method based on deep learning that fuses rPPG signals from multiple locations. When the patient's head is covered or occluded, this method extracts rPPG signals from the patient's limbs (such as the arms and legs). Combining dynamic convolution and asynchronous knowledge distillation techniques, this method ensures high-precision and robust physiological signal monitoring even under complex occlusion conditions, avoiding the limitation of being unable to acquire facial rPPG signals.
[0069] See also Figure 1 As shown, a multi-site rPPG fusion physiological signal monitoring method specifically includes the following:
[0070] Step 1) Collect skin videos of multiple parts of the current patient;
[0071] Step 2) Preprocessing the collected skin videos of multiple parts to eliminate environmental noise and interference;
[0072] Step 3) editing the physiological region prompt information (BioRegionPrompt) for indicating the source region of the collected multi-region skin video signals;
[0073] Step 4) The pre-processed multi-site skin video and the edited physiological site prompt information are used as input data and fed into a multi-site rPPG fusion physiological signal extraction model (BRPDNet) based on an improved hyperconvolutional neural network and trained using asynchronous distillation technology for processing;
[0074] The BRPDNet model uses input physiological site information to adjust the attention weights of skin videos at different locations, dynamically optimizes feature extraction capabilities for specific regions, extracts rPPG pulse waveform features from the skin video space at each location, and combines this with time series analysis to ensure the spatiotemporal integrity and stability of the rPPG pulse waveform features. Finally, it performs a weighted fusion of the rPPG pulse waveform features extracted from each location to generate a comprehensive rPPG waveform signal.
[0075] Step 5) The extracted integrated rPPG pulse waveform signal is post-processed, and the physiological signal values including heart rate and respiratory rate are calculated through frequency domain analysis methods (such as fast Fourier transform FFT). A comprehensive report containing these physiological data is then output. These physiological signal data can be fed back to medical equipment operators in real time for clinical monitoring and decision support.
[0076] See also Figure 2 As shown, the method of the present invention can provide high-precision multi-site rPPG signal detection even when a patient is inside a medical device such as a CT scan and their head or upper body is covered. This significantly enhances the real-time and reliability of physiological signal monitoring. This technology will provide strong technical support for a wide range of medical applications, including clinical diagnosis, surgical monitoring, and long-term health management.
[0077] As a preferred embodiment of the present invention, in step 1), the method for collecting rPPG signals at multiple locations is as follows:
[0078] When a patient enters a large medical device such as a CT scanner, their face and upper body are often covered. This makes traditional physiological signal monitoring methods that rely on facial rPPG signals ineffective. To address this problem, the present invention adopts an rPPG signal acquisition solution based on the limbs.
[0079] Specifically, a camera device (such as a near-infrared camera or a high-frame-rate RGB camera) is used to shoot the patient's exposed limbs to obtain video images of the patient's limb skin. These limb skin video images can reflect subtle changes in the light reflected by the patient's limb skin, so as to capture blood flow information (pulse waveform characteristics or blood pressure volume pulse signals), thereby completing the rPPG signal acquisition of the limbs, ensuring that even if the face is covered, effective physiological signals can be extracted from the limbs.
[0080] At this stage, it is crucial to ensure the stability of camera equipment configuration and signal acquisition, especially in medical environments, where the calibration of camera equipment and environmental light management will directly affect the quality of rPPG signal acquisition.
[0081] As a preferred embodiment of the present invention, in step 2), the rPPG signal must first be preprocessed after acquisition, as changes in lighting, equipment movement, or slight patient movement in the medical environment can interfere with the signal. This preprocessing focuses on segmenting and enhancing the skin regions in the multi-region skin video to ensure the stability and accuracy of the subsequent rPPG pulse waveform feature extraction.
[0082] The skin region segmentation process is to extract a mask including the face and limbs.
[0083] Skin area enhancement processing is to filter and remove artifacts of multiple skin parts videos by using image processing technology to eliminate environmental noise and interference, ensuring that the multiple skin parts videos input into the BRPDNet model have sufficient signal quality.
[0084] Preprocessing methods can use a variety of filtering techniques, such as bandpass filtering, to remove high- and low-frequency noise and retain the frequency range related to the pulse. In this step, motion artifacts can also be eliminated through smoothing filtering and artifact removal techniques to improve signal reliability.
[0085] As a preferred embodiment of the present invention, in step 3), the method for editing the physiological part prompt information is as follows:
[0086] The physiological part prompt information is a physiological part code edited by using One-Hot Encoding to edit the physiological part prompt word input by the user, or is a vector corresponding to each of the multiple part visual features converted by the CLIP model from the physiological part prompt word input by the user.
[0087] Physiological part prompt words are prompt texts with part keywords (such as face, left arm, right arm, left hand, right hand, left leg, right leg, left foot, right foot, etc.) manually entered by users based on the characteristics of different body parts in the collected multi-part skin videos.
[0088] Each of the physiological part codes or the vectors corresponds to a single or combined physiological part (such as face, left arm, right arm, left hand, right hand, left leg, right leg, left foot, right foot, etc.), which is used to indicate the source part of the skin video signal of multiple parts and guide the attention of these parts by weighting.
[0089] As a preferred embodiment of the present invention, in step 4), see Figure 3 As shown in the figure, the present invention designs and builds a multi-region rPPG fusion physiological signal extraction model (BRPDNet) capable of dynamically adapting to multi-region rPPG signals. The architecture of the BRPDNet model mainly consists of an input layer, a multi-layer convolutional layer (ConvLayer), a hyper-conv module (Hyper-convModule), a physiological region prompt module (BioRegionPrompt Module), a time shift module (TSM), a dynamic attention layer (DynamicAttention Layer), a multi-layer feature weighting module, a dense layer (Dense), and an output layer.
[0090] The multi-layer convolutional layer acts as an encoder, and together with the super-convolution module, time shift module, and dynamic attention layer, it constitutes the feature extraction module of the BRPDNet model. The multi-layer feature weighting module acts as a decoder. Their functions are as follows:
[0091] The input layer is responsible for receiving the collected multi-site skin videos and performing preliminary standardization processing. Multi-site skin videos are optical sensor signals from different parts of the patient's body (such as arms and legs). The input size of each signal will vary depending on the resolution and sampling frequency of the camera equipment. Therefore, the preliminary standardization process can ensure that the multi-site skin videos collected from different parts can be processed in the same feature space. The multi-site skin videos are skin videos obtained by shooting multiple different parts of the patient using a camera equipment. It is also responsible for receiving physiological site prompt words input by the user, such as face, left arm, right arm, left hand, right hand, left leg, right leg, left foot, right foot, etc.
[0092] Multiple convolutional layers (ConvLayer) are responsible for extracting spatial features from the input multi-site skin video. These spatial features reflect local variations in light reflection from the skin. Convolutional layers are the core of the BRPDNet model. Each layer uses multiple convolution kernels to scan the input multi-site skin video to extract low-level features such as edges and texture. By stacking multiple layers of convolution, the BRPDNet model is able to gradually capture high-level features in the multi-site skin video, such as characteristic variations in the pulse waveform.
[0093] The Hyper-convModule is responsible for dynamically adjusting the size and receptive field of the convolution kernel according to the body part of the multi-part skin video, so as to improve the BRPDNet model's ability to capture the details of the multi-part skin video during the feature extraction process, ensuring that the BRPDNet model can efficiently extract the unique signal features of each part of the skin video. Figure 4 As shown in the figure, superconvolution is an improved convolution operation that allows the convolution kernel to dynamically adjust its size and receptive field at runtime to adapt to skin videos from different body parts. By adjusting the size and receptive field of the convolution kernel, superconvolution can improve the BRPDNet model's ability to capture signal details during feature extraction, especially in extracting differentiated signals between different body parts.
[0094] The time shift module (TSM) is responsible for injecting the time series of multi-site skin videos into the multi-layer convolutional layers, enabling the BRPDNet model to capture the temporal features in multi-site skin videos. It is used to capture the dynamic fluctuations of multi-site skin videos and perform time series analysis on spatial features, thereby obtaining stable and accurate spatiotemporal features, namely rPPG pulse waveform features. Multi-site skin videos are time series data, so the introduction of time series ensures that the BRPDNet model has better processing capabilities for the time-dependent characteristics of multi-site skin videos. For example, the BRPDNet model can better capture the changing patterns of pulse waveforms at different time points.
[0095] The Dynamic Attention Layer automatically adjusts the BRPDNet model's attention to different regions within the multi-region skin video based on its characteristics. This helps the model prioritize high-quality signal areas while downplaying noisy regions. In multi-region skin videos, signal quality can vary from region to region. For example, some regions may be significantly affected by motion artifacts or lighting variations, while others may maintain relatively stable signals. The Dynamic Attention Layer prioritizes high-quality signal areas while downplaying noisy regions.
[0096] BioRegionPromptModule, see Figure 5 As shown, it is responsible for converting the physiological part prompt words (such as left arm, right arm, left hand, right hand, left leg, right leg, left foot, right foot, etc.) input by the user into vectors that match the visual features of multiple parts through the CLIP model, or using one-hot encoding to edit the physiological part prompt words input by the user into physiological part codes, and passing the vectors or the physiological part codes as input to the super-convolution module, so as to provide a basis for the super-convolution module to identify the body part sources of the input multi-part skin video, and help the super-convolution module dynamically adjust the convolution kernel size and receptive field, thereby enhancing the BRPDNet model's attention to specific body parts.
[0097] These physiological site cues not only influence the superconvolutional module but also, when combined with asynchronous distillation during BRPDNet model training, further optimize feature extraction within the distillation network. By learning effective feature extraction methods from the teacher model, the distillation network enhances the student model (BRPDNet)'s learning of key site features, enabling the BRPDNet model to assign higher weights to certain sites during signal processing. The introduction of asynchronous distillation ensures efficient learning in complex environments, reducing computational complexity while improving the accuracy and robustness of multi-site signal extraction.
[0098] The multi-layer feature weighting module is responsible for fusing and weighting the rPPG pulse waveform features extracted by the feature extraction module at various locations through a weighted gate layer (WeightedGateLayer) according to a weighting mechanism. This ensures that the BRPDNet model can effectively integrate rPPG pulse waveform features from multiple locations and appropriately adjusts the weighting based on the quality of the rPPG pulse waveform features. The weighting mechanism weights the rPPG pulse waveform features of each location based on their quality and time series. Specifically, rPPG pulse waveform features with more stable and higher signal quality receive a greater weight in the model output, while rPPG pulse waveform features with greater noise or instability receive a smaller weight. This weighting mechanism can significantly improve the accuracy of the final physiological parameter estimation.
[0099] The dense layer is responsible for weighting and nonlinearly mapping the feature vectors output by the feature weighting module. This process essentially integrates multidimensional features into a compact representation, namely the integrated rPPG waveform signal. This allows the BRPDNet model to convert the spatiotemporal features of the input image into one-dimensional continuous signal features, ultimately outputting a waveform signal related to the target physiological signal (such as heart rate or respiratory rate) for subsequent calculation. In the BRPDNet model, convolutional features from multiple regions may form a high-dimensional vector. The dense layer compresses these high-dimensional features into a low-dimensional global representation H, reducing computational effort and improving the robustness of the BRPDNet model. With the help of the dense layer, the BRPDNet model can remove redundant information and retain only the features most valuable for heart rate or respiratory rate estimation.
[0100] The features of each part are represented as a tensor Fi, where i represents a different part (e.g., face, arm, etc.). Each Fi is a high-dimensional feature vector representing the spatiotemporal information extracted from that part. The function of the dense layer is to integrate these features into the final global feature vector H. Assuming the weight matrix of the dense layer is W, the operation of the dense layer can be expressed as:
[0101] (1);
[0102] In formula (1), σ is a nonlinear activation function (such as ReLU), which is used to introduce nonlinear features.
[0103] The output layer is responsible for outputting a one-dimensional visualized comprehensive rPPG waveform signal.
[0104] As a preferred embodiment of the present invention, in step 4), the training method of the BRPDNet model of the present invention is as follows:
[0105] 1) Construction of the dataset;
[0106] A certain number of skin videos of multiple parts (such as arms and legs) in clinical environments of past patients are collected, and the corresponding PPG pulse waveform signal true values and physiological parameter true values including heart rate, blood oxygen, and respiratory rate are obtained synchronously to form an rPPG dataset that contains temporally and spatially aligned multi-part skin videos, paired with PPG pulse waveform signal true values and physiological parameter true values.
[0107] To address the problem of facial skin being easily affected by facial expressions, head movement, and changing lighting conditions in rPPG applications, and to enhance signal stability and robustness, this paper designed and collected a dataset of rPPG signals from multiple locations and states, including the face, arms, and legs. This dataset covers subjects of varying ages, skin tones, and health conditions, ensuring stable signal monitoring even in complex medical equipment environments. Leg rPPG signals, in particular, are less affected by movement and lighting, offering greater stability and reliability, making them a reliable source of physiological signals.
[0108] During the research of this invention, we cooperated with the Fifth Affiliated Hospital of Sun Yat-sen University to collect a new dataset and integrated public facial datasets such as PURE, UBFC-rPPG, VIPL-HR, and UCLA-rPPG. Figure 6 As shown, Figure 6 Figure (a) shows an example from the PURE dataset, and Figure (b) shows an example from the UBFC-rPPG dataset. We also included multi-site PPG and respiratory monitoring data from 30 adults and newborns at Kiang Wu Hospital in Macau, as well as 30 clinical case data from Zhuhai People's Hospital.
[0109] It should be noted that the facial dataset not only improves the generalization of the facial rPPG model for the teacher model, but also serves as a soft label for distillation, providing guidance on rPPG signal labels for multiple parts of the student model. After adding RTT, the teacher's pred is equivalent to the student's label.
[0110] 2) Model training;
[0111] To further improve the efficiency and accuracy of signal extraction, this paper introduces an innovative BioRegionPrompt distillation network framework during model training. This framework dynamically adapts rPPG signal feature extraction from different locations by combining BioRegionPrompt and hyperconvolution techniques. This enables the method to focus on extracting signals from the limbs in scenarios susceptible to facial expressions, head movement, and lighting conditions, significantly improving the model's adaptability and signal extraction accuracy.
[0112] See also Figure 3and Figure 7 As shown in the figure, the BRPDNet model is trained using the constructed rPPG signal dataset and asynchronous distillation technology in the distillation network. The training process of the BRPDNet model adopts a staged guidance method, with the goal of enhancing the BRPDNet model's ability to extract multi-site skin video features. During training, multi-site skin videos are used as input, the true value of the PPG pulse waveform signal is used as the primary label, and the true value of the physiological parameter is used as the supplementary label. The BRPDNet model learns the mapping relationship between the video frames of multi-site skin videos and the corresponding PPG pulse waveform true value features, so that the BRPDNet model can accurately extract the rPPG pulse waveform features of each site based on the input multi-site rPPG signals, and output the comprehensive rPPG waveform signal after fusion weighting.
[0113] In the BRPDNet model, different parts of the body exhibit distinct skin characteristics (such as color, texture, and reflectivity), as well as signal signatures of blood flow changes. Using physiological site cues to guide feature weighting for different parts enables the BRPDNet model to assign attention weights to specific parts (such as "face" or "arm") based on these cues. However, in standard distillation, the target and non-target class information of the teacher model are transmitted simultaneously, preventing the student model from refining its focus on specific parts when optimizing its target.
[0114] The asynchronous distillation technology of the present invention decouples the optimization process of different body part information and dynamically adjusts the learning weights of each body part signal, allowing efficient extraction of signal features from each body part. The signal importance of each body part is considered equal. The BRPDNet model adjusts the optimization process based on the body part prompt information, ensuring that the signal features of all bodies are fully learned, while also supporting the independent prediction and fusion of multi-body part signals. Facial signals, due to their strong reflective properties and large data volume, provide more guidance in BRPDNet model training. However, because facial signals are obscured in scenes such as CT equipment, the asynchronous distillation technology of the present invention enhances the learning and optimization of signals from body parts other than the face. In the absence of facial signals, the BRPDNet model can still independently and accurately predict the physiological signal features of body parts other than the face. During the multi-body part signal fusion process, the model dynamically adjusts the weighting method for each body part signal to ensure that the physiological signal features output during the fusion process are consistent or synchronized with the facial signals in terms of timing and amplitude. The features of "non-target" parts are asynchronously optimized with weights in a different cycle, allowing the student model to maintain a certain understanding of the features of non-target parts without interfering with the main target features, thereby achieving efficient separation and optimization of multi-body part information.
[0115] See also Figure 8 As shown, Figure 8The first row is an example of the UBFC dataset, the second row is a schematic diagram of the feature learning hotspot of the mask area of the teacher model under the distillation network, the third row is a schematic diagram of the feature learning hotspot of the mask area of the student model under the distillation network, and the fourth row is the pulse feature waveform prediction output corresponding to the example of the UBFC dataset.
[0116] Different from traditional distillation, in asynchronous distillation, the present invention decouples the loss weight α and 1-α into two independent weights α target and β non-target , and apply these weights to the target class and non-target class respectively. The specific formula is as follows:
[0117] L=α target L target +β non-target L non-target (2);
[0118] In formula (2),
[0119] L target It is the loss for the target class, which mainly depends on the hard loss L of the real label hard and guidance from teacher models.
[0120] L non-target It is the loss of non-target classes, which mainly depends on the soft label loss L output by the teacher model. soft .
[0121] In addition, in the BRPDNet model, ordinary distillation requires large computing resources for the simultaneous optimization of all parts. The asynchronous distillation of the present invention divides the learning process of different parts into different cycles, which not only reduces the learning complexity of the student model, but also ensures that the characteristics of each part are paid attention to. For the lightweight multi-part rPPG fusion physiological signal extraction model, this update method of asynchronous distillation divided by part and prompt word can effectively cover the feature learning needs of multiple parts while saving calculations, thereby improving the performance of the lightweight model in multiple parts.
[0122] In summary, in the BRPDNet model that combines multiple parts such as the face and limbs and incorporates physiological part cues, asynchronous distillation has the following significant advantages over conventional distillation:
[0123] (1) Optimize the parts by parts to improve the model's adaptability to the characteristics of different parts.
[0124] (2) The weighted attention stability guided by the prompt word effectively improves the generalization effect of the prompt word in multiple parts.
[0125] (3) Computational resource optimization reduces the learning burden of student models and enables efficient multi-part lightweight design.
[0126] These advantages enable asynchronous distillation in the BRPDNet model to more accurately identify and separate blood flow signal features from different parts of the body, thereby improving the accuracy of heart rate estimation and the generalization ability of the model. This is especially important in applications on real-time devices or mobile devices.
[0127] This paper built the model using Python 3.9 and the PyTorch framework. The Adam optimizer with an initial learning rate of 0.0001 was used. The regularization step size was set to 1, and the gamma was set to 0.75. The number of training rounds was set to 50, and the batch size was set to 1. The model was trained and tested on an Intel Xeon W2245 CPU @ 3.90 GHz and 4 x NVIDIA RTX A1000 processors.
[0128] 3) Model optimization;
[0129] The present invention uses loss functions such as mean absolute error (MAE), root mean square error (RMSE) and signal-to-noise ratio (SNR) to optimize the model parameters of the BRPDNet model.
[0130] Mean absolute error (MAE) is used to evaluate the average absolute difference between the predicted and true values of rPPG signals. For rPPG signals, MAE helps measure the accuracy of heart rate and other physiological parameters. The smaller the MAE, the more accurate the model's predictions. The MAE is calculated as follows:
[0131] (3);
[0132] In formula (3), n is the number of samples, xi and yi are the predicted value and true value of the model respectively.
[0133] Root mean square error (RMSE) is used to evaluate the root mean square of the squared errors between the predicted and true values of the rPPG signal. RMSE provides a more sensitive error metric than MAE because it gives greater weight to larger errors. For rPPG signals, RMSE helps detect biases in heart rate estimation. A smaller RMSE indicates better model performance. The RMSE is calculated as follows:
[0134] (4);
[0135] In formula (4), n is the number of samples, xi and yi are the predicted value and true value of the model respectively.
[0136] The signal-to-noise ratio (SNR) is used to evaluate the ratio of signal strength to noise strength in rPPG signals. For image frame sequences, SNR measures the signal clarity and noise level during imaging, ensuring high-quality images. For rPPG signals, SNR helps assess the reliability and stability of signal detection. A higher SNR indicates higher signal quality and lower noise. The SNR calculation formula is as follows:
[0137] (5);
[0138] In formula (5), P signal Indicates the signal strength, P noise Indicates the noise intensity of the signal.
[0139] Through analysis and verification of these rPPG signal datasets, the BRPDNet model performed outstandingly in quality and stability, effectively supporting the fusion application of rPPG signals from multiple parts of the body.
[0140] In a CT equipment environment, the mean absolute error, the root mean square error, and the signal-to-noise ratio are used to calculate the error for signals of specific parts other than the face, and finally the total error is weighted to improve the generalization ability of the overall model and its ability to adapt to complex environments.
[0141] Based on the above-constructed rPPG signal dataset and the trained BRPDNet model, the present invention proposes a method for multi-site rPPG signal fusion to comprehensively monitor hemodynamic changes and help doctors better assess patients' long-term and short-term physiological health status.
[0142] See also Figure 7 As shown in the figure, multiple skin videos of the current patient are collected and input into the trained BRPDNet model after preprocessing including image enhancement and semantic segmentation. At the same time, the physiological part prompt words corresponding to the multiple skin parts videos are manually input and converted into physiological part codes or vectors.
[0143] The BRPDNet model first uses the input physiological site codes or vectors to adjust attention weights for different skin sites in the video, dynamically optimizing feature extraction capabilities for specific regions. It extracts rPPG pulse waveform features from the skin video space for each site, and combines them with time series analysis to ensure the spatiotemporal integrity and stability of these rPPG pulse waveform features. The BRPDNet model's asynchronous distillation process further optimizes feature learning, allowing the model to distinguish between target regions and reduce the impact of irrelevant regions on model performance. This approach enables the model to accurately extract physiological signals across diverse sites and complex scenarios. The BRPDNet model then fuses and weights the rPPG pulse waveform features from each site to generate a composite rPPG waveform signal. This process optimizes the signal contributions of different sites through a multi-layer feature weighting module, ensuring the final physiological parameters are highly accurate and robust.
[0144] Finally, an adaptive algorithm based on illumination variation is used to post-process the synthesized rPPG pulse waveform signal to correct for signal distortion caused by varying light intensity. Specifically, based on the extracted brightness value, cross-attention is performed on the real and imaginary components, and finally, an inverse fast Fourier transform is performed to convert it into a brightness-adapted rPPG waveform output. Frequency domain analysis is then used to calculate the physiological signal values, including heart rate and respiratory rate, on the post-processed rPPG waveform signal. A comprehensive report containing this physiological data is then generated. This physiological data can be provided to medical equipment operators in real time for clinical monitoring and decision support.
[0145] In addition, to measure the consistency between the BRPDNet model test results and subjective evaluation, this paper uses the PCC (Pearson correlation coefficient), also known as the Pearson correction coefficient. The PCC coefficient evaluates the model's prediction accuracy by calculating the linear correlation between two variables. The PCC coefficient ranges from -1 to 1. When the PCC value is 1, it indicates that the two sets of data are completely positively correlated. The calculation formula for PCC is as follows:
[0146] (6);
[0147] In formula (6), xi and yi are the model predicted value and the true value respectively. and In rPPG signal detection, PCC can measure the correlation between the predicted and true values of physiological parameters such as heart rate, ensuring the accuracy and consistency of signal detection.
[0148] The PTT experiment used contact PPG sensors at multiple locations to measure the blood pulse transit time at different body sites. The time delay between these sites was calculated using sliding cross-correlation. As shown in Table 1, the PTT results showed significant differences in blood pulse transit time between different sites, with the maximum delay between the left triceps and right ankle being 51.23 milliseconds.
[0149] Table 1 PTT experiment comparison results
[0150]
[0151] Table 1 shows significant differences in the time it takes for blood to travel from the heart to different sites. Specifically, sites closer to the heart (such as the right forearm, right triceps, and back of the neck) have shorter delays, while sites farther from the heart (such as the left knee, left ankle, and right ankle) have longer delays. This suggests that blood transport is more affected at more distant sites. Furthermore, negative values indicate that blood reaches the control site before the reference site, while positive values indicate that blood reaches the control site after the reference site.
[0152] Meanwhile, the rPTT experiment used high-frame-rate video to remotely measure pulse transit time at different locations. As shown in Table 2, the rPPG experiment found a maximum delay of 46.75 milliseconds, occurring between the face and the left leg. The rPPG results demonstrate that despite the presence of noise in the rPPG signal, high-frame-rate video can effectively estimate remote pulse transit time, providing reliable data support for non-contact physiological monitoring.
[0153] Table 2 rPPG experimental comparison results
[0154]
[0155] The tests and the above experiments prove that the BRPDNet model provided by the present invention has the following advantages:
[0156] 1. Contactless multi-site signal fusion to make up for signal loss in coverage scenarios:
[0157] When a patient enters a medical device such as a CT scan, the face and upper body are partially blocked. The present invention achieves reliable monitoring of physiological signals by extracting rPPG signals from the limbs (such as the left and right arms, left and right legs, etc.).
[0158] The experimental data in Table 2 show that the maximum rPTT delay between the left leg and the face is 46.75 milliseconds. These results prove that using signal fusion of the limbs can effectively reduce the impact of facial occlusion on physiological signal detection.
[0159] 2. Accuracy and consistency are significantly improved:
[0160] Ablation experiments demonstrate that, compared to other technologies such as PhysNet and RhythmMamba, the inclusion of dynamic convolution kernel generation significantly improves the accuracy of signal prediction by reducing the mean average effect (MAE) of rPPG signal detection from 1.64 to 1.55 and increasing the mean average correlation (PCC) to 0.76. This approach utilizes hyperconvolution to dynamically generate convolution kernels tailored to specific physiological locations by integrating cues from physiological regions. This adaptive convolution kernel mechanism enhances the model's sensitivity to input signals, significantly improving its stability under complex lighting and coverage conditions.
[0161] 3. Asynchronous distillation and knowledge distillation to reduce computational complexity:
[0162] Through asynchronous knowledge distillation, more effective feature extraction and dynamic adjustment of convolution kernels are achieved, enabling the model to screen features more flexibly in complex scenarios.
[0163] Compared with TSCAN and DeepPhys, BRPDNet's inference time is only 0.29 seconds per time, while still maintaining close accuracy, indicating that this method is extremely efficient and practical in environments with limited computing resources.
[0164] 4. Ability to adapt to noise interference and complex environment:
[0165] Under complex conditions such as changing lighting, facial occlusion, and even head movement, the present invention significantly reduces the impact of environmental changes on detection results by combining limb signal detection. In particular, the present invention demonstrates stronger anti-interference capabilities in low-light and high-light conditions.
[0166] For example, in the UBFC-rPPG test set, the MAE of the proposed method is 1.57 and the PCC is 0.73, which are much better than PhysNet and RhythmMamba, indicating the stability and accuracy of the proposed model in various complex environments.
[0167] In addition to simple physiological signal monitoring results, the present invention can also combine other information such as CT images to provide a comprehensive health status assessment. This data can be used for clinical diagnosis, monitoring patient status during surgery, or providing continuous monitoring support for long-term health management.
[0168] The present invention also provides a computer device comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus, and the memory is used to store at least one executable instruction, wherein the executable instruction enables the processor to perform operations corresponding to the above-mentioned multi-site rPPG fusion physiological signal monitoring method.
[0169] The present invention also provides a computer-readable storage medium, which stores at least one executable instruction, and the executable instruction enables a processor to perform operations corresponding to the above-mentioned multi-site rPPG fusion physiological signal monitoring method.
[0170] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A multi-site rPPG fusion physiological signal monitoring method, characterized in that: include: Step 1) Collect skin videos of multiple parts of the current patient; Step 2) Preprocessing the collected skin videos of multiple parts to eliminate environmental noise and interference; Step 3) editing physiological site prompt information for indicating the source sites of the collected multi-site skin video signals; Step 4) The pre-processed multi-site skin video and the edited physiological site prompt information are used as input data and fed into a multi-site rPPG fusion physiological signal extraction model trained based on an improved hyperconvolutional neural network and asynchronous distillation technology for processing; The multi-site rPPG fusion physiological signal extraction model adjusts the attention weight of skin videos of different sites based on the input physiological site prompt information, dynamically optimizes the regional feature extraction capability, extracts the rPPG pulse waveform features of the skin video space of each site, and combines it with time series analysis to ensure the spatiotemporal integrity and stability of the rPPG pulse waveform features. Finally, the rPPG pulse waveform features extracted from each site are weighted and fused to generate a comprehensive rPPG waveform signal. The training method of the multi-site rPPG fusion physiological signal extraction model is: 1) Collect multiple skin videos of patients in clinical settings and simultaneously obtain the corresponding true PPG pulse waveform signals and true physiological parameters including heart rate, blood oxygen, and respiratory rate. This forms an rPPG dataset that contains the temporally aligned multi-site skin videos, the true PPG pulse waveform signals, and the true physiological parameters. 2) Using the constructed rPPG signal dataset, asynchronous distillation technology is employed in a distillation network to train the multi-site rPPG fusion physiological signal extraction model. The model training process adopts a staged guidance approach. During training, multi-site skin videos are used as input, with the true values of PPG pulse waveform signals as primary labels and the true values of physiological parameters as supplementary labels. The model learns the mapping relationship between the features of the video frames of the multi-site skin videos and the corresponding true values of the PPG pulse waveforms, and outputs a comprehensive rPPG waveform signal after fusion and weighting. 3) Use loss function to optimize model parameters; Step 5) The extracted integrated rPPG pulse waveform signal is post-processed, and the physiological signal values including heart rate and respiratory rate are calculated through frequency domain analysis methods, and a comprehensive report containing these physiological data is output.
2. The multi-site rPPG fusion physiological signal monitoring method according to claim 1, characterized in that: The multi-site skin video is a skin video obtained by capturing multiple different parts of the patient through a camera device including a near-infrared camera or a high-frame-rate RGB camera. The multi-site skin video can reflect subtle changes in light reflected by the skin of multiple different parts of the patient in order to capture blood flow information.
3. The multi-site rPPG fusion physiological signal monitoring method according to claim 1, characterized in that: The physiological part prompt information is a physiological part code generated by editing the physiological part prompt word input by the user using one-hot encoding, or is a vector corresponding to each of the multiple part visual features converted by the CLIP model into the physiological part prompt word input by the user; The physiological part prompt words are prompt words with part keywords manually input by the user based on the characteristics of different body parts in the multi-part skin video collected; Each of the physiological part codes or the vectors corresponds to a separate or combined physiological part, and is used to indicate the source parts of the multi-part skin video signals and to guide the attention of these parts by weighting.
4. The multi-site rPPG fusion physiological signal monitoring method according to claim 1, characterized in that: The architecture of the multi-site rPPG fusion physiological signal extraction model includes an input layer, a feature extraction module, a physiological site prompt module, a feature weighting module and a dense layer; The input layer is responsible for receiving the collected multi-site skin videos and performing preliminary standardization processing. The multi-site skin videos are skin videos obtained by shooting multiple different parts of the patient using a camera device; it is also responsible for receiving physiological site prompt words input by the user; The feature extraction module consists of multiple convolutional layers, superconvolution modules, time translation modules and dynamic attention layers. It is responsible for extracting the rPPG pulse waveform features of each part from the skin video of multiple parts, and obtaining a comprehensive rPPG waveform signal after fusion and weighting for subsequent physiological signal analysis. The multi-layer convolutional layer is the core of the model, responsible for extracting spatial features from multiple skin videos. These spatial features reflect local changes in the skin's reflected light. Each convolutional layer uses multiple convolution kernels to scan multiple skin videos to extract low-level features such as edges and textures. By superimposing multiple layers of convolution, the model can gradually capture high-level features in multiple skin videos, including characteristic changes in pulse waveforms. The hyperconvolution module is responsible for dynamically adjusting the size and receptive field of the convolution kernel based on the body part of the multi-part skin video, thereby improving the model's ability to capture details of the multi-part skin video during the feature extraction process, ensuring that the model can efficiently extract the unique signal characteristics of each skin part video; The time shift module is responsible for injecting the time series of multiple skin videos into the multi-layer convolutional layer, enabling the model to capture the temporal features in the multiple skin videos, capturing the dynamic fluctuations of the multiple skin videos, and performing time series analysis on the spatial features, thereby obtaining stable and accurate spatiotemporal features, namely rPPG pulse waveform features. Introducing the time series ensures that the model has better processing capabilities for the time-dependent characteristics of multiple skin videos, enabling the model to better capture the changing patterns of the pulse waveform at different time points. The dynamic attention layer is responsible for automatically adjusting the model's attention to different areas in the multi-part skin video based on the characteristics of the multi-part skin video; The physiological part prompt module is responsible for converting the physiological part prompt word input by the user into a vector that matches the visual features of multiple parts through the CLIP model, or editing the physiological part prompt word input by the user into a physiological part code using one-hot encoding, and passing the vector or the physiological part code as input to the superconvolution module to provide a basis for the superconvolution module to identify the body part source of the input multi-part skin video; The feature weighting module is responsible for fusing and weighting the rPPG pulse waveform features of various parts extracted by the feature extraction module according to the weighting mechanism through the weighted gate layer to ensure that the model can effectively fuse the rPPG pulse waveform features of multiple parts and adjust the weighting ratio according to the quality of the rPPG pulse waveform features; The weighting mechanism weights the rPPG pulse waveform features of each site based on the quality and time sequence of the rPPG pulse waveform features of each site; The dense layer is responsible for weighting and nonlinear mapping the feature vector output by the feature weighting module, integrating the multi-dimensional features into a compact representation, namely the comprehensive rPPG waveform signal.
5. The multi-site rPPG fusion physiological signal monitoring method according to claim 1, characterized in that: The asynchronous distillation technology decouples the optimization process of different body part information and dynamically adjusts the learning weights of the signals of each body part, so that the signal features of each body part can be efficiently extracted. The signal importance of each body part is considered to be uniform and equal. The model adjusts the optimization process according to the body part prompt information to ensure that the signal features of all bodies are fully learned, while also supporting the independent prediction and fusion of multi-body part signals. Facial signals provide more guidance in model training due to their strong reflective characteristics and large data volume. However, since facial signals are obscured in CT equipment scenes, the asynchronous distillation technology enhances the learning and optimization of signals of parts other than the face, so that even in the absence of facial signals, the model can still independently and accurately predict the physiological signal features of parts other than the face. During the multi-body part signal fusion process, the model dynamically adjusts the weighting method for each body part signal to ensure that the physiological signal features output during the fusion process are consistent or synchronized with the facial signals in terms of timing and amplitude. The features of non-target parts are asynchronously optimized with weights of another cycle, so that the student model can maintain an understanding of the features of non-target parts without interfering with the main target features, thereby achieving efficient separation and optimization of multi-body part information.
6. The multi-site rPPG fusion physiological signal monitoring method according to claim 1, characterized in that: The loss function includes mean absolute error, root mean square error and signal-to-noise ratio; wherein, The mean absolute error is used to evaluate the mean absolute difference between the predicted value and the true value of the rPPG signal; The root mean square error is used to evaluate the root mean square of the square error between the predicted value and the true value of the rPPG signal; The signal-to-noise ratio is used to evaluate the ratio of signal intensity to noise intensity in the rPPG signal; In a CT equipment environment, the mean absolute error, the root mean square error, and the signal-to-noise ratio are used to calculate the error for signals of parts other than the face, and finally the total error is weighted to improve the generalization ability of the overall model and its ability to adapt to complex environments.
7. The multi-site rPPG fusion physiological signal monitoring method according to claim 1, characterized in that: An adaptive algorithm based on illumination changes is used to post-process the integrated rPPG pulse waveform signal to correct the signal distortion caused by changes in light intensity. Specifically, according to the extracted brightness value, cross-attention is performed on the real and imaginary parts respectively, and finally an inverse fast Fourier transform is performed to convert it into a brightness-adapted rPPG waveform signal output.
8. A computer device, characterized in that: include: A processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus, and the memory is used to store at least one executable instruction, wherein the executable instruction causes the processor to perform an operation corresponding to the multi-site rPPG fusion physiological signal monitoring method according to any one of claims 1 to 7.
9. A computer storage medium, characterized in that The computer-readable storage medium stores at least one executable instruction, which enables the processor to perform operations corresponding to the multi-site rPPG fusion physiological signal monitoring method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Non-contact multi-modal physiological signal detection method based on self-supervision and lifelong learning
CN115497143A
Non-contact physiological signal detection method and device based on fusion feature enhancement
CN117694845A