A remote physiological signal estimation method and system based on diffusion model

By introducing a diffusion model-based method in remote physiological signal detection technology, combining multi-scale spatiotemporal maps for feature extraction and fusion, the problem of insufficient anti-interference ability and model generalization ability in the existing technology is solved, and higher detection accuracy and robustness are achieved.

CN119670022BActive Publication Date: 2025-05-06HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510186180.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-06
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

The existing remote physiological signal detection technology based on facial videos has shortcomings in anti-interference ability, model generalization ability and computing resource requirements, especially when facing different populations, lighting changes and facial movements, the detection accuracy and stability are low.

Method used

Using a remote physiological signal estimation method based on diffusion model, through the forward diffusion and backward diffusion process, we learn how to extract and restore physiological signals from noise signals, and combine multi-scale spatiotemporal maps for feature extraction and fusion to improve the accuracy and robustness of detection.

Benefits of technology

It improves the accuracy and robustness of remote physiological signal detection, enhances anti-interference ability, reduces the demand for computing resources, is suitable for more application scenarios, and improves the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119670022B_ABST
    Figure CN119670022B_ABST
Patent Text Reader

Abstract

The present invention discloses a remote physiological signal estimation method and system based on a diffusion model, and relates to the technical field of non-contact physiological signal detection. The training process of the diffusion model is: forward diffusion of the original rPPG signal to obtain a noise rPPG signal; feature extraction of the face video to obtain a multi-scale space-time graph; fusion of the noise rPPG signal and the multi-scale space-time graph; backward diffusion of the fused data to obtain denoised data; output of the predicted rPPG signal according to the denoised data; model training by minimizing the difference between the predicted rPPG signal and the original rPPG signal; remote physiological signal estimation using the trained diffusion model. The present invention improves the accuracy of rPPG signal estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of non-contact physiological signal detection, and in particular to a remote physiological signal estimation method and system based on a diffusion model. Background Art

[0002] Currently, in the field of physiological signal detection, the main technologies can be divided into contact and non-contact methods. Contact methods include pulse diagnosis, stethoscopes, electrocardiographs, and smart bracelets, which obtain physiological signals by direct contact with the human body. Although these methods have high accuracy, they are cumbersome to operate, not convenient enough, and are not applicable in some application scenarios. In contrast, non-contact methods are gradually gaining more attention, such as remote physiological signal detection based on facial video (rPPG, remote Photoplethysmography). This technology captures changes in light absorption of the skin through video to infer physiological parameters such as heart rate. The advantage of this method is that it does not require direct contact with the object being detected, has low cost and a wide range of application scenarios, but its detection accuracy is relatively low and is affected by external factors such as light and movement.

[0003] Existing physiological signal detection technologies based on facial videos mainly include the following different model architectures:

[0004] 1. Based on non-end-to-end models: This type of model usually requires complex feature preprocessing steps to extract information from channels such as RGB and YUV, and uses feature decoupling strategies to improve the accuracy of rPPG estimation. Although this type of model has high accuracy, the process is complex and the real-time performance is poor.

[0005] 2. Based on end-to-end models: In recent years, end-to-end models such as the video Transformer architecture have begun to emerge. These models can adaptively aggregate local and global spatiotemporal features to enhance the representation of rPPG and reduce intermediate preprocessing steps. However, such models are often computationally intensive and require high computing power on the device.

[0006] 3. Unsupervised model based on contrastive learning: This type of model uses the spatial similarity of different regions and the temporal similarity of rPPG signals in a short period of time for feature learning. Unsupervised learning reduces the dependence on labeled data, but its model stability and accuracy still need to be improved.

[0007] Although the existing non-contact physiological signal detection technology has made significant progress, there are still some shortcomings. First, the anti-interference ability is limited. Light changes, facial occlusion, head movement, etc. will significantly affect the accuracy of the detection results. For example, facial occlusions such as hair, beards, glasses, and changes in light conditions will lead to inaccurate capture of skin color changes. Secondly, there are large differences in different races, ages and skin characteristics. When facing different groups of people, the generalization ability of the existing technology model is limited. The optical properties of different skins have different detection effects on rPPG signals. At the same time, the computing resource requirements are high, especially the end-to-end model based on 3D CNN (three-dimensional convolutional neural network). Although it can improve the accuracy of signal detection, its computational complexity is extremely high and is not suitable for scenarios with limited resources. Summary of the invention

[0008] In order to overcome the above-mentioned defects in the prior art, the present invention provides a remote physiological signal estimation method based on a diffusion model, which improves the accuracy of rPPG signal estimation.

[0009] To achieve the above object, the present invention adopts the following technical solutions, including:

[0010] A remote physiological signal estimation method based on diffusion model,

[0011] The training process of the diffusion model is as follows:

[0012] S11, obtaining a sample set, where the sample includes a face video and an original rPPG signal corresponding to the face video;

[0013] S12, forward diffusion is performed on the original rPPG signal, i.e., noise is added to obtain a noisy rPPG signal;

[0014] S13, extracting features from the face video to obtain a multi-scale spatiotemporal graph;

[0015] S14, fusing the noisy rPPG signal and the multi-scale spatiotemporal graph to obtain fused data;

[0016] S15, performing back diffusion, i.e., denoising, on the fused data to obtain denoised data;

[0017] S16, outputting a predicted rPPG signal based on the denoised data;

[0018] S17, performing model training by minimizing the difference between the predicted rPPG signal and the original rPPG signal to obtain a trained diffusion model;

[0019] The trained diffusion model is used to estimate long-range physiological signals as follows:

[0020] S21, inputting a face video to be estimated and generating a noise rPPG signal;

[0021] S22, extracting features from the estimated face video to obtain a multi-scale spatiotemporal graph;

[0022] S23, fusing the generated noise rPPG signal and the multi-scale spatiotemporal graph to obtain fused data;

[0023] S24, performing denoising processing on the fused data to obtain denoised data;

[0024] S25, outputting a predicted rPPG signal based on the denoised data.

[0025] Preferably, in step S12, the expression of forward diffusion is as follows:

[0026] ;

[0027] in, x 0 is the original rPPG signal; x t For the t Noise rPPG signal after forward diffusion; t is the number of forward diffusions, t =1,2,..., N , and finally obtained N Noise rPPG signal after forward diffusion x N ; q ( x t | x 0) indicates a given initial state x In the case of 0, the state x t The conditional probability distribution of ; is the attenuation factor, In the diffusion process, as the diffusion times t decreases gradually with the increase of For noise.

[0028] Preferably, in step S13 and step S22, feature extraction is performed on the face video, and the specific method is as follows:

[0029] Divide the face video into multiple segments, and get F Frame video image;

[0030] Perform face detection on video images and divide the faces into n A region of interest is ROI;

[0031] Will nROIs are combined to obtain a combination of a single ROI and different ROIs, a total of 2 n - 1 combination;

[0032] For each combination, the average pixel value of all ROIs in the combination in each color channel is calculated; there are 6 color channels in total, namely RGB and YUV;

[0033] Finally, the dimension is 6×(2 n -1)× F The vector of is used as a multi-scale space-time graph.

[0034] Preferably, in step S14 and step S23, the noisy rPPG signal and the multi-scale spatiotemporal graph are fused, and the specific method is as follows:

[0035] The dimension of the noisy rPPG signal is 1× F , the dimension of the noisy rPPG signal is expanded to 1×(2 n -1)× F , the expanded noisy rPPG signal is spliced ​​with the multi-scale spatiotemporal graph in the channel dimension to obtain the fused data, the dimension of the fused data is 7×(2 n -1)× F .

[0036] Preferably, in step S15, during the training process, the fusion data is subjected to denoising processing, i.e., backward diffusion, using a denoising module, wherein the denoising module uses a mixSTE, i.e., a hybrid space-time encoder, which is internally composed of c layers, each layer includes a Spatial Transformer block and a Temporal Transformer block connected in sequence; the Spatial Transformer block is used to extract the spatial features of the fused data, and the Temporal Transformer block is used to extract the spatiotemporal features of the fused data; after mixSTE, a Resnet residual connection is used; the denoising expression during training is as follows:

[0037] ;

[0038] ;

[0039] in, y N To fuse the data, zis the feature vector after backward diffusion, that is, the denoised data; mixSTE(·) is the processing function of the mixed space-time encoder; SpatialTransformer(·) is the processing function of the Spatial Transformer block; TemporalTransformer(·) is the processing function of the Temporal Transformer block.

[0040] Preferably, the Spatial Transformer block is processed as follows:

[0041] First, the input data is the fusion data y N Using the linear mapping function, we get the three parameters of the spatial self-attention mechanism: Q s 、K s 、V s :

[0042] ;

[0043] Then, the fusion data is calculated through the self-attention mechanism Attention y N Spatial characteristics S :

[0044] ;

[0045] The Temporal Transformer block is processed as follows:

[0046] First, use the linear mapping function on the spatial feature S to obtain the three parameters of the temporal self-attention mechanism, namely Q t 、K t 、V t :

[0047] ;

[0048] Then, the fusion data is calculated through the self-attention mechanism Attention y N The spatiotemporal characteristics of ST :

[0049] ;

[0050] Spatial and temporal characteristics ST That is:

[0051] .

[0052] Preferably, in step S24, during the estimation process, the fusion data is denoised using a denoising module, and the denoising is performed iteratively. K After backward diffusion, the denoised data is finally output; the denoising expression in the estimation process is as follows:

[0053] z k = z k-1 +mixSTE( z k-1 );

[0054] in, k is the number of iterations of back diffusion, k =1,2,..., K ; z k For the k The eigenvector of the backward diffusion, z 0 is the fusion data, and finally the K The eigenvector of the backward diffusion z K , which is the denoised data.

[0055] Preferably, in step S12 and step 21, the noise rPPG signal is Gaussian noise that obeys a standard normal distribution.

[0056] The present invention also provides a remote physiological signal estimation system based on a diffusion model, characterized in that it is applicable to the above-mentioned remote physiological signal estimation method based on a diffusion model, and the system includes: a forward diffusion module, an MSTmap calculation module, a fusion module, a denoising module, a prediction module and a sampling module; the working process of the system is divided into two parts, namely a training process and an estimation process, which are specifically as follows:

[0057] During training:

[0058] A forward diffusion module is used to add noise to the original rPPG signal to obtain a noisy rPPG signal;

[0059] MSTmap calculation module, used to process face videos to obtain multi-scale spatiotemporal maps containing visual features;

[0060] A fusion module is used to fuse the noisy rPPG signal with the multi-scale spatiotemporal graph to obtain fused data;

[0061] A denoising module is used to denoise the fused data to obtain denoised data;

[0062] A prediction module, used to predict the rPPG signal based on the denoised data output;

[0063] During the estimation process:

[0064] A sampling module, used for random sampling to generate noise rPPG signals;

[0065] MSTmap calculation module, used to process the estimated face video to obtain a multi-scale spatiotemporal map containing visual features;

[0066] A fusion module is used to fuse the noisy rPPG signal with the multi-scale spatiotemporal graph to obtain fused data;

[0067] A denoising module is used to denoise the fused data to obtain denoised data;

[0068] The prediction module is used to output the predicted rPPG signal based on the denoised data.

[0069] The present invention also provides a computer program product, which includes a computer program / instruction. When the computer program / instruction is executed by a processor, the remote physiological signal estimation method based on the diffusion model is implemented.

[0070] The advantages of the present invention are:

[0071] (1) The present invention introduces a diffusion model for remote physiological signal detection. The diffusion model has stronger generalization ability because it can learn the distribution of target data. Through the diffusion model, the estimation of rPPG signals becomes more realistic and effective, especially it can better utilize the periodicity and distribution characteristics of rPPG signals. In addition, the diffusion model of the present invention can better cope with external environmental changes, such as lighting changes and facial movements, and improve the robustness of the overall detection.

[0072] (2) The introduction of the diffusion model makes the estimation of rPPG signals more realistic, better reflects the real physiological signals, and improves the authenticity of rPPG signal estimation.

[0073] (3) The diffusion model is more adaptable in dealing with problems such as illumination changes and head movement, thereby improving the stability and robustness of detection and enhancing anti-interference capabilities.

[0074] (4) Compared with the traditional 3D CNN model, the diffusion model is more optimized in terms of computing resource usage, has higher computational efficiency, and is suitable for more application scenarios.

[0075] (5) A multi-scale spatiotemporal map, or MSTmap, is used instead of directly inputting face videos. When calculating MSTmap, a face recognition tool is used to detect faces. The model can act directly on the effective area, greatly increasing the accuracy of the model.

[0076] (6) The MSTmap calculation introduces an additional YUV channel on top of the original RGB channel, which can enrich the extracted feature information, allowing the model to be better trained and reasoned.

[0077] (7) The MSTmap calculation arranges and combines the ROI (region of interest) of the face, which can enrich the spatial information, so that more information can be analyzed in the subsequent denoising module to obtain more accurate results.

[0078] (8) The noisy rPPG signal is fused with the MSTmap on the channel. The noisy rPPG signal is one-dimensional, while the MSTmap color channel has multiple dimensions. Therefore, the MSTmap still occupies the main part and will not be contaminated by the noisy rPPG signal, thereby affecting the model prediction.

[0079] (9) The noisy rPPG signal is fused with the MSTmap by splicing. The numerical value of the noisy rPPG signal and the numerical value of the MSTmap are not changed. Only the dimensions are connected together. The feature meanings are independent, so that the model can distinguish well. It can not only learn the data distribution of the rPPG signal from the noisy rPPG signal restoration process, but also use this distribution to guide the mapping of the MSTmap to the real rPPG signal. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] Figure 1 It is a schematic diagram of a remote physiological signal estimation system based on a diffusion model of the present invention.

[0081] Figure 2 The present invention is a flowchart of a remote physiological signal estimation method based on a diffusion model.

[0082] Figure 3 Schematic diagram of facial feature points. DETAILED DESCRIPTION

[0083] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0084] The related technologies involved in the present invention are introduced as follows:

[0085] 1. Basic principles of diffusion model

[0086] The diffusion model gradually transforms the data from the initial distribution to a simple Gaussian distribution through a random process, and then reverses the process to generate the data. This process consists of two main stages: forward diffusion (adding noise) and backward diffusion (removing noise).

[0087] Forward diffusion process: from the original data x Starting from 0, noise is gradually introduced to generate a series of increasingly noisy data through multiple steps. x 1, x 2,..., x T This process is performed via a Gaussian transformation, with the noise increase at each step being predefined and calculable.

[0088] Backward diffusion process: from the noisiest data x T Start by removing noise gradually and finally restore to close to the original data x Each denoising step uses a conditional probability model to predict the state of the data with less noise.

[0089] 2. Mathematical representation of the diffusion model

[0090] Transition Kernel: used to define how to go from one step to the next, for example, using a Gaussian distribution model Indicates how to pass data at each step x t Introducing noise.

[0091] Parameterization and Hyperparameters: β t It is a hyperparameter that controls the amount of noise introduced and is usually set before training. The goal of the whole process is to transform the data into a standard normal distribution (Gaussian distribution) by gradually adding noise.

[0092] Inverse process: In the reverse process, the model used for denoising (such as U-net) predicts the state of the denoised data based on the current noisy data. This is usually achieved by minimizing the difference between the predicted data and the actual data.

[0093] 3. Application of diffusion model in physiological signal detection

[0094] The application of diffusion model in remote physiological signal (rPPG signal) detection is based on its ability to process highly noisy data and recover accurate information. By applying the diffusion model to video data, remote physiological signals (rPPG signals) that are disturbed by various external factors (such as lighting changes and facial movements) can be effectively extracted.

[0095] Data preprocessing: The raw signals extracted from the video frames are first subjected to a forward diffusion process to transform them into a more standardized form, which helps to isolate and remove non-physiological noise components.

[0096] Signal reconstruction: During the back diffusion process, the model learns how to gradually remove noise and ultimately reconstruct a physiological signal that is close to the real thing.

[0097] Depend on Figure 1 As shown, a remote physiological signal estimation system based on a diffusion model of the present invention includes: a forward diffusion module, an MSTmap calculation module, a fusion module, a denoising module, a prediction module and a sampling module. The working process of the whole system is divided into two parts, namely a training process and an estimation process.

[0098] During training:

[0099] The forward diffusion module uses a Gaussian mixture noise model to simulate the diffusion of the rPPG signal, so that the rPPG signal is gradually transformed into a more random state and converted into a noisy rPPG signal, which helps the subsequent diffusion model learn how to effectively remove noise.

[0100] The MSTmap calculation module processes the face video and calculates the multi-scale spatiotemporal map, which involves extracting the motion, structure and texture information in the video to provide the necessary visual features for subsequent processing.

[0101] The fusion module fuses the noisy rPPG signal with the multi-scale spatiotemporal map to generate a fused data that combines physiological and visual information and provides input for the denoising module.

[0102] The denoising module (backward diffusion module) uses a deep learning network (such as the Transformer structure) to learn how to restore the fused data to a noise-free state, remove the noise introduced by the forward diffusion, and obtain the denoised data.

[0103] Prediction module,The final rPPG signal prediction is completed by processing the denoised data,,ensuring that the output predicted rPPG signal is as close to the real physiological,state as possible.

[0104] During the estimation process:

[0105] The sampling module generates a noisy rPPG signal by randomly sampling from a Gaussian distribution, providing the model with an initial noisy rPPG signal for starting the back diffusion process.

[0106] The MSTmap calculation module processes the input face video to be estimated and calculates a multi-scale spatiotemporal map, which involves extracting motion, structure and texture information in the video to provide necessary visual features for subsequent processing.

[0107] The fusion module combines the noisy rPPG signal with the multi-scale spatiotemporal map to generate a fused data that combines physiological and visual information and provides input to the denoising module.

[0108] Denoising module,During the estimation process, the denoising module is not only used once, but is reused for multiple iterations. Through continuous denoising and reconstruction processes, the final noise-free state is gradually refined and restored, and the final denoised data is output.

[0109] Prediction module,The final rPPG signal prediction is completed through the denoised data, and the predicted rPPG signal is output.

[0110] Depend on Figure 2 As shown, a remote physiological signal estimation method based on a diffusion model of the present invention first inputs the collected sample data to train the diffusion model during the training process, and then uses the trained diffusion model to estimate the remote physiological signal. During the estimation process, the input is the face video of the person to be detected, and the output is the corresponding predicted rPPG signal.

[0111] The training process of the diffusion model is as follows:

[0112] S11, obtaining a sample set, wherein the sample includes a face video and a label corresponding to the face video, that is, an original rPPG signal;

[0113] S12, the original rPPG signal is input into the forward diffusion module for forward diffusion, which is used to add noise and transform the ordered original rPPG signal into a disordered noise rPPG signal, gradually transitioning from an initial state with almost no noise to a state mainly composed of noise. This process provides a rich data set for model training, including all intermediate states from no noise to high noise, which is crucial for the subsequent reverse process (i.e., denoising process) because it relies on these intermediate states. x t Let's learn how to gradually restore to the original noise-free state. The mathematical expression of forward diffusion is as follows:

[0114] ;

[0115] in,x 0 represents the original rPPG signal; x t Indicates t Noise rPPG signal after forward diffusion; t represents the number of forward diffusions, t =1,2,..., N , and finally obtained N Noise rPPG signal after forward diffusion x N In this embodiment N =1000; q ( x t | x 0) indicates a given initial state x In the case of 0, the state x t The conditional probability distribution of x 0 to x t the evolution process; is a predefined attenuation factor, It gradually decreases during the diffusion process. Controls the proportion of the original rPPG signal retained in each step of forward diffusion; Represents the standard normal distribution Random noise extracted from Determines the proportion of noise in each step of forward diffusion. t increase, Gradually increase.

[0116] S13, inputting the face video corresponding to the original rPPG signal into the MSTmap calculation module to calculate the multi-scale spatiotemporal map, specifically in the following way:

[0117] Divide the input face video into F The frame is a segment, and each segment is divided into multiple segments with an interval of 15 frames, that is, each segment is F Frame video image, in this embodiment, F =300;

[0118] The facial detection tool OpenFace is used to detect faces and extract 68 facial feature points, such as Figure 3 As shown, based on the facial feature points, six ROIs (Region Of Interester) of the face are defined, namely the forehead, left cheek, right cheek, left triangle area, right triangle area and chin.

[0119] Will nROIs are combined to obtain a combination of a single ROI and different ROIs, a total of 2 n -1 combination, in this embodiment n =6.

[0120] For each combination, the average pixel value of all ROIs in the combination in each color channel is calculated. There are 6 color channels in total, namely RGB and YUV.

[0121] Finally, the dimension is 6×(2 n -1)× F The vector of is used as a multi-scale space-time graph. In this embodiment, the dimension of the multi-scale space-time graph is 1×63×300.

[0122] S14, inputting the noisy rPPG signal obtained in step S12 and the multi-scale spatiotemporal graph obtained in step S13 into a fusion module for fusion to obtain fused data.

[0123] Noisy rPPG signal after forward diffusion x N The dimension is 1×300. First, the noise rPPG signal x N The dimension is expanded to 1×63×300, and the expanded noisy rPPG signal is connected with the multi-scale spatiotemporal graph in the channel dimension using the concat function to obtain the fused data y N , fused data y N The dimensions are 7×63×300.

[0124] S15, inputting the fused data into a denoising module for denoising processing, i.e., performing backward diffusion, to obtain denoised data, thereby obtaining a stable and pure (i.e., noise-free) rPPG feature vector.

[0125] The denoising module uses mixSTE (hybrid space-time encoder), which is internally composed of c layers, each of which includes a Spatial Transformer block and a Temporal Transformer block connected in sequence. c =6.

[0126] mixSTE is a hybrid space-time encoder that extracts features in spatial and temporal dimensions and fuses multiple layers of embedded features to better understand the spatiotemporal relationship in the data.

[0127] The Spatial Transformer block is a neural network module based on the Transformer architecture, which can enhance the model's ability to process spatial transformations and extract spatial features, and is used to extract fused data. y N spatial characteristics.

[0128] The Temporal Transformer block is a neural network module based on the Transformer architecture that can effectively capture dependencies in the temporal dimension and is used to extract fused data. y N spatiotemporal characteristics.

[0129] After mixSTE, a Resnet residual connection is used to ensure that the network can maintain stable training and efficient feature extraction even when performing complex spatial and temporal feature extraction. The mathematical expression of the denoising module during training is as follows:

[0130] ;

[0131] in, z is the feature vector after back diffusion, that is, the denoised data, z The dimension is 7×63×300; mixSTE(·) is the processing function of the hybrid space-time encoder.

[0132] ;

[0133] Among them, SpatialTransformer(·) is the processing function of the Spatial Transformer block, and the processing method is:

[0134] First, the input data is the fusion data y N Using the linear mapping function, we get the three parameters of the spatial self-attention mechanism: Q s 、K s 、V s :

[0135] ;

[0136] Then, the fusion data is calculated through the self-attention mechanism Attention y N Spatial characteristics S :

[0137] ;

[0138] TemporalTransformer(·) is the processing function of the Temporal Transformer block, and the processing method is:

[0139] First, use the linear mapping function on the spatial feature S to obtain the three parameters of the temporal self-attention mechanism, namely Q t 、K t 、V t :

[0140] ;

[0141] Then, the fusion data is calculated through the self-attention mechanism Attention y N The spatiotemporal characteristics of ST :

[0142] ;

[0143] Spatial and temporal characteristics ST That is:

[0144] .

[0145] S16, the denoised data obtained from step S15, i.e., the feature vector after back diffusion z Input into the prediction module and finally get the predicted rPPG signal.

[0146] First, the denoised data, i.e., the feature vector after back diffusion, is z (Dimension is 7×63×300) Average pooling is performed on the second dimension to obtain intermediate features (dimension is 7×1×300), and then a Linear layer is passed to reduce the channel dimension to finally obtain the predicted rPPG signal (dimension is 1×1×300).

[0147] S17, performing model training by minimizing the difference between the predicted rPPG signal and the original rPPG signal to obtain a trained diffusion model.

[0148] The trained diffusion model is used to estimate the remote physiological signal. The estimation process is as follows:

[0149] S21, using the sampling module from a standard normal distribution An initial noisy rPPG signal is sampled from the Gaussian noise X N , and then use the randomly generated noise rPPG signalX N As the starting point of the reverse diffusion process. The mathematical expression of sampling is as follows:

[0150] ;

[0151] Among them, the randomly generated noise rPPG signal X N The dimension is 1×300.

[0152] S22, according to the method of step S13, the input face video is input into the MSTmap calculation module to calculate the multi-scale spatiotemporal map.

[0153] S23, according to step S14, the randomly generated noise rPPG signal X N The multi-scale spatiotemporal graph is input into the fusion module for fusion to obtain fused data Y N .

[0154] S24, the fusion data Y N Input into the denoising module for denoising to obtain denoised data.

[0155] The denoising method in the estimation process is different from that in the training process. The feature vector output after denoising by the denoising module in the estimation process is used as input data again and input into the denoising module for denoising. K The final output is obtained after backward diffusion iterations. The expression is as follows:

[0156] z k = z k-1 +mixSTE( z k-1 );

[0157] mixSTE( z k-1 )=TemporalTransformer(SpatialTransformer( z k-1 ))× c ;

[0158] in, k is the number of iterations of back diffusion, k =1,2,..., K ; z k For the k The eigenvector of the backward diffusion, z0 is fused data, z 0= Y N , and finally obtained K The eigenvector of the backward diffusion z K , that is, the denoised data, z K The dimensions are 7×63×300.

[0159] In this embodiment, the number of iterations of the backward diffusion is K =5, the number of iterations of forward diffusion N =1000, from Y 1000 Start back diffusion, the feature vector of each back diffusion is equivalent to:

[0160] .

[0161] S25, according to the method of step S16, K The eigenvector of the backward diffusion z K Input into the prediction module and finally get the predicted rPPG signal.

[0162] The present invention uses a multi-scale spatiotemporal map, i.e., MSTmap, instead of directly inputting the video, which can produce the following effects:

[0163] First, when calculating MSTmap, face recognition tools are used to detect faces, which greatly increases the accuracy of model recognition. When the video is directly input, the model cannot distinguish the effective area of ​​the picture, so it may mistake the background picture for the face. Using face recognition tools to detect faces, the model can directly act on the effective area, which naturally increases the accuracy.

[0164] Secondly, MSTmap introduces an additional YUV channel in addition to the original RGB channel. These three channels have been confirmed to be effective by previous studies. Therefore, the introduction of the additional three channels can enrich the extracted feature information, allowing the model to be better trained and inferred.

[0165] Finally, MSTmap arranges and combines the face ROI (Region Of Interester), and the original 6 regions are combined to produce 63 different combinations. These 63 combinations can enrich the spatial information, so that the SpatialTransformer in the mixSTE layer in the subsequent Denoiser can parse more information and obtain more accurate results.

[0166] The present invention fuses the rPPG signal into the color channel dimension of the MSTmap, which can produce the following effects:

[0167] The noisy rPPG signal is fused with the MSTmap on the channel. The noisy rPPG signal is one-dimensional, while the MSTmap color channel has 6 dimensions, so the MSTmap still occupies the main part and will not be contaminated by the noisy rPPG signal, thus affecting the prediction of the model.

[0168] At the same time, the noisy rPPG signal is fused using concat and MSTmap. The values ​​of the noise rPPG signal and the MSTmap are not changed. They are just connected in dimension and the feature meanings are independent, so that the model can distinguish well. It can not only learn the data distribution of the rPPG signal from the noise rPPG signal restoration process, but also use this distribution to guide the mapping of MSTmap to the real rPPG signal.

[0169] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A remote physiological signal estimation method based on a diffusion model, characterized in that: The training process of the diffusion model is as follows: S11, obtaining a sample set, where the sample includes a face video and an original rPPG signal corresponding to the face video; S12, forward diffusion is performed on the original rPPG signal, i.e., noise is added to obtain a noisy rPPG signal; S13, extracting features from the face video to obtain a multi-scale spatiotemporal graph; S14, fusing the noisy rPPG signal and the multi-scale spatiotemporal graph to obtain fused data; S15, performing back diffusion, i.e., denoising, on the fused data to obtain denoised data; S16, outputting a predicted rPPG signal based on the denoised data; S17, performing model training by minimizing the difference between the predicted rPPG signal and the original rPPG signal to obtain a trained diffusion model; The trained diffusion model is used to estimate long-range physiological signals as follows: S21, inputting a face video to be estimated and generating a noise rPPG signal; S22, extracting features from the estimated face video to obtain a multi-scale spatiotemporal graph; S23, fusing the generated noise rPPG signal and the multi-scale spatiotemporal graph to obtain fused data; S24, performing denoising processing on the fused data to obtain denoised data; S25, outputting a predicted rPPG signal based on the denoised data; In step S15, during the training process, the fused data is denoised using a denoising module, i.e., backward diffusion. The denoising module uses a mixSTE, i.e., a hybrid space-time encoder, which consists of c layers, each layer includes a Spatial Transformer block and a Temporal Transformer block connected in sequence; the Spatial Transformer block is used to extract the spatial features of the fused data, and the Temporal Transformer block is used to extract the spatiotemporal features of the fused data; after mixSTE, a Resnet residual connection is used; the denoising expression during training is as follows: ; ; in, y N To fuse the data, z is the feature vector after backward diffusion, that is, the denoised data; mixSTE(·) is the processing function of the mixed space-time encoder; SpatialTransformer(·) is the processing function of the Spatial Transformer block; TemporalTransformer(·) is the processing function of the Temporal Transformer block.

2. The remote physiological signal estimation method based on diffusion model according to claim 1, characterized in that: In step S12, the expression of forward diffusion is as follows: ; in, x 0 is the original rPPG signal; x t For the t Noise rPPG signal after forward diffusion; t is the number of forward diffusions, t =1,2,..., N , and finally obtained N Noise rPPG signal after forward diffusion x N ; q ( x t | x 0) indicates a given initial state x In the case of 0, the status x t The conditional probability distribution of ; is the attenuation factor, In the diffusion process, as the diffusion times t decreases gradually with the increase of For noise.

3. The remote physiological signal estimation method based on diffusion model according to claim 1, characterized in that: In step S13 and step S22, feature extraction is performed on the face video, and the specific method is as follows: Divide the face video into multiple segments, and get F Frame video image; Perform face detection on video images and divide the faces into n A region of interest is ROI; Will n ROIs are combined to obtain a combination of a single ROI and different ROIs, a total of 2 n - 1 combination; For each combination, the average pixel value of all ROIs in the combination in each color channel is calculated; there are 6 color channels in total, namely RGB and YUV; Finally, the dimension is 6×(2 n -1)× F The vector of is used as a multi-scale space-time graph.

4. The remote physiological signal estimation method based on diffusion model according to claim 3, characterized in that: In step S14 and step S23, the noisy rPPG signal and the multi-scale spatiotemporal graph are fused, and the specific method is as follows: The dimension of the noisy rPPG signal is 1× F , the dimension of the noisy rPPG signal is expanded to 1×(2 n -1)× F , the expanded noisy rPPG signal is spliced ​​with the multi-scale spatiotemporal graph in the channel dimension to obtain the fused data, the dimension of the fused data is 7×(2 n -1)× F .

5. The remote physiological signal estimation method based on diffusion model according to claim 1, characterized in that: The Spatial Transformer block is processed as follows: First, the input data is the fusion data y N Using the linear mapping function, we get the three parameters of the spatial self-attention mechanism: Q s 、K s 、V s : ; Then, the fusion data is calculated through the self-attention mechanism Attention y N Spatial characteristics S : ; The Temporal Transformer block is processed as follows: First, use the linear mapping function on the spatial feature S to obtain the three parameters of the temporal self-attention mechanism, namely Q t 、K t 、 V t : ; Then, the fusion data is calculated through the self-attention mechanism Attention y N The spatiotemporal characteristics of ST : ; Spatial and temporal characteristics ST That is: 。 6. The remote physiological signal estimation method based on diffusion model according to claim 1, characterized in that: In step S24, during the estimation process, the fusion data is denoised using a denoising module, and the denoising process is repeated. K After backward expansion, the denoised data is finally output; the denoising expression in the estimation process is as follows: z k = z k-1 +mixSTE( z k-1 ); in, k is the number of iterations of back diffusion, k =1,2,..., K ; z k For the k The eigenvector of the backward diffusion, z 0 is the fusion data, and finally the K The eigenvector of the backward diffusion z K , which is the denoised data.

7. The remote physiological signal estimation method based on diffusion model according to claim 1, characterized in that: In step S12 and step S21, the noise rPPG signal is Gaussian noise that obeys a standard normal distribution.

8. A remote physiological signal estimation system based on a diffusion model, characterized in that: A remote physiological signal estimation method based on a diffusion model applicable to any one of claims 1 to 7, the system comprising: a forward diffusion module, an MSTmap calculation module, a fusion module, a denoising module, a prediction module and a sampling module; the working process of the system is divided into two parts, namely a training process and an estimation process, as shown below: During training: A forward diffusion module is used to add noise to the original rPPG signal to obtain a noisy rPPG signal; MSTmap calculation module, used to process face videos to obtain multi-scale spatiotemporal maps containing visual features; A fusion module is used to fuse the noisy rPPG signal with the multi-scale spatiotemporal graph to obtain fused data; A denoising module is used to denoise the fused data to obtain denoised data; A prediction module, used to predict the rPPG signal based on the denoised data output; During the estimation process: A sampling module, used for random sampling to generate noise rPPG signals; MSTmap calculation module, used to process the estimated face video to obtain a multi-scale spatiotemporal map containing visual features; A fusion module is used to fuse the noisy rPPG signal with the multi-scale spatiotemporal graph to obtain fused data; A denoising module is used to denoise the fused data to obtain denoised data; The prediction module is used to output the predicted rPPG signal based on the denoised data.

9. A computer program product, characterized in that It includes a computer program / instruction, which, when executed by a processor, implements a remote physiological signal estimation method based on a diffusion model as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Non-contact heart rate measurement method based on space-time attention network and input optimization

    CN113343821A

  • End-to-end remote heart rate detection method based on channel enhanced space-time attention network

    CN114912487A