Pulse analysis method, device and equipment and storage medium
Pulse analysis is carried out through the space-time state-space model of image acquisition equipment and dual characteristics, and the problems of inconvenience of finger-clip-type oxygen meter and insufficient real-time video monitoring are solved, and non-contact and immediate feedback pulse monitoring is realized, reducing equipment costs and resource consumption.
Patent Information
- Application Number
- CN202510410817.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-08
AI Technical Summary
The existing finger-clip-type oxygen meter has inconvenience during use, especially for the elderly and people with limited hand movements. The video-based pulse waveform monitoring has problems such as real-time and insufficient resource consumption.
The face image frame is obtained through the image acquisition device, the initial state space is set, and the pulse analysis is performed using the space-time state space model with dual characteristics, so as to realize contactless monitoring and reduce equipment costs and resource consumption.
It realizes contactless monitoring of instant feedback pulse values, reduces equipment costs, improves user comfort and monitoring convenience, and is especially suitable for the elderly and people with limited hand movements.
Smart Images

Figure CN120278982A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of medical and health monitoring, and particularly to a method, device, equipment and storage medium for pulse analysis. Background Art
[0002] Finger clip oximeters are commonly used to measure physiological signals of pulse waves, such as blood oxygen saturation and heart rate. This device works by inserting a finger into a small clip that contains a light emitter and a detector, and evaluates the blood oxygen level and other physiological parameters by monitoring the change of light passing through the tissue.
[0003] Although finger clip oximeters provide a relatively convenient way to monitor these important physiological indicators, their usage also has certain limitations. First, putting a finger into the clip may be inconvenient, especially during long-term use, which may cause discomfort or a sense of restraint to the user. Second, for some people, such as the elderly or individuals with limited hand mobility, using a finger clip oximeter may be more difficult because they may have difficulty accurately inserting their finger into the clip or keeping their finger still during use to ensure the accuracy of the measurement. Finally, if there are wounds or other skin problems on the finger, using a finger clip oximeter may cause unnecessary pressure or pain. Therefore, although finger clip oximeters are an effective monitoring tool, the contact method in actual use brings inconvenience to some users.
[0004] Therefore, how to improve the comfort and adaptability of using finger clip oximeters, especially in the case of long-term use, special populations or skin problems, and ensure accuracy and convenience, is a technical problem that needs to be urgently solved by those skilled in the art. Summary of the Invention
[0005] Based on the above problems, this application provides a method, device, equipment and storage medium for pulse analysis, which can overcome the differences in visible light and infrared image modalities, improve the accuracy of image registration and fusion, optimize the fusion effect and enhance the final image quality.
[0006] The embodiments of this application disclose the following technical solutions:
[0007] A method for pulse analysis, the method includes:
[0008] Obtaining a face image frame to be analyzed through an image acquisition device, and setting an initial state space;
[0009] Input the face image frame to be analyzed and the initial state space into a pre-constructed pulse analysis model for pulse analysis, to obtain a first predicted pulse value and an updated state space; the pre-constructed pulse analysis model is obtained by training a model based on a spatio-temporal state space model with dual characteristics; the initial state space is used to provide historical background and dynamic context information for the face image frame to be analyzed when the pulse analysis model performs pulse analysis; the updated state space is used as the initial state space for the next face image frame to be analyzed; the next face image frame is the next face image frame to be analyzed obtained by an image acquisition device; the acquisition time of the next face image frame is later than that of the face image frame to be analyzed;
[0010] Output the first predicted pulse value and the updated state space.
[0011] In a possible implementation manner, the construction process of the pre-constructed pulse analysis model includes:
[0012] Obtain a historical video segment, and extract face video frames in the historical video segment to obtain historical face video frames;
[0013] For each of the historical face video frames, obtain a historical true-value pulse value synchronized with the historical face video frame;
[0014] Input multiple historical face video frames into the spatio-temporal state space model for model training to obtain the pulse analysis model;
[0015] Use the multiple historical true-value pulse values to optimize the pulse analysis model.
[0016] In a possible implementation manner, the layer structure of the spatio-temporal state space model with dual characteristics includes a first projection layer, a first spatio-temporal state space module, a second spatio-temporal state space module, a third spatio-temporal state space module, a fourth spatio-temporal state space module, a weighted summation module, a downsampling module, and a prediction head; the input end of the first spatio-temporal state space module is connected to the output end of the first projection layer; the input end of the second spatio-temporal state space module is connected to the output end of the first spatio-temporal state space module; the input end of the third spatio-temporal state space module is connected to the output end of the second spatio-temporal state space module; the input end of the fourth spatio-temporal state space module is connected to the output end of the third spatio-temporal state space module; the output end of the fourth spatio-temporal state space module is connected to the weighted summation module; the weighted summation module is connected to the downsampling module; the downsampling module is connected to the prediction head; the spatio-temporal state space module has dual characteristics;
[0017] Among them, the first projection layer is used to perform a linear transformation on the multiple historical face video frames to obtain first linear transformation features; the first spatio-temporal state space module is used to perform temporal normalization processing and spatio-temporal dual processing on the first linear transformation features to obtain a first dual state space representation; the second spatio-temporal state space module is used to perform temporal normalization processing and spatio-temporal dual processing on the first dual state space representation to obtain a second dual state space representation; the third spatio-temporal state space module is used to perform temporal normalization processing and spatio-temporal dual processing on the second dual state space representation to obtain a third dual state space representation; the fourth spatio-temporal state space module is used to perform temporal normalization processing and spatio-temporal dual processing on the third dual state space representation to obtain a fourth dual state space representation; the weighted summation module is used to perform a weighted summation on the fourth dual state space representation and the first linear transformation features to obtain a comprehensive feature representation; the downsampling module is used to perform dimensionality reduction processing on the comprehensive feature representation to obtain a dimensionality-reduced feature representation; the prediction head is used to perform a prediction on the dimensionality-reduced feature representation to obtain a pulse value prediction result; the pulse value prediction result includes a plurality of second predicted pulse values; the pulse value prediction result is used to compare with the plurality of historical true value pulse values for model optimization.
[0018] In a possible implementation manner, each of the spatio-temporal state space modules includes a temporal normalization module, a first spatio-temporal dual module, and a second spatio-temporal dual module; the output end of the temporal normalization module is connected to the input end of the first spatio-temporal dual module; the output end of the first spatio-temporal dual module is connected to the input end of the second spatio-temporal dual module; the output end of the second spatio-temporal dual module is connected to the weighted summation module;
[0019] Among them, the temporal normalization module is used to extract the temporal correlation features between the respective historical face video frames, and align and standardize the multiple historical face video frames based on the temporal correlation features to obtain a normalized video; the first spatio-temporal dual module is used to perform a first spatial two-dimensional convolution, a temporal causal convolution, a state space equivalent decoupling, and a second spatial two-dimensional convolution on the normalized video in combination with the temporal correlation features to obtain an initial dual state space representation; the second spatio-temporal dual module is used to perform a first spatial two-dimensional convolution, a temporal causal convolution, a state space equivalent decoupling, and a second spatial two-dimensional convolution on the initial dual state space representation in combination with the temporal correlation features to obtain the fourth dual state space representation.
[0020] In a possible implementation, each of the spatio-temporal dual modules includes a first spatial two-dimensional convolutional layer, a second projection layer, a temporal causal convolutional layer, a first activation function layer, a second activation function layer, a dual state space layer, a matrix multiplication layer, a normalization layer, a third projection layer, and a second spatial two-dimensional convolutional layer; the input ends of the first spatial two-dimensional convolutional layer and the third projection layer are both connected to the output end of the temporal normalization module; the output end of the first spatial two-dimensional convolutional layer is connected to the input end of the second projection layer; the output end of the second projection layer is respectively connected to the input ends of the temporal causal convolutional layer, the first activation function layer, and the dual state space layer; the output end of the temporal causal convolutional layer is connected to the input end of the second activation function layer; the output end of the first activation function layer is connected to the input end of the matrix multiplication layer; the output end of the dual state space layer is connected to the input end of the matrix multiplication layer; the output end of the second activation function layer is connected to the input end of the dual state space layer; the output end of the matrix multiplication layer is connected to the input end of the normalization layer; the output end of the normalization layer is connected to the input end of the third projection layer; the output end of the third projection layer is connected to the input end of the second spatial two-dimensional convolutional layer;
[0021] Among them, the first spatial two-dimensional convolutional layer is used to perform spatial two-dimensional convolution on the normalized video to obtain a spatio-temporal feature map; the second projection layer is used to perform a linear transformation on the spatio-temporal feature map to obtain a second linearly transformed feature; the temporal causal convolutional layer is used to perform temporal causal convolution on the second linearly transformed feature to obtain a temporal feature; the first activation function layer is used to activate the temporal feature to obtain an activated temporal feature; the second activation function layer is used to activate the transformed feature to obtain an activated transformed feature; the dual state space layer is used to perform a spatio-temporal dual operation on the activated transformed feature and the transformed feature to obtain a dual state space representation; the matrix multiplication layer is used to perform a matrix multiplication operation on the dual state space representation and the activated temporal feature to obtain a fused feature; the normalization layer is used to normalize the fused feature to obtain a normalized feature; the third projection layer is used to perform a non-linear transformation on the normalized feature based on the normalized video to obtain a non-linearly transformed feature; the second spatial two-dimensional convolutional layer is used to perform spatial two-dimensional convolution on the non-linearly transformed feature to obtain the dual state space representation.
[0022] In a possible implementation, the extraction steps of the temporal normalization module sequentially include pixel change residual calculation, standard deviation calculation of the pixel change residual, and normalization operation on each frame of the image.
[0023] A pulse analysis device, the device includes:
[0024] A first acquisition unit, configured to acquire a face image frame to be analyzed;
[0025] A setting unit for setting an initial state space;
[0026] A pulse analysis unit for inputting the face image frame to be analyzed and the initial state space into a pre - constructed pulse analysis model for pulse analysis, obtaining a first predicted pulse value and an updated state space; the pre - constructed pulse analysis model is obtained by training a spatio - temporal state space model with dual characteristics; the initial state space is used to provide historical background and dynamic context information for the face image frame to be analyzed when the pulse analysis model performs pulse analysis; the updated state space is used as the initial state space for the next face image frame to be analyzed; the next face image frame is the next face image frame to be analyzed acquired by an image acquisition device; the acquisition time of the next face image frame is later than that of the face image frame to be analyzed;
[0027] An output unit for outputting the first predicted pulse value and the updated state space.
[0028] In a possible implementation manner, the device further includes:
[0029] A second acquisition unit for acquiring a historical video segment;
[0030] An extraction unit for extracting face video frames from the historical video segment to obtain historical face video frames;
[0031] An integration unit for, for each of the historical face video frames, acquiring a historical true - value pulse value synchronized with the historical face video frame and annotating the historical true - value pulse value on the historical face video frame;
[0032] A model training unit for inputting multiple historical face video frames annotated with historical true - value pulse values into the spatio - temporal state space model for model training to obtain the pulse analysis model.
[0033] A pulse analysis device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, when the processor executes the computer program, implementing the pulse analysis method as described above.
[0034] A computer - readable storage medium storing instructions, when the instructions run on a terminal device, causing the terminal device to execute the pulse analysis method as described above.
[0035] Compared with the prior art, the present application has the following beneficial effects:
[0036] The present application provides a method, apparatus, device, and storage medium for pulse analysis. Specifically, when implementing the pulse analysis method provided by the embodiments of the present application, first, a face image to be analyzed is obtained through an image acquisition device, and an initial state space is set. The initial state space is used to provide historical background and dynamic context information for the face image frame to be analyzed when performing pulse analysis in the pulse analysis model. Then, the image frame and the initial state space are input into a pre-constructed pulse analysis model for analysis to obtain a first predicted pulse value and an updated state space. The pulse analysis model is obtained by training based on a spatio-temporal state space model with dual characteristics, which can effectively handle temporal and spatial changes and extract minute dynamic features related to the pulse. Finally, the obtained first predicted pulse value and the updated state space are output to achieve accurate prediction of the pulse. The present application analyzes a single face image to obtain pulse data points, thereby achieving stronger real-time performance and being able to provide pulse value feedback immediately. The present application is trained through a spatio-temporal state space model with dual characteristics, jointly modeling temporal and spatial changes, effectively extracting minute blood flow dynamic features from a single frame image, and accurately predicting the pulse waveform. Even without continuous video frames, the model can still calculate accurate pulse data, has stronger real-time performance, and can provide immediate feedback of the pulse value. In addition, compared with a finger clip oximeter, the present application realizes non-contact physiological monitoring, avoids the discomfort of traditional devices, and is particularly suitable for the elderly and people with limited hand mobility. Different from traditional methods that require special equipment and complex algorithms, the present application only needs a common image acquisition device (such as a mobile phone camera) for pulse analysis, reducing the equipment cost and improving the convenience and popularity of the technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] To more clearly illustrate the technical solutions in the embodiments or the prior art, the following briefly introduces the drawings required for describing the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0038] Figure 1 It is a flowchart of a method for a pulse analysis method provided by an embodiment of the present application;
[0039] Figure 2 It is a flowchart of a method for constructing a pulse analysis model provided by an embodiment of the present application;
[0040] Figure 3 It is a schematic structural diagram of a spatio-temporal state space model provided by an embodiment of the present application;
[0041] Figure 4Schematic diagram of a spatio-temporal state space module provided by an embodiment of the present application;
[0042] Figure 5 Schematic diagram of a spatio-temporal dual module provided by an embodiment of the present application;
[0043] Figure 6 Schematic diagram of a pulse analysis device provided by an embodiment of the present application. Detailed implementation manners
[0044] To facilitate the understanding of the technical solutions provided by the embodiments of the present application, the background art related to the embodiments of the present application will be described first.
[0045] Although finger clip oximeters can conveniently monitor blood oxygen saturation and heart rate, their clamping methods may cause discomfort to users, especially for the elderly or those with limited hand mobility.
[0046] However, in recent years, remote physiological measurement methods based on face videos have gradually attracted attention and made some progress to a certain extent. This method monitors the pulse waveform by analyzing face videos, enabling non-contact measurement and avoiding physical contact with traditional finger clip oximeters. Although the video-based measurement method has advantages in non-contact monitoring, it still has some significant limitations. Different from oximeters, video-based pulse waveform monitoring cannot provide real-time data feedback. Usually, this method needs to record a long video first, and then process the video frames through a complex algorithm model to extract the pulse waveform. This is because the pulse wave is a tiny periodic signal, and only through a long video clip can the noise be effectively removed and the original signal be retained. Due to the weakness of the pulse wave, it is difficult to capture obvious and stable physiological changes in a single video frame. Therefore, in order to accurately extract the pulse waveform, it is necessary to analyze the subtle changes between multiple consecutive frames and remove the interference of noise such as ambient light changes and head movements.
[0047] In addition, video-based pulse wave monitoring also faces the problem of high algorithm memory occupancy. Since a large amount of video frame data needs to be processed, this requires the system to have a large storage space and computing power, which may cause performance bottlenecks in practical applications. Therefore, although this method provides a new direction for non-contact physiological monitoring, there are still significant deficiencies in terms of real-time performance and resource consumption.
[0048] To solve this problem, an embodiment of the present application provides a method, apparatus, device, and storage medium for pulse analysis. First, a face image frame to be analyzed can be obtained through an image acquisition device, and an initial state space is set. The initial state space is used to provide historical background and dynamic context information for the face image frame to be analyzed when performing pulse analysis in the pulse analysis model. Subsequently, these face image frames to be analyzed and the initial state space are input into a pre-constructed pulse analysis model for pulse analysis, thereby obtaining a first predicted pulse value and an updated state space. The pre-constructed pulse analysis model here is obtained by training a model based on a spatio-temporal state space model with dual characteristics. Finally, the obtained first predicted pulse value and the updated state space are output, thus completing the entire pulse analysis process. The pulse analysis model obtained by training a model based on a spatio-temporal state space model with dual characteristics in the present application jointly models the changes in time and space. This model design enables it to effectively extract minute blood flow dynamic features from the input single-frame image and perform accurate pulse prediction. Even without a large number of consecutive video frames, the model can deduce accurate pulse waveform data, so it has stronger real-time performance and can provide instant feedback of the pulse value. This method greatly reduces the latency, enabling more efficient monitoring in practical applications. At the same time, since only a single-frame image needs to be analyzed, the present application avoids the high memory occupancy and computational power requirements caused by the need to process a large amount of video frame data in traditional methods, thus alleviating the problem of performance bottlenecks. This means that the system can perform efficient pulse monitoring with lower resource consumption and has a higher energy efficiency ratio. In addition, compared with traditional finger clip oximeters, the present application can achieve non-contact physiological monitoring, avoiding the discomfort caused by traditional devices, especially suitable for the elderly and people with limited hand mobility. This non-contact monitoring method not only improves the user's comfort but also provides convenience for long-term monitoring. Compared with traditional methods that require dedicated devices and complex algorithms to process video data, the present application only needs to use a common image acquisition device (such as a mobile phone camera) to obtain a face image and perform subsequent pulse analysis. This not only reduces the device cost but also makes this technology more convenient and easier to popularize. In summary, the present application has significant advantages in terms of real-time performance, resource consumption, user comfort, and device cost, greatly improving the convenience and practicality of pulse monitoring.
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0050] See Figure 1, This figure is a method flowchart of a pulse analysis method provided by an embodiment of the present application. As Figure 1 shown, the pulse analysis method may include steps S101 - S103:
[0051] S101: Obtain a face image frame to be analyzed through an image acquisition device and set an initial state space.
[0052] To implement pulse measurement based on a single - frame image, first use a device, an image acquisition device (such as a mobile phone camera, glasses camera, or camera), to obtain a face image frame to be analyzed. Specifically, the image acquisition device (such as a smartphone camera) takes a picture containing a face, and this picture will be used as input data to extract minute blood flow dynamic features of the face, and then to predict and analyze the pulse value.
[0053] After capturing the image frame, set an initial state space for this image frame. The initial state space is a starting state prepared for subsequent processing by the pulse analysis model. It carries historical data and background information from the previous moment. This initial state space provides necessary context support for the pulse analysis model, enabling the model to understand the temporal dependence relationship between image frames and the dynamic changes in blood flow. By providing a clear starting point for each frame of the image, the initial state space ensures that the model can combine historical data with the content of the current image frame to more accurately predict the pulse and make dynamic adjustments, thereby improving the accuracy and reliability of the analysis. This step is crucial in the pulse analysis process, ensuring the coherence of information transfer during model processing and enabling effective tracking of the pulse dynamics of consecutive frames.
[0054] The main task of the image acquisition device is to obtain clear and stable images that contain the facial features of the person to be analyzed. Since the pulse waveform is related to minute changes in the color of the facial skin, the accuracy and resolution of the device are crucial for the accuracy of subsequent analysis.
[0055] S102: Input the face image frame to be analyzed and the initial state space into a pre - constructed pulse analysis model for pulse analysis, and obtain a first predicted pulse value and an updated state space.
[0056] The face image frame to be analyzed and the initial state space are input into a pre-constructed pulse analysis model for pulse analysis. The model is trained based on a spatio-temporal state space model with dual characteristics and can accurately predict the pulse of the input data. Through analysis, the model outputs two results: the first predicted pulse value, which is the pulse prediction value of the current input data, and the updated state space, which combines the pulse information in the image frame with other relevant data. This updated state space will be input as the initial state space of the next face image frame to be analyzed, providing a basis for further analysis. At the same time, the next face image frame to be analyzed is obtained by an image acquisition device, and its acquisition time is later than that of the current face image frame to be analyzed, ensuring that the pulse prediction process is synchronized with the actual time and providing continuous input data for the model for real-time pulse monitoring.
[0057] It should be noted that by training with a spatio-temporal state space model with dual characteristics, the changing characteristics of time and space can be fully captured, combined with the dynamic change law of the pulse waveform, so that the pulse analysis model can extract minute blood flow dynamic characteristics from the input "single-frame" video frame. Through the joint modeling of spatio-temporal information, the model can effectively capture the subtle changes in the pulse waveform. Even in the absence of continuous video frames, it can still deduce the pulse waveform data from a single-frame image. This is because the spatio-temporal model can make full use of the local spatial information and potential time series characteristics in the image, so as to achieve accurate prediction of the pulse waveform.
[0058] See Figure 2 , Figure 2 is a method flowchart of a method for constructing a pulse analysis model provided by an embodiment of the present application, which can be specifically implemented through steps A1 - A3:
[0059] A1: Obtain historical video segments, and extract face video frames from the historical video segments to obtain historical face video frames.
[0060] First, extract face video frames from historical video segments. These video segments are previously recorded images containing faces, and each extracted frame image is a face video frame, which contains visual information related to the pulse.
[0061] A2: For each of the historical face video frames, obtain historical ground truth pulse values synchronized with the historical face video frames.
[0062] For each extracted historical face video frame, historical ground truth pulse values synchronized with it need to be obtained. These pulse values are accurately measured pulse data, usually obtained through traditional medical devices or sensors.
[0063] A3: Input multiple of the historical face video frames into the spatio-temporal state space model for model training to obtain the pulse analysis model.
[0064] Next, input multiple of the extracted face video frames into a spatio-temporal state space model for training. These historical face video frames contain visual information related to the pulse, and the spatio-temporal state space model can capture and analyze the time-varying features in the video frames. Through this training process, the model can learn the change patterns related to the pulse in the face images and finally generate a model that can accurately analyze and predict the pulse.
[0065] A4: Use the multiple historical true pulse values to optimize the pulse analysis model.
[0066] After the model training is completed, use multiple historical true pulse values to optimize the pulse analysis model. By comparing with the actual pulse data, adjust the model parameters to reduce the error between the predicted value and the true value, and improve the accuracy and robustness of the model. The optimization process is usually completed through techniques such as backpropagation and gradient descent to ensure that the model can accurately predict the pulse value.
[0067] See Figure 3 , Figure 3 which is a schematic structural diagram of a spatio-temporal state space model provided by an embodiment of the present application. The layer structure of the spatio-temporal state space model with dual characteristics includes multiple key modules that work together to achieve efficient and accurate prediction of pulse analysis.
[0068] Specifically, the structure includes a first projection layer, 4 identical spatio-temporal state space modules (i.e., the first spatio-temporal state space module, the second spatio-temporal state space module, the third spatio-temporal state space module, and the fourth spatio-temporal state space module), a weighted summation module (i.e., the ⊕ in Figure 3 ), a downsampling module (i.e., the DS module in Figure 3 ), and a prediction head. Each module plays an important role in the pulse prediction process. The 4 identical spatio-temporal state space modules are connected in sequence. The output end of the last spatio-temporal state space module (i.e., the fourth spatio-temporal state space module) connected in sequence is connected to the weighted summation module. The weighted summation module is connected to the downsampling module. The downsampling module is connected to the prediction head. The spatio-temporal state space module has dual characteristics.
[0069] First projection layer: First, multiple historical face video frames will undergo a linear transformation through the first projection layer to obtain the first linearly transformed features. This process is mainly used to extract the spatio-temporal features in the video frames and provide a meaningful information representation for subsequent processing.
[0070] Spatio-temporal State Space Module: Four identical spatio-temporal state space modules receive the output from the first projection layer. Through temporal normalization processing and spatio-temporal duality processing, a dual state space representation of the output is obtained. The dual characteristics of the spatio-temporal state space module can effectively capture the spatio-temporal change laws in video frames and enhance the dynamic characteristics of the model.
[0071] Weighted Summation Module: The output of the fourth spatio-temporal state space module will be combined with the first linear transformation feature by the weighted summation module to obtain a comprehensive feature representation. The process of weighted summation is to better fuse features for subsequent dimensionality reduction and prediction.
[0072] Downsampling Module: The comprehensive feature representation after weighted summation will enter the downsampling module for dimensionality reduction processing. This step can reduce the feature dimension, retain important information, and reduce computational complexity in preparation for the final prediction.
[0073] Prediction Head: Finally, the dimensionality-reduced features will be input into the prediction head for pulse value prediction. The prediction head outputs multiple second predicted pulse values based on the input dimensionality-reduced features. These prediction results will be compared with multiple historical true pulse values as the basis for model optimization to further improve the prediction accuracy.
[0074] The entire model can achieve efficient and accurate pulse prediction through the dual characteristics of the spatio-temporal state space module, combined with the processing of modules such as weighted summation and downsampling, and self-optimize by comparing with historical true pulse values.
[0075] See Figure 4 , Figure 4 which is a schematic structural diagram of a spatio-temporal state space module provided by an embodiment of this application. The architecture of the spatio-temporal state space module further details its components. The spatio-temporal state space module specifically includes a temporal normalization module, a first spatio-temporal duality module, and a second spatio-temporal duality module. The output end of the temporal normalization module is connected to the input end of the first spatio-temporal duality module. The output end of the first spatio-temporal duality module is connected to the input end of the second spatio-temporal duality module. The output end of the second spatio-temporal duality module is connected to the weighted summation module.
[0076] The functions of each module in the pulse prediction process are as follows:
[0077] Temporal Normalization Module: Its main function is to process the temporal correlation features between historical face video frames. By analyzing the temporal information between video frames, the temporal normalization module aligns and standardizes multiple historical video frames to obtain a normalized video sequence. This process ensures the consistency and alignment of video frames in the time dimension and provides accurate input for subsequent spatio-temporal feature extraction.
[0078] The first spatiotemporal dual module: This module receives the output from the time normalization module and processes it in combination with the temporal correlation features. The specific steps include:
[0079] First spatial 2D convolution: extract features in the spatial dimension and understand the static information in the video frame.
[0080] Temporal Causal Convolution: Processes information in the time dimension and ensures that the causal relationship of the time series is fully utilized through causal convolution.
[0081] State space equivalent decoupling: The core of this step is to decouple the spatiotemporal features through the state space model and find the independent relationship between different time and space dimensions.
[0082] Secondary spatial 2D convolution: further extracts spatial features and helps capture spatial changes more finely.
[0083] After these steps, the initial dual state space representation is obtained, which is the module’s summary representation of the spatiotemporal features of the input video.
[0084] Second space-time dual module:
[0085] The structure is similar to the first space-time dual module. The second space-time dual module also combines the time series correlation features for processing, but the processing object at this time is the initial dual state space representation. The processing steps include: first spatial two-dimensional convolution, temporal causal convolution, state space equivalent decoupling and secondary spatial two-dimensional convolution.
[0086] The output of this module is the fourth dual state space representation, which is another representation of spatiotemporal features.
[0087] Finally, the fourth dual state space representation is weighted and summed with the first linear transformation feature to obtain a combined comprehensive feature representation. This combined representation contains the spatiotemporal information from the two spatiotemporal dual modules and can more comprehensively reflect the spatiotemporal characteristics of historical video frames.
[0088] This architecture, by combining multiple convolution operations in time and space and state-space decoupling methods, can deeply mine the complex spatiotemporal patterns in video frames and provide more accurate input features for the final pulse prediction.
[0089] In a possible implementation, the extraction steps of the time normalization module sequentially include calculating pixel change residuals, calculating standard deviations of the pixel change residuals, and performing normalization operations on each frame of the image.
[0090] Pixel change residual calculation: The goal of this stage is to calculate the pixel changes between adjacent frames. This is usually done by computing the pixel differences between each video frame and its previous frame to obtain the residual of pixel changes. In this way, the temporal differences between video frames can be captured, and the subtle changes between frames can be extracted to help capture dynamic information.
[0091] Standard deviation calculation of residuals: By calculating the standard deviation of the pixel change residuals, the temporal normalization module can further quantify the magnitude of changes between video frames. The standard deviation can provide the degree of dispersion of the residual data, indicating the regions and intensities of changes in the video frames. This stage helps to determine which regions have significant changes and which regions have minor changes, making the subsequent processing more targeted.
[0092] Normalization operation on each frame of image: Finally, the temporal normalization module performs a normalization process on each frame. This operation ensures that the data ranges of all video frames are consistent, avoiding biases caused by differences in brightness or contrast between different frames. The normalization operation usually includes adjusting the pixel values of each frame of image to a unified scale for subsequent spatio-temporal feature extraction.
[0093] Through these steps, the temporal normalization module can effectively perform temporal correlation processing on the input video frames and ensure that the temporal information of each frame is effectively utilized in subsequent analysis. This temporal normalization processing makes the subsequent spatio-temporal feature extraction more accurate and reduces the interference of external factors (such as light changes) on the prediction results.
[0094] See Figure 5 , Figure 5A schematic structural diagram of a spatio-temporal dual module provided by an embodiment of the present application. The spatio-temporal dual module specifically includes multiple levels such as spatio-temporal feature extraction, activation, linear transformation, and convolution operations. It is a combination of a multi-layer convolutional neural network (CNN) and temporal causal convolution, aiming to process video data or spatio-temporal sequence data. Each of the spatio-temporal dual modules includes a first two-dimensional spatial convolution layer, a second projection layer, a temporal causal convolution layer, a first activation function layer, a second activation function layer, a dual state space layer, a matrix multiplication layer, a normalization layer, a third projection layer, and a second two-dimensional spatial convolution layer; the input ends of the first two-dimensional spatial convolution layer and the third projection layer are both connected to the output end of the temporal normalization module; the output end of the first two-dimensional spatial convolution layer is connected to the input end of the second projection layer; the output end of the second projection layer is respectively connected to the input ends of the temporal causal convolution layer, the first activation function layer, and the dual state space layer; the output end of the temporal causal convolution layer is connected to the input end of the second activation function layer; the output end of the first activation function layer is connected to the input end of the matrix multiplication layer; the output end of the dual state space layer is connected to the input end of the matrix multiplication layer; the output end of the second activation function layer is connected to the input end of the dual state space layer; the output end of the matrix multiplication layer is connected to the input end of the normalization layer; the output end of the normalization layer is connected to the input end of the third projection layer; the output end of the third projection layer is connected to the input end of the second two-dimensional spatial convolution layer.
[0095] The first two-dimensional spatial convolution layer: used to perform two-dimensional convolution operations on the input normalized video in the space, so as to extract spatial features and generate spatio-temporal feature maps.
[0096] The second projection layer: performs a linear transformation on the spatio-temporal feature map and maps it to another feature space.
[0097] The temporal causal convolution layer: uses temporal causal convolution to capture the temporal dynamic features in the video. Causal convolution is usually used to process time series data to ensure that information only propagates from the past to the future.
[0098] The first activation function layer: performs non-linear activation on the features after temporal causal convolution to increase the expression ability of the network.
[0099] The second activation function layer: activates the features transformed by the second projection layer to further enhance the learning ability of the model.
[0100] The dual state space layer: performs spatio-temporal dual operations, converts spatio-temporal features into dual space representations, and may be used to model the dependencies between space and time.
[0101] Matrix multiplication layer: Performs matrix multiplication on the dual state space representation and the activation time features to obtain fused features, which may be a key step for the model to fuse information.
[0102] Normalization layer: Normalizes the fused features to ensure stability and effectiveness during network training.
[0103] Third projection layer: Performs a non-linear transformation on the normalized features to obtain a more complex feature representation.
[0104] Second spatial two-dimensional convolutional layer: Convolves the features after non-linear transformation to further extract spatial features, which may be used for the final decision or output.
[0105] This network architecture is very complex and suitable for processing data that requires simultaneous consideration of spatial and temporal information, such as videos or time-series data. The operations of each layer and the connections between layers reflect the common feature extraction, information fusion, activation, and transformation processes in deep learning models.
[0106] In a possible implementation, the dual state space is divided into two modes: a structured state space (which can be updated as part of the model parameters and exists in matrix form), and a discrete state space (which exists as an independent, propagable external parameter of the model during inference, in vector form, and is input into the model together with a single-frame video).
[0107] S103: Output the first predicted pulse value and the updated state space.
[0108] After the model processes the input image frame and the initial state space, it outputs the predicted pulse value and the updated state space. This process realizes the function of extracting pulse information from a single image frame, improving the real-time performance and accuracy of detection.
[0109] Based on the content of S101 - S103, the working process of the pulse analysis model includes three main steps: First, obtain a face image frame to be analyzed through an image acquisition device and set an initial state space; then, input the face image frame and the initial state space into a pre - constructed pulse analysis model for analysis to obtain a first predicted pulse value and an updated state space. The pulse analysis model is trained based on a spatio - temporal state space model with dual characteristics; finally, output the first predicted pulse value and the updated state space. Through training with a spatio - temporal state space model with dual characteristics, the obtained pulse analysis model of the present application can capture changes in both time and space simultaneously, thus effectively extracting minute blood flow dynamic features from a single - frame image and making accurate pulse predictions. Even without continuous video frames, the model can calculate accurate pulse waveform data, with stronger real - time performance and immediate feedback capabilities. Compared with a finger - clip oximeter, this method realizes non - contact physiological monitoring, reduces the discomfort of users, and is especially suitable for the elderly and people with limited hand mobility. In addition, it only needs to use ordinary image acquisition devices (such as a mobile phone camera) for face image acquisition and analysis, reducing equipment costs, and the technology is more convenient and easier to popularize.
[0110] See Figure 6 , Figure 6 is a schematic structural diagram of a pulse analysis device provided by an embodiment of the present application. As Figure 6 shown, the pulse analysis device includes:
[0111] A first acquisition unit 601, configured to acquire a face image frame to be analyzed;
[0112] A setting unit 602, configured to set an initial state space;
[0113] A pulse analysis unit 603, configured to input the face image frame to be analyzed and the initial state space into a pre - constructed pulse analysis model for pulse analysis to obtain a first predicted pulse value and an updated state space; the pre - constructed pulse analysis model is obtained through model training based on a spatio - temporal state space model with dual characteristics; the initial state space is used to provide historical background and dynamic context information for the face image frame to be analyzed when the pulse analysis model performs pulse analysis; the updated state space is used as the initial state space for the next face image frame to be analyzed; the next face image frame is the next face image frame to be analyzed acquired through an image acquisition device; the acquisition time of the next face image frame is later than that of the face image frame to be analyzed;
[0114] An output unit 604, configured to output the first predicted pulse value and the updated state space.
[0115] In a possible implementation, the device further includes:
[0116] A second acquisition unit, configured to acquire historical video segments;
[0117] An extraction unit, configured to extract face video frames from the historical video segments to obtain historical face video frames;
[0118] A synthesis unit, configured to, for each of the historical face video frames, acquire a historical true value pulse value synchronized with the historical face video frame, and label the historical true value pulse value on the historical face video frame;
[0119] A model training unit, configured to input a plurality of the historical face video frames labeled with historical true value pulse values into the spatio-temporal state space model for model training to obtain the pulse analysis model.
[0120] In a possible implementation, the layer structure of the spatio-temporal state space model with dual characteristics includes a first projection layer, a first spatio-temporal state space module, a second spatio-temporal state space module, a third spatio-temporal state space module, a fourth spatio-temporal state space module, a weighted summation module, a downsampling module, and a prediction head; the input end of the first spatio-temporal state space module is connected to the output end of the first projection layer; the input end of the second spatio-temporal state space module is connected to the output end of the first spatio-temporal state space module; the input end of the third spatio-temporal state space module is connected to the output end of the second spatio-temporal state space module; the input end of the fourth spatio-temporal state space module is connected to the output end of the third spatio-temporal state space module; the output end of the fourth spatio-temporal state space module is connected to the weighted summation module; the weighted summation module is connected to the downsampling module; the downsampling module is connected to the prediction head; the spatio-temporal state space module has dual characteristics;
[0121] Among them, the first projection layer is used to perform a linear transformation on the multiple historical face video frames to obtain a first linear transformation feature; the first spatio-temporal state space module is used to perform temporal normalization processing and spatio-temporal dual processing on the first linear transformation feature to obtain a first dual state space representation; the second spatio-temporal state space module is used to perform temporal normalization processing and spatio-temporal dual processing on the first dual state space representation to obtain a second dual state space representation; the third spatio-temporal state space module is used to perform temporal normalization processing and spatio-temporal dual processing on the second dual state space representation to obtain a third dual state space representation; the fourth spatio-temporal state space module is used to perform temporal normalization processing and spatio-temporal dual processing on the third dual state space representation to obtain a fourth dual state space representation; the weighted summation module is used to perform a weighted summation on the fourth dual state space representation and the first linear transformation feature to obtain a comprehensive feature representation; the downsampling module is used to perform dimensionality reduction processing on the comprehensive feature representation to obtain a dimensionality reduction feature representation; the prediction head is used to perform a prediction on the dimensionality reduction feature representation to obtain a pulse value prediction result; the pulse value prediction result includes a plurality of second predicted pulse values; the pulse value prediction result is used to be compared with the plurality of historical true pulse values for model optimization.
[0122] In a possible implementation manner, each of the spatio-temporal state space modules includes a temporal normalization module, a first spatio-temporal dual module, and a second spatio-temporal dual module; the output end of the temporal normalization module is connected to the input end of the first spatio-temporal dual module; the output end of the first spatio-temporal dual module is connected to the input end of the second spatio-temporal dual module; the output end of the second spatio-temporal dual module is connected to the weighted summation module;
[0123] Among them, the temporal normalization module is used to extract the temporal correlation features between the respective historical face video frames, and align and standardize the multiple historical face video frames based on the temporal correlation features to obtain a normalized video; the first spatio-temporal dual module is used to perform a first spatial two-dimensional convolution, a temporal causal convolution, a state space equivalent decoupling, and a second spatial two-dimensional convolution on the normalized video in combination with the temporal correlation features to obtain an initial dual state space representation; the second spatio-temporal dual module is used to perform a first spatial two-dimensional convolution, a temporal causal convolution, a state space equivalent decoupling, and a second spatial two-dimensional convolution on the initial dual state space representation in combination with the temporal correlation features to obtain the fourth dual state space representation.
[0124] In a possible implementation, each of the spatio-temporal dual modules includes a first spatial two-dimensional convolutional layer, a second projection layer, a temporal causal convolutional layer, a first activation function layer, a second activation function layer, a dual state space layer, a matrix multiplication layer, a normalization layer, a third projection layer, and a second spatial two-dimensional convolutional layer; the input ends of the first spatial two-dimensional convolutional layer and the third projection layer are both connected to the output end of the temporal normalization module; the output end of the first spatial two-dimensional convolutional layer is connected to the input end of the second projection layer; the output end of the second projection layer is respectively connected to the input ends of the temporal causal convolutional layer, the first activation function layer, and the dual state space layer; the output end of the temporal causal convolutional layer is connected to the input end of the second activation function layer; the output end of the first activation function layer is connected to the input end of the matrix multiplication layer; the output end of the dual state space layer is connected to the input end of the matrix multiplication layer; the output end of the second activation function layer is connected to the input end of the dual state space layer; the output end of the matrix multiplication layer is connected to the input end of the normalization layer; the output end of the normalization layer is connected to the input end of the third projection layer; the output end of the third projection layer is connected to the input end of the second spatial two-dimensional convolutional layer;
[0125] Among them, the first spatial two-dimensional convolutional layer is used to perform spatial two-dimensional convolution on the normalized video to obtain a spatio-temporal feature map; the second projection layer is used to perform a linear transformation on the spatio-temporal feature map to obtain a second linearly transformed feature; the temporal causal convolutional layer is used to perform temporal causal convolution on the second linearly transformed feature to obtain a temporal feature; the first activation function layer is used to activate the temporal feature to obtain an activated temporal feature; the second activation function layer is used to activate the transformed feature to obtain an activated transformed feature; the dual state space layer is used to perform a spatio-temporal dual operation on the activated transformed feature and the transformed feature to obtain a dual state space representation; the matrix multiplication layer is used to perform a matrix multiplication operation on the dual state space representation and the activated temporal feature to obtain a fused feature; the normalization layer is used to normalize the fused feature to obtain a normalized feature; the third projection layer is used to perform a non-linear transformation on the normalized feature based on the normalized video to obtain a non-linearly transformed feature; the second spatial two-dimensional convolutional layer is used to perform spatial two-dimensional convolution on the non-linearly transformed feature to obtain the dual state space representation.
[0126] In a possible implementation, the extraction steps of the temporal normalization module sequentially include pixel change residual calculation, standard deviation calculation of the pixel change residual, and normalization operation on each frame of the image.
[0127] In addition, an embodiment of the present application further provides a pulse analysis device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the pulse analysis method described above is implemented.
[0128] In addition, an embodiment of the present application further provides a computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a terminal device, the terminal device is enabled to execute the pulse analysis method described above.
[0129] An embodiment of the present application provides a pulse analysis device. First, a first acquisition unit 301 is used to acquire a face image frame to be analyzed, and an initial state space is set by a setting unit 602. The initial state space is used to provide historical background and dynamic context information for the face image frame to be analyzed when performing pulse analysis in a pulse analysis model. A pulse analysis unit 603 inputs the face image frame to be analyzed and the initial state space into a pre-constructed pulse analysis model for pulse analysis, obtaining a first predicted pulse value and an updated state space. Among them, the pre-constructed pulse analysis model is obtained by training based on a spatio-temporal state space model with dual characteristics. Then, an output unit 604 is used to output the first predicted pulse value and the updated state space. The pulse analysis model in the present application is trained through a spatio-temporal state space model with dual characteristics, jointly modeling the changes in time and space. Therefore, it can extract tiny blood flow dynamic features from a single-frame image and perform accurate pulse prediction. Even without a large number of consecutive video frames, accurate pulse waveform data can be calculated. This makes the model have stronger real-time performance, can instantaneously feedback the pulse value, and greatly reduces the latency. Compared with traditional finger clip oximeters, the present application realizes non-contact physiological monitoring, avoiding the discomfort brought by traditional devices, and is especially suitable for the elderly and people with limited hand movement. In addition, only a common image acquisition device (such as a mobile phone camera) is needed to acquire a face image and perform subsequent analysis, which not only reduces the device cost but also enhances the convenience and popularity of the technology.
[0130] The above has introduced in detail a pulse analysis method, device, equipment, and storage medium provided by the present application. The embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
[0131] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of a single item (one) or multiple items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0132] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent in such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.
Claims
1. A method for analyzing a pulse, characterized in that, The method includes: Obtaining a face image frame to be analyzed by an image acquisition device and setting an initial state space; Inputting the face image frame to be analyzed and the initial state space into a pre-constructed pulse analysis model for pulse analysis to obtain a first predicted pulse value and an updated state space; the pre-constructed pulse analysis model is obtained by training a spatio-temporal state space model with dual characteristics; the initial state space is used to provide historical background and dynamic context information for the face image frame to be analyzed when the pulse analysis model performs pulse analysis; the updated state space is used as the initial state space for the next face image frame to be analyzed; the next face image frame is the next face image frame to be analyzed obtained by the image acquisition device; the acquisition time of the next face image frame is later than that of the face image frame to be analyzed; Outputting the first predicted pulse value and the updated state space.
2. The method according to claim 1, wherein The construction process of the pre-constructed pulse analysis model includes: Obtaining a historical video segment and extracting face video frames in the historical video segment to obtain historical face video frames; For each of the historical face video frames, obtaining a historical true-value pulse value synchronized with the historical face video frame; Inputting multiple historical face video frames into the spatio-temporal state space model for model training to obtain the pulse analysis model; Using the multiple historical true-value pulse values to optimize the pulse analysis model.
3. The method according to claim 2, wherein The layer structure of the spatio-temporal state space model with dual characteristics includes a first projection layer, a first spatio-temporal state space module, a second spatio-temporal state space module, a third spatio-temporal state space module, a fourth spatio-temporal state space module, a weighted summation module, a downsampling module, and a prediction head; the input end of the first spatio-temporal state space module is connected to the output end of the first projection layer; the input end of the second spatio-temporal state space module is connected to the output end of the first spatio-temporal state space module; the input end of the third spatio-temporal state space module is connected to the output end of the second spatio-temporal state space module; the input end of the fourth spatio-temporal state space module is connected to the output end of the third spatio-temporal state space module; the output end of the fourth spatio-temporal state space module is connected to the weighted summation module; the weighted summation module is connected to the downsampling module; the downsampling module is connected to the prediction head; the spatio-temporal state space module has dual characteristics; Among them, the first projection layer is used to perform a linear transformation on the multiple historical face video frames to obtain first linear transformation features; the first spatio-temporal state space module is used to perform temporal normalization processing and spatio-temporal duality processing on the first linear transformation features to obtain a first dual state space representation; the second spatio-temporal state space module is used to perform temporal normalization processing and spatio-temporal duality processing on the first dual state space representation to obtain a second dual state space representation; the third spatio-temporal state space module is used to perform temporal normalization processing and spatio-temporal duality processing on the second dual state space representation to obtain a third dual state space representation; the fourth spatio-temporal state space module is used to perform temporal normalization processing and spatio-temporal duality processing on the third dual state space representation to obtain a fourth dual state space representation; the weighted summation module is used to perform a weighted summation on the fourth dual state space representation and the first linear transformation features to obtain a comprehensive feature representation; the downsampling module is used to perform dimensionality reduction processing on the comprehensive feature representation to obtain a dimensionality reduction feature representation; the prediction head is used to perform a prediction on the dimensionality reduction feature representation to obtain a pulse value prediction result; the pulse value prediction result includes a plurality of second predicted pulse values; the pulse value prediction result is used to be compared with the plurality of historical true value pulse values for model optimization.
4. The method according to claim 3, wherein Each of the spatio-temporal state space modules includes a temporal normalization module, a first spatio-temporal duality module, and a second spatio-temporal duality module; the output end of the temporal normalization module is connected to the input end of the first spatio-temporal duality module; the output end of the first spatio-temporal duality module is connected to the input end of the second spatio-temporal duality module; the output end of the second spatio-temporal duality module is connected to the weighted summation module; Among them, the temporal normalization module is used to extract the temporal correlation features between the respective historical face video frames, and align and standardize the multiple historical face video frames based on the temporal correlation features to obtain a normalized video; the first spatio-temporal duality module is used to perform a first spatial two-dimensional convolution, a temporal causal convolution, a state space equivalent decoupling, and a second spatial two-dimensional convolution on the normalized video in combination with the temporal correlation features to obtain an initial dual state space representation; the second spatio-temporal duality module is used to perform a first spatial two-dimensional convolution, a temporal causal convolution, a state space equivalent decoupling, and a second spatial two-dimensional convolution on the initial dual state space representation in combination with the temporal correlation features to obtain the fourth dual state space representation.
5. The method according to claim 4, characterized in that Each of the spatio-temporal dual modules includes a first spatial two-dimensional convolutional layer, a second projection layer, a temporal causal convolutional layer, a first activation function layer, a second activation function layer, a dual state space layer, a matrix multiplication layer, a normalization layer, a third projection layer, and a second spatial two-dimensional convolutional layer; the input ends of the first spatial two-dimensional convolutional layer and the third projection layer are both connected to the output end of the temporal normalization module; the output end of the first spatial two-dimensional convolutional layer is connected to the input end of the second projection layer; the output end of the second projection layer is respectively connected to the input ends of the temporal causal convolutional layer, the first activation function layer, and the dual state space layer; the output end of the temporal causal convolutional layer is connected to the input end of the second activation function layer; the output end of the first activation function layer is connected to the input end of the matrix multiplication layer; the output end of the dual state space layer is connected to the input end of the matrix multiplication layer; the output end of the second activation function layer is connected to the input end of the dual state space layer; the output end of the matrix multiplication layer is connected to the input end of the normalization layer; the output end of the normalization layer is connected to the input end of the third projection layer; the output end of the third projection layer is connected to the input end of the second spatial two-dimensional convolutional layer; Among them, the first spatial two-dimensional convolutional layer is used to perform spatial two-dimensional convolution on the normalized video to obtain a spatio-temporal feature map; the second projection layer is used to perform a linear transformation on the spatio-temporal feature map to obtain a second linearly transformed feature; the temporal causal convolutional layer is used to perform temporal causal convolution on the second linearly transformed feature to obtain a temporal feature; the first activation function layer is used to activate the temporal feature to obtain an activated temporal feature; the second activation function layer is used to activate the transformed feature to obtain an activated transformed feature; the dual state space layer is used to perform a spatio-temporal dual operation on the activated transformed feature and the transformed feature to obtain a dual state space representation; the matrix multiplication layer is used to perform a matrix multiplication operation on the dual state space representation and the activated temporal feature to obtain a fused feature; the normalization layer is used to normalize the fused feature to obtain a normalized feature; the third projection layer is used to perform a non-linear transformation on the normalized feature based on the normalized video to obtain a non-linearly transformed feature; the second spatial two-dimensional convolutional layer is used to perform spatial two-dimensional convolution on the non-linearly transformed feature to obtain the dual state space representation.
6. The method according to claim 4, wherein The extraction steps of the temporal normalization module sequentially include pixel change residual calculation, standard deviation calculation of the pixel change residual, and normalization operation on each frame of the image.
7. A pulse analysis device, characterized in that, The device includes: A first acquisition unit, configured to acquire a face image frame to be analyzed; A setting unit, configured to set an initial state space; A pulse analysis unit for inputting the face image frame to be analyzed and the initial state space into a pre-constructed pulse analysis model for pulse analysis, to obtain a first predicted pulse value and an updated state space; the pre-constructed pulse analysis model is obtained by training a spatio-temporal state space model with dual characteristics; the initial state space is used to provide historical background and dynamic context information for the face image frame to be analyzed when the pulse analysis model performs pulse analysis; the updated state space is used as the initial state space for the next face image frame to be analyzed; the next face image frame is the next face image frame to be analyzed obtained by an image acquisition device; the acquisition time of the next face image frame is later than that of the face image frame to be analyzed; An output unit for outputting the first predicted pulse value and the updated state space.
8. The device according to claim 7, characterized in that, The device further includes: A second acquisition unit for acquiring a historical video segment; An extraction unit for extracting face video frames from the historical video segment to obtain historical face video frames; A synthesis unit for, for each of the historical face video frames, acquiring a historical true-value pulse value synchronized with the historical face video frame, and annotating the historical true-value pulse value on the historical face video frame; A model training unit for inputting a plurality of the historical face video frames annotated with historical true-value pulse values into the spatio-temporal state space model for model training to obtain the pulse analysis model.
9. A pulse analysis device, characterized in that, Including: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the pulse analysis method according to any one of claims 1-6 is implemented.
10. A computer-readable storage medium, characterized in that, Instructions are stored in the computer-readable storage medium, and when the instructions are run on a terminal device, the terminal device is caused to execute the pulse analysis method according to any one of claims 1-6.