A method, apparatus, device, medium, and program product for extracting rPPG signals.
By performing channel enhancement and spatiotemporal fusion feature extraction on rPPG signals, combined with an attention mechanism, the problem of rPPG signals being susceptible to noise interference is solved. This achieves improved signal accuracy and stability while reducing computational costs, making it suitable for non-contact heart rate measurement.
Patent Information
- Application Number
- CN202411852998.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-12-16
AI Technical Summary
Existing rPPG signal extraction methods are easily affected by video noise and have weak signals in non-contact heart rate measurement, resulting in insufficient measurement accuracy and stability, as well as high computational costs, making them difficult to promote in practical applications.
By performing channel enhancement on the original video, Boolean images and chromatograms are generated, spatiotemporal fusion features are extracted, and time displacement and difference frame stitching are used. Combined with the time-channel attention and space-time attention mechanisms in 3D convolutional networks, feature sequences at different frame rates are segmented, and rPPG sub-signals at different frame rates are fused, reducing computational costs while improving signal accuracy.
While reducing computational costs, it significantly improves the extraction accuracy and stability of rPPG signals, making it suitable for non-contact heart rate measurement in various scenarios.
Smart Images

Figure CN119672612B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of rPPG signal extraction, and particularly to a method, apparatus, device, medium, and program product for extracting rPPG signals. Background Technology
[0002] Heart rate is an important indicator of human health. Currently, commonly used traditional methods for measuring heart rate include electrocardiogram (ECG) and photoplethysmography (PPG). Both methods calculate heart rate by acquiring ECG or PPG signals. However, since both technologies require direct contact with the skin for measurement, they have limitations in certain scenarios where contact is not suitable (such as skin burns, newborn monitoring, sleep monitoring, etc.). Remote photoplethysmography (rPPG) offers a non-contact alternative. The heartbeat causes blood circulation, resulting in changes in blood volume in the microvascular bed beneath the skin, causing periodic color fluctuations in areas such as the face. rPPG remotely analyzes these color fluctuations to extract the periodic physiological signals, i.e., the rPPG signal. This means that even using a regular webcam, rPPG signals can be easily acquired to calculate heart rate information. Compared to traditional contact methods, rPPG technology is more economical and convenient, and suitable for various scenarios. However, rPPG signals are susceptible to interference from various strong noises in video, and the signal itself is weak, posing a significant challenge to the accuracy and stability of heart rate measurement. Summary of the Invention
[0003] The purpose of this invention is to provide at least one method, apparatus, device, medium, and program product for extracting rPPG signals. This application provides an rPPG signal extraction method comprising: acquiring an original video from which rPPG signals to be extracted; performing channel enhancement on the original video to obtain a channel-enhanced video; extracting spatiotemporal fusion features of the rPPG signal from the channel-enhanced video; segmenting the spatiotemporal fusion features into feature sequences at different frame rates; extracting rPPG sub-signals of the original video at different frame rates from each of the feature sequences; and determining the rPPG signal of the original video based on the rPPG sub-signals at each frame rate. This method improves the accuracy of the extracted rPPG signal while reducing computational costs.
[0004] To address the aforementioned technical problems, this application proposes five aspects.
[0005] In a first aspect, this application provides a method for extracting rPPG signals, comprising: acquiring an original video from which rPPG signals to be extracted; performing channel enhancement on the original video to obtain a channel-enhanced video; extracting spatiotemporal fusion features of the rPPG signals from the channel-enhanced video; segmenting the spatiotemporal fusion features into feature sequences at different frame rates; extracting rPPG sub-signals of the original video at different frame rates from each of the feature sequences; and determining the rPPG signals of the original video based on the rPPG sub-signals at each frame rate.
[0006] In some embodiments, performing channel enhancement on the original video to obtain a channel-enhanced video includes: generating a Boolean image and a chromatogram based on the original video; and generating the channel-enhanced video based on the Boolean image, the chromatogram, and the original video.
[0007] In some embodiments, extracting the spatiotemporal fusion features of the rPPG signal from the channel-enhanced video includes: obtaining multiple time-shifted videos by performing a time-shifting operation on the channel-enhanced video; determining difference frames between two adjacent time-shifted videos based on the multiple time-shifted videos; stitching the multiple difference frames together; and extracting the spatiotemporal fusion features of the rPPG signal from the stitching result.
[0008] In some embodiments, extracting rPPG sub-signals of the original video at different frame rates from each of the feature sequences includes: performing the following operations for multiple feature sequences at any frame rate: determining the spatiotemporal attention features of each of the feature sequences based on each of the feature sequences; and determining the rPPG sub-signals of the original video at the current frame rate based on each of the spatiotemporal attention features.
[0009] In some embodiments, determining the spatiotemporal attention features of each of the feature sequences based on each of the feature sequences includes: performing the following operations for any feature sequence: determining the attention weight of the feature sequence based on the feature sequence; determining the temporal channel attention features of the feature sequence based on the attention weight and the feature sequence; determining the spatiotemporal attention weight of the feature sequence based on the temporal channel attention features; and determining the spatiotemporal attention features of the feature sequence based on the spatiotemporal attention weight and the temporal channel attention features.
[0010] In some embodiments, determining the rPPG signal of the original video based on the rPPG sub-signals at each frame rate includes: determining the self-attention weight of the original video based on the spatiotemporal fusion feature; and fusing the rPPG sub-signals at different frame rates based on the self-attention weight to determine the rPPG signal of the original video.
[0011] Secondly, this application provides an rPPG signal extraction apparatus, comprising: a first acquisition module for acquiring an original video from which the rPPG signal to be extracted is to be acquired; a first enhancement module for performing channel enhancement on the original video to obtain a channel-enhanced video; a first fusion module for extracting spatiotemporal fusion features of the rPPG signal from the channel-enhanced video; a first segmentation module for segmenting the spatiotemporal fusion features into feature sequences at different frame rates; a first extraction module for extracting rPPG sub-signals of the original video at different frame rates from each of the feature sequences; and a first determination module for determining the rPPG signal of the original video based on the rPPG sub-signals at each frame rate.
[0012] Thirdly, this application proposes a computer electronic production device, characterized in that it includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of any of the methods in the first aspect.
[0013] Fourthly, this application proposes a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of the method described in any one of the first aspects.
[0014] Fifthly, this application proposes a computer program product comprising a computer program that, when executed by a processor, implements the steps of the method described in any of the first aspects.
[0015] This application discloses a method for extracting rPPG signals, comprising: acquiring an original video from which rPPG signals are to be extracted; performing channel enhancement on the original video to obtain a channel-enhanced video; extracting spatiotemporal fusion features of the rPPG signals from the channel-enhanced video; segmenting the spatiotemporal fusion features into feature sequences at different frame rates; extracting rPPG sub-signals of the original video at different frame rates from each feature sequence; and determining the rPPG signals of the original video based on the rPPG sub-signals at each frame rate. This method improves the accuracy of the extracted rPPG signals while reducing computational costs. Attached Figure Description
[0016] One or more embodiments are illustrated by way of example with reference to the accompanying drawings, and these illustrative descriptions do not constitute a limitation on the embodiments.
[0017] Figure 1 The main flowchart of an rPPG signal extraction method provided in this application embodiment;
[0018] Figure 2 A schematic diagram of a channel enhancement process provided in an embodiment of this application;
[0019] Figure 3 A schematic diagram illustrating the process of extracting spatiotemporal fusion features provided in this application embodiment;
[0020] Figure 4 This is a schematic diagram illustrating a process for extracting channel attention features, provided in an embodiment of this application.
[0021] Figure 5 A schematic diagram illustrating the process of extracting spatiotemporal attention features provided in an embodiment of this application;
[0022] Figure 6 This is a schematic diagram of the structure of a neural network model provided in an embodiment of this application;
[0023] Figure 7 This is a main structural block diagram of an rPPG signal extraction device provided in an embodiment of this application. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details have been provided in the various embodiments of this application to help readers better understand this application. However, the technical solutions claimed in this application can be implemented even without these technical details and various changes and modifications based on the following embodiments. The division of the various embodiments below is for the convenience of description and should not constitute any limitation on the specific implementation of this application. The various embodiments can be combined with and referenced by each other without contradiction.
[0025] Heart rate is an important indicator of human health. Currently, commonly used traditional methods for measuring heart rate include electrocardiogram (ECG) and photoplethysmography (PPG). Both methods calculate heart rate by acquiring ECG or PPG signals. However, since both technologies require direct contact with the skin for measurement, they have limitations in certain scenarios where contact is not suitable (such as skin burns, newborn monitoring, sleep monitoring, etc.). Remote photoplethysmography (rPPG) offers a non-contact alternative. The heartbeat causes blood circulation, resulting in changes in blood volume in the microvascular bed beneath the skin, causing periodic color fluctuations in areas such as the face. rPPG remotely analyzes these color fluctuations to extract the periodic physiological signals, i.e., the rPPG signal. This means that even using a regular webcam, rPPG signals can be easily acquired to calculate heart rate information. Compared to traditional contact methods, rPPG technology is more economical and convenient, and suitable for various scenarios. However, rPPG signals are susceptible to interference from various strong noises in video, and the signal itself is weak, posing a significant challenge to the accuracy and stability of heart rate measurement.
[0026] In recent years, significant progress has been made in rPPG research. Especially driven by deep learning, many data-driven rPPG signal extraction methods have emerged. However, suppressing noise such as motion artifacts, video compression artifacts, and illumination variations remains a core challenge for rPPG methods. To improve the robustness of rPPG signal extraction, early researchers employed manual modeling to separate noise from useful rPPG signals, extracting cleaner rPPG signals. While this method is effective in dealing with noise, its limitations are also significant. In data-driven methods, some studies have attempted to combine prior knowledge from manual modeling to spatially transform RGB videos to improve signal extraction performance. However, because this direct transformation may lead to the loss of original video features, its effectiveness is still constrained by the limitations of traditional methods. Furthermore, these methods fail to fully leverage the advantages of deep neural networks in automatic feature extraction.
[0027] While deep learning-based methods have achieved significant performance improvements over manual modeling, their high computational cost limits their practical application feasibility. To improve model efficiency, a two-stream network structure is proposed, accelerating model inference through differential frames after downsampling. However, since this method uses a 2D convolutional network, it cannot simultaneously process the temporal and spatial features in face videos, resulting in less than ideal performance. Furthermore, an attempt is made to simplify the computationally expensive PhysNet network and downsample the input face video in the spatial dimension to reduce network depth and input data volume, thereby improving signal extraction speed. However, this operation inevitably affects the model's stability and final performance.
[0028] To address the aforementioned technical problems, this invention proposes a method for extracting rPPG signals. The implementation details of the bandwidth determination method in this embodiment are described below. The following content is only for ease of understanding and is not essential for implementing this solution.
[0029] Example 1:
[0030] like Figure 1 As shown, this application provides a method for extracting rPPG signals. The method is implemented in electronic production equipment, which can be a server, mobile terminal, computer, cloud platform, etc. The data processing functionality provided in this application embodiment can be implemented by the processor of the electronic production equipment calling program code, wherein the program code can be stored in a computer storage medium. The rPPG signal extraction method includes:
[0031] Step S1: Obtain the original video from which the rPPG signal to be extracted is obtained.
[0032] Step S2: Perform channel enhancement on the original video to obtain a channel-enhanced video.
[0033] In some embodiments, step S2, "performing channel enhancement on the original video to obtain a channel-enhanced video," includes:
[0034] Step S21: Generate a Boolean image and a chromatogram based on the original video.
[0035] Step S22: Generate the channel-enhanced video based on the Boolean image, the chromatogram, and the original video.
[0036] Prior knowledge-based POS methods can effectively reduce motion-induced noise. This method generates a dual-channel image for extracting the rPPG signal by projecting the RGB image (original video image) onto a plane orthogonal to the vector [1,1,1]. However, because this method directly performs spatial transformation on the original RGB image, some information in the original video frame is lost. For example... Figure 2 As shown, in this application, the original RGB video is converted into a two-channel POS image (Boolean image) and a CHROM image (chromatogram). The conversion method is shown in the following formula.
[0037]
[0038] Here, X is the input for channel expansion. Utilizing prior knowledge from manual modeling, the input X is multiplied by this matrix to obtain the channel-enhanced X', i.e., the channel-enhanced video. Through steps S21-S22, the enhanced video features are added to the channel dimension, effectively separating useful features from noise, thereby significantly improving the representational capability of the rPPG signal.
[0039] Step S3: Extract the spatiotemporal fusion features of the rPPG signal from the channel-enhanced video.
[0040] Since facial color changes are extremely subtle, effectively representing the rPPG signal in the enhanced video sequence becomes a key challenge. Previous methods effectively removed inherent noise such as background noise and skin color variations by calculating difference frames over time. This is because the rPPG signal primarily originates from minute color changes between adjacent frames in a face video, and difference frames, by subtracting pixel values from adjacent frames, help the neural network more accurately capture these changes, thus improving the extraction of the rPPG signal.
[0041] In some embodiments, step S3, "extracting the spatiotemporal fusion features of the rPPG signal from the channel-enhanced video," includes:
[0042] Step S31: Obtain multiple time-shifted videos by performing a time-shifting operation on the enhanced video of the channel.
[0043] Step S32: Determine the difference frames between two adjacent time-shifted videos based on the multiple time-shifted videos.
[0044] Step S33: Perform difference frame stitching on multiple difference frames, and extract the spatiotemporal fusion features of the rPPG signal from the stitching result.
[0045] like Figure 3 As shown, assuming the dimensions of the channel-enhanced video are N×7×T×128×128, where N is the batch size, 7 is the number of channels, T is the number of video frames, and 128 is the spatial resolution of the video frames. First, the channel-enhanced video is time-shifted to generate x(t-1), x(t), and x(t+1), and the corresponding difference frames D1 and D2 are calculated in chronological order. Next, these difference frames are concatenated along the channel dimension to form a feature representation containing multi-scale information. Then, these features are fed into two sub-networks, Stem1 and Stem2, to extract the main rPPG signal waveform features. Finally, through a feature fusion step, the features from different channels are integrated to further enhance the representation capability of the rPPG signal. The process can be represented by the following formula:
[0046] x fusion =Stem2(Stem1(Concat(D1,D2)))
[0047] Where, x fusion This is a spatiotemporal fusion feature. Step S3 enhances the model's ability to perceive rPPG signals and effectively reduces noise interference, ensuring that the extracted signals are more stable and accurate.
[0048] Step S4: Segment the spatiotemporal fusion features into feature sequences with different frame rates.
[0049] Since extracting rPPG signals using low frame rate video sometimes yields better results, this application constructs a temporal feature pyramid to capture multi-scale rPPG signal features from videos at different frame rates. Specifically, this application processes spatiotemporal fusion features into slices at various frame rates, such as 30FPS, 15FPS, and lower. Because different frame rate slices are used, the spatiotemporal fusion features can form different numbers of feature sequences at different frame rate levels. Then, each frame rate level's feature sequence is processed by a dedicated rPPG signal extraction module to capture multi-scale features in the time series, ensuring full utilization of information from different time scales.
[0050] Step S5: Extract rPPG sub-signals of the original video at different frame rates from each of the feature sequences.
[0051] In some embodiments, step S5, "extracting rPPG sub-signals of the original video at different frame rates from each of the said feature sequences," includes:
[0052] Step S51: Determine the spatiotemporal attention features of each of the feature sequences based on each of the feature sequences.
[0053] In some embodiments, step S51, "determining the spatiotemporal attention features of each of the feature sequences based on each of the feature sequences," includes:
[0054] Step S511: Determine the attention weight of the feature sequence based on the feature sequence.
[0055] Step S512: Determine the temporal channel attention features of the feature sequence based on the attention weights and the feature sequence.
[0056] Step S513: Determine the spatiotemporal attention weights of the feature sequence based on the temporal channel attention features.
[0057] Step S514: Determine the spatiotemporal attention features of the feature sequence based on the spatiotemporal attention weights and the temporal channel attention features.
[0058] Step S52: Determine the rPPG sub-signal of the original video at the current frame rate based on each of the spatiotemporal attention features.
[0059] To fully extract rPPG signals from videos, this application designs an ingenious rPPG signal extraction module. This module mainly consists of two parts, each introducing different attention mechanisms to capture signal features more comprehensively. The first part comprises a 3D convolutional layer, temporal-channel attention (CTA), and a normalization layer. CTA focuses on feature information in the channel and temporal dimensions, effectively capturing subtle signals that change over time in different color channels. This helps the model accurately separate features reflecting physiological information while suppressing errors introduced by noise or lighting variations. The second part consists of a 3D convolutional layer, temporal-spatial attention (STA), and a normalization layer. STA focuses on the spatial and temporal dimensions, capturing subtle signal fluctuations in the facial region over time and effectively perceiving periodic changes in facial pixels, which is crucial for enhancing the stability of rPPG signals. The main difference between these two parts lies in the different attention mechanisms introduced. By combining CTA and STA in the extraction module, we can fully utilize the channel, spatial, and temporal features of the video. This design improves the model's ability to perceive rPPG signals, enabling more stable and efficient signal extraction in complex scenarios.
[0060] Temporal channel attention aims to optimize the weight differences of rPPG signals across different channels in face videos. This module is designed to simultaneously extract attention weights in both temporal and channel dimensions to achieve feature correlation between the two dimensions. Specifically... Figure 4 As shown. Since its main goal is to capture information in the temporal and channel dimensions, the input features are globally averaged in the spatial dimension. Then, the processed features are convolved using three different dilation rates. For an input of shape C×T×H×W, a C×T×1×1 output result Di is obtained. Its calculation formula is as follows:
[0061] D i =DilConv i (Ave 3,4 (CT in ))
[0062] Dilated convolutions, by expanding the receptive field, help capture a wider range of contextual information. This application sets the dilation rates of the dilated convolutions to 1, 2, and 4. The resulting features D1, D2, and D4 are concatenated along the first dimension, and then attention weights are calculated using a group convolution with a group size of 3, a 1×1×1 convolution kernel, and a sigmoid function. Finally, the attention weights are multiplied by the original input to obtain the temporal channel attention feature CTout, where CT... in Given the input feature sequence, its calculation formula is as follows:
[0063] W ct=σ(Conv(concat(D1,D2,D4)))
[0064]
[0065] Facial features in videos change over time; therefore, the spatiotemporal attention module aims to jointly extract attention weights from both spatial and temporal dimensions to more accurately capture these dynamic changes. For example... Figure 5 As shown, this module comprises two branches: the first branch performs average pooling along the channel dimension to extract spatiotemporal features; the second branch focuses on learning temporal characteristics by performing average pooling along both the channel and spatial dimensions. The outputs of both branches are processed by convolution and the sigmoid activation function, and then multiplied element-wise to generate the final temporal and spatial joint attention weight Wst. The specific calculation formula is as follows:
[0066]
[0067] The first branch has a 3×3×3 kernel size, used to capture spatiotemporal features; the second branch has a 1×1×1 kernel size, specifically for extracting temporal features. Finally, the generated spatiotemporal attention weights Wst are multiplied element-wise with the input feature STin (channel attention feature) to obtain the final spatiotemporal attention feature STout. The specific calculation formula is as follows:
[0068]
[0069] After each feature sequence passes through the temporal channel attention module and the temporal-spatial attention module, the spatiotemporal attention features of that feature sequence can be obtained. By fusing the spatiotemporal attention features of each feature sequence at each frame rate, the rPPG signal at that frame rate can be obtained, which is the rPPG sub-signal of the original video at that frame rate.
[0070] Step S6: Determine the rPPG signal of the original video based on the rPPG sub-signals at each frame rate.
[0071] In some embodiments, step S6, "determining the rPPG signal of the original video based on the rPPG sub-signals at each frame rate," includes:
[0072] Step S61: Determine the self-attention weight of the original video based on the spatiotemporal fusion features.
[0073] Step S62: Based on the self-attention weights, fuse the rPPG sub-signals of different frame rates to determine the rPPG signal of the original video.
[0074] Since rPPG sub-signals at different frame rates have different characteristics, in order to obtain a higher precision rPGG signal, it is necessary to fuse rPPG sub-signals at different frame rates. During the fusion process, the proportion of rPPG sub-signals at different frame rates needs to be considered. Therefore, it is necessary to first determine the self-attention weight of the original video. In this application, the self-attention weight of the original video is determined by global pooling of the spatiotemporal fusion features. Then, the rPPG sub-signals corresponding to each frame rate are fused according to the self-attention weight to obtain a high precision rPPG signal.
[0075] Because rPPG signals are inherently weak and easily interfered with by various strong noises in video, it is difficult to obtain high-precision rPPG signals using existing technologies. Moreover, obtaining high-precision rPPG signals requires complex calculations and a huge amount of computation. However, this application achieves high-precision rPPG signals with a relatively low amount of computation through the method described above.
[0076] In addition, the method of this application can also be implemented using a neural network model, as in some embodiments, such as Figure 6 As shown, the neural network model includes a channel enhancement module, a feature fusion module, and a temporal feature pyramid rPPG signal extraction module. To address the strong noise and weak signal characteristics in the video, the channel enhancement module uses POS and CHROM to enhance video features. These enhanced video features are added to the channel dimension, effectively separating useful features from noise, thus significantly improving the representational ability of the rPPG signal. To further improve the signal extraction effect, we designed a feature fusion module, which enhances the network's perception of the rPPG signal by analyzing differential frames. Finally, we downsample the features in the temporal dimension to obtain multi-scale feature representations to construct a temporal feature pyramid, extracting rPPG signals from multiple scales, and fusing them to obtain the final rPPG signal. In the rPPG signal extraction module, our proposed Channel-Temporal Attention (CTA) and Spatial-Temporal Attention (STA) mechanisms can focus on the features of the rPPG signal in multiple dimensions.
[0077] Example 2:
[0078] Based on the foregoing embodiments, this application provides an rPPG signal extraction device. The various modules and units included in the device can be implemented by a processor in a computer device; of course, they can also be implemented by specific logic circuits. In the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0079] like Figure 7 As shown, an rPPG signal extraction device includes: a first acquisition module 1, a first enhancement module 2, a first fusion module 3, a first segmentation module 4, a first extraction module 5, and a first determination module 6.
[0080] The first acquisition module 1 is used to acquire the original video from which the rPPG signal to be extracted is to be acquired. The first enhancement module 2 is used to perform channel enhancement on the original video to obtain a channel-enhanced video. The first fusion module 3 is used to extract the spatiotemporal fusion features of the rPPG signal from the channel-enhanced video. The first segmentation module 4 is used to segment the spatiotemporal fusion features into feature sequences at different frame rates. The first extraction module 5 is used to extract the rPPG sub-signals of the original video at different frame rates from each of the feature sequences. The first determination module 6 is used to determine the rPPG signal of the original video based on the rPPG sub-signals at each frame rate.
[0081] The modules in the aforementioned rPPG signal extraction device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the device in hardware form or independently of it, or stored in the memory of the processing device in software form, so that the processor can call and execute the operations corresponding to each module. It should be noted that the module division in this embodiment is illustrative and only represents a logical functional division; in actual implementation, there may be other division methods.
[0082] Example 3:
[0083] Thirdly, this application provides a computer electronic production apparatus, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of any of the methods described in the first aspect.
[0084] The memory and processor are connected via a bus, which can include any number of interconnecting buses and bridges, connecting various circuits of one or more processors and memories. The bus can also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.
[0085] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory is used to store data used by the processor during operation.
[0086] Example 4:
[0087] Fourthly, this application proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in any of the first aspects.
[0088] Example 5:
[0089] Fifthly, this application proposes a computer program product, including a computer program / instructions, characterized in that, when the computer program is executed by a processor, it implements the steps of the method described in any one of the first aspects.
[0090] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0091] Those skilled in the art will understand that the above embodiments are specific embodiments for implementing this application, and in practical applications, various changes can be made to them in form and detail without departing from the spirit and scope of this application.
Claims
1. A method for extracting rPPG signals, characterized in that, include: Obtain the original video from which the rPPG signal to be extracted; The original video is enhanced by channel boosting to obtain a channel-enhanced video; The process of performing channel enhancement on the original video to obtain a channel-enhanced video includes: Generate a Boolean image and a chromatogram based on the original video; The channel-enhanced video is generated based on the Boolean image, the chromatogram, and the original video. Spatiotemporal fusion features of rPPG signals are extracted from the enhanced video of the channel; The spatiotemporal fusion features are segmented into feature sequences with different frame rates; Extract rPPG sub-signals of the original video at different frame rates from each of the aforementioned feature sequences; The step of extracting rPPG sub-signals of the original video at different frame rates from each of the feature sequences includes: For multiple feature sequences at any frame rate, perform the following operations: The spatiotemporal attention features of each of the aforementioned feature sequences are determined based on each of the aforementioned feature sequences; The rPPG sub-signal of the original video at the current frame rate is determined based on each of the aforementioned spatiotemporal attention features; The rPPG signal of the original video is determined based on the rPPG sub-signals at each frame rate.
2. The method according to claim 1, characterized in that, The spatiotemporal fusion features extracted from the channel-enhanced video for the rPPG signal include: Multiple time-shifted videos are obtained by performing a time-shifting operation on the enhanced video of the channel. The difference frames between two adjacent time-shifted videos are determined based on multiple time-shifted videos; Multiple difference frames are stitched together, and the spatiotemporal fusion features of the rPPG signal are extracted from the stitching result.
3. The method according to claim 1, characterized in that, Determining the spatiotemporal attention features of each of the feature sequences based on each of the feature sequences includes: For any feature sequence, perform the following operation: The attention weights of the feature sequences are determined based on the feature sequences. The temporal channel attention features of the feature sequence are determined based on the attention weights and the feature sequence. The spatiotemporal attention weights of the feature sequence are determined based on the temporal channel attention features; The spatiotemporal attention features of the feature sequence are determined based on the spatiotemporal attention weights and the temporal channel attention features.
4. The method according to claim 1, characterized in that, Determining the rPPG signal of the original video based on the rPPG sub-signals at each frame rate includes: The self-attention weights of the original video are determined based on the spatiotemporal fusion features. The rPPG sub-signals at different frame rates are fused according to the self-attention weights to determine the rPPG signal of the original video.
5. An rPPG signal extraction device, characterized in that, include: The first acquisition module is used to acquire the original video of the rPPG signal to be extracted; The first enhancement module is used to enhance the channels of the original video to obtain a channel-enhanced video; The first enhancement module is also used to generate Boolean images and chromatograms based on the original video; The channel-enhanced video is generated based on the Boolean image, the chromatogram, and the original video. The first fusion module is used to extract the spatiotemporal fusion features of the rPPG signal from the channel-enhanced video; The first segmentation module is used to segment the spatiotemporal fusion features into feature sequences with different frame rates; The first extraction module is used to extract rPPG sub-signals of the original video at different frame rates from each of the feature sequences; The first extraction module is further configured to determine the spatiotemporal attention features of each of the feature sequences based on each of the feature sequences; The rPPG sub-signal of the original video at the current frame rate is determined based on each of the aforementioned spatiotemporal attention features; The first determining module is used to determine the rPPG signal of the original video based on the rPPG sub-signals at each frame rate.
6. A computer electronic production equipment, characterized in that, It includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 4.
8. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method as claimed in any one of claims 1-4.
Citation Information
Patent Citations
Device and method for obtaining pulse transit time and / or pulse wave velocity information of a subject
CN105792742A
RPPG signal preprocessing method and system for video heart rate detection
CN113243900A