Video instantaneous heart rate detection method based on time-frequency spectrum conversion and image enhancement network

CN118212568BActive Publication Date: 2026-08-21HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410416513.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-08
Publication Date
2026-08-21
Estimated Expiration
2044-04-08

AI Technical Summary

Technical Problem

[0005]以上方法通常是获得血容量脉冲信号或心率值,但所获取的心率值往往是一段时间内的平均心率值,缺乏瞬时心率信息的获取;即便是获取血容量脉冲信号,后续的频谱分析由于所获得的血容量脉冲信号质量不佳,往往也是获得平均心率值,难以提取瞬时心率值

Benefits of technology

[0043]1、本发明将时频谱转换和图像增强网络相结合,一方面通过时频谱转换挖掘数据的时空-频域特征,使人脸视频的瞬时心率提取成为可能,另一方面通过图像增强网络对原始时频谱进行增强,从而提高了人脸视频瞬时心率提取的准确性;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118212568B_ABST
    Figure CN118212568B_ABST
Patent Text Reader

Abstract

The application discloses a video instantaneous heart rate detection method based on time-frequency spectrum conversion and an image enhancement network, and comprises the following steps: 1, selecting a face region of interest from a video to obtain a target BVP time sequence; 2, converting the target BVP time sequence into an original time-frequency spectrum through time-frequency spectrum conversion; 3, establishing an image enhancement network to enhance the original time-frequency spectrum and obtain an enhanced time-frequency spectrum; 4, training the image enhancement network; and 5, extracting an instantaneous heart rate value from the enhanced time-frequency spectrum. The application combines time-frequency spectrum conversion and the image enhancement network, extracts a time-frequency spectrum containing instantaneous heart rate information from a face video, can improve the accuracy of instantaneous heart rate detection, and thus provides a new method for non-contact instantaneous heart rate detection based on a video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of non-contact heart rate detection, specifically a method for extracting instantaneous heart rate values ​​from face videos based on time-spectrum transformation and image enhancement networks. Background Technology

[0002] Heart rate is a vital sign for the human body and is of great significance in fields such as disease diagnosis, health monitoring, and affective computing. Heart rate detection methods are generally divided into contact and non-contact methods. Contact methods rely on specific sensors to contact the subject's skin, and long-term monitoring may cause discomfort. Video-based non-contact heart rate detection technology is gradually becoming popular due to the advantages of inexpensive cameras, widespread use, and the fact that it does not require wearing a mask.

[0003] To improve the accuracy of heart rate detection, existing methods are generally classified into four categories: signal decomposition methods, blind source separation methods, model-based methods, and deep learning-based methods that have emerged with the development of artificial intelligence technology. The first three methods manually extract features from raw data to calculate heart rate based on certain algorithms or models, lacking generalization ability. Deep learning-based methods, due to their excellent data processing capabilities and self-learning abilities, have gradually become the mainstream approach in this field.

[0004] Existing deep learning-based methods are mostly divided into two categories: end-to-end and feature decoder types. The former directly establishes a mapping from video frames to blood volume pulse signals or heart rate values. This type of method heavily relies on the learning ability of neural networks and is prone to training burdens when noise interference is present. The latter mainly extracts latent features from video frames through preprocessing and other methods, and then obtains the blood volume pulse signal or heart rate value through a decoder. The input features of the decoder can be obtained in various ways, including leveraging the network's powerful learning ability to mine latent features in the data, simultaneously mining the temporal and spatial features of the data to obtain spatiotemporal features, and using traditional color difference models to obtain a preliminarily denoised blood volume pulse signal as the network input.

[0005] The methods described above typically obtain blood volume pulse signals or heart rate values. However, the obtained heart rate values ​​are often averaged over a period of time, lacking instantaneous heart rate information. Even when blood volume pulse signals are obtained, subsequent spectral analysis often yields average heart rate values ​​due to the poor quality of the obtained pulse signals, making it difficult to extract instantaneous heart rate values. However, in real-world applications of rPPG, there is an increasing need to detect more refined physiological parameters, such as stress detection and emotion perception, with a greater focus on instantaneous heart rate detection. Therefore, there is an urgent need for a method capable of accurately extracting instantaneous heart rate values. Summary of the Invention

[0006] This invention aims to address the shortcomings of existing technologies by proposing a video instantaneous heart rate detection method based on time-spectrum conversion and image enhancement networks. The goal is to extract instantaneous heart rate values ​​from face videos and improve the accuracy of instantaneous heart rate detection, thereby promoting the application of non-contact heart rate detection methods in fields such as medicine.

[0007] To achieve the above-mentioned objectives, the present invention adopts the following technical solution:

[0008] The present invention provides a video instantaneous heart rate detection method based on time-spectrum transformation and image enhancement networks, characterized by the following steps:

[0009] Step 1: Select a sub-region of interest from multiple face regions in the subject's video image and obtain the target BVP sequence set X = {X} [1] ,X [2] ,...,X [q] ,...,X [Q]}, X [q] Let X represent the target BVP sequence of the q-th sub-region of interest, and let X be... [q] =[X1 [q] X2 [q] ,...,X j [q] ,...,X J [q] ], where X j [q] Let Y represent the j-th color difference value in the target BVP sequence of the q-th sub-region of interest; Q < M; let Y be the label of X;

[0010] Step 2: Convert the target BVP sequence set X into the original time-spectrum set. in, This represents the original time-frequency spectrum after the q-th normalization.

[0011] Will All normalized original time spectra are spliced ​​together to form a multi-channel time spectrum. Where W represents the width of the spectrum when there are multiple channels, and H represents the height of the spectrum when there are multiple channels;

[0012] Step 3: Establish an image enhancement network comprising a downsampling module, a self-attention module, and an upsampling module, and train it to obtain a trained image enhancement model for multi-channel time-frequency spectrum processing. Feature extraction is performed to obtain the enhanced time-frequency spectrum S. out ;

[0013] Step 4: Extract instantaneous heart rate value;

[0014] For S outDivide the data along the time axis into L segments, B = [B1, B2, ..., B]. l ,...,B L ]; among them, B l S represents the strong time spectrum out The l-th segment in; for the l-th segment B l All time points on the timeline The frequency axis coordinates corresponding to the maximum pixel value along the frequency axis direction were extracted for each. Thus, the set of frequency values ​​f = {f [1] ,f [2] ,...,f [l] ,...f [L]};in, f represents the τ-th time point in the l-th segment. t [l] f represents the frequency axis coordinate corresponding to the maximum pixel value at the τ-th time point in the l-th segment along the frequency axis. [l] This represents the frequency axis coordinate corresponding to the maximum pixel value along the frequency axis at all time points of the l-th segment;

[0015] The mean vector is obtained by averaging all frequency axis coordinates in f. in, Indicates the l-th segment B l The average frequency axis coordinates corresponding to the maximum pixel value along the frequency axis at all time points; for f mean After unit conversion, L heart rate values ​​are obtained.

[0016] The video instantaneous heart rate detection method based on time-spectrum conversion and image enhancement network described in this invention is characterized in that step 1 is performed as follows:

[0017] Step 1.1: Use the facial feature point detection method to identify the face region in the first frame of the subject's facial video image, and select M sub-regions of interest from the face region in the first frame. Then use the feature point tracking algorithm to identify and locate the face regions in the remaining frames of the facial video, and obtain M sub-regions of interest for each frame of the facial video image.

[0018] Step 1.2: Calculate the mean pixel value of each sub-region of interest in the j-th frame of the facial video image, thereby obtaining the mean pixel value of each sub-region of interest in the J-th frame of the facial video image, and forming a time series of the mean pixel value of each sub-region of interest; let the time series of the mean pixel value of any single sub-region of interest in the J-th frame of the facial video image include: the time series of the red channel of the single sub-region of interest in the J-th frame of the facial video image R = [R1, R2, ..., R j ,...,R JThe time series of the green channel of a single sub-region of interest in a J-frame facial video image, G = [G1, G2, ..., G]. j ,...,G J The time series B = [B1, B2, ..., B] of the blue channel of a single sub-region of interest in the J-frame facial video image. j ,...,B J ]; where R j G represents the pixel mean of the red channel of a single sub-region of interest in the j-th frame of the facial video image. j B represents the pixel mean of the green channel of a single region of interest in the j-th frame of the facial video image. j This represents the average pixel value of the blue channel in a single sub-region of interest in the j-th frame of the facial video image; j = 1, 2, ..., J; J represents the total number of frames in the facial video image;

[0019] Perform chromatic aberration transformation on the R, G, and B values ​​of a single region of interest in a J-frame facial video image to obtain the target BVP sequence P = [P1, P2, ..., P...]. j ,...,P J ]; where P j This represents the j-th color difference value of a single sub-region of interest;

[0020] Step 1.3: Calculate the signal-to-noise ratio (SNR) of the target BVP sequences for the M sub-regions of interest and sort them in descending order. Then, select the target BVP sequences corresponding to the top Q sub-regions of interest, denoted as set X = {X...} [1] ,X [2] ,…,X [q] ,…,X [Q]}

[0021] The original time-frequency spectrum set in step 2 It is obtained by following these steps:

[0022] Step 2.1: Perform normalized bandpass filtering on the filtered target BVP sequence set X to obtain the normalized bandpass filtered target BVP sequence set. in, This represents the target BVP sequence after normalized bandpass filtering for the q-th sub-region of interest;

[0023] Step 2.2, for Perform continuous wavelet transform to obtain Continuous wavelet transform spectrum (WT) in the time domain [q] (a,b); where a represents the scaling factor of the continuous wavelet transform spectrum and b represents the time shift factor of the continuous wavelet transform spectrum;

[0024] Step 2.3, WT [q] (a,b) is transformed from the time domain to the frequency domain to obtain Frequency domain spatial continuous wavelet transform spectrum

[0025] Step 2.4, Calculation The phase partial derivative is obtained. instantaneous frequency ω [q] (a,b);

[0026] Step 2.5, WT [q] (a,b) Transform from the time-scale domain to the time-frequency domain to obtain the continuous wavelet transform spectrum WT in the time-frequency domain. [q] (ω [q] (a,b),b);

[0027] For WT [q] (ω [q] The interval (a,b),b) belongs to the interval The values ​​are compressed and recombined to obtain the original temporal spectrum S of the q-th sub-region of interest. [q] ;in, WT [q] (ω [q] The k-th center frequency of (a,b),b); Δω represents WT [q] (ω [q] The parameters of the compressed recombination interval of (a,b),b);

[0028] Step 2.6: Following the process of steps 2.2 to 2.5, obtain the original time-frequency spectrum set S = {S [1] ,S [2] ,...,S [q] ,...,S [Q]};

[0029] Step 2.7: Perform modulo and clipping processing on S to obtain the original time-frequency spectrum set after modulo and clipping. in, This represents the original temporal spectrum after modulo clipping of the q-th sub-region of interest;

[0030] Step 2.8, for the set Each original temporal spectrum after modulo cropping is normalized using min-max normalization, and the image dimensions are adjusted to W×H to obtain the normalized set of original temporal spectra.

[0031] Step 3 is performed as follows:

[0032] Step 3.1: The downsampling module sequentially includes: an input layer, D downsampling layers, and P linear projection layers; wherein, the input layer includes: a first convolutional layer, a first set of normalization layers, and a first activation layer; each downsampling layer includes: G second set of normalization layers, G second convolutional layers, and a second activation layer; each linear projection layer includes: a third convolutional layer and a first dropout layer;

[0033] The input is processed by the downsampling module, and after feature extraction and downsampling dimensionality reduction, a downsampled feature map is obtained. Where C1 represents the number of channels in the downsampled feature map, F1 represents the width of the downsampled feature map, and T1 represents the height of the downsampled feature map;

[0034] Step 3.2: The self-attention module includes: T self-attention layers and 1 output normalization layer; wherein, each self-attention layer includes: a normalization unit and a self-attention unit;

[0035] The input is processed by the self-attention module, and after dynamic attention weight allocation, a global feature map is obtained. Where C2 represents the number of channels in the global feature map, F2 represents the width of the global feature map, and T2 represents the height of the global feature map;

[0036] Step 3.3: The upsampling module includes one input normalization layer, D upsampling layers and one output layer. Each upsampling layer includes two activation convolutional layers and an upsampling interpolation layer. The d-th downsampling layer and the (D-d+1)-th upsampling layer are connected in a skip connection.

[0037] The input to the upsampling module undergoes upsampling and dimensionality enhancement, followed by fully connected mapping, to obtain the enhanced time-frequency spectrum.

[0038] Step 3.4, based on Y and Construct a binary cross-entropy loss function and a structural similarity loss function;

[0039] Step 3.5: Train the image enhancement network using the ADAM optimizer, calculate the gradient of the loss function to update the network parameters, until the maximum number of iterations is reached or the loss function converges, thereby obtaining the trained image enhancement model.

[0040] The present invention provides an electronic device, including a memory and a processor, wherein the memory is used to store a program that supports the processor in executing the video instantaneous heart rate detection method, and the processor is configured to execute the program stored in the memory.

[0041] The present invention discloses a computer-readable storage medium on which a computer program is stored, wherein the computer program, when executed by a processor, performs the steps of the video instantaneous heart rate detection method.

[0042] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0043] 1. This invention combines time-spectrum transformation and image enhancement network. On the one hand, it mines the spatiotemporal-frequency domain features of the data through time-spectrum transformation, making it possible to extract the instantaneous heart rate of face videos. On the other hand, it enhances the original time spectrum through image enhancement network, thereby improving the accuracy of instantaneous heart rate extraction from face videos.

[0044] 2. This invention extracts the time spectrum of heart rate frequency changes over time from multiple regions of interest in a face video, obtaining multiple spatiotemporal-frequency features that jointly contain instantaneous heart rate information, reducing interference caused by noise errors and improving the feasibility of instantaneous heart rate extraction.

[0045] 3. This invention establishes an image enhancement network to enhance the original time spectrum. Through the network's high-resolution feature extraction capability and global information modeling capability, it learns the distribution patterns of various signal features in the time spectrum, thereby improving the image quality of the original time spectrum and increasing the accuracy of instantaneous heart rate extraction. Attached Figure Description

[0046] Figure 1 This is a flowchart of the video instantaneous heart rate detection method based on time-spectrum conversion and image enhancement network of the present invention;

[0047] Figure 2 This is a schematic diagram of the network downsampling module structure of the present invention;

[0048] Figure 3 This is a schematic diagram of the network self-attention module structure of the present invention;

[0049] Figure 4 This is a schematic diagram of the network sampling module structure of the present invention;

[0050] Figure 5 This is a schematic diagram of the network jump connection module structure of the present invention. Detailed Implementation

[0051] In this embodiment, as Figure 1 As shown, a video instantaneous heart rate detection method based on time-spectrum transformation and image enhancement networks includes the following steps:

[0052] Step 1: Select the region of interest from multiple face regions in the subject's video image and obtain the target BVP sequence;

[0053] Step 1.1: Using the facial feature point detection method, identify the face region in the first frame of the subject's facial video image, and select M sub-regions of interest from the face region in the first frame. Then, use the feature point tracking algorithm to identify and locate the face regions in the remaining frames of the facial video, and obtain M sub-regions of interest for each frame of the facial video image. In this embodiment, the facial video is segmented with a window length of 10 seconds and a step size of 1 second to obtain several video samples; M = 100.

[0054] Step 1.2: Calculate the mean pixel value of each sub-region of interest in the j-th frame of the facial video image, thereby obtaining the mean pixel value of each sub-region of interest in the J-th frame of the facial video image, and forming a time series of the mean pixel value of each sub-region of interest; let the time series of the mean pixel value of any single sub-region of interest in the J-th frame of the facial video image include: the time series of the red channel of the single sub-region of interest in the J-th frame of the facial video image R = [R1, R2, ..., R j ,...,R J The time series of the green channel of a single sub-region of interest in a J-frame facial video image, G = [G1, G2, ..., G]. j ,...,G J The time series B = [B1, B2, ..., B] of the blue channel of a single sub-region of interest in the J-frame facial video image. j ,...,B J ]; where R j G represents the pixel mean of the red channel of a single sub-region of interest in the j-th frame of the facial video image. j B represents the pixel mean of the green channel of a single region of interest in the j-th frame of the facial video image. j This represents the average pixel value of the blue channel in a single sub-region of interest in the j-th frame of the facial video image; j = 1, 2, ..., J; J represents the total number of frames in the facial video image;

[0055] Perform chromatic aberration transformation on the R, G, and B values ​​of a single region of interest in a J-frame facial video image to obtain the target BVP sequence P = [P1, P2, ..., P...]. j ,...,P J ]; where P j This represents the j-th color difference value of a single sub-region of interest.

[0056] Step 1.3: Calculate the signal-to-noise ratio (SNR) of the target BVP sequences for the M sub-regions of interest and sort them in descending order. Then, select the target BVP sequences corresponding to the top Q sub-regions of interest, denoted as set X = {X...} [1] ,X [2] ,…,X [q] ,…,X[Q]}, where X [q] Let X represent the target BVP sequence of the selected q-th sub-region of interest, and let X be... [q] =[X1 [q] X2 [q] ,...,X j [q] ,...,X J [q] ], where X j [q] Let X represent the j-th color difference value in the target BVP sequence of the q-th sub-region of interest; Q < M; Let X = {X [1] ,X [2] ,…,X [q] ,…,X [Q] The label for} is denoted as Y; in this embodiment, Q = 3 and M = 100.

[0057] Step 2: Convert the target BVP sequence into the original time spectrum;

[0058] Step 2.1: Perform normalized bandpass filtering on the filtered target BVP sequence set X to obtain the normalized bandpass filtered target BVP sequence set. in, This represents the target BVP sequence after normalized bandpass filtering for the q-th sub-region of interest; The specific method of obtaining it is shown in equation (1):

[0059]

[0060] Where mean(·) represents the calculation of the average value, and std(·) represents the calculation of the standard deviation.

[0061] Step 2.2, for Perform a continuous wavelet transform, and obtain the result through equation (2). Continuous wavelet transform spectrum (WT) in the time domain [q] (a,b):

[0062]

[0063] In equation (2), a represents the scaling factor and b represents the time shift factor. This represents the mother wavelet function.

[0064] Step 2.3, WT [q] (a,b) is transformed from the time domain to the frequency domain, and obtained through equation (3). Frequency domain spatial continuous wavelet transform spectrum

[0065]

[0066] In equation (3), and In equation (1) and The Fourier transform result.

[0067] Step 2.4, calculate using equation (4) The phase partial derivative is obtained. instantaneous frequency ω [q] (a,b):

[0068]

[0069] Step 2.5, WT [q] (a,b) Transform from the time-scale domain to the time-frequency domain to obtain the continuous wavelet transform spectrum WT in the time-frequency domain. [q] (ω [q] (a,b),b); WT is obtained through equation (5) [q] (ω [q] The interval (a,b),b) belongs to the interval The values ​​are compressed and recombined to obtain the original temporal spectrum S of the q-th sub-region of interest. [q] ;

[0070]

[0071] in, WT [q] (ω [q] The k-th center frequency of (a,b),b); Δω represents WT [q] (ω [q] The compressed reconstruction interval parameters of (a,b),b) are: Since signal processing in practical applications consists of discrete values, the parameters in equation (5) are as follows: a k Let Δa represent the k-th discrete scaling factor. k =a k -a k-1 .

[0072] Step 2.6: Following the process of steps 2.2 to 2.5, obtain the original time-frequency spectrum set S = {S [1] ,S [2] ,...,S [q] ,...,S [Q]};

[0073] Step 2.7: Due to edge effects in the time-spectrum conversion result, i.e., the left and right ends of the original time-spectrum may become blurred due to algorithmic issues, the original time-spectrum set S = {S [1] ,S[2] ,...,S [q] ,...,S [Q] The first and last seconds are trimmed, leaving only the segment from 1 to 9 seconds in between. Simultaneously, since the heart rate range is 0.65–2.5 Hz, this range is also trimmed along the frequency axis. The original time spectrum is then modulo-diminutive and normalized to 0–1 using a max-min normalization method, resulting in the original time spectrum set after modulo-trimming. in, This represents the original temporal spectrum after modulo clipping of the q-th sub-region of interest.

[0074] Step 2.8, for the set The original temporal spectrum after each modulus clipping in the data Normalization was performed using the min-max normalization method, and the image dimensions were adjusted to W×H to obtain the normalized original time-frequency spectrum set. in, This represents the original time-frequency spectrum after the q-th normalization.

[0075] Will All normalized original time spectra are spliced ​​together to form a multi-channel time spectrum. W represents the width of the spectrum when there are multiple channels, and H represents the height of the spectrum when there are multiple channels; in this embodiment, W = 224 and H = 224.

[0076] Step 3: Establish the image enhancement network, which includes a downsampling module, a self-attention module, and an upsampling module, used to process the multi-channel time-frequency spectrum. Feature extraction calculations are performed to obtain the enhanced time spectrum.

[0077] Step 3.1, the structural diagram of the downsampling module is as follows: Figure 2 As shown, it includes an input layer, D downsampling layers, and P linear projection layers; the input layer includes: an N1×N1 first convolutional layer, a first set of normalization layers, and a first activation layer; the D downsampling layers are denoted as Down1, Down2, ..., Down d Down D Let P linear projection layers be denoted as Proj1, Proj2, ..., Proj p ,...,Proj P Among them, Down d Proj represents the d-th downsampling layer. pLet P represent the p-th linear projection layer; P linear projection layers are connected to D downsampling layers; each downsampling layer includes: G second group normalization layers, G N2×N2 second convolutional layers, and a second activation layer; each linear projection layer includes: an N3×N3 third convolutional layer and a first dropout layer; in this embodiment, D=3, P=1, G=3, N1=7, N2=1, N3=1;

[0078] The first downsampling layer (Down1) has 256 channels, the second downsampling layer (Down2) has 512 channels, and the third downsampling layer (Down3) has 1024 channels.

[0079] In the input downsampling module, after feature extraction and downsampling dimensionality reduction, a downsampled feature map is obtained. Where C1 represents the number of channels in the downsampled feature map, F1 represents the width of the downsampled feature map, and T1 represents the height of the downsampled feature map; in this embodiment, C1 = 256, F1 = 28, and T1 = 28.

[0080] Step 3.2, Schematic diagram of the self-attention module structure as shown below Figure 3 As shown, it includes T self-attention layers and 1 output normalization layer. The T self-attention layers are denoted as Atten1, Atten2, ..., Atten... t ,...,Atten T A single output normalization layer is denoted as LN; where Atten t Let represent the t-th self-attention layer; each self-attention layer includes: a normalization unit and a self-attention unit; each normalization unit includes: two third-group normalization layers, two first fully connected layers and a second dropout layer; each self-attention unit includes: a q-type fully connected layer, a k-type fully connected layer, a v-type fully connected layer, a first output layer, and a softmax layer; the output normalization layer includes: a first normalization layer; the t-th self-attention layer is followed by one output normalization layer;

[0081] The input is processed by the self-attention module, and after dynamic attention weight allocation, a global feature map is obtained. Where C2 represents the number of channels in the global feature map, F2 represents the width of the global feature map, and T2 represents the height of the global feature map; in this embodiment, T = 6; C2 = 512, F2 = 14, and T2 = 14.

[0082] Step 3.3, Schematic diagram of the upsampling module structure as shown below Figure 4 As shown, it includes one input normalization layer, D upsampling layers, and one output layer. The input normalization layer is denoted as BN, and the U upsampling layers are denoted as Up1, Up2, ..., Up.d ,...,Up D One output layer is denoted as Out; where Up d Let d represent the d-th upsampling layer; the input normalization layer includes: an N4×N4 fourth convolutional layer, a first batch of normalization layers, and a third activation layer; each upsampling layer includes: two activation convolutional layers and an upsampling interpolation layer; each activation convolutional layer includes: an N5×N5 fifth convolutional layer, a second batch of normalization layers, and a fourth activation layer; the output layer includes: a sixth convolutional layer; wherein, there is a skip connection between the d-th downsampling layer and the (D-d+1)-th upsampling layer; in this embodiment, the skip connection path structure is illustrated as follows. Figure 5 As shown, there are three jump connection paths: The first path concatenates the output of the third downsampling layer Down3 in the downsampling module with the output of the input normalization layer in the upsampling module in terms of channel count, and uses it as the input of the first upsampling layer Up1 in the upsampling module; The second path concatenates the output of the second downsampling layer Down2 in the downsampling module with the output of the first upsampling layer Up1 in the upsampling module in terms of channel count, and uses it as the input of the second upsampling layer Up2 in the upsampling module; The third path concatenates the output of the first downsampling layer Down1 in the downsampling module with the output of the second upsampling layer Up2 in the upsampling module in terms of channel count, and uses it as the input of the third upsampling layer Up3 in the upsampling module; D=3, N4=3, N5=3;

[0083] The first upsampling layer Up1 has 1024 channels, the second upsampling layer Up2 has 512 channels, and the third upsampling layer Up3 has 256 channels.

[0084] In the input upsampling module, after upsampling and dimensionality enhancement and fully connected mapping, the enhanced time spectrum is obtained. In this embodiment, D = 3, N4 = 3, and N5 = 3.

[0085] Step 3.4, based on Y and Construct a binary cross-entropy loss function (BCELoss) and a structural similarity loss function (SSIMLoss);

[0086] Specifically, the pixel values ​​in the dominant frequency range of the original time-frequency spectrum are close to 1, while the pixel values ​​in the non-dominant frequency range are close to 0. Therefore, image enhancement of the original time-frequency spectrum is essentially a binary classification task of image segmentation, and BCELoss has wide applications in this binary classification task. Therefore, the method of this invention uses BCELoss to establish the loss function of the image enhancement network. This loss function l BCE (x,y) is defined as shown in equations (6) and (7):

[0087]

[0088]

[0089] In equations (6) and (7), x n y represents the nth pixel value of the spectrum during enhancement. n This represents the nth pixel value in the spectrum when the label is applied, and σ(·) represents the sigmoid function. BCELoss, ω represents the value of the nth pixel. n This represents the weighting coefficient of the nth BCELoss. Additionally, BCELoss provides a positive sample weighting option (pos-weight) to address the imbalance in the distribution of positive and negative samples. Since the number of pixels (positive samples) in the dominant heart rate frequency range of the time spectrum is smaller than that in the non-dominant heart rate frequency range (negative samples), resulting in a relatively unbalanced distribution, an appropriate value for pos-weight can significantly improve the generation quality of the time spectrum. In this embodiment, pos-weight is set to 5.

[0090] Furthermore, since the time-frequency spectrum contains various signal features, such as heart rate signal features and noise signal features, and these different signal features are mainly manifested in structural features, the structural features of noise signals, such as frequency domain distribution and energy distribution, will differ from those of heart rate signals. Therefore, in order for the image enhancement network to learn the structural features of different signals within the time-frequency spectrum and to distinguish noise signals from heart rate signals based on these features, ultimately learning the frequency change trend, the method of this invention uses SSIMLoss to establish the loss function of the image enhancement network; the formula for SSIMLoss is... SSIM (x,y) is defined as shown in equations (8) and (9):

[0091]

[0092]

[0093] In equations (8) and (9), μ x μ represents the pixel mean of the spectrum during enhancement. y σ represents the pixel mean of the spectrum when the label is applied. xσ represents the standard deviation of pixel values ​​in the enhanced spectrum. y C1 and C2 represent the standard deviation of pixel values ​​in the time spectrum of the tag, and are offset constants to prevent the numerator or denominator from being zero. In practical applications, the time spectrum is divided into n blocks. This represents the SSIMLoss of the nth block.

[0094] Step 3.5: Train the image enhancement network using the ADAM optimizer, calculate the gradient of the loss function to update the network parameters, until the maximum number of iterations is reached or the loss function converges, thereby obtaining a trained image enhancement model, which is used to enhance the original temporal spectrum;

[0095] Step 4: Extract instantaneous heart rate value;

[0096] Multi-channel time spectrum The image is input into a trained image enhancement model for processing, resulting in an enhanced temporal spectrum S with dimensions W×H. out ; For S out Divide the data along the time axis into L segments, B = [B1, B2, ..., B]. l ,...,B L ]; among them, B l S represents the strong time spectrum out The l-th segment in; for the l-th segment B l All time points on the timeline The frequency axis coordinates corresponding to the maximum pixel value along the frequency axis direction were extracted for each. Thus, the set of frequency values ​​f = {f [1] ,f [2] ,...,f [l] ,...f [L]};in, f represents the τ-th time point in the l-th segment. t [l] f represents the frequency axis coordinate corresponding to the maximum pixel value at the τ-th time point in the l-th segment along the frequency axis. [l] This represents the frequency axis coordinate corresponding to the maximum pixel value along the frequency axis at all time points of the l-th segment.

[0097] The mean vector is obtained by averaging all frequency axis coordinates in f. in, Indicates the l-th segment B l The average frequency axis coordinates corresponding to the maximum pixel value along the frequency axis at all time points; for f meanAfter unit conversion, L heart rate values ​​are obtained. In this embodiment, since the time spectrum length used as network input is clipped to 8 seconds, L = 8, that is, 1 heart rate value is obtained every 1 second, thereby achieving the purpose of instantaneous heart rate detection.

[0098] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the above-described method, and the processor is configured to execute the program stored in the memory.

[0099] In this embodiment, a computer-readable storage medium stores a computer program, which is executed by a processor to perform the steps of the above method.

[0100] To verify the effectiveness of the method of this invention, performance validation was performed on the publicly available datasets UBFC-RPPG and PURE. The UBFC-RPPG dataset contains videos of 42 subjects in real-world environments. Subjects were asked to play a time-sensitive mathematical game to maintain changes in heart rate (HR). The PURE dataset contains 60 videos from 10 subjects (8 men and 2 women). Each subject performed six different head movements, including stabilization, speaking, slow translation, fast translation, small rotation, and moderate rotation. This invention uses conventional evaluation metrics to assess the instantaneous heart rate detection performance, specifically including mean absolute error (HR). MAE (mean absolute error, MAE), root mean square error (HR) RMSE (root mean square error, RMSE) and Pearson correlation coefficient HR P (Pearson's Correlation Coefficient, P).

[0101] Table 1. In-dataset test results for PURE dataset

[0102]

[0103] Table 2. Test results of PURE to UBFC-RPPG across datasets.

[0104]

[0105] Table 1 presents the results of the proposed method and several other methods on the PURE database. As can be seen from Table 1, the proposed method achieves good results in all three metrics. Specifically, the proposed method reduces RMSE and MAE by 39% and 33%, respectively, and improves P by 2%. While the comparative methods utilize spatiotemporal features, the lack of frequency domain features makes it difficult to perform instantaneous HR detection on the generated BVP signal. The proposed method, however, considers both the spatiotemporal and frequency domain features of the HR signal and performs image enhancement on the time-spectrum of the original BVP signal, enabling it to perform instantaneous HR detection.

[0106] Table 2 presents the test results of the proposed method across datasets from PURE to UBFC-RPPG. It can be seen that the proposed method outperforms other comparative methods. Specifically, the proposed method reduces RMSE and MAE by 18% and 23%, respectively, and improves P by 2%. This is because we extracted the spatiotemporal and frequency domain features of the original data. The time-spectrum format effectively reduces the distribution differences between different datasets, allowing the network to focus on the time-frequency features of the data and largely avoid noise interference, thus achieving good generalization. Simultaneously, by combining with an image enhancement network, the network can accurately learn the latent features of the data, thereby generating an enhanced time-spectrum that accurately depicts the instantaneous changes in heart rate.

Claims

1. A video instantaneous heart rate detection method based on time-spectrum transformation and image enhancement network, characterized in that, The procedure is as follows: Step 1: Select sub-regions of interest from multiple face regions in the subject's video images and obtain a set of target BVP sequences. , Indicates the first The target BVP sequence of each sub-region of interest, and ,in, Indicates the first The target BVP sequence of the nth sub-region of interest Individual color difference values; ; make The label is recorded as ; The number of sub-regions of interest; Step 2: Set the target BVP sequence Convert to the original time-frequency spectrum set ;in, Indicates the first The normalized original time spectrum; Will All normalized original time spectra are spliced ​​together to form a multi-channel time spectrum. Where W represents the width of the spectrum when using multiple channels, and H represents the height of the spectrum when using multiple channels; Step 3: Establish an image enhancement network comprising a downsampling module, a self-attention module, and an upsampling module, and train it to obtain a trained image enhancement model for multi-channel time-frequency spectrum processing. Feature extraction is performed to obtain the enhanced time-frequency spectrum. ; Step 4: Extract instantaneous heart rate value; right Divided according to the time axis A fragment ;in, Represents the strong time spectrum The first in The segment; for the first segment A fragment All time points on the timeline The frequency axis coordinates corresponding to the maximum pixel value along the frequency axis direction were extracted for each. Thus, a set of frequency values ​​is obtained. ;in, Indicates the first The first segment At a certain point in time, Indicates the first The first segment The frequency axis coordinates corresponding to the maximum pixel value at each time point along the frequency axis. Indicates the first The frequency axis coordinates corresponding to the maximum pixel value along the frequency axis at all time points of each segment; right All frequency axis coordinates are averaged to obtain the mean vector. ;in, Indicates the first A fragment The average frequency axis coordinate corresponding to the maximum pixel value along the frequency axis at all time points; for After performing a unit transformation, we obtain Heart rate values.

2. The video instantaneous heart rate detection method based on time-spectrum transformation and image enhancement network according to claim 1, characterized in that, Step 1 is performed as follows: Step 1.1: Using the facial landmark detection method, identify the face region in the first frame of the subject's facial video image, and select from the face region in the first frame. The algorithm identifies and locates facial regions in each frame of the facial video using a feature point tracking algorithm, thus obtaining the facial video image for each frame. One sub-region of interest; Step 1.2, calculate the first... The average pixel value of each sub-region of interest in a frame of facial video image is obtained. The pixel mean of each sub-region of interest in a frame of facial video image is used to construct a time series of the pixel mean of each sub-region of interest; let... The time series of pixel mean values ​​for any single sub-region of interest in a frame-by-frame facial video image includes: Time series of the red channel of a single sub-region of interest in a frame of facial video image , Time series of the green channel of a single sub-region of interest in a frame of facial video image as well as Time series of the blue channel of a single sub-region of interest in a frame of facial video image ;in, Indicates the first The pixel mean of the red channel in a single sub-region of interest in a frame of facial video image. Indicates the first The pixel mean of the green channel in a single region of interest in a frame of facial video image. Indicates the first The pixel mean of the blue channel in a single sub-region of interest in a frame of facial video image; ; This indicates the total number of frames in the facial video image; right A single sub-region of interest in a frame of facial video image , , Perform chromatic aberration transformation to obtain the target BVP sequence for a single sub-region of interest. ;in, The first subregion of interest represents the first subregion of interest. Individual color difference values; Step 1.3, Calculation The signal-to-noise ratio (SNR) of the target BVP sequences in each sub-region of interest is calculated and sorted in descending order to select the top-ranked sequences. The target BVP sequences corresponding to each sub-region of interest are denoted as set. .

3. The video instantaneous heart rate detection method based on time-spectrum transformation and image enhancement network according to claim 2, characterized in that, The original time-frequency spectrum set in step 2 It is obtained by following these steps: Step 2.1: Perform normalized bandpass filtering on the filtered target BVP sequence set X to obtain the normalized bandpass filtered target BVP sequence set. ,in, Indicates the first Normalized bandpass filtered target BVP sequence for each sub-region of interest; Step 2.2, for Perform continuous wavelet transform to obtain Continuous wavelet transform spectrum in the time domain Where a represents the scaling factor of the continuous wavelet transform spectrum and b represents the time shift factor of the continuous wavelet transform spectrum; Step 2.3, Transforming from the time domain to the frequency domain, we obtain Frequency domain spatial continuous wavelet transform spectrum ; Step 2.4, Calculation The phase partial derivative is obtained. instantaneous frequency ; Step 2.5, Transforming from the time-scale domain to the time-frequency domain yields the continuous wavelet transform spectrum in the time-frequency domain. ; right The middle belongs to the interval The value is compressed and recombined to obtain the first... Original time-frequency spectra of each sub-region of interest ;in, express The kth center frequency; express The parameters of the compressed recombination interval; Step 2.6: Following the process of steps 2.2 to 2.5, obtain the original time-frequency spectrum set. ; Step 2.7, for The original time-frequency spectrum set is obtained by performing modulus extraction and clipping. ;in, Indicates the first The original temporal spectrum after modulating and cropping the sub-regions of interest; Step 2.8, for the set Each original temporal spectrum after modulo cropping is normalized using min-max normalization, and the image dimensions are adjusted to [value missing]. Thus, the normalized original time-frequency spectrum set is obtained. .

4. The video instantaneous heart rate detection method based on time-spectrum transformation and image enhancement network according to claim 3, characterized in that, Step 3 is performed as follows: Step 3.1: The downsampling module includes, in sequence: an input layer, Each downsampling layer and Each linear projection layer comprises: a first convolutional layer, a first set of normalization layers, and a first activation layer; each downsampling layer includes: The second group of normalization layers, Each linear projection layer includes a second convolutional layer and a second activation layer; each convolutional layer includes a third convolutional layer and a first dropout layer. The input is processed by the downsampling module, and after feature extraction and downsampling dimensionality reduction, a downsampled feature map is obtained. ;in, This indicates the number of channels in the downsampled feature map. This indicates the width of the downsampled feature map. Indicates the height of the downsampled feature map; Step 3.2, the self-attention module includes: Each self-attention layer consists of one self-attention layer and one output normalization layer; wherein, each self-attention layer includes: one normalization unit and one self-attention unit; The input is processed by the self-attention module, and after dynamic attention weight allocation, a global feature map is obtained. ;in, This represents the number of channels in the global feature map. This represents the width of the global feature map. Indicates the height of the global feature map; Step 3.3: The upsampling module includes one input normalization layer. The system consists of one upsampling layer and one output layer. Each upsampling layer includes two activation convolutional layers and one upsampling interpolation layer. The downsampling layer and the first Skip connections between upsampling layers; The input to the upsampling module undergoes upsampling and dimensionality enhancement, followed by fully connected mapping, to obtain the enhanced time-frequency spectrum. ; Step 3.4, based on and Construct a binary cross-entropy loss function and a structural similarity loss function; Step 3.5: Train the image enhancement network using the ADAM optimizer, calculate the gradient of the loss function to update the network parameters, until the maximum number of iterations is reached or the loss function converges, thereby obtaining the trained image enhancement model.

5. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports the processor in executing any of the video instantaneous heart rate detection methods of claims 1-4, and the processor is configured to execute the program stored in the memory.

6. A computer-readable storage medium storing a computer program thereon, characterized in that, The computer program, when run by a processor, performs the steps of the video instantaneous heart rate detection method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Video-based human body heart rate and facial blood volume accurate detection method and system

    CN111626182A

  • DeepFake detection method and device, computer equipment and storage medium

    CN115100722A