Method for training rppg signal extraction model and rppg signal extraction method

CN118568486BActive Publication Date: 2026-09-29TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410490747.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-23
Publication Date
2026-09-29
Estimated Expiration
2044-04-23

AI Technical Summary

Technical Problem

这要求用户在拍摄面部视频的同时佩戴接触式传感器,例如穿戴式、指夹式等设备,导致收集数据标注的过程十分繁琐且受试者体验极差,难以形成大规模数据,且为了保证面部视频与PPG信号的同步收集还需开发独立系统,不仅产生了高昂的数据收集成本,更重要的是会由于PPG信号标签本身存在的误差导致算法性能受限

Benefits of technology

[0032]在本发明实施例中,无需PPG信号标签即可进行模型训练,摆脱了现有技术对信号标签的依赖性,能够大幅减少数据收集成本,同时后续可以基于训练的模型从面部视频中提取rPPG信号,实现非接触式生命体征监测,使得生命体征监测更加便捷。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118568486B_ABST
    Figure CN118568486B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a kind of rPPG signal extraction model training method and rPPG signal extraction method, comprising: face video slice is carried out spatial enhancement, obtain multiple with the rPPG signal frequency same positive sample of face video slice, and generate positive sample set with it, face video slice is carried out frequency enhancement, obtain multiple negative sample with different rPPG signal frequency, and generate negative sample set with it, input rPPG signal extraction model to positive sample set and negative sample set, obtain positive sample rPPG signal set and negative sample rPPG signal set, then update the parameter of rPPG signal extraction model with it, constantly repeat the above operation, until meet training stop condition, complete training, subsequent rPPG signal is extracted from face video based on the model of training, realize non-contact vital sign monitoring. In this way, model training can be carried out without PPG signal label, and signal label dependency is got rid of, and non-contact vital sign monitoring can be realized simultaneously.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of rPPG signal extraction technology, and in particular to an rPPG signal extraction model training method and an rPPG signal extraction method. Background Technology

[0002] Current training schemes for remote photoplethysmography (rPPG) signal extraction models typically have high label dependence. They require collecting facial video and simultaneously recorded photoplethysmography (PPG) signals, and then modeling the mapping relationship between facial video and PPG signal labels using PPG signals as labels during model training. This necessitates users wearing contact sensors, such as wearable or finger-clip devices, while capturing facial video, making data collection and labeling extremely cumbersome and resulting in a poor user experience. This hinders the generation of large-scale datasets. Furthermore, ensuring simultaneous collection of facial video and PPG signals requires developing a separate system, leading to high data collection costs and, more importantly, limitations in algorithm performance due to inherent errors in the PPG signal labels. This dependence restricts the technology's user-friendliness and incurs significant data collection costs, limiting its widespread adoption and promotion in practical applications. Summary of the Invention

[0003] This invention provides an rPPG signal extraction model training method and an rPPG signal extraction method.

[0004] In a first aspect, embodiments of the present invention provide a method for training an rPPG signal extraction model, the method comprising:

[0005] Spatial augmentation is performed on facial video slices to obtain multiple positive samples with the same rPPG signal frequency as the facial video slices, and a set of positive samples is generated based on this.

[0006] Frequency enhancement is performed on facial video slices to obtain multiple negative samples with different rPPG signal frequencies, and a negative sample set is generated from these samples.

[0007] Input the positive sample set and the negative sample set into the rPPG signal extraction model to obtain the positive sample rPPG signal set and the negative sample rPPG signal set.

[0008] The parameters of the rPPG signal extraction model are updated based on the positive sample rPPG signal set and the negative sample rPPG signal set;

[0009] Repeat the above steps until the training stops.

[0010] In some possible implementations of the first aspect, spatial augmentation is performed on the facial video slice to obtain multiple positive samples with the same frequency as the rPPG signal of the facial video slice, including:

[0011] Different spatial enhancement mechanisms were used to transform facial video slices frame by frame, resulting in multiple positive samples with the same rPPG signal frequency as the facial video slices.

[0012] In some possible implementations of the first aspect, frequency enhancement is performed on facial video slices to obtain multiple negative samples with different rPPG signal frequencies, including:

[0013] Randomly sample multiple different target frequency ratios within a preset frequency ratio range;

[0014] Based on multiple different target frequency ratios, the facial video slices are frequency modulated to obtain multiple negative samples with different rPPG signal frequencies.

[0015] In some possible implementations of the first aspect, facial video slices are frequency modulated according to multiple different target frequency ratios to obtain multiple negative samples with different rPPG signal frequencies, including:

[0016] The facial video slices are processed by a 3D convolutional layer, a 3D residual module, and a downsampling module to generate multi-scale visual features. At each scale, the target frequency ratio and the visual features corresponding to the scale are input into a frequency modulation block to modulate the rPPG signal frequency. The spatial dimension is unified by an upsampling module. Finally, the facial video slices are reconstructed by a 3D convolutional layer to obtain negative samples corresponding to the target frequency ratio.

[0017] Among the possible implementations of the first aspect, the rPPG signal extraction model includes: feature coding backbone network and hybrid expert model.

[0018] In some possible implementations of the first aspect, the positive sample set and the negative sample set are input into the rPPG signal extraction model to obtain the positive sample rPPG signal set and the negative sample rPPG signal set, including:

[0019] Input the positive and negative sample sets into the rPPG signal extraction model;

[0020] The feature encoding backbone network encodes the features of the input samples to generate corresponding visual feature maps;

[0021] The hybrid expert model uniformly divides the generated visual feature map into L sub-regions in the spatial dimension, extracts effective information from each sub-region, and uses it as the corresponding local expert, thus obtaining the local experts E1,…,E for each of the L sub-regions.l ,…,E L Through an internal spatiotemporal gating network, local experts E1,…,E are provided with... l ,…,E L Different spatiotemporal weights are assigned, and different local experts are aggregated in this way to obtain the rPPG signal corresponding to the sample;

[0022] A set of positive sample rPPG signals is generated based on the rPPG signal corresponding to each positive sample, and a set of negative sample rPPG signals is generated based on the rPPG signal corresponding to each negative sample.

[0023] In some possible implementations of the first aspect, the parameters of the rPPG signal extraction model are updated based on the sets of positive and negative rPPG signals, including:

[0024] Based on the set of positive sample rPPG signals, the set of negative sample rPPG signals, and a preset loss function, the loss value is calculated, and the parameters of the rPPG signal extraction model are updated accordingly.

[0025] Among some possible implementations of the first aspect, the loss function includes:

[0026] Frequency contrast loss function, frequency ratio consistency loss function, cross-video frequency consistency loss function, and video reconstruction loss function.

[0027] Secondly, embodiments of the present invention provide an rPPG signal extraction method, the method comprising:

[0028] Acquire the facial video to be processed;

[0029] The facial video is input into the rPPG signal extraction model to obtain the corresponding rPPG signal. The rPPG signal extraction model is obtained based on the above rPPG signal extraction model training method.

[0030] In some possible implementations of the second aspect, the method further includes:

[0031] The rPPG signal was analyzed to obtain the corresponding vital signs information.

[0032] In this embodiment of the invention, model training can be performed without PPG signal tags, eliminating the dependence on signal tags in existing technologies and significantly reducing data collection costs. Furthermore, the trained model can be used to extract rPPG signals from facial videos to achieve non-contact vital sign monitoring, making vital sign monitoring more convenient.

[0033] It should be understood that the description in the Summary of the Invention is not intended to limit the key or essential features of the embodiments of the present invention, nor is it intended to restrict the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0034] The above and other features, advantages, and aspects of the various embodiments of the present invention will become more apparent from the accompanying drawings and the following detailed description. The drawings are provided for a better understanding of the invention and are not intended to limit the invention. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0035] Figure 1 The flowchart illustrates a training method for an rPPG signal extraction model provided by an embodiment of the present invention.

[0036] Figure 2 An overall framework diagram provided by an embodiment of the present invention is shown;

[0037] Figure 3 A schematic diagram of a frequency enhancement submodule provided in an embodiment of the present invention is shown;

[0038] Figure 4 This diagram illustrates an rPPG signal extraction model provided by an embodiment of the present invention.

[0039] Figure 5 A flowchart of an rPPG signal extraction method provided by an embodiment of the present invention is shown;

[0040] Figure 6 A structural diagram of an exemplary electronic device capable of implementing embodiments of the present invention is shown. Detailed Implementation

[0041] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0042] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0043] To address the problems in the background technology, this invention provides an rPPG signal extraction model training method and an rPPG signal extraction method. Specifically, spatial enhancement is performed on facial video slices to obtain multiple positive samples with the same rPPG signal frequency as the facial video slices, and a positive sample set is generated. Frequency enhancement is performed on the facial video slices to obtain multiple negative samples with different rPPG signal frequencies, and a negative sample set is generated. The positive and negative sample sets are input into the rPPG signal extraction model to obtain a positive sample rPPG signal set and a negative sample rPPG signal set. The parameters of the rPPG signal extraction model are then updated based on these sets. The above operations are repeated until the training stopping condition is met, and the training is completed. Subsequently, the trained model extracts rPPG signals from facial videos to achieve non-contact vital sign monitoring.

[0044] In this way, model training can be performed without PPG signal labels, eliminating the dependence on signal labels in existing technologies and significantly reducing data collection costs. Furthermore, the trained model can be used to extract rPPG signals from facial videos to achieve non-contact vital sign monitoring, making vital sign monitoring more convenient.

[0045] The following detailed description, with reference to the accompanying drawings and specific embodiments, illustrates an rPPG signal extraction model training method and an rPPG signal extraction method provided by the present invention.

[0046] Figure 1 The flowchart of an rPPG signal extraction model training method provided by an embodiment of the present invention is shown, as follows: Figure 1 As shown, the rPPG signal extraction model training method 100 may include the following steps:

[0047] S110, Spatial enhancement is performed on the facial video slice to obtain multiple positive samples with the same rPPG signal frequency as the facial video slice, and a positive sample set is generated based on these samples.

[0048] In some embodiments, different spatial enhancement mechanisms, such as image rotation (0°, 90°, 180°, 270°) and image flipping (horizontal, vertical), can be used to transform the facial video slice frame by frame to obtain multiple positive samples with the same rPPG signal frequency as the facial video slice, and thus generate a positive sample set.

[0049] The principle behind the above operation is that these spatial enhancement mechanisms do not change the image color information, but only change the spatial position of the pixels. Therefore, the skin color change process caused by heartbeat in the facial video slice will not be disturbed, that is, the rPPG signal frequency will not change.

[0050] S120, frequency enhancement is performed on the facial video slices to obtain multiple negative samples with different rPPG signal frequencies, and a negative sample set is generated from these samples.

[0051] In some embodiments, multiple different target frequency ratios can be randomly sampled within a preset frequency ratio range. Based on the multiple different target frequency ratios, the facial video slices are frequency modulated to obtain multiple negative samples with different rPPG signal frequencies, and a negative sample set is generated accordingly.

[0052] For example, facial video slices can be processed using 3D convolutional layers (3D Conv), 3D residual modules (3D RB), and downsampling modules to generate multi-scale visual features. At each scale, the target frequency ratio and the corresponding visual features are input into a frequency modulation block (FMB) to modulate the rPPG signal frequency. An upsampling module unifies the spatial dimension, and finally, the facial video slices are reconstructed using 3D convolutional layers to obtain negative samples corresponding to the target frequency ratio. Because the target frequency ratios are different, the negative samples obtained based on the above operations have different rPPG signal frequencies.

[0053] S130, input the positive sample set and the negative sample set into the rPPG signal extraction model to obtain the positive sample rPPG signal set and the negative sample rPPG signal set.

[0054] For example, the rPPG signal extraction model may include a feature encoding backbone network and a hybrid expert model. To this end, a set of positive samples and a set of negative samples can be input into the rPPG signal extraction model. The feature encoding backbone network encodes the features of the input samples to generate corresponding visual feature maps. Then, the hybrid expert model uniformly divides the generated visual feature maps into L sub-regions in the spatial dimension, extracts effective information from each sub-region, and uses it as the corresponding local expert, thus obtaining the local experts E1,…,E for each of the L sub-regions. l ,…,E L Then, through an internal spatiotemporal gating network, local experts E1,…,E are provided with... l ,…,E L Different spatiotemporal weights are assigned to aggregate different local experts to obtain the rPPG signal corresponding to the sample. Finally, a set of positive sample rPPG signals is generated based on the rPPG signal corresponding to each positive sample, and a set of negative sample rPPG signals is generated based on the rPPG signal corresponding to each negative sample.

[0055] S140, Update the parameters of the rPPG signal extraction model based on the positive sample rPPG signal set and the negative sample rPPG signal set.

[0056] In some embodiments, a loss value can be calculated based on a set of positive rPPG signals, a set of negative rPPG signals, and a preset loss function, and the parameters of the rPPG signal extraction model can be updated accordingly.

[0057] It is worth noting that the loss functions here can include: frequency contrast loss function, frequency ratio consistency loss function, cross-video frequency consistency loss function, and video reconstruction loss function.

[0058] S150, repeat the above operation continuously until the training stop condition is met.

[0059] The training stopping condition can be either the number of iterations reaching a preset threshold or the loss value being less than a preset threshold; there are no restrictions on this.

[0060] The following is in conjunction with the appendix Figure 2-4 The rPPG signal extraction model training method 100 provided in this embodiment of the invention will be described in detail in the form of functional modules, as follows:

[0061] (1) Data enhancement module, including frequency enhancement submodule and spatial enhancement submodule.

[0062] (1.1) Spatial enhancement submodule: Different spatial enhancement mechanisms, such as image rotation (0°, 90°, 180°, 270°) and image flipping (horizontal, vertical), are used to transform the facial video slices frame by frame to obtain multiple positive samples with the same rPPG signal frequency as the facial video slices, and a positive sample set is generated based on this.

[0063] In each iteration of model training, two spatial augmentation mechanisms are randomly selected and applied to facial video slices to obtain positive samples with the same rPPG signal frequency as the facial video slices, and a set of positive samples is generated accordingly.

[0064] (1.2) Frequency Enhancement Submodule: This module is designed as a pyramid structure. It processes facial video slices through a three-dimensional convolutional layer (3D Conv), a three-dimensional residual block (3D RB), and a downsampling module to generate multi-scale visual features. At each scale, the frequency of the rPPG signal is modulated by a frequency modulation block (FMB). After frequency modulation, the spatial dimension is unified by an upsampling module. Finally, the facial video slices are reconstructed through a 3D convolutional layer to obtain negative samples.

[0065] Frequency Modulation Block (FMB): The input to the FMB is visual features s at a certain scale. a The ratio of the target frequency to r i The aim is to make s aThe rPPG signal frequency is from f a Modulated to r i ×f a Specifically, firstly, s is processed using 3D global average pooling (3D GAP) and 1D convolution to... a By compressing the spatial and channel dimensions, a coarse signal z is obtained, which can be considered an approximate expression of the rPPG signal. To change the frequency of z, r... i The signal is copied to dimension T and concatenated with z. A one-dimensional residual block (1D RB) with convolutional layers and ReLU nonlinear activation function is then designed to perform local frequency transformation on the signal z. Furthermore, a bidirectional LSTM (BiLSTM) layer is used to achieve global frequency transformation and frequency error cancellation of z from the complete time series. The frequency modulation vector output by the BiLSTM is shown. Each value in can be viewed as a multiplier of the component at the corresponding time in z, which is also equivalent to s. a The multiplier of the corresponding frame at that time. Then, m i Multiply by s a It is possible to reduce s at the original feature scale. a The frequency of the rPPG signal in the middle is from f a Modulated to r i ×f a This generates negative samples. It is worth noting that in order to obtain r i m is calculated from the cascaded features of z. i The required transformation is nonlinear, that is, the nonlinear transformation of the signal is achieved by using nonlinear activation functions such as ReLU, Sigmoid and Hyperbolic tangent function (Tanh) in 1D RB and BiLSTM layers. This is determined by the nature of signal frequency modulation.

[0066] Overall, for the generation of the negative sample set, the frequency enhancement submodule randomly samples k target frequencies within the range of (0.3, 0.8) ∪ (1.2, 1.7), with a proportion R = {r1, ..., r...}. k}, and use it as input to generate x a k negative samples The frequencies of the rPPG signals it contains are {r1f} a ,…,r k f a}

[0067] (2) rPPG signal extraction model, also known as regional expert aggregation module, includes feature coding backbone network and hybrid expert model.

[0068] (2.1) Feature encoding backbone network: The input facial video is feature encoded using a 3D ResNet residual network to generate a visual feature map.

[0069] (2.2) Hybrid Expert Model: The visual feature map is uniformly divided into L sub-regions in the spatial dimension, and effective information is extracted from each sub-region and defined as a local expert E. l Specifically, firstly, two 3D Reinforcing Blocks (RBs) are used to encode the cardiac features of each region. Simultaneously, a region attention block is embedded between the two 3D RBs to capture information from skin regions rich in cardiac features. Then, a 3D GAP layer and a 1D Conv layer are used to project the features onto a one-dimensional vector. This yields a set of local experts E1,…,E corresponding to L sub-regions. l ,…,E L Then, a spatiotemporal gating network was built. The results from different experts are then fused into the final estimated rPPG signal. The spatiotemporal gating network is explained below:

[0070] Given that some facial areas better express cardiac characteristics during systole, while others are more sensitive to cardiac impulses during diastole, therefore, any local expert E l Different weights should be assigned at different times. Therefore, a spatiotemporal gating network was designed. Can be local experts E1,…,E l ,…,E L Different spatiotemporal weights are assigned and feature fusion is achieved. It includes two 3D Conv layers, one 3D GAP layer, one 1D Conv layer, and one normalized exponential function (Softmax) layer. Output L weight vectors G1,…,G corresponding to L sub-regions. l ,…,G L And each vector has a dimension of 1×T, that is The Softmax calculation method involves normalizing the weights of different experts at each time step, i.e. Finally, each G l Its corresponding expert E l Element-wise multiplication and weighted summation are performed to obtain the final rPPG signal:

[0071]

[0072] The signal extraction module was applied to the positive sample set X respectively. p ={x p} and negative sample set X n ={x n Each sample in} generates a corresponding set of positive sample signals. and negative sample signal set These signals have different frequency distributions, which can be used as a basis for optimizing the model parameters through subsequent model optimization modules.

[0073] (3) Model optimization module, used to calculate the loss value based on the loss function and update the model parameters accordingly. The loss functions include: frequency comparison loss function, frequency ratio consistency loss function, cross-video frequency consistency loss function, and video reconstruction loss function.

[0074] (3.1) Frequency contrast loss function Training a model on unlabeled data based on a contrastive learning mechanism hinges on bringing the set of positive sample signals (Y) in the feature space closer together. p ={y p}), while pushing it away from the negative sample signal set (Y) in the feature space. n ={y n By comparing the frequency distributions of the two samples, the model is guided to focus on the commonalities and differences in frequency between the positive and negative sample sets, helping the rPPG signal extraction model to discover and encode periodic cardiac features in the video. Therefore, the frequency contrast loss is defined based on the InfoNCE loss function commonly used in contrastive learning mechanisms.

[0075]

[0076] Where d(·,·) is the mean square error (MSE) of the power spectral density (PSD) of the two signals. This distance function can be used to measure the difference in the frequency distribution of the signals, and τ is a hyperparameter constant.

[0077] (3.2) Frequency Proportion Consistency Loss Function This loss function constrains any pair of positive and negative sample signals (y p and y n The frequency ratio between these two values ​​should be consistent with the corresponding frequency ratio r of the input frequency enhancement submodule. This loss... The formula is:

[0078]

[0079] Where P(·) is the main frequency of the signal, which is calculated by performing a Fast Fourier Transform (FFT) on the given signal and selecting the frequency component with the maximum power. and Since they all come from positive samples, they should theoretically have the same frequency. Represents negative sample signal With positive sample signal The ratio of the main frequency, if this ratio equals r i This indicates that the frequency enhancement submodule successfully generated negative samples with the target rPPG signal frequency. Therefore, this loss function helps constrain the frequency enhancement submodule to achieve its frequency enhancement function and further improve the frequency enhancement effect.

[0080] (3.3) Cross-video frequency consistency loss function The loss function is designed based on prior knowledge that the frequency of the rPPG signal will not change rapidly in the short term. Considering the input video x... a The signal should have a similar frequency distribution to the signals of its nearest neighbor slices, therefore x can be... a The nearest neighbor slices were also calculated using the rPPG signal extraction model to obtain the nearest neighbor slice signal set Y. b ={y b}, and designed cross-video frequency consistency loss. To constrain the positive sample signal set Y p and Y b Frequency consistency between them:

[0081]

[0082] Where J is x a The total number of nearest neighbor slices, d(·,·) is the MSE of the PSDs of the two signals. This loss function uses contextual information from the facial video to correct signal estimation bias and enhance the periodicity of the estimated signal.

[0083] The total loss function is a linear combination of the above loss function and the video reconstruction loss function:

[0084]

[0085] Based on the rPPG signal extraction model training method provided in the embodiments of the present invention, the embodiments of the present invention also provide an rPPG signal extraction method, such as... Figure 5 As shown, the rPPG signal extraction method 500 may include the following steps:

[0086] S510, acquire the facial video to be processed.

[0087] S520 inputs facial video into the rPPG signal extraction model to obtain the corresponding rPPG signal.

[0088] The rPPG signal extraction model is obtained based on the rPPG signal extraction model training method described above.

[0089] Furthermore, the rPPG signal can be analyzed to obtain corresponding vital sign information, such as heart rate, respiratory rate, and heart rate variability.

[0090] According to embodiments of the present invention, at least the following technical effects are achieved:

[0091] By using a self-supervised learning framework, the accuracy and generalization ability of vital sign signals such as heart rate, respiratory rate, and heart rate variability have been improved, computational costs have been reduced, and new possibilities have been brought to the field of medical and health monitoring.

[0092] This invention can effectively estimate heart rate, respiratory rate, and heart rate variability from facial video without the need for synchronously recorded PPG signals, improving detection accuracy and making vital sign monitoring more convenient. At the same time, since this invention does not rely on synchronously recorded PPG signals, users do not need to wear any contact devices during the detection process, which protects user privacy to a certain extent and makes the monitoring process more natural and seamless.

[0093] It will help promote the development of telemedicine and personal health monitoring technologies, especially in the fields of home healthcare, emergency medical response and public health monitoring, where it has broad application prospects.

[0094] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0095] Figure 6 A structural diagram of an exemplary electronic device capable of implementing embodiments of the present invention is shown. Electronic device 600 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0096] like Figure 6As shown, the electronic device 600 may include a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the electronic device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0097] Multiple components in electronic device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of displays, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows electronic device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0098] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as method 100 or method 500. For example, in some embodiments, method 100 or method 500 may be implemented as a computer program product, including a computer program tangibly contained in a computer-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of method 100 or method 500 described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform method 100 or method 500 by any other suitable means (e.g., by means of firmware).

[0099] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.

[0100] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for training an rPPG signal extraction model, characterized in that, The method includes: Spatial enhancement of facial video slices includes: using at least one spatial enhancement mechanism among image rotation, image horizontal flipping, and image vertical flipping to transform the facial video slices frame by frame, obtaining multiple positive samples with the same rPPG signal frequency as the facial video slices, and generating a set of positive samples accordingly. Frequency enhancement of facial video slices includes: randomly sampling multiple different target frequency ratios within a preset frequency ratio range, frequency modulation of facial video slices according to multiple different target frequency ratios, obtaining multiple negative samples with different rPPG signal frequencies, and generating a negative sample set accordingly. The positive sample set and the negative sample set are input into the rPPG signal extraction model to obtain the positive sample rPPG signal set and the negative sample rPPG signal set. The parameters of the rPPG signal extraction model are updated based on the positive sample rPPG signal set and the negative sample rPPG signal set. Repeat the above steps until the training stops.

2. The method according to claim 1, characterized in that, The process involves frequency modulation of facial video slices based on multiple different target frequency ratios to obtain multiple negative samples with different rPPG signal frequencies, including: The facial video slices are processed by a 3D convolutional layer, a 3D residual module, and a downsampling module to generate multi-scale visual features. At each scale, the target frequency ratio and the visual features corresponding to the scale are input into a frequency modulation block to modulate the rPPG signal frequency. The spatial dimension is unified by an upsampling module. Finally, the facial video slices are reconstructed by a 3D convolutional layer to obtain the negative samples corresponding to the target frequency ratio.

3. The method according to claim 1, characterized in that, The rPPG signal extraction model includes: a feature coding backbone network and a hybrid expert model.

4. The method according to claim 3, characterized in that, The step of inputting the positive sample set and the negative sample set into the rPPG signal extraction model to obtain the positive sample rPPG signal set and the negative sample rPPG signal set includes: The positive sample set and the negative sample set are input into the rPPG signal extraction model; The feature encoding backbone network encodes the features of the input samples to generate corresponding visual feature maps. The hybrid expert model uniformly divides the generated visual feature map into spatial dimensions. Each sub-region is then analyzed, and relevant information is extracted from each sub-region, serving as the corresponding local expert to obtain... Each sub-region corresponds to a local expert. Through an internal spatiotemporal gating network, local experts Different spatiotemporal weights are assigned, and different local experts are aggregated in this way to obtain the rPPG signal corresponding to the sample; A set of positive sample rPPG signals is generated based on the rPPG signal corresponding to each positive sample, and a set of negative sample rPPG signals is generated based on the rPPG signal corresponding to each negative sample.

5. The method according to claim 1, characterized in that, The step of updating the parameters of the rPPG signal extraction model based on the positive sample rPPG signal set and the negative sample rPPG signal set includes: Based on the set of positive rPPG signals, the set of negative rPPG signals, and a preset loss function, a loss value is calculated, and the parameters of the rPPG signal extraction model are updated accordingly.

6. The method according to claim 5, characterized in that, The loss function includes: Frequency contrast loss function, frequency ratio consistency loss function, cross-video frequency consistency loss function, and video reconstruction loss function.

7. A method for extracting rPPG signals, characterized in that, The method includes: Acquire the facial video to be processed; The facial video is input into the rPPG signal extraction model to obtain the corresponding rPPG signal, wherein the rPPG signal extraction model is obtained based on the rPPG signal extraction model training method according to any one of claims 1-6.

8. The method according to claim 7, characterized in that, The method further includes: The rPPG signal was analyzed to obtain the corresponding vital signs information.

Citation Information

Patent Citations

  • Face living body detection method and system and medium

    CN114612965A

  • Multi-modal fusion learning method and system for sleep management

    CN117789931A