A training method, device, equipment and medium for RPPG signal generation model

By converting light-skinned face videos into dark-skinned videos and using rPPG generation model to generate rPPG signals, determining the contrast loss for model training, the problem of robustness and poor measurement effects of remote physiological measurements on people with different skin color is solved, achieving wider applicability and higher measurement accuracy.

CN117197870BActive Publication Date: 2025-05-06CHONGQING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311203161.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-18
Publication Date
2025-05-06
Estimated Expiration
2043-09-18

AI Technical Summary

Technical Problem

The existing remote physiological measurement technology has poor robustness and poor measurement effects on people with different skin tones, mainly because the public data set is biased towards light skin tones and cannot effectively expand to dark skin tones. The video compression problem affects pixel information and reduces measurement accuracy.

Method used

By obtaining light-skinned face videos, converting them into dark-skinned face videos using the partition migration model, combining the constructed rPPG generation model to generate rPPG signals, and determining the contrast loss based on the rPPG signal, the rPPG generation model is trained to improve the robustness and applicability of the model.

Benefits of technology

It effectively solves the problem that the measurement results are biased towards light skin color groups, expands the applicability of the model to multiple skin color, and improves the robustness and measurement accuracy of remote physiological measurements, especially when processing compressed videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117197870B_ABST
    Figure CN117197870B_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of remote physiological measurement, and specifically relates to a training method, device, equipment and medium for an rPPG signal generation model. The application obtains a light-skinned face video; inputs the light-skinned face video into a grid migration model to obtain a dark-skinned face video; generates an rPPG signal from the light-skinned face video and the dark-skinned face video through a constructed rPPG generation model; determines the contrast loss according to the rPPG signal; and trains the rPPG generation model according to the contrast loss. This solves the problem that most of the current public data sets are composed of light-skinned groups, and some data sets do not even have dark-skinned groups, which will cause the measurement results to be completely biased towards the light-skinned group and cannot be expanded to more different groups.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of remote physiological measurement, and specifically relates to a training method, device, equipment and medium for an rPPG signal generation model. Background Art

[0002] Remote physiological measurement is an important task in the field of computer vision. The goal of remote physiological measurement is to measure the physiological indicators of the human body, such as heart rate, heart rate variability and respiratory rate, based on face videos. Face videos contain pixels, and the vascular information of the human body can be obtained based on the two-color reflection model, and then these physiological information can be measured. At present, remote physiological measurement is mainly divided into: traditional methods, deep learning supervised methods and deep learning unsupervised methods. These different methods can be distinguished by different categories, different technologies, and whether or not to use label standards.

[0003] Among them, problems such as facial skin color and video compression will lead to poor robustness of remote physiological measurement and significant differences in the results on different populations. For example, most methods are currently trained on public data sets, and most of the current public data sets are composed of light-skinned groups. Some data sets do not even have dark-skinned groups, which will cause the measurement results to be completely biased towards light-skinned groups and cannot be expanded to more different groups. In addition, the compression problem of the video will seriously affect the pixel information in the video, thereby making the measurement effect worse. Summary of the invention

[0004] In response to the above technical problems, the present invention proposes a training method, device, equipment and medium for an rPPG signal generation model. The present application obtains a light-skinned face video; inputs the light-skinned face video into a grid migration model to obtain a dark-skinned face video; generates an rPPG signal from the light-skinned face video and the dark-skinned face video through a constructed rPPG generation model; determines the contrast loss according to the rPPG signal; and trains the rPPG generation model according to the contrast loss. This solves the problem that most of the current public data sets are composed of light-skinned groups, and some data sets do not even have dark-skinned groups, which will cause the measurement results to be completely biased towards the light-skinned group and cannot be expanded to more different groups.

[0005] In order to solve the above technical problems, the technical solution adopted by the present invention includes four aspects.

[0006] In a first aspect, a training method for an rPPG signal generation model is provided, characterized in that it includes: obtaining a light-skinned face video; inputting the light-skinned face video into a grid migration model to obtain a dark-skinned face video; generating an rPPG signal from the light-skinned face video and the dark-skinned face video through a constructed rPPG generation model; determining a contrast loss based on the rPPG signal; and training the rPPG generation model based on the contrast loss.

[0007] In some embodiments, inputting the light skin face video into the grid migration model to obtain the dark skin face video includes: decomposing the light skin face video frame by frame to obtain a light skin video frame set; inputting each frame of the light skin video frame set into the grid migration model to obtain a dark skin video frame set; and synthesizing each frame of the dark skin video frame set into the dark skin face video.

[0008] In some embodiments, generating rPPG signals from the light-skinned face video and the dark-skinned face video through the constructed rPPG generation model includes: generating rPPG blocks from the light-skinned face video and the dark-skinned face video through the rPPG generation model; and determining rPPG signals based on the rPPG blocks, wherein the rPPG block is a collection of multiple rPPG signals in time and space dimensions.

[0009] In some embodiments, determining the contrast loss based on the rPPG signal includes: converting the rPPG signal into a picture data set, wherein the picture data set includes a plurality of picture data, and each of the picture data represents a different position on the face; and determining the contrast loss based on the picture data set.

[0010] In some embodiments, determining the contrast loss based on the picture dataset includes: generating a positive loss function and a negative loss function based on the target picture dataset of the target rPPG block and the picture datasets of other rPPG blocks; and determining the contrast loss based on the positive loss function and the negative loss function.

[0011] In some embodiments, determining the rPPG signal according to the rPPG block includes: cyclically extracting rPPG sample blocks at any spatial position and with a preset duration from the rPPG block as the rPPG signal.

[0012] In some embodiments, generating a positive loss function and a negative loss function according to the picture data set of the plurality of rPPG blocks is implemented by the following formula:

[0013] Among them, LossPositive is a positive loss function, p i is the i-th image data in the target image dataset, p j is the jth image data in the target image dataset, and i is not equal to j, N is the number of image data in the target image dataset, p' i is the i-th image data in other image datasets, p' j is the jth image data in other image datasets, and p' i and p' j Belong to the same image dataset, Loss Negetive is a negative loss function.

[0014] In the second aspect, the present application proposes a training device for an rPPG signal generation model, comprising: a first acquisition module, used to acquire a light-skinned face video; a first execution module, used to input the light-skinned face video into a grid migration model to obtain a dark-skinned face video; a second execution module, used to generate an rPPG signal from the light-skinned face video and the dark-skinned face video through a constructed rPPG generation model; a first determination module, used to determine the contrast loss based on the rPPG signal; and a first training module, used to train the rPPG generation model based on the contrast loss.

[0015] In a third aspect, the present application proposes an electronic device, comprising: a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the method described in the first aspect is performed.

[0016] In a fourth aspect, the present application proposes a storage medium, which stores a computer program that can be executed by one or more processors, and the computer program can be used to implement the method described in the first aspect.

[0017] Beneficial effects of the invention: The invention obtains a light-skinned face video; inputs the light-skinned face video into a grid migration model to obtain a dark-skinned face video; generates an rPPG signal from the light-skinned face video and the dark-skinned face video through a constructed rPPG generation model; determines the contrast loss according to the rPPG signal; and trains the rPPG generation model according to the contrast loss. This solves the problem that most of the current public data sets are composed of light-skinned groups, and some data sets do not even have dark-skinned groups, which will cause the measurement results to be completely biased towards light-skinned groups and cannot be expanded to more different groups. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The scope of the present disclosure may be better understood by reading the following detailed description of exemplary embodiments in conjunction with the accompanying drawings. The drawings included are:

[0019] Figure 1 An overall flow chart of a training method for an rPPG signal generation model provided in an embodiment of the present application;

[0020] Figure 2 A structural block diagram of a training device for an rPPG signal generation model provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.

[0022] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0023] If similar descriptions of "first\second\third" appear in the application documents, the following instructions will be added. In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0025] Embodiment 1:

[0026] Currently, remote physiological measurements are mainly divided into: traditional methods, deep learning supervised methods and deep learning unsupervised methods. These different methods can be distinguished by different categories, different technologies, and whether or not to use label standards.

[0027] Problems such as facial skin color and video compression can lead to poor robustness of remote physiological measurements and significant differences in the results on different populations. For example, most methods are currently trained on public datasets, and most of the current public datasets are composed of light-skinned groups. Some datasets do not even have dark-skinned groups, which will cause the measurement results to be completely biased towards light-skinned groups and cannot be expanded to more different groups. In addition, the compression problem of the video will seriously affect the pixel information in the video, making the measurement effect worse.

[0028] In view of the problems existing in the prior art, such as Figure 1 As shown, the present application provides a training method for an rPPG signal generation model, and the method is applied to an electronic device, and the electronic device can be a server, a mobile terminal, a computer, a cloud platform, etc. The function implemented by the device data processing provided in the embodiment of the present application can be implemented by calling a program code by a processor of the electronic device, wherein the program code can be stored in a computer storage medium, and the training method for the rPPG signal generation model includes:

[0029] Step S1: Obtain a light-skinned face video.

[0030] Step S2: Input the light-skinned face video into the grid migration model to obtain the dark-skinned face video.

[0031] Since the existing remote physiological measurements of the face are all trained based on light-skinned people as the basic training data, the existing remote physiological measurements of the face, especially the extraction of rPPG signals, have a narrow scope of application and cannot be applied to dark-skinned people. Therefore, in order to make the generation of rPPG signals suitable for people of all skin colors, the existing light-skinned face videos are converted into dark-skinned face videos in this application.

[0032] In some embodiments, step S2 of “inputting the light-skinned face video into the grid migration model to obtain the dark-skinned face video” includes:

[0033] Step S21: Decompose the light-skin face video frame by frame to obtain a light-skin video frame set.

[0034] Step S22: inputting each frame of the light skin color video frame set into the grid migration model to obtain a dark skin color video frame set.

[0035] Step S23: synthesizing the frames of the dark skin color video frame set into the dark skin color face video.

[0036] Style transfer is a very hot and important topic in the field of image processing, and it plays an important role in synthetic videos and artistic creation. This technology aims to combine the style of one image with the content of another image to create a new image with unique visual effects. The basic idea of ​​style transfer is to adjust the pixel values ​​of the input image through an optimization algorithm so that it matches the reference image in both content and style. In the early research on style transfer, neural network-based methods gradually dominated. By using pre-trained deep neural networks, high-level feature representations of images can be captured, thereby achieving more accurate transfer of style and content. A key component of this method is the loss function, which combines content loss and style loss to guide the optimization process to generate synthetic images. Style transfer technology has been widely used in fields such as artistic creation, image generation, and image stylization, and style transfer has also shown great potential in video processing, image restoration, and visual enhancement.

[0037] CyleGAN is currently the most popular style transfer model. It can achieve good style transfer effects by training based on the principle of single-cycle consistency. At the same time, we are the first to apply style transfer to the field of remote physiological measurement. In the specific implementation, we first decompose the light-skinned face video into multiple PNG format images frame by frame in chronological order, and obtain the light-skinned video frame set. Since the environment in each original frame is relatively complex, including the influence of lighting changes, body movement, etc., we fine-tune the excellent model multi-task convolutional neural network in the field of face recognition and adjust it to 256×256, so that each image is cropped to 256×256 size as the input of CycleGAN. Finally, CycleGAN converts each image in the light-skinned video frame set into the corresponding dark-skinned image, and then obtains the dark-skinned video frame set, and restores the dark-skinned video frame set to the dark-skinned face video frame by frame in chronological order.

[0038] Step S3: generating rPPG signals from the light-skinned face video and the dark-skinned face video using the constructed rPPG generation model.

[0039] In order to make remote physiological measurement applicable to people with more skin colors, this application will use light-skinned face videos and dark-skinned face videos to generate rPPG signals through the constructed rPPG generation model.

[0040] In some embodiments, step S3 "generating rPPG signals from the light-skinned face video and the dark-skinned face video using the constructed rPPG generation model" includes:

[0041] Step S31: generating rPPG blocks from the light-skinned face video and the dark-skinned face video through the rPPG generation model.

[0042] Step S32: determining an rPPG signal according to the rPPG block, wherein the rPPG block is a collection of multiple rPPG signals in the spatiotemporal dimension.

[0043] The rPPG generation model constructed in this application includes: a convolution module, an average pooling module, an interpolation module and a maximum pooling module.

[0044] The convolution module passes the 128×128×3× frame face video through a 5×5 convolution, which doubles the number of channels to improve the spatial information expression ability of the rPPG generation module. The result of the convolution is then normalized and activated to further extract the receptive field features. After that, average pooling is performed through the average pooling module, and adaptive average pooling is used at the end of the network to downsample along the spatial dimension, so that the spatial dimension length of the output result can be controlled, so that the rPPG generation model can output a spatiotemporal rPPG block with a shape of T×S×S×1, where S is the length of the spatial dimension. The generated rPPG block is actually a collection of rPPG signals in the spatiotemporal dimension. For the rPPG block, we can use U to represent the rPPG block, U∈R T×S×S .

[0045] Therefore, in this application, if you want to obtain a remote physiological measurement model suitable for people with different skin colors, you need an rPPG generation model that can accurately obtain rPPG signals from people with different skin colors. In order to enable the rPPG generation model to generate accurate rPPG signals for people with different skin colors, it is necessary to train the rPPG generation network with light-skinned face videos and dark-skinned face videos at the same time in this application. During the training process of the application, the light-skinned face videos and dark-skinned face videos are mixed, randomly shuffled and divided, and divided into training sets and test sets in a ratio of 8:2, where videos corresponding to different faces cannot overlap in these sets. Then the rPPG generation network is trained and tested according to the training set and the test set.

[0046] In some embodiments, step S32 “determine the rPPG signal according to the rPPG block” comprises:

[0047] Step S321: cyclically extracting rPPG sample blocks at any spatial position and with a preset duration from the rPPG block as the rPPG signal.

[0048] Since the rPPG block is a collection of rPPG signals in the spatiotemporal dimension, during the training phase, we can obtain rPPG signals from the rPPG block by spatiotemporal sampling.

[0049] When obtaining rPPG signals from the rPPG block, we can get the rPPG signal U(·, x, y) at any spatial position. Further, for time sampling, we can sample a short time interval from U(·, x, y), so that the final spatiotemporal sampling is U(t→t+Δt, x, y), where x and y are spatial positions, t is the start time, and Δt is the length of the time interval. Therefore, for an rPPG block, we will loop over all spatial positions and sample K rPPG signal segments at a randomly selected start time t for each spatial position, where N is the set hyperparameter. Therefore, for an rPPG block, S×S×K, a total of N rPPG signal segments, can be obtained for contrastive learning.

[0050] After the rPPG generation model is trained, in the testing phase, the rPPG blocks can be directly averaged in the spatial dimension to obtain the rPPG signal.

[0051] Step S4: Determine contrast loss according to the rPPG signal.

[0052] During the generation model and model training process, it is necessary to use the loss function to adjust the parameters of the rPPG generation model. Therefore, the rPPG signal is used to determine the contrast loss in this application, and then the rPPG generation model is trained according to the contrast loss.

[0053] Therefore, in some embodiments, step S4 “determining contrast loss according to the rPPG signal” includes:

[0054] Step S41: converting the rPPG signal into an image data set, wherein the image data set includes a plurality of image data, and each of the image data represents a different position on the face.

[0055] Step S42: Determine the contrast loss according to the image data set.

[0056] For a video, we can get an rPPG block U, a set of rPPG samples [u 1 ,…,u N ] and the corresponding PSD[p 1 , ..., p N For other videos, we can obtain an rPPG block U′, a set of rPG samples [u′ 1 ,…,u' N ] and the corresponding PSD[p' 1 ,…,p' N]. Different from some previous unsupervised methods that directly use rPPG signals as samples for contrastive learning, we convert rPPG signals into PSD and use PSD as samples. This is because the conversion to PSD can use an important prior knowledge, that is, the human heart rate range is between 40 and 250bpm, so only PSD from 0.66hz to 4.16hz can be used to remove some redundant signals. Then use PSD samples to determine the contrast loss of the rPPG generation model.

[0057] In general, the training of most remote physiological measurement models is limited to supervised methods, and must rely on real labels that match face videos to train. However, there are currently few public datasets and a lack of large-scale datasets. Therefore, most methods only perform well on a single dataset and cannot be generalized to other datasets and real-world scenarios. Due to the requirements of contrastive learning unsupervised methods, contrastive learning methods require complex and diverse samples for contrastive learning. In the absence of real labels, the only thing contrastive learning methods rely on is face videos as samples, but current methods cannot obtain more samples, and the samples are severely biased. As a result, the current unsupervised methods have poor performance, lagging behind supervised methods, and are more difficult to apply in real-world scenarios.

[0058] Therefore, in some embodiments, step S42 of “determining the contrast loss according to the image dataset” includes:

[0059] Step S421: Generate a positive loss function and a negative loss function according to the target image data set of the target rPPG block and the image data sets of other rPPG blocks.

[0060] Step S422: Determine the contrast loss according to the positive loss function and the negative loss function.

[0061] The core of contrast loss is to pull the PSD of the same video together and push different PSDs apart, which will be specifically constrained by positive loss function and negative loss. Positive loss function: According to the important prior knowledge of rPPG signals, the rPPG signals of different regions of the same face are similar, and the PSDs converted from rPPG signals of similar time segments are similar. Therefore, we can draw an important conclusion that the PSDs converted from the rPPG segments obtained by spatiotemporal sampling from the same rPPG block should be similar. Negative loss function: According to the cross-video rPPG dissimilarity, the rPPG signals obtained from different face videos are different, and the corresponding converted PSDs are also different. Therefore, we can get that the corresponding converted PSDs of rPPG segments obtained by spatiotemporal sampling from different rPPG blocks should be different. Therefore, based on this principle, we aggregate the PSDs obtained from the same face video and push the PSDs of different faces away, thereby constraining the training process of contrastive learning and making the model training more balanced and stable. The final contrast loss is the sum of the positive loss function and the negative loss function. The overall contrast loss function encourages the model to search on supported frequencies to discover visual features of strong periodic signals. Therefore, in order to explain the expression well in this application, one of the rPPG blocks is set as the target rPPG block, and then the image dataset obtained according to the target rPPG block is the target dataset, and the contrast loss obtained is based on the contrast loss of the target rPPG block.

[0062] Therefore, in some embodiments, the present invention generates a positive loss function and a negative loss function according to the picture data set of multiple rPPG blocks through the following formula:

[0063]

[0064]

[0065] Among them, Loss Positive is a positive loss function, p i is the i-th image data in the target image dataset, p j is the jth image data in the target image dataset, and i is not equal to j, N is the number of image data in the target image dataset, p' i is the i-th image data in other image datasets, p' j is the jth image data in other image datasets, and p' i and p' j Belong to the same image dataset, Loss Negetive is a negative loss function.

[0066] Step S5: training the rPPG generation model according to the contrast loss.

[0067] The present application obtains the contrast loss of the target rPPG block through steps S4-S422, so that the rPPG generation model can be trained according to the contrast loss. Compared with the traditional remote physiological measurement method, the final rPPG generation model does not require manual provision of prior knowledge. Relying on the powerful learning ability of the neural network, the model has strong performance and more accurate measurement results in the physiological measurement task of face videos.

[0068] Compared with other deep learning remote physiological measurement methods, the present invention is reasonably designed and expanded based on 3D CNN, so that 3D CNN can generate rPPG blocks containing spatiotemporal information, and a large number of rPPG signal samples can be obtained from them by spatiotemporal sampling for comparative learning, breaking the limitation that previous deep models only generate single rPPG signals. At the same time, the method of combining style transfer and comparative learning is used to generate more diverse and rich samples for comparative learning, and then train 3D CNN, which enhances the robustness of the model for different populations and different data sets. Compared with other excellent remote physiological measurement methods, it has better performance and generalization ability; it also has good performance for challenges such as compressed video; and it significantly reduces the parameters of the model and reduces the possibility of overfitting, meeting the lightweight requirements.

[0069] In this application, the face videos in the current public data set are expanded by style transfer technology to generate corresponding dark-skinned group face videos, thereby generating a synthetic data set, which will be used together with the original data set for comparative learning. At the same time, a 3D CNN-based network is designed to generate an rPPG generation model containing spatiotemporal information rPPG blocks, so that a large number of rPPG samples can be obtained from a face video, enhancing the representation of rPPG information. At the same time, the contrast loss function is used to constrain the contrast learning process to train the overall robustness and measurement capabilities of the network. Among them, the encoder-decoder structure is applied to the rPPG block generation model part, which greatly reduces the training cost and solves the problem of not being able to obtain more rPPG information from the face video. Therefore, the more abundant samples generated by style transfer enable comparative learning from more people and more samples, solving the problems of poor robustness and low measurement accuracy of remote physiological measurements.

[0070] Therefore, the method of the present application solves the measurement difficulties caused by the high training cost, narrow scope of application, large differences in measurement results for different populations, and small and poor samples of unsupervised methods in most remote physiological measurement methods in the prior art.

[0071] Embodiment 2:

[0072] Based on the foregoing embodiments, the embodiments of the present application provide a training device for an rPPG signal generation model. The modules included in the device and the units included in the modules can be implemented by a processor in a computer device; of course, they can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU, Central Processing Unit), a microprocessor (MPU, Microprocessor Unit), a digital signal processor (DSP, Digital Signal Processing) or a field programmable gate array (FPGA, Field Programmable Gate Array), etc.

[0073] like Figure 2 As shown, a training device for an rPPG signal generation model includes: a first acquisition module 1, a first execution module 2, a second execution module 3, a first determination module 4 and a first training module 5.

[0074] The first acquisition module 1 is used to acquire a light-skinned face video. The first execution module 2 is used to input the light-skinned face video into a grid migration model to obtain a dark-skinned face video. The second execution module 3 is used to generate an rPPG signal from the light-skinned face video and the dark-skinned face video through a constructed rPPG generation model. The first determination module 4 is used to determine the contrast loss based on the rPPG signal. The first training module 5 is used to train the rPPG generation model based on the contrast loss.

[0075] In some embodiments, the first execution module 2 includes: a third execution module, a fourth execution module and a fifth execution module.

[0076] The third execution module is used to decompose the light skin face video frame by frame to obtain a light skin video frame set. The fourth execution module is used to input each frame of the light skin video frame set into the grid migration model to obtain a dark skin video frame set. The fifth execution module is used to synthesize each frame of the dark skin video frame set into the dark skin face video.

[0077] In some embodiments, the second execution module 3 includes: a sixth execution module and a second determination module.

[0078] The sixth execution module is used to generate rPPG blocks from the light-skinned face video and the dark-skinned face video through the rPPG generation model. The second determination module is used to determine the rPPG signal according to the rPPG block, wherein the rPPG block is a collection of multiple rPPG signals in the spatiotemporal dimension.

[0079] In some embodiments, the first determination module 4 includes: a seventh execution module and a third determination module.

[0080] The seventh execution module is used to convert the rPPG signal into an image data set, wherein the image data set includes a plurality of image data, and each of the image data has a different position on the face. The third determination module is used to determine the contrast loss according to the image data set.

[0081] In some embodiments, the third determination module includes: an eighth execution module and a fourth determination module.

[0082] The eighth execution module is used to generate a positive loss function and a negative loss function according to the target image data set of the target rPPG block and the image data sets of other rPPG blocks. The fourth determination module is used to determine the contrast loss according to the positive loss function and the negative loss function.

[0083] In some embodiments, the second determination module includes: a ninth execution module.

[0084] The ninth execution module is used to cyclically extract an rPPG sample block with any spatial position and a preset duration from the rPPG block as the rPPG signal.

[0085] Each module in the training device of the above-mentioned rPPG signal generation model can be implemented in whole or in part by software, hardware and a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the device in the form of hardware, or can be stored in the memory in the processing device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules. It should be noted that the division of modules in the embodiment of the present application is schematic, which is only a logical function division, and there may be other division methods in actual implementation.

[0086] Embodiment 3:

[0087] The third aspect provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of a training method for an rPPG signal generation model are implemented.

[0088] Embodiment 4:

[0089] The fourth aspect provides a storage medium, which stores a computer program that can be executed by one or more processors, and the computer program can be used to implement the steps of any one of the training methods for an rPPG signal generation model in the first aspect.

[0090] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0091] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned sequence numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.

[0092] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0093] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0094] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0095] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0096] A person of ordinary skill in the art can understand that: all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM, Read Only Memory), disks or optical disks, etc. Various media that can store program codes.

[0097] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can essentially or in other words, the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for a controller to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0098] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A training method for an rPPG signal generation model, characterized in that: include: Get videos of light-skinned faces; Inputting the light-skinned face video into the grid migration model to obtain a dark-skinned face video; Generate rPPG signals from the light-skinned face video and the dark-skinned face video through the constructed rPPG generation model; determining contrast loss based on the rPPG signal; Training the rPPG generation model according to the contrast loss; The generating rPPG signals from the light-skinned face video and the dark-skinned face video by using the constructed rPPG generation model includes: Generate rPPG blocks from the light-skinned face video and the dark-skinned face video through the rPPG generation model; Determine the rPPG signal according to the rPPG block, wherein the rPPG block is a collection of a plurality of the rPPG signals in a spatiotemporal dimension; The determining of contrast loss according to the rPPG signal comprises: Converting the rPPG signal into an image data set, wherein the image data set includes a plurality of image data, and each of the image data represents a different position on the face; Determining the contrast loss according to the image dataset; The determining the contrast loss according to the image data set includes: Generate a positive loss function and a negative loss function according to a target picture data set of a target rPPG block and the picture data sets of other rPPG blocks; The positive loss function and the negative loss function are generated according to the picture data set of the multiple rPPG blocks by the following formula: Among them, Loss Positive is a positive loss function, p i is the i-th image data in the target image dataset, p j is the jth image data in the target image dataset, and i is not equal to j, N is the number of image data in the target image dataset, p' i is the i-th image data in other image datasets, p' j is the jth image data in other image datasets, and p' i and p' j Belong to the same image dataset, Loss Negetive is a negative loss function; The contrast loss is determined according to the positive loss function and the negative loss function.

2. The method according to claim 1, characterized in that The step of inputting the light-skinned face video into the grid migration model to obtain the dark-skinned face video comprises: Decomposing the light-skinned face video frame by frame to obtain a light-skinned video frame set; Inputting each frame of the light skin video frame set into the grid migration model to obtain a dark skin video frame set; The frames of the dark skin color video frame set are synthesized into the dark skin color face video.

3. The method according to claim 1, characterized in that The determining of the rPPG signal according to the rPPG block comprises: An rPPG sample block at any spatial position and with a preset duration is cyclically intercepted from the rPPG block as the rPPG signal.

4. A training device for an rPPG signal generation model, characterized in that: include: The first acquisition module is used to acquire light-skinned face videos; A first execution module is used to input the light-skinned face video into a grid migration model to obtain a dark-skinned face video; A second execution module is used to generate rPPG signals from the light-skinned face video and the dark-skinned face video through the constructed rPPG generation model; A first determination module, configured to determine contrast loss according to the rPPG signal; The first determination module includes: a seventh execution module and a third determination module; The seventh execution module is used to convert the rPPG signal into an image data set, wherein the image data set includes a plurality of image data, and each of the image data represents a different position on the face; The third determination module is used to determine the contrast loss according to the image data set; The third determination module includes: an eighth execution module and a fourth determination module; An eighth execution module is used to generate a positive loss function and a negative loss function according to a target picture data set of a target rPPG block and the picture data sets of other rPPG blocks; The eighth execution module generates a positive loss function and a negative loss function according to the picture data set of the plurality of rPPG blocks through the following formula: Among them, Loss Positive is a positive loss function, p i is the i-th image data in the target image dataset, p j is the jth image data in the target image dataset, and i is not equal to j, N is the number of image data in the target image dataset, p' i is the i-th image data in other image datasets, p' j is the jth image data in other image datasets, and p' i and p' j Belong to the same image dataset, Loss Negetive is a negative loss function; The fourth determination module is used to determine the contrast loss according to the positive loss function and the negative loss function; The first training module is used to train the rPPG generation model according to the contrast loss.

5. An electronic device, characterized in that: include: A memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the method according to any one of claims 1 to 3 is executed.

6. A storage medium, characterized in that: The computer program stored in the storage medium can be executed by one or more processors, and the computer program can be used to implement the method according to any one of claims 1 to 3.