Systems and methods for training remote photoplethysmography models
A self-supervised contrastive learning approach trains remote PPG models using unlabeled video clips, addressing the challenge of data availability in existing systems and enhancing the accuracy and efficiency of heart rate estimation.
Patent Information
- Application Number
- JP2022037239
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-05-11
- Filing Date
- 2022-03-10
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2042-03-10
AI Technical Summary
Existing remote photoplethysmography (PPG) systems require highly specialized separation algorithms or supervised deep learning approaches, necessitating labeled data that are difficult to obtain, limiting their effectiveness in training models for estimating heart rate from unlabeled video.
A self-supervised contrastive learning method is employed to train a remote PPG model using unlabeled video clips of human faces, utilizing a series of neural network architectures and loss functions to adjust model weights, allowing the model to output heart rate signals without requiring labeled data.
The method enables accurate estimation of heart rate and saliency signals from unlabeled video, reducing the need for expensive labeled data and improving the efficiency and accessibility of remote PPG systems.
Smart Images

Figure 0007679786000019 
Figure 0007679786000020 
Figure 0007679786000021
Abstract
Description
[Technical field]
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 161,799, filed March 16, 2021, entitled "SYSTEM AND METHOD FOR LEARNING TO ESTIMATE HEART RATE FROM UNLABELED VIDEO," the entire contents of which are incorporated herein by reference.
[0002] The subject matter described herein relates generally to systems and methods for training video information models, such as remote photoplethysmography ("PPG") models. [Background technology]
[0003] The background description provided generally presents the context of the disclosure. The inventors' work to the extent that it can be described in this background, and aspects of the description that may not qualify as prior art at the time of filing, are not admitted, expressly or impliedly, to be prior art against the present technology.
[0004] There are several different methods for measuring the physiological state of a human. For example, electrocardiography involves placing electrodes on the human skin. These electrodes detect small electrical changes caused by myocardial polarization, followed by repolarization during each heartbeat.
[0005] Conventional PPG is an optical technique used to detect blood volumetric changes in peripheral blood circulation. These conventional PPG systems use low-intensity infrared light that propagates through physiological tissue and is absorbed by bone, skin pigment, and venous and arterial blood. Because light is absorbed more strongly by blood than by surrounding tissue, changes in blood flow can be detected by changes in light intensity by a PPG sensor placed on the subject's skin.
[0006] Recently, systems have been developed to perform remote PPG. Unlike ECG and traditional PPG systems, these systems do not require a light source and / or sensor to be placed on the subject's skin. Using standard imaging devices such as cameras, remote PPG systems can remotely measure changes in light transmitted or reflected from the body due to volumetric changes in blood flow. To properly process the imaging information, remote PPG systems must either use highly specialized separation algorithms or utilize models trained with supervised deep learning approaches. Summary of the Invention
[0007] This section provides a general summary of the disclosure but is not an exhaustive description of its entire scope or all of its features.
[0008] In one embodiment, a system for training a remote PPG model to output a PPG signal of a subject based on a subject video clip of the subject includes a processor and a memory in communication with the processor. The memory can include a training module having instructions that, when executed by the processor, cause the processor to train the remote PPG model in a self-supervised contrastive learning manner using unlabeled video clips having a sequence of images of a human face.
[0009] In another embodiment, a method for training a remote PPG model to output a PPG signal of a subject based on a subject video clip of the subject includes training the remote PPG model in a self-supervised contrastive learning method using unlabeled video clips having a sequence of images of a human face.
[0010] In yet another embodiment, a non-transitory computer-readable medium includes instructions that, when executed by a processor, cause the process to train a remote PPG model in a self-supervised contrastive learning manner using unlabeled video clips having a sequence of images of a human face. As before, the remote PPG model is configured to output a subject PPG signal based on the subject video clips of the subject.
[0011] Further areas of applicability and various ways of enhancing the disclosed technology will become apparent from the description provided. The description and specific examples in this summary are intended for purposes of illustration only and are not intended to limit the scope of the present disclosure. [Brief description of the drawings]
[0012] The accompanying drawings, which are incorporated in and form a part of the specification, illustrate various systems, methods, and other embodiments of the disclosure. It will be appreciated that the boundaries of the illustrated elements in the drawings (e.g., boxes, groups of boxes, or other shapes) represent one embodiment of the boundaries. In some embodiments, one element may be designed as multiple elements, or multiple elements may be designed as one element. In some embodiments, an element shown as an internal component of another element may be realized as an external component, and vice versa. Additionally, elements may not be drawn to scale.
[0013] [Figure 1] FIG. 1 illustrates an example remote PPG system that can determine a subject's heart rate from a video clip and, optionally, a saliency signal indicating the portion of the video clip where the PPG signal from the subject is strongest. [Diagram 2] FIG. 1 illustrates a remote PPG training system for training a model utilized by a remote PPG system. [Diagram 3]FIG. 1 illustrates a process flow showing how the remote PPG training system trains a model. [Figure 4] FIG. 1 illustrates a method for training a model for use by a remote PPG system. [Diagram 5] FIG. 1 illustrates a method for determining losses using self-supervised contrastive learning to train a model utilized by a remote PPG system. [Figure 6] FIG. 1 illustrates an optional method for using supervised learning to fine-tune a model utilized by a remote PPG system. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0014] A system and method for training a model to be used by a remote PPG system is described. The system and method trains the model in a self-supervised contrastive learning manner. This has an advantage over other systems in that it can use unlabeled video clips to train the model. Current systems use self-supervised deep learning techniques that require labeled video clips that are not available or are difficult to develop.
[0015] To better understand how the system and method train a model for a remote PPG system, a brief description of one example of a remote PPG system is provided. Referring to FIG. 1, a remote PPG system 10 is illustrated receiving a video clip 12 including an image of a face 11 of a subject 13. The video clip 12 can be a sequenced video clip captured by an imaging device such as a camera. The camera utilized to capture the video clip 12 can be any type of camera capable of capturing an image, such as a conventional camera, a high definition camera, an infrared camera, etc. The video clip 12 can be a stored video clip provided to the remote PPG system 10, or a real-time video clip 12 captured in real-time and provided to the remote PPG system 10.
[0016] Upon receiving the video clip 12, the remote PPG system 10 can utilize information extracted from the video clip 12 and output a heart rate signal 16. Optionally, the remote PPG system 10 can also output a saliency signal 18 indicating where in the video clip 12 the PPG signal is strongest. For example, the saliency signal 18 indicates that the remote PPG signal is strongest at the forehead 15 of the subject 11. Information from the saliency signal can be utilized to determine how well the remote PPG system 10 is performing and / or if there is any problem with the video clip 12 provided to the remote PPG system 10. For example, if the saliency signal 18 indicates a very weak PPG signal, the subject 11 can be repositioned with respect to the camera to improve the strength of the saliency signal 18, resulting in a more accurate heart rate signal 16 predicted by the remote PPG system 10.
[0017] The remote PPG system 10 can take any of a number of different forms. In this example, the remote PPG system 10 includes one or more processors 20. The processor 20 may be a single processor or may be multiple processors operating in coordination. Furthermore, the processor 20 may be located within the remote PPG system 10 or may be located remotely from the remote PPG system 10. For example, the processor 20 may be located partially or entirely in a separate external device, such as a cloud computing device, or in a remote device.
[0018] The processor 20 can be in communication with a memory 30 that includes a PPG module 32. The PPG module 32 includes instructions that cause the processor to interpret the video clip 12 and output the heart rate signal 16 and / or the saliency signal 18. Additionally, it is shown that the remote PPG system 10 can also include a data store 40 that can be combined with the memory 30 or that can be separate and apart from the memory 30. Here, the data store 40 includes a PPG model 42 that includes one or more model weights 44.
[0019] The PPG model 42 can be any one of a number of different neural networks that can interpret the video clip 12, process the video clip 12, and output the heart rate signal 16 and / or the saliency signal 18. The model weights 44 can be one or more model weights that, based on their adjustment, affect how the PPG module 42 interprets the video clip 12 and outputs the heart rate signal 16 and the saliency signal 18. In general, the model weights 44 are parameters of the PPG model 42 and are used in layers of the PPG model 42. As such, adjustments to the model weights 44 can result in an optimization of the PPG model 42 such that the PPG model 42 outputs a more accurate heart rate signal 16 and / or saliency signal 18.
[0020] The PPG model 42 can take any one of a number of different forms and can use any one of a number of different models having different layers arranged in a number of different ways. In one example, the encoder and decoder of the PPG model 42 can be detailed as shown in the diagram below. [Table 1] [Table 2]
[0021] The core of the above PPG model 42 includes a series of eight 3D convolutions with kernel of (3, 3, 3), 64 channels, and ELU activation. This allows the PPG model 42 to learn spatio-temporal features for the input video clip. Average pooling and batch normalization are also adopted between layers. The PPG model 42 also adopts upsampling interpolation (x4) and 3D convolution with kernel (3, 1, 1). This upsampling step is repeated twice to eliminate aliasing in the output. Then, the PPG model 42 performs adaptive average pooling to collapse the spatial dimension and generate a one-dimensional signal. A final one-dimensional convolution is applied to convert the 64 channels and output a single-channel PPG.
[0022] The PPG model 42 that determines the saliency signal 18 can follow a Resnet-18 structure with pre-trained ImageNetweights truncated after layer 1. Each BasicBlock consists of a 3x3 convolution, a 2D batch normalization, a ReLU, a 3x3 convolution, a 2D batch normalization, an addition at the BasicBlock input (permanent), and a final ReLU. The diagram below illustrates the architecture of this network. [Table 3]
[0023] Here again, it should be understood that the PPG model 42 can take any one of a number of different forms, and not necessarily those specifically illustrated and described above.
[0024] As mentioned above, the PPG model 42 is trained with a self-supervised contrastive learning method using unlabeled video clips. Figure 2 shows one example of a remote PPG training system 100 for training the PPG model 42. The remote PPG training system 100 can be incorporated within the remote PPG system 10 of Figure 1 or can be separate, as shown.
[0025] As shown, the remote PPG training system 100 includes one or more processors 110. As such, the processor 110 can be part of the remote PPG training system 100 or the remote PPG training system 100 can access the processor 110 through a data bus or other communication path. In one or more embodiments, the processor 110 is an application specific integrated circuit configured to implement the functions associated with the training module 122. Generally, the processor 110 is an electronic processor, such as a microprocessor, that can perform various functions as described herein. In one embodiment, the remote PPG training system 100 includes a memory 120 that stores the training module 122. The memory 120 can be a random access memory ("RAM"), a read only memory ("ROM"), a hard disk drive, a flash memory, or other suitable memory for storing the training module 122. Training module 122 is, for example, computer readable instructions that, when executed by processor 110 , cause processor 110 to perform various functions disclosed herein, i.e., train PPG model 42 .
[0026] Additionally, in one embodiment, the remote PPG training system 100 includes a data storage device 130, which in one embodiment is an electronic data structure, such as a database, stored in the memory 120 or other memory and comprised of routines executable by the processor 110 to analyze, present, and organize the stored data. Thus, in one embodiment, the data storage device 130 stores data used by the training module 122 in performing various functions. In one embodiment, the data storage device 130 includes the PPG model 42 to be trained.
[0027] As explained above, the PPG model 42 can include one or more model weights 44 that are adjusted when the PPG model 42 is trained by the remote PPG training system 100. The data storage 130 can also include unlabeled video clips 132 used to train the PPG model 42. The unlabeled video clips 132 can be a single video clip, or multiple video clips. As explained later in this description, the PPG model 42 can be trained in a self-supervised contrastive learning method that utilizes unlabeled data in the form of the unlabeled video clips 132. The unlabeled video clips 132 can be similar to the video clips 12 shown in FIG. 1, where the unlabeled video clips 132 include a subject's face.
[0028] With respect to the training module 122, the training module 122 includes instructions that cause the processor 110 to train the PPG model 42 using the unlabeled video clips 132. To better illustrate how this occurs, assume that the PPG model 42 trains the unlabeled video clips 132 (x a 3, which illustrates a process flow 200 showing how the PPG training system 100 is trained using the labelled video clips 132 (xa ) a bounding box around the subject's face can be estimated. The processor 110 can add an additional 50% scale buffer to the bounding box before extracting the frame, which may be a 192x128 frame. Of course, it should be understood that the frame scaling and size may vary depending on the application. The bounding box position can be updated in subsequent frames of the unlabeled video clip 132 if the outer new unbuffered bounding box is a larger bounding box. The bounding box can be cropped and reduced to 64x64 resolution.
[0029] In one example, the training module 122 may provide the processor 110 with a warped, unlabeled video clip 204 (x s a ), we use unlabeled video clip 132 (x a ) is passed through the saliency sampler 202 to obtain the unlabeled video clip 132 (x a ) contains instructions to warp the unlabeled video clip 132 (x a It should be understood that passing the saliency sampler 202 may be an optional step that does not necessarily have to occur. The use of the saliency sampler 202 provides transparency as to what the PPG model 42 is learning and can have the advantage of warping the input images to spatially enhance task salient regions before feeding the input images into the PPG model 42.
[0030] In one example, the saliency sampler 202 can use a pre-trained ResNet-18 truncated after the convolution 2x block. Optionally, two additional loss terms are added:
number
number
number
[0031] Here again, as mentioned above, the use of the saliency sampler 202 is optional and need not be utilized. Therefore, the warped unlabeled video clip 204 (x s a ) would simply be an unlabeled video clip 132 (x a ), and the training module 122 may provide the processor 110 with the anchor PPG signal 230 (y a ) to generate the warped unlabeled video clip 204 (x s a ) (If the saliency sampler 202 is not utilized, then simply an unlabeled video clip 132 (x a ) can be passed through the PPG model 42.
[0032] The training module 122 transmits to the processor 110 the negative sample clips 206 (X s n ) to generate a first frequency ratio 210 (R f ) to generate a warped unlabeled video clip 204 (x s a) can be resampled. Again, if the saliency sampler 202 is not utilized, the warped unlabeled video clip 204 (x s a ) instead of resampling the unlabeled video clip 132 (x a ) resampling can be performed. The first frequency ratio is 210 (R f ) may be between 66% and 80%. The training module 122 also transmits to the processor 110 a negative PPG signal 232 (y n ), negative sample clip 206 (X s n ) can be passed through the PPG model 42.
[0033] Negative PPG signal 232(y n ), the training module 122 may then provide the processor 110 with a positive PPG signal 234 (y p ) to generate a second frequency ratio 233 (R f -1 ), a trilinear resampler 208 is used to generate a negative PPG signal 232 (y n ) may be performed. The second frequency ratio may be the inverse of the first frequency. The training module 122 may cause the processor 110 to adjust the model weights 44 of the PPG model 42 based on the calculated first loss. The first calculated loss may be a resampling of the anchor PPG signal 230 (y a ), negative PPG signal 232(y n ), and positive PPG signal 234 (y p ) is determined using a first loss function that takes into account
[0034] As explained above, the PPG model 42 is trained in a self-supervised contrastive learning method using a self-supervised triplet loss function 240. In this example, the training module 122 instructs the processor 110 to receive the anchor PPG signal 230 (y a ) based on the anchor subset view 242(f a), negative PPG signal 232(y n ) based on the negative subset view 246(f n ), and a positive PPG signal 234 (y p ) based on the positive subset view 244(f p ) is generated. a ), negative subset view 246(f n ), and the positive subset view 244(f p ) may have one or more views that correspond to each other. a ), negative PPG signal 232(y n ), and positive PPG signal 234 (y p ), the training module 122 informs the processor 110 of the length V L V N This forces the assumption that the heart rate is relatively stable within a window, so the signal in each view should look the same.
[0035] The training module 122 can cause the processor 110 to calculate the distance between all the combinations. In this example, the training module 122 can cause the processor 110 to calculate the distance between the corresponding anchor subset views 242 (f a ) and positive subset view 244(f p ) based on the distance between the positive total (P tot ) and determine the corresponding anchor subset view 242 (f a ) and negative subset view 246(f n ) based on the distance between the negative totals (N tot ) to generate the loss. tot ) and negative total (N tot ) is determined. Optionally, the loss is calculated based on the total number of views (V 2 N) to zoom in / out.
[0036] The training module 122 performs contrastive training on the multi-view triplet loss to obtain the total positives (P tot ) and negative total (N tot ) signals can use the power spectral density mean square error as a distance metric between the signals. The training module 122 can have the processor 110 calculate the power spectral density for each signal and remove all frequencies outside the appropriate heart rate range of 40 to 250 beats per minute ("bpm"). The training module 122 can have the processor 110 normalize each to have a sum value of 1 and calculate the mean square error between them. Finally, the training module 122 can have the processor 110 remove the inappropriate power ratios as a validation metric during contrast training. The training module 122 can have the processor 110 first calculate the power spectral density and divide it into appropriate frequencies (40-250 bpm) and inappropriate frequencies. The training module 122 can have the processor 110 divide the power in the inappropriate region by the total power. The inappropriate power ratios can be used as an unsupervised measure of signal quality.
[0037] Using the loss, the training module 122 can cause the processor 110 to adjust the model weights 44 of the PPG model 42. Thus, the process flow 200 illustrates a method for training the PPG model 42 using unlabeled video clips 132 in a self-supervised contrastive learning method. The use of unlabeled training clips is advantageous in that it does not require expensive, hard-to-obtain interpreted data to train the PPG model 42. Thus, the PPG model 42 can be trained with only unlabeled training clips.
[0038] However, in addition to training the PPG model 42 with unlabeled training clips, it is also possible to fine-tune the PPG model 42 by additionally training the PPG model 42 in a supervised manner. For example, the process flow 200 may also include training the ground truth PPG signal 254
number
number
[0039] The Pearson correlation loss function can be used as the basic supervised loss and validation metric. It is scale invariant but assumes perfect time synchronization between the ground truth and the observed data. Otherwise, the PPG model42 must be able to learn the time offset, assuming the offset is constant.
[0040] The signal-to-noise ratio loss function is calculated by using the ground truth PPG signal instead of the full PPG signal.
number
[0041] The maximum cross-correlation ("MCC") loss can be calculated in the frequency domain as follows:
number
[0042] Additionally, the training module 122 may provide the processor 110 with an anchor PPG signal 230 (y a ) and ground truth PPG signal 254
number
number
number
[0043] The training module 122 can cause the processor 110 to filter the cross-correlation to remove cross-correlation outside the range of expected heart rate (40-250 bpm) to generate a filtered spectrum of the cross-correlation. The training module 122 can cause the processor 110 to take the inverse FFT of the filtered spectrum of the cross-correlation to filter the anchor PPG signal 230 (y a ) and ground truth PPG signal 254
number
[0044] Referring to Figure 4, a method 300 for training a remote PPG system is shown. The method 300 is described in terms of the remote PPG training system 100 of Figure 2 and the process flow 200 of Figure 2. However, it should be understood that this is just one example for implementing the method 300. Although the method 300 is discussed in combination with the remote PPG training system 100, it should be appreciated that the method 300 is not limited to being implemented within the remote PPG training system 100, but rather within one example of a system in which the method 300 may be implemented.
[0045] In step 302, the training module 122 instructs the processor 110 to generate a warped, unlabeled video clip 204 (x s a ), we use unlabeled video clip 132 (x a ) through the saliency sampler 202 to obtain the unlabeled video clip 132 (x a ) in the 3D space. As explained above, this step is optional.
[0046] In step 304, the training module 122 instructs the processor 110 to generate a negative sample clip (x s n ) to generate a first frequency ratio 210 (R f ) to generate a warped unlabeled video clip 204 (x s a ) or unlabeled video clip 132(x a ) resampling can be performed.
[0047] In step 306, the training module 122 instructs the processor 110 to receive the anchor PPG signal 230 (y a ) to generate the warped unlabeled video clip 204 (x s a ) or unlabeled video clip 132(x a ) can be passed through the PPG model 42. Similarly, in step 308, the training module 122 also causes the processor 110 to n ), negative sample clip 206 (x s n ) can be passed through the PPG model 42.
[0048] In step 310, the training module 122 instructs the processor 110 to receive the positive PPG signal 234 (y p ) to generate a second frequency ratio 233 (Rf -1 ) to obtain a negative PPG signal 232 (y n ) may be performed. As explained above, the second frequency ratio may be the inverse of the first frequency ratio.
[0049] In step 312, the training module 122 can cause the processor 110 to determine a first loss using the first loss function. The first loss is then used to adjust the model weights 44 of the PPG model 42, as shown in step 314. The steps for determining the first loss are shown in FIG. 5 and described below.
[0050] In step 316, which may be optional, the training module 122 may have the processor 110 determine a second loss using a second loss function. Typically, the second loss is used to fine-tune the PPG model 42 and may utilize labeled clips. Steps for determining the second loss are shown in FIG. 6 and described below. In step 318, which may also be optional, the training module 122 may have the processor 110 adjust the model weights 44 of the PPG model 42 based on the second loss. The method 300 may then end or may return to step 302.
[0051] With respect to determining the first loss in step 312, reference is now made to FIG. 5, which illustrates steps for determining the first loss. In step 312A, the training module 122 instructs the processor 110 to receive the anchor PPG signal 230 (y a ) based on the anchor subset view 242(f a ), negative PPG signal 232(y n ) based on the negative subset view (f n ) 246, and the positive PPG signal 234 (y p ) based on the positive subset view 244(f pAs explained above, the anchor subset view 242 (f a ), negative subset view (f n ) 246, and the positive subset view 244(f p ) can have one or more views that correspond to each other.
[0052] In step 312B, the training module 122 instructs the processor 110 to generate a corresponding anchor subset view 242 (f a ) and positive subset view 244(f p ) based on the distance between the positive total (P tot Similarly, in step 312C, the training module 122 can cause the processor 110 to determine the corresponding anchor subset views 242 (f a ) and the negative subset view (f n )246 based on the distance between the negative total (N tot ) can be determined.
[0053] In step 312D, the training module 122 instructs the processor 110 to calculate the positive total (P tot ) and negative total (N tot ) can be determined. The first loss is the difference between the views (V 2 N ) can be scaled by the total number of
[0054] With respect to determining the second loss in step 316, reference is now made to FIG. 6, which illustrates the steps of determining the second loss. In step 316A, the training module 122 instructs the processor 110 to receive the anchor PPG signal 230 (y a ) and ground truth PPG signal 254
number
number
[0055] In step 316B, the training module 122 instructs the processor 110 to receive the anchor PPG signal 230 (y a ) and ground truth PPG signal 254
number
[0056] In step 316, the training module 122 can cause the processor 110 to filter the cross-correlation to remove cross-correlation outside the range of expected heart rate (40-250 bpm) to generate a filtered spectrum of the cross-correlation. In step 316, the training module 122 can cause the processor 110 to take the inverse FFT of the filtered spectrum of the cross-correlation to filter the anchor PPG signal 230 (y a ) and ground truth PPG signal 254
number
[0057] Detailed embodiments are disclosed herein. However, the disclosed embodiments are intended to be examples only. Therefore, the specific structural and functional details disclosed herein should not be construed as limiting, but are intended only as a basis for the claims and as a representative basis to teach one skilled in the art to variously employ the aspects herein in substantially any suitable detailed structure. Furthermore, the terms and phrases used herein are not intended to be limiting, but rather to provide an understandable description of possible implementations. Although various embodiments are shown in Figures 1-6, the embodiments are not limited to the illustrated structures or applications.
[0058] According to various embodiments, the flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of code, comprising one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions shown in the blocks may occur in a different order than that shown in the figures. For example, two blocks shown in succession may be executed substantially simultaneously, or the blocks may be executed in the reverse order, depending on the functions involved.
[0059] The above-mentioned systems, components, and / or processes can be realized in hardware or a combination of hardware and software, either centralized in one processing system or distributed where different elements are spread across several interconnected processing systems. Any kind of processing system or other device adapted to perform the methods described herein is suitable. A typical combination of hardware and software can be a processing system having a computer usable program that, when loaded and executed, controls the processing system such that the processing system performs the methods described herein. The systems, components, and / or processors can also be embedded in a computer readable storage device, such as a computer program product or other data program storage device that is readable by a machine and tangibly contains a program of instructions executable by the machine to perform the methods and processes described herein. These elements can also be embedded in an application product that comprises all features enabling the implementation of the methods described herein and that, when loaded into a processing system, can perform these methods.
[0060] Additionally, the arrangements described herein may take the form of a computer program product embodied in one or more computer readable media that contain, e.g., store, computer readable program code. Any combination of one or more computer readable media may be utilized. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. The phrase "computer readable storage medium" refers to a non-transitory storage medium. The computer readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, device, or any suitable combination of the foregoing. More specific examples of computer readable storage media may include the following (a non-exhaustive list): portable computer diskette, hard disk drive ("HDD"), solid state drive ("SSD"), ROM, erasable programmable read only memory (EPROM or flash memory), portable compact disk read only memory ("CD-ROM"), digital versatile disk (DVD), optical storage device, magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0061] Generally, a module as used herein includes routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular data types. In a further aspect, a memory generally stores the above-mentioned modules. The memory associated with a module may be a buffer or cache contained within a processor, RAM, ROM, flash memory, or other suitable electronic storage medium. In a further aspect, a module as contemplated by the present disclosure is implemented as an application specific integrated circuit ("ASIC") that is a hardware component of a system on a chip ("SoC"), as a programmable logic array ("PLA"), or as other suitable hardware component that includes a defined configuration set (e.g., instructions) to perform the disclosed functions.
[0062] The program code contained on the computer readable medium may be transmitted using any suitable medium, including, but not limited to, wireless, wireline, fiber optic, cable, RF, etc., or any suitable combination of the above. Computer programs for performing operations for aspects of the present arrangement may be implemented using Java. TMThe program code may be written in any combination of one or more programming languages, including object-oriented programming languages such as, for example, Smalltalk, C++, and the like, and traditional procedural programming languages such as the "C" programming language or similar programming languages. The program code may run entirely on the user's computer, partially on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server, as a stand-alone software package. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network ("LAN") or a wide area network ("WAN"), or may be connected to an external computer (e.g., through the Internet using an Internet Service Provider).
[0063] As used herein, the term "a" is defined as one or more than one. As used herein, the term "multiple" is defined as two or more than two. As used herein, the term "another" is defined as at least a second or more. As used herein, the terms "including" and / or "having" are defined as comprising (i.e., open language). As used herein, the phrase "and at least one of" refers to and is inclusive of any and all possible combinations of one or more of the associated listed items. By way of example, the phrase "at least one of A, B, and C" includes only A, only B, only C, or any combination thereof (e.g., AB, AC, BC, and ABC).
[0064] The aspects herein may be embodied in other forms without departing from the spirit or essential attributes thereof, and reference should accordingly be made to the following claims, rather than the foregoing specification, as indicating their scope. The invention disclosed in this specification includes the following aspects. [Aspect 1] 1. A system for training a remote photoplethysmography ("PPG") model that outputs a subject PPG signal based on a subject video clip of a subject, the system comprising: A processor; a memory in communication with the processor, the memory having a training module having instructions that, when executed by the processor, cause the processor to train a remote PPG model in a self-supervised contrastive learning manner using an unlabeled video clip having a sequence of images of a human face; system. [Aspect 2] 2. The system of claim 1, wherein the remote PPG model is configured to output a saliency signal indicative of a portion of the subject video clip where the PPG signal is strongest based on the subject video clip. [Aspect 3] The training module may include instructions that, when executed by the processor, cause the processor to: resampling the unlabeled video clip at a first frequency ratio to generate a negative sample clip; passing the unlabeled video clip through the remote PPG model to generate an anchor PPG signal; passing the negative sample clip through the remote PPG model to generate a negative PPG signal; resampling the negative PPG signal at a second frequency ratio that is the inverse of the first frequency ratio to generate a positive PPG signal; adjusting one or more model weights of the remote PPG model based on a first loss calculated using a first loss function that takes into account the anchor PPG signal, the negative PPG signal, and the positive PPG signal. 2. The system of embodiment 1, further comprising instructions. Aspect 4 The training module may include instructions that, when executed by the processor, cause the processor to: warping the unlabeled video clip by passing the unlabeled video clip through a saliency sampler; 4. The system of embodiment 3, further comprising instructions. Aspect 5 The training module may include instructions that, when executed by the processor, cause the processor to: generating an anchor subset view based on the anchor PPG signal, a negative subset view based on the negative PPG signal, and a positive subset view based on the positive PPG signal, the anchor subset view, the negative subset view, and the positive subset view having one or more views corresponding to each other; determining a positive sum based on a distance between the corresponding anchor subset view and the positive subset view; determining a negative sum based on a distance between the corresponding anchor subset view and the negative subset view; determining a difference between the positive total and the negative total to generate the first loss; 4. The system of embodiment 3, further comprising instructions. Aspect 6 The training module may include instructions that, when executed by the processor, cause the processor to: adjusting the one or more model weights of the remote PPG model based on a second loss using a second loss function that uses the anchor PPG signal and a ground truth PPG signal. 4. The system of embodiment 3, further comprising instructions. Aspect 7 7. The system of embodiment 6, wherein the second loss function includes at least one of a Pearson correction loss function, a signal-to-noise ratio loss function, and a maximum cross-correlation loss function. Aspect 8 The training module may include instructions that, when executed by the processor, cause the processor to: subtracting an average of the anchor PPG signal and the ground truth PPG signal from the anchor PPG signal and the ground truth PPG signal; performing a Fast Fourier Transform of the anchor PPG signal and the ground truth PPG signal and determining a cross-correlation between the anchor PPG signal and the ground truth PPG signal by multiplying one of the anchor PPG signal and the ground truth PPG signal by a conjugate of the other; filtering the cross-correlation to remove cross-correlation outside a range of expected heart rate rates to generate a filtered spectrum of the cross-correlation; determining a final cross-correlation by taking the inverse of the fast Fourier transform of the filtered spectrum of the cross-correlation and dividing by the standard deviation of the anchor PPG signal and the ground truth PPG signal; the maximum value of the final cross-correlation is the second loss; 8. The system of embodiment 7, further comprising instructions. Aspect 9 1. A method for training a remote photoplethysmography ("PPG") model that outputs a subject PPG signal based on a subject video clip of a subject, the method comprising: The method includes training the remote PPG model in a self-supervised contrastive learning method using unlabeled video clips having a sequence of images of human faces. Aspect 10 10. The method of claim 9, wherein the remote PPG model is configured to output a saliency signal based on the subject video clip that indicates portions of the subject video clip where the PPG signal is strongest. Aspect 11 resampling the unlabeled video clip at a first frequency ratio to generate a negative sample clip; passing the unlabeled video clip through the remote PPG model to generate an anchor PPG signal; passing the negative sample clip through the remote PPG model to generate a negative PPG signal; resampling the negative PPG signal at a second frequency ratio that is the inverse of the first frequency ratio to generate a positive PPG signal; adjusting one or more model weights of the remote PPG model based on a first loss calculated using a first loss function that takes into account the anchor PPG signal, the negative PPG signal, and the positive PPG signal; 10. The method of embodiment 9, further comprising: Aspect 12 12. The method of claim 11, further comprising warping the unlabeled video clip by passing the unlabeled video clip through a saliency sampler. Aspect 13 generating an anchor subset view based on the anchor PPG signal, a negative subset view based on the negative PPG signal, and a positive subset view based on the positive PPG signal, the anchor subset view, the negative subset view, and the positive subset view having one or more views corresponding to each other; determining a positive total based on a distance between the corresponding anchor subset view and the positive subset view; determining a negative total based on a distance between the corresponding anchor subset view and the negative subset view; determining a difference between the positive total and the negative total to generate the first loss; 12. The method of embodiment 11, further comprising: Aspect 14 12. The method of claim 11, further comprising adjusting the one or more model weights of the remote PPG model based on a second loss using a second loss function that uses the anchor PPG signal and a ground truth PPG signal. 。 Aspect 15 15. The method of embodiment 14, wherein the second loss function includes at least one of a Pearson correction loss function, a signal-to-noise ratio loss function, and a maximum cross-correlation loss function. Aspect 16 subtracting an average of the anchor PPG signal and the ground truth PPG signal from the anchor PPG signal and the ground truth PPG signal; determining a cross-correlation between the anchor PPG signal and the ground truth PPG signal by performing a Fast Fourier Transform of the anchor PPG signal and the ground truth PPG signal and multiplying one of the anchor PPG signal and the ground truth PPG signal by the conjugate of the other; filtering the cross-correlation to remove cross-correlation outside a range of expected heart rate rates to generate a filtered spectrum of the cross-correlation; determining a final cross-correlation by taking the inverse of the Fast Fourier Transform of the filtered spectrum of the cross-correlation and dividing by the standard deviation of the anchor PPG signal and the ground truth PPG signal; Further comprising: 16. The method of claim 15, wherein a maximum value of the final cross-correlation is the second loss. Aspect 17 A non-transitory computer-readable medium having instructions, The instructions, when executed by a processor, cause the processor to train a remote photoplethysmography ("PPG") model in a self-supervised contrastive learning manner using unlabeled video clips having a sequence of images of a human face, the remote PPG model being configured to output a subject PPG signal based on the subject video clips of the subject. Non-transitory computer-readable medium. Aspect 18 A non-transitory computer-readable medium as described in aspect 17, wherein the remote PPG model is configured to output a saliency signal based on the subject video clip, the saliency signal indicating a portion of the subject video clip where the PPG signal is strongest. Aspect 19 The instructions, when executed by the processor, cause the processor to: resampling the unlabeled video clip at a first frequency ratio to generate a negative sample clip; passing the unlabeled video clip through the remote PPG model to generate an anchor PPG signal; passing the negative sample clip through the remote PPG model to generate a negative PPG signal; resampling the negative PPG signal at a second frequency ratio that is the inverse of the first frequency ratio to generate a positive PPG signal; the anchor PPG signal, the negative PPG signal, and the positive PPG signal adjusting one or more model weights of the remote PPG model based on a first loss calculated using a first loss function that takes into account the signal;
[0033] 18. The non-transitory computer-readable medium of embodiment 17. Aspect 20 20. The non-transitory computer-readable medium of claim 19, further comprising instructions that, when executed by the processor, cause the processor to warp the unlabeled video clip by passing the unlabeled video clip through a saliency sampler.
Claims
1. 1. A system for training a remote photoplethysmography ("PPG") model that outputs a subject PPG signal based on a subject video clip of a subject, the system comprising: A processor; a memory in communication with the processor, the memory having a training module having instructions that, when executed by the processor, cause the processor to train the remote PPG model in a self-supervised contrastive learning manner using unlabeled video clips having a sequence of images of human faces; the remote PPG model is a neural network that outputs a heart rate signal or a saliency signal based on an input subject video clip; the remote PPG model is configured to output a saliency signal based on the subject video clip, the saliency signal being indicative of portions of the subject video clip where a PPG signal is strongest; The training module may include instructions that, when executed by the processor, cause the processor to: resampling the unlabeled video clip at a first frequency ratio to generate a negative sample clip; passing the unlabeled video clip through the remote PPG model to generate an anchor PPG signal; passing the negative sample clip through the remote PPG model to generate a negative PPG signal; resampling the negative PPG signal at a second frequency ratio that is the inverse of the first frequency ratio to generate a positive PPG signal; adjusting one or more model weights of the remote PPG model based on a first loss calculated using a first loss function that takes into account the anchor PPG signal, the negative PPG signal, and the positive PPG signal. further comprising the instructions: system.
2. The training module may include instructions that, when executed by the processor, cause the processor to: warping the unlabeled video clip by passing the unlabeled video clip through a saliency sampler; The system of claim 1 further comprising instructions.
3. The training module may include instructions that, when executed by the processor, cause the processor to: generating an anchor subset view based on the anchor PPG signal, a negative subset view based on the negative PPG signal, and a positive subset view based on the positive PPG signal, the anchor subset view, the negative subset view, and the positive subset view having one or more views corresponding to each other; determining a positive sum based on a distance between the corresponding anchor subset view and the positive subset view; determining a negative sum based on a distance between the corresponding anchor subset view and the negative subset view; determining a difference between the positive total and the negative total to generate the first loss; The system of claim 1 further comprising instructions.
4. The training module may include instructions that, when executed by the processor, cause the processor to: adjusting the one or more model weights of the remote PPG model based on a second loss using a second loss function that uses the anchor PPG signal and a ground truth PPG signal; The system of claim 1 further comprising instructions.
5. 5. The system of claim 4, wherein the second loss function comprises at least one of a Pearson correction loss function, a signal-to-noise ratio loss function, and a maximum cross-correlation loss function.
6. The training module may include instructions that, when executed by the processor, cause the processor to: subtracting an average of the anchor PPG signal and the ground truth PPG signal from the anchor PPG signal and the ground truth PPG signal; determining a cross-correlation between the anchor PPG signal and the ground truth PPG signal by performing a Fast Fourier Transform of the anchor PPG signal and the ground truth PPG signal and multiplying one of the anchor PPG signal and the ground truth PPG signal by a conjugate of the other; filtering the cross-correlation to remove cross-correlation outside a range of expected heart rate rates to generate a filtered spectrum of the cross-correlation; determining a final cross-correlation by taking the inverse of the fast Fourier transform of the filtered spectrum of the cross-correlation and dividing by the standard deviation of the anchor PPG signal and the ground truth PPG signal; the maximum value of the final cross-correlation is the second loss; The system of claim 5 further comprising instructions.
7. 1. A method for training a remote photoplethysmography ("PPG") model that outputs a subject PPG signal based on a subject video clip of a subject, the method comprising: training, by a processor, the remote PPG model in a self-supervised contrastive learning manner using unlabeled video clips having a sequence of images of human faces; the remote PPG model is a neural network that outputs a heart rate signal or a saliency signal based on an input subject video clip; the remote PPG model is configured to output a saliency signal based on the subject video clip, the saliency signal being indicative of portions of the subject video clip where a PPG signal is strongest; resampling the unlabeled video clip at a first frequency ratio to generate a negative sample clip; passing the unlabeled video clip through the remote PPG model to generate an anchor PPG signal; passing the negative sample clip through the remote PPG model to generate a negative PPG signal; resampling the negative PPG signal at a second frequency ratio that is the inverse of the first frequency ratio to generate a positive PPG signal; adjusting one or more model weights of the remote PPG model based on a first loss calculated using a first loss function that takes into account the anchor PPG signal, the negative PPG signal, and the positive PPG signal; [0033] method.
8. The method of claim 7 , further comprising warping the unlabeled video clip by passing the unlabeled video clip through a saliency sampler.
9. generating an anchor subset view based on the anchor PPG signal, a negative subset view based on the negative PPG signal, and a positive subset view based on the positive PPG signal, the anchor subset view, the negative subset view, and the positive subset view having one or more views that correspond to each other; determining a positive total based on a distance between the corresponding anchor subset view and the positive subset view; determining a negative total based on a distance between the corresponding anchor subset view and the negative subset view; determining a difference between the positive total and the negative total to generate the first loss; The method of claim 7 further comprising:
10. 8. The method of claim 7, further comprising adjusting the one or more model weights of the remote PPG model based on a second loss using a second loss function that uses the anchor PPG signal and a ground truth PPG signal.
11. The method of claim 10 , wherein the second loss function comprises at least one of a Pearson correction loss function, a signal-to-noise ratio loss function, and a maximum cross-correlation loss function.
12. subtracting an average of the anchor PPG signal and the ground truth PPG signal from the anchor PPG signal and the ground truth PPG signal; determining a cross-correlation between the anchor PPG signal and the ground truth PPG signal by performing a Fast Fourier Transform of the anchor PPG signal and the ground truth PPG signal and multiplying one of the anchor PPG signal and the ground truth PPG signal by the conjugate of the other; filtering the cross-correlation to remove cross-correlation outside a range of expected heart rate rates to generate a filtered spectrum of the cross-correlation; determining a final cross-correlation by taking the inverse of the Fast Fourier Transform of the filtered spectrum of the cross-correlation and dividing by the standard deviation of the anchor PPG signal and the ground truth PPG signal; Further comprising: The method of claim 11 , wherein the maximum value of the final cross-correlation is the second loss.
13. A non-transitory computer-readable medium having instructions, The instructions, when executed by a processor, cause the processor to train a remote photoplethysmography ("PPG") model in a self-supervised contrastive learning manner using unlabeled video clips having a sequence of images of a human face, the remote PPG model being configured to output a subject PPG signal based on the subject video clips of a subject; the remote PPG model is a neural network that outputs a heart rate signal or a saliency signal based on an input subject video clip; the remote PPG model is configured to output a saliency signal based on the subject video clip, the saliency signal indicative of portions of the subject video clip where a PPG signal is strongest; The instructions, when executed by the processor, cause the processor to: resampling the unlabeled video clip at a first frequency ratio to generate a negative sample clip; passing the unlabeled video clip through the remote PPG model to generate an anchor PPG signal; passing the negative sample clip through the remote PPG model to generate a negative PPG signal; resampling the negative PPG signal at a second frequency ratio that is the inverse of the first frequency ratio to generate a positive PPG signal; the anchor PPG signal, the negative PPG signal, and the positive PPG signal adjusting one or more model weights of the remote PPG model based on a first loss calculated using a first loss function that takes into account the signal; [0033] Non-transitory computer-readable medium.
14. 14. The non-transitory computer-readable medium of claim 13, further comprising instructions that, when executed by the processor, cause the processor to warp the unlabeled video clip by passing the unlabeled video clip through a saliency sampler.
Citation Information
Patent Citations
Remote monitoring of vital signs
JP2014527863A
Apparatus, system and method for pulse detection
JP2019508116A
Optimizing Policy Controllers for Robotic Agents Using Image Embedding
JP2020530602A