Generating physiologically realistic avatars for training of non-contact models to recover physiological properties
By generating physiologically realistic avatar videos and training models using synthetic data, the challenge of collecting high-quality physiological data was addressed, enabling robust model training and evaluation under different conditions.
Patent Information
- Application Number
- CN202180043487.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-19
- Filing Date
- 2021-04-06
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2041-04-06
AI Technical Summary
Collecting high-quality physiological data presents challenges, including the high cost of recruiting and equipping participants, the lack of data on motion, lighting variations, or appearance types in datasets, and the leakage of identity and health information, resulting in fragile training models that lack broad applicability.
By training a physiological sensing system with synthetic data, and generating physiologically realistic avatar videos through a computer graphics pipeline, the system simulates various movements and lighting changes, avoiding physical involvement and identity exposure, and utilizes cloud computing for scalable data generation.
It provides robust training data, enabling the evaluation of model performance under different conditions, and improves the accuracy and broad applicability of non-contact physiological measurements.
Smart Images

Figure CN115802943B_ABST
Abstract
Description
Background Technology
[0001] Collecting high-quality physiological data presents numerous challenges. First, recruiting and equipping participants is often expensive and requires advanced technical expertise, severely limiting the potential number. This is particularly true for imaging-based methods, which require recording and storing video content. Second, the training datasets already collected may not contain the type of motion, lighting changes, or appearance in the application context. Therefore, models trained on this data may be fragile and fail to generalize well. Third, the data can reveal the identity of the subject and / or sensitive health information. This is exacerbated for imaging methods by the fact that most video recording datasets include the subject's face in some or all frames. It is with regard to these and other general considerations that the aspects disclosed herein are made. Furthermore, while relatively specific problems may be discussed, it should be understood that these examples should not be limited to addressing specific problems identified in the background art or elsewhere in this disclosure. Summary of the Invention
[0002] According to examples in this disclosure, synthetic data can be used to train physiological sensing systems, thereby circumventing the challenges associated with recruiting and equipping participants, limited training data encompassing various types of motion, lighting variations, or appearances, and identity protection. Once the computer graphics pipeline is in place, the generation of synthetic data is more scalable than recorded video because computation is relatively inexpensive and readily available using cloud computing. Furthermore, rare events or underrepresented populations in videos can be simulated given a proper understanding of the statistical characteristics of the events or a set of examples. Additionally, synthetic datasets do not need to contain facial or physiological signals similar to any particular individual. Finally, parametric simulations will systematically alter certain variables of interest (e.g., motion speed or lighting intensity within the video), which is useful for training more robust methods and evaluating performance under different conditions.
[0003] According to examples in this disclosure, high-fidelity physiological simulations can be used to enhance training data that can be used to improve non-contact physiological measurements.
[0004] According to at least one example of this disclosure, a method for generating a video sequence including a physiologically realistic avatar is provided. The method may include receiving an albedo for the avatar, modifying a subsurface skin color associated with the albedo based on physiological data associated with physiological characteristics, rendering the avatar based on the albedo and the modified subsurface skin color, and synthesizing frames of a video, the frames of which contain the avatar.
[0005] According to at least one example of this disclosure, a system for training a machine learning model using video sequences including physiologically realistic avatars is provided. The system may include a processor and a memory storing instructions, which, when executed by the processor, cause the processor to receive a request from a requesting entity to train a machine learning model to detect physiological characteristics, receive a plurality of video segments, wherein one or more of the video segments contain a synthetic physiologically realistic avatar generated using physiological characteristics, train the machine learning model using the plurality of video segments, and provide the trained model to the requesting entity.
[0006] According to at least one example of this disclosure, a computer-readable medium is provided. The computer-readable medium includes instructions that, when executed by a processor, cause the processor to receive a request to recover physiological characteristics from a video clip, obtain a machine learning model trained with training data containing physiologically realistic avatars generated from the physiological characteristics, receive the video clip, identify metrics associated with the physiological characteristics from the video clip using the trained machine learning model, and provide an assessment of the physiological characteristics to the requesting entity based on the metrics.
[0007] Any one of the foregoing aspects in combination with any other aspect of one or more of the foregoing aspects. Any one of the one or more aspects described herein.
[0008] This synopsis is provided to present a simplified selection of concepts, which will be further described in the detailed description below. This synopsis is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Additional aspects, features, and / or advantages of the examples will be set forth in part in the description which follows, and will be apparent in part from the description, or may be learned by practice of this disclosure. Attached Figure Description
[0009] The following diagram illustrates non-restrictive and non-exhaustive examples.
[0010] Figure 1 First details of an example of generating a physiologically realistic avatar video according to this disclosure are depicted;
[0011] Figure 2 Details of video frames involving rendering and compositing, including physiologically realistic avatars, according to examples of this disclosure are depicted;
[0012] Figure 3 A second detail is depicted regarding the generation of physiologically realistic avatar videos according to an example of this disclosure;
[0013] Figure 4 Details involving the training of a machine learning model according to examples of this disclosure are described;
[0014] Figure 5 An example rendering depicts an embodiment of physiological data signals according to an example of this disclosure;
[0015] Figure 6 Details of an example according to this disclosure involving the use of a trained machine learning model to recover physiological signals are described;
[0016] Figure 7 Example diagrams depicting waveforms and power spectra for isolating or otherwise determining physiological signals according to examples of this disclosure are provided.
[0017] Figure 8 Details of physiologically realistic video and / or model generators according to examples of this disclosure are depicted;
[0018] Figure 9 A method for generating physiologically realistic avatar videos, according to examples of this disclosure, is described;
[0019] Figure 10 Methods involving training machine learning models, based on examples of this disclosure, are described;
[0020] Figure 11 Methods involving the generation and / or positioning of physiologically realistic avatar videos, according to examples of this disclosure, are described;
[0021] Figure 12 The present disclosure describes a method for recovering physiological signals using a trained machine learning model, according to examples of this disclosure.
[0022] Figure 13 A block diagram depicting the physical components (e.g., hardware) of a computing device that can implement aspects of this disclosure;
[0023] Figure 14A The illustration shows a first example of a computing device in which aspects of this disclosure can be practiced;
[0024] Figure 14B A second example of a computing device that can practice aspects of this disclosure is illustrated; and
[0025] Figure 15 The illustration shows at least one aspect of the architecture of a system for processing data according to an example of this disclosure. Detailed Implementation
[0026] In the following detailed description, reference is made to the accompanying drawings, which form a part of the description, in which specific embodiments or examples are illustrated by way of illustration. These aspects may be combined, other aspects may be utilized, and structural changes may be made without departing from this disclosure. Embodiments may be practiced as methods, systems, or devices. Therefore, embodiments may take the form of hardware implementations, entirely software implementations, or implementations combining software and hardware aspects. Accordingly, the following detailed description should not be considered limiting, and the scope of this disclosure is defined by the appended claims and their equivalents.
[0027] Photoplethysmography (PPG) is a non-invasive method for measuring peripheral hemodynamics and vital signs, such as blood volume pulse (BVP), via light reflected or transmitted through the skin. While traditional PPG sensors are used through skin contact, recent studies have shown that digital imagers can be used even at a distance from the body, offering several unique benefits. First, for subjects with delicate skin (e.g., infants in the NICU, burn patients, or the elderly), contact sensors could damage their skin, cause discomfort, and / or increase their susceptibility to infection. Second, the ubiquitous presence of cameras (found on many tablets, PCs, and mobile phones) allows for inconspicuous, ubiquitous health monitoring. Third, unlike traditional contact-based measurement devices (e.g., smartwatches), remote cameras allow for spatial mapping of pulse signals, which can be used to approximate pulse wave velocity and capture spatial patterns of peripheral hemodynamics.
[0028] While non-contact photoplethysmography (iPPG) measurements (also known as imaging photoplethysmography) offer many benefits, this approach is particularly susceptible to various environmental factors, posing challenges to related research. For example, recent research has focused on making iPPG measurements more robust to dynamic lighting and motion, as well as characterizing and resisting the effects of video compression. Historically, iPPG methods have typically relied on unsupervised methods (e.g., Independent Component Analysis (ICA) or Principal Component Analysis (PCA)) or handcrafted separation algorithms. More recently, supervised neural models have been proposed, offering state-of-the-art performance in the context of heart rate measurement. These performance gains are often a direct result of the model scaling well with the amount of training data; however, as with many machine learning tasks, the quantity and diversity of available data quickly become limiting factors.
[0029] As mentioned earlier, collecting high-quality physiological data is challenging for several reasons. First, recruiting and equipping participants is often expensive and requires advanced technical expertise, severely limiting the potential number. This is especially true for imaging-based methods, which require recording and storing video content. Second, the training datasets already collected may not contain the types of motion, lighting variations, or appearances needed to train the model. Consequently, models trained on this data may be fragile and fail to generalize well. Third, the data can reveal the subject's identity and / or sensitive health information. This is exacerbated for imaging methods, as most video-recorded datasets contain the subject's face in some or all frames.
[0030] Based on the examples in this disclosure, synthetic data can be used to train iPPG systems to overcome the aforementioned challenges. Synthetic data, which is more scalable than recorded video, can be generated using a graph pipeline. Furthermore, generating synthetic data is computationally relatively inexpensive and can be performed using cloud computing. Rare events or underrepresented populations can be simulated in videos, and such simulated videos do not need to contain facial or physiological signals similar to any particular individual. Moreover, parametric simulations will provide a way to systematically change certain variables of interest (e.g., motion speed or lighting intensity within the video), which is useful for training more robust models and evaluating model performance under different conditions.
[0031] Camera-based vital sign measurements using photoplethysmography involve capturing subtle color variations within skin pixels. The graphical simulation first assumes a light source with a constant spectral composition but varying intensity. Therefore, the red, green, and blue (RGB) values of the k-th skin pixel in the image sequence can then be defined by a time-varying function:
[0032] C k (t)=I(t)·(v s (t)+v d (t))+v n (t) Formula 1
[0033] C k (t)=I(t)·(v s (t)+v abs (t)+v sub (t))+v n (t) Formula 2 Where C k (t) represents a vector of RGB color channel values; I(t) is the luminance intensity level, which varies with the light source and the distance between the light source, skin tissue, and camera; I(t) is modulated by two components in the DRM: specular (gloss) reflection v s(t), specular reflection and diffuse reflection of the skin surface v d (t). Diffuse reflection is further divided into two parts: light absorption in skin tissue v abs (t) and subsurface scattering v sub (t); v n (t) represents the quantization noise of the camera sensor. I(t), v s (t) and v n (t) can be decomposed into a static part and a time-dependent part through a linear transformation:
[0034] v d (t)=u d ·d0+(u abs +u sub Formula 4: p(t)
[0035] Where u d This represents the unit color vector of skin tissue; d0 represents a fixed reflectance; v abs (t) and v sub (t) represents the relative pulse intensity caused by changes in blood volume, changes in hemoglobin and melanin absorption, and changes in subsurface scattering, respectively; p(t) represents BVP.
[0036] v s (t)=u s Formula 5: (s0+Φ(m(t),p(t)))
[0037] Where u s is the unit color vector of the light source spectrum; s0 and Φ(m(t), p(t)) represent the static and variable parts of specular reflection; m(t) represents all non-physiological changes, such as light source flicker, head rotation, facial expressions and movements (e.g., blinking, smiling).
[0038] I(t)=I0·(1+Ψ(m(t),p(t))) Formula 6 Where I0 is the stationary portion of luminance intensity, and I0·Ψ(m(t), p(t)) is the intensity variation observed by the camera. The interaction between physiological motion Φ(·) and non-physiological motion Ψ(·) is typically a complex nonlinear function. The stationary components from specular and diffuse reflection can be combined into a single component representing the reflection from stationary skin:
[0039] u c ·c0=u s ·s0+u d Formula 7 (d0)
[0040] Where u cLet c represent the unit color vector of skin reflection and c0 represent the reflection intensity. Substituting (3), (4), (5), and (6) into (1) produces:
[0041] C k (t)=I0·(1+Ψ(m(t),p(t)))· Formula 8
[0042] (u c ·c0+u s ·Φ(m(t),p(t))+(u abs +u sub )·p(t))+v n (t)
[0043] Since the time-varying components are several orders of magnitude smaller than the stationary components in Equation 7, any product between the changing terms can be ignored to make C k Approximately:
[0044] C k (t)≈u c ·I0·c0+u c ·I0·c0·Ψ(m(t),p(t))+u s ·I0·Φ(m(t),p(t))+(u abs +u sub )·I0·p(t)+v n (t) Formula 9
[0045] To synthesize data for physiological measurement methods, it is desirable to create skin with RGB variations as a function of p(t). A principled two-way scattering distribution function (BSDF) shader is used, employing subsurface color and subsurface radius parameters, u p u abs and u sub The components can be captured. Specular reflection is controlled by specular reflection parameters. Therefore, for a given pulse signal p(t), the appearance of the skin over time can be synthesized. Furthermore, changes in the skin's appearance, as well as various other variations, can be synthesized, representing noise sources for vital sign measurements. Data synthesized in this way is highly valuable for improving the versatility of camera-based vital sign measurement algorithms.
[0046] For any of the video-based physiological measurement methods, the task is to start from C k (t) Extract p(t). Use a machine learning model to capture C in Equation 8. k The motivation for the relationship between p(t) and p(t) is that neural models can capture more complex relationships than handcrafted separation or source separation algorithms (e.g., ICA, PCA), which ignore p(t) within Φ(·) and Ψ(·) and assume Ck There is a linear relationship between p(t) and p(t).
[0047] High-fidelity facial avatars and physiology-based animation models (based on the above) were generated to simulate facial videos with realistic blood flow (pulse) signals. These videos were then used to train a neural model to recover blood volume pulse (BVP) from video sequences. The resulting model was tested on real-world video benchmark datasets.
[0048] To achieve a physiologically realistic appearance for the synthesized avatars, photoplethysmography (PPG) waveform recordings can be used. For example, various PPG and respiration datasets with different contact PPG recordings and sampling frequencies from different individuals can be used. PPG recordings from different subjects can be used to synthesize multiple avatars. The synthesized video can be of any length, such as short sequences (nine 10-second sequences); therefore, only a small portion of the PPG recordings can be used.
[0049] To train the machine learning model, a realistic model of facial blood flow was synthesized. Therefore, blood flow can be simulated by adjusting the properties of the physically based shading material used to render the avatar's face. Specifically, the albedo component of the material is a texture map transferred from a high-quality 3D facial scan. In some instances, facial hair has been removed from these textures, allowing skin properties to be easily manipulated (3D hair can be added later in this process). Specular reflection effects are controlled by a roughness map to make certain parts of the face (e.g., the lips) brighter than others.
[0050] As blood flows through the skin, the skin's composition changes, leading to variations in subsurface color. These skin color variations can be manipulated using subsurface color parameters. Weights for these parameters can be derived from the absorption spectrum of hemoglobin and typical frequency bands from digital cameras. Therefore, subsurface color parameters can be varied across all skin pixels (but not non-skin pixels) on an albedo map. The albedo map can be an image texture without any shadows or highlights. Furthermore, variations in subsurface scattering are captured by changing the subsurface radius for each channel as blood volume changes. A spatially weighted subsurface scattering texture, capturing variations across the skin layers of the face, is used. The BSDF subsurface radius for the RGB channels can be changed using the same weighting as described above. Empirically, these parameters can be used with synthetic data to train camera-based vital sign measurements. Changing subsurface scattering alone, without altering the subsurface color, may be too subtle and may not reproduce the effect of BVP on reflected light observed in real-world video. Alternatively or additionally, color spaces other than RGB can be used. For example, color spaces that include luminance and chrominance channels (e.g., YUV, Y'UV, YCrCb, Y'CrCb) can be used. Similarly, hue, saturation, and value (HSV) color spaces can be used.
[0051] By precisely specifying the types of variations appearing in the data, machine learning systems can be trained to be robust to forms of variation encountered in the real world. Many different system variations can be used for aspects disclosed herein, such as facial appearance, head movement, facial expressions, and environment. For example, faces with fifty different appearances can be synthesized. For each face, the skin material can be configured as an albedo texture randomly picked from an albedo set. To model wrinkle-scale geometry, a matching high-resolution displacement map transmitted from scan data can be applied. Skin type is particularly important in imaging PPG measurements; therefore, an approximate skin type distribution for a face can contain a non-uniform distribution, but represents a more balanced distribution than in existing imaging PPG datasets. Since motion is one of the largest sources of noise in imaging PPG measurements, a set of rigid head movements can be simulated to enhance training examples capturing these conditions. In particular, the head can smoothly rotate around the vertical axis at angular velocities of 0, 10, 20, and 30 degrees / second, similar to head movements; facial expression movements are also a common source of noise in PPG measurements. To simulate facial expressions, the video can be synthesized using smiling, blinking, and mouth opening (similar to speaking), some of the most common facial expressions seen in everyday life. Smiling and blinking can be applied to the face using blend shapes, and the mouth can be opened by rotating the jawbone using linear blending skinning. The face can be rendered in different image-based environments to create realistic variations in the face's background appearance and lighting. Static backgrounds and backgrounds with motion can be used. In some instances, even more realistic facial occlusion is included.
[0052] Figure 1Details are depicted according to examples of this disclosure concerning the generation and subsequent training of a machine learning model for the detection of physiological signals using physiologically realistic synthetic avatars. Specifically, physiologically realistic avatar generator 116 can generate synthetic videos of physiologically realistic avatars 120; these synthetic videos can then be used to train an end-to-end learning model 136 (such as a neural model) to recover or identify specific physiological responses from video sequences. Physiologically realistic avatar generator 116 can synthesize physiologically realistic avatars based on physiological data 104, appearance data 108, and parameter data 112. Physiological data 104 can contain one or more signals indicating a physiological response, condition, or signal. For example, physiological data 104 can correspond to blood volume pulse measurements based on real human records, such as blood volume pulse waveforms. As another example, physiological data 104 can correspond to respiratory rate / waveform, cardiac condition indicated by waveforms or measurements (such as atrial fibrillation), and / or oxygen saturation levels. For example, physiological data 104 may correspond to cardiac impulse plethysmography (BCG) and / or respiratory data, and may be an impulse ECG waveform or a respiratory waveform. As another example, physiological data 104 may be an optical plethysmography waveform. Optical plethysmography waveforms can be generated by optical sensors to measure changes in blood volume in a non-contact manner. Optical plethysmography provides useful physiological information for assessing cardiovascular function, and PPG signals are typically measured using transmission and reflection methods, which sense light passing through or reflected from tissue. In some examples, physiological data 104 may be used to assess one or more conditions, such as, but not limited to, peripheral artery disease, Raynaud's phenomenon, systemic sclerosis, and Takayasu arteritis. The PPG waveform also varies with respiratory patterns. For example, the amplitude, frequency, and baseline of the PPG waveform are modulated by respiration. Physiological data 104 may contain other waveforms, measurements, or others, and may be from different individuals. In some examples, the waveform may be a record of various lengths and sampling rates.
[0053] Appearance data 108 may contain skin material with an albedo texture selected randomly. In some examples, the albedo component of the material is a texture map transferred from a high-quality 3D facial scan. As mentioned above, in some examples, facial hair has been removed from these textures, allowing skin properties to be easily manipulated. Specular effects can be controlled to make certain parts of the face (e.g., lips) brighter than others. In some examples, high-resolution displacement map wrinkle-scale geometry transferred from the scan data can be applied. Skin type can also be selected randomly. For example, the skin type can be selected from one of six Fitzpatrick skin types. The Fitzpatrick skin type (or photo type) depends on the amount of melanin in the skin. This is determined by body color (white, brown, or black skin) and exposure to ultraviolet radiation (tanning). Fitzpatrick skin types may include: 1. Pale skin; 2. Fair skin; 3. Darker white skin; 4. Light brown skin; 5. Brown skin; 6. Dark brown or black skin. In some examples, skin type classifications other than Fitzpatrick skin type can be utilized.
[0054] Parameter data 112 may contain parameters that affect the avatar and / or light transmission and reflection. For example, parameter data 112 may include facial expressions, head movements, background lighting, environment, etc. Since motion is one of the largest sources of noise in imaging PPG measurements, rigid head movements can be used to enhance training examples capturing such conditions. The head can rotate around the vertical axis at different angular velocities (such as 0, 10, 20, and 30 degrees / second). Similarly, to simulate expressions, videos of smiling, blinking, opening the mouth (similar to speaking), and / or other common facial expressions exhibited in daily life can be synthesized. Smiling and blinking can be applied to the face using a set of blended shapes; the mouth can be opened by rotating the jawbone with linear blending skin. Furthermore, different environments can be utilized to render the avatar to create a realistic avatar in both background appearance and facial lighting. In some examples, video sequences describing physiologically realistic avatars may contain static backgrounds. Alternatively or additionally, the background may contain motion or avatar occlusion that more closely resembles challenging real-life conditions. Parameter data 112 may also include environmental conditions; for example, parameter data 112 may include temperature, time of day, weather such as wind, rain, snow, etc.
[0055] Physiological data 104, appearance data 108, and parameter data 112 can be provided to a physiologically realistic avatar generator 116. The physiologically realistic avatar generator 116 can use a bidirectional scattering distribution function (BSDF) shader to render the physiologically realistic avatar and combine the physiologically realistic image with a background. Furthermore, a synthetic video of the physiologically realistic avatar 120 can be generated. The synthetic video of the physiologically realistic avatar 120 can contain, for example, various video sequences depicting different physiologically realistic avatars 122 and 124. In some examples, physiologically realistic video sequences and / or physiologically realistic avatars can be stored in a physiologically realistic avatar video repository 128. One or more of the physiologically realistic avatars 122 can be labeled as training data 123. Examples of training labels include, but are not limited to, blood volume, pulse, and / or peripheral artery disease. Therefore, when using synthetic videos to train a machine learning model, labels can identify one or more features of the video as training and / or test / validation data. The synthetic video of the physiologically realistic avatar 120 can be fed to an end-to-end learning model 136, such as a convolutional attention network (CAN), to evaluate the impact of the synthetic data on the quality of the physiological signals 140 recovered from the video sequence. Furthermore, in addition to the real human video 132, the synthetic video of the physiologically realistic avatar 120 can also be used to train the end-to-end learning model 136.
[0056] CAN utilizes motion and appearance representations learned jointly through an attention mechanism. The method primarily consists of a two-branch convolutional neural network. The motion branch allows the network to distinguish intensity variations caused by noise, such as motion, from subtle feature intensity variations caused by physiological characteristics (such as blood flow). The motion representation is the difference between two consecutive video frames. To reduce noise caused by ambient lighting and variations in the distance from the face to the lighting source, frame differences based on a skin reflection model are first normalized. Normalization is applied to the video sequence by subtracting the pixel mean and dividing by the standard deviation. The appearance representation captures regions in the image that contribute strong iPPG signals. Via the attention mechanism, the appearance representation guides the motion representation and helps distinguish the iPPG signal from other noise sources. The input frames are similarly normalized by subtracting the mean and dividing by the standard deviation.
[0057] Once trained with physiologically realistic avatars and / or real human videos 132, the end-to-end learning model 136 can be used to evaluate video information from subject 148. Subject 148 can be equipped to compare physiological signals provided by gold standards, contact and / or non-contact measurement devices or sensors with recovered physiological signals 152 from the same participant. Thus, the two physiological signals can be compared to each other to determine the effectiveness of the end-to-end learning model 136. Upon finding that the end-to-end learning model 136 is effective and / or has the desired accuracy, the trained model, including model structure and model weights, can be stored in a physiological model repository 156, making the trained model usable for recovering physiological signals from different participants or subjects.
[0058] Figure 2 Additional details regarding the synthesis of physiologically realistic avatar video frames according to examples of this disclosure are depicted. More specifically, albedo 204 can be selected; the selection of albedo 204 may correspond to a texture map transmitted from a high-quality 3D facial scan. Albedo can be randomly selected or selected to represent a specific population. Albedo may lack facial hair, thus skin characteristics can be easily manipulated. Other parameters affecting appearance via appearance parameter 208 (such as, but not limited to, skin color / type, hair, mirror effect, wrinkles, etc.) are added. Skin type can be randomly selected or selected to represent a specific population. For example, skin type can be selected from one of six Fitzpatrick skin types; however, those skilled in the art will understand that classifications other than Fitzpatrick skin types can be used.
[0059] When blood flows through the skin, the skin's composition changes, resulting in changes in subsurface color. Therefore, changes in skin color tone can be manipulated using subsurface color parameters, including but not limited to a base subsurface skin color 212, subsurface skin color weights 220, and subsurface skin scattering parameters 228. The weights for the subsurface skin color weights 220 can be derived from the absorption spectrum of hemoglobin and from typical frequency bands of an example digital camera. For example, the example camera can provide colors based on the following frequency bands: red: 550-700 nm; green: 400-650 nm; blue: 350-550 nm. The subsurface skin color weights 220 can include weights for one or more color channels and can be applied to physiological data 216, which can be the same as or similar to the previously described physiological data 104 and can include one or more signals indicating a physiological response, condition, or signal. For example, physiological data 216 can correspond to blood volume pulse measurements based on real human records, such as blood volume pulse waveforms. As another example, physiological data 216 may correspond to respiratory rate / waveform, cardiac condition indicated by waveform or measurement (e.g., atrial fibrillation), and / or oxygen saturation level. For example, physiological data 216 may correspond to cardiac impulse recording (BCG) and may be an impulse cardiac recording waveform. As another example, physiological data 216 may be a photoplethysmography waveform. In some examples, physiological data 216 may be based on signal measurements from actual humans or may be synthesized based on known physiological signal characteristics indicating physiological responses, conditions, or signals. The weighted physiological data signal resulting from the application of subsurface skin color weights 220 may be added to the base subsurface skin color 212, resulting in subsurface skin color 224 comprising multiple color channels. Subsurface skin color 224 may be provided to shader 232. In some examples, subsurface skin color weights 220 may be applied to all pixels identified as facial pixels on the albedo map; subsurface skin color weights 220 may not be applied to non-skin pixels.
[0060] Furthermore, the subsurface radius for color channels can be manipulated to capture changes in subsurface scattering as physiological characteristics, such as blood volume, change. A spatially weighted subsurface scattering texture is used, employing a subsurface scattering radius texture that captures variations across the skin layers of the face. The subsurface radius for RGB channels can be altered by using weights that are the same as or similar to the subsurface skin color weight 220.
[0061] In some examples, extrinsic parameter 210 can alter the skin tone and color. Extrinsic parameter 210 can include parameters that affect the avatar and / or light transmittance and reflectance. For example, extrinsic parameter 210 can include facial expressions, head movements, background lighting, environment, etc. Since motion is one of the largest sources of noise in imaging PPG measurements, rigid head movements can be used to enhance training examples capturing such conditions. The head can rotate around the vertical axis at different angular velocities (such as 0, 10, 20, and 30 degrees / second). Similarly, to simulate expressions, the video can be synthesized with smiling, blinking, opening the mouth (such as speaking), and / or other common facial expressions exhibited in daily life. Smiling and blinking can be applied to the face using a blended shape set; the mouth can be opened by rotating the jawbone using linear blending skinning.
[0062] In some examples, the physiological processes being modeled cause changes in color and motion; therefore, motion weights 222 can be applied to the physiological data 216 to account for pixel movement and translation caused at least in part by the physiological data 216. For example, a region, portion, or area represented by one or more pixels can move from a first position in a first frame to a second position in a second frame. Thus, motion weights can provide a mechanism for identifying and / or addressing specific pixels in an input image that move or translate at least in part due to physiological characteristics. For example, blood flowing through veins, arteries, and / or under the skin can cause the veins, arteries, and / or skin to twist in one or more directions. Motion weights 222 can account for such movement or translation and, in some instances, can be represented as vectors.
[0063] In the example, shader 232 can provide an initial rendering of one or more pixels of the avatar based on extrinsic parameters 210, appearance parameters 208, subsurface skin color 224, subsurface skin scattering parameters 228, motion weights, and physiological data 216. Of course, other parameters can also be considered. Shader 232 can be a program running in a graphics pipeline that provides instructions to a computer processing unit (such as a graphics processing unit) on how to render one or more pixels. In the example, shader 232 can be a principled two-way scattering distribution function (BSDF) shader, which determines the probability that a particular ray of light will be reflected (scattered) at a given angle.
[0064] The image rendered by shader 232 can be an avatar for a specific video frame. In some examples, a background 236 can be added to the avatar so that the avatar appears in front of the image. In some examples, the background 236 can be static; in some examples, the background 236 can be dynamic. Further, in some examples, a foreground object contained in the background 236 can occlude part of the avatar. Frame sequence 240 can be composited at 240, resulting in a video sequence. Such frames can be combined with a video compositor configured to apply backgrounds and / or combine multiple frames or images into a video sequence. In some examples, the background 236 can be rendered by shader 232 along with the avatar.
[0065] Figure 3 Additional details are depicted regarding an example of a physiologically realistic video generator 304 according to this disclosure, which is configured to render and synthesize physiologically realistic video sequences containing physiologically realistic avatars. The physiologically realistic video generator 304 may be a computing device and / or a dedicated computing device specifically configured to render and synthesize video. The physiologically realistic video generator 304 may comprise multiple devices and / or utilize a portion of cloud infrastructure to divide one or more portions of the rendering and / or compositing task among different devices.
[0066] The physiologically realistic video generator 304 may include a physiologically realistic shader 308 and a frame synthesizer 312.
[0067] Physiological realism shader 308 may be the same as or similar to shader 232, and may provide initial rendering of one or more pixels of the avatar based on appearance parameters 320, albedo 324, physiological data 328, subsurface skin parameters 332, background 336, and extrinsic parameters 340. Of course, other parameters may also be considered by physiological realism shader 308. Appearance parameters 320 may be the same as or similar to appearance parameters 208; albedo 324 may be the same as or similar to albedo 204; physiological data 328 may be the same as or similar to physiological data 216; subsurface skin parameters 332 may be the same as or similar to basic subsurface skin color 212, subsurface skin color weight 220, and subsurface skin scattering parameters 228; background 336 may be the same as or similar to background 236, and extrinsic parameters 340 may be the same as or similar to extrinsic parameters 210.
[0068] The image rendered by the physiological realism shader 308 can be an avatar exhibiting a specific physiological response based on physiological data 328, and can be rendered as frames of video as discussed previously. In some examples, the avatar can be rendered in front of a background 336, such that the avatar appears in front of the image. In some examples, the background 336 can be static; in some examples, the background 336 can be dynamic. Furthermore, in some examples, foreground objects contained in the background 336 can occlude part of the avatar. The frames generated by the physiological realism shader 308 can be provided to a frame synthesizer 312 for compositing frames and assembling them into a video sequence. The synthesized video can then be provided to a physiological realism avatar video library 316, which can be the same as or similar to the physiological realism avatar video library 128.
[0069] Synthetic videos can be labeled or tagged before being stored; alternatively or additionally, synthetic videos can be stored in a location or repository associated with specific tags. Examples of tags include, but are not limited to, blood volume, pulse, and / or peripheral artery disease. Therefore, when using synthetic videos to train machine learning models, tags can identify one or more features of the video as training and / or testing / validation data.
[0070] Figure 4Additional details are depicted regarding an example according to this disclosure involving training a machine learning architecture 404 based on training data 408 to construct a machine learning model 442, the training data 404 comprising synthetic physiologically realistic videos from a synthetic video repository 412 and human videos from a human video repository 416. The synthetic physiologically realistic videos can be output from a frame synthesizer 312 as previously described and can be labeled with training tags to identify one or more physiological characteristics. Examples of training tags include, but are not limited to, blood volume, pulse, and / or peripheral artery disease. The human videos are videos of real individuals. In some examples, the machine learning architecture 404 utilizes training data comprising synthetic physiologically realistic videos, human videos, or a combination of both. The machine learning architecture 404 can be stored as processor-executable instructions in a file such that when a set of algorithms associated with the machine learning architecture 404 is executed by a processor, a machine learning model 442 comprising various layers and optimization functions and weights is constructed. That is, the various layers of the architecture including the machine learning architecture 404 can be iteratively trained using the training data 408 to recover the physiological signals 440 present in the training data 408. The individual layers of the machine learning structure 404 can be trained to identify and acquire physiological signals 440. After many iterations or rounds, a configuration of the machine learning structure 404 with the minimum error associated with the iteration (e.g., various layers and weights associated with one or more layers) can be utilized as a machine learning model 442, wherein the structure of the machine learning model can be stored in a model file 444, and the weights associated with one or more layers and / or configurations can be stored in a model weight file 448.
[0071] According to the examples in this disclosure, the machine learning architecture 404 may contain two paths: a first path associated with the motion model 424 and a second path associated with the appearance model 432. The architecture of the motion model 424 may contain, for example, nine layers with 128 hidden units. Furthermore, average pooling and hyperbolic tangent may be used as activation functions. The last layer of the motion model 424 may contain linear activation units and mean squared error (MSE) loss. The architecture of the appearance model 432 may be the same as that of the motion model 424, but without the last three layers (e.g., layers 7, 8, and 9).
[0072] Motion model 424 allows machine learning structure 404 to distinguish intensity changes caused by noise in motion, for example, subtle feature intensity variations derived from physiological characteristics. Motion representations are computed from the input differences between two consecutive video frames 420 (e.g., C(t) and C(t+1)). Ambient lighting may be uneven on the face, and the lighting distribution varies with the distance from the face to the light source and may be influenced by supervised learning methods. Therefore, to reduce these sources of lighting noise, the frame difference is first normalized at 428 using AC / DC normalization based on a skin reflection model. Normalization can be applied once across the entire video sequence by subtracting the pixel mean and dividing by the standard deviation. Furthermore, one or more layers (layers 1 through 5) can be convolutional layers of different or the same size and can be used to identify various feature maps utilized during training of machine learning structure 404. In the example, normalized difference 428 may correspond to normalized differences for three color channels (e.g., red, green, and / or blue channels). The individual layers of motion model 424 may contain feature maps of various sizes and color channels and / or various convolutions.
[0073] The appearance model 432 allows the machine learning structure 404 to learn which regions in an image are likely reliable for computing strong physiological signals (such as iPPG signals). The appearance model 432 can generate representations from the texture and color information of the input video frames. The appearance model 432 guides the motion representations to recover iPPG signals from the individual regions contained in the input image and further distinguish them from other noise sources. The appearance model 432 can use a single image or video frame as input. That is, a single frame of video or image 436 can be used as input to the various layers (layers 1–6).
[0074] Once trained, the machine learning structure 404 can be output as a machine learning model 442, wherein the structure of the machine learning structure 404 can be stored in a model file 444, and the individual weights of the machine learning model are stored in a model weight file 448. Although described as a specific deep learning implementation, it should be understood that the machine learning structure can be modified, adapted, or otherwise altered to achieve maximum accuracy associated with detecting physiological signals such as blood volume and pulse.
[0075] Figure 5Examples illustrating how an input physiological signal 504 (such as a pulse signal) can be rendered into a physiologically realistic avatar and affect the RGB pixel values in the resulting video frame 508 are described. For example, the input physiological signal 504 could correspond to blood volume based on a real human record; the input physiological signal 504 could be a blood volume pulse waveform. As another example, the physiological signal 504 could correspond to respiratory rate / waveform, cardiac condition indicated by a waveform or measurement (such as atrial fibrillation), and / or oxygen saturation level. For example, the physiological signal 504 could correspond to cardiac impaction imaging (BCG) data and could be a cardiac impaction imaging waveform. As another example, the physiological signal 504 could be a photoplethysmography waveform. The physiological signal 504 can contain other waveforms, measurements, or others, and can come from different individuals. In some examples, the waveform can be a record of various lengths and various sampling rates.
[0076] like Figure 5 As depicted, the avatar 506 representing the physiological signal 504 can be generated in multiple video frames 508 of the video sequence, as previously described. The scan line 510 of the avatar 506 is drawn over time in 512. The corresponding RGB pixels (red 520, blue 524, and green 528) are drawn over time in Figure 516. As shown in Figure 516, the decay waveform corresponding to the pulse signal can be identified from the RGB pixels. Alternatively or additionally, color spaces other than RGB can be used. For example, color spaces containing luminance and chrominance channels (e.g., YUV, Y'UV, YCrCb, Y'CrCb) can be used. Similarly, the hue, saturation, and value (HSV) color space can be used.
[0077] Figure 6 Additional details are depicted of a system 600 for recovering physiological signals from a video sequence using a trained machine learning model, according to examples of this disclosure. System 600 may include a subject or patient 604 within the field of view of a camera 608. The subject 604 or patient may be a real person and may or may not exhibit one or more physiological characteristics. For example, system 600 may utilize a trained machine learning model (such as machine learning model 620) to recover or determine one or more physiological characteristics of the subject 604. Physiological characteristics may be an assessment of pulse rate and / or cardiovascular function. As non-limiting examples, peripheral artery disease, Raynaud's phenomenon, systemic sclerosis, and Takayasu arteritis may be assessed. While the examples provided herein primarily concern cardiovascular function, other physiological characteristics and / or conditions may also be considered.
[0078] Camera 608 can correspond to any camera capable of capturing or photographing multiple images. In some examples, camera 608 can capture a sequence of frames 612 or images at a specific frame rate. Example frame rates may include, but are not limited to, 32 frames per second or 16 frames per second. Of course, other frame rates are also considered herein. The camera can provide a video containing a sequence of frames 612 to physiological measurement device 616. Physiological measurement device 616 can be a computing device or other device capable of executing machine learning model 620. In some examples, physiological measurement device 616 can be distributed among multiple computing devices and / or can utilize cloud infrastructure. In some examples, physiological measurement device 616 may include a service, such as a web service, that receives a sequence of frames and provides recovered physiological signals, such as heart rate.
[0079] The physiological measurement device 616 can execute a machine learning model 620 to process the sequence of frames 612. In the example, the machine learning model 620 can utilize model / structure data 624 to create or generate a model structure. The model structure can be the same as or similar to a machine learning structure trained with one or more video sequences. For example, the model / structure data 624, when executed or run by the physiological measurement device 616, can generate a structure similar to... Figure 4 The model structure of the machine learning structure. Model weight data 628 can be used to weight one or more parts or features of the newly created model structure determined during the machine learning process. Thus, the machine learning model 620 can receive a sequence of frames 612 and process the frames to recover or identify physiological signals 632. The physiological measurement device 616 can further process the recovered physiological signals 632 to output or provide physiological measurements or assessments 636, such as pulse rate. In some examples, the physiological assessment may correspond to a similarity measurement to a predicted training label (such as condition). In some examples, the physiological measurement or assessment 636 may be stored in a repository or provided to the subject 604 or the subject 604's caregiver.
[0080] Figure 7 An example of a waveform corresponding to a recovered pulse from a machine learning model (such as machine learning model 620 as an example according to this disclosure) is depicted. For example, one or more physiological signals can be recovered from machine learning model 620 to a recovered physiological signal 632. The example recovered physiological signal may include... Figure 7The signal depicted. That is, a machine learning model (such as machine learning model 620) can be trained using only real human videos; machine learning model 620 can recover waveforms, such as waveform 704, where power spectral analysis indicates the dominant frequency occurs at 92 beats per minute (BPM), as shown in Figure 708. Such recovered waveforms and pulses are consistent with one or more readings of the contact sensor depicted by waveform 720 and the resulting pulse shown in Figure 724. Furthermore, machine learning models (such as machine learning model 620) can use physiologically realistic avatar videos (such as... Figure 1 Those depicted in 122 and 124) and real human videos (such as Figure 1 The machine learning model 620 is trained using the videos depicted. It can recover waveforms, such as waveform 712, where power spectral analysis indicates, as shown in Figure 716, that the dominant frequency occurs at 92 beats per minute (BPM). This recovered waveform and pulse are consistent with one or more readings from the contact sensor depicted by waveform 720 and the resulting pulse shown in Figure 724.
[0081] As another example, a machine learning model (such as machine learning model 620) can be trained using only real human videos; machine learning model 620 can recover waveforms, such as waveform 728, where power spectrum analysis indicates the dominant frequency occurs at 69 beats per minute (BPM), as shown in Figure 732. Such recovered waveforms and pulses do not match the pulses depicted by waveform 736 and one or more readings from the contact sensor, as shown in Figure 740. Furthermore, machine learning models (such as machine learning model 620) can use physiologically realistic avatar videos (such as... Figure 1 (as shown in 122 and 124) and real human videos (such as...) Figure 1 The machine learning model 620 is trained using the waveform (as depicted in Figure 736) where power spectral analysis indicates the dominant frequency occurs at 92 beats per minute (BPM), as shown in Figure 740. This recovered waveform and pulse are consistent with one or more readings from the contact sensor depicted in waveform 744 and the resulting pulse shown in Figure 748. Therefore, by using realistic and physiologically accurate synthetic video, the recovered physiological signal can correspond more accurately, or otherwise better, to the gold standard of non-contact measurement.
[0082] Figure 8A system 800 according to an example of this disclosure is depicted, detailing a physiologically realistic video generator and / or a generator for training a machine learning model with synthetic physiologically realistic videos. That is, a device such as computing device 804 can interact with physiologically realistic video / model generator 824 to retrieve and / or generate one or more of synthetic physiologically realistic video sequences and / or physiological models in real-time or near real-time. For example, a user operating computing device 804 might want training data to train their own machine learning model to recover physiological signals. In some examples, a user operating computing device 804 might want to obtain a machine learning model for integration into a physiological measurement system. As another example, a user operating computing device 804 might want to anonymize existing physiological data (such as physiological waveforms) so that the avatar exhibits physiological characteristics rather than those of an actual human. In some examples, a user operating computing device 804 might want an avatar exhibiting physiological characteristics generated by a specific cinematic effect. Thus, a user can utilize physiologically realistic video / model generator 824 to obtain such synthetic physiologically realistic video sequences and / or models.
[0083] Users can browse one or more of the physiologically realistic avatar video repository 816 and / or physiological model repository 812 via computing device 804 for synthesized physiologically realistic video sequences and / or models. For example, if a user cannot locate the desired synthesized physiologically realistic video sequence and / or model, the user can select one or more parameters via user interface 820. Users can submit such parameters using a submit feature or button 852, causing the parameters to be provided to the physiologically realistic video / model generator 824 via network 808. The physiologically realistic video / model generator 824 can then continue to interact with... Figures 1-3 A consistent approach is used to generate synthetic physiologically realistic video sequences using the physiological shader 858 and frame synthesizer 862 of the physiologically realistic video generator 856. In some examples, the physiologically realistic video / model generator 824 may further utilize the machine learning model 860 to generate one or more physiological models. The physiologically realistic video / model generator 824 may then provide the synthetic physiologically realistic video sequence and / or physiological model to one or more of the repositories 812 or 816 and / or to a user operating the computing device 804 as a synthetic physiologically realistic video sequence 868 and / or physical model 872.
[0084] In some examples, a user operating computing device 804 can provide physiological data 864, causing a synthetic physiologically realistic video sequence based on physiological data 864 to be generated. For example, physiological data 864 can be obtained using gold standards, contact and / or non-contact measurement devices or sensors, or it can be recovered physiological signals, such as recovered physiological data 140. Physiologically realistic video / model generator 824 can generate a synthetic physiologically realistic video sequence based on physiological data 864 and provide the synthetic physiologically realistic video sequence 868 to the user via network 808.
[0085] Figure 9 Details of a method 900 for generating physiologically realistic avatar videos according to an example of this disclosure are described. The general sequence of the steps of method 900 is as follows: Figure 9 As shown. Typically, method 900 begins at 904 and ends at 932. Method 900 may contain more or fewer steps, or the steps may be arranged in a specific order. Figure 9 The steps shown are different. Method 900 can be executed as a set of computer-executable instructions that are executed by a computer system and encoded or stored on a computer-readable medium. Furthermore, method 900 can be executed by gates or circuits associated with a processor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), system-on-a-chip (SoC), or other hardware device. In the following, method 900 will be referred to in conjunction with… Figures 1-8 Explain using descriptions of systems, components, modules, software, data structures, user interfaces, etc.
[0086] The method begins at 904, where the process can proceed to 908. At 908, physiological data can be received. The physiological data may contain one or more signals indicative of a physiological response, condition, or signal. For example, the physiological data may correspond to a blood volume pulse measurement based on a real human record, such as a blood volume pulse waveform. As another example, the physiological data may correspond to respiratory rate / waveform, cardiac condition indicated by a waveform or measurement (such as atrial fibrillation), and / or oxygen saturation level. For example, the physiological data may correspond to a cardiac impaction plethysmography (BCG) and may be a BCG waveform. As another example, the physiological data may be a photoplethysmography waveform. In some examples, the physiological data may be used to assess one or more conditions, such as, but not limited to, peripheral artery disease, Raynaud's phenomenon, systemic sclerosis, and Takayasu arteritis. The physiological data may contain other waveforms, measurements, or other data and may be from different individuals. In some examples, the waveforms may be records of various lengths and sampling rates.
[0087] Method 900 can proceed to 912, where the received physiological data is adjusted by subsurface skin color weights or otherwise modified. When blood flows through the skin, the skin's composition changes, resulting in changes in subsurface color. Therefore, skin color tone changes can be manipulated using subsurface color parameters, including but not limited to the base subsurface skin color, subsurface skin color weights, and subsurface skin scattering parameters. Weights for the subsurface skin color weights can be derived from the absorption spectrum of hemoglobin and typical frequency bands from an example digital camera. Subsurface skin color weights can include weights for one or more color channels and can be applied to the physiological data signal.
[0088] The method can proceed to 916, where the underlying subsurface skin color, which forms the base color beneath the skin, can be modified based on weighted physiological data to obtain the subsurface skin color. In 920, albedo can be selected. Albedo can correspond to a texture map transferred from a high-quality 3D facial scan. Albedo can be randomly selected or selected to represent a specific population. Albedo may lack facial hair, making skin characteristics easily manipulated. Skin type can be randomly selected or selected to represent a specific population. For example, skin type can be selected from six Fitzpatrick skin types. Fitzpatrick skin type (or photo type) depends on the amount of melanin in the skin. In 922, method 900 can generate or otherwise interpret motion variations at least in part due to the physiological data. In some examples, the modeled physiological processes result in color and motion variations; therefore, motion weights (such as motion weight 222) can be applied to the physiological data to interpret pixel movement and pixel translation at least in part due to the physiological data. Method 900 can then proceed to 924, where the physiologically realistic avatar can be rendered based on albedo and subsurface skin color. In some examples, additional parameters (such as appearance parameters, other subsurface skin parameters, motion weights, and extrinsic parameters) may affect the rendering of the avatar. In some examples, the avatar can be rendered by a physiologically realistic shader (such as the physiologically realistic shader 308 described previously). Since the physiological signals received in 908 can be temporal, multiple images of the avatar shifted over time can be rendered.
[0089] Method 900 can proceed to 928, where multiple images of the avatar, time-shifted, can be composited to form a physiologically realistic avatar video of a predetermined length. In some examples, a static or dynamic background can be composited with the rendered avatar. Method 900 can then proceed to 932, where the physiologically realistic avatar video can be stored in a physiologically realistic avatar video repository, such as the physiologically realistic avatar video repository 316 described previously. The physiologically realistic avatar video can be labeled or tagged with training labels before being stored; alternatively or additionally, the physiologically realistic avatar can be stored in a location or repository associated with a specific training label. Examples of training labels include, but are not limited to, blood volume, pulse, and / or peripheral artery disease. Thus, when using the physiologically realistic avatar video to train a machine learning model, the training labels can identify one or more features of the video as training and / or test / validation data. Method 900 can then end at 936.
[0090] Figure 10 Details of a method 1000 for training a machine learning architecture according to an example of this disclosure are described. The general order of the steps in method 1000 is as follows: Figure 10 As shown. Typically, method 1000 begins at step 1004 and ends at step 1020. Method 1000 may contain more or fewer steps, or the order of the steps may be different. Figure 10 The steps are shown in the diagram. Method 1000 can be executed as a set of computer-executable instructions that are executed by a computer system and encoded or stored on a computer-readable medium. Furthermore, method 1000 can be executed by gates or circuits associated with a processor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), system-on-a-chip (SoC), or other hardware device. In the following, method 1000 will be referred to in conjunction with… Figures 1-9 Explain using descriptions of systems, components, modules, software, data structures, user interfaces, etc.
[0091] The method begins at 1004, where the process can proceed to 1008. At 1008, training data can be received. The training data received at 1008 may contain physiologically realistic avatar videos; in some examples, these videos may have already been synthesized according to the previously discussed method 1000. At 1012, one or more videos containing human participants can be received. That is, the machine learning architecture can benefit from training data utilizing both physiologically realistic avatar videos and videos of actual human participants. At 1016, the machine learning architecture can be trained using both types of videos.
[0092] For example, a machine learning architecture can contain two paths, such as Figure 4The discussion focuses on two paths. The first path can be associated with a motion model, and the second path with an appearance model. The architecture of the motion model can contain various layers and hidden units, and can include average pooling and hyperbolic tangent, which can be used as activation functions. The architecture of the appearance model can be the same as or similar to that of the motion model. The motion model allows the machine learning structure to distinguish intensity changes caused by noise, such as motion from subtle feature intensity changes caused by physiological characteristics. Motion representations are computed from the input differences of two consecutive video frames (e.g., C(t) and C(t+1)). The appearance model allows the machine learning structure to learn which regions in an image are likely reliable for computed strong physiological signals (e.g., iPPG signals). The appearance model can generate representations from one or more of the texture and color information of the input video frames. The appearance model guides the motion representations to recover iPPG signals from the various regions contained in the input image and further distinguish them from other noise sources. The appearance model can take frames from a single image or video as input.
[0093] As part of the training process, the recovered physiological signals can be compared with known or valid physiological signals. Once satisfactory accuracy is achieved, the machine learning structure can be output as a machine learning model in 1020, where the structure of the machine learning model can be stored in a model file, and the individual weights in the machine learning model are stored in a location associated with the weight file. Once the model is generated, method 1000 can end in 1024.
[0094] Figure 11 Details of a method 1100 for identifying and / or generating physiologically realistic avatar videos for a requester, according to an example of this disclosure, are described. The general sequence of the steps of method 1100 is as follows: Figure 11 As shown. Typically, method 1100 begins at 1104 and ends at 1132. Method 1100 may contain more or fewer steps, or the steps may be arranged in a specific order. Figure 11 The steps shown are different. Method 1100 can be executed as a set of computer-executable instructions that are executed by a computer system and encoded or stored on a computer-readable medium. Furthermore, method 1100 can be executed by gates or circuits associated with a processor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), system-on-a-chip (SoC), or other hardware device. Hereinafter, method 1100 will be referred to in conjunction with... Figures 1-10 Explain using descriptions of systems, components, modules, software, data structures, user interfaces, etc.
[0095] The method begins at 1104, where the process can proceed to 1108. At 1108, selection of one or more physiological characteristics can be received. For example, a user can interact with a user interface (such as user interface 820) to select the physiological characteristic to be represented by the physiologically realistic avatar. Such characteristics can include conditions, traits, or signals that the avatar wants to represent. As another example, a physiological characteristic could be, for example, pulse rate or, for example, an avatar with atrial fibrillation. At 1112, the user can interact with user interface 820 to select one or more parameters. For example, these parameters can include, but are not limited to, appearance parameters 828, albedo 832, physiologically realistic data signals 836, subsurface skin parameters 840, background 744, and / or other external parameters 848, as previously discussed. Figure 3 and Figure 8 As described. At 1116, physiological data can be received. In the example, the physiological data received at 1116 can be received from a physiological data repository. For example, if a user wishes for their avatar to exhibit a pulse rate of 120 beats per minute, the physiological data corresponding to that pulse rate can be obtained from the repository. In some examples, the user can upload physiological data at 1116. That is, a user operating the computing device can provide physiological data so that the synthesized physiologically realistic video sequence is based on that data.
[0096] Method 1100 can then move to 1120, where the physiologically realistic avatar video clip can be generated based on one or more physiological characteristics and one or more physiological parameters. That is, the physiologically realistic avatar can be generated or rendered in real time, such that the physiological characteristics, parameters, and physiological data are specific to the rendered avatar. In some examples, physiological characteristics cause changes in both color and motion; therefore, motion weights can be applied to the physiological data to account for pixel movement and pixel translation caused at least in part by the physiological data. Multiple images of the avatar can be generated, such that the images can be composited into the physiologically realistic avatar video along with the background. At 1124, the physiologically realistic avatar video can be stored in a physiologically realistic avatar video repository, such as a physiologically realistic avatar video library 128. Parts 1116, 1120, and 1124 of method 1100 can be optional, since instead of generating the physiologically realistic avatar video based on one or more characteristics, parameters, and physiological data, existing physiologically realistic avatar videos that meet user-specified criteria can be located and provided to the requester. Therefore, at 1128, method 1100 can provide the requester with the requested video, either a live video as previously described, or a pre-existing video. Method 1100 can end at 1132.
[0097] Figure 12Details of a method 1200 for recovering physiological signals from a video using a machine learning model trained on a synthetic physiological realism model, according to an example of this disclosure, are described. The general sequence of the steps of method 1200 is as follows: Figure 12 As shown. Typically, method 1200 begins at 1204 and ends at 1232. Method 1200 may contain more or fewer steps, or the order of the steps may be arranged differently. Figure 12 The order shown is different. Method 1200 can be executed as a set of computer-executable instructions that are executed by a computer system and encoded or stored on a computer-readable medium. Furthermore, method 1200 can be executed by gates or circuits associated with a processor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), system-on-a-chip (SoC), or other hardware device. References will be incorporated herein by reference. Figures 1-11 Method 1200 is explained by describing the system, components, modules, software, data structures, user interface, etc.
[0098] The method begins at 1204, where the process can proceed to 1208. At 1208, multiple images can be received. These multiple images may correspond to one or more video frames containing a human subject; in some examples, the multiple images are video clips depicting a human subject. The subject or patient can be a real person and may or may not exhibit one or more physiological characteristics. A camera can be used to capture the multiple images. At 1212, the multiple images can be provided to a physiological measurement device. The physiological measurement device can be a computing device or service (such as a web service) that receives the multiple images and acquires or identifies physiological signals, such as heart rate. At 1216, the physiological measurement device can execute a machine learning model to process the multiple images. In an example, the machine learning model can utilize model / structure data to create or generate a model structure. The model structure can be the same as or similar to a machine learning structure trained with one or more video sequences. For example, the model / structure data, when executed or run by the physiological measurement device, can generate something similar to... Figure 4 The model structure of a machine learning architecture. Model weight data can be used to weight one or more parts or features of a newly created model structure determined during the machine learning process. Therefore, in a 1220 machine learning model, multiple images can be received and processed to recover or identify physiological signals.
[0099] Once the physiological signal has been recovered, the physiological measurement device at 1224 can further process the recovered physiological signal to output or provide a physiological measurement or assessment. The physiological measurement may be a rate, such as pulse rate. In some examples, the physiological assessment may correspond to a similarity measurement to a predictive label (such as condition). In some examples, the physiological measurement or assessment may be output at 1228 and stored in a repository, or provided to the subject or the subject's caregiver.
[0100] Figures 13-15 The associated description provides a discussion of various operating environments in which aspects of this disclosure can be practiced. However, regarding Figures 13-15 The devices and systems shown and discussed are for illustrative purposes and are not intended to limit the wide range of computing device configurations that can be used to practice the aspects of this disclosure described herein.
[0101] Figure 13 This is a block diagram illustrating the physical components (e.g., hardware) of computing device 1300, with which various aspects of this disclosure can be practiced. The computing device components described below are applicable to the computing device described above. In a basic configuration, computing device 1300 may include at least one processing unit 1302 and system memory 1304. Depending on the configuration and type of computing device, system memory 1304 may include, but is not limited to, volatile memory (e.g., random access memory), non-volatile memory (e.g., read-only memory), flash memory, or any combination of such memory.
[0102] System memory 1304 may contain an operating system 1305 and one or more program modules 1306 suitable for running software applications 1307, such as, but not limited to, a machine learning model 1324, a machine learning architecture 1326, and a physiologically realistic avatar video generator 1325. Machine learning model 1324 may, but is not limited to, be related to at least the provisions of this disclosure. Figures 1 to 12 The machine learning models 144 and 442 described are the same or similar. The physiologically realistic avatar video generator 1325 may, but is not limited to, be consistent with at least the provisions of this disclosure. Figures 1 to 12 The physiologically realistic video generator 304 described is the same as or similar to the one described herein. The machine learning structure 1326 may, but is not limited to, those relating to at least the present disclosure. Figures 1 to 12 The end-to-end learning models 136 and 404 described are the same or similar. For example, operating system 1305 may be adapted to control the operation of computing device 1300.
[0103] Furthermore, embodiments of this disclosure can be practiced in conjunction with graphics libraries, other operating systems, or any other applications, and are not limited to any application or system. This basic configuration is in Figure 13 The components within the dashed lines 1308 are shown. The computing device 1300 may have additional features or functions. For example, the computing device 1300 may also include additional data storage devices (removable and / or non-removable), such as disks, optical discs, or magnetic tapes. Such additional storage... Figure 13 The image shows a removable storage device 809 and a non-removable storage device 1310.
[0104] As described above, several program modules and data files can be stored in system memory 1304. When executed on at least one processing unit 1302, program module 1306 can perform processing including, but not limited to, one or more aspects as described herein. Other program modules that can be used according to aspects of this disclosure may include email and contact applications, word processing applications, spreadsheet applications, database applications, PowerPoint presentation applications, drawing or computer-aided applications, and / or one or more components supported by the system described herein.
[0105] Furthermore, embodiments of this disclosure can be practiced in circuits including discrete electronic components, packaged or integrated electronic chips containing logic gates, circuits utilizing microprocessors, or on a single chip containing electronic components or a microprocessor. For example, embodiments of this disclosure can be practiced via a system-on-a-chip (SOC), wherein... Figure 13 Each or more of the components shown can be integrated onto a single integrated circuit. Such a SoC device can include one or more processing units, graphics units, communication units, system virtualization units, and various application functions, all integrated (or “programmed”) onto a chip substrate as a single integrated circuit. When operating via the SoC, the capabilities described herein regarding the client switching protocol can be operated via application-specific logic integrated with other components of the computing device 1300 on the single integrated circuit (chip). Embodiments of this disclosure can also be practiced using other techniques capable of performing logical operations, such as AND, OR, and NOT, including but not limited to mechanical, optical, fluid, and quantum technologies. Furthermore, embodiments of this disclosure can be practiced within a general-purpose computer or in any other circuit or system.
[0106] The computing device 1300 may also have one or more input devices 1312, such as a keyboard, mouse, pen, voice or speech input device, touch or swipe input device, etc. Multiple output devices 1314A (such as a monitor, speaker, printer, etc.) may also be included. An output 1314B corresponding to a virtual display may also be included. The above devices are examples, and other devices may be used. The computing device 1300 may include one or more communication connections 1316 that allow communication with other computing devices 1350. Examples of suitable communication connections 1318 include, but are not limited to, radio frequency (RF) transmitters, receivers, and / or transceiver circuitry; universal serial bus (USB), parallel and / or serial ports.
[0107] The term "computer-readable medium" as used herein can include computer storage media. Computer storage media can include volatile and non-volatile media, removable and non-removable media, such as computer-readable instructions, data structures, or program modules, implemented using any method or technology for storing information. System memory 1304, removable storage device 1309, and non-removable storage device 1310 are examples of computer storage media (e.g., memory storage). Computer storage media can include RAM, ROM, electrically erasable read-only memory (EEPROM), flash memory or other storage technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape, magnetic tape, disk storage or other magnetic storage devices, or any other article of manufacture that can be used to store information and is accessible by computing device 1300. Any such computer storage medium may be part of computing device 1300. Computer storage media does not contain carrier waves or other propagated or modulated data signals.
[0108] Communication media can be implemented by computer-readable instructions, data structures, program modules, or other data in modulated data signals (such as carrier waves or other transmission mechanisms), and can include any information transmission medium. The term "modulated data signal" can describe a signal having one or more characteristics that are set or altered in a manner that encodes information within the signal. By way of example and not limitation, communication media can include wired media, such as wired networks or direct wired connections, and wireless media, such as acoustic, radio frequency (RF), infrared, and other wireless media.
[0109] Figure 14A and 14B The illustration depicts a computing device or mobile computing device 1400, such as a mobile phone, smartphone, wearable computer (e.g., smartwatch), tablet computer, laptop computer, etc., with which aspects of this disclosure can be practiced. Reference Figure 14AOne aspect of the mobile computing device 1400 for implementation is shown. In a basic configuration, the mobile computing device 1400 is a handheld computer with both input and output elements. The mobile computing device 1400 typically includes a display 1405 and one or more input buttons 1410, which allow the user to enter information into the mobile computing device 1400. The display 1405 of the mobile computing device 1400 can also be used as an input device (e.g., a touchscreen display). If included, optional side input elements 1415 allow other user input. The side input elements 1415 can be any other type of rotary switch, button, or manual input element. Alternatively, the mobile computing device 1400 can include more or fewer input elements. For example, in some aspects, the display 1405 may not be a touchscreen. In yet another alternative aspect, the mobile computing device 1400 is a portable telephone system, such as a cellular phone. The mobile computing device 1400 may also include an optional keypad 1435. The optional keypad 1435 can be a physical keypad or a “soft” keypad generated on the touchscreen display. In various aspects, the output elements include a display 1405 for displaying a graphical user interface (GUI), a visual indicator 1431 (e.g., a light-emitting diode), and / or an audio transducer 1425 (e.g., a speaker) for displaying a graphical user interface (GUI). In some aspects, the mobile computing device 1400 includes a vibration transducer for providing haptic feedback to a user. In yet another aspect, the mobile computing device 1400 includes input and / or output ports, such as audio inputs (e.g., a microphone jack), audio outputs (e.g., a headphone jack), and video outputs (e.g., a High Definition Multimedia Interface (HDMI) port), for sending signals to or receiving signals from external sources.
[0110] Figure 14B This is a block diagram illustrating the architecture of one aspect of a computing device, server, or mobile computing device. That is, the mobile computing device 1400 can incorporate aspects of a system (1402) (e.g., architecture). The system 1402 can be implemented as a "smartphone" capable of running one or more applications (e.g., browser, email, calendar, contact manager, messaging client, game, and media client / player). In some aspects, the system 1402 is integrated as a computing device, such as integrating a personal digital assistant (PDA) and a wireless phone.
[0111] One or more applications 1466 may be loaded into memory 1462 and run on or associated with operating system 1464. Examples of applications include telephone dialers, email programs, personal information management (PIM) programs, word processing programs, spreadsheet programs, internet browser programs, messaging programs, and / or one or more components supported by the system described herein. System 1402 also includes a non-volatile storage area 1468 within memory 1462. Non-volatile storage area 1468 may be used to store persistent information that should not be lost when system 1402 is powered off. Applications 1466 may use and store information in non-volatile storage area 1468, such as emails or other information from email applications, such as emails or other messages used by email applications. A synchronization application (not shown) also resides on system 1402 and is programmed to interact with a corresponding synchronization application residing on the host computer to keep the information stored in non-volatile storage area 1468 synchronized with the corresponding information stored on the host computer. It should be understood that other applications can be loaded into memory 1462 and run on the mobile computing device 1400 described herein (e.g., machine learning model 1323 and physiologically realistic avatar video generator 1325, etc.).
[0112] System 1402 has a power supply 1470, which can be implemented as one or more batteries. The power supply 1470 may further include an external power source, such as an AC adapter or an electric docking station for replenishing or recharging the batteries.
[0113] System 1402 may also include a radio interface layer 1472 that performs functions of transmitting and receiving radio frequency communications. Radio interface layer 1472 facilitates wireless connectivity between system 1402 and the "external world" via a communications operator or service provider. Transmissions to and from radio interface layer 1472 are conducted under the control of operating system 1464. In other words, communications received by radio interface layer 1472 can be propagated to application 1466 via operating system 1464, and vice versa.
[0114] A visual indicator 1420 may be used to provide visual notifications, and / or an audio interface 1474 may be used to generate auditory notifications via an audio transducer 1425. In the illustrated configuration, the visual indicator 1420 is a light-emitting diode (LED) and the audio transducer 1425 is a speaker. These devices may be directly coupled to a power supply 1470 such that when activated, they remain on for a duration specified by the notification mechanism, even if the processor 1460 and other components may be turned off to conserve battery power. The LED may be programmed to remain on indefinitely until the user takes action to indicate the device's power-on status. The audio interface 1474 is used to provide and receive audible signals to and from the user. For example, in addition to being coupled to the audio transducer 1425, the audio interface 1474 may also be coupled to a microphone to receive audible input, such as facilitating telephone conversations. According to aspects of this disclosure, the microphone may also be used as an audio sensor to facilitate control of notifications, as will be described below. System 1402 may further include a video interface 1476, which enables an in-vehicle camera to operate to record still images, video streams, etc.
[0115] The mobile computing device 1400 implementing system 1402 may have additional features or functions. For example, the mobile computing device 1400 may also include additional data storage devices (removable and / or non-removable), such as disks, optical discs, or magnetic tapes. Such additional storage... Figure 14B The non-volatile storage region 1468 is illustrated.
[0116] Data / information generated or captured by mobile computing device 1400 and stored via system 1402 can be locally stored on mobile computing device 1400, as described above, or the data can be stored on any number of storage media that can be accessed by the device via radio interface layer 1472 or via a wired connection between mobile computing device 1400 and a separate computing device associated with mobile computing device 1400 (e.g., a server computer in a distributed computing network such as the Internet). It should be understood that such data / information can be accessed via wireless interface layer 1472 or via a distributed computing network through mobile computing device 1400. Similarly, such data / information can be easily transferred between computing devices for storage and use via known data / information transmission and storage devices (including email and collaborative data / information sharing systems).
[0117] Figure 15 The illustration depicts one aspect of a system architecture for processing data received at a computing system from a remote source, such as a personal computer 1504, a tablet computing device 1506, or a mobile computing device 1508 as described above. Content displayed at server device 1502 can be stored in different communication channels or other storage types.
[0118] In some respects, one or more of the machine learning architecture 1526, the machine learning model 1520, and the physiologically realistic avatar video generator 1524 may be employed by the server device 1502. The machine learning model 1520 may, but is not limited to, be related to at least the provisions of this disclosure. Figure 1 to Figure 1 The machine learning models 144 and 442 described in 4 are the same or similar. The physiologically realistic avatar video generator 1524 may, but is not limited to, be similar to at least those disclosed herein. Figure 1 to Figure 1 The physiologically realistic video generator 304 described in 4 is the same as or similar to the one described in this disclosure. The machine learning architecture 1526 may, but is not limited to, those relating to at least the present disclosure. Figure 1 to Figure 1 The end-to-end learning models 136 and 404 described in 4 are the same or similar. Server device 1502 can provide data from or to client computing devices (such as personal computers 1504, tablet computing devices 1506, and / or mobile computing devices 1508 (e.g., smartphones)) via network 1512. As an example, the aforementioned computer system can be implemented in personal computers 1504, tablet computing devices 1506, and / or mobile computing devices 1508 (e.g., smartphones). Any of these embodiments of the computing device can obtain content from storage 1516, in addition to receiving graphical data that can be preprocessed at the graphical origin system or post-processed at the receiving computing system. Content storage may include a physiological model repository 1532, a physiologically realistic avatar video repository 1536, and / or physiological measurements 1540.
[0119] Figure 15 An exemplary mobile computing device 1500 is illustrated, capable of performing one or more aspects of the present disclosure. Furthermore, the aspects and functions described herein can operate on a distributed system (e.g., a cloud-based computing system), where application functions, memory, data storage and retrieval, and various processing functions can remotely operate on each other over a distributed computing network (such as the Internet or an intranet). Various types of user interfaces and information can be displayed via an in-vehicle computing device display or via a remote display unit associated with one or more computing devices. For example, various types of user interfaces and information can be displayed and interacted with on a wall, where various types of user interfaces and information are projected. Interaction with multiple computing systems that can practice embodiments of the invention includes keystroke input, touchscreen input, voice or other audio input, gesture input, wherein the associated computing device is equipped with detection (e.g., camera) functionality for capturing and interpreting user gestures used to control functions of the computing device, etc.
[0120] The phrases “at least one,” “one or more,” “or,” and “and / or” are open-ended expressions that both connect and separate elements in their operation. For example, each of the expressions “at least one of A, B, and C,” “one of A, B, or C,” “one or more of A, B, and C,” “A, B, and / or C,” and “A, B, or C” means A alone, B alone, C alone, A and B together, A and C together, B and B together, or A, B, and C together.
[0121] The term "a" or "an" entity refers to one or more of the entities. Therefore, the terms "a" (or "an"), "one or more," and "at least one" are used interchangeably herein. It should also be noted that the terms "comprising," "including," and "having" are used interchangeably.
[0122] As used herein, the term "automatic" and its variations refer to any process or operation that is typically continuous or semi-continuous, completed without significant human input when it is executed. However, a process or operation can be automatic even if its execution uses significant or insignificant human input, if that input is received prior to its execution. Human input is considered significant if it influences how the process or operation is executed. Human input that consents to the execution of a process or operation is not considered "significant."
[0123] Any of the steps, functions, and operations discussed in this article can be performed continuously and automatically.
[0124] Exemplary systems and methods disclosed herein have been described in conjunction with computing devices. However, to avoid unnecessarily obscuring this disclosure, several known structures and apparatuses have been omitted from the foregoing description. This omission should not be construed as limiting. Specific details have been set forth to provide an understanding of this disclosure. However, it should be understood that this disclosure can be practiced in a variety of ways beyond the specific details herein.
[0125] Furthermore, while the exemplary aspects illustrated herein depict various components of a co-located system, some components of the system may reside in remote portions of a remote, distributed network (e.g., a local area network (LAN) and / or the Internet) or within a dedicated system. Therefore, it should be understood that system components may be combined into one or more devices, such as servers, communication equipment, or co-located on specific nodes of a distributed network, such as analog and / or digital telecommunications networks, packet-switched networks, or circuit-switched networks. As can be understood from the foregoing description, and for computational efficiency reasons, system components can be arranged anywhere within a distributed component network without affecting the operation of the system.
[0126] Furthermore, it should be understood that the various links of the connecting element can be wired or wireless links, or any combination thereof, or any other known or later-developed element(s) capable of supplying data to and / or transmitting data from the connecting element. These wired or wireless links can also be secure links and capable of transmitting encrypted information. For example, the transmission medium used as the link can be any suitable carrier for electrical signals, including coaxial cables, copper wires, and optical fibers, and can take the form of sound waves or light waves, such as those generated during radio wave and infrared data communication.
[0127] While the flowcharts have been discussed and described for specific sequences of events, it should be understood that changes, additions, and omissions to the sequence can occur without substantially affecting the operation of the disclosed configurations and aspects.
[0128] Several variations and modifications of this disclosure may be used. Some features of this disclosure may be provided without providing others.
[0129] In another configuration, the systems and methods of this disclosure may combine a dedicated computer, a programmable microprocessor or microcontroller and peripheral integrated circuit(s) elements, an ASIC or other integrated circuit, a digital signal processor, hardwired electronic or logic circuits (such as discrete component circuits), a programmable logic device or gate array (such as a PLD, PLA, FPGA, PAL), any similar device, etc. Generally, any device(s) capable of implementing the methods shown herein can be used to implement various aspects of this disclosure. Exemplary hardware that can be used with this disclosure includes computers, handheld devices, telephones (e.g., cellular, internet, digital, analog, hybrid, etc.), and other hardware known in the art. Some of these devices include processors (e.g., single or multiple microprocessors), memory, non-volatile memory, input devices, and output devices. Furthermore, alternative software implementations, including but not limited to distributed processing or component / object distributed processing, parallel processing, or virtual machine processing, may also be constructed to implement the methods described herein.
[0130] In another configuration, the disclosed method can be readily implemented using software from an object-oriented or object-based software development environment that provides portable source code usable on various computer or workstation platforms. Alternatively, the disclosed system can be implemented partially or entirely in hardware using standard logic circuitry or very-large-scale integration (VLSI) designs. Whether to implement the system according to this disclosure using software or hardware depends on the system's speed and / or efficiency requirements, specific functionality, and the particular software or hardware system or microprocessor or microcomputer system utilized.
[0131] In yet another configuration, the disclosed methods can be implemented in part in software, which can be stored on a storage medium and executed on a programmed general-purpose computer in cooperation with a controller and memory, a dedicated computer, a microprocessor, etc. In these examples, the systems and methods of this disclosure can be implemented as programs embedded in a personal computer (e.g., applets, etc.). This can be implemented as a resource residing on a server or computer workstation, or as a routine embedded in a dedicated measurement system, system component, etc. The system can also be implemented by physically incorporating the system and / or method into the software and / or hardware system.
[0132] If described herein, this disclosure is not limited to standards and protocols. Other similar standards and protocols not mentioned herein exist and are included in this disclosure. Furthermore, the standards and protocols mentioned herein, as well as other similar standards and protocols not mentioned herein, are regularly superseded by faster or more efficient equivalents having substantially the same functionality. Such alternative standards and protocols having the same functionality are considered equivalents included in this disclosure.
[0133] According to at least one example of this disclosure, a method for generating a video sequence including a physiologically realistic avatar is provided. The method may include receiving an albedo for the avatar, modifying a subsurface skin color associated with the albedo based on physiological data associated with physiological characteristics, rendering the avatar based on the albedo and the modified subsurface skin color, and synthesizing frames of a video, the frames of which contain the avatar.
[0134] According to at least one aspect of the above method, physiological data changes over time, and the method further includes modifying a subsurface skin color associated with albedo based on the physiological data at a first time, rendering an avatar based on the albedo and the modified subsurface skin color associated with the physiological data at the first time, synthesizing frames of a first video, the frames of the first video including the avatar, the avatar being rendered based on the albedo and the modified subsurface skin color associated with the physiological data at the first time, modifying the subsurface skin color associated with the albedo based on the data at a second time, rendering the avatar based on the albedo and the modified subsurface skin color associated with the physiological data at the second time, and synthesizing frames of a second video, the second frame of the video including the avatar, the avatar being rendered based on the albedo and the modified subsurface skin color associated with the physiological data at the second time. According to at least one aspect of the above method, the method includes modifying multiple color channels with a weighting factor specific to the physiological data, and modifying the subsurface skin associated with albedo with the multiple color channels. According to at least one aspect of the above method, the method includes changing the subsurface radius for one or more of the multiple color channels based on a weighting factor specific to the physiological data. According to at least one aspect of the above method, the method includes training a machine learning model with multiple synthesized frames containing an avatar. According to at least one aspect of the above method, the method includes training a machine learning model with multiple videos containing a human subject. According to at least one aspect of the above method, the method includes receiving multiple video frames depicting a human subject, and recovering physiological signals based on the trained machine learning model. According to at least one aspect of the above method, the video frames contain an avatar against a dynamic background. According to at least one aspect of the above method, the method includes receiving physiological data from a requesting entity, synthesizing frames of a video containing an avatar substantially in real time, and providing the video frames to the requesting entity. According to at least one aspect of the above method, the physiological characteristic is blood volume and pulse. According to at least one aspect of the above method, the method includes frames of a synthesized video with training labels specific to physiological characteristics.
[0135] According to at least one example of this disclosure, a system for training a machine learning model using video sequences containing physiologically realistic avatars is provided. The system may include a processor and a memory storing instructions that, when executed by the processor, cause the processor to receive a request from a requesting entity to train a machine learning model to detect physiological characteristics, receive multiple video segments, wherein one or more of the video segments contain synthetic physiologically realistic avatars generated using physiological characteristics, train the machine learning model using the multiple video segments, and provide the trained model to the requesting entity.
[0136] According to at least one aspect of the aforementioned system, the instructions, when executed by a processor, cause the processor to receive a second plurality of video segments, including one or more video segments depicting a human having the said physiological characteristic, and to train a machine learning model using the plurality of video segments and the second plurality of video segments. According to at least one aspect of the aforementioned system, the physiological characteristic is blood volume and pulse. According to at least one aspect of the aforementioned system, one or more video segments are labeled with training tags based on the said physiological characteristic. According to at least one aspect of the aforementioned system, the instructions, when executed by a processor, cause the processor to receive the second video segments, identify the physiological characteristic from the second video segments using the trained model, and provide an assessment of the physiological characteristic to the requesting entity.
[0137] According to at least one example of this disclosure, a computer-readable medium is provided. The computer-readable medium includes instructions that, when executed by a processor, cause the processor to receive a request to recover physiological characteristics from a video clip, obtain a machine learning model trained with training data including physiologically realistic avatars generated from the physiological characteristics, receive the video clip, identify metrics associated with the physiological characteristics from the video clip using the trained machine learning model, and provide an assessment of the physiological characteristics to the requesting entity based on the metrics.
[0138] According to at least one example of the aforementioned computer-readable medium, the instructions, when executed by a processor, cause the processor to receive albedo for an avatar, modify subsurface skin color associated with the albedo based on physiological data associated with physiological characteristics, render the avatar based on the albedo and the modified subsurface skin color, synthesize frames of a video containing the avatar, and train a machine learning model using the frames of the synthesized video. According to at least one example of the aforementioned computer-readable medium, the assessment of the physiological characteristic is pulse rate. According to at least one example of the aforementioned computer-readable medium, the received video clip depicts a human subject.
[0139] At least one aspect of the aforementioned system may include instructions that cause the processor to use a tree-based classifier to identify covariates affecting the quality metric based on features contained in the first and second telemetry data. At least one aspect of the aforementioned system may include instructions that cause the processor to stratify the first and second groups of devices using a subset of the identified covariates that is greater than a threshold. At least one aspect of the aforementioned system may include instructions that cause the processor to provide the quality metric to the display device closest to the predicted quality metric.
[0140] In various configurations and aspects, this disclosure includes components, methods, processes, systems, and / or apparatuses substantially as depicted and described herein, including various combinations, sub-combinations, and subsets thereof. Those skilled in the art, upon understanding this disclosure, will understand how to make and use the systems and methods disclosed herein. In various configurations and aspects, this disclosure includes providing apparatus and processes in the absence of items not depicted and / or described herein, or in its various configurations or aspects, including in the absence of such items that may have been used in previous apparatuses or processes, for example, to improve performance, achieve ease of use, and / or reduce implementation costs.
Claims
1. A method for generating a video sequence containing a physiologically realistic avatar, the method comprising: Receive albedo for an avatar, wherein the albedo characterization includes a texture map of skin pixels associated with the avatar; The subsurface skin color associated with the skin pixel of the albedo is modified based on the subsurface skin color weight applied to physiological data associated with physiological characteristics, wherein the subsurface skin color weight is based on the absorption spectrum associated with physiological data of one or more color channels; The avatar is rendered based on the albedo and the modified subsurface skin color; and Frames of a composite video, wherein the frames of the video contain the avatar.
2. The method according to claim 1, wherein the physiological data changes over time, the method further includes: The subsurface skin color associated with the albedo is modified in the first instance based on both the baseline subsurface skin color and the subsurface skin color weight applied to the physiological data; The avatar is rendered at the first time based on the albedo and the modified subsurface skin color associated with the physiological data; The first frame of the synthesized video contains the avatar rendered at the first time based on the albedo and the modified subsurface skin color associated with the physiological data; The subsurface skin color associated with the albedo is modified based on the physiological data at a second time point; The avatar is rendered at the second time based on the albedo and the modified subsurface skin color associated with the physiological data; and The second frame of the synthesized video contains the avatar rendered at the second time based on the albedo and the modified subsurface skin color associated with the physiological data.
3. The method according to claim 1, further comprising: Modify the physiological data using a weighting factor specific to the physiological data; as well as The subsurface skin associated with the albedo is modified using the modified physiological data.
4. The method of claim 3, further comprising changing the subsurface radius for the one or more color channels based on the weighting factor for the physiological data.
5. The method of claim 1, further comprising training a machine learning model using a plurality of synthesized frames containing the avatar.
6. The method of claim 5, further comprising training the machine learning model with a plurality of videos including human subjects.
7. The method according to claim 6, further comprising: Receive multiple video frames depicting human subjects; as well as Physiological signals are recovered based on the trained machine learning model.
8. The method of claim 1, wherein the frames of the video include the avatar in front of the dynamic background.
9. The method according to claim 1, further comprising: Receive the physiological data from the requesting entity; Frames of the video containing the avatar are synthesized in near real-time; as well as Provide the requesting entity with frames of the video.
10. The method of claim 1, wherein the physiological characteristic is blood volume and pulse.
11. The method of claim 1, further comprising labeling video segments containing frames of the synthesized video with training labels for the physiological characteristics.
12. A system for training a machine learning model using video sequences containing physiologically realistic avatars, the system comprising: processor; as well as A memory storing instructions that, when executed by the processor, cause the processor to: Receive a request from the requesting entity to detect physiological characteristics using the trained machine learning model; Receive a plurality of video clips, wherein one or more of the plurality of video clips contain a synthetic physiologically realistic avatar generated using the physiological characteristics, wherein the albedo characterization for the physiologically realistic avatar includes a texture map of skin pixels associated with the avatar, and wherein the subsurface skin color associated with the skin pixels of the albedo is modified based on a subsurface skin color weight applied to physiological data associated with the physiological characteristics, and wherein the subsurface skin color weight is based on an absorption spectrum associated with physiological data of one or more color channels; The machine learning model is trained using the multiple video clips; as well as The trained model is provided to the requesting entity.
13. The system of claim 12, further comprising instructions that, when executed by the processor, cause the processor to: Receive a second plurality of video clips, wherein one or more of the second plurality of video clips depict a human having the said physiological characteristics; and The machine learning model is trained using the plurality of video clips and the second plurality of video clips.
14. The system of claim 12, wherein the physiological characteristic is blood volume and pulse.
15. The system of claim 12, wherein one or more video segments of the plurality of video segments are labeled with trained tags based on the physiological characteristics.
16. The system of claim 12, further comprising instructions that, when executed by the processor, cause the processor to: Receive the second video clip; Using the trained model, physiological characteristics are identified from the second video segment; and Provide the requesting entity with an assessment of the physiological characteristics.
17. A computer-readable medium comprising instructions that, when executed by a processor, cause the processor to: Receive requests to recover physiological characteristics from video clips; A machine learning model trained with training data is obtained, the training data including a synthetic physiologically realistic avatar generated with the physiological characteristics and an avatar corresponding to the synthetic physiologically realistic avatar, wherein the albedo characterization for the physiologically realistic avatar includes a texture map of skin pixels associated with the avatar, and wherein the subsurface skin color associated with the skin pixels of the albedo is modified based on the subsurface skin color applied to the physiological data associated with the physiological characteristics according to the subsurface skin color weight, and wherein the subsurface skin color weight is based on the absorption spectrum associated with the physiological data of one or more color channels; Receive the video clip; The trained machine learning model is used to identify metrics associated with the physiological characteristics from the video clips; as well as An assessment of the physiological characteristics is provided to the requesting entity based on the metric.
18. The computer-readable medium of claim 17, wherein the instructions, when executed by the processor, cause the processor to: Receive the albedo for the avatar; The avatar is rendered based on the albedo and the modified subsurface skin color; Frames of a composite video, wherein the frames of the video contain the avatar; and The machine learning model is trained using frames from the synthesized video.
19. The computer-readable medium of claim 17, wherein the assessment of the physiological characteristic is pulse rate.
20. The computer-readable medium of claim 17, wherein the received video segment depicts a human subject.
Citation Information
Patent Citations
Image processing apparatus and image processing method
US20150154790A1
Self-adaptive matrix completion for heart rate estimation from face videos under realistic conditions
US20170367590A1
Avatar digitization from a single image for real-time rendering
US20180374242A1
Augmenting Virtual Reality Video Games With Friend Avatars
US20190099675A1