Photorealistic talking faces from audio
By using machine learning models to predict the geometric shapes and textures of 3D speaking faces based on audio signals, the problem of lack of fidelity and personalization of 3D face generation in the prior art is solved, and high fidelity and personalized 3D face generation is achieved, which is suitable for a variety of application scenarios.
Patent Information
- Application Number
- CN202180011913.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-01-29
- Filing Date
- 2021-01-29
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-01-29
AI Technical Summary
The prior art is difficult to generate high fidelity 3D speaking faces, especially when only audio input is conditioned, and existing methods generally lack geometric information and personalization, and fixed viewpoint 2D methods limit applications.
Using machine-learning facial geometry prediction model and facial texture prediction model, facial geometry and texture are predicted based on audio signals, and combined them to generate a three-dimensional facial mesh model.
It realizes the generation of photo-level realistic 3D speaking faces from audio signals, improving the realistic and personalization of the generated faces, and is suitable for a variety of application scenarios such as VR, games and video editing.
Smart Images

Figure CN115004236B_ABST
Abstract
Description
[0001] Related Applications
[0002] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 62 / 967,335, filed on January 29, 2020, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present disclosure generally relates to synthesizing images of speaking faces from audio signals. More specifically, the present disclosure relates to a framework for generating photorealistic three-dimensional (3D) speaking faces conditioned in some examples solely on audio input, and related methods for optionally inserting the generated faces into existing videos or virtual environments. Background Art
[0004] "Talking head" videos consisting of close-ups of a talking person are widely used in news broadcasts, video blogs, online courses, etc. Other modalities with similar frame compositions focused on faces include face-to-face live chats and 3D game avatars.
[0005] The importance of talking head synthesis has led to a variety of approaches in the research literature. Many recent techniques use methods that regress facial motion from audio and use this method to warp a single reference image of the desired subject. These methods can inherit the realism of the reference photo. However, the results can lack geometry information and personalization, and do not necessarily reproduce 3D facial articulation and appearance with high fidelity. They also typically do not incorporate lighting changes, and fixed viewpoint 2D approaches limit possible applications.
[0006] Another body of research predicts 3D facial meshes from audio. These methods are directly applicable to VR, games, and other applications that require dynamic viewpoints, and dynamic lighting is also easy to implement. However, visual realism is often limited by what can be obtained with real-time 3D rendering, so only gaming-quality results are achieved.
[0007] Other recent papers have proposed techniques for generating talking head videos by transferring facial features (such as landmarks or fused shape parameters) from videos of different narrators to a target subject. These techniques produce particularly impressive results, however they require videos of proxy actors. Furthermore, while text-based editing does not require human actors, it relies on the availability of temporally aligned transcripts. Summary of the invention
[0008] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or may be learned from the description, or may be learned through practice of the embodiments.
[0009] One example aspect of the present disclosure relates to a computing system for generating a speaking face from an audio signal. The computing system may include one or more processors and one or more non-transitory computer-readable media that collectively store: a machine-learned facial geometry prediction model configured to predict facial geometry based on data describing an audio signal including speech; a machine-learned facial texture prediction model configured to predict facial texture based on data describing the audio signal including the speech; and instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations may include obtaining data describing an audio signal including speech; using the machine-learned facial geometry prediction model to predict the facial geometry based at least in part on the data describing the audio signal; using the machine-learned facial texture prediction model to predict the facial texture based at least in part on the data describing the audio signal; and combining the facial geometry with the facial texture to generate a three-dimensional facial mesh model.
[0010] These and other features, aspects and advantages of various embodiments of the present disclosure will be better understood with reference to the following description and appended claims.The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and, together with the description, serve to explain relevant principles. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] A detailed discussion of embodiments for those of ordinary skill in the art is set forth in the specification with reference to the accompanying drawings, in which:
[0012] Figure 1A A block diagram of an example system for generating a dynamically textured 3D surface mesh from audio is shown according to an example embodiment of the present disclosure.
[0013] Figure 1B A block diagram of an example system for inserting a generated facial mesh into a target video to create a synthetic talking head video from a new audio input is shown according to an example embodiment of the present disclosure.
[0014] Figure 2 Depicted are results of an example technique for normalizing training data according to an example embodiment of the present disclosure.
[0015] Figure 3A A block diagram of an example system for training a facial geometry prediction model for machine learning according to an example embodiment of the present disclosure is shown.
[0016] Figure 3B A block diagram of an example system for training a facial texture prediction model for machine learning according to an example embodiment of the present disclosure is shown.
[0017] Figure 4 Depicted is an example of a talking face integrated into a virtual environment according to an example embodiment of the present disclosure.
[0018] Figure 5A -C depicts a block diagram of an example computing system according to an example embodiment of the present disclosure.
[0019] Figure 6 Depicted is a flow chart of an example method of generating a speaking face from audio according to an example embodiment of the present disclosure.
[0020] Fig. 7A A flow chart is depicted of an example method of training a machine-learned facial geometry prediction model according to an example embodiment of the present disclosure.
[0021] Figure 7B Depicted is a flow chart of an example method for training a machine-learned facial texture prediction model according to an example embodiment of the present disclosure.
[0022] Reference numerals repeated in various figures are intended to identify like features in the various embodiments. DETAILED DESCRIPTION
[0023] In general, the present disclosure relates to systems and methods for generating a photorealistic 3D speaking face (e.g., a 3D textured mesh model of a face) conditioned in some embodiments solely on audio input. Specifically, some example embodiments include and use a machine learning facial geometry prediction model to predict facial geometry based on an audio signal, and include and use a machine learning facial texture prediction model to predict facial texture based on an audio signal. The predicted geometry and texture may be combined to obtain a 3D mesh model of the face. Additionally, the present disclosure provides associated methods for inserting the generated face into an existing video or virtual environment.
[0024] In some embodiments, the machine learning model used in the present disclosure can be trained on video data, including, for example, by decomposing faces from videos into a normalized space that decouples 3D geometry, head pose, and texture. This allows the prediction problem to be separated into regression on 3D facial shape and corresponding 2D texture atlas, as described above.
[0025] Additional aspects of the disclosure provide for improved quality of generated faces. As one example, to stabilize temporal dynamics, some embodiments of the disclosure utilize autoregressive methods that condition the model on its previous visual state. As another example, facial lighting may be performed by the model using audio-independent 3D texture normalization. These techniques significantly improve the realism of the generated sequences, providing results that are superior to prior art lip sync systems.
[0026] There are a large number of different uses or applications for the generated speaking faces. As examples, applications enabled by the proposed framework include: personalized and voice-controlled photorealistic talking games or virtual reality (VR) avatars; automatic translation of videos into different languages (e.g., lip syncing for translation and dubbing of videos in new languages); general video editing (e.g., inserting new audio / voice content in educational videos); and compression in multimedia communications (by sending only the audio signal (and in some embodiments, reference images) and reconstructing visual aspects from the audio when needed). Thus, in some example uses, the 3D information can be used to essentially edit 2D videos, producing photorealistic results. Alternatively, the 3D mesh can be used for 3D games and VR.
[0027] More specifically, aspects of the present disclosure utilize machine learning techniques to train models that predict the shape and appearance of faces from instantaneous audio input. These models provide a practical framework applicable to a variety of scenarios and also produce sufficiently realistic results for real-world applications. To this end, various example implementations exhibit the following optional features:
[0028] Audio as driving input: Some embodiments of the present disclosure use audio as driving input, which gives the flexibility to use the proposed technology with spoken input or synthesized text-to-speech (TTS) audio. Using audio directly also simplifies data preparation and model architecture, because synchronized audio and video frame pairs can be used directly as training data without any additional processing. On the other hand, using text, phonemes, and visemes requires additional feature extraction and time alignment steps.
[0029] 3D Decomposition: A 3D face detector (an example is described in Kartynnik et al., Real-time facial surface geometry from monocular video on mobile GPUs, Third Workshop on Computer Vision for AR / VR, Long Beach, CA, 2019) is used to obtain the pose and triangular mesh of the face of the speaker in the video. This information enables the decomposition of the face into a normalized 3D mesh and texture atlas, thereby decoupling head pose from speech-induced facial deformations such as lip movement and teeth / tongue appearance. Models can be trained to predict facial geometry and texture from audio in this normalized domain. This approach has two benefits: (1) the number of degrees of freedom that the model must handle is greatly reduced (for speech-related features), which allows plausible models to be generated even from relatively short videos. (2) The model predicts a full 3D speaking face rather than just a 2D image, which extends its applicability beyond video to games and VR, while also improving the quality and flexibility of video resynthesis.
[0030] Personalized Models: Instead of building a single generic model to be applied across different people, personalized speaker-specific models can be trained. While generic models have their advantages, such as easy reuse, they require larger training sets to fully capture the individual motion style of every possible speaker. On the other hand, personalized models can easily incorporate such person-specific traits by learning the model from videos of a specific speaker during training. Note that once trained, such a model can still be used across different videos of the same speaker.
[0031] Temporally consistent photorealistic synthesis: Example implementations include model architectures that use an encoder-decoder framework that computes embeddings from audio spectrograms and decodes them into 3D geometry and texture. In one example, a facial geometry prediction model can predict face geometry, which can be represented as mesh vertex deformations relative to a reference mesh. Similarly, a facial texture prediction model can predict facial appearance around a lip region, which can be represented as a difference map to a reference texture atlas.
[0032] In some embodiments, to further achieve temporal smoothing, an autoregressive framework can be used that adjusts texture generation on both the audio as well as previously generated texture outputs, resulting in a visually stable sequence. Additionally, when resynthesizing a video by blending the predicted face into the target video, it is important to be consistent with the target face lighting. In some embodiments, this can be achieved by incorporating a 3D normalized fixed texture atlas into the model(s), which is uncorrelated with the audio signal and acts as a proxy for the instantaneous lighting.
[0033] The systems and methods of the present disclosure provide many technical effects and benefits. One example technical effect is the ability to convert arbitrary talking head video clips into a normalized space that disentangles pose, geometry, and texture, which simplifies model architecture and training and enables general high-quality results even with limited training data.
[0034] Another example technical effect is a novel method of capturing facial illumination via audio-independent 3D texture normalization and an autoregressive texture prediction model for temporally smooth video synthesis. Thus, the techniques described herein enable the generation of significantly more realistic images of speaking faces from audio.
[0035] Additional example technical effects are an end-to-end framework for training speaker-specific audio-to-face models that can be learned from a single video of the subject, and alignment, blending, and re-rendering techniques for employing them in video editing, translation, and 3D environments. The result is a photorealistic video or 3D face driven solely by audio.
[0036] Another example technical effect and benefit provided by the techniques described herein is the ability to "compress" a video of a speaker into only an audio signal while still being able to recreate a photorealistic, representation of the visual aspects of the video. Specifically, a video may contain both audio data and visual data. Because the techniques of the present disclosure enable a photorealistic image of a speaking face to be (re)created from audio alone, the video may be compressed by maintaining only the audio portion of the video (potentially along with a small number (e.g., 1) of reference images), which will greatly reduce the amount of data required to store and / or transmit the video. Then, when a visual image of a speaking face is desired, the techniques described herein may be employed to create an image from the audio signal. In this way, the amount of data required to be able to store and / or transmit a video of a speaking face may be significantly reduced. For example, this compression scheme may be of great benefit in video conferencing / chatting use cases, particularly where network bandwidth is limited.
[0037] US Provisional Patent Application No. 62 / 967,335, which is incorporated into and forms a part of the present disclosure, describes exemplary embodiments and experimental uses of the systems and methods described herein.
[0038] Referring now to the accompanying drawings, example embodiments of the present disclosure will be discussed in greater detail.
[0039] Example technique for generating talking faces
[0040] This section describes an example method for predicting a dynamic 3D face model from audio input. This section first discusses an example method for extracting training data from an input video (or multiple), and then describes in detail an example neural network architecture and training method for predicting 3D geometry and associated texture.
[0041] In some embodiments, the audio channels from the training videos can be extracted and transformed into frequency domain spectrograms. For example, these audio spectrograms can be calculated using a short-time Fourier transform (STFT) with a Hann window function with a sliding window of 30ms width and 10ms interval. These STFTs can be aligned with the video frames and stacked across time to create a 256×96 complex spectrogram image centered on each video frame. One or more machine learning models can then predict the facial geometry and texture of each frame based on the audio spectrograms.
[0042] To detect faces in training videos and obtain 3D facial features, a facial landmark detector can be used. Various facial landmark detectors (also called three-dimensional face detectors) are known and available in the art. An example of a facial landmark detector is described in Real-time facial surface geometry from monocular video on mobile GPUs by Kartynnik et al. at the 3rd AR / VR Computer Vision Workshop in Long Beach, California in 2019. This video-based face tracker detects 468 facial features in 3D, with a depth (Z) component hallucinated based on deep learning; these are interchangeably referred to as features or vertices. Some embodiments of the present disclosure define a fixed, predefined triangulation of these features, and represent any dynamic changes in facial geometry entirely through mesh vertex displacements rather than through mesh topology changes.
[0043] Example Techniques for Normalizing Training Data
[0044] This section describes an example method for normalizing input facial data. An example goal is to remove the effects of head movement, and work with normalized facial geometry and texture. Both training and inference can be done in this normalized space, which greatly reduces the degrees of freedom that the model must handle, and as shown in U.S. Provisional Patent Application No. 62 / 967,335, a few minutes (typically 2-5 minutes) of video footage of the target person is often sufficient to train the model to achieve high-quality results.
[0045] Example pose normalization
[0046] First, a frame of the input video can be selected as a reference frame, and its corresponding 3D facial landmarks can be selected as reference points. The choice of frame is not critical; any frame in which the face is sufficiently frontal and the resolution is acceptable is suitable. Using the reference points, a reference cylindrical coordinate system with a vertical axis can be defined so that most facial points are equidistant from the axis. The face size can then be normalized so that the average distance to the axis is equal to 1. The facial points can be projected onto this reference cylinder, creating a 2D map of the surface of the reference face, which can be used to "unwrap" its texture.
[0047] Next, for each frame of the training video, a 3D facial point in the upper, more rigid part of the face can be selected and aligned with the corresponding point in the normalized reference. As an example, Umeyama's algorithm (Shinji Umeyama, Least-squares estimation of transformation parameters between two point patterns, IEEE Trans. Pattern Anal. Mach. Intell., 13(4): 376-380, 1991) can be used to estimate the rotation R, translation t, and scale c in 3D. The points p used for tracking provide registered, normalized 3D facial points suitable for training facial geometry prediction models.
[0048] In some embodiments, to train a texture model, these normalized points, now registered to a reference cylindrical texture domain, can be used to create two projections of the texture for each face: (a) a “moving atlas,” created by projecting the moving normalized points onto a reference cylinder as texture coordinates and rendering the associated triangles in 2D; thus, the mouth texture resembles a frontal view, where the facial features move with speech; and (b) a “fixed atlas,” created by mapping each video triangle texture to a corresponding reference triangle using the corresponding reference triangle's texture coordinates, so the facial features are frozen in the positions defined by the reference.
[0049] Figure 2 shows the effect of this normalization; as shown, the head pose is removed. While the moving atlas (third column) is more suitable for training lip shape and internal mouth appearance as a function of speech, the fixed atlas (fourth column) is useful for extracting lighting information independently of speech, because the positions of changing features such as the mouth and eyes are fixed and can be easily masked out for lighting extraction.
[0050] Example Light Normalization
[0051] Another example aspect involves normalizing the frontalized texture atlas to remove illumination variations caused primarily by head motion or changing illumination. An example illumination normalization algorithm of the present disclosure works in two stages. It first exploits facial symmetry to spatially normalize a reference atlas R, removing specular reflections and illumination variations that extend across the face. It then performs temporal normalization across video frames, which transforms the atlas F of each frame to match the illumination of R. The resulting atlas has a more uniform albedo-like appearance that is consistent across frames.
[0052] The temporal normalization algorithm is described first, as it is a core component that is also used during spatial normalization. The algorithm may assume that the two textures F and R are pre-aligned geometrically. However, any non-rigid facial motion, such as from speech, may result in different texture coordinates and, therefore, misalignment between R and F. Therefore, a computing system implementing the algorithm first warps R to align it with the texture coordinates of F, employing the same triangle-based warping algorithm used for front-endization.
[0053] Assuming R and F are aligned, the computing system implementing the algorithm can estimate a mapping of the transformation F to match the illumination of R. This mapping can include a smooth multiplicative pixel-by-pixel gain G in the luminance domain, followed by a global channel-by-channel gain and bias mapping {a, b} in the RGB domain. The resulting normalized texture F can be obtained by the following steps: n :
[0054] (1)(F y ; F u ; F v ) = RGBtoYUV(F);
[0055] (2)
[0056] (3)
[0057] (4)F n =aF l +b;
[0058] Gain Estimation: To estimate the gain G, observe that a pair of corresponding pixels at the same position k in F and R should have the same underlying appearance, modulo any changes in illumination, since they are geometrically aligned. If fully satisfied, this albedo constancy assumption yields a gain G at pixel k k =R k / F k Note, however, that (a) G is a smoothly varying illumination map, and (b) albedo constancy may occasionally be violated, for example in non-skin pixels such as the mouth, eyes, and nostrils, or in cases of sharp skin deformations, such as nasolabial folds. In some embodiments, this can be accomplished by first filtering the image in a larger patch p centered around k. k The example embodiments of the present disclosure may estimate G k Formulated to minimize the error:
[0059]
[0060] Where W is a per-pixel weight image. Example embodiments may use iterative reweighted least squares (IRLS) to solve the error. In particular, example embodiments may uniformly initialize the weights and then update them after each (i-th) iteration as:
[0061]
[0062] where T is the temperature parameter. The weights and gains can converge in 5-10 iterations; for a 256×256 atlas, some embodiments use T=0:1 and a patch size of 16×16 pixels. In some embodiments, with a large error E k Pixels with a low weight can receive a low weight and have their gain value implicitly interpolated from neighboring pixels with higher weights.
[0063] In some embodiments, to estimate the global color transform {a, b} in closed form, the computing system may minimize ∑ k W k ‖R k -aF k -b‖ 2 , where W k are now fixed to the weights estimated above.
[0064] Reference Normalization: This section discusses how to exploit facial symmetry to spatially normalize the reference atlas. Some example implementations first use the above algorithm to estimate the gain G between the reference R and its mirror image R′. mThis gain represents the illumination variation between the left and right halves of the face. To obtain a reference with uniform illumination, the computing system may calculate the symmetrization gain G s =max(G m ,G m′ ), where G m′ It's G m The normalized reference is then G m Note that the weighting scheme makes the technique robust to inherent asymmetries on the face, since any inconsistent pairs of pixels will be weighted downward during gain estimation, thereby preserving those asymmetries.
[0065] Specular Removal: Some example implementations remove specular reflections from faces before normalizing the reference and video frames, as they are not properly modeled as multiplicative gains and also result in duplicate specular reflections on the reference due to symmetrization. Some example implementations model specular image formation as:
[0066] I=α+(1-α)*I c
[0067] where I is the observed image, α is the specular alpha map, and I c is the underlying clean image without specular surfaces. Some example embodiments first compute a mask, where α>0, as pixels whose minimum value across RGB channels in smoothed I exceeds ninety percent of the intensity across all skin pixels in I. Some example implementations use the face mesh topology to identify skin pixels and limit computation to skin pixels. Then, some example embodiments estimate a pseudo-clean image by hole-filling masked pixels from neighboring pixels and use it to estimate Then, the final clean image is I c =(I-α) / (1-α). Note that the soft alpha computation gracefully handles any erroneous overestimation of the specular mask.
[0068] Example Techniques for Audio to Facial Geometry Synthesis
[0069] Some example embodiments of the present disclosure use complex Fourier spectrograms directly as input, thus simplifying the overall algorithm. Specifically, in some example embodiments, the time-shifted complex spectrogram can be represented as a 256×96×2 (frequency×time×real / imaginary) input vector of a 12-layer deep encoder network, where the first 6 layers apply 1D convolutions in frequency (kernel 3×1, stride 2×1), and the subsequent 6 layers apply 1D convolutions in time (kernel 1×3, stride 1×2), all with leaky ReLU activations, intuitively corresponding to phoneme detection and excitation, respectively. The resulting latent space has 256 dimensions. In some embodiments, an additional single dimension from a blink detector can be added to enable blinks to be detected during training and generated on demand during inference. This is followed by a decoder, and an example decoder can include two fully connected layers with 150 and 1404 units and linear excitation. These can be considered as mappings of speech to linear "mixed shape" facial representations with 468 vertices (1404=468×3 coordinates). Some example embodiments also include a dropout layer between each of the above layers. In some embodiments, the last layer may be initialized using PCA on the vertex training data. An example loss function includes L2 vertex position loss; regularization loss; and / or speed loss.
[0070] Example Techniques for Audio to Texture Synthesis
[0071] This section describes an example framework for learning a function G that maps from a domain S of audio spectrograms to a domain T of moving texture atlas images; G: S → T. In some embodiments, for the purpose of texture prediction, the atlas can be cropped to a region around the lips (e.g., to a 128×128 region), and references to texture in this section mean the cropped atlas. Figure 3B An example of a texture model and training pipeline is shown.
[0072] The input at time t is a complex spectrum, And the output is the difference map Δ t , which is added to the reference atlas I r , to obtain the predicted texture atlas
[0073] Some embodiments of the present disclosure follow an encoder-decoder architecture for implementing G(·). First, the spectrogram can be processed through a series of convolutional layers to produce a latent code, Where N L is the latent code dimension. Next, the latent code is distributed in space and progressively upsampled using convolutional and interpolation layers to generate a textured output. The model(s) implementing G can be trained to minimize the combined loss, R = R pix +αR mom, which consists of:
[0074]
[0075] Among them, A t corresponds to S t is the ground truth value of , and d is the pixel-level distance metric, and
[0076]
[0077] where μ(·) and σ(·) are the mean and standard deviation, and and is obtained by applying a binary mask M to the corresponding atlas, which zeros out the mouth region, leaving only skin pixels.
[0078] Pixel loss R pix Aims to maintain pixel-level similarity between the predicted texture and the ground truth texture. Example variants of d(·) can include l1 loss, structural similarity loss (SSIM), and gradient difference loss (GDL) (Mathieu et al., Deep multi-scale video prediction beyond mean square error, ICLR, 2016).
[0079] Moment loss term R mom Encourage the first and second moments of the distribution of skin pixels to match. A soft constraint is imposed to respect the overall illumination of the reference frame and make the training less sensitive to illumination changes over time. Masking the mouth region ensures that appearance changes of the oral cavity due to speech do not affect the moment computation.
[0080] Another example aspect relates to a blended shape decoder. For example, to animate CGI characters using audio, some example embodiments may optionally include another decoder in the network that predicts the blended shape coefficients B in addition to the geometry and texture. t For training, one can start from vertices V by fitting the vertices to an existing blend shape basis via optimization or using a pre-trained model t These blend shapes are derived. Some example embodiments may use a single fully connected layer to extract the audio code from Prediction coefficient And use l1 loss to train it to encourage sparse coefficients.
[0081] Example Technique for Autoregressive Texture Synthesis
[0082] Predicting speaking faces from audio can suffer from ambiguities caused by changes in facial expressions during speech or even during silence. In the latter case, for example, the model may map subtle noise in the audio channel to different expressions, resulting in disturbing jitter artifacts.
[0083] Although some embodiments of the present disclosure do not explicitly model facial expressions, this problem can be alleviated by incorporating memory into the network. The current output of the network (at time t) can be not only t is conditioned on, and can be based on the predicted map generated at the previous time step is adjusted. is encoded as a latent code, For example, a cascade of 3×3 convolutions with a stride of 2 pixels is used. and Can be combined and passed to the decoder network to generate the current texture
[0084] Note that in some cases, the previously predicted atlas is not available during training unless it is modeled as a true recurrent network. However, the network can be satisfactorily trained by using a technique called "teacher forcing", where the ground truth atlas from previous frames is used as prediction input during training. This autoregressive (AR) approach significantly improves the temporal consistency of the synthesized results.
[0085] Example Technique for Joint Texture and Spectrogram Reconstruction
[0086] Some example implementations of the framework described so far do not explicitly enforce the ability to reconstruct the input spectrogram from the latent domain. While such a constraint is not strictly required for the inference of lip shape, it can help regularization and generalization by forcing the latent domain to span the manifold of valid spectrograms. To achieve this, some embodiments of the present disclosure include an additional audio decoder that extracts the audio from the latent domain for generating the lip shape. The same shared latent code Reconstruct the input spectrogram. About the predicted spectrogram The additional autoencoder loss R ae Given by:
[0087]
[0088] Example Techniques for Matching Target Lighting
[0089] In order to blend the synthesized texture back into the target video (see Section 3.5), it is desired that the synthesis is consistent with the illumination of the target face. The function mapping G:S→T does not contain any such illumination information. The moment loss R momA soft constraint is imposed to respect the global illumination of the reference frame. However, the instantaneous illumination on the target face may differ significantly from the reference and also vary over time. This may lead to inconsistent results even when using advanced techniques such as Poisson blending (Perez et al., Poisson image editing, ACM Trans. Graph., 22(3):313-318, July 2003).
[0090] This problem can be solved by using a (i.e. uncropped) fixed atlas. Solved as a proxy light map. Similar to the moment loss calculation, it can mask the The eye and mouth areas are removed to leave only skin pixels. The intensity of the skin pixels on is independent of the input spectrogram and varies mainly due to illumination or occlusion. where M is a binary mask that encodes a measure of instantaneous illumination. Therefore, it can be called an illumination map. Next, we use the illumination encoder network E light right Encoding to generate light codes, Note that in some embodiments, Before the light is fed into the network, it can be The masked reference atlas is subtracted from the original image to treat the reference as neutral (zero) illumination.
[0091] In some embodiments, instead of or in addition to a lighting map, a transformation matrix can be used as a proxy for lighting.
[0092] Finally, all three potential codes (spectrum diagram), (previously predicted map) and (illumination) can be combined and passed to the joint vision decoder, such as Figure 3B As shown, to generate the output texture. The entire framework can be trained end-to-end using the combined loss:
[0093] R=R pix +α1R mom +α2R ae ,(4)
[0094] Among them, α1 and α2 control the importance of moment loss and spectrogram autoencoder loss respectively.
[0095] Example technique for 3D meshes based on predicted geometry and texture
[0096] Previous sections have detailed examples of how both texture and geometry can be predicted. However, since the predicted texture is a “moving atlas”, i.e. a projection onto a reference cylinder, it will typically have to be back-projected onto the actual mesh in order to use it for the 3D head model. Fortunately, this can be achieved by simply projecting the corresponding predicted vertices onto the reference cylinder and using their 2D positions as new texture coordinates, without any resampling. Note that using a moving atlas plus reprojection has two additional advantages: (a) this can mask small differences between predicted vertices and predicted texture; and (b) this leads to a more uniform texture resolution across the mesh, since the sizes of the triangles in the synthesized atlas closely correspond to their surface area in the mesh. Combined with the predefined triangle topology, the result is a fully textured 3D face mesh driven by the audio input, such as Figure 1A In some embodiments, the input audio source may be encoded into an encoded representation using an audio encoder before 2D texture prediction and 3D vertex prediction.
[0097] Example technique for inserting predicted face meshes into videos
[0098] The normalized transformation from video to reference is reversible and can therefore be used to insert the audio-generated face into the target video, thereby synthesizing talking head videos, e.g. Figure 1B As shown in the flowchart.
[0099] More specifically, given a target video, when synthesizing a face from a new audio track, lighting and facial pose can be extracted for each frame and employed during texture synthesis and 3D rendering, respectively. In some embodiments, only the areas of the lower face affected by speech are rendered, for example, below the mid-nose point. This is because some example current texture models do not generate varying eye gaze or blinks, and will therefore result in dull eyes on the upper face. However, one caveat is that the areas below the upper face and chin of the target frame will not necessarily coincide with the newly generated face. In particular, if in the target frame the original mouth is opened wider than in the synthesized frame, simply rendering the new face into the frame may result in a double chin.
[0100] Therefore, each target frame can be preprocessed by warping the image region below the original chin to match the expected new chin position. To avoid seams at boundary regions, a gradually blending blend can be created between the original and new facial geometries, and the original face in the target frame can be warped according to the blended geometry. Finally, Poisson blending (Perez et al., Poisson image editing, ACM Trans. Graph., 22(3):313-318, July 2003) can be used to remove any remaining chromatic aberrations and blend the rendered facial view into the warped target frame.
[0101] Example Method
[0102] Figure 6 Depicted is a flow chart of an example method 600 of generating a speaking face from audio, according to an example embodiment of the present disclosure.
[0103] At 602 , a computing system may obtain data describing an audio signal including speech.
[0104] In some embodiments, the audio signal is a separate audio signal from the visual representation of the speech. In other embodiments, the audio signal is associated with the visual representation of the speech.
[0105] In some implementations, the audio signal comprises recorded human audio utterances. In some implementations, the audio signal comprises synthesized text-to-speech audio generated from text data.
[0106] At 604, the computing system may predict facial geometry using a machine learning facial geometry prediction model.
[0107] At 606 , the computing system may predict facial texture using a machine learning facial texture prediction model.
[0108] In some embodiments, the machine-learned facial texture prediction model is an autoregressive model that receives as input, for each iteration of a plurality of iterations, a previous iteration prediction of the machine-learned facial texture prediction model.
[0109] In some embodiments, the predicted facial texture is a combination of a difference map predicted by a machine learning facial texture prediction model and a reference texture atlas.
[0110] In some implementations, the machine-learned facial geometry prediction model and the machine-learned facial texture prediction model are personalized models specific to a speaker of speech included in the audio signal.
[0111] In some embodiments, facial geometry predicted based at least in part on data describing an audio signal is predicted in a normalized three-dimensional space associated with a three-dimensional mesh; and facial texture predicted based at least in part on data describing an audio signal is predicted in a normalized two-dimensional space associated with a two-dimensional texture atlas.
[0112] At 608, the computing system may combine the facial geometry and the facial texture to generate a three-dimensional facial mesh model.
[0113] At 610 , the computing system may insert the facial mesh model into the two-dimensional video and / or the three-dimensional virtual environment.
[0114] For example, the facial mesh model can be inserted into a two-dimensional target video to generate a composite video. For example, inserting a three-dimensional face mesh model into a two-dimensional target video can include: obtaining a two-dimensional target video; detecting a target face in the two-dimensional target video; aligning the three-dimensional face mesh with the target face at a target location; and / or rendering the three-dimensional face mesh within the two-dimensional target video at the target location to generate a composite video.
[0115] In some embodiments, inserting the three-dimensional face mesh model into the two-dimensional target video may include: generating a fixed atlas from the two-dimensional target video; and / or providing the fixed atlas to a machine learning facial texture prediction model as a proxy illumination map.
[0116] In some embodiments, detecting the target face may include: using a three-dimensional face detector to obtain a pose and a triangular mesh of the target face in the video; and / or decomposing the target face into a three-dimensional normalized space associated with the three-dimensional mesh and a two-dimensional normalized space associated with the two-dimensional texture atlas. In some embodiments, the facial geometry predicted based at least in part on the data describing the audio signal is predicted in the normalized three-dimensional space associated with the three-dimensional mesh. In some embodiments, the facial texture predicted based at least in part on the data describing the audio signal is predicted in the normalized two-dimensional space associated with the two-dimensional texture atlas.
[0117] Fig. 7A A flow chart depicting an example method 700 for training a machine learning facial geometry prediction model according to an example embodiment of the present disclosure.
[0118] At 702 , a computing system may obtain a training video including visual data and audio data, wherein the visual data depicts a speaker and the audio data includes speech uttered by the speaker.
[0119] At 704 , the computing system may apply a three-dimensional facial landmark detector to the visual data to obtain three-dimensional facial features associated with the speaker's face.
[0120] At 706 , the computing system may predict facial geometry based at least in part on the data describing the audio data using a machine learning facial geometry prediction model.
[0121] At 708, the computing system can evaluate a loss term that compares the facial geometry predicted by the machine learning facial geometry model to the three-dimensional facial features generated by the three-dimensional facial landmark detector.
[0122] At 710, the computing system may modify one or more values of one or more parameters of the machine learning facial geometry prediction model based at least in part on the loss term.
[0123] Figure 7B A flow chart depicting an example method 750 for training a machine learning facial texture prediction model according to an example embodiment of the present disclosure. The method 750 may be performed separately from the method 700, or simultaneously / in conjunction with the method 700.
[0124] At 752 , the computing system may obtain a training video including visual data and audio data, wherein the visual data depicts a speaker and the audio data includes speech uttered by the speaker.
[0125] At 754 , the computing system may apply a three-dimensional facial landmark detector to the visual data to obtain three-dimensional facial features associated with the speaker's face.
[0126] At 756 , the computing system may project the training video onto a reference shape based on the three-dimensional facial features to obtain a training facial texture.
[0127] At 758, the computing system may predict facial texture based at least in part on the data describing the audio data using a machine learning facial texture prediction model.
[0128] In some embodiments, the method may further include generating a fixed atlas from the training video; and / or inputting the fixed atlas into a machine learning facial texture prediction model to be used as a proxy illumination map. In some embodiments, generating the fixed atlas may include: projecting the training video onto a reference shape using fixed reference facial coordinates; and / or masking pixels corresponding to eye and mouth regions.
[0129] At 760, the computing system may evaluate a loss term comparing the facial texture predicted by the machine learning facial texture model to the training facial textures.
[0130] At 762, the computing system may modify one or more values of one or more parameters of the machine learning facial texture prediction model based at least in part on the loss term.
[0131] Sample Application
[0132] So far, the proposed method for creating 3D speaking faces from audio input has been described. This section discusses some example applications of the technique. The method of generating fully textured 3D geometry enables a wider range of applications compared to purely image-based or 3D-only techniques.
[0133] Example photorealistic talking faces for gaming and VR
[0134] There is an increasing demand for similar-looking avatars in modern multiplayer online games and virtual reality (VR) to make the gaming environment more social and engaging. While such avatars can be driven by a video feed from a webcam (at least for a seated experience), the ability to generate a 3D speaking face from audio alone eliminates the need for any auxiliary camera equipment and as a side effect protects home privacy. Furthermore, it can reduce bandwidth and (combined with speech translation) even allow players to interact regardless of their language. Figure 4 An audio-only generated 3D face is shown integrated into a demo game. In this case, the model was trained from about six minutes of offline webcam footage of the subject.
[0135] Figure 4 : Screenshot of the mobile app, where a talking face driven only by audio is integrated into the demo game. Since a full 3D face model is generated, the face can be rendered from any viewpoint during gameplay.
[0136] Video Editing, Translation and Dubbing
[0137] Another important class of applications is the re-synthesis of video content. Using the techniques described in this article, a given video of a subject can be modified to match a new soundtrack. This can be used in a variety of scenarios:
[0138] Video creation and editing: New content can be inserted to update or enhance online courses, or to correct errors, without the tedious and sometimes impossible process of reshooting the entire video in its original condition. Instead, subjects only need to record new audio of the edited portion and apply our synthesis to modify the corresponding video segment. Extrapolating further, existing videos can be used only as a generic background to create completely new and different content driven by audio or text, enabling speech-to-video or text-to-video systems.
[0139] Video Translation and Dubbing: Even though some of the example models used for experiments were trained primarily on English videos, they were empirically shown to be surprisingly robust to both two different languages as well as TTS audio at inference time. Using available transcription or speech recognition systems to obtain subtitles, and subsequently using text-to-speech systems to generate audio, example implementations can automatically translate and lip-sync existing videos into different languages. Combined with appropriate video retiming and voice cloning, the resulting videos look quite convincing. Notably, in contrast to narrator-driven techniques, the thus-implemented approach for video transcription does not require a human actor in the loop, and is therefore immediately scalable across languages.
[0140] Additional example use cases
[0141] Many additional use cases or applications are possible. One additional example is a 2D or 3D cartoon talking avatar driven by audio. For example, additional layers can be used to map predicted geometry to control knobs of an animated character, such as blend shapes.
[0142] Another example application is video compression for face chat and / or converting an audio call to a speaking face. For example, a computing system (e.g., a receiving computing system) can reconstruct a face from the audio and (if necessary) other metadata (such as expression, lighting, etc.).
[0143] Another example application is to generate visualizations of a virtual assistant. For example, the computing system may be operated to provide a face to the assistant, which may be shown as a visual display such as Google Home. Expressions may also be added.
[0144] Example devices and systems
[0145] Figure 5A A block diagram of an example computing system 100 is depicted, in accordance with an example embodiment of the present disclosure. System 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 communicatively coupled via a network 180.
[0146] The user computing device 102 may be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop computer), a mobile computing device (e.g., a smart phone or tablet computer), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.
[0147] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and may be one processor or multiple processors operably connected. The memory 114 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 114 may store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.
[0148] In some implementations, the user computing device 102 may store or include one or more machine learning models 120. For example, the machine learning model 120 may be or may otherwise include various machine learning models, such as a neural network (e.g., a deep neural network) or other types of machine learning models, including nonlinear models and / or linear models. The neural network may include a feedforward neural network, a recurrent neural network (e.g., a long short-term memory recurrent neural network), a convolutional neural network, or other forms of neural networks. Figure 1A-4 An example machine learning model 120 is discussed.
[0149] In some implementations, one or more machine learning models 120 may be received from a server computing system 130 via a network 180, stored in a user computing device memory 114, and then used or otherwise implemented by one or more processors 112. In some implementations, a user computing device 102 may implement multiple parallel instances of a single machine learning model 120 (e.g., to execute in parallel across multiple instances).
[0150] Additionally or alternatively, one or more machine learning models 140 may be included in or otherwise stored and implemented by a server computing system 130 that communicates with the user computing device 102 according to a client-server relationship. For example, the machine learning models 140 may be implemented by the server computing system 140 as part of a web service (e.g., a face synthesis service). Thus, one or more models 120 may be stored and implemented at the user computing device 102, and / or one or more models 140 may be stored and implemented at the server computing system 130.
[0151] The user computing device 102 may also include one or more user input components 122 that receive user input. For example, the user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display screen or a touchpad) that is sensitive to the touch of a user input object (e.g., a finger or a stylus). A touch-sensitive component may be used to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other means by which a user can provide user input.
[0152] The server computing system 130 includes one or more processors 132 and memory 134. The one or more processors 132 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and may be one processor or multiple processors operably connected. The memory 134 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 may store data 136 and instructions 138 that are executed by the processor 132 to cause the server computing system 130 to perform operations.
[0153] In some implementations, server computing system 130 includes or is otherwise implemented by one or more server computing devices. Where server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.
[0154] As described above, the server computing system 130 may store or otherwise include one or more machine learning models 140. For example, the model 140 may be or may otherwise include a variety of machine learning models. Example machine learning models include neural networks or other multi-layer nonlinear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Figure 1A-4 Discuss example model 140 .
[0155] User computing device 102 and / or server computing system 130 may train models 120 and / or 140 via interaction with training computing system 150 communicatively coupled via network 180. Training computing system 150 may be separate from server computing system 130 or may be part of server computing system 130.
[0156] The training computing system 150 includes one or more processors 152 and a memory 154. The one or more processors 152 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and may be one processor or multiple processors operably connected. The memory 154 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc. and combinations thereof. The memory 154 may store data 156 and instructions 158 that are executed by the processor 152 to cause the training computing system 150 to perform operations. In some embodiments, the training computing system 150 includes or is otherwise implemented by one or more server computing devices.
[0157] The training computing system 150 may include a model trainer 160 that trains the machine learning models 120 and / or 140 stored at the user computing device 102 and / or the server computing system 130 using various training or learning techniques (e.g., back propagation of errors). For example, a loss function may be back propagated through the model(s) to update one or more parameters of the model(s) (e.g., based on the gradient of the loss function). Various loss functions may be used, such as mean squared error, likelihood loss, cross entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques may be used to iteratively update parameters over multiple training iterations.
[0158] In some implementations, performing back propagation of errors may include performing truncated back propagation through time.The model trainer 160 may perform a variety of generalization techniques (eg, weight decay, dropout, etc.) to improve the generalization capabilities of the model being trained.
[0159] In particular, model trainer 160 can train machine learning models 120 and / or 140 based on a set of training data 162. Training data 162 can include, for example, existing videos depicting speech.
[0160] In some implementations, if the user has provided consent, the training examples may be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 may be trained by the training computing system 150 on user-specific data received from the user computing device 102. In some cases, this process may be referred to as personalizing the model.
[0161] The model trainer 160 includes computer logic for providing the desired functionality. The model trainer 160 can be implemented in hardware, firmware, and / or software that controls a general purpose processor. For example, in some embodiments, the model trainer 160 includes a program file stored on a storage device, loaded into a memory, and executed by one or more processors. In other embodiments, the model trainer 160 includes one or more sets of computer executable instructions stored in a tangible computer readable storage medium, such as a RAM hard disk or an optical or magnetic medium.
[0162] Network 180 may be any type of communications network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and may include any number of wired or wireless links. In general, communications over network 180 may be carried via any type of wired and / or wireless connection using a variety of communications protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).
[0163] Figure 5A An example computing system that can be used to implement the present disclosure is shown. Other computing systems may also be used. For example, in some embodiments, the user computing device 102 may include a model trainer 160 and a training data set 162. In such embodiments, the model 120 may be trained and used locally at the user computing device 102. In some such embodiments, the user computing device 102 may implement the model trainer 160 to personalize the model 120 based on user-specific data.
[0164] Figure 5B Depicted is a block diagram of an example computing device 10 that performs in accordance with an example embodiment of the present disclosure. Computing device 10 may be a user computing device or a server computing device.
[0165] Computing device 10 includes multiple applications (e.g., applications 1 to N). Each application contains its own machine learning library and machine learning model. For example, each application can include a machine learning model. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc.
[0166] like Figure 5B As shown, each application can communicate with multiple other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some embodiments, each application can communicate with each device component using an API (e.g., a public API). In some embodiments, the API used by each application is specific to the application.
[0167] Figure 5C Depicted is a block diagram of an example computing device 50 that performs in accordance with an example embodiment of the present disclosure. Computing device 50 may be a user computing device or a server computing device.
[0168] The computing device 50 includes a plurality of applications (e.g., applications 1 to N). Each application communicates with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can communicate with the central intelligence layer (and the models stored therein) using an API (e.g., a public API across all applications).
[0169] The central intelligence layer includes multiple machine learning models. For example, Figure 5C As shown, a corresponding machine learning model (e.g., model) can be provided for each application and managed by the central intelligence layer. In other embodiments, two or more applications can share a single machine learning model. For example, in some embodiments, the central intelligence layer can provide a single model (e.g., a single model) for all applications. In some embodiments, the central intelligence layer is included in the operating system of the computing device 50 or is otherwise implemented by the operating system of the computing device 50.
[0170] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized data repository for computing devices 50. Figure 5C As shown, the central device data layer can communicate with multiple other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some embodiments, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0171] Additional disclosure
[0172] The techniques discussed herein refer to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from these systems. The inherent flexibility of computer-based systems allows for various possible configurations, combinations, and divisions of tasks and functions between and among components. For example, the processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0173] Although the present subject matter has been described in detail with respect to various specific example embodiments of the present subject matter, each example is provided by way of explanation rather than limitation of the present disclosure. Those skilled in the art, after obtaining an understanding of the foregoing, can easily produce changes, variations and equivalents to these embodiments. Therefore, the present subject matter disclosure does not exclude such modifications, variations and / or additions including to the present subject matter, which are apparent to those of ordinary skill in the art. For example, a feature shown or described as a part of an embodiment can be used together with another embodiment to produce yet another embodiment. Therefore, the present disclosure is intended to cover these changes, variations and equivalents.
Claims
1. A computing system for generating a speaking face from an audio signal, the computing system comprising: one or more processors; as well as One or more non-transitory computer-readable media that collectively store: a machine learning facial geometry prediction model configured to predict facial geometry based on data describing an audio signal including speech; a machine learning facial texture prediction model configured to predict facial texture based on data describing an audio signal including speech; and The instructions, when executed by the one or more processors, cause the computing system to perform operations comprising: obtaining data describing an audio signal including speech; using the machine learning facial geometry prediction model to predict facial geometry based at least in part on data describing an audio signal; using the machine learning facial texture prediction model to predict the facial texture based at least in part on data describing an audio signal; and The facial geometry and facial texture are combined to generate a three-dimensional facial mesh model.
2. The computing system of claim 1, wherein: The audio signal comprises a separate audio signal separate from the visual representation of the speech.
3. The computing system of claim 1, wherein: The data describing an audio signal comprises a spectrogram of the audio signal.
4. The computing system of claim 1, wherein: The facial geometry predicted based at least in part on the data describing the audio signal is predicted within a normalized three-dimensional space associated with the three-dimensional mesh; as well as The facial texture predicted based at least in part on data describing the audio signal is predicted within a normalized two-dimensional space associated with the two-dimensional texture atlas.
5. The computing system of claim 1, wherein: The operations also include inserting the three-dimensional facial mesh model into the two-dimensional target video to generate a composite video.
6. The computing system of claim 5, wherein: Inserting a 3D facial mesh model into a 2D target video involves: Obtaining the two-dimensional target video; Detect target faces in 2D target videos; aligning the three-dimensional facial mesh with the target face at the target location; and The three-dimensional facial mesh is rendered within the two-dimensional target video at a target location to generate a composite video.
7. The computing system of claim 5, wherein: Inserting a 3D facial mesh model into a 2D target video involves: Generating a fixed atlas from a two-dimensional target video; and The fixed atlas is fed into a machine-learned facial texture prediction model as a proxy illumination map.
8. The computing system of claim 6, wherein: Detecting the target face includes: Using a 3D face detector to obtain the pose and triangle mesh of a target face in the video; and Decomposing the target face into a three-dimensional normalized space associated with the three-dimensional mesh and a two-dimensional normalized space associated with the two-dimensional texture map; wherein the facial geometry predicted based at least in part on data describing the audio signal is predicted within the normalized three-dimensional space associated with the three-dimensional mesh; and Wherein the facial texture predicted at least in part based on data describing the audio signal is predicted within the normalized two-dimensional space associated with the two-dimensional texture atlas.
9. The computing system of claim 1, wherein: The operations also include rendering the three-dimensional facial mesh in the three-dimensional virtual environment.
10. The computing system of claim 1, wherein: The audio signal comprises recorded human audio utterances or synthesized text-to-speech audio generated from text data.
11. The computing system of claim 1 , further comprising: Perform illumination normalization on facial texture predicted by a machine learning facial texture prediction model.
12. The computing system of claim 1, wherein: The machine learning facial geometry prediction model and the machine learning facial texture prediction model include personalized models specific to a speaker of speech included in the audio signal.
13. The computing system of claim 1, wherein: The machine learning facial texture prediction model comprises an autoregressive model that, for each iteration of a plurality of iterations, receives as input a previous iteration prediction of the machine learning facial texture prediction model.
14. The computing system of claim 1, wherein: The predicted facial texture comprises a combination of a difference map predicted by a machine learning facial texture prediction model and a reference texture atlas.
15. A computer-implemented method for learning to generate a three-dimensional facial mesh from a training video, the method comprising: obtaining, by a computing system including one or more computing devices, a training video including visual data and audio data, wherein the visual data depicts a speaker and the audio data includes speech uttered by the speaker; applying, by a computing system, a three-dimensional facial landmark detector to the visual data to obtain three-dimensional facial features associated with the speaker's face; Projecting the training video onto the reference shape by a computing system based on the three-dimensional facial features to obtain a training facial texture; predicting, by a computing system and using a machine learning facial geometry prediction model, facial geometry based at least in part on data describing the audio data; Predicting, by a computing system and using a machine learning facial texture prediction model, facial texture based at least in part on the data describing the audio data; modifying, by the computing system, one or more values of one or more parameters of the machine-learned facial geometry prediction model based at least in part on a first loss term that compares facial geometry predicted by the machine-learned facial geometry prediction model to three-dimensional facial features generated by a three-dimensional facial landmark detector; as well as One or more values of one or more parameters of the machine-learned facial texture prediction model are modified, by the computing system, based at least in part on a second loss term that compares a facial texture predicted by the machine-learned facial texture prediction model to training facial textures.
16. The computer-implemented method of claim 15, further comprising: generating a fixed atlas from the training video; as well as The fixed atlas is fed into a machine learning facial texture prediction model to serve as a proxy illumination map.
17. The computer-implemented method of claim 16, wherein: Generating a fixed graph includes: projecting the training video onto a reference shape using fixed reference facial coordinates; and Mask out pixels corresponding to the eye and mouth regions.
18. The computer-implemented method of claim 15, wherein: The machine learning facial texture prediction model includes an autoregressive model that receives as input, for each iteration of a plurality of iterations, a previous iteration prediction of the machine learning facial texture prediction model.
19. The computer-implemented method of claim 15, wherein: The predicted facial texture comprises a combination of a difference map predicted by a machine learning facial texture prediction model and a reference texture atlas.
20. One or more non-transitory computer-readable media storing instructions that, when executed by a computing system comprising one or more computing devices, cause the computing system to perform operations comprising: obtaining, by a computing system, a training video comprising visual data and audio data, wherein the visual data depicts a speaker and the audio data comprises speech uttered by the speaker; applying, by a computing system, a three-dimensional facial landmark detector to the visual data to obtain three-dimensional facial features associated with the speaker's face; Projecting the training video onto the reference shape by a computing system based on the three-dimensional facial features to obtain a training facial texture; Predicting, by a computing system and using a machine learning facial texture prediction model, facial texture based at least in part on the data describing the audio data; evaluating, by a computing system, a loss term comparing a facial texture predicted by the machine learning facial texture prediction model to training facial textures; and One or more values of one or more parameters of the machine learning facial texture prediction model are modified, by the computing system, based at least in part on the loss term.
Citation Information
Patent Citations
Photo-realistic synthesis of three dimensional animation with facial features synchronized with speech
US20120280974A1
Producing realistic talking Face with Expression using Images text and voice
US20190197755A1