Three-dimensional face animation from speech
Patent Information
- Authority / Receiving Office
- TW · TW
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-03-04
- Publication Date
- 2022-10-01
Smart Images

Figure TWG2TA000877808_001 
Figure TWG2TA000877808_002 
Figure TWG2TA000877808_003
Abstract
Description
[Technical Field]
[0001] This disclosure generally pertains to the field of generating three-dimensional computer models of individuals captured by video capture. More specifically, this disclosure pertains to generating three-dimensional (3D), full-face animations of individuals with their own speech in video capture.
[0002] Cross-reference to related applications
[0003] This disclosure is made pursuant to and claims priority to U.S. Provisional Application No. 63 / 161848 filed March 16, 2021, by Alexander Richard et al., entitled “MESH TALK: 3D FACE ANIMATION FROM SPEECH USING CROSS-MODALITY DISENTANGLEMENT”, under 35 USC §119(e), which is incorporated herein by reference in its entirety for all purposes. [Previous Technology]
[0004] Existing methods of audio-driven facial animation exhibit strange or static upper facial animations, fail to produce accurate and seemingly reasonable pronunciations, or rely on personalized models that limit their scalability. [Summary of the Invention]
[0005] In a first specific example, a computer implementation method includes: identifying an audio-related facial feature from an individual's audio capture; generating a first mesh for a lower portion of the individual's face based on the audio-related facial feature; and identifying an expression-like facial feature of the individual. The computer implementation method also includes: generating a second mesh for an upper portion of the individual's face based on the expression-like facial feature; forming a composite mesh using the first mesh and the second mesh; and determining a loss value for the composite mesh based on a ground truth image of the individual. The computer implementation method also includes: generating a three-dimensional model of the individual's face using the composite mesh based on the loss value; and providing the three-dimensional model of the individual's face to a display in a client device running an immersive reality application including the individual.
[0006] In a second specific embodiment, a system includes: a memory storing a plurality of instructions; and one or more processors configured to execute the instructions to cause the system to perform operations. The operations include: identifying an audio-related facial feature from an individual's audio capture; generating a first mesh of a lower portion of the individual's face based on the audio-related facial feature; and identifying an expression-like facial feature of the individual. The operations also include: generating a second mesh of an upper portion of the individual's face based on the expression-like facial feature; forming a composite mesh using the first mesh and the second mesh; and determining a loss value of the composite mesh based on a baseline image of the individual. The operations also include: generating a three-dimensional model of the individual's face using the composite mesh based on the loss value; and providing the three-dimensional model of the individual's face to a display in a client device running an immersive reality application including the individual.
[0007] In a third specific example, a computer implementation method includes: determining a first correlation value of a facial feature based on an audio waveform from a first person; generating a first mesh of a lower portion of a face based on the facial feature and the first correlation value; updating the first correlation value based on a difference between the first mesh of the first person and a reference truth image; and providing a three-dimensional model of the voice-animated face to an immersive reality application accessed by a user terminal device based on the difference between the first mesh of the first person and the reference truth image.
[0008] In another specific example, a non-transitory computer-readable media storage instruction, when executed by a processor, causes a computer to perform a method. The method includes: identifying an audio-related facial feature from audio capture of an individual; generating a first mesh of a lower portion of the individual's face based on the audio-related facial feature; and identifying an expression-like facial feature of the individual. The method also includes: generating a second mesh of an upper portion of the individual's face based on the expression-like facial feature; forming a composite mesh using the first mesh and the second mesh; and determining a loss value of the composite mesh based on a reference image of the individual. The method further includes: generating a three-dimensional model of the individual's face using the composite mesh based on the loss value; and providing the three-dimensional model of the individual's face to a display in a client device running an immersive reality application including the individual.
[0009] In yet another specific example, a system includes components for storing one type of instructions and for executing the instructions to perform one type of method, the method including: identifying an audio-related facial feature from an individual's audio capture; generating a first mesh for a lower portion of the individual's face based on the audio-related facial feature; and identifying an expression-like facial feature of the individual. The method also includes: generating a second mesh for an upper portion of the individual's face based on the expression-like facial feature; forming a composite mesh using the first mesh and the second mesh; and determining a loss value for the composite mesh based on a reference image of the individual. The method further includes: generating a three-dimensional model of the individual's face using the composite mesh based on the loss value; and providing the three-dimensional model of the individual's face to a display in a client device running an immersive reality application including the individual.
Implementation Method
[0023] In the following detailed description, numerous specific details are set forth to provide a full understanding of this disclosure. However, it will be apparent to those skilled in the art that specific examples of this disclosure can be practiced without some of these specific details. In other instances, well-known structures and techniques have not been shown in detail so as not to obscure this disclosure.
[0024] General Overview
[0025] Voice-driven facial animation presents a challenging technical problem for several applications, such as facial animation in computer games, e-commerce, immersive virtual reality (VR) telepresence, and other augmented reality (AR) applications. The requirements for voice-driven facial animation vary depending on the application. Applications such as speech therapy or entertainment (e.g., Animoji effects or AR effects) can use animations with lower precision / realism. Conversely, in film production, movie dubbing, and driven avatars for e-commerce applications or immersive telepresence, voice animation needs to be highly realistic, plausible, and must provide intelligibility comparable to that of a natural speaker. The human visual system has evolved to understand subtle facial movements and expressions. Therefore, insufficiently animated faces lacking realistic coordinated pronunciation or lip-syncing are considered distracting to users and detrimental to the commercial success of devices or applications.
[0026] There is a significant degree of dependence between speech and facial gestures. This dependence has been utilized by audio-driven facial animation methods developed in computer vision and graphics. With the development of deep learning technology, some audio-driven facial animation techniques are based on large corpora of paired audio and mesh data, using personalized methods trained in a supervised manner. Some of these methods achieve high-quality lip-sync animation and synthesize seemingly reasonable upper facial movements solely from the audio. However, to obtain the required training data, high-quality visual-based motion capture of the user is needed, making these methods highly impractical for consumer-facing applications in real-world scenarios. Some methods involve generalization or averaging of different identities and thus can animate any user based on a given audio stream and a static, neutral 3D scan of the user. While such methods are indeed feasible in real-world scenarios, they often exhibit odd or static upper facial animation because the audio does not encode all the variations of facial expressions. Therefore, typical audio-driven facial animation models attempt to learn a one-to-many mapping, meaning that for each input there are multiple seemingly plausible outputs. This often produces overly smoothed results (e.g., unnatural, anomalous, or obviously artificial), especially in areas of the face that are only slightly related to or even unrelated to the audio signal.
[0027] To address these technical problems arising in the fields of computer networks, computer simulations, and immersive reality applications, specific examples disclosed herein include techniques such as audio-driven facial animation methods, which enable highly realistic motion synthesis of the entire face and are also generalized to unseen identities. Therefore, machine learning applications include a classification latent space for facial animation that distinguishes between audio-related and audio-unrelated information. For example, eye closure may not be limited to a specific lip shape. The latent space is trained based on a novel cross-modal loss that promotes accurate upper facial reconstruction independent of audio input and accurate mouth region dependent only on the provided audio input. This distinguishes between the motion of the lower and upper facial regions and prevents oversmoothing. Motion synthesis is based on an autoregressive sampling strategy using an audio conditioning temporal model via the learned classification latent space. Our method ensures highly accurate lip movements while also sampling seemingly reasonable animations of facial features unrelated to the audio signal, such as blinking and eyebrow movements.
[0028] It is desirable to animate arbitrary neutral facial meshes using only speech, as this is faster (e.g., an audio waveform of less than 1 second is sufficient). Because speech does not encode all forms of facial expressions, such as blinking and similar features, many speech-irrelevant facial features exist in the face. This results in the odd or static upper facial animations exhibited by most existing audio-driven methods. To overcome this technical problem, specific examples disclosed herein include a classification latent space for facial expressions stored in a training database. At inference time, some specific examples perform autoregressive sampling of a speech-modulated temporal model via the classification latent space, which ensures accurate lip movements and simultaneously synthesizes seemingly plausible animations of speech-irrelevant facial parts. The classification latent space may include the following features: 1) Classification: The space is segmented according to the learned categories. 2) Expression: The latent space may be able to encode a variety of facial expressions, including rare facial phenomena such as blinking. And 3) Semantic distinction: It is expected that at least partially, speech-related and speech-unrelated information can be distinguished, for example, eye closing should not be limited to a given lip shape or mouth posture.
[0029] Furthermore, specific examples disclosed herein include retargeting configurations in which a 3D speech animation model trained on one or more individuals is smoothly applied to different individuals. In some specific instances, the 3D speech animation model disclosed herein can be used to dub speech from a given individual into multilingual speech from one or more different individuals.
[0030] Instance System Architecture
[0031] Figure 1 illustrates an instance architecture 100 suitable for accessing a 3D speech animation engine, according to some specific examples. Architecture 100 includes servers 130 that are communicatively coupled to a client device 110 and at least one database 152 via a network 150. One of the servers 130 is configured to manage memory including instructions that, when executed by a processor, cause the server 130 to perform at least some steps of the methods disclosed herein. In some specific examples, the processor is configured to control a graphical user interface (GUI) for a user accessing one of the client devices 110 of the 3D speech animation engine. The 3D speech animation engine can be configured to train a machine learning model for executing a specific application. Therefore, the processor may include a dashboard tool configured to display components and graphical results to the user via the GUI. For load balancing purposes, multiple servers 130 may manage memory including instructions to one or more processors, and multiple servers 130 may manage historical logs and a database 152 including multiple training archives for a 3D speech animation engine. Furthermore, in some specific instances, multiple users of client devices 110 may access the same 3D speech animation engine to run one or more machine learning models. In some specific instances, a single user with a single client device 110 may train multiple machine learning models to run in parallel on one or more servers 130. Therefore, client devices 110 may communicate with each other via network 150 and by accessing one or more servers 130 and the resources located therein. In some specific instances, at least one or more client devices 110 may include a headset for virtual reality (VR) applications or smart glasses for augmented reality (AR) applications, as disclosed herein. In this regard, the head-mounted unit or smart glasses can be paired with a smartphone for wireless communication with AR / VR applications installed on the smartphone, and using the smartphone, the head-mounted unit or smart glasses can communicate with the server 130 via the network 150.
[0032] Server 130 may include any device having a suitable processor, memory, and communication capabilities for managing the 3D voice animation engine (including multiple tools associated therewith). The 3D voice animation engine may be accessed via network 150 by various client devices 110. Client devices 110 may be, for example, desktop computers, mobile computers, tablet computers (e.g., including e-book readers), mobile devices (e.g., smartphones or PDAs), or any other device having a suitable processor, memory, and communication capabilities for accessing the 3D voice animation engine on one or more of servers 130. Network 150 may include, for example, any or more of a Local Area Network (LAN), a Wide Area Network (WAN), the Internet, and the like. In addition, network 150 may include, but is not limited to, any or more of the following topologies, including bus networks, star networks, ring networks, mesh networks, star bus networks, tree or hierarchical networks, and the like.
[0033] Figure 2 is a block diagram 200 illustrating an example server 130 and client device 110 from architecture 100 according to certain configurations of this disclosure. Client device 110 and server 130 are communicatively coupled via network 150 through respective communication modules 218-1 and 218-2 (hereinafter collectively referred to as "communication module 218"). Communication module 218 is configured to interface with network 150 to send and receive information such as data, requests, responses, and commands to other devices via network 150. Communication module 218 may be, for example, a modem or Ethernet card. Users can interact with client device 110 via input device 214 and output device 216. Input device 214 may include a mouse, keyboard, indicator, touchscreen, microphone, joystick, wireless joystick, and the like. Output device 216 may be a screen display, a touch screen, a speaker, or the like. Client device 110 may include memory 220-1 and a processor 212-1. Memory 220-1 may include an application 222 and a GUI 225 configured to run on client device 110 and coupled to input device 214 and output device 216. Application 222 may be downloaded by a user from server 130 and may be managed by server 130. In some specific instances, client device 110 may include a head-mounted array or smart glasses, and application 222 may include an immersive reality environment in an AR / VR application, as disclosed herein. During the execution of application 222, client device 110 and server 130 may transmit data packets 227-1 and 227-2 between each other via communication module 218 and network 150. For example, client device 110 may provide data packet 227-1, which includes voice signals or sound files from the user, to server 130. Therefore, server 130 may provide data packet 227-2, which includes a 3D animated model of the user, to client device 110 based on the voice signals or sound files from the user.
[0034] Server 130 includes memory 220-2, processor 212-2, and communication module 218-2. Hereinafter, processors 212-1 and 212-2 and memory 220-1 and 220-2 will be collectively referred to as "processor 212" and "memory 220," respectively. Processor 212 is configured to execute instructions stored in memory 220. In some specific instances, memory 220-2 includes a 3D voice animation engine 232. The 3D voice animation engine 232 may share features and resources or be provided to the GUI 225, including multiple tools associated with training and using 3D model animations of faces for immersive reality applications including voice. Users can access the 3D voice animation engine 232 via application 222 installed in memory 220-1 of user terminal device 110. Therefore, application 222 can be installed on server 130 and execute instruction codes and other routines provided by server 130 via any of a plurality of tools. The execution of application 222 can be controlled by processor 212-1.
[0035] In this regard, the 3D speech animation engine 232 can be configured to create, store, update, and maintain the multi-peak encoder 240, as disclosed herein. The multi-peak encoder 240 may include an audio encoder 242, a facial expression encoder 244, a convolution tool 246, and a synthesis encoder 248. The 3D speech animation engine 232 may also include a synthesis decoder 248. In some specific instances, the 3D speech animation engine 232 may access one or more machine learning models stored in a training database 252. The training database 252 includes training archives and other data files that can be used by the 3D speech animation engine 232 when training machine learning models based on input from a user via application 222. Furthermore, in some specific instances, at least one or more training archives or machine learning models may be stored in either memory 220. A user of the client device 110 may access the training archives via application 222.
[0036] Audio encoder 242 identifies audio-related facial features based on a classification scheme learned through training to generate a first grid for the lower portion of an individual's face. To do this, audio encoder 242 is able to identify the intensity and frequency of a sound wave waveform or a portion thereof from the individual's audio capture. Audio capture may include portions of the individual's speech captured in real time by an AR / VR application (e.g., application 222) or collected during a training session and stored in a training database 252. Audio encoder 242 may also correlate the intensity and frequency of the sound wave waveform with the geometry of the lower portion of the individual's face (e.g., the mouth and lips, and portions of the jaw and cheeks). Facial expression encoder 244 identifies facial features resembling expressions of the individual to generate a second grid for the upper portion of the individual's face. Thus, facial expression encoder 244 may randomly select facial features resembling expressions based on previous samples of facial expressions from multiple individuals. In this regard, multiple individual facial expressions collected during training sessions involving reading text or being in a dialogue with a second individual can be stored in training database 252 and accessed by facial expression encoder 244. In some specific instances, facial expression encoder 244 correlates upper facial features with speech features captured from the individual's audio input.
[0037] The convolution tool 246 may be part of a convolutional neural network (CNN) configured to reduce the dimensionality of multiple neural network layers in a 3D animation model. In some specific instances, the convolution tool 246 provides temporal convolutions (e.g., tCNN) for 3D animation of an individual's face based on speech. In some specific instances, the convolution tool 246 provides autoregressive convolutions, where labels generated in other layers of the neural network are fed back to previous layers to improve class scanning in the CNN. The synthesizer 248 generates a synthetic mesh of the individual's full face using a first mesh provided by the audio encoder 242 and a second mesh provided by the facial expression encoder 244. Thus, the synthesizer 248 continuously and smoothly incorporates the lip shape in the first mesh provided by the audio encoder 242 into the eye closure in the second mesh provided by the facial expression encoder 244 across the individual's face. In some specific instances, the synthesizer 248 may include additional skip connections that utilize the inductive bias of the CNN to handle computationally limited capacity.
[0038] The 3D voice animation engine 232 also includes a multi-peak decoder 250, which is configured to generate a three-dimensional model of an individual’s face using a synthetic mesh and provides the three-dimensional model of the individual’s face to a display on a user device 110 running application 222 (e.g., including an immersive reality application for the individual).
[0039] The 3D voice animation engine 232 may include algorithms trained for the specific purpose of the engine and the tools included therein. The algorithms may include machine learning or artificial intelligence algorithms utilizing any linear or nonlinear algorithms, such as neural network algorithms or multivariate regression algorithms. In some specific instances, the machine learning model may include neural networks (NN), convolutional neural networks (CNN), generative adversarial neural networks (GAN), deep reinforcement learning (DRL) algorithms, deep recurrent neural networks (DRNN), classical machine learning algorithms such as random forests, k-nearest neighbor (KNN) algorithms, k-means clustering algorithms, or any combination thereof. More generally, the machine learning model may include any machine learning model involving training and optimization steps. In some specific instances, the training database 252 may include a training archive used to modify coefficients according to the desired results of the machine learning model. Therefore, in some specific instances, the 3D speech animation engine 232 is configured to access the training database 252 to retrieve files and archives as input for machine learning models. In some specific instances, at least a portion of the 3D speech animation engine 232, the tools contained therein, and the training database 252 may be hosted on different servers accessible by server 130.
[0040] Figure 3 illustrates a block diagram of the speech-animated mapping 300 from a neutral face mesh 327 and a speech signal 328 to an expression face mesh 351, based on some specific examples. The synthesis encoder 348 includes a fusion block 330, which maps the sequence of input animated face meshes 329 (expression signals) and speech signals 328 to the encoded expression 341 in the classification latent space 340 via the synthesis encoder 348. The decoder 350 animates the neutral face mesh 327 from the encoded expression 341.
[0041] To achieve high realism, in some specific instances, the mapping 300 is trained using a dataset of multiple individuals and available data including eyelids, facial hair, or eyebrows, thus displaying highly realistic full-face motion from speech by any identity. In some specific instances, an internal dataset of 250 individuals is used for training, each of whom is reading a total of 50 speech-balanced sentences. Speech signals 328 are captured at 30 frames per second, and a face mesh is tracked from 80 synchronized cameras surrounding the individual's head (see neutral face mesh 327 and animated face mesh 329). In some specific instances, the face mesh may include 6,172 vertices with high-level detail, including eyelids, upper facial structure, and different hairstyles. In some specific instances, the data is equivalent to 13 hours of paired audiovisual data, or 1.4 million frames of tracked 3D face mesh. Mapping 300 can be trained on the first 40 sentences of 200 individuals, and the remaining 10 sentences of the remaining 50 individuals are used as the validation set (10 individuals) and the test set (40 individuals). In some specific instances, a subset of 16 individuals from this dataset can be used as a baseline for comparison with Mapping 300. The data is stored in a database (see Training Database 252).
[0042] In some specific instances, the speech signal 328 is recorded at 16 kHz. For each tracked grid, a Mel spectrogram is generated, comprising a 600 ms audio segment starting 500 ms before and ending 100 ms after each individual visual frame. In some specific instances, the speech signal 328 comprises 80-dimensional Mel spectrogram features collected every 10 ms using 1,024 frequency ranges and a window size of 800 for the basic Fourier transform.
[0043] To train the classification latent space 340, x1:T = (x1, ..., xT), xt ∈ RV×3 is a series of T face meshes 329, each with V vertices. Further, a1:T = (a1, ..., aT), at ∈ RD is a series of T speech segments 328, each having D samples aligned with the corresponding (visual) frame t. Additionally, the template mesh 327 can be represented as h ∈ RV×3.
[0044] To achieve high performance, a large classification latent space 340 is desired. However, this can result in an infeasible large number of categories C for a single latent classification layer. Therefore, in some specific instances, a smaller number H of latent classification heads 335 for C-type categories is modeled. This allows for a large expression space with a fairly small number of categories because the number of configurations of the classification latent space 340 is CH, and therefore grows exponentially in units of H. In some specific instances, values C=128 and H=64 are sufficient to obtain accurate results for immediate applications.
[0045] The mapping from the facial expression and audio input signals to the multi-head classification latent space is implemented by an encoder (e.g., fusion block 330), which maps from the space of audio sequence 328 and facial expression sequence 329 to T×H×C dimensional encoding as follows: (1)
[0046] In some specific instances, the continuous value encoding in Equation 1 is transformed into a classification representation via the Gumbel-softmax transform through each latent classification head, (2)
[0047] such that each classification component in time step t and in the potential classification head h is assigned one of the C classification labels, ct,hϵ{1、……、C}. The complete encoding function for the classification (see Equation 2) can then be expressed as follows.
[0048] The animation of the input template mesh 327 (h) is implemented by the decoder 350 (D), as follows: (3)
[0049] It maps the encoded expression 341 onto the template mesh 327(h). The decoder 350 generates an animation sequence 351() of the face mesh, which looks like the person represented by the template mesh 327(h), but moves according to the expression code c1:T,1:H.
[0050] During training, the baseline truth correspondence can be used when: (a) the template grid 327, speech signal 328, and facial expression signal 329 are from the same individual, and (b) the desired output from the decoder 350 (e.g., animation sequence 351) equals the facial expression input 329 (e.g., x1:T, see above). To complete the training, some specific instances include a cross-modality loss function L, which ensures that information from the two input modalities (e.g., speech signal 328 and facial expression signal 329) is used to classify the latent space 340. Let x1:T and a1:T be the given facial expression sequence 329 and speech sequence 328, respectively. Further let hx represent the template grid 327 of the individual represented in signal x1:T. In some specific instances, the decoder 350 produces two different reconstructions instead of a single reconstruction: (4)(5)
[0051] wherein and are randomly sampled expression and audio sequences from the training database (e.g., training database 252). In some specific instances, given a correct audio sequence but a random expression sequence, it is a reconstruction, and given a correct expression sequence but a random audio sequence, it is a reconstruction. Therefore, the cross-modal loss LxMod can then be defined as: (6)
[0052] The mask assigns high weight to the vertex v on the upper face and low weight to the vertices surrounding the mouth. Similarly, it assigns high weight to the vertex v surrounding the mouth and low weight to the other vertices.
[0053] In some specific instances, the cross-modal loss LxMod promotes accurate upper face reconstruction independent of audio input 328, and therefore accurate mouth region reconstruction based on audio independent of expression sequence 329. Since blinking is a fast and sparse event that affects only a few vertices, some specific instances include the loss Leyelid, which emphasizes eyelid vertices during training, as follows: (7)
[0054] Wherein is a binary mask, which has zeros for one of the eyelid vertices and zeros for the other vertices. Therefore, the final loss function L can be optimized as: L = LxMod + Leeyelid. In some specific instances, equal weighting of the two terms (LxMod and Leeyelid) actually works well. Therefore, other specific instances may include different weightings between the LxMod and Leeyelid losses.
[0055] In some specific instances, the audio encoder 342 includes a four-layer, one-dimensional (1D) temporal convolutional network. In some specific instances, the expression encoder 344 may include three fully connected layers followed by a single long short-term memory (LSTM) layer to capture temporal dependencies. The fusion block 330 may include a three-layer perceptron. The decoder 350(D) may include an additional skip connection architecture. This architecture generalizes bias to prevent excessive divergence of the network from the template mesh 327. In the bottleneck layer, the expression codes c1:T,1:H are concatenated with the encoded expression 341. In some specific instances, the bottleneck layer is followed by two LSTM layers to model the temporal dependencies between frames, followed by three fully connected layers that remap the representation to the vertex space. The expression input x1:T includes a target signal that will minimize the loss function at the output of the decoder 350 (see Equations 6 and 7) by means of a sequence including the audio signal 328 in the classification latent space 340 and the face mesh 329. This method avoids the problem found in many multimodal methods, which often neglect "weaker" modes (e.g., audio, which is typically less data-intensive).
[0056] In some specific instances, the audio signal 328 may be omitted when training the classification latent space 340. The finite capacity of the classification latent space 340 and the inductive bias of the audio decoder 342 (e.g., skip connections therein) ensure that even in this case, sufficient information from the template geometry is used. In some specific instances, this setup also produces low reconstruction errors as shown in Table 1. In some specific instances, it is desirable to avoid strong entanglement between eye movements and mouth shapes in the latent representation of accurate lip shapes, while simultaneously producing temporally consistent and seemingly plausible upper facial movements. Table 1
[0057] To quantify this effect (“complexity”), given the classification latent representation 340(c1:T,1:H) of the test set data, the complexity can be calculated as follows: (8)
[0058] Equation 8 is the inverse geometric mean of the probabilities of the latent representations under model 300. Intuitively, low complexity means that model 300 has only a small number of latent classes h to choose from at each prediction step, while high complexity means that the model is less certain which classification representation to choose next. A complexity of 1 would mean that the autoregressive model is fully determinate, for example, the latent embedding is fully defined by the modulated audio input. This may not occur frequently in practice due to the presence of facial movements unrelated to the audio. In some specific instances (see Table 1, third column), training a classification latent space 340 from both audio and facial inputs produces a stronger and more confident model 300 than learning a latent space from only facial expression input.
[0059] The training loss of the decoder (Equations 6 and 7) determines how model 300 utilizes different input modalities (audio / facial expressions). Since the facial expression input (facial expression 329) is sufficient for accurate reconstruction, a simple loss on the output grid will cause model 300 to ignore the audio input, and the result is similar to the case above where no audio is given as encoder input (see Table 1, Columns 1 and 2). The cross-modal loss LxMod (Equation 6) provides an effective solution by promoting model 300 to learn accurate lip shapes even when the facial expression input is changed by different random expressions. Similarly, upper facial motion is promoted to remain accurate independent of audio input. The cross-modal loss does not affect the expressiveness of the learned latent space (see Table 1, Column 3), for example, the reconstruction error is small for all forms of latent space variation, and it positively affects the complexity of the autoregressive model (see Equation 8).
[0060] Figure 4 illustrates a block diagram of an autoregressive model 400 including preselected markers 405, based on some specific examples. When the audio input 428 is used alone to drive the template mesh (e.g., mesh 327), the expression input x1:T is unavailable. Given only one modality, the missing information not inferred from the audio input 428 is synthesized. Thus, some specific examples include an autoregressive temporal model 400 on a classification latent space 440. The audio signal 428 is encoded by an audio encoder 442, and a head reader prepares a classification encoding space 440, which is scanned along the temporal direction by audio conditioning latent codes 435. Audio conditioning latent codes 435 are sampled for each position ct,h in the classification latent expression space 440, where the autoregressive block 445 can access the preselected markers 405.
[0061] The autoregressive temporal model 400 allows sampling of the classification latent space 440 to produce seemingly plausible expressions consistent with the audio input 428. According to Bayes' rule, the probability of the latent embedding c1:T,1:H given the audio input a1:T can be decomposed into (9).
[0062] Equation 9 includes the temporal causality in the decomposition, that is, the category ct,h at time t depends only on the current and past audio information a≤t, and not on the future content a1:T. In some specific instances, the autoregressive block 445 includes a temporal CNN with four convolutional layers that expand incrementally along the time axis. In some specific instances, the convolutions are masked so that for the prediction of ct,h, the model can only access information from all category heads in the past c<t,1:H, and in the current time step ct,<h, it can access information from the aforementioned category heads (see the blocks before selected block 405 in the timeline). To train the autoregressive block 445, the audio encoder 442 maps the facial expressions and audio sequences (x1:T, a1:T) in the training set to their classification embeddings (see Equation 1). The autoregressive block 445 is optimized using teacher forcing and cross-entropy loss on the latent classification labels. During inference, an autoregressive time model 400 is used to sequentially sample and classify the facial expression code at each position ct,h.
[0063] Figure 5 illustrates a graph 500 of a latent classification space 540 (e.g., classification spaces 340 and 440) with a classifier that clusters based on facial expression input for some specific examples. Graph 500 includes a lower face mesh 521A, a composite mesh 521B, and an upper face mesh 521C (collectively referred to below as "face mesh 521") in the latent classification space 540. The composite mesh 521B successfully incorporates upper face motion and lip synchronization from different input modalities. In some specific examples, the classification latent space 540 may be better than a continuous latent space to reduce computational complexity. In some specific examples, a continuous latent space can provide higher rendering realism.
[0064] Cross-modal disambiguation produces a structured classification latent space 540, where each input modality has a different effect on the face mesh 521. In some specific instances, model 500 produces two distinct sets of latent representations, Saudio and Sepr. Saudio contains latent codes obtained by fixing the facial expression input to a facial expression encoder (e.g., facial expression encoders 244 and 344) and altering the audio signal (lower face mesh 521A). Similarly, Sepr contains latent codes obtained by fixing the audio signal and altering the facial expression input (upper face mesh 521C). In the extreme case of perfect cross-modal disambiguation, Saudio and Sepr form two non-overlapping clusters 521A and 521C. A separating hyperplane 535 fitted to points in Saudio∪Sexpr facilitates the 2D projection of the observations. It should be noted that only minimal leakage exists between the clusters formed by Saudio and Sepr.
[0065] Figures 6A and 6B illustrate model results 600 based on specific examples, demonstrating the influence of audio and facial input on the lower face mesh 621A, the upper face mesh 621C, and the composite meshes 621B-1 (continuous) and 621B-2 (classification), which are collectively referred to below as "face mesh 621" and "composite mesh 621B". Face mesh 621 includes lower face vertices 610A, upper face vertices 610C, and transition vertices 610B (collectively referred to below as "vertices 610"). Face mesh 621 indicates which face vertices move the most by the latent representations within the clusters of Saudio (e.g., lower face vertex 610A), within the clusters of Sepr (e.g., upper face vertex 610C), and near the decision boundary (e.g., transition vertex 610B). While audio primarily controls the mouth region (e.g., lower face mesh 621A) and facial expression controls the upper face mesh 621C, latent representations near the decision boundary influence face vertices in all regions (vertex 610B), reflecting intuitive concepts that some representations of upper facial expressions are related to speech such as raised eyebrows. For example, in some specific instances, the loss LxMod (see Equation 6) makes clear cross-modal discrimination between upper and lower face movements. Additionally, it is noteworthy that, in addition to its influence on the lips and chin, audio also has a significant impact on the eyebrow region (see vertex 611A).
[0066] Figure 6B illustrates the variation of vertex 610B of the composite face mesh 621B. Note how the upper face motion shrinks toward the average expression of the continuous space 621B-1 (with only a few vertex movements, see vertex 611B-1), while the classification space 621B-2 allows for rich and varied upper face motions (see vertex 611B-2).
[0067] To maintain the stochastic nature of the continuous space (see grid 621B-1), the model predicts the mean and variance of each frame from which subsequent sampled representations originate. At inference time, for example, an autoregressive model predicts the mean and variance from the audio input and all past latent representations. The next embedding is then sampled from these mean and variance predictions. In some specific instances, the lip error and overall vertex error of the continuous space grid 621B-1 are larger than those of the classification latent space (see Table 2). Table 2
[0068] To evaluate the quality of lip synchronization achieved by the specific examples disclosed herein, the lip error of a single frame can be the maximum error of the lip apex, and the average total frame error is reported in the test set. Because the upper lip and corners of the mouth move much less than the lower lip, the average total lip apex error tends to mask inaccurate lip shapes, although the maximum lip apex error per frame is better correlated with perceived quality. Table 3 illustrates the lip apex errors of different models disclosed herein, including Character Animation of Speech Operations (VOCA), where deep-speech features include variations of Mel spectrograms, and models disclosed herein (e.g., Model 300 and Autoregressive Convolutional Model 400). Table 3 shows the average lower lip error per frame achieved by the autoregressive convolutional model. Table 3
[0069] As revealed in this paper, the quality of the model is entirely independent of the chosen moderating identity. As revealed in this paper, Table 4 compares the perceptual evaluation results from different models, where the baseline truth was judged by a total of 100 participants after three sub-tasks: a full-face comparison, a lip-synchronization comparison using only the area between the jaw and nose, and an upper-face comparison using the area of the face from the nose upwards. For each column, 400 pairs of short segments were evaluated, each containing a sentence spoken by an individual from a self-test set. Participants could choose to like one segment over another, or rate them equally well. Table 4
[0070] Figure 7 illustrates different facial expressions of different individuals 727-1, 727-2, and 727-3 (hereinafter collectively referred to as "individuals 727") under the same spoken expression 728, based on some specific examples. The spoken expression 728 is an English sentence comprising three phonetic parts (A, B, and C). Therefore, the 3D speech animation engine disclosed herein generates facial animations 751A-1, 751B-1, and 751C-1 for individual 727-1 and phonetic parts A, B, and C, respectively; generates facial animations 751A-2, 751B-2, and 751C-2 for individual 727-2; and generates facial animations 751A-3, 751B-3, and 751C-3 (hereinafter collectively referred to as "facial animation 751") for individual 727-3.
[0071] The lip shape is consistent with the individual speech parts A, B and C in individual 727. In addition, for each sequence, such as sequences 751A-1, 751A-2, 751A-3 ("Sequence 751A"); sequences 751B-1, 751B-2, 751B-3 ("Sequence 751B"); and sequences 751C-1, 751C-2, 751C-3 ("Sequence 751C"), unique and different upper facial movements such as eyebrow raising and blinking are produced separately.
[0072] Figure 8 illustrates, based on specific examples, facial animations 851A-1, 851B-1, and 851C-1 for individual 827-1 (hereinafter collectively referred to as "Face Animation 851-1"), and facial animations 851A-2, 851B-2, and 851C-2 for individual 827-2 (hereinafter collectively referred to as "Face Animation 851-2", and together referred to as "Face Animation 851" for individual 827). Retargeting is the process of mapping facial movements from one identity's face to the face of another identity. A typical application is in videos or computer games, where actors animate faces that are not their own.
[0073] The facial animation 851 is obtained from the speech portions derived from different individuals 827A, 827B and 827C by the 3D speech animation engine (see model 300) as disclosed herein. It can be seen that the facial animation 851 maintains the common features of different individuals, such as the shape of the lips, eye closure and eyebrow level of the neutral expression.
[0074] The template mesh used in the model belongs to the target individual 827. The 3D speech animation engine synthesizes the audio and the original animated face mesh into a classification latent code and decodes it into a face animation 851. In some specific instances, the face animation 851 can be obtained without an autoregressive model (e.g., autoregressive model 400).
[0075] Figure 9 illustrates the adjustments ("grid dubbing") made to facial expressions 951A-1, 951B-1, 951C-1, 951D-1, and 951E-1 (hereinafter collectively referred to as "English facial expressions 951-1") based on some specific examples of audio input 927-1-English and 927-2-Spanish (hereinafter collectively referred to as "multilingual audio input 927"); and 951A-2, 951B-2, 951C-2, 951D-2, and 951E-2 (hereinafter collectively referred to as "Spanish facial expressions 951-2").
[0076] In some specific instances, such as the 3D speech animation engine disclosed herein (see 3D speech animation engine 232), it can be applied to dubbing video for speech translation into multilingual audio input 927 that is completely consistent with lip movements in the original language. Facial expressions 951-1 and 951-2 (hereinafter collectively referred to as "facial expressions 951") have matching lip movements in the multilingual audio input 927 while maintaining the integrity of upper facial movements. Therefore, the 3D speech animation engine resynthesizes the lip movements in the new language 927-2. Because the classification latent space is lateral across modalities (see meshes 521 and 621), the lip movements are applied to the audio segment 927-2, but maintain overall upper facial movements such as blinking from the original segment (see lower facial meshes 521A and 621A with upper facial meshes 521C and 621C).
[0077] Figure 10 is a flowchart illustrating the steps of a method 1000 for embedding a 3D voice animation model into a virtual reality environment according to some specific examples. In some specific examples, method 1000 may be executed at least in part by a processor executing instructions in a client device or server as disclosed herein (see processor 212 and memory 220, client device 110 and server 130). In some specific examples, at least one or more of the steps in method 1000 may be executed by an application installed on the client device or a 3D voice animation engine including a multi-peak encoder and a multi-peak decoder (e.g., application 222, 3D voice animation engine 232, multi-peak encoder 240 and multi-peak decoder 250). Users can interact with the application in the client device via input and output elements as disclosed herein and a GUI (see input device 214, output device 216 and GUI 225). Multi-peak encoders may include audio encoders, facial expression encoders, convolution tools, and synthesis encoders as disclosed herein (e.g., audio encoder 242, facial expression encoder 244, convolution tool 246, and synthesis encoder 248). In some specific instances, methods consistent with this disclosure may include at least one or more steps of method 1000, which are performed in different orders, simultaneously, semi-simultaneously, or temporally overlapping.
[0078] Step 1002 includes identifying audio-related facial features from the individual's audio capture. In some specific instances, step 1002 further includes receiving the individual's audio capture from the virtual reality headset. In some specific instances, step 1002 further includes identifying the intensity and frequency of the audio capture from the individual, and relating the amplitude and frequency of the audio waveform to the geometry of the lower portion of the individual's face.
[0079] Step 1004 includes generating a first grid of the lower portion of the individual's face based on audio-related facial features. In some specific instances, step 1004 further includes adding blinking or eyebrow movements of the individual.
[0080] Step 1006 includes identifying facial features resembling facial expressions of an individual. In some specific instances, step 1006 further includes randomly selecting facial features resembling facial expressions based on previous sampling of facial expressions of multiple individuals. In some specific instances, step 1006 further includes associating upper facial features with speech features captured from the individual's audio. In some specific instances, step 1006 further includes using random sampling of facial expressions of multiple individuals collected during a training session in which a second individual is reading text or in a dialogue.
[0081] Step 1008 includes generating a second mesh of the upper portion of an individual's face based on facial features resembling an expression. In some specific instances, step 1008 further includes accessing a three-dimensional model of the face of an individual with a neutral expression.
[0082] Step 1010 includes forming a composite mesh using the first mesh and the second mesh. In some specific instances, step 1010 includes merging the lip shape in the first mesh into the eye closure in the second mesh across the individual's face.
[0083] Step 1012 includes determining the loss value of the synthetic mesh based on the baseline truth image of the individual.
[0084] Step 1014 includes generating a 3D model of the face of an individual using a synthetic mesh based on the loss value.
[0085] Step 1016 includes providing a 3D model of the individual's face to a display on a user device running an immersive reality application that includes the individual. In some specific instances, step 1016 includes receiving audio capture of the individual and image capture of the individual's face, and generating a second mesh using the image capture.
[0086] Figure 11 is a flowchart illustrating the steps of a method 1100 for training a 3D model to create a real-time 3D voice animation of an individual, according to some specific examples. In some specific examples, method 1000 may be executed at least in part by a processor executing instructions in a client device or server as disclosed herein (see processor 212 and memory 220, client device 110 and server 130). In some specific examples, at least one or more of the steps in method 1100 may be executed by an application installed on the client device or a 3D voice animation engine including a multi-peak encoder and a multi-peak decoder (e.g., application 222, 3D voice animation engine 232, multi-peak encoder 240 and multi-peak decoder 250). Users may interact with the application on the client device via input and output elements as disclosed herein and a GUI (see input device 214, output device 216 and GUI 225). Multi-peak encoders may include audio encoders, facial expression encoders, convolution tools, and synthesis encoders as disclosed herein (e.g., audio encoder 242, facial expression encoder 244, convolution tool 246, and synthesis encoder 248). In some specific instances, methods consistent with this disclosure may include at least one or more steps of method 1100, which are performed in different orders, simultaneously, semi-simultaneously, or temporally overlapping.
[0087] Step 1102 includes determining a first correlation value of facial features based on the audio waveform from the first body. In some specific instances, step 1102 further includes determining a second correlation value of upper facial features. In some specific instances, step 1102 includes identifying facial features based on the intensity and frequency of the audio waveform.
[0088] Step 1104 includes generating a first mesh for the lower portion of the face based on facial features and a first correlation value. In some specific instances, step 1104 further includes generating a second mesh for the upper portion of the face based on upper facial features and a second correlation value, and forming a composite mesh using the first mesh and the second mesh.
[0089] Step 1106 includes updating the first correlation value based on the difference between the first grid of the first volume and the reference truth image.
[0090] Step 1108 includes providing a 3D model of a voice-animated face to an immersive reality application accessed by a user device based on the difference between the first mesh of the first person and the reference truth image. In some specific instances, step 1108 further includes forming a 3D model of a voice-animated face using a synthetic mesh. In some specific instances, step 1108 includes determining a loss value of the first mesh based on the reference truth image of the first person. In some specific instances, step 1108 includes updating a first correlation value of the facial features based on the audio waveform from the second person.
[0091] Hardware Overview
[0092] Figure 12 is a block diagram illustrating an exemplary computer system 1200, which can implement the client and server of Figures 1 and 2, as well as the methods of Figures 10 and 11. In some cases, the computer system 1200 can be implemented using hardware or a combination of hardware and software that is dedicated to a server, integrated into another entity, or distributed across multiple entities.
[0093] Computer system 1200 (e.g., client 110 and server 130) includes a bus 1208 or other communication mechanism for transmitting information, and a processor 1202 (e.g., processor 212) coupled to bus 1208 for processing information. By way of example, computer system 1200 may implement one or more processors 1202. Processor 1202 may be a general-purpose microprocessor, microcontroller, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), programmable logic device (PLD), controller, state machine, gate control logic, discrete hardware component, or any other suitable entity capable of performing computation or other manipulations of information.
[0094] In addition to hardware, computer system 1200 may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, protocol stack, database management system, operating system, or a combination of one or more of the following stored in memory 1204 (e.g., memory 220), such as random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable PROM (EPROM), etc., temporary storage, hard disk, removable disk, CD-ROM, DVD, or any other suitable storage device coupled to bus 1208 for storing information and instructions to be executed by processor 1202. Processor 1202 and memory 1204 may be supplemented by or incorporated into a dedicated logic circuit system.
[0095] These instructions may be stored in memory 1204 and implemented in one or more computer program products, such as computer program instructions encoded on a computer-readable medium, by any method known to those skilled in the art, for computer system 1200 to execute or control the operation of the computer system. These instructions include, but are not limited to, the following computer languages: data-oriented languages (e.g., SQL, dBase), system languages (e.g., C, Objective-C, C++, Assembly), architecture languages (e.g., Java, .NET), and application languages (e.g., PHP, Ruby, Perl, Python). Instructions can also be implemented in computer languages, such as array languages, feature-oriented languages, assembly languages, production languages, command-line interface languages, compiled languages, parallel languages, waveform bracket languages, data stream languages, data structure languages, declarative languages, esoteric languages, extended languages, fourth-generation languages, functional languages, interactive mode languages, interpreted languages, repetitive languages, serial-based languages, small languages, logic-based languages, machine languages, macro languages, metaprogramming languages, multiparadigm languages, numerical analysis, non-English languages, object-oriented classification languages, object-oriented prototype-based languages, field-rule languages, programming languages, reflection languages, rule-based languages, instruction code processing languages, stack-based languages, synchronization languages, syntax processing languages, visual languages, Wrth languages, and XML-based languages. Memory 1204 can also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by processor 1202.
[0096] The computer programs discussed herein may not necessarily correspond to files in a file system. Programs may be stored in a portion of a file that holds other programs or data (e.g., stored in one or more instruction codes in a markup language file), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., a file that stores portions of one or more modules, subroutines, or code). Computer programs may be deployed to execute on a single computer or on multiple computers located in one location or distributed across multiple locations and interconnected by a communication network. The programs and logic flows described in this specification may be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and producing output.
[0097] The computer system 1200 further includes a data storage device 1206, such as a magnetic disk or optical disk, coupled to bus 1208 for storing information and instructions. The computer system 1200 may be coupled to various devices via input / output module 1210. Input / output module 1210 may be any input / output module. An exemplary input / output module 1210 includes a data port such as a USB port. Input / output module 1210 is configured to connect to communication module 1212. An exemplary communication module 1212 (e.g., communication module 218) includes a network interface card, such as an Ethernet card and a modem. In some configurations, input / output module 1210 is configured to connect to a plurality of devices, such as input device 1214 (e.g., input device 214) and / or output device 1216 (e.g., output device 216). Example input device 1214 includes a keyboard and a pointing device, such as a mouse or trackball, through which a user can provide input to computer system 1200. Other types of input devices 1214 may also be used to provide interaction with the user, such as tactile input devices, visual input devices, audio input devices, or brain-computer interface devices. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and any form of input can be received from the user, including acoustic input, voice input, tactile input, or brainwave input. Example output device 1216 includes a display device for displaying information to the user, such as a liquid crystal display (LCD) monitor.
[0098] According to one embodiment of this disclosure, the computer system 1200 can be used to implement the client 110 and the server 130 in response to the processor 1202 executing one or more sequences of one or more instructions contained in the memory 1204. These instructions may be read into the memory 1204 from another machine-readable medium, such as a data storage device 1206. Execution of the instruction sequence contained in the main memory 1204 causes the processor 1202 to execute the program steps described herein. One or more processors in a multiprocessor configuration may also be used to execute the instruction sequence contained in the memory 1204. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions to implement the various embodiments of this disclosure. Therefore, the embodiments of this disclosure are not limited to any particular combination of hardware circuitry and software.
[0099] Various embodiments of the subject matter described in this specification can be implemented in a computing system that includes back-end components, such as a data server, or middleware components, such as an application server, or front-end components, such as a client computer having a graphical user interface or web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end components, middleware components, or front-end components. The components of the system can be interconnected via any form or medium of digital data communication, such as a communication network. A communication network (e.g., network 150) may include, for example, any or more of a LAN, WAN, Internet, and the like. Furthermore, a communication network may include, but is not limited to, any or more of, the following topologies, including bus networks, star networks, ring networks, mesh networks, star bus networks, tree or hierarchical networks, or the like. A communication module may be, for example, a modem or an Ethernet card.
[0100] Computer system 1200 may include a client and a server. The client and server are generally geographically separated and typically interact via a communication network. The relationship between the client and server is generated by computer programs running on individual computers and having a client-server relationship with each other. Computer system 1200 may be, for example, but not limited to, a desktop computer, a laptop computer, or a tablet computer. Computer system 1200 may also be embedded in another device, for example, but not limited to, a mobile phone, a PDA, a mobile audio player, a Global Positioning System (GPS) receiver, a video game console, and / or a set-top box.
[0101] As used herein, the terms "computer-readable storage media" or "computer-readable media" refer to any one or more media that participate in providing instructions to the processor 1202 for execution. Such media can take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical discs or magnetic disks, such as data storage device 1206. Volatile media include dynamic memory, such as memory 1204. Transmission media include coaxial cables, copper wires, and optical fibers, including wires forming bus 1208. Common forms of machine-readable media include, for example, floppy disks, floppy disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, DVDs, any other optical media, punch cards, paper tapes, any other physical media with a perforated pattern, RAM, PROMs, EPROMs, FLASH EPROMs, any other memory chips or cartridges, or any other media that can be read by a computer. Machine-readable storage media can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a component of a substance that affects machine-readable transmission signals, or a combination of one or more of these.
[0102] To illustrate the interchangeability of hardware and software, various exemplary blocks, modules, components, methods, operations, instructions, and algorithms have been described in general terms of their functionality. Whether this functionality is implemented as hardware, software, or a combination of both depends on the specific application and design constraints imposed on the overall system. Those skilled in the art can implement the described functionality in different ways for each specific application.
[0103] As used herein, the phrase "at least one of" preceding a series of items, separating any one of those items with the terms "and" or "or," modifies the entire list, rather than each member of the list (i.e., each item). The phrase "at least one of" does not require selection of at least one item; in fact, the phrase allows for the inclusion of at least one of any of those items and / or at least one of any combination of those items and / or at least one of each of those items. By way of example, the phrases "at least one of A, B, and C" or "at least one of A, B, or C" each refer to only A, only B, or only C; any combination of A, B, and C; and / or at least one of each of A, B, and C.
[0104] With respect to the use of the terms "comprising," "having," or the like in the implementation or claims, this term is intended to be inclusive in a manner similar to how the term "including" is interpreted when "comprising" is used as a transitional word in the claims. The word "illustrative" is used herein to mean "serving as an example, instance, or illustration." Any specific example described herein as "illustrative" should not be construed as being better or more advantageous than other specific examples.
[0105] Unless specifically stated otherwise, references to elements in the singular are not intended to mean "one and only one," but rather "one or more." All structural and functional equivalents of elements in various configurations described throughout this disclosure, known or later to those skilled in the art, are expressly incorporated herein by reference and are intended to be covered by the subject matter. Furthermore, nothing disclosed herein is intended to be used exclusively by the public, whether or not it is directly described in the foregoing description. No element of a clause should be interpreted in accordance with paragraph 6 of 35 USC §112, unless the element is explicitly described using the phrase "component for..." or, in the case of a method clause, using the phrase "step for...".
[0106] Although this specification contains numerous details, such details should not be construed as limiting the scope of what may be claimed, but rather as a description of specific embodiments of the subject matter. Certain features described in this specification in the context of separate concrete examples may also be implemented in combination in a single concrete example. Conversely, various features described in the context of a single concrete example may also be implemented separately or in any suitable sub-combination in multiple concrete examples. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this manner, one or more features from the claimed combination may be removed from that combination in some circumstances, and the claimed combination may be for sub-combinations or variations thereof.
[0107] The subject matter of this specification has been described with respect to a particular configuration, but other configurations may be implemented and are within the scope of the following claims. For example, although operations are depicted in a particular order in the drawings, this should not be construed as requiring the specific order shown or sequential order of these operations, or the execution of all depicted operations to achieve the desired result. The actions described in the claims may be performed in different orders and still achieve the desired result. As an example, the program depicted in the drawings may not require the specific order shown or sequential order to achieve the desired result. In some environments, multitasking and parallel processing may be advantageous. Furthermore, the separation of the various system components in the configurations described above should not be construed as requiring this separation in all configurations, and it should be understood that the described program components and systems may be substantially integrated together in a single software product or packaged into multiple software products. Other variations are within the scope of the following claims. 100: Architecture 110: Client Device / Client 130: Server 150: Network 152: Database 200: Block Diagram 214: Input Device 216: Output Device 212-1: Processor 212-2: Processor 218-1: Communication Module 218-2: Communication Module 220-1: Memory 220-2: Memory 222: Application 225: GUI 227-1: Data Packet 227-2: Data Packet 232: 3D Speech Animation Engine 240: Multi-peak Encoder 242: Audio Encoder 244: Facial Expression Encoder 246: Convolution Tool 248: Synthesis Encoder 250: Multi-peak Decoder 252: Training Database 300: Mapping / Model 327: Neutral Face Mesh / Template Mesh / Mesh 328: Speech Signal / T Speech Segment / Audio Sequence / Speech Sequence / Audio Input / Audio Signal 329: Animated Face Mesh / T Face Mesh / Expression Sequence / Expression Signal / Expression Input / Face Mesh / Face Expression 330: Fusion Block 335: Latent Classification Head 340: Classification Latent Space 341: Encoded Expression 34 2: Audio Encoder 344: Expression Encoder / Facial Expression Encoder 348: Synthesis Encoder 350: Decoder 351: Facial Expression Mesh / Animation Sequence 400: Autoregressive Model / Autoregressive Temporal Model / Autoregressive Convolutional Model 405: Preselected Labels / Selected Blocks 428: Audio Input / Audio Signal 435: Audio Conditioning Latent Code 440: Classification Latent Space / Classification Latent Expression Space / Classification Space 442: Audio Encoder 445: Autoregressive Block 500: Graph / Model 521A: Lower Facial Mesh / Non-overlapping Cluster 521B: Synthesis Mesh 521C: Upper Facial Mesh / Non-overlapping Cluster 535: Separating Hyperplane 540: Classification Latent Space 600: Model Results610A: Lower face vertex 610B: Transformed vertex / Vertex 610C: Upper face vertex 611A: Vertex 611B-1: Vertex 611B-2: Vertex 621A: Lower face mesh 621B-1: Composite mesh / Continuous space / Mesh / Continuous space mesh 621B-2: Composite mesh / Classification space 621C: Upper face mesh 727-1: Individual 727-2: Individual 727-3: Individual 728: Verbal expression 751A-1: Facial animation / Sequence 751A-2: Facial animation / Sequence 7 51A-3: Facial Animation / Sequence 751B-1: Facial Animation / Sequence 751B-2: Facial Animation / Sequence 751B-3: Facial Animation / Sequence 751C-1: Facial Animation / Sequence 751C-2: Facial Animation / Sequence 751C-3: Facial Animation / Sequence 827: Individual 827-1: Individual 827-2: Individual 827A: Individual 827B: Individual 827C: Individual 851: Facial Animation 851A-1: Facial Animation 851A-2: Facial Animation 851B-1: Facial Animation 851 B-2: Facial Animation 851C-1: Facial Animation 851C-2: Facial Animation 927-1: Audio Input 927-2: Audio Input / New Language / Audio Clip 951A-1: Facial Expressions 951A-2: Facial Expressions 951B-1: Facial Expressions 951B-2: Facial Expressions 951C-1: Facial Expressions 951C-2: Facial Expressions 951D-1: Facial Expressions 951D-2: Facial Expressions 951E-1: Facial Expressions 951E-2: Facial Expressions 1000: Method 1002: Step 1 004: Step 1006: Step 1008: Step 1010: Step 1012: Step 1014: Step 1016: Step 1100: Method 1102: Step 1104: Step 1106: Step 1108: Step 1200: Computer System 1202: Processor 1204: Memory 1206: Data Storage Device 1208: Bus 1210: Input / Output Module 1212: Communication Module 1214: Input Device 1216: Output Device A: Voice Section B: Voice Section C: Voice Section [Simplified Explanation of the Diagram]
[0010] [Figure 1] illustrates an instance architecture suitable for providing 3D facial animation from voice to an immersive reality environment based on some specific examples.
[0011] [Figure 2] is a block diagram illustrating an example server and client from the architecture of Figure 1 according to certain aspects of this disclosure.
[0012] [Figure 3] illustrates a block diagram of mapping facial meshes and speech signals to a classified facial expression space based on some specific examples.
[0013] [Figure 4] illustrates a block diagram of an autoregressive model based on some specific examples, including pre-selected labels.
[0014] [Figure 5] illustrates the visualization of the potential space clustered based on facial expression input based on some specific examples.
[0015] [Figures 6A] to [Figures 6B] illustrate the effects of audio and facial expression input on a face grid based on some specific examples.
[0016] [Figure 7] illustrates different facial expressions for different identities under the same verbal expression, based on some specific examples.
[0017] [Figure 8] illustrates the retargeting of neutral facial expressions from different identities based on some specific examples, such as lip shape, eye closure and eyebrow level.
[0018] [Figure 9] illustrates the adjustment of facial expressions based on audio languages (English / Spanish) according to some specific examples.
[0019] [Figure 10] is a flowchart illustrating the steps in a method for using a 3D model of a voice-animated human face in an immersive reality application, based on some specific examples.
[0020] [Figure 11] is a flowchart illustrating the steps in a method for generating a 3D model of a human face animated by voice, based on some specific examples.
[0021] [Figure 12] is a block diagram illustrating an example computer system, which can be used to implement the client and server of Figures 1 and 2, as well as the methods of Figures 10 and 11.
[0022] In the drawings, unless otherwise stated, elements referred to by the same or similar markings have the same or similar features and descriptions.
Claims
1. A computer-implemented method comprising: identifying an audio-related facial feature from an individual's audio capture; generating a first mesh for a lower portion of the individual's face based on the audio-related facial feature; identifying an expression-like facial feature of the individual; generating a second mesh for an upper portion of the individual's face based on the expression-like facial feature; forming a composite mesh using the first mesh and the second mesh; determining a loss value of the composite mesh based on a reference image of the individual; generating a three-dimensional model of the individual's face using the composite mesh based on the loss value; and providing the three-dimensional model of the individual's face to a display in a client device running an immersive reality application including the individual.
2. The computer implementation method of claim 1 further includes receiving the audio capture of the individual from a virtual reality headset group.
3. The computer implementation method of claim 1, wherein identifying an audio-related facial feature includes identifying an intensity and a frequency of the audio captured from the individual, and relating an amplitude and a frequency of an audio waveform to a geometry of the lower portion of the individual's face.
4. The computer implementation method of claim 1, wherein generating the first grid includes the blinking or eyebrow movement of one of the individuals.
5. The computer implementation method of claim 1, wherein identifying one of the facial features resembling an expression of an individual comprises randomly selecting the facial feature resembling an expression based on prior sampling of one of the facial expressions of multiple individuals.
6. The computer implementation method of claim 1, wherein identifying an individual’s facial expression feature includes associating an upper facial feature with a speech feature captured from the individual’s audio.
7. The computer implementation method of claim 1, wherein identifying one of the facial features resembling an expression of an individual comprises randomly sampling one of a plurality of individual facial expressions collected during a training session of one of a second individual reading a text or being in a conversation.
8. The computer implementation method of claim 1, wherein generating a second mesh includes access to a three-dimensional model of the face of the individual having a neutral expression.
9. The computer implementation method of claim 1, wherein forming a composite mesh includes the face of the individual, and successively merging a lip shape in the first mesh into an eye closure in the second mesh.
10. The computer implementation method of claim 1, further comprising receiving the audio capture of the individual and an image capture of one of the individual's faces, and generating the second grid comprising using the image capture.
11. A system comprising: one or more processors; and a memory storing instructions that, when executed by the one or more processors, cause the system to: identify an audio-related facial feature from audio capture of an individual; generate a first mesh for a lower portion of a face of the individual based on the audio-related facial feature; identify an expression-like facial feature of the individual; generate a second mesh for an upper portion of a face of the individual based on the expression-like facial feature; form a composite mesh using the first mesh and the second mesh; determine a loss value of the composite mesh based on a reference truth image of the individual; generate a three-dimensional model of the face of the individual using the composite mesh based on the loss value; and provide the three-dimensional model of the face of the individual to a display in a client device running an immersive reality application including the individual.
12. The system of claim 11, wherein the one or more processors further execute instructions to receive the audio capture of the individual from a virtual reality headset group.
13. The system of claim 11, wherein, in order to identify one of the facial features resembling an expression of an individual, the one or more processors execute instructions to randomly select the facial feature resembling an expression based on a previous sample of one of the facial expressions of a plurality of individuals.
14. The system of claim 11, wherein, in order to identify one of the individual's facial features resembling an expression, the one or more processors execute instructions to correlate an upper facial feature with one of the speech features captured from the individual's audio.
15. The system of claim 11, wherein, in order to identify one of the facial features resembling an individual's facial expressions, the one or more processors execute instructions to randomly sample one of a plurality of individual facial expressions, which are collected during a training session of one of a second individual in a dialogue or while reading a text.
16. A computer-implemented method comprising: determining a first correlation value of a facial feature based on an audio waveform from a first person; generating a first mesh of a lower portion of a face based on the facial feature and the first correlation value; updating the first correlation value based on a difference between the first mesh of the first person and a reference truth image; and providing a three-dimensional model of the voice-animated face to an immersive reality application accessed by a user terminal device based on the difference between the first mesh of the first person and the reference truth image.
17. The computer implementation method of claim 16 further includes: determining a second correlation value of an upper facial feature; generating a second mesh of an upper portion of the face based on the upper facial feature and the second correlation value; forming a composite mesh using the first mesh and the second mesh; and forming the three-dimensional model of the face animated by speech using the composite mesh.
18. The computer implementation method of claim 16, wherein determining a first correlation value of a facial feature includes identifying the facial feature based on the intensity and frequency of an audio waveform.
19. The computer implementation method of claim 16 further includes determining a loss value of the first grid based on a reference truth image of the first body.
20. The computer implementation method of claim 16 further includes updating the first correlation value of a facial feature based on an audio waveform from a second body.