Robust facial animation of video and audio
Through a two-stage semi-supervised machine learning model, combining video and audio inputs, realistic facial animation is generated, which solves the jitter and delay problems of animation solutions in the prior art, improves the natural motion performance of the user avatar and reduces computing resources and energy consumption.
Patent Information
- Application Number
- CN202480004981.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-25
- Filing Date
- 2024-01-25
- Publication Date
- 2025-07-04
AI Technical Summary
In existing online platforms, the animation scheme of user avatars has problems such as motion jitter, delay, exaggeration or minimization, and the computing resources and energy consumption are high, making it difficult to achieve realistic real-time facial animation.
Using a two-stage semi-supervised machine learning model, combining video and audio input, generates realistic facial animation through a modular mix of video FACS weights and audio FACS weights.
It improves the stability and realistic nature of facial animation, reduces computing resources and energy consumption, and enhances the natural sports performance of the user avatar.
Smart Images

Figure CN120266167A_ABST
Abstract
Description
Cross - Reference to Related Applications
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 440,993, filed on January 25, 2023, with the title "ROBUST FACIAL ANIMATION FROM VIDEO USING NEURAL NETWORKS", the entire content of which is incorporated herein by reference. Technical Field
[0002] Embodiments generally relate to computer - based virtual experiences, and more particularly, to methods, systems, and computer - readable media for real - time robust facial animation of video. Background Art
[0003] Some online platforms (e.g., gaming platforms, media exchange platforms, etc.) allow users to connect with each other, interact with each other (e.g., within a game), create games, and share information with each other over the Internet. Users of online platforms can participate in multiplayer game environments or virtual environments (e.g., three - dimensional environments), design customized game environments, design characters and avatars, decorate avatars, exchange virtual items / objects with other users, communicate with other users using audio or text messaging, etc. Environments such as the metaverse or multiverse environments can also enable participating users to share, sell, or trade the objects they create with other users.
[0004] Interactions between users can use an interaction interface that includes a display of the user's avatar. Animating an avatar typically involves the user inputting requested poses, actions, and other similar pre - configured animation details and presenting the animation based on the user's input. This traditional solution has drawbacks.
[0005] The background art description provided herein is intended to introduce the background of the present disclosure. What the current inventors have done, insofar as it is described in this background section, and aspects of the specification that may not constitute prior art at the time of filing, whether expressly or implicitly, shall not be regarded as prior art to the present disclosure. Summary of the Invention
[0006] Embodiments of the present application relate to the real-time automatic creation of robust facial animations based on video and audio. According to one aspect, a computer-implemented method includes: receiving a plurality of input video frames; receiving a plurality of input audio frames and a mixing term, wherein the plurality of input audio frames includes audio associated with the plurality of input video frames; obtaining video Facial Action Coding System (FACS) weights from a first trained machine learning model based on the input video frames; obtaining audio FACS weights from a second trained machine learning model based on the input audio frames; combining the video FACS weights and the audio FACS weights to obtain final FACS weights, wherein the combination is at least partially based on the mixing term; and outputting the final FACS weights to drive the facial animation of a 3D model.
[0007] Various embodiments and variations of the computer-implemented method are disclosed.
[0008] In some embodiments, the first trained machine learning model and the second trained machine learning model are trained in a two-stage semi-supervised training process.
[0009] In some embodiments, the first trained machine learning model includes at least one encoder and at least three task-specific decoders.
[0010] In some embodiments, the first task-specific decoder among the at least three task-specific decoders is used to output a predicted head pose, the second task-specific decoder among the at least three task-specific decoders is used to output the probability that a face is visible in an input video frame among the received plurality of input video frames, and the third task-specific decoder among the at least three task-specific decoders is used to output facial key points.
[0011] In some embodiments, at least one of the at least three task-specific decoders includes a causal convolutional layer applied in the time dimension.
[0012] In some embodiments, the second trained machine learning model includes at least one encoder and at least two task-specific decoders.
[0013] In some embodiments, the first task-specific decoder among the at least two task-specific decoders is used to output audio FACS weights, and the second task-specific decoder among the at least two task-specific decoders is used to output the mixing term.
[0014] According to another aspect, a system is provided. The system includes: a memory storing instructions thereon; and a processing device coupled to the memory, the processing device configured to access the memory and execute the instructions, wherein the instructions cause the processing device to perform operations, the operations including: receiving a plurality of input video frames; receiving a plurality of input audio frames and a mixing term, wherein the plurality of input audio frames includes audio associated with the plurality of input video frames; obtaining video Facial Action Coding System (FACS) weights from a first trained machine learning model based on the input video frames; obtaining audio FACS weights from a second trained machine learning model based on the input audio frames; combining the video FACS weights and the audio FACS weights to obtain final FACS weights, wherein the combining is at least partially based on the mixing term; and outputting the final FACS weights to drive facial animation of a 3D model.
[0015] Various embodiments and variations of the system are disclosed.
[0016] In some embodiments, the first trained machine learning model and the second trained machine learning model are trained in a two-stage semi-supervised training process.
[0017] In some embodiments, the first trained machine learning model includes at least one encoder and at least three task-specific decoders.
[0018] In some embodiments, the first task-specific decoder among the at least three task-specific decoders is configured to output a predicted head pose, the second task-specific decoder among the at least three task-specific decoders is configured to output a probability that a face is visible in an input video frame among the received plurality of input video frames, and the third task-specific decoder among the at least three task-specific decoders is configured to output facial key points.
[0019] In some embodiments, at least one of the at least three task-specific decoders includes a causal convolutional layer applied in the temporal dimension.
[0020] In some embodiments, the second trained machine learning model includes at least one encoder and at least two task-specific decoders.
[0021] In some embodiments, the first task-specific decoder among the at least two task-specific decoders is configured to output audio FACS weights, and the second task-specific decoder among the at least two task-specific decoders is configured to output the mixing term.
[0022] According to another aspect, a non-transitory computer-readable medium is provided. Instructions are stored on the non-transitory computer-readable medium, and the instructions, when executed by a processing device, cause the processing device to perform operations including: receiving a plurality of input video frames; receiving a plurality of input audio frames and a mixing term, wherein the plurality of input audio frames includes audio associated with the plurality of input video frames; obtaining video Facial Action Coding System (FACS) weights from a first trained machine learning model based on the input video frames; obtaining audio FACS weights from a second trained machine learning model based on the input audio frames; combining the video FACS weights and the audio FACS weights to obtain final FACS weights, wherein the combining is at least partially based on the mixing term; and outputting the final FACS weights to drive facial animation of a 3D model.
[0023] Various embodiments and variations of the non-transitory computer-readable medium are disclosed.
[0024] In some embodiments, the first trained machine learning model and the second trained machine learning model are trained in a two-stage semi-supervised training process.
[0025] In some embodiments, the first trained machine learning model includes at least one encoder and at least three task-specific decoders.
[0026] In some embodiments, the first task-specific decoder among the at least three task-specific decoders is used to output a predicted head pose, the second task-specific decoder among the at least three task-specific decoders is used to output the probability of facial visibility in the received input video frames, and the third task-specific decoder among the at least three task-specific decoders is used to output facial key points.
[0027] In some embodiments, at least one of the at least three task-specific decoders includes a causal convolutional layer applied in the temporal dimension.
[0028] In some embodiments, the second trained machine learning model includes at least one encoder and at least two task-specific decoders, and wherein the first task-specific decoder among the at least two task-specific decoders is used to output audio FACS weights, and the second task-specific decoder among the at least two task-specific decoders is used to output the mixing term.
[0029] According to yet another aspect, parts, features, and implementation details of the above systems, methods, and non-transitory storage media can be combined to form additional aspects, including aspects that omit and / or modify some or a portion of individual components or features, including additional components or features and / or other modifications; and all such modifications are within the scope of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a diagram of an example network environment according to some embodiments.
[0031] Figure 2 is a diagram of a facial animation engine according to some embodiments.
[0032] Figure 3 is a diagram of a video animation component according to some embodiments.
[0033] Figure 4 is a diagram of an audio animation component according to some embodiments.
[0034] Figure 5A and Figure 5B shows a flowchart of an example method for training parts of a video animation component according to some embodiments.
[0035] Figure 6A and Figure 6B shows a flowchart of an example method for training parts of an audio animation component according to some embodiments.
[0036] Figure 6C shows an example training environment for training parts of an audio animation component according to some embodiments.
[0037] Figure 7 is a flowchart of an example method for real-time robust facial animation based on video and audio according to some embodiments.
[0038] Figure 8 is a block diagram of an example computing device that can be used to implement one or more features described herein. DETAILED DESCRIPTION
[0039] One or more embodiments described herein relate to real-time robust animation based on video and audio. Features can include automatically creating an animation of a three-dimensional (3D) avatar based on input video and input audio received from a client device.
[0040] The features described herein provide for automatically detecting a face in a video, automatically detecting / recognizing facial actions based on phonemes in audio, regressing parameters for animating a 3D avatar based on the detected face and facial actions, and creating an animation of the 3D avatar based on the above parameters. A trained model receives audio and video inputs and outputs audio and video facial action coding system (FACS) weights. A modular mixing component combines the audio FACS weights and the video FACS weights to obtain the final FACS weights for facial animation.
[0041] The trained model can be deployed on a client device for use by users who wish to automatically create animations for their associated avatars. The client device can also be configured to communicate operably with an online platform (e.g., a virtual experience (VE) platform), where the avatar associated with the client device can be richly animated for presentation in a communication interface (e.g., a video chat), a virtual experience (e.g., a richly animated face on a representative virtual body), an animated video sent to other users (e.g., a recording of an animated avatar sent via a chat function or other feature), and other parts of the online platform.
[0042] Online virtual experience platforms (also referred to as “user-generated content platforms” or “user-generated content systems”) provide users with a variety of ways to interact with each other. For example, users of an online virtual experience platform can create experiences, games, or other content or resources (e.g., characters, graphics, items, etc. for use in a game in a virtual world) within the platform.
[0043] Users of an online virtual experience platform can cooperate towards a common goal in a game or game creation, share various virtual items, send electronic messages to each other, etc. Users of an online virtual experience platform can interact with the environment and play games, e.g., including characters (avatars) or other game objects and mechanisms. The online virtual experience platform can also allow users of the platform to communicate with each other. For example, users of an online virtual experience platform can communicate with each other using voice messages, text messaging, video messaging, or a combination of the above (e.g., via a voice chat). Some online virtual experience platforms can provide a virtual three-dimensional environment where users can use their own avatars or virtual representations to represent themselves.
[0044] To help enhance the entertainment value of an online virtual experience platform, the platform can provide a facial animation engine to automatically animate avatars. The facial animation engine can allow users to request or select animation options, including, for example, the animation of the face or body of an avatar based on a real-time video feed and / or real-time audio sent from a client device.
[0045] For example, a user can allow an application on a user device associated with the online virtual experience platform to access the camera and microphone. The video created by the camera can be parsed to extract poses or other information that helps animate the avatar based on the extracted poses. Additionally, the audio captured by the microphone can be parsed to extract facial movements. The user can also enhance the facial animation by inputting directional controls to move other body parts or exaggerate facial poses.
[0046] In some embodiments discussed herein, where user data (e.g., a user's image, a user's video, a user's audio, user population information, user behavior data on a platform, user search history, items purchased and / or viewed, user friendships on a platform, etc.) can be obtained or used, options are provided to the user to control whether and how such information is collected, stored, or used. That is, the embodiments discussed herein collect, store, and / or use user information when receiving explicit user authorization and in compliance with applicable regulations.
[0047] The user can control whether to allow a program or function to collect user information about that particular user or other users associated with the program or function. Options are presented (e.g., via a user interface) to each user for whom information is to be collected, to allow the user to exert control over the information collection related to that user, providing permission or authorization as to whether to collect information and which portions of the information are to be collected. Additionally, certain data can be modified in one or more ways before storage or use to remove personally identifiable information. As an example, a user's identity can be modified (e.g., by substituting with a pseudonym, numerical values, etc.) such that personally identifiable information cannot be determined. In another example, a user's geographical location can be generalized to a larger area (e.g., city, zip code, state, country, etc.).
[0048] As described herein, automatically detecting faces in a video, automatically detecting / identifying phonemes in audio for facial movements, regressing parameters for animating a 3D avatar based on the detected faces and facial movements, and creating an animation of the 3D avatar based on the above parameters can be provided by a trained model. Model training can include two or more stages. For example, first, a part of the model (e.g., an encoder) is trained, and then the output of the trained encoder is used to train another part of the model (e.g., one or more decoders). Such two-stage training can overcome the drawbacks associated with typical automatic animation, including jitter, latency, exaggerated movements, minimized movements, unrealistic movements, etc. The technical effects and advantages of two-stage training can include reducing the training cycle, thereby improving energy consumption (e.g., by reducing the training time, reducing the energy usage of the computing resources used for training), reducing the number of computing cycles (e.g., by reducing the training time, thereby reducing the total number of computing cycles), improving storage requirements (e.g., by using synthetic data and real data, parts of the training data can be created during training rather than stored), etc.
[0049] As further described herein, a deployed trained model can be used to provide automatic animation of avatars that are both realistic and stable. For example, realistic animation can refer to animation having natural movements that are consistent between facial actions and audio. For example, stable animation can refer to the absence of jerky transitions between consecutive frames of a frame sequence of the animation. The trained model can overcome disadvantages associated with typical automatic animation, including jitter, latency, exaggerated movements, minimized movements, unrealistic movements, and the like. Technical effects and advantages of the deployed trained model can include improved energy consumption of a client device (e.g., by deploying a trained model with adjustable level-of-detail, computational cycles can be reduced, and thus energy usage of the client device can be reduced), improved energy consumption of a server device (e.g., by deploying a trained model with improved efficiency in generating FACS output, computational cycles can be reduced, and thus energy usage of the server device can be reduced), and the like.
[0050] These and other advantages of the present disclosure will be apparent from the included specification and the related drawings. Turning now to Figure 1 , an example system architecture in which a model can be trained and / or deployed is described in detail. Figure 1 : Example system architecture
[0051] Figure 1 FIG. 11 shows an example network environment 100 in accordance with some embodiments of the present disclosure. The network environment 100 (also referred to herein as a “system”) includes an online virtual experience platform 102, a first client device 110, and a second client device 116 (generally referred to herein as “client devices 110 / 116”), all connected via a network 122. The online virtual experience platform 102 can include a virtual experience (VE) engine 104, one or more virtual experiences 105, a communication engine 106, a facial animation engine 107, and a data store 108, among others. The client device 110 can include a virtual experience application 112. The client device 116 can include a virtual experience application 118. A user 114 can use the client device 110 to interact with the online virtual experience platform 102, and a user 120 can use the client device 116 to interact with the online virtual experience platform 102.
[0052] The network environment 100 is provided for illustration. In some embodiments, the network environment 100 can include the same, fewer, more, or different elements configured in the same or different manner as Figure 1 shown.
[0053] In some embodiments, network 122 may include a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or a wide area network (WAN)), a wired network (e.g., Ethernet), a wireless network (e.g., an 802.11 network, a Wi-Fi network, or a wireless LAN (WLAN)), a cellular network (e.g., a long term evolution (LTE) network), routers, hubs, switches, server computers, or a combination thereof.
[0054] In some embodiments, data store 108 may be a non-transitory computer-readable memory (e.g., random access memory), a cache, a drive (e.g., a hard disk drive), a flash drive, a database system, or another type of component or device capable of storing data. Data store 108 may also include multiple storage components (e.g., multiple drives or multiple databases) that may also span multiple computing devices (e.g., multiple server computers).
[0055] In some embodiments, online virtual experience platform 102 may include a server having one or more computing devices (e.g., a cloud computing system, a rack server, a server computer, a physical server cluster, a virtual server, etc.). In some embodiments, the server may be included in online virtual experience platform 102, may be a stand-alone system, or may be part of another system or platform.
[0056] In some embodiments, online virtual experience platform 102 may include one or more computing devices (e.g., rack servers, router computers, server computers, personal computers, mainframe computers, laptop computers, tablet computers, desktop computers, etc.), a data store (e.g., a hard disk, a memory, a database), a network, software components, and / or hardware components that may be used to perform operations on online virtual experience platform 102 and provide users with access to online virtual experience platform 102. Online virtual experience platform 102 may also include a website (e.g., one or more web pages) or application backend software that may be used to provide users with access rights to the content provided by online virtual experience platform 102. For example, users may access online virtual experience platform 102 using virtual experience applications 112 / 118 on client devices 110 / 116, respectively.
[0057] In some embodiments, the online virtual experience platform 102 may include a social network that provides connections among users or a user-generated content system that allows users (e.g., end users or consumers) to communicate with other users through the online virtual experience platform 102, where the communication may include voice chat (e.g., synchronous and / or asynchronous voice communication), video chat (e.g., synchronous and / or asynchronous video communication), or text chat (e.g., synchronous and / or asynchronous text-based communication). In some embodiments of the present disclosure, a "user" may represent a single individual. However, other embodiments of the present disclosure cover a "user" as an entity controlled by a group of users or an automated source (e.g., a creative user). For example, a group of individual users united as a community or group in a user-generated content system may be regarded as a "user".
[0058] In some embodiments, the online virtual experience platform 102 may be a virtual game platform. For example, the game platform may provide single-player or multi-player games to a user community, where the user community may access (e.g., user-generated games or other games) or interact with the games through the network 122 using the client devices 110 / 116. In some embodiments, the games (also referred to herein as "video games", "online games", or "virtual games") may be, for example, two-dimensional (2D) games, three-dimensional (3D) games (e.g., 3D user-generated games), virtual reality (VR) games, or augmented reality (AR) games. In some embodiments, users may search for games and game items and participate in games with other users in one or more games. In some embodiments, games may be played in real time with other users of the games.
[0059] In some embodiments, other collaboration platforms may be used instead of or in addition to the online virtual experience platform 102 in combination with the robust animation features described herein. For example, social network platforms, video chat platforms, messaging platforms, user content creation platforms, virtual meeting platforms, etc. may be used together with the robust animation features described herein to facilitate the fast, robust, and accurate representation of a user's facial movements onto a virtual avatar based on input video and / or audio.
[0060] In some embodiments, "gameplay" may refer to the interaction of one or more players within a game or experience (e.g., VE105) using a client device (e.g., 110 and / or 116), or the interaction presented on a display or other output device of the client device 110 or 116.
[0061] One or more virtual experiences 105 are provided by an online virtual experience platform. In some embodiments, the virtual experience 105 can include an electronic file that can be executed or loaded using software, firmware, or hardware for presenting virtual content (e.g., digital media items) to an entity. In some embodiments, the virtual experience application 112 / 118 can be executed and, in combination with the virtual experience engine 104, render the virtual experience 105. In some embodiments, the virtual experience 105 can have a set of common rules or common goals, and the environment of the virtual experience 105 shares the set of common rules or common goals. In some embodiments, different virtual experiences can have rules or goals that are different from each other. Similarly, or alternatively, some virtual experiences can have no goals at all and are intended for users to interact with each other in any social way.
[0062] In some embodiments, a virtual experience can have one or more environments (also referred to herein as "game environments" or "virtual environments") in which multiple environments can be linked. Examples of environments can be three-dimensional (3D) environments. One or more environments of the virtual experience 105 can be collectively referred to herein as a "world" or "virtual world" or "virtual universe" or "metaverse". An example of a world can be the 3D world of the virtual experience 105. For example, a user can build a virtual environment that links to another virtual environment created by another user. The characters of a virtual experience can cross virtual boundaries and enter adjacent virtual environments.
[0063] It should be noted that 3D environments or 3D worlds use graphics that use a three-dimensional representation of geometric data representing virtual content (or at least present the content as 3D content whether or not a 3D representation of geometric data is used). 2D environments or 2D worlds use graphics that use a two-dimensional representation of geometric data representing virtual content.
[0064] In some embodiments, the online virtual experience platform 102 can host one or more virtual experiences 105 and can allow users to interact with the virtual experiences 105 (e.g., search for games, VE-related content, or other content) using the virtual experience applications 112 / 118 of the client devices 110 / 116. Users of the online virtual experience platform 102 (e.g., 114 and / or 120) can play the virtual experiences 105, create the virtual experiences 105, interact with the virtual experiences 105, or build the virtual experiences 105, search for the virtual experiences 105, communicate with other users, create and build objects (e.g., also referred to herein as "items" or "game objects" or "virtual game items") of the virtual experiences 105, and / or search for objects. For example, when generating user-generated virtual items, a user can create characters, decorations for the characters, one or more virtual environments for an interactive experience, or structures used in the virtual experience 105, etc.
[0065] In some embodiments, a user may purchase, sell, or trade virtual objects, such as in-platform currency (e.g., virtual currency), with other users of the online virtual experience platform 102. In some embodiments, the online virtual experience platform 102 may transmit virtual content to virtual experience applications (e.g., 112, 118). In some embodiments, virtual content (also referred to herein as "content") may refer to any data or software instructions associated with the online virtual experience platform 102 or virtual experience applications (e.g., virtual objects, experiences, user information, videos, images, commands, media items, etc.).
[0066] In some embodiments, a virtual object (e.g., also referred to herein as an "item" or "object" or "virtual game item") may refer to an object used, created, shared, or otherwise depicted in the virtual experience application 105 of the online virtual experience platform 102 or the virtual experience applications 112 or 118 of the client devices 110 / 116. For example, virtual objects may include parts, models, characters, tools, weapons, clothing, buildings, vehicles, currency, flora, fauna, components of the foregoing (e.g., windows of a building), etc.
[0067] It should be noted that the online virtual experience platform 102 that provides the hosted virtual experience 105 is for illustration and not limitation. In some embodiments, the online virtual experience platform 102 may host one or more media items, which may include communication messages from one user to one or more other users. Media items may include, but are not limited to, digital videos, digital movies, digital photos, digital music, audio content, melodies, website content, social media updates, e-books, e-magazines, digital newspapers, digital audiobooks, e-journals, web blogs, real simple syndication (RSS) feeds, electronic comic books, software applications, etc. In some embodiments, a media item may be an electronic file that may be executed or loaded using software, firmware, or hardware for presenting the digital media item to an entity.
[0068] In some embodiments, the virtual experience 105 can be associated with a specific user or group of users (e.g., a private experience), or can be widely available to users of the online virtual experience platform 102 (e.g., a public experience). In some embodiments, in the case where the online virtual experience platform 102 associates one or more virtual experiences 105 with a specific user or group of users, the online virtual experience platform 102 can use user account information (e.g., user account identifiers such as a username and password) to associate the specific user with the virtual experience 105. Similarly, in some embodiments, the online virtual experience platform 102 can use developer account information (e.g., developer account identifiers such as a username, and a password) to associate a specific developer or group of developers with the virtual experience 105.
[0069] In some embodiments, the online virtual experience platform 102 or the client devices 110 / 116 can include a virtual experience engine 104 or a virtual experience application 112 / 118. The virtual experience engine 104 can include a virtual experience application similar to the virtual experience application 112 / 118. In some embodiments, the virtual experience engine 104 can be used for the development or execution of the virtual experience 105. For example, the virtual experience engine 104 can include a rendering engine (a "renderer") for 2D, 3D, VR, or AR graphics, a physics engine, a collision detection engine (and collision response), a sound engine, a scripting function, an animation engine, an artificial intelligence engine, a network function, a streaming function, a storage management function, a threading function, a scene graph function, or support for animated videos, and other functions. The components of the virtual experience engine 104 can generate commands (e.g., rendering commands, collision commands, physics commands, etc.) that assist in calculating and rendering the virtual experience. In some embodiments, the virtual experience applications 112 / 118 of the client devices 110 / 116 can work independently and / or in cooperation with the virtual experience engine 104 of the online virtual experience platform 102.
[0070] In some embodiments, the online virtual experience platform 102 and the client devices 110 / 116 respectively execute virtual experience engines (104, 112, and 118). The online virtual experience platform 102 using the virtual experience engine 104 can execute some or all of the virtual experience engine functions (e.g., generating physics commands, rendering commands, etc.), or divert some or all of the virtual experience engine functions to the virtual experience engine 104 of the client device 110. In some embodiments, there can be different ratios between the virtual experience engine functions executed by each virtual experience 105 on the online virtual experience platform 102 and the virtual experience engine functions executed on the client devices 110 and 116.
[0071] For example, the virtual experience engine 104 of the online virtual experience platform 102 can be used to generate physical commands in the case of a collision occurring between at least two game objects, while additional virtual experience engine functions (e.g., generating rendering commands) can be diverted to the client device 110. In some embodiments, the ratio of the virtual experience engine functions executed on the online virtual experience platform 102 and the client device 110 can be changed based on interactivity conditions (e.g., dynamically). For example, if the number of users participating in a particular virtual experience 105 exceeds a threshold number, the online virtual experience platform 102 can execute one or more virtual experience engine functions previously executed by the client device 110 or 116.
[0072] For example, a user can interact with the virtual experience 105 on the client devices 110 and 116, and can send control instructions (e.g., user input, such as right, left, up, down, user selection, or character position and speed information, etc.) to the online virtual experience platform 102. After receiving the control instructions from the client devices 110 and 116, the online virtual experience platform 102 can send interaction instructions (e.g., the position and speed information of the character participating in the virtual experience, or commands, such as rendering commands, collision commands, etc.) to the client devices 110 and 116 based on the control instructions. For example, the online virtual experience platform 102 can perform one or more logical operations on the control instructions (e.g., using the virtual experience engine 104) to generate interaction instructions for the client devices 110 and 116. In other instances, the online virtual experience platform 102 can transfer one or more control instructions from one client device 110 to other client devices (e.g., 116) participating in the virtual experience 105. The client devices 110 and 116 can use the instructions and render the experience to present on the displays of the client devices 110 and 116.
[0073] In some embodiments, the control instructions can refer to instructions indicating actions in the experience of the user's character or avatar. For example, the control instructions can include user input for controlling actions in the experience, such as right, left, up, down, user selection, gyroscope position and orientation data, force sensor data, etc. The control instructions can include character position and speed information. In some embodiments, the control instructions are directly sent to the online virtual experience platform 102. In other embodiments, the control instructions can be sent from the client device 110 to another client device (e.g., 116), where the other client device uses the local virtual experience engine 104 to generate play instructions. The control instructions can include instructions for playing voice communication messages or other sounds of another user on an audio device (e.g., speakers, headphones, etc.), instructions for moving the character or avatar, and other instructions.
[0074] In some embodiments, the interaction or game instructions may refer to instructions that allow the client device 110 (or 116) to render a virtual experience (such as a multiplayer game). The instructions may include one or more of user input (e.g., control instructions), character position and speed information, or commands (e.g., physical commands, rendering commands, collision commands, etc.). As described more fully herein, other instructions may include facial animation instructions extracted from an input video of the user's face to guide the animation of the representative virtual face of the virtual avatar in real time. Thus, while the interaction instructions may include input for the user to directly control some body movements of the character, the interaction instructions may also include poses extracted from the user's video.
[0075] In some embodiments, a character (or generally, a virtual avatar) is composed of components, where one or more of these components can be selected by the user and these components are automatically connected together to assist the user in editing. One or more characters (also referred to herein as "avatars" or "models") can be associated with the user, where the user can control the character to facilitate the user's interaction with the virtual experience 105. In some embodiments, the character may include components such as body parts (e.g., hair, arms, legs, etc.) and accessories (e.g., t-shirts, glasses, decorative images, tools, etc.). In some embodiments, the customizable body parts of the character include head types, body part types (arms, legs, torso, and hands), face types, hair types, and skin types, etc. In some embodiments, the customizable accessories include clothing (e.g., shirts, pants, hats, shoes, glasses, etc.), weapons, or other tools.
[0076] In some embodiments, the user can also control the size of the character (e.g., height, width, or depth) or the size of the components of the character. In some embodiments, the user can control the proportion of the character (e.g., blocky, anatomical, etc.). It can be noted that in some embodiments, the character may not include character objects (e.g., body parts, etc.), but the user can (in the absence of character objects) control the character to facilitate the user's interaction with the game (e.g., a puzzle game, where there are no rendered character game objects, but the user still controls the character to control in-game actions).
[0077] In some embodiments, a component (such as a body part) can be a basic geometric shape such as a block, cylinder, sphere, etc., or some other basic shapes such as a wedge, ring, tube, channel, etc. In some embodiments, the creator module can publish the user's role for other users of the online virtual experience platform 102 to view or use. In some embodiments, the user can use a user interface (e.g., a developer interface) and create, modify, or customize a role, other virtual experience objects, virtual experience 105, or virtual experience environment with or without a script (or with or without an application programming interface (API)). It should be noted that, for illustrative purposes and not limitation, the role is described as having a humanoid form. It can also be noted that the role can have any form, such as a vehicle, an animal, an inanimate object, or other creative forms.
[0078] In some embodiments, the online virtual experience platform 102 can store the user-created role in the data store 108. In some embodiments, the online virtual experience platform 102 maintains a role directory and an experience directory that can be presented to the user through the virtual experience engine 104, virtual experience 105, and / or client device 110 / 116. In some embodiments, the experience directory includes images of different experiences stored on the online virtual experience platform 102. In addition, the user can select a role from the role directory (e.g., a role created by the user or other users) to participate in the selected experience. The role directory includes images of the roles stored on the online virtual experience platform 102. In some embodiments, one or more roles in the role directory may have been created or customized by the user. In some embodiments, the selected role can have role settings that define one or more components of the role.
[0079] In some embodiments, the user's role can include a configuration of components, where the configuration and appearance of the components and more generally the appearance of the role can be defined by role settings. In some embodiments, at least part of the role settings of the user's role can be selected by the user. In other embodiments, the user can select a role with default role settings or role settings selected by other users. For example, the user can select a default role from the role directory with predefined role settings, and the user can further customize the default role by changing some role settings (e.g., adding a shirt with a custom logo). The online virtual experience platform 102 can associate the role settings with a specific role.
[0080] In some embodiments, client device 110 or 116 may include a computing device such as a personal computer (PC), a mobile device (e.g., a laptop computer, a mobile phone, a smartphone, a tablet computer, or a netbook computer), an Internet TV, a game console, etc. In some embodiments, client device 110 or 116 may also be referred to as a "user device". In some embodiments, one or more client devices 110 or 116 may be connected to the online virtual experience platform 102 at any given moment. It should be noted that the number of client devices 110 or 116 provided is for illustration purposes only and not for limitation. In some embodiments, any number of client devices 110 or 116 may be used.
[0081] In some embodiments, each client device 110 or 116 may respectively include an instance of the virtual experience application 112 or 118. In one embodiment, the virtual experience application 112 or 118 may allow a user to use and interact with the online virtual experience platform 102, e.g., search for a specific experience or other content, control a virtual character in a virtual game hosted by the online virtual experience platform 102, or view or upload content (e.g., virtual experience 105, images, video items, web pages, documents, etc.). In one example, the virtual experience application may be a web application (e.g., an application that operates in conjunction with a web browser) that can access, retrieve, present, or navigate content provided by a web server (e.g., virtual characters in a virtual environment, etc.). In another example, the virtual experience application may be a native application (e.g., a mobile application, an app, a program) that is installed on the client device 110 or 116 and executed locally, and allows the user to interact with the online virtual experience platform 102. The virtual experience application may render, display, or present content (e.g., web pages, user interfaces, media viewers) to the user. In an embodiment, the virtual experience application may also include an embedded media player (e.g., Flash player) embedded in a web page.
[0082] According to aspects of the present disclosure, the virtual experience application 112 / 118 may be an online virtual experience platform application for a user to build, create, edit, upload content to the online virtual experience platform 102, and interact with the online virtual experience platform 102 (e.g., play the virtual experience 105 hosted by the online virtual experience platform 102 and interact with the virtual experience 105). Thus, the online virtual experience platform 102 may provide the virtual experience application 112 / 118 to the client device 110 or 116. In another example, the virtual experience application 112 / 118 may be an application downloaded from a server.
[0083] In some embodiments, a user may log in to the online virtual experience platform 102 via a virtual experience application. The user may access the user account by providing user account information (e.g., username and password), wherein the user account is associated with one or more roles that can be used to participate in one or more virtual experiences 105 of the online virtual experience platform 102.
[0084] Generally speaking, if appropriate, functions described as being performed by the online virtual experience platform 102 may also be performed by the client device 110 or 116 or the server in other embodiments. Additionally, functions attributed to a particular component may be performed by different or multiple components operating together. The online virtual experience platform 102 may also be accessed as a service provided to other systems or devices via a suitable application programming interface (API), and is thus not limited to use within a website.
[0085] In some embodiments, the online virtual experience platform 102 may include a communication engine 106. In some embodiments, the communication engine 106 may be a system, application, or module that allows the online virtual experience platform 102 to provide video communication functions to users, and may be a function that allows users to experience virtual chat or virtual video conferencing using the online virtual experience platform 102 and their associated virtual representations. For example, a user may design and build a virtual avatar and use the virtual avatar via the chat function provided by the communication engine 106.
[0086] In various embodiments, the platform 102 may provide a chat function and / or avatar animation function via the communication engine 106 and other components such as an animation engine or a facial animation engine. In these and other embodiments, the communication engine 106 may animate the avatar representing the user by automatically generating FACS weights, and the user is presented as the animated avatar while using the chat function. For example, the communication engine 106 may send audio data and FACS weights instead of sending video via the network 122. In response to receiving the audio data and FACS weights, the receiving VE application 112 may animate the avatar and present the animated avatar in a graphical user interface (GUI) or another interface. Thus, the user can "chat via the avatar" instead of being presented in a live video.
[0087] In some embodiments, the online virtual experience platform 102 may include a facial animation engine 107. In some embodiments, the facial animation engine 107 may be a system, application, or module that implements facial detection (e.g., via video analysis), facial motion detection (e.g., via audio analysis), and a regression model to create robust real-time animation of a user's avatar or character face based on audio and video signals. The animation may be based on the user's actual face and spoken voice, and thus may include smiles, blinks, winks, frowns, head poses, and other poses extracted from the input video of the user's face and the audio of phrases spoken by the user. Although illustrated as being executed directly on the online virtual experience platform 102, it should be understood that facial detection, audio processing, and the regression model may be implemented, for example, on each client device 110, 116, or on other devices.
[0088] The facial animation engine 107 in combination with the communication engine 106 may provide some or all of the above-described avatar chat functionality. For example, the facial animation engine 107 may receive input video and input audio and output associated FACS weights to the communication engine 106 (or the VE engine 104). The communication engine 106 may then (or substantially simultaneously) send the FACS weights and user audio data for presentation via a receiving chat interface and associated components.
[0089] In some embodiments, the facial animation engine 107 may also be used to render realistic avatar animation in a virtual experience 105, in a 3D environment during a game or virtual event, etc. Thus, although described in some embodiments as being related to the chat functionality, the features provided by the facial animation engine 107 may also be used to animate avatars in scenarios such as in a game, in the rendering of a scene (e.g., a video game cutscene or an animated video), in a virtual experience (e.g., to enhance user enjoyment and engagement and improve the immersive experience provided by the 3D environment), etc.
[0090] In the following, reference Figure 2 、 Figure 3 and Figure 4 details various functions and components associated with the facial animation engine 107. Figure 2 : Facial Animation Engine
[0091] Figure 2 is a diagram of the facial animation engine 107 according to some embodiments. As shown, the facial animation engine 107 includes a video animation component 204. The video animation component 204 may be a software component deployed on a computing device that includes one or more models trained to provide a robust facial animation output, i.e., FACS weights s, determined from input video frames 202 v。
[0092] An input video frame 202 can be received from a camera, an image capture device, or other devices / technologies (e.g., such as stored video, recreated video, synthetic video, etc.). The input video frame can be sent to the video animation component 204. Thereafter, the video animation component 204 can determine and output FACS weights s v 。For example, the video animation component 204 can output FACS weights based on the video and thus can include facial actions such as actions of the eyebrows, eyes, eyelids, nose, ears, cheeks, chin, neck, lips, mouth, etc.
[0093] The facial animation engine 107 also includes an audio animation component 210. The audio animation component 210 can be a software component deployed on a computing device, and the software component includes one or more models that are trained to provide a robust facial animation output determined from the input audio 208, i.e., FACS weights s a 。
[0094] An input audio 208 can be received from a microphone, an audio capture device, or other devices / technologies (e.g., such as stored audio, recreated audio, synthetic audio, etc.). The input audio 208 can be sent to the audio animation component 210. Thereafter, the audio animation component 210 can determine and output FACS weights s a and an audio signal α. For example, the audio animation component 210 can output FACS weights based on the audio and / or phonemes and thus can include facial actions such as actions of the lips, mouth, tongue, chin, etc. The FACS weights s a can also include values representing the chin moving down, the lips puckering, the lips stretching, and other values distinguishable from the audio information. In addition, the audio animation component 210 can determine whether the speaker is vocalizing and adjust and / or output α representing the detected speech and / or utterance.
[0095] The facial animation engine 107 can also include a modular mixing component 206. The modular mixing component 206 can be deployed on a computing device and is used to mix or fuse s v FACS weights and s a FACS weights based on a weighted linear function and / or the speech activity signal α. For example, in some embodiments, the modular mixing component 206 can mix or fuse s v FACS weights and s a FACS weights according to: Equation 1: FACS i = FACS v,i (1 - ∝)+∝(w a,i FACSa,i +w v,i FACS v,i )
[0096] In Equation 1, the result of mixing or fusing is represented by FACS i and is output by the video animation component as FACS v,i and output by the audio animation component as FACS a,i is obtained. Note that i represents the index value of a single frame or group of frames of the input video. Additionally, α represents the voice activity signal, w a,i and w v,i are the mixing weights of the audio result and the video result, respectively. Note that if no speech is detected in the input audio 208, the audio animation component 210 outputs α as zero and drives the final FACS weight 212 according to the prediction of the video animation component 204. If voice activity is detected, the value of α changes from 0 to 1, and the final FACS weight 212 becomes a weighted combination of the results of the video animation component 204 and the audio animation component 210.
[0097] Hereinafter, the operations and components associated with the video animation component 204 are described in detail with reference to Figure 3 FIG. Figure 3 : Video animation component
[0098] Figure 3 is a diagram of the video animation component 204 according to some embodiments. As shown, the video animation component 204 is used to receive an input video frame 302. The input video frame can be provided to an initial face candidate network 304. In some embodiments, the input video frame can also be directed such that the face candidate network 304 (e.g., elements 310 and 320) is bypassed when a face is detected.
[0099] The face candidate network 304 can include a candidate network 306 (labeled P-Net) and a refinement network 308 (labeled R-Net). The P-Net 306 can be used to determine face candidates. For example, the P-Net 306 can be used to determine whether a person's face is within the video frame of the input video frame 302.
[0100] The R-Net 308 can be used to filter the determined face candidates received from the P-Net 306. The R-Net 308 can also be used to refine the determined face candidates within an appropriate bounding box. When the user's face is detected and the detected face is refined within the bounding box, the face candidate network 304 can be bypassed in a logical OR (OR) function 310. Bypassing the face candidate network 304 can increase the speed of outputting the video FACS weights s v by reducing the overall computational cycles.
[0101] The P-Net 306 and / or the R-Net 308 can be part of a multi-task cascaded convolutional network (MTCNN) of the video animation component 204. In this regard, the P-Net 306 and the R-Net 308 can be the initial stages of the MTCNN.
[0102] When refining and / or bypassing the face candidate network 304 within the bounding box, two decision blocks 312 and 320 of two levels of detail can be operated in an interlocking manner. For example, for the first level of detail, the decision block 312 can direct the refined input video frame and the bounding box to the B-Net 314, and the B-Net 314 represents the basic level of detail. For example, for the second level of detail, the decision block 320 can direct the refined input video frame and the bounding box to the H-Net 314, and the H-Net 314 represents a higher level of detail than the first level of detail. Additionally, in some embodiments, the B-Net 314 and the H-Net 322 can be used to provide a highest level of detail that is higher than both the first level of detail and the second level of detail.
[0103] The B-Net 314 (also referred to as the Basic-Net) includes a feature encoder 316 and up to four decoders 318. The four decoders 318 include a keypoint decoder D l 、a FACS decoder D vl 、a head pose decoder D z 、and a face probability decoder D p . In some embodiments, the feature encoder 316 is a convolutional neural network. In some embodiments, the keypoint decoder D l 、the FACS decoder D vl 、the head pose decoder D z 、and the face probability decoder D p are relatively small time-aware neural networks for inputting the high-level features output by the feature encoder 316.
[0104] It should be noted that different from the convolution applied in the spatial dimension, the causal convolution layer of the decoder 318 is applied in the time dimension. In this regard, the decoder 318 (and 326, described below) allows implicit learning of the filtering function, reducing jitter while maintaining responsiveness in the FACS weight output. Additionally, the relatively small size of the decoder allows the size of the feature encoder 316 to be reduced without affecting the output quality, because the availability of time information allows the decoder 318 to compensate for the lower-quality features output by the feature encoder 316.
[0105] Note that the keypoint decoder D l 、the FACS decoder D vl 、the head pose decoder D z 、and the face probability decoder D p are trained on a continuous sequence of video data rather than on a single image, and thus can be trained to associate the temporal relationships between input video frames 302. The keypoint decoder D l can be used to decode face keypoints for use as input to the H-Net 322. When bypassing the H-Net 322 (e.g., in the logical OR operation 328 and the switch / judgment box 320), the FACS decoder D vl can be used to generate FACS weights as output. The head pose decoder D z can be used to output a head pose signal z v to assist in achieving a realistic head pose in the animation of the avatar. Additionally, the face probability decoder D p can be used to operate the judgment box 312 to bypass the face candidate network 304.
[0106] Other forms of bypass blocks can be implemented to replace the specific bypass components shown to assist in bypassing the face candidate network 304 and / or the B-Net 314 and / or the H-Net 322.
[0107] The H-Net 322 (also known as the HiFi-Net) includes a feature encoder 324 and a decoder 326. The decoder 326 is the FACS decoder D vh . In some embodiments, the feature encoder 324 is a convolutional neural network. In some embodiments, the FACS decoder D vh is a relatively small temporal-aware neural network for inputting the high-level features output by the feature encoder 324.
[0108] The B-Net 314 and the H-Net 322, as well as the corresponding encoders 316, 324, differ in terms of input resolution and the overall size of the associated convolutional neural network. For example, compared to the H-Net 322 and the associated encoder 324, the B-Net 314 and the associated encoder 316 have a smaller input resolution. Thus, in operation, a higher-resolution output can be obtained from the H-Net 322 compared to the B-Net 314. Accordingly, different client devices with different computing capabilities can be used to implement the face tracking features as described herein, where, through appropriate bypassing, low-end devices can use the B-Net 314 to provide automatic and robust animation, while high-end devices are capable of implementing either the B-Net 314 or the H-Net 322.
[0109] Additionally, B-Net 314 and H-Net 322, along with the corresponding encoders 316, 324, can be used to operate on each frame (e.g., 30 frames-per-second (fps)) or downsample the input and operate at a lower frame rate (e.g., 15 fps). In some cases, generating the output at 15 fps may reduce the visual quality. However, in some embodiments, for the input frames "skipped" when the encoder is operating at 15 fps, the system can use the feature vectors from the associated encoder of the previous frame as input, which enables extrapolating the FACS output and generating results that are substantially similar or nearly identical to the output of running the encoder at 30 fps. It should be noted that such interpolation can be achieved through the causal convolutional architecture described above in the form of providing temporal filtering.
[0110] As shown, the video animation component 204 is used to output the video animation FACS weights s v . As described above with reference to the modular hybrid component 206 and Equation 1, the video animation FACS weights s v can be used for robust facial animation with or without the audio animation FACS weights. In embodiments where audio is available and / or selected by the user, the audio animation component 210 is operable to provide the audio animation FACS weights s a as output. Figure 4 : Audio animation component
[0111] Figure 4 is a diagram of the audio animation component 210 according to some embodiments. As shown, the audio animation component 210 can be used to receive the input audio 402. The input audio 402 can be captured by a microphone or, in some embodiments, synthesized for non-verbal users.
[0112] The input audio 402 can be processed by the feature extraction component 404 to extract features with a first hop length and a first window length. In some embodiments, the feature extraction component is used to extract twenty-six features with a hop length of 16 ms and a window length of 16 ms. Depending on any particular embodiment, other variations can be applied.
[0113] In some embodiments, five consecutive frames of the extracted features are packed together to form an input packet for the audio network 406. In some embodiments, the audio network 406 is arranged to include a feature encoder 408 and at most two decoders 410, 412. The feature encoder 408 may include stacked causal convolutional layers, followed by batch normalization. Since the causal convolutional layers are arranged to examine (e.g., in the temporal sense) a series of values from the past, the causal convolutional layers can expand the receptive field without introducing latency. Additional layers may be arranged to operate along the time dimension to preserve the temporal embedding. The additional layers may be recurrent neural networks, such as long short-term memory (LSTM) layers.
[0114] The decoder 410 may be used to decode the audio FACS weights s a , while the decoder 412 may be used to provide the audio signal α. Thus, compared to larger decoders for multiple tasks, the decoders 410 and 412 can be dedicated to specific tasks, thereby improving efficiency and reducing computational costs.
[0115] The video animation component 204 and the audio animation component 210 are trained in a multi-stage / multi-level training process, in which the associated encoders are first trained, and then the decoders are trained while freezing / fixing the weights of the encoders based on the previous training data. Hereinafter, the training process of the video animation component 204 is described with reference to FIG. 5, and the training process of the audio animation component 210 is described with reference to FIG. 6. Figure 5: Training Video Animation Component
[0116] Figure 5A and Figure 5B FIG. 500 shows a flow chart of an example method for training the various parts of the video animation component 204 according to some embodiments. In some embodiments, the method 500 may be implemented, for example, on a server system (e.g., an online virtual experience platform 102 as Figure 1 shown). In some embodiments, some or all of the method 500 may be implemented on a system (e.g., one or more client devices 110 and 116 as Figure 1 shown), and / or on a server system and one or more client systems. In the described example, the implementation system includes one or more processors or processing circuits, and one or more storage devices (e.g., a database or other accessible storage device). In some embodiments, different components of one or more servers and / or clients may execute different blocks or other parts of the method 500. The method 500 starts at block 502.
[0117] At block 502, facial keypoints are obtained from real video training data and synthetic video training data. For example, it is not straightforward to directly train a FACS model in a supervised manner from synthetic data because the domain gap between synthetic data and actual real faces is large. Therefore, the feature representations learned by a network trained on a synthetic dataset are different from those required for real data. Consequently, a model trained only on a synthetic dataset cannot generalize. Thus, both the real dataset and the synthetic dataset are used for training, and facial keypoints are obtained in the first stage of encoder training. Although facial keypoints are not used in the regression step described above with reference to Figure 3 they assist in training the model to learn a representation that is effective for both real data and synthetic data. After block 502 is block 504.
[0118] At block 504, the training data is provided to the encoder model for training. For example, real video frames and synthetic video frames are provided as inputs to the relevant encoder of the video animation component 204. After block 504 is block 506.
[0119] At block 506, the obtained facial keypoints are provided to the keypoint decoder D l . In this way, the keypoint decoder D l can be trained simultaneously with the encoder. After block 506 is block 508.
[0120] At block 508, the encoder and / or the keypoint decoder are adjusted based on two or more loss terms. For example, two or more terms can be linearly combined so that the encoder and / or the keypoint decoder can be jointly trained.
[0121] The first loss term can be L LMK , where L LMK is the location loss on the keypoints. The root mean square error of the regression location can be used for the location loss. The second loss term can be L CON , where L CON is the consistency loss on the keypoints. The consistency loss on the keypoints can encourage the keypoint predictions to be equivariant under different transformations and allow the use of real images with unannotated keypoints. Based on the corresponding outputs and the computed loss terms, the internal weights of each of the encoder and the keypoint decoder can be adjusted. After block 508 is block 510.
[0122] At block 510, it is determined whether all the training data has been input, or whether training has otherwise been completed. For example, completing a threshold number of training epochs or inputting a threshold amount of images can indicate the completion of the training process. Other thresholds or conditions during training can also apply. If training is complete, then after block 510 can be block 512. Otherwise, after block 510 is block 504, where additional training data is input to the encoder.
[0123] At block 512, if training is completed, the internal weights of the encoder are frozen, and the second stage of training can commence at block 514. After block 514 is block 516.
[0124] At block 516, the output of the trained encoder is obtained. For example, an additional training set or other training data can be provided as input to the trained encoder. The encoder output can be obtained for training the individual decoders. After block 516 is block 518.
[0125] At block 518, the encoder output is provided to the untrained decoders (e.g., FACS decoder D vl , head pose decoder D z , and face probability decoder D p ). For example, the high-level encoded features output by the trained encoder are input into each decoder. After block 518 is block 520.
[0126] At block 520, the decoder is adjusted based on three or more loss terms. For example, three or more loss terms can be linearly combined. The first loss term can be L POS , L POS is the positional loss on the FACS weights. Mean squared error can be used as the L POS loss term. The second loss term can be L VEL , L VEL is the velocity loss. By promoting the smoothness of dynamic expressions, using L VEL helps reduce jitter. The third loss term can be L ACC , L ACC is the acceleration regularization loss. A regularization term for acceleration is added to reduce FACS weight jitter (e.g., keeping the weights low to maintain responsiveness). After block 520 is block 522.
[0127] At block 522, it is determined whether all the training data has been input, or whether training has otherwise been completed. For example, completing a threshold number of training rounds or inputting a threshold amount of encoded features can indicate the completion of the training process. Other thresholds or conditions during training can also apply. If training is completed, block 524 can follow block 522. Otherwise, after block 522 is block 516, where additional training data (e.g., encoded features from the trained encoder) is input into the decoder.
[0128] At block 524, if training is completed, the internal weights of the decoder are frozen, and the trained encoder and decoder are output and / or deployed.
[0129] As described above, the video animation component 204 and associated sub-components are trained using a two-stage training method. As described below, the audio animation component 210 is also trained using a two-stage method. Figure 6: Training Audio Animation Component
[0130] Figure 6A and Figure 6B FIG. 600 is a flow diagram of an example method for training portions of an audio animation component according to some embodiments. In some embodiments, method 600 may be implemented, for example, on a server system (e.g., an online virtual experience platform 102 as shown in Figure 1 ). In some embodiments, some or all of method 600 may be implemented on a system (e.g., one or more client devices 110 and 116 as shown in Figure 1 ), and / or on a server system and one or more client systems. In the described example, the implementing system includes one or more processors or processing circuits, and one or more storage devices (e.g., a database or other accessible storage device). In some embodiments, different components of one or more servers and / or clients may execute different blocks or other portions of method 600. Method 600 begins at block 602.
[0131] At block 602, an audio sample with time-aligned labels is obtained. Block 602 is followed by block 604.
[0132] At block 604, the training data is augmented. For example, the audio sample is augmented with various randomly selected noises and pitch transformations. To improve the model's robustness to different types of noise, the input audio sample may be augmented by adding randomly generated white noise or by selecting pre-recorded environmental noises (e.g., street, restaurant, wind, or rain).
[0133] To simulate people speaking in different virtual rooms, impulse response convolution may add reverb to the input audio. Additionally, pitch shifting, gain shifting, and speed changes may also be used to further augment the audio sample. Block 604 is followed by block 606.
[0134] At block 606, the augmented audio sample is provided to an encoder and an additional phoneme layer or phoneme decoder. For example, referring to Figure 6C , during an initial training phase labeled 680, the augmented audio sample 670 may be provided to encoder 408 and an additional phoneme decoder 675.
[0135] Although the phoneme decoder 675 is not a permanent component or is not used in operation, the phoneme prediction task can make the output clearer or more consistent. The intuition is that there is a strong correlation between phonemic speech and visual speech units (e.g., chin, lip, mouth movements). Therefore, training the phoneme recognition task is beneficial for internal representations that can be converted into continuous viseme sequences, represented by FACS curves. In some embodiments, the phoneme decoder 675 is implemented as a fully connected layer after embedding to estimate phoneme labels. Return to Figure 6A , box 606 is followed by box 608.
[0136] At box 608, the encoder is adjusted based on the loss term. For example, the unique connectionist temporal classification (CTC) loss can be calculated and used to adjust the encoder. Box 608 is followed by box 610.
[0137] At box 610, it is determined whether all the training data has been input or whether training has otherwise been completed. For example, completing a threshold number of training rounds or inputting a threshold amount of images can indicate the completion of the training process. Other thresholds or conditions during training are also applicable. If training is completed, box 610 can be followed by box 612. Otherwise, box 610 is followed by box 604, where additional training data is input to the encoder.
[0138] At box 612, if training is completed, the internal weights of the encoder are frozen, and the second stage of training can begin at box 614. Box 614 is followed by box 616.
[0139] At box 616, the output of the trained encoder is obtained. For example, an additional training set or other training data can be provided as input to the trained encoder. The encoder output can be obtained for training the decoder. Box 616 is followed by box 618.
[0140] At box 618, the encoder output (e.g., the phoneme-aware representation of the input data) is provided to the untrained decoder. For example, referring to Figure 6C , the second stage 685 of training can include training the decoders 410 and 412 using the phoneme-aware encoding features provided by the trained encoder 408 based on the input audio sample 680. It should be noted that the training data 670 includes audio samples from multiple speakers, while the training data 680 is based on a single speaker. In this way, the encoder is not biased towards the voice of a single speaker. Return to Figure 6B , box 618 is followed by box 620.
[0141] At block 620, the decoder is adjusted based on three or more loss terms. For example, three or more loss terms can be linearly combined. The first loss term can be L POS , L POS which is a positional loss on the FACS weights. Mean squared error can be used as the L POS loss term. The second loss term can be L VEL , L VEL which is a velocity loss. By encouraging smoothness of the dynamic expression, using L VEL helps reduce jitter. The third loss term can be L VAD , L VAD which is a cross-entropy loss. Note that L POS and L VEL in the audio network training can be implemented as smooth L1 loss, while in the video network training, the same loss terms are L2 loss. After block 620 is block 622.
[0142] At block 622, it is determined whether all the training data has been input, or whether training has otherwise been completed. For example, completion of a threshold number of training epochs or input of a threshold amount of encoded features can indicate completion of the training process. Other thresholds or conditions during training can also apply. If training is complete, then after block 622 can be block 624. Otherwise, after block 622 is block 616, where additional training data (e.g., encoded features from the trained encoder) is input to the decoder.
[0143] At block 624, if training is complete, the internal weights of the decoder are frozen, and the trained encoder and decoder are output and / or deployed.
[0144] As described above, both the video animation component and the audio animation component can be trained using a two-stage method, thus achieving a more robust animated FACS weight output. In the following, reference Figure 7 is made to the functions and operations associated with the deployed trained model. Figure 7 :Animating an avatar with a trained model
[0145] Figure 7 is a flowchart of an example method 700 for real-time robust facial animation based on video and audio according to some embodiments. In some embodiments, method 700 can be implemented, for example, on a server system (e.g., an online virtual experience platform 102 as Figure 1 shown). In some embodiments, some or all of method 700 can be in a system (e.g., as Figure 1implemented on one or more of the client devices 110 and 116 shown, and / or implemented on the server system and one or more client systems. In the example described, the implementation system includes one or more processors or processing circuits, and one or more storage devices (e.g., a database or other accessible storage device). In some embodiments, different components of one or more servers and / or clients may perform different blocks or other portions of method 700.
[0146] To provide avatar animation, the face can be detected from the input video, and face key points, head pose, tongue state, etc. can be determined and used to animate the face of the corresponding avatar. Additionally, audio can be received from a microphone, and face actions can be determined from phonemes. Before performing face detection or analysis, the user is provided with an indication that such technology will be used for avatar animation. If the user refuses permission, face animation based on video and / or audio is turned off (e.g., a default animation can be used, or the animation can be based on other user-permitted inputs, such as audio and / or text input provided by the user). The video and / or audio provided by the user is specifically used for avatar animation and is not stored. The user can turn off video analysis, audio analysis, and animation generation at any time. Additionally, face detection is performed to detect the position of the face in the video; face recognition is not performed. If the user permits the use of video analysis and audio analysis for avatar animation, method 700 begins at block 702 and block 706.
[0147] At block 702, an input video frame is received from a user device. For example, the user's face can be captured on a camera or image capture device. In some embodiments, the user can select an option in the interface to allow the capture of an image. In these and other examples, the user can also opt out of automatic animation of the video and / or audio. After block 702 is block 704.
[0148] At block 704, video FACS weights are obtained from a trained machine learning model. For example, the video FACS weights can be obtained from the video animation component 204. The video FACS weights can be based on a first level of detail (e.g., using only the B-Net 314), a second level of detail (e.g., using only the H-Net 322), or a third level of detail (using both the B-Net 314 and the H-Net 322).
[0149] At block 706, an input audio frame and a mixing term (e.g., an active audio input signal α, an active speech signal, a non-mute audio signal, and / or others) are received from the user device. The value of the mixing term can increase from zero (no audio or speech) to one. The mixing term was described in detail above with reference to Equation 1. After block 706 is block 708.
[0150] At block 708, audio FACS weights are obtained from a trained machine learning model. For example, the audio FACS weights may be obtained from the audio animation component 210. After blocks 704 and 708 may be block 710.
[0151] At block 710, the audio FACS weights and the video FACS weights are combined with a modular blending component. The combination may include linearly combining the audio FACS weights and the video FACS weights. Additionally, the combination may be based on a blending term. For example, if the user is not actively speaking, the blending term may be close to (or be) zero such that the audio FACS weights are not combined. Additionally, for example, as described in detail above with reference to Equation 1, if the user is actively speaking, the blending term may be used to combine the audio FACS weights with the video FACS weights. After block 710 is block 712.
[0152] At block 712, the combined FACS weights are output as final FACS weights for animating a user's avatar, character rig, 3D model, or any other animatable structure. For example, the final FACS weights may be a blend or fusion after linearly combining the video FACS weights and the audio FACS weights. In this way, even when the user's face is partially occluded while the user is actively speaking, the avatar can be animated. For example, audio can be used to generate FACS weights to provide lip, chin, and / or mouth movements based solely on the audio. Additionally, even if the user is not actively speaking, despite the lack of emitted speech, the video FACS weights can be used to animate the avatar's eyes, mouth, head, etc. based on actual movements. These features and others provide a robust facial animation framework that overcomes many drawbacks and provides technical effects and advantages including: reduced computational cost, increased energy efficiency, improved storage usage, improved bandwidth usage, etc.
[0153] Blocks 702 through 712 may be executed (or repeated) in a different order than described above, and / or one or more blocks may be omitted, modified, combined with other blocks, supplemented with blocks, etc. Method 700 may be executed on a server (e.g., 102) and / or a client device (e.g., 110 or 116). Additionally, according to any desired implementation, portions of method 700 may be combined and executed sequentially or in parallel.
[0154] As described above, the techniques for robust facial animation include implementing a trained face detection model, an audio model, and a trained regression model on a client device. The models may output video FACS weights and audio FACS weights, which are linearly combined in a modular blending component based on a blending term to create final FACS weights for animating a user's avatar, character rig, 3D model, or other animatable structure.
[0155] In the following, referenceFigure 8 Provide a more detailed description of various computing devices that can be used to implement Figures 1 to 4 the different devices and components shown in
[0156] Figure 8 is a block diagram of an example computing device 800 that can be used to implement one or more features described herein. In one example, the device 800 can be used to implement a computer device (e.g., Figure 1 102, 110, and / or 116 of
[0157] ), and perform the appropriate method embodiments described herein. The computing device 800 can be any suitable computer system, server, or other electronic or hardware device. For example, the computing device 800 can be a mainframe computer, a desktop computer, a workstation, a portable computer, or an electronic device (portable device, mobile device, cellular phone, smart phone, tablet computer, television, set-top box, personal digital assistant (PDA), media player, gaming device, wearable device, etc.). In some embodiments, the device 800 includes a processor 802, a memory 804, an input / output (I / O) interface 806, and an audio / video input / output device 814 (e.g., a display screen, a touch screen, display goggles or glasses, an audio speaker, a microphone, etc.).
[0158] 。The memory 804 is typically provided in the device 800 for access by the processor 802 and can be any suitable processor-readable storage medium, e.g., random access memory (RAM), read-only memory (ROM), electrical erasable read-only memory (EEPROM), flash memory, etc. The memory 804 is suitable for storing instructions for execution by the processor and is separate from and / or integrated with the processor 802. The memory 804 can store software for operation by the processor 802 on the server device 800, including an operating system 808, applications 810, and associated data 812. In some embodiments, the applications 810 can include instructions enabling the processor 802 to perform the functions described herein, e.g., some or all of the methods in FIGS. 5 to Figure 7 In some embodiments, as described herein, the applications 810 can also include one or more trained models for generating robust real-time animations based on input video.
[0159] For example, the memory 804 can include software instructions for an application 810 that can provide an animated avatar based on the facial movements of a user captured on a camera (or another device) and audio captured via a microphone (or another device) within an online virtual experience platform (e.g., 102). Any software in the memory 804 can optionally be stored in any other suitable storage location or computer-readable medium. Additionally, the memory 804 (and / or other connected storage devices) can store instructions and data used in the features described herein. The memory 804 and any other type of memory (disk, optical disc, magnetic tape, or other tangible media) can be considered "memory" or "storage device".
[0160] The I / O interface 806 can provide functionality enabling the server device 800 to interface with other systems and devices. For example, network communication devices, storage devices (e.g., memory and / or data storage area 108), and input / output devices can communicate via the interface 806. In some embodiments, the I / O interface can be connected to an interface device including input devices (keyboard, pointing device, touch screen, microphone, camera, scanner, etc.) and / or output devices (display device, speaker device, printer, motor, etc.).
[0161] For ease of illustration, Figure 8A box is shown for each of the processor 802, the memory 804, the I / O interface 806, and the software blocks 808 and 810, and the database 812. These boxes may represent one or more processors or processing circuits, operating systems, memories, I / O interfaces, applications, and / or software modules. In other embodiments, the device 800 may not have all of the components shown and / or may have other types of elements including elements alternative to or in addition to those shown herein. Although the online virtual experience platform 102 is described as performing the operations as described in some embodiments herein, any suitable component or combination of components of the online virtual experience platform 102 or a similar system, or any suitable one or more processors associated with such a system, may perform the described operations.
[0162] User devices may also implement and / or be used in conjunction with the features described herein. Example user devices may be computer devices that include some components similar to those of the device 800, such as the processor 802, the memory 804, and the I / O interface 806. Operating systems, software, and applications suitable for client devices may be provided in the memory and used by the processor. The I / O interface for a client device may be connected to network communication devices as well as input and output devices, such as a microphone for capturing sound, a camera for capturing images or video, an audio speaker device for outputting sound, a display device for outputting images or video, or other output devices. For example, the display device within the audio / video input / output device 814 may be connected to (or included in) the device 800 to display pre-processed and post-processed images as described herein, where such display device may include any suitable display device, such as an LCD, LED, or plasma display screen, CRT, television, monitor, touch screen, 3-D display, projector, or other visual display device. Some embodiments may provide an audio output device, such as a synthesized voice for voice output or reading text.
[0163] The methods, blocks, and / or operations described herein may be performed in an order different from that shown or described, and / or, where appropriate, may be performed simultaneously (in part or in whole) with other blocks or operations. Some blocks or operations may be performed on a portion of the data and then, for example, again on another portion of the data. In different embodiments, not all of the described blocks and operations need to be performed. In some embodiments, the blocks and operations may be performed multiple times, in different orders, and / or at different times in the method, etc.
[0164] In some embodiments, some or all of the methods may be implemented on a system such as one or more client devices. In some embodiments, one or more of the methods described herein may be implemented, for example, on a server system and / or on both a server system and a client system. In some embodiments, different components of one or more servers and / or clients may perform different blocks, operations, or other portions of the methods.
[0165] One or more of the methods described herein (e.g., methods 500, 600, 680, 685, and / or 700) may be implemented by computer program instructions or code executable on a computer. For example, the code may be implemented by one or more digital processors (e.g., a microprocessor or other processing circuitry) and may be stored on a computer program product including a non-transitory computer-readable medium (e.g., a storage medium), such as a magnetic, optical, electromagnetic, or semiconductor storage medium, including semiconductor or solid-state memory, magnetic tape, removable computer floppy disk, random access memory (RAM), read-only memory (ROM), flash memory, rigid disk, optical disk, solid-state storage drive, etc. The program instructions may also be embodied in and provided as an electronic signal, for example, in the form of software as a service (SaaS) delivered from a server (e.g., a distributed system and / or a cloud computing system). Optionally, one or more of the methods may be implemented in hardware (e.g., logic gates, etc.) or in a combination of hardware and software. Example hardware may be a programmable processor (e.g., a field-programmable gate array (FPGA), a complex programmable logic device), a general-purpose processor, a graphics processor, an application specific integrated circuit (ASIC), etc. One or more of the methods may be performed as part of or a component of an application running on a system, or as an application or software running together with other applications and an operating system.
[0166] One or more of the methods described herein can be run in a stand-alone program executable on any type of computing device, a program running on a web browser, or a mobile application ("app") running on a mobile computing device (e.g., a phone, smartphone, tablet, wearable device (watch, armband, jewelry, headgear, goggles, glasses, etc.), laptop, etc.). In one example, a client / server architecture can be used, e.g., a mobile computing device (as a client device) sends user input data to a server device and receives final output data from the server for output (e.g., for display). In another example, all computations can be performed within a mobile application (and / or other applications) on the mobile computing device. In yet another example, computations can be split between the mobile computing device and one or more server devices.
[0167] Although described with respect to specific embodiments herein, these specific embodiments are for illustration only and not limitation. The concepts illustrated in the examples can be applied to other examples and embodiments.
[0168] In cases where certain embodiments discussed herein may obtain or use user data (e.g., images of the user, videos of the user, audio of the user, user population information, user behavior data on the platform, user search history, items purchased and / or viewed, user friendships on the platform, etc.), options are provided to the user to control whether and how such information is collected, stored, or used. That is, the embodiments discussed herein collect, store, and / or use user information upon receiving explicit user authorization and in compliance with applicable regulations.
[0169] The user can control whether a program or feature is allowed to collect user information about that particular user or other users associated with the program or feature. Options are presented (e.g., via a user interface) to each user for whom information is to be collected to allow the user to exert control over the information collection related to that user, providing permission or authorization as to whether information is collected and which portions of the information are to be collected. Additionally, certain data can be modified in one or more ways before storage or use to remove personally identifiable information. As an example, a user's identity can be modified (e.g., by substituting with a pseudonym, numerical value, etc.) such that personally identifiable information cannot be determined. In another example, a user's geographical location can be generalized to a larger area (e.g., city, zip code, state, country, etc.).
[0170] Note that the functional blocks, operations, features, methods, devices, and systems described in this disclosure may be integrated or divided into different combinations of systems, devices, and functional blocks known to those skilled in the art. Any suitable programming language and programming technique may be used to implement the routines of a particular implementation. Different programming techniques may be employed, e.g., procedural or object-oriented. The routines may be executed on a single processing device or multiple processors. Although steps, operations, or calculations are presented in a particular order, the order may be changed in different particular implementations. In some implementations, multiple steps or operations shown as sequential in this specification may be executed simultaneously.
[0171] More details and descriptions of the drawings are provided below in the context, and then the claims directed to one or more aspects of this disclosure are provided.
Claims
1. A computer-implemented method, comprising: Receiving a plurality of input video frames; Receiving a plurality of input audio frames and a mixing term, wherein the plurality of input audio frames includes audio associated with the plurality of input video frames; Obtaining video Facial Action Coding System (FACS) weights from a first trained machine learning model based on the input video frames; Obtaining audio FACS weights from a second trained machine learning model based on the input audio frames; Combining the video FACS weights and the audio FACS weights to obtain final FACS weights, wherein the combination is at least partially based on the mixing term; and Outputting the final FACS weights to drive facial animation of a 3D model.
2. The computer-implemented method according to claim 1, wherein, The first trained machine learning model and the second trained machine learning model are trained in a two-stage semi-supervised training process.
3. The computer-implemented method according to claim 1, wherein, The first trained machine learning model includes at least one encoder and at least three task-specific decoders.
4. The computer-implemented method according to claim 3, wherein: The first task-specific decoder of the at least three task-specific decoders is configured to output a predicted head pose, The second task-specific decoder of the at least three task-specific decoders is configured to output the probability of facial visibility in an input video frame among the received plurality of input video frames, and The third task-specific decoder of the at least three task-specific decoders is configured to output facial key points.
5. The computer-implemented method according to claim 4, wherein, At least one of the at least three task-specific decoders includes a causal convolutional layer applied in the time dimension.
6. The computer-implemented method according to claim 1, wherein, The second trained machine learning model includes at least one encoder and at least two task-specific decoders.
7. The computer-implemented method according to claim 6, wherein, The first task-specific decoder of the at least two task-specific decoders is configured to output the audio FACS weights, and the second task-specific decoder of the at least two task-specific decoders is configured to output the mixing term.
8. A system, comprising: A memory having instructions stored thereon; And A processing device coupled to the memory, wherein the processing device is configured to access the memory and execute the instructions, and wherein the instructions cause the processing device to perform operations, the operations including: Receiving a plurality of input video frames; Receiving a plurality of input audio frames and a mixing term, wherein the plurality of input audio frames includes audio associated with the plurality of input video frames; Obtaining video Facial Action Coding System (FACS) weights from a first trained machine learning model based on the input video frames; Obtaining audio FACS weights from a second trained machine learning model based on the input audio frames; Combining the video FACS weights and the audio FACS weights to obtain final FACS weights, wherein the combination is at least partially based on the mixing term; and Outputting the final FACS weights to drive facial animation of a 3D model.
9. The system according to claim 8, wherein, The first trained machine learning model and the second trained machine learning model are trained in a two-stage semi-supervised training process.
10. The system according to claim 8, wherein, The first trained machine learning model includes at least one encoder and at least three task-specific decoders.
11. The system according to claim 10, wherein: the first task - specific decoder among the at least three task - specific decoders is configured to output a predicted head pose; the second task - specific decoder among the at least three task - specific decoders is configured to output a probability that a face is visible in an input video frame among the received plurality of input video frames; and the third task - specific decoder among the at least three task - specific decoders is configured to output facial key points.
12. The system according to claim 11, wherein, At least one of the at least three task - specific decoders includes a causal convolutional layer applied in the temporal dimension.
13. The system according to claim 8, wherein The second trained machine - learning model includes at least one encoder and at least two task - specific decoders.
14. The system according to claim 13, wherein The first task - specific decoder among the at least two task - specific decoders is configured to output the audio FACS weights, and the second task - specific decoder among the at least two task - specific decoders is configured to output the mixing term.
15. A non - transitory computer - readable medium storing instructions that, when executed by a processing device, cause the processing device to perform operations, the operations including: receiving a plurality of input video frames; receiving a plurality of input audio frames and a mixing term, wherein the plurality of input audio frames includes audio associated with the plurality of input video frames; obtaining video Facial Action Coding System (FACS) weights from a first trained machine - learning model based on the input video frames; obtaining audio FACS weights from a second trained machine - learning model based on the input audio frames; combining the video FACS weights and the audio FACS weights to obtain final FACS weights, wherein the combining is at least partially based on the mixing term; and outputting the final FACS weights to drive facial animation of a 3D model.
16. The non-transitory computer-readable medium according to claim 15, wherein, The first trained machine - learning model and the second trained machine - learning model are trained in a two - stage semi - supervised training process.
17. The non-transitory computer-readable medium according to claim 15, wherein, The first trained machine - learning model includes at least one encoder and at least three task - specific decoders.
18. The non - transitory computer - readable medium according to claim 17, wherein: the first task - specific decoder among the at least three task - specific decoders is configured to output a predicted head pose; the second task - specific decoder among the at least three task - specific decoders is configured to output a probability that a face is visible in an input video frame among the received input video frames; and the third task - specific decoder among the at least three task - specific decoders is configured to output facial key points.
19. The non-transitory computer-readable medium according to claim 18, wherein, At least one of the at least three task - specific decoders includes a causal convolutional layer applied in the temporal dimension.
20. The non-transitory computer-readable medium according to claim 15, wherein, The second trained machine - learning model includes at least one encoder and at least two task - specific decoders, and wherein the first task - specific decoder among the at least two task - specific decoders is configured to output the audio FACS weights, and the second task - specific decoder among the at least two task - specific decoders is configured to output the mixing term.