Automatic generation of personalized avatars from 2D images
The system addresses the inefficiencies of manual avatar creation by using a backend server and client-side compact models to generate high-quality avatars from 2D images, improving processing speed and user engagement.
Patent Information
- Application Number
- JP2025574530
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-05
- Filing Date
- 2024-08-16
- Publication Date
- 2026-08-25
AI Technical Summary
Conventional methods for creating personalized avatars from 2D images require significant manual effort and expertise, and mobile devices often lack the computing resources to efficiently generate high-quality avatars due to long rendering times and reduced graphics quality.
A system that leverages a backend server for complex processing and deploys user-customized compact models on client devices to create personalized avatars, using machine learning models to generate 3D meshes from input images and user prompts, with options for image-based or text-based generation.
Reduces computing resource usage on client devices, improves processing efficiency, and enhances user engagement through intelligent prompts for style preferences, enabling rapid generation of high-quality avatars.
Smart Images

Figure 2026528681000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Application No. 63 / 533,106, filed on August 16, 2023, titled "FULLY GENERATIVE AVATARS", U.S. Provisional Application No. 63 / 598,936, filed on November 14, 2023, titled "AUTOMATIC PERSONALIZED AVATAR GENERATION FROM 2D IMAGES", and U.S. Provisional Application No. 63 / 656,571, filed on June 5, 2024, titled "AUTOMATIC PERSONALIZED AVATAR GENERATION FROM 2D IMAGES", the entire contents of each of which are incorporated herein by reference.
[0002] Embodiments generally relate to computer - based virtual experiences, and more specifically, to methods, systems, and computer - readable media for automatically generating personalized avatars from two - dimensional (2D) images.
Background Art
[0003] Some online platforms (e.g., gaming platforms, media exchange platforms, etc.) allow users to connect with each other, interact with each other (e.g., within games), create games, and share information with each other over the internet. Users of online platforms may participate in multiplayer gaming environments or virtual environments (e.g., 3D environments), design custom gaming environments, design characters and avatars, decorate avatars, exchange virtual items / objects with other users, and communicate with other users using voice or text messaging. Users interacting with each other may use interactive interfaces that include the presentation of their avatars. Customizing avatars to reproduce a user's facial features may traditionally involve allowing users to create complex three-dimensional (3D) meshes using computer-aided design tools or other tools. For example, creating an avatar or character involves a series of intricate stages, each requiring significant manual effort, expertise, and proficiency in specific software tools. These stages encompass tasks such as shaping the mesh, applying textures, setting up rigging and skinning, defining the structural framework, and segmenting components. Such traditional solutions have drawbacks, and several implementations have been conceived with these points in mind.
[0004] The background information provided herein is intended to present the context of this disclosure. The research of the inventors named herein is not considered prior art to this disclosure, either explicitly or implicitly, to the extent that the research is described in this background art section, nor any other description that would otherwise be considered prior art at the time of filing. [Overview of the project] [Means for solving the problem]
[0005] The implementations described herein relate to methods, systems, apparatus, and computer-readable media for generating personalized avatars for a user based on one or more 2D images of a face. Automatic generation can be facilitated by a generation component deployed on a virtual experience server and configured to generate a data structure representing a 3D mesh of the personalized avatar in response to the input of one or more 2D images. This data structure can be converted into a polygon mesh and automatically fitted and rigged onto the head portion of the avatar data model.
[0006] In yet another aspect, parts, features, and details of systems, methods, apparatus, and non-temporary computer-readable media may be combined to form additional aspects, including additional components or features, and / or other modifications, which omit and / or modify some or more individual components or features, all of which are within the scope of the disclosure. [Brief explanation of the drawing]
[0007] [Figure 1] This is a diagram illustrating an exemplary network environment with several implementation configurations. [Figure 2A] This block diagram illustrates an exemplary generation pipeline for the generation component shown in Figure 1, based on several implementation configurations. [Figure 2B] This block diagram illustrates an example of individual components of the pipeline shown in Figure 2A, based on several implementation configurations. [Figure 2C] This block diagram illustrates another example of the individual components of the pipeline shown in Figure 2A, based on several implementation configurations. [Figure 2D] This block diagram illustrates another example of the individual components of the pipeline shown in Figure 2A, based on several implementation configurations. [Figure 3A]This block diagram illustrates an example of an exemplary training architecture for the pipeline shown in Figure 2A, based on several implementation configurations. [Figure 3B] This block diagram illustrates an example of a training architecture for the exemplary components shown in Figure 2B, based on several implementation configurations. [Figure 3C] This block diagram illustrates an example of an exemplary training architecture for the component shown in Figure 2C, based on several implementation configurations. [Figure 3D] This block diagram illustrates an example of a training architecture for the components shown in Figure 2D, based on several implementation configurations. [Figure 3E] This block diagram illustrates an example of a training architecture for a conditional model, showing several implementation variations. [Figure 4A] This is a block diagram of a mapping network in several implementation forms. [Figure 4B] This is a block diagram of an exemplary content mapping decoder for the mapping network shown in Figure 4A, representing several implementation configurations. [Figure 4C] This is a block diagram of another exemplary content mapping decoder for the mapping network shown in Figure 4A, with several implementation configurations. [Figure 5] This is a block diagram of an exemplary opacity decoder in several implementation forms. [Figure 6] This is a block diagram of an exemplary conditional model, showing several implementation variations. [Figure 7] This is a flowchart illustrating exemplary methods for training the exemplary generation pipeline shown in Figure 2A, using several implementation configurations. [Figure 8] This is a flowchart illustrating exemplary methods for training an unconditional model with a GAN generator, using several implementation configurations. [Figure 9]This is a flowchart illustrating exemplary methods for training an unconditional model with a Dual-GAN generator, using several implementation configurations. [Figure 10] This is a flowchart illustrating several implementation methods for automatically generating personalized avatars from 2D images. [Figure 11] This is a flowchart illustrating another exemplary method for automatically generating personalized avatars from 2D images, using several implementation configurations. [Figure 12] This block diagram illustrates an exemplary computing device that may be used to implement one or more of the features described herein, in several implementation forms. [Modes for carrying out the invention]
[0008] The following detailed description refers to the accompanying drawings, which form part of the detailed description. Similar symbols in the drawings typically indicate similar components unless the context indicates otherwise. The exemplary implementations described in the detailed description, drawings, and claims are not intended to limit. Other implementations may be used, and other modifications may be made without departing from the spirit or scope of the subject matter presented herein. All aspects of the disclosure, as generally described herein and illustrated in the drawings, may be arranged, configured, replaced, combined, separated, and designed in a variety of different configurations, all as intended herein.
[0009] References in this specification to "several implementations," "one implementation," or "exemplary implementation" indicate that the implementations described may possess certain features, structures, or characteristics, but not all implementations may necessarily include those features, structures, or characteristics. Furthermore, such phrases do not necessarily refer to the same implementation. Moreover, when a particular feature, structure, or characteristic is described in relation to one implementation, that feature, structure, or characteristic may be affected in relation to other implementations, whether explicitly described or not.
[0010] In some embodiments, systems and methods are provided for the automatic generation of personalized avatars from 2D images. Features may include automatically creating a 3D mesh of the avatar's face based on one or more input images and text prompt responses received from a client device, and rigging the 3D mesh into an avatar data model for display within a virtual experience. The automatic generation may include the generation of one or more feature vectors by one or more machine learning models, and the generation is based on the input image of the face as well as the user's responses / prompts to one or more prompts presented to the user. The one or more prompts may be prompts for describing the features, style, and other attributes that the user desires on which the personalized avatar will be based.
[0011] In some implementations, one or more input images depict a human face. In these implementations, the user is presented with guidance on how the images are stored (e.g., temporarily in memory, for one day, etc.) and how they are used (for a specific purpose such as generating a 3D mesh of an avatar's face). The user may proceed to create a 3D mesh by providing an input image or may choose not to provide an input image. If the user chooses not to provide an input image, 3D mesh generation is performed without using an input image (e.g., based on stock images, based on text prompts, or using other techniques that do not use input images). The input image is used only for 3D mesh generation and is discarded (deleted from memory and storage) after generation is complete. The input image-based 3D mesh generation feature is provided only in jurisdictions with legal authority where various operations related to the use of input images depicting faces and the use of faces in avatar generation are permitted in accordance with local regulations. In jurisdictions where such use is not permitted (e.g., where the storage of face data is not permitted), face-based 3D mesh generation is not implemented. Further, face-based 3D mesh generation is implemented for a specific set of users, such as adults (e.g., 18 years or older, 21 years or older, etc.) who provide legal consent regarding the use of input images and face data. The user is provided with an option to delete the 3D mesh generated from the input image, and in response to the user selecting the option, the 3D mesh (and any associated data, including the input image if it is stored) is deleted from memory and storage. The face information from the input image is used specifically for 3D mesh generation and is used in accordance with the contractual conditions when such information is obtained.
[0012] For example, the user may want to generate a caricature or other stylized drawing of their face. In these examples, the response / prompt to one or more prompts may include "caricature", "comic", and / or other responses / prompts.
[0013] For example, a user may want to change their personal style to match a popular or other style. In these examples, the response / prompt to one or more prompts may include "popular culture", "rock star", and / or other responses / prompts.
[0014] For example, a user may want to change their appearance to avoid or emphasize certain personal characteristics. In these examples, the response / prompt to one or more prompts may include "curly hair", "bald head", "big eyes", and / or other responses / prompts.
[0015] For example, a user may want to change their appearance so that it matches or mimics the style of a movie or TV series. In these examples, the response / prompt to one or more prompts may include "Poppy", "Addams Family", and / or other responses / prompts.
[0016] For example, a user may want a plurality of different features to be considered so that the automatically generated avatar depicts a combination of the plurality of different features. In these examples, the response / prompt to one or more prompts may include "increase hair", "reduce face size", "make skin like a cartoon", and / or other responses / prompts.
[0017] The features described herein provide for the automatic generation of a vector representing a face detected in a 2D image, the automatic generation of a response / prompt to one or more prompts and / or one or more vectors (e.g., conditional density sampling vectors) based on the detected face, the generation of a 3D mesh based on the generated vector, and the automatic rigging of the 3D mesh onto an avatar data model and / or a skeleton frame.
[0018] The generative component is trained to accurately generate feature vectors and 3D meshes, including conditional density sampling. Training may include independent training of a 2D generative component configured to output a computer-generated (CG) representation of a face, and training of a 2D-to-3D generative component configured to generate a 3D polygon mesh based on the CG representation. In some implementations, the 2D generative component may be replaced by a sequence of selectable 2D CG images for a user-selected face to begin with. In some implementations, the 2D-to-3D generative component may include a generative network formed using a conditional model and an unconditional model. In some implementations, the unconditional model may include a Generative Adversarial Network (GAN) generator, which includes at least one GAN model. In some implementations, the unconditional model may include a dual-GAN generator, which includes at least two GAN models.
[0019] In some implementations, training may include independently training a neural network configured to output a computer-generated representation of the user's face, training a style encoder to generate a vector from conditional density sampling of desired style characteristics / answers / prompts for one or more prompts, and training a GAN generator component. The trained model can then be used to automatically generate an avatar for that user using an output 3D mesh created from the output provided by the GAN generator.
[0020] In some implementations, training may include independently training a neural network configured to output a computer-generated representation of the user's face, training a style encoder to generate a vector from conditional density sampling of desired style characteristics / answers / prompts for one or more prompts, and training a dual-GAN generator component. The trained model can then be used to automatically generate an avatar for that user using an output 3D mesh created from the output provided by the dual-GAN generator.
[0021] The trained models may be deployed to server or client devices for use by users who request to have avatars automatically created based on their personalized preferences. Client devices may be configured to communicate with online platforms, such as virtual experience platforms, thereby allowing their associated avatars to be richly animated for presentation in communication interfaces (e.g., video chat), within virtual experiences (e.g., customized faces on representative virtual bodies), in animated videos transmitted to other users (e.g., by sending recordings or renderings of avatars through chat or other functions), and within other parts of the online platform.
[0022] Online virtual experience platforms (also known as "user-generated content platforms" or "user-generated content systems") provide various ways for users to interact with each other. For example, users of an online virtual experience platform can create experiences, games, or other content or resources (such as characters, graphics, or items for gameplay in a virtual world) within the platform.
[0023] Users of an online virtual experience platform can collaborate towards common goals in games or game creation, share various virtual items, and send electronic messages to each other. Users of an online virtual experience platform can interact with the environment and play games that include, for example, characters (avatars) or other game objects and mechanisms. An online virtual experience platform can enable users of the platform to communicate with each other. For example, users of an online virtual experience platform can communicate with each other using voice messages (e.g., via voice chat), text messaging, video messaging, or a combination of the above. Some online virtual experience platforms can provide a virtual 3D environment in which users can express themselves using their own avatars or virtual representations.
[0024] To enhance the entertainment value of online virtual experience platforms, the platform may provide a generation component that facilitates the automatic generation of avatars based on user preferences. This generation component may allow users to request features for generation, respond to prompts, or select options, for example, by providing plain text descriptions of desired features for the automatically generated avatar.
[0025] For example, a user can grant an application on their device associated with an online virtual experience platform access to their camera. Images captured by the camera can be interpreted to extract features or other information that facilitate the generation of a basic 2D representation of the face based on extracted gestures. Similarly, the user can augment the generation through responses to / inputs to one or more prompts. The stylized drawing of the face can then be generated as a 3D mesh usable within the virtual experience platform (or other platform) to create a personalized avatar.
[0026] However, conventional solutions are limited in their ability to automatically generate customized or personalized avatars because many mobile client devices lack sufficient computing resources. For example, many users may use portable computing devices (e.g., ultralight portables, tablets, mobile phones, etc.) that lack the computing power to quickly interpret facial features and prompts and accurately create customized computer-generated images. In these situations, many auto-generated components suffer from drawbacks including long rendering times (e.g., ranging from 30 minutes to several hours), reduced graphics quality (e.g., when attempting to shorten rendering times), and other issues.
[0027] Thus, while some users may acquire mobile or computing devices with sufficient processing power to handle robust generative models through conventional solutions, many users lack this experience due to using some other suitable device that lacks sufficient computing resources for complex computer vision processing.
[0028] In these scenarios, the various implementations described herein leverage a backend server that can provide processing power for more complex generative tasks, while one or more user-customized compact models are deployed on a client device to create personalized encodings of facial features from input 2D images. Thus, exemplary embodiments offer technical benefits, including reduced use of computing resources on the client device, improved data processing flow from the client device to the server deploying trained models and generative components, improved user engagement through intelligent prompts for style preferences, and other technical advantages and effects revealed throughout this disclosure.
[0029] Figure 1: System Architecture Figure 1 illustrates exemplary network environments 100 in several implementations of the present disclosure. The network environment 100 (also referred to herein as the “System”) includes an online virtual experience platform 102, a first client device 110, and a second client device 116 (collectively referred to herein as “Client Devices 110 / 116”), all connected via a network 122. The online virtual experience platform 102 may, among other things, include a virtual experience (VE) engine 104, one or more virtual experiences 105, a generation component 107, and a data store 108. Client device 110 may include a virtual experience application 112. Client device 116 may include a virtual experience application 118. Users 114 and 120 can interact with the online virtual experience platform 102 using client devices 110 and 116, respectively.
[0030] The network environment 100 is provided for illustrative purposes. In some implementations, the network environment 100 may include elements configured in the same, fewer, more, or different ways as shown in Figure 1.
[0031] In some implementations, network 122 may include a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or a wide area network (WAN)), a wired network (e.g., an Ethernet network), a wireless network (e.g., an 802.11 network, a Wi-Fi® network, or a wireless LAN (WLAN)), a cellular network (e.g., a Long-Term Evolution (LTE) network), a router, a hub, a switch, a server computer, or any combination thereof.
[0032] In some implementations, the datastore 108 may be non-temporary computer-readable memory (e.g., random access memory), a cache, a drive (e.g., a hard drive), a flash drive, a database system, or another type of component or device capable of storing data. The datastore 108 may also include multiple storage components (e.g., multiple drives or multiple databases) that may further span multiple computing devices (e.g., multiple server computers).
[0033] In some implementations, the online virtual experience platform 102 may include servers having one or more computing devices (e.g., cloud computing systems, rack-mount servers, server computers, clusters of physical servers, virtual servers, etc.). In some implementations, the servers may be included in the online virtual experience platform 102, be a separate system, or be part of another system or platform.
[0034] In some implementations, the online virtual experience platform 102 may include one or more computing devices (such as rack-mount servers, router computers, server computers, personal computers, mainframe computers, laptop computers, tablet computers, and desktop computers), data stores (e.g., hard disks, memory, and databases), networks, software components, and / or hardware components that can be used to run on the online virtual experience platform 102 and to provide users with access to the online virtual experience platform 102. The online virtual experience platform 102 may also include a website (e.g., one or more web pages) or application backend software that can be used to provide users with access to content provided by the online virtual experience platform 102. For example, users may access the online virtual experience platform 102 using virtual experience applications 112 / 118 on client devices 110 / 116, respectively.
[0035] In some implementations, the online virtual experience platform 102 may include a type of social network or user-generated content system that provides user-to-user connectivity, enabling users (e.g., end users or consumers) to communicate with other users through the online virtual experience platform 102, and communication may include voice chat (e.g., synchronous and / or asynchronous voice communication), video chat (e.g., synchronous and / or asynchronous video communication), or text chat (e.g., synchronous and / or asynchronous text-based communication). In some implementations of this disclosure, “user” may be represented as a single individual. However, other implementations of this disclosure may include the fact that “user” (e.g., creating a user) is an entity controlled by a collection of users or an automated source. For example, a collection of individual users federated as a community or group in a user-generated content system may be considered a “user.”
[0036] In some implementations, the online virtual experience platform 102 may be a virtual gaming platform. For example, the gaming platform may provide single-player or multiplayer games to a community of users who can access or interact with games (e.g., user-generated games or other games) using client devices 110 / 116 via the network 122. In some implementations, games (also referred to herein as “video games,” “online games,” or “virtual games”) may be, for example, two-dimensional (2D) games, three-dimensional (3D) games (e.g., 3D user-generated games), virtual reality (VR) games, or augmented reality (AR) games. In some implementations, users may search for games and game items and participate in gameplay with other users in one or more games. In some implementations, games may be played in real time with other users of the game.
[0037] In some implementations, other collaboration platforms may be used in place of or in addition to the online virtual experience platform 102, along with the generation functions described herein. For example, social networking platforms, video chat platforms, messaging platforms, user content creation platforms, virtual conferencing platforms, etc., can be used with the generation functions described herein to facilitate the rapid, robust, and accurate generation of personalized virtual avatars.
[0038] In some implementations, “gameplay” may refer to the interactive interaction between one or more players using client devices (e.g., 110 and / or 116) within a game or experience (e.g., VE105), or the presentation of such interactive interaction on the display or other output device of client device 110 or 116.
[0039] One or more virtual experiences 105 are provided by an online virtual experience platform. In some implementations, a virtual experience 105 may include electronic files that can be executed or loaded using software, firmware, or hardware configured to present virtual content (e.g., digital media items) to entities. In some implementations, a virtual experience application 112 / 118 may be executed, and the virtual experience 105 may be rendered in relation to a virtual experience engine 104. In some implementations, a virtual experience 105 may have a common set of rules or common goals, and the environment of the virtual experience 105 may share a common set of rules or common goals. In some implementations, different virtual experiences may have different rules or goals from one another. Similarly, or alternatively, some virtual experiences may have no goals at all, intended to be interactive interactions between users in any social manner.
[0040] In some implementations, a virtual experience may have one or more environments (also referred to herein as “gaming environments” or “virtual environments”) to which multiple environments can be linked. An example of an environment may be a three-dimensional (3D) environment. One or more environments of virtual experience 105 may collectively be referred herein as a “world,” “virtual world,” “virtual universe,” or “metaverse.” An example of a world may be the 3D world of virtual experience 105. For example, a user may build a virtual environment linked to another virtual environment created by another user. A character in a virtual experience may also cross virtual boundaries to enter adjacent virtual environments.
[0041] Note that a 3D environment or 3D world uses graphics that employ a three-dimensional representation of geometric data to represent virtual content (or presents content in a way that makes it appear as 3D content, regardless of whether a three-dimensional representation of geometric data is used). A 2D environment or 2D world uses graphics that employ a two-dimensional representation of geometric data to represent virtual content.
[0042] In some implementations, the online virtual experience platform 102 can host one or more virtual experiences 105 and enable users to interact with the virtual experiences 105 using virtual experience applications 112 / 118 on client devices 110 / 116 (for example, to search for games, VE-related content, or other content). Users of the online virtual experience platform 102 (for example, 114 and / or 120) can play, create, interact with, or build virtual experiences 105, search for virtual experiences 105, communicate with other users, and create, build, and / or search for objects of virtual experiences 105 (for example, also referred to herein as “items,” “game objects,” or “virtual game items”). For example, when creating a user-generated virtual item, the user may build, among other things, a character, decorations for a character, one or more virtual environments for an interactive experience, or structures used in virtual experience 105.
[0043] In some implementations, users may buy, sell, or trade virtual objects, such as in-platform currency (e.g., virtual currency), with other users of the online virtual experience platform 102. In some implementations, the online virtual experience platform 102 may transmit virtual content to virtual experience applications (e.g., 112, 118). In some implementations, virtual content (also referred to herein as “content”) may refer to any data or software instructions associated with the online virtual experience platform 102 or a virtual experience application (e.g., virtual objects, experiences, user information, videos, images, commands, media items, etc.).
[0044] In some implementations, a virtual object (for example, also referred to herein as “item,” “object,” or “virtual game item”) may refer to an object used, created, shared, or otherwise depicted in a virtual experience application 105 of the online virtual experience platform 102, or in a virtual experience application 112 or 118 of a client device 110 / 116. For example, a virtual object may include parts, models, characters, tools, weapons, clothing, buildings, vehicles, currency, flora, fauna, and the aforementioned components (for example, windows of a building).
[0045] It should be noted that the online virtual experience platform 102 hosting the virtual experience 105 is provided for illustrative purposes only, and not as an extension. In some implementations, the online virtual experience platform 102 may host one or more media items that may contain communication messages from one user to one or more other users. Media items may include, but are not limited to, digital videos, digital movies, digital photographs, digital music, audio content, melodies, website content, social media updates, ebooks, e-magazines, digital newspapers, digital audiobooks, e-journals, web blogs, Real Simple Syndication (RSS) feeds, e-comic books, and software applications. In some implementations, media items may be electronic files that can be executed or loaded using software, firmware, or hardware configured to present digital media items to entities.
[0046] In some implementations, a virtual experience 105 may be associated with a specific user or a specific group of users (e.g., a private experience) or may be made widely available to users of the online virtual experience platform 102 (e.g., a public experience). In some implementations, if the online virtual experience platform 102 associates one or more virtual experiences 105 with a specific user or group of users, the online virtual experience platform 102 may use user account information (e.g., a user account identifier such as a username and password) to associate the specific user with the virtual experience 105. Similarly, in some implementations, the online virtual experience platform 102 may use developer account information (e.g., a developer account identifier such as a username and password) to associate a specific developer or group of developers with the virtual experience 105.
[0047] In some implementations, the online virtual experience platform 102 or client devices 110 / 116 may include a virtual experience engine 104 or virtual experience applications 112 / 118. The virtual experience engine 104 may include virtual experience applications similar to the virtual experience applications 112 / 118. In some implementations, the virtual experience engine 104 may be used to develop or run a virtual experience 105. For example, the virtual experience engine 104 may include, among others, a rendering engine ("renderer") for 2D, 3D, VR, or AR graphics, a physics engine, a collision detection engine (and collision response), a sound engine, scripting capabilities, an animation engine, an artificial intelligence engine, networking capabilities, streaming capabilities, memory management capabilities, threading capabilities, scene graph capabilities, or video support for cinematics. Components of the virtual experience engine 104 may generate commands (e.g., rendering commands, collision commands, physics commands, etc.) that help compute and render the virtual experience. In some implementations, the virtual experience applications 112 / 118 of the client devices 110 / 116 may operate independently, in cooperation with the virtual experience engine 104 of the online virtual experience platform 102, or in a combination of both.
[0048] In some implementations, both the online virtual experience platform 102 and client devices 110 / 116 run virtual experience engines (104, 112, and 118, respectively). The online virtual experience platform 102, using virtual experience engine 104, may run some or all of the virtual experience engine functions (e.g., generating physical commands, rendering commands, etc.) or offload some or all of the virtual experience engine functions to the virtual experience engine 104 of client device 110. In some implementations, each virtual experience 105 may have a different ratio between the virtual experience engine functions running on the online virtual experience platform 102 and the virtual experience engine functions running on client devices 110 and 116.
[0049] For example, the virtual experience engine 104 of the online virtual experience platform 102 may be used to generate physics commands when there is a collision between at least two game objects, while additional virtual experience engine functions (e.g., generating rendering commands) may be offloaded to the client device 110. In some implementations, the ratio of virtual experience engine functions running on the online virtual experience platform 102 and the client device 110 may be changed (e.g., dynamically) based on interactivity conditions. For example, if the number of users participating in the virtual experience 105 exceeds a threshold number, the online virtual experience platform 102 may execute one or more virtual experience engine functions previously executed by the client device 110 or 116.
[0050] For example, a user may be interactively manipulating the virtual experience 105 on client devices 110 and 116, and may send control commands (e.g., user input such as right, left, up, down, user selection, or character position and velocity information) to the online virtual experience platform 102. After receiving control commands from client devices 110 and 116, the online virtual experience platform 102 may send interaction commands (e.g., position and velocity information of characters participating in the virtual experience, or commands such as rendering commands or collision commands) to the client devices 110 and 116 based on the control commands. For example, the online virtual experience platform 102 may perform one or more logical operations on the control commands (e.g., using the virtual experience engine 104) to generate interaction commands for client devices 110 and 116. In other cases, the online virtual experience platform 102 may pass one or more control commands from one client device 110 to another client device (e.g., 116) participating in the virtual experience 105. Client devices 110 and 116 may use instructions to render the experience to be presented on the displays of client devices 110 and 116.
[0051] In some implementations, control commands may refer to commands that indicate an in-experience action for the user's character or avatar. For example, control commands may include user inputs to control in-experience actions, such as right, left, up, down, user selection, gyroscope position and orientation data, and force sensor data. Control commands may also include character position and velocity information. In some implementations, control commands are sent directly to the online virtual experience platform 102. In other implementations, control commands may be sent from client device 110 to other client devices (e.g., 116), which use the local virtual experience engine 104 to generate play commands. Control commands may include commands to play voice communication messages or other sounds from the user through an audio device (e.g., speaker, headphones), commands to move the character or avatar, and other commands.
[0052] In some implementations, interaction or play commands may refer to commands that enable the client device 110 (or 116) to render the movement of elements of a virtual experience, such as a multiplayer game. Commands may include one or more of the following: user input (e.g., control commands), character position and velocity information, or commands (e.g., physics commands, rendering commands, collision commands, etc.). Other commands may include face animation commands extracted through analysis of face input video to direct the animation of a representative virtual face of a virtual avatar in real time. Thus, interaction commands may include user input for direct control of some bodily movement of the character, but interaction commands may also include gestures extracted from the user's video.
[0053] In some implementations, a character (or virtual object in general) is constructed from components that can be selected by the user, one or more of which are automatically combined to assist the user in editing. One or more characters (also referred to herein as “avatars” or “models”) are associated with the user, and the user can control the characters to facilitate the user’s interaction with the virtual experience 105. In some implementations, a character may include components such as body parts (e.g., hair, arms, legs, etc.) and accessories (e.g., T-shirts, glasses, decorative images, tools, etc.). In some implementations, customizable character body parts may include, among others, head type, body part type (arms, legs, torso, and hands), face type, hair type, and skin type. In some implementations, customizable accessories may include clothing (e.g., shirts, trousers, hats, shoes, glasses, etc.), weapons, or other tools. Some or all of these components may be automatically generated using the features described herein.
[0054] In some implementations, the user may also control the character's scale (e.g., height, width, or depth) or the scale of the character's components. In some implementations, the user may also control the character's proportions (e.g., mass, anatomical, etc.). In some implementations, the character may not include a character object (e.g., body parts), but it should be noted that the user may still control the character (without a character object) to facilitate user interaction with the game (e.g., a puzzle game where there is no rendered character game object, but the user still controls the character to control in-game actions).
[0055] In some implementations, components such as body parts may be primitive geometric shapes such as blocks, cylinders, or spheres, or other primitive shapes such as wedges, tori, tubes, or channels. In some implementations, the creator module may expose the user's character for viewing or use by other users of the online virtual experience platform 102. In some implementations, creating, modifying, or customizing characters, other virtual objects, virtual experiences 105, or virtual environments may be done by the user using a user interface (e.g., a developer interface), with or without scripting (or with or without an application programming interface (API)). Note that, for illustrative purposes rather than limitation, characters are described as having human form. Furthermore, note that characters may have any form, such as vehicles, animals, inanimate objects, or other creative forms.
[0056] In some implementations, the online virtual experience platform 102 may store user-created characters in the data store 108. In some implementations, the online virtual experience platform 102 maintains a character catalog and an experience catalog that can be presented to the user via the virtual experience engine 104, the virtual experience 105, and / or client devices 110 / 116. In some implementations, the experience catalog contains images of different experiences stored in the online virtual experience platform 102. In addition, the user may select a character (for example, a character created by the user or another user) from the character catalog and participate in the selected experience. The character catalog contains images of characters stored in the online virtual experience platform 102. In some implementations, one or more characters in the character catalog may be created or customized by the user. In some implementations, the selected character may have character settings that define one or more of the character's components.
[0057] In some implementations, a user's character may include component configurations, and the configuration and appearance of these components, and more generally the appearance of the character, may be defined by character settings and / or personalized settings, as described herein. In some implementations, at least part of a user's character settings may be selected by the user. In other implementations, a user may select a character that has default character settings or other user-selected character settings. For example, a user may select a default character from a character catalog with predefined character settings, and the user may further customize the default character by changing some of the character settings (e.g., adding a customized logo to a shirt). Character settings may be associated with a particular character by the online virtual experience platform 102.
[0058] In some implementations, the user may input character settings and / or style and / or personal preferences as one or more responses / prompts to one or more prompts. These responses / prompts may be used by a trained neural network and a trained style encoder to create feature vectors for automatically generating a 3D mesh representing those character settings and / or style and / or personal preferences by a GAN generator and / or dual-GAN generator, as described herein.
[0059] In some implementations, the user may select from a catalog of CG images as a starting point or reference point when generating a customized avatar. The CG images are displayed to the user in VE application 112 or 118, and the user may select at least one CG image. The user may then be presented with one or more prompts for customizing the resulting avatar from the selected CG image. For example, one or more prompts such as "What hairstyle do you prefer?" or "Is there a celebrity you would like your avatar to resemble?" may be displayed. The user may type natural language text as a response to any prompt, which is used by system 100 to generate appropriate feature vectors and / or select additional CG images from the image repository to generate feature vectors representing those characteristics that match the user's responses / prompts. These feature vectors and the selected CG images are used by a 2D-to-3D generation component, which may automatically generate a mesh that fits onto the user's avatar, displaying the characteristics based on the selected CG images and the user's responses / prompts.
[0060] In some implementations, the user may also be allowed access to the camera and take a series of "selfies" (e.g., one or more images including their face). A trained neural network may then create a computer graphics (CG) image based on the selfies for input into a 2D-to-3D generation component. Note that in this example, user responses to prompts may also be used. The generated CG image of the face, as well as any user responses to prompts, are input into the 2D-to-3D generation component, which can automatically generate a mesh to fit onto the user's avatar, displaying the CG image of the face and characteristics based on the user's responses / prompts.
[0061] In some implementations, client devices 110 or 116 may each include computing devices such as personal computers (PCs), mobile devices (laptops, mobile phones, smartphones, tablet computers, or netbooks), network-connected televisions, and game consoles. In some implementations, client devices 110 or 116 may also be referred to as “user devices.” In some implementations, one or more client devices 110 or 116 may connect to the online virtual experience platform 102 at any given moment. Note that the number of client devices 110 or 116 is provided as an example, not an exemption. In some implementations, any number of client devices 110 or 116 may be used.
[0062] In some implementations, each client device 110 or 116 may contain an instance of a virtual experience application 112 or 118. In one implementation, the virtual experience application 112 or 118 may enable the user to use and interact with the online virtual experience platform 102, such as searching for a specific experience or other content, controlling a virtual character in a virtual game hosted by the online virtual experience platform 102, or browsing or uploading content such as virtual experiences 105, images, video items, web pages, or documents. In one example, the virtual experience application may be a web application (e.g., an application that works in conjunction with a web browser) that can access, retrieve, present, or navigate content provided by a web server (e.g., a virtual character in a virtual environment). In another example, the virtual experience application may be a native application (e.g., a mobile application, app, or program) installed and running locally on the client device 110 or 116, enabling the user to interact with the online virtual experience platform 102. A virtual experience application may render, display, or present content (e.g., a web page, user interface, media viewer) to a user. In one implementation, a virtual experience application may also include an embedded media player (e.g., a Flash® player) embedded within a web page.
[0063] In aspects of this disclosure, the virtual experience application 112 / 118 may be an online virtual experience platform application that allows users to build, create, edit, and upload content to the online virtual experience platform 102, as well as to interact with the online virtual experience platform 102 (for example, to play and interact with a virtual experience 105 hosted by the online virtual experience platform 102). As such, the virtual experience application 112 / 118 may be provided to a client device 110 or 116 by the online virtual experience platform 102. In another example, the virtual experience application 112 / 118 may be an application downloaded from a server.
[0064] In some implementations, a user may log in to the online virtual experience platform 102 via a virtual experience application. The user may access a user account by providing user account information (e.g., username and password), and the user account may be associated with one or more characters available to participate in one or more virtual experiences 105 of the online virtual experience platform 102.
[0065] In general, functions described as being performed by the online virtual experience platform 102 may also be performed by client devices 110 or 116, or by a server, in other implementations where appropriate. In addition, functions attributed to a particular component may be performed by different components or multiple components working together. The online virtual experience platform 102 may be accessed as a service provided to other systems or devices through an appropriate application programming interface (API), and is therefore not limited to use on a website.
[0066] In some implementations, the online virtual experience platform 102 may include a generation component 107. In some implementations, the generation component 107 may be a system, application, pipeline, and / or module that leverages several trained models to automatically generate a user's avatar or character face based on selected images, personal images, personalized preferences, character settings, and / or responses / prompts. In some implementations, the generation component 107 may include a pipeline having a 2D to 3D generation AI component. In some implementations, the generation component 107 may include trained conditional models and trained unconditional models. In some implementations, the generation component 107 may include a trained style encoder and a GAN generator. In some implementations, the generation component 107 may include a trained style encoder and / or a content encoder, as well as a dual-GAN generator.
[0067] The generation by the generation component 107 may, in some implementations, be based on user selection of CG images provided by the system 100. For example, multiple CG images may be displayed for user selection. The user may make different style selections, provide responses to style prompts, and / or provide other personal preferences.
[0068] The generation by the generation component 107 may also be based on the user's actual face and, as such, may include several features extracted from the face (e.g., via an input image), but may also include features and styles defined in the user's preferences and / or responses / prompts. In situations where these and several other implementations described herein may acquire or use user data (e.g., user images, user demographics, user responses to prompts, user self-descriptions, etc.), the user is provided with options to control whether, and how, such information is collected, stored, or used. That is, the implementations described herein, after obtaining explicit user authorization, will collect, store, and / or use user information in accordance with applicable regulations.
[0069] Furthermore, users are provided with control over whether a program or feature collects user information about that particular user or other users associated with that program or feature. Each user from whom information should be collected is presented with options (e.g., via the user interface) that allow the user to control the information collection related to that user and provide permission or authorization for whether information is collected and which parts of the information are collected. In addition, certain data may be modified in one or more ways before being stored or used, so that personally identifiable information is removed.
[0070] Although the trained model described is exemplified to run directly on the online virtual experience platform 102, it should be understood that in some implementations, it may run on each client device 110, 116, for example. Furthermore, in some implementations, one trained model may be deployed on a client device while another trained model is deployed on platform 102.
[0071] In the following sections, the exemplary components of generated component 107 will be explained in more detail with reference to Figures 2A to 2E.
[0072] Figure 2A: Generation pipeline Figure 2A is a block diagram illustrating exemplary generation pipelines for the generation component 107 in Figure 1, in several implementation configurations. The generation pipeline is referred to as pipeline 200, but in some implementation configurations, it can be used interchangeably as generation component 107.
[0073] The pipeline 200 may include three or more stages / components. For example, the pipeline 200 may include a 2D generation AI stage 202, a 2D to 3D generation AI stage 204, a 3D mesh to avatar stage 206, and / or an optional morphing stage 208. Generally, a set of user responses / prompts to a series of images 221 and various prompts 224 may be provided as input to the pipeline 200. The pipeline 200 may output a completed avatar 265. The output avatar 265 may be available on platform 102 and / or other platforms as a fully animable avatar, avatar head, and / or avatar head, body, clothing, accessories, etc.
[0074] In some implementations, one or more of these stages are trained independently of the others. In some implementations, the stages are trained at least partially in parallel (or in conjunction). In some implementations, the stages are trained in parallel (or in conjunction).
[0075] The 2D Generative AI Stage 202 may include one or more models trained to output a CG image generated based on image 221 or a combination of image 221 and user response / prompt 224. The CG image may be output by the 2D Generative AI Stage and input by the 2D-to-3D Generative AI Stage 204.
[0076] In some implementations, the 2D generation AI stage 202 may be omitted entirely, and user selection of multiple CG images may be used instead. For example, if the user does not wish to provide an image of their face, or if the user prefers to start with a different exemplary image, the user may select computer-generated images to use for avatar generation. For example, several different CG images may be displayed in the user interface. The user may then select images that can be used as CG images input by stage 204.
[0077] The 2D-to-3D generation AI stage 204 may include one or more models trained to output polygon meshes representing CG images (for example, provided by stage 202 or user selection). For example, stage 204 may generate feature vectors, generate data structures (such as a three-plane or three-grid) based on the feature vectors, and generate polygon meshes based on the generated data structures. The data structures may also include polygon meshes depending on the configuration of the trained models deployed within stage 204. For example, if a GAN generator is deployed, the GAN generator may be trained to output a three-plane or three-grid that can be directly converted into a high-fidelity 3D polygon mesh. For example, if a rendering layer is integrated with the GAN generator, the rendering layer may directly output polygon meshes based on its internal three-plane or three-grid representations.
[0078] In some implementations, the data structure may also include two or more triplanes or lattices. For example, if a dual-GAN generator is deployed, each GAN generator within the dual-GAN generator may output individual triplanes or lattices representing different parts of the head (e.g., bald and hair). In this example, semantic operations may be used to create two or more polygon meshes resembling "building blocks." For example, if one triplane represents baldness and the second triplane represents hair, both a bald mesh and a non-bald mesh may be output. In some implementations, both a bald mesh and a mesh with only hair may also be output. Thus, stage 204 may be configured to provide multiple different meshes that can be selected to create avatars that are uniquely customized by removing hair features present in the input image and can be replaced with either baldness or different hairstyles and / or accessories.
[0079] A 3D polygon mesh is output by the 2D-to-3D generation stage 204 and can be input into the avatar stage 206 via the 3D mesh.
[0080] The Avatar Stage 206 from 3D Mesh may include one or more models trained to automatically fit and rig polygon meshes onto the avatar head and / or body. In this configuration, Stage 206 is configured to automatically position, fit, and rig the polygon meshes onto the avatar's body or skeleton. For example, the output polygon mesh 245 may be input to a multi-view texturing algorithm for creating a texturized 3D mesh that fits for rigging onto the avatar head data model, with the topology fitting algorithm adjusting the mesh to a specific avatar head topology. The texturized 3D mesh may then be provided as input to an automatic rigging algorithm for properly rigging the texturized 3D mesh onto the avatar head as Output 265. In some implementations, Stage 206 may also take two meshes output from Stage 204 and combine or assemble them to create a single mesh (for example, combining a bald head mesh with a hair mesh). In this method, unique, user-customizable hairstyles can be implemented on auto-generated avatars with reduced computation cycles compared to other conventional solutions that output a mesh that essentially combines hair and head, leaving no easy way to remove or replace the hair. In some implementation forms, both meshes are combined in stage 204, and stage 206 only provides fitting and rigging of the previously combined mesh onto the avatar's head and / or body.
[0081] In some implementations, the avatar may be morphed to different scales and / or further configured by the user via a morphing stage 208. If implemented, the morphing stage 208 may provide the user with customization options for user selection and automatic implementation of the selected customization options in the avatar output 265.
[0082] In the following sections, various exemplary components for stages 202, 204, and 206 are described in detail with reference to Figures 2B, 2C, and 2D.
[0083] Figure 2B: Unconditional model and conditional model Figure 2B is a block diagram illustrating examples of individual components of the exemplary pipeline shown in Figure 2A, in several implementation configurations.
[0084] In the simplified pipeline 200 in Figure 2B, different exemplary components that may be implemented in stages 202, 204, and 206 are illustrated in several implementation forms. Note that, according to some implementations, the entire stage 202 may be replaced with a user selection stage in which the user can choose a computer-generated exemplary image instead of providing their own image.
[0085] In an alternative example, stage 202 may include a personalization encoder 220 and a neural network 220 for generating a 2D image from an image 221 provided by the user. The personalization encoder 220 may be a machine learning model trained to output personal features of a face. For example, the personalization encoder may receive a captured input 2D image 221 of a face. The input 2D image 221 may include three or more images in some implementations. In some implementations, the input 2D image may include more or fewer images. In some implementations, the input 2D image may include three images taken (relatively) from the left perspective, front perspective, and right perspective of the face. In some implementations, the images used to train the personalization encoder 220 may also include computer-generated images from the same left perspective, front perspective, and right perspective.
[0086] The personalization encoder 220 can be trained to output feature vectors of facial features. After training, the output of the trained personalization encoder is used to train a neural network 222 that communicates operably with the personalization encoder 220.
[0087] The neural network 222 may include any suitable neural network that can be trained to output a 2D computer-generated image 225 of a face based on a response to a user prompt 224, an input base 2D image 223 (for example, a plain CG face or a CG face with few features provided as a base template), and / or an input image 221 encoded by the personalization encoder 220. In some implementations, the input base 2D image 223 is a computer-generated reference image. Note that in some implementations, the input base 2D image 223 may be omitted, and the neural network 222 may be trained to output a 2D computer-generated image 225 based on the user prompt 224 and the input image 221 encoded by the personalization encoder 220. Other modifications may also be applicable.
[0088] The neural network 222 can be further trained to incorporate features described in user responses / prompts 224 such that these features form at least a portion of the CG image 225. For example, user prompts / responses such as “thin face,” “long hair,” “big horns,” or others can be used in 2D image generation so that the CG image includes a thin face with long hair and big horns depicted therein. Given a large number of possible user responses to a prompt, it will be readily apparent that a Large-Scale Language Model (LLM) or other natural language processing model can be used to identify images containing such features (e.g., labeled features or self-descriptions of images) from multiple images (e.g., from an image repository of a virtual experience platform). The identified images can be provided along with image 223 so that the relevant features are implemented in the CG image 225. Other implementation forms may also be applicable, such as omissions where pre-coded style vectors and user prompts with those features are omitted entirely, and the output computer-generated image 225 includes one based on the input image 221 and / or base image 223.
[0089] The computer-generated 2D image output from face 225 can be passed to stage 204 for a 2D-to-3D generation task.
[0090] Stage 204 may include a conditional model 241 trained to output at least one feature vector encoding the features identified in the CG image 225. For example, conditional sampling of different features may be encoded by parameterization associated with the conditional model 241, and even with an unconditional model 240. An exemplary conditional model is illustrated in Figure 6, which will be described in more detail below.
[0091] The unconditional model 241 can be trained to output a polygon mesh 245 based on the input feature vector provided by the conditional model 241 and any parameterization encoded therein. In some implementations, the polygon mesh may be based on a direct transformation of three planes and / or three grids. In some implementations, the polygon mesh 245 can be rendered by a rendering component and / or a marching cube algorithm.
[0092] The output polygon mesh 245 can be automatically placed on the avatar's body or skeleton and passed to stage 206 for fitting and rigging. For example, the output polygon mesh 245 can be input to a topology fitting algorithm 262 to adjust the mesh to the topology of a particular avatar head. In some implementations, the topology fitting algorithm may include the operation of an auto-retopology method, an auto-retopology algorithm, and / or an auto-retopology program. For example, auto-retopology may be based on an INSTANTMESHES operation in some implementations.
[0093] The adjusted mesh can be provided as input to a multi-view texturing algorithm 264 to create a texturized 3D mesh that fits for rigging into an avatar head data model. For example, the multi-view texturing algorithm 264 may operate to exclude certain facial features, flatten the adjusted mesh, and / or texturize the flattened mesh.
[0094] Next, the textured 3D mesh may be provided as input to an automatic rigging algorithm 266 for properly rigging the textured 3D mesh onto the avatar head, as output 265. The automatic rigging algorithm 266 may re-inflate the textured mesh (if it is not yet fully inflated) and automatically fit and rig the mesh onto the avatar head. Then, in some implementations, the properly fitted and rigged avatar head may be output as output 265, or it may be rigged onto a suitable avatar body before output.
[0095] In some implementations, the conditional model 241 may include one or more encoders, and the unconditional model may include one or more GAN generators. Additional exemplary components of stage 204 are described in detail below with reference to Figures 2C and 2D.
[0096] Figure 2C: GAN generator Figure 2C is a block diagram illustrating another example of individual components of the exemplary pipeline 200 in Figure 2A, in several implementation forms. Note that stages 202 and 206 are illustrated here as containing components similar to those in Figure 2B. Therefore, any further explanation of the same components is omitted for brevity.
[0097] As illustrated, stage 204 may include a style encoder 242 configured to receive an output computer-generated 2D image of face 225. The style encoder 242 may be trained to output conditional density samples of style features based on responses to user prompts 224. For example, the style encoder 242 may identify a parameterization of an embodiment of a generative adversarial network (GAN) generator 244, so that the GAN generator 244 outputs data representing a 3D mesh usable for rigging on the avatar body.
[0098] The GAN generator 244 may include a generative adversarial network trained on millions of computer-generated input images so that the output 3D mesh contains personalized features from the computer-generated input images. In this manner, the GAN generator 244 may be trained to take a style feature vector from a style encoder 242 as input and output a 3D representation of an input 2D image 225 that contains the features encoded in the style feature vector.
[0099] The GAN generator 244 may output data representing a 3D mesh (e.g., a three-plane or three-grid) of the user's new avatar face, based on parameterization encoded in style feature vectors. The output data may be provided as input to a multiplane renderer 246 to create a scalar field representing the 3D mesh. The output scalar field may be provided as input to a trained mesh generation algorithm 248 (e.g., a ray marching algorithm, a renderer, or another algorithm) to output a polygon mesh 245 of isosurfaces represented by the scalar field.
[0100] It should be noted that the style encoder 242 may represent a conditional model, while the GAN generator 244, renderer 246, and mesh generation algorithm 248 may represent an unconditional model. In some implementations, the unconditional model may be trained first and then frozen for training of the conditional model. In some implementations, other training methods may be used.
[0101] A single GAN generator 244 may be sufficient to generate a high-fidelity three-plane representation of a face depicted in a CG image, but some features may be difficult to customize based on a single three-plane output. For example, hair features, in particular, may be difficult to customize directly based on a single three-plane representation. However, as described below, a dual-GAN generator can be deployed to generate two different three-planes or three-grids, each representing a different part of a face depicted in a CG image (e.g., parted hair and baldness).
[0102] Figure 2D: Dual-GAN generator Figure 2D is a block diagram illustrating another example of individual components of the exemplary pipeline in Figure 2A, in several implementation forms. Note that stages 202 and 206 are illustrated here as containing components similar to those in Figures 2B and 2C. Therefore, any further explanation of the same components is omitted for brevity.
[0103] As illustrated, stage 204 may include a style and / or content encoder 243 configured to receive an output computer-generated 2D image of face 225. The style / content encoder 243 may be trained to output two or more vectors representing a conditional density sampling of style and content features based on the response to user prompts 224 and the input CG image 225. For example, the style / content encoder 243 may also identify a parameterization of a dual-GAN generator 271, thereby outputting data representing a 3D mesh usable for rigging on an avatar body. For example, the dual-GAN generator 271 may be configured in some implementations as two GAN-based networks communicating with a single discriminator configured to receive outputs from each of the two GAN-based networks.
[0104] Two or more feature vectors output by the style / content encoder 243 may separately represent features of the CG image 225, such as hair and head. In this manner, different representations of the CG image may be processed separately to create a dual polygon mesh 246. In one implementation, the first feature vector of the two or more feature vectors is input to the first GAN generator of the dual-GAN generator 271, and the second feature vector of the two or more feature vectors is input to the second GAN generator of the dual-GAN generator 271.
[0105] The dual-GAN generator 271 is trained to take two different feature vectors output by the style / content encoder 242 as input and output two different data representations or data structures that can be directly converted into separate 3D meshes. In one implementation, the two different data structures include three planes or three grids. In one implementation, the first of the two different data structures represents a bald head based on the face contained in the CG image 225, and the second of the two different data structures represents hair based on the face contained in the CG image 225. In such a case, in some implementations, avatars with both bald and hairy heads can be created.
[0106] The data structure output by the dual-GAN generator 271 can be input to the opacity decoder 273. The opacity decoder 273 can be configured to output both a set of color values and a set of density for each position within each volume of the different data structures output by the dual-GAN generator. An exemplary opacity decoder is illustrated in Figure 5, which will be described in more detail below.
[0107] The opacity decoder 273 can provide outputs to both low-resolution and high-resolution networks to generate both high-resolution and low-resolution image outputs. For example, the ray marching algorithm 248 may represent a low-resolution network configured to output a low-resolution image. In some implementations, the ray marching algorithm 248 outputs a low-resolution image representing only bald heads. Furthermore, for example, the difference renderer 275 and the super-resolution neural network 277 may represent a high-resolution network configured to output a high-resolution image. In one implementation, the combination of the difference renderer 275 and the super-resolution neural network 277 outputs a high-resolution image of hair. In some implementations, a feature vector containing encoded hair features is also directly input to the super-resolution neural network 277.
[0108] Both the high-resolution and low-resolution output images are assembled into a polygon mesh so that a dual polygon mesh 246 is output by stage 204. One or both of the dual polygon meshes can be assembled by stage 206 to create either a bald or non-bald avatar as output 265 (for example, by placing a hair mesh on top of a bald mesh). Note that although two polygon meshes are generated, only a single avatar is output in this example. In other implementations, both bald and non-bald avatars may also be output. Depending on the parameterization and training of the dual-GAN generator 271, other modifications may also be applicable, including separating other features into different three-plane and three-grid configurations.
[0109] The training methods and architecture for training individual components in Stage 204 of the 2D to 3D transition will be described in detail below, with reference to Figures 3A to 3E.
[0110] Figure 3A: Basic training architecture Figure 3A is a block diagram illustrating an example of an exemplary training architecture for the exemplary pipeline shown in Figure 2A, in several implementation forms. As illustrated in Figure 3A, the training architecture 300 may include a 2D-to-3D generative AI stage 204 (the object to be trained), a discriminator 312 that operably communicates with stage 204, a content latent representation 331, a mapping network 333 that operably communicates with the content latent representation 331 and stage 204, and a number of training images 302 that are provided to stage 204 during training.
[0111] In one implementation, stage 204 receives multiple training images 302 as input. An output data structure representing a face depicted in one of the training images is generated by stage 204 and provided to the classifier 312. The output data structure is based on the training image and the style latent representation output by the mapping network 333. In some implementations, the style latent representation is based on a content latent representation 331, which may also include noise. The style latent representation is configured to capture style details and improve the style quality of the output.
[0112] The classifier 312 is configured to compare the output data structure from stage 204 with the input training image and output a Boolean value. The Boolean value is based on the classifier 312's determination that the output data value represents a "true" or "false" representation of the input image. If the Boolean value is "true", the classifier 312 is adjusted to improve its ability to distinguish the provided output from the training image. If the Boolean value is "false", stage 204 is adjusted to improve the configuration of the output data structure so that it better represents the input image and / or more accurately depicts the features of the input image.
[0113] Training in Stage 204 may be repeated over one or more training epochs for a large number (e.g., millions) of training images 302. In some implementations, training may be considered complete when a certain number of training epochs have been completed. In some implementations, training may be considered complete when convergence is expected or predicted for Stage 204. Note that under some circumstances, it may be difficult to confirm convergence, and the expectation or prediction of convergence may be based on the calculation of classifier outputs approaching 50%. In some implementations, training may be considered complete when the sampling window of classifier outputs approaches 50% on average. This can improve the probability that Stage 204 is not being trained with randomly selected classifier outputs (e.g., when the classifier is essentially "heads or tails" for its Boolean output).
[0114] The training of an unconditional model with a defined adversarial generative network is described in detail below, with reference to Figure 3B.
[0115] Figure 3B: Training architecture of the unconditional model Figure 3B is a block diagram illustrating an example of an exemplary training architecture for the exemplary components of Figure 2B in several implementation forms. As illustrated in Figure 3B, the training architecture 320 may include an unconditional model 240 (the object of training), a discriminator 312 that operably communicates with the model 240, a content latent representation 331, a mapping network 333 that operably communicates with the content latent representation 331 and the model 240, and a number of training images 302 that are provided to the conditional model 241 during training.
[0116] In some implementations, the conditional model 241 receives multiple training images 302 as input. The conditional model 241 may be configured to provide a style vector based on the training images 302 to the mapping network 333. The mapping network 331 may receive a content latent representation 331. The mapping network 333 may generate a style latent representation. For example, the content latent representation 331 may contain noise in some implementations.
[0117] The mapping network 331 provides the style latent representation and style vectors to the GAN encoder 321 of model 240. The output of the GAN encoder 321 is provided to the renderer 323, and the rendered output of the renderer 323 is provided to the GAN decoder 325 of model 240.
[0118] An output data structure representing a face depicted in one of multiple training images is generated by the GAN decoder 325 and provided to the classifier 312. The output data structure is based on a style vector and a style latent representation output by the mapping network 333.
[0119] The classifier 312 is configured to compare the output data structure from the model 240 with the input training image and output a Boolean value. The Boolean value is based on the classifier 312's determination that the output data structure represents a "true" or "false" representation of the input image. If the Boolean value is "true", the classifier 312 is adjusted to improve its ability to distinguish the provided output from the training image. If the Boolean value is "false", the model 240 is adjusted (for example, by adjusting the encoder 321 and / or decoder 325) to improve the configuration of the output data structure so that it better represents the input image and / or more accurately depicts the features of the input image.
[0120] Training of Model 240 can be repeated over one or more training epochs with millions of training images 302. In some implementations, training may be terminated when a certain number of training epochs have been completed. In some implementations, training of Model 240 may be terminated when convergence is expected or predicted. Note that under some circumstances, it may be difficult to confirm convergence, and the expectation or prediction of convergence may be based on the calculation of classifier outputs approaching 50%. In some implementations, training may be terminated when the sampling window of classifier outputs approaches 50% on average. This can improve the probability that Stage 204 is not training with randomly selected classifier outputs (for example, when the classifier is essentially "heads or tails" for its Boolean output).
[0121] The training of a GAN generator, as configured in Figure 2C, will be explained in detail below, with reference to Figure 3C.
[0122] Figure 3C: Training architecture of the GAN generator Figure 3C is a block diagram illustrating an example of an exemplary training architecture of the exemplary components of Figure 2C in several implementation forms. As illustrated in Figure 3C, the training architecture 330 may include a GAN generator 244 (the object to be trained), a multiplane renderer 246, and a marching cube algorithm 247. The architecture 330 further includes a discriminator 335 that operably communicates with the GAN generator 244. The architecture 330 also includes a content latent representation 331, a mapping network 333 that operably communicates with the content latent representation 331 and the GAN generator 244, and a set of training images 302 provided to the conditional model 241 during training.
[0123] In some implementations, the conditional model 241 receives multiple training images 302 as input. The conditional model 241 may be configured to provide a style vector based on the training images 302 to the mapping network 333. The mapping network 331 may receive a content latent representation 331. The mapping network 333 may generate a style latent representation. For example, the content latent representation 331 may contain noise in some implementations.
[0124] The mapping network 331 provides style latent representations and style vectors to the GAN generator 244. The GAN generator 244 generates an output data structure representing faces depicted in the training images among multiple images. The multiplane renderer 246 receives the output data structure and creates a scalar field representing a 3D mesh (e.g., through ray marching or other methods). In some implementations, the multiplane renderer 246 may also include multilayer perceptron layers to additionally output color values and density. The output scalar field, along with the color and density values, may be provided as input to a marching cube algorithm 247 (or another rendering algorithm such as a ray marching algorithm) to output a polygon mesh of isosurfaces represented by the scalar field.
[0125] The classifier 335 is configured to compare the output polygon mesh with the input training image and output a Boolean value. The Boolean value is based on the classifier 335's determination that the polygon mesh contains a "true" representation of the input image or a "false" representation of the input image. If the Boolean value is "true", the classifier 335 is adjusted to improve the discrimination between the provided output and the training image. If the Boolean value is "false", the GAN generator 244 is adjusted to improve the form of the output data structure used to create the polygon mesh so that it better represents the input image and / or more accurately depicts the features of the input image.
[0126] The training of the GAN generator 244 may be repeated over one or more training epochs for a large number (e.g., millions) of training images 302. In some implementations, training may be terminated when a certain number of training epochs have been completed. In some implementations, training of the GAN generator 244 may be terminated when convergence is expected or predicted. Note that under some circumstances, it may be difficult to confirm convergence, and the expectation or prediction of convergence may be based on the calculation of discriminator outputs approaching 50%. In some implementations, training may be terminated when the sampling window of discriminator outputs approaches 50% on average. This can improve the probability that the GAN generator 244 is not trained with randomly selected discriminator outputs (e.g., when the discriminator is essentially "heads or tails" for its Boolean output).
[0127] The training of the dual-GAN generator is described in detail below, with reference to Figure 3D.
[0128] Figure 3D: Training architecture of a dual-GAN generator Figure 3D is a block diagram illustrating an example of an exemplary training architecture of the exemplary components shown in Figure 2D, in several implementation forms. As illustrated in Figure 3D, the training architecture 340 may include a dual-GAN generator 271 (the object to be trained), an opacity decoder 273, a differential renderer 275, a super-resolution neural network 277, and a ray marching algorithm 248. The architecture 340 may also include a discriminator 335 that operably communicates with the dual-GAN generator 271, a content latent representation 331, a mapping network 333 that operably communicates with the content latent representation 331 and the dual-GAN generator 271, and a set of training images 302 provided to the conditional model 241 during training.
[0129] In some implementations, the conditional model 241 receives multiple training images 302 as input. The conditional model 241 may be configured to provide style and content vectors based on the training images 302 to the mapping network 333. The mapping network 333 may generate style latent representation vectors and content latent representation vectors based on the content latent representation 331. For example, the content latent representation 331 may contain noise in some implementations.
[0130] In some implementations, the mapping network 333 may include regularization based on Kullback-Leibler divergence (KLD regularization) to improve the unraveling between hair and head. For example, a variational autoencoder (VAE) may be implemented in the mapping network 333 such that, instead of directly outputting content and style vectors, mean and variance vectors are learned and then used to sample style and content vectors. The sampled style and content vectors are then used for training (e.g., to improve spatial embedding). This approach improves the unraveling and output quality during training of the dual-GAN generator 271.
[0131] In addition, by dropping the conditions, the style vectors remain independent of the camera parameters associated with the training input images. Instead, null embeddings of the camera parameters are learned and can then be used by the renderer and during inference. Furthermore, to improve parameter conditioning, noisy, quantized camera parameter conditions can be implemented in the mapping network 333.
[0132] For example, extreme camera angles relative to training images can be quantized to avoid inaccuracies related to these extreme angles. According to some implementations, the camera angle can vary such that yaw can vary in the range of [-120 degrees, 120 degrees] and pitch in the range of [-30 degrees, 30 degrees]. Other modifications and / or enhancements to camera parameters related to training images may be applicable.
[0133] The mapping network 331 provides the style latent representation, content latent representation, sampled content vector, and sampled style vector to the dual-GAN generator 271. The dual-GAN generator 271 includes a first GAN generator 342 and a second GAN generator 344. Since the two GAN generators are being trained and are aimed at producing different outputs representing both bald and hairy heads, several improvements are implemented compared to the training method associated with a single GAN generator.
[0134] The first GAN generator 342 may be configured as a whole-head three-plane generator G0 configured to generate a bald three-plane P0. The first GAN generator 342 generates a whole-head style vector
[0135]
number
[0136] This is conditioned during training.
[0137] The second GAN generator 344 may be configured as a hair three-plane generator G1 configured to generate hair three-planes P1. The second GAN generator 344 generates head and hair style vectors.
[0138]
number
[0139] Conditioned during training by combination. Head and hair style vectors.
[0140]
number
[0141] This can produce an output that allows both the head and hair to match properly for downstream processing (for example, assembling the two final meshes with little difficulty).
[0142] Additional "hairless" embed ω null This can be learned, and this represents the absence of hair (i.e.,
[0143]
number
[0144] (When this happens). Therefore, the dual-GAN generator 271 is given by Equation 1
[0145]
number
[0146] It can be represented by:
[0147] In Equation 1, G0 and G1 are such that in G1 (
[0148]
number
[0149] It shares the same architecture except for an extra input layer configured to handle ) dimensions.
[0150]
number
[0151] too
[0152]
number
[0153] It should be noted that this is supplied to G1, which allows G1 to recognize the geometric shape of the head, thereby ensuring that the generated hair plane P1 matches the bald head P0.
[0154] Each of the three generated planes contains 32 channel dimensions, which are passed to the opacity decoder 273. In some implementations, the opacity decoder 273 can be formed as a multilayer perceptron. The opacity decoder 273 is configured to output both color and density data (e.g., RGB and σ) of the head and hair composition in each voxel arrangement of the three associated planes.
[0155] For hair, the opacity decoder 273 implements a maximum function across multiple density candidates from the three planes representing the hair (see, for example, Figure 5). For color data, the opacity decoder 273 implements the combination of all outputs from each of the three planes (for example, from both of the three planes) fed into the fully connected layer, resulting in 32 dimensions for the color data.
[0156] The output from the opacity decoder 273 is fed to the differential renderer 275 (for high-resolution images) and the ray marching algorithm 248 (for low-resolution images). The low-resolution output is subject to a condition defined by a sampled style vector of the head only (e.g., not the hair). This condition can be injected through modulated convolution.
[0157] The high-resolution output is first processed by the differential renderer 275. Optionally, the differential renderer 275 may receive camera angle data associated with the input image so that the orientation of the renderer 275's output matches the orientation of the input image. However, it should be noted that the camera parameters are not used during inference, as learned null embeddings are used in place of the camera parameters during inference.
[0158] The output from the differential renderer 275 is received by the super-resolution neural network 277. The super-resolution neural network 277 takes the combined style vectors of both the head and hair. This combination can be injected using arbitrary style transfers.
[0159] The outputs of both the ray marching algorithm 248 and the super-resolution neural network 277 are fed to the classifier 335. The classifier 335 may also receive a Boolean signal "is-bald" represented by signal 346. Furthermore, the classifier 335 may receive randomly "dropped" camera parameters to avoid the classifier 335 associating fake images with camera parameters. In such a case, null embeddings are learned to be used in place of the camera parameters during inference. In addition, in some implementation forms, Gaussian noise (or other suitable noise) may be added to the camera parameters in addition to the randomly dropped camera parameters to help ensure that the classifier has no way of associating camera parameters with input samples. Through these features, a single classifier 335 can be used to train both the first and second GAN generators 342, 344, but can accurately distinguish between hair-only outputs and bald outputs. Furthermore, in order to force a untanglement between bald heads and hair, a regularization loss may be added to the hair three-plane output from the second GAN generator 344 so that it converges to null content.
[0160] The classifier 335 is configured to compare the output polygon mesh (both bald and hairy) with the input training image and output a Boolean value. The Boolean value is based on the classifier 335's determination that the selected polygon mesh contains either a "true" representation of the input image or a "false" representation of the input image, with or without hair. If the Boolean value is "true," the classifier 335 is adjusted to improve its ability to distinguish the provided output from the training image. If the Boolean value is "false," the dual-GAN generator 271 is adjusted to improve the form of the output data structure used to create the polygon mesh so that it better represents the input image and / or more accurately depicts the features of the input image.
[0161] It should be noted that several implementation choices may be made to control output quality, including avoiding hue shifts between bald and hair, thereby improving computational efficiency and reducing the visual distinction between the hair and bald planes. For example, a low-resolution network (e.g., Marching Cube 248) accepts conditions only for the bald plane. In addition, conditions for the super-resolution neural network 277 are head and hair style and content vectors. Furthermore, the discriminator 335 is passed both low-resolution and high-resolution outputs (e.g., to prevent color differences between high-resolution and low-resolution outputs). Furthermore, when extracting density, RGB and / or color content may be disabled in some situations such as when unraveling the head and hair planes. Furthermore, the opacity decoder 273 implements an unbiased fully connected layer to ensure that hair color information does not leak into the final color information for baldness. These and other modifications may be appropriate for any implementation form.
[0162] Training of the dual-GAN generator 271 can be repeated over one or more training epochs for a large number (e.g., millions) of training images 302. In some implementations, training may be terminated when a certain number of training epochs have been completed. In some implementations, training may be terminated when convergence is expected or predicted for the dual-GAN generator 271, while also having properly unraveled hair and head planes. Note that under some circumstances, it may be difficult to confirm convergence, and the expectation or prediction of convergence may be based on calculations of discriminator outputs approaching 50%. In some implementations, training may be terminated when the sampling window of discriminator outputs approaches 50% on average. This can improve the probability that the dual-GAN generator 271 is not trained with randomly selected discriminator outputs (e.g., when the discriminator is essentially "heads or tails" for its Boolean output).
[0163] After training an unconditional model, including a GAN generator or a dual-GAN generator, the unconditional model can be frozen so that training of a conditional model can begin. Training of a conditional model is described in more detail below with reference to Figure 3E.
[0164] Figure 3E: Training architecture of the conditional model Figure 3E is a block diagram illustrating an example of a typical training architecture for a conditional model in several implementation forms. As illustrated, a frozen unconditional model 240 operably communicates with a conditional model 241 and one or more loss functions 312.
[0165] The frozen unconditional model 240 may include any of the exemplary unconditional models presented above, including a GAN-based model, a single GAN generator, and / or a dual-GAN generator. Furthermore, the conditional model 241 may include any of the exemplary conditional models presented above, including a style encoder, a style / content encoder, and others. Additional exemplary conditional encoders are presented with reference to Figure 6 and are described more fully below.
[0166] The conditional model is trained by inputting a series of training images 302. The conditional model 241 generates at least one feature vector as an output in response to each training image. At least one feature vector is received by the frozen unconditional model 240.
[0167] Based on the output of the frozen unconditional model 240, one or more losses can be calculated by the loss function 312. Based on the calculated losses, the conditional model 241 is adjusted, and training is repeated on new training images from a set of images.
[0168] To improve the training and functionality of the conditional model 241, multiple losses may be implemented in the loss function 312. For example, appropriate losses may include L1-smooth loss, identity loss, learned perceptual image patch similarity (LPIPS) loss, and / or others.
[0169] In several implementations, the L1-smooth loss reconstruction of the input image pair with the generated super-resolution neural network output is implemented.
[0170] In some implementations, a reconstruction L1-smooth loss is implemented for a downscaled input image versus a low-resolution output (e.g., from the Marching Cubes algorithm). Note that this L1-smooth loss can help further ensure that the colors do not diverge between the low-resolution and high-resolution outputs.
[0171] In some implementations, identity loss is implemented by facial recognition algorithms or trained facial recognition models.
[0172] In several implementations, an LPIPS loss is implemented for input images versus generated bald superresolution outputs. For this loss, a ray marching algorithm may be performed over half the depth, receiving hair segmentation and weighting the loss according to the hair mask. Such loss calculations allow for finer matching of haired heads versus bald heads based on the same training images.
[0173] In several implementations, LPIPS loss is implemented for the input image paired with the output of the generated super-resolution neural network.
[0174] In some implementations, a GAN loss is applied to the input image versus the generated super-resolution neural network output. This loss can avoid the generation of "flat heads" or distortions in the scalp area of the generated bald head.
[0175] Training may be performed in architecture 350 in several implementations until a threshold number of individual training cycles are completed, until a certain number of training epochs based on different sets of training images are completed, until the loss is reduced to below a threshold, and / or until the loss is minimized.
[0176] As explained above, training of an unconditional model may be performed before training of a conditional model. Furthermore, in training some unconditional models, a mapping network may be used (with or without a VAE) to provide feature vectors to the GAN generator or dual-GAN generator during training. Below, exemplary mapping decoders as well as examples of appropriate mapping networks are explained in more detail with reference to Figures 4A-4C.
[0177] Figure 4: Content Mapping Figure 4A is a block diagram of the mapping network 333 in several implementation forms. As illustrated, the mapping network 333 may include a content mapping network 401 configured to output a content latent representation vector 402. The mapping network 33 may also include a content mapping decoder 403 configured to receive the content latent representation vector 402 and produce an output of a latent representation, denoted 404, to be added to a GAN generator under training.
[0178] It should be noted that conventional GAN generators configured to generate 3D meshes have an existing drawback: details such as animal horns, animal ears, robot features, and other non-human features are often difficult to capture. However, as shown in Figures 4B and 4C, the content latent representation 402 can be decoded so that spatial details are captured more accurately and effectively. In this way, the latent representation 404 added to the GAN generator during training is more likely to accurately reproduce features that cannot be associated with the average human face, thereby improving the accuracy of many different types of faces for use in avatar creation.
[0179] Figure 4B is a block diagram of an exemplary content mapping decoder 403 of the mapping network in Figure 4A, in several implementation configurations. As shown, in this implementation configuration, the content mapping decoder 403 is configured as a multilayer perceptron with Fourier features. For example, the content latent representation vector 402 is decoded through layers 405, 407, 409, and 411.
[0180] For example, first, the content latent representation vector 402 is transformed into a Fourier feature using a Fourier embedding 405. This allows for the capture of high-frequency information about the mesh and texture. The Fourier feature is then decoded by a fully connected layer 407 and a Gaussian error linear unit 409 before the final fully connected layer 411 maps the output of the GELU layer 409 to the appropriate spatial resolution of the intended GAN generator block.
[0181] It should be noted that, according to some implementations, each of layers 405, 407, 409, and 411 can be divided into two separate decoder networks operating in parallel. In these implementations, separate outputs from different GELU layers can provide latent representations 404 for different GAN generators in the dual-GAN generator.
[0182] Alternatively, a deep content decoder network can be implemented, as shown in Figure 4C.
[0183] Figure 4C is a block diagram of another exemplary content mapping decoder 403 of the mapping network in Figure 4A, in several implementation configurations. As shown, the content mapping decoder 403 may also be implemented with a reshape layer 421, an upconvert layer 423, and a fully connected layer 425 to generate a latent representation 404.
[0184] For example, the reshape layer 421 may reshape the content latent representation vector 402 into a latent representation embedding of a defined spatial dimension. This may further include a single N-layer convolutional network as a decoder for the content latent representation embedding to be upconverted.
[0185] Subsequently, a series of upconversion layers 423 gradually decode the content latent representation embedding until the final fully connected layer 425 outputs the latent representation 404. In some implementations, a convolutional neural network architecture is used for layers 421, 423, and 425.
[0186] It should be noted that, according to some implementations, each of layers 421, 423, and 425 can be divided into two separate decoder networks operating in parallel. In these implementations, the separate outputs from the layers can provide latent representations 404 for different GAN generators in the dual-GAN generator.
[0187] As described above, during training, the dual-GAN generator may output two different three-planes. These different three-planes can be fed into an opacity decoder so that density and color data are provided to render both high-resolution and low-resolution outputs. Below, an exemplary opacity decoder that can be used to extract color and density data from the three-planes output by the dual-GAN generator is described in detail with reference to Figure 5.
[0188] Figure 5: Exemplary opacity decoder Figure 5 is a block diagram of an exemplary opacity decoder 273 in several implementation configurations. As illustrated, the opacity decoder 273 may receive the output provided by a dual-GAN generator as input. For example, the output 501 from the first GAN generator and the output 502 from the second GAN generator may be input to the opacity decoder 273.
[0189] During operation and for each arrangement within the three-plane volume being processed, the opacity decoder samples the features that are input to the fully connected layers. For example, two fully connected layers 503 and 505 receive their respective outputs 501 and 502 separately. The fully connected layers 503 and 505 separate the outputs 501 and 502 into chunks, with each of the first chunks being transmitted to fully connected layers 512 and 513, and each of the second chunks being transmitted to the combined layer 507 to be combined. The combined features are provided to the fully connected layer 509, where the color value 511 is extracted.
[0190] The fully connected layers 512 and 513 operate to extract density information, and the maximum function 515 operates to provide the maximum extracted density 517 as an output.
[0191] It should be noted that the exemplary opacity decoder shown in Figure 5 can be implemented using a dual-GAN generator, such as dual-GAN generator 271.
[0192] As described above, conditional models can be implemented in several forms to provide feature vectors to the GAN generator. In some implementations, style encoders and / or style / content encoders can be used as conditional models. In some implementations, other conditional models may be appropriate. For example, a conditional model based on a vision transformer is described in detail with reference to Figure 6.
[0193] Figure 6: Exemplary Conditional Model Figure 6 is a block diagram of an exemplary conditional model 241 in several implementation forms. As illustrated, the conditional model 241 may include a vision transformer (ViT) backbone 601 configured to receive a CG image 225. The ViT backbone 601 may have all pooling layers removed to preserve high-frequency information. Furthermore, it should be noted that the ViT backbone 601, as arranged and configured in Figure 6, ensures the regression of both style vectors and content vectors using the same transformer (for example, two different types of vectors are regression in conditional inference). The ViT backbone 601 is configured to generate an embedding sequence with dimensions (1024 × 577) from the image 225.
[0194] Subsequently, the first fully connected and transposed layer 603 receives the embedding sequence. The embedding sequence is processed and transposed by the fully connected layer, resulting in an output of dimension (577 × 512).
[0195] The transformed embeddings are further processed by a second fully connected and transposed layer 605. The sequence of transformed embeddings is processed and transposed by a fully connected layer, resulting in an output of dimension (512 × 54).
[0196] Finally, the output from layer 605 is split by the splitting layer 607 into separate style vectors 610 and content vectors 611, respectively. Note that the style vectors 610 and content vectors 611 may also be referred to as “first and second style vectors,” “first and second feature vectors,” and similar phrases, without departing from the scope of this disclosure.
[0197] As detailed above, separate exemplary components that can be used together in different implementations of the 2D-to-3D generation stage 204 of the generation pipeline 200 are described in detail. Furthermore, separate training architectures for different components are exemplified and described in detail above. Below, the training methods are presented with reference to Figures 7-9.
[0198] Figures 7-9: Training methods Figure 7 is a flowchart of exemplary method 700 for training the exemplary generation pipeline of Figure 2A in several implementation forms. In some implementation forms, method 700 may be implemented, for example, on a server system, such as an online virtual experience platform 102 as shown in Figure 1, and / or according to the training architecture presented above. In some implementation forms, some or all of method 700 may be implemented on a system such as one or more client devices 110 and 116 as shown in Figure 1, and / or both on the server system and one or more client systems. In the examples described, the implementing system comprises one or more processors or processing circuits, and one or more storage devices such as a database or other accessible storage. In some implementation forms, different components of one or more servers and / or clients may execute different blocks or parts of method 700.
[0199] Method 700 begins with block 702. Block 702 involves training an unconditional model. For example, training may include training an unconditional model as described in detail with reference to Figures 3A, 3B, 3C, and 3D. Training may include training based on at least one training epoch and / or a set of training images. Block 704 follows block 702.
[0200] In block 704, it is determined whether the training of the unconditional model is complete. If the training of the conditional model is complete, block 706 follows block 704; otherwise, block 702 follows block 704, and training continues.
[0201] In block 706, the parameters of the unconditional model are frozen. Block 708 follows block 706.
[0202] In block 708, the training of the conditional model is performed. For example, the training may include training a conditional model as described in detail above with reference to Figure 3E. Block 710 follows block 708.
[0203] In block 710, it is determined whether the training of the conditional model is complete. Training may continue until one or more loss functions are minimized and / or other training thresholds are met. If the training of the conditional model is complete, block 712 follows block 710; otherwise, block 708 follows block 710, and training continues.
[0204] In block 712, once both the unconditional and conditional models have been trained, the models can be stored and / or deployed, for example, on the virtual experience platform 102 and / or system 100. Other platforms and systems can also deploy the conditional and unconditional models and provide fully animable avatars from them.
[0205] In some implementations, such as the examples illustrated in Figures 2C and 3C, the GAN generator can be deployed in an unconditional model. Figure 8 is a flowchart of an exemplary method 800 for training an unconditional model with a GAN generator, in several implementations. In some implementations, method 800 can be implemented, for example, on a server system, such as an online virtual experience platform 102 as shown in Figure 1, and / or according to the training architecture presented above. In some implementations, part or all of method 800 can be implemented on a system such as one or more client devices 110 and 116 as shown in Figure 1, and / or both on the server system and one or more client systems. In the examples described, the implementing system comprises one or more processors or processing circuits, and one or more storage devices such as a database or other accessible storage. In some implementations, different components of one or more servers and / or clients can execute different blocks or parts of method 800.
[0206] Method 800 begins with block 802. In block 802, a 3D mesh is generated by a GAN generator based on the training data. For example, the GAN generator 244 may generate three planes that are converted into a suitable 3D mesh. Block 804 follows block 802.
[0207] In block 804, the classifier determines whether the generated 3D mesh is a valid representation of the training data. For example, classifier 335 may provide a Boolean output in the decision of block 804. If the output is true or yes, block 804 is followed by block 808. If the output is false or no, block 804 is followed by block 810.
[0208] In block 808, the classifier is updated to improve its ability to distinguish between the generated 3D mesh and the training data. While in block 810, the GAN generator is updated to improve its output that deceives the classifier. Block 812 follows both blocks 808 and 810.
[0209] In block 812, it is determined whether training is complete. For example, training may be stopped by considering whether a threshold number of training images have been processed, whether a certain number of training epochs have been completed, and / or whether the GAN generator has converged. If training is complete, block 812 is followed by block 814; otherwise, block 812 is followed by block 802, and training continues, generating additional 3D meshes.
[0210] In block 814, the trained GAN generator can be saved for use and / or deployed.
[0211] As explained above, training a GAN generator may involve using a discriminator to converge the GAN generator during training. In one example of training a dual-GAN generator, a single discriminator may be used. Below, a method for training a dual-GAN generator using a single discriminator is described with reference to Figure 9.
[0212] Figure 9 is a flowchart of an exemplary method 900 for training an unconditional model having a dual-GAN generator, in several implementation forms. In some implementation forms, method 900 may be implemented, for example, on a server system, such as an online virtual experience platform 102 as shown in Figure 1, and / or according to the training architecture presented above. In some implementation forms, part or all of method 900 may be implemented on a system such as one or more client devices 110 and 116 as shown in Figure 1, and / or both on the server system and one or more client systems. In the examples described, the implementing system comprises one or more processors or processing circuits, and one or more storage devices such as a database or other accessible storage. In some implementation forms, different components of one or more servers and / or clients may execute different blocks or parts of method 900.
[0213] Method 900 begins with block 902. In block 902, two 3D meshes are generated by a dual-GAN generator based on the training data. For example, the dual-GAN generator 271 may generate two different three planes, which are then converted into their respective 3D meshes. Block 904 follows block 902.
[0214] In block 904, a single classifier determines whether both generated 3D meshes are valid representations of the training data. For example, classifier 335 may be configured to compare both bald and hairy heads, as described in detail above. The use of a single classifier can offer several advantages, including, among others, better unraveled outputs from each GAN generator.
[0215] In this example, classifier 335 may provide a Boolean output in the decision of block 904 based on both generated 3D meshes. When determining whether a hairy head has been determined, classifier 335 may operate as described above with reference to Figure 8. When determining whether a bald head has been determined, signal 346 may be passed to classifier 335 so that an appropriate decision is made between the features of a bald head. If the output is true or yes, block 904 is followed by block 908. If the output is false or no, block 904 is followed by block 910.
[0216] In block 908, the classifier is updated to improve its ability to distinguish between the generated 3D mesh and the training data. While in block 910, the dual-GAN generator is updated to improve its output that deceives the classifier. Block 912 follows both blocks 908 and 910.
[0217] In block 912, it is determined whether training is complete. For example, training may be stopped if a threshold number of training images have been processed, if a certain number of training epochs have been completed, and / or if the dual-GAN generator has converged and includes a proper untangle between hairy and bald heads. If training is complete, block 912 is followed by block 914; otherwise, block 92 is followed by block 902, and training continues, generating additional 3D meshes.
[0218] In block 914, the trained dual-GAN generator can be saved for use and / or deployed.
[0219] Deployed GAN generators, such as GAN generator 244 and dual-GAN generator 271, can be used to automatically generate avatars. Below, a more detailed explanation of automatically generating avatars using the trained models described above is presented.
[0220] Figures 10-11: Method for automatically generating avatars Figure 10 is a flowchart illustrating an exemplary method for automatically generating personalized avatars from 2D images in several implementation forms. In some implementation forms, Method 1000 may be implemented, for example, on a server system, such as an online virtual experience platform 102 as shown in Figure 1, and / or using a generation pipeline configured as shown in Figure 2C. In some implementation forms, part or all of Method 1000 may be implemented on a system such as one or more client devices 110 and 116 as shown in Figure 1, and / or on both the server system and one or more client systems. In the examples described, the implementing system comprises one or more processors or processing circuits, and one or more storage devices such as a database or other accessible storage. In some implementation forms, different components of one or more servers and / or clients may execute different blocks or parts of Method 1000.
[0221] To provide avatar generation, faces obtained from an input image may be detected and / or processed by one or more trained machine learning models. Before performing face detection or analysis, the user is provided with instructions indicating that such techniques will be used for avatar generation. If the user denies permission, automatic generation based on the input image is turned off (for example, a default character is used, or no avatar is generated at all). Images provided by the user are used specifically for avatar generation and are not stored. The user can turn off image analysis and automatic avatar generation at any time. Furthermore, face detection may be performed to encode facial features in the image, but face recognition is not performed. If the user permits the use of analysis for avatar generation, method 1000 begins at block 1102.
[0222] Block 1002 receives a set of input 2D images that capture a face and user responses / prompts. For example, the images may be two-dimensional (2D). For example, the images may include a left perspective image, a front view image, and a right perspective image. For example, the input images and associated perspectives / viewpoints may be based on computer-generated training images provided for training the generation component 107. In some implementations, the set of input images includes two or more images (for example, of the user's face or another face), and this method includes capturing two or more images of a face in an image sensor that operably communicates with a user device, receiving permission from the user to transmit two or more images to the virtual experience platform, and transmitting two or more images from the user device to the virtual experience platform. Block 1004 follows Block 1102.
[0223] In block 1004, a 2D representation of a face is generated by a trained neural network. The 2D representation is based on input images, but may be entirely computer-generated. In this scheme, the 2D representation is used not only as the base for an avatar's face, but in some implementations, it may be extended to include personalized features. For example, the generation of the 2D may include receiving a set of input images by a personalization encoder, outputting a feature vector of the face depicted in the set of input images by the personalization encoder, the feature vector being unique to the face and different from the feature vectors of other faces, receiving the feature vector by a neural network, and outputting a 2D representation based on the feature vector by the neural network. Block 1008 follows block 1004.
[0224] In block 1008, the style vector is encoded based on the user's responses / prompts to a 2D representation and / or one or more prompts (e.g., "What styles do you like?", "What style do you want?"). The style vector may be encoded by a trained style encoder, which is trained to output conditional density sampling based on the output of a trained GAN network or a trained GAN generator such as generator 244. For example, if no user responses / prompts are provided, the style vector may be based on features detected in the 2D representation. If one or more user responses / prompts are provided, the style vector may include features extracted based on the 2D representation as well as the user responses / prompts.
[0225] For example, encoding a style vector may include receiving a 2D representation with a style encoder, identifying the parameterization of a GAN with a style encoder, and encoding a conditional sampling vector by conditional density sampling of style features based on one or more user-provided answers, 2D representations, and parameterizations with a style encoder. Block 1008 is followed by block 1010.
[0226] In block 1010, a 3D mesh is generated by a trained GAN generator based on the input style vector / conditional sampling vector. For example, generating a mesh may involve the GAN receiving a conditional sampling vector and the GAN outputting a three-plane data structure representing the 3D mesh in response to receiving the conditional sampling vector. Block 1012 follows block 1010.
[0227] In block 1012, the generated 3D mesh is automatically fitted and rigged onto an avatar data model and / or avatar skeleton for deployment on a virtual experience platform or another platform. For example, automatic fitting and rigging may include adjusting the 3D mesh to a specific head topology using a topology fitting algorithm, texturing the adjusted mesh to create a textured mesh using a multi-view texturing algorithm, and rigging the textured mesh onto the avatar data model using an automatic rigging algorithm.
[0228] In some implementations, this method may also include fitting body features onto an avatar. For example, this method may include extracting body features associated with faces from a set of input images and applying the extracted body features to a rigged avatar data model containing a textured mesh to create an animable full-body avatar for the user / face represented in the set of input images. Other modifications may also be applicable.
[0229] As described above, Method 1000 may be a computer-based method that includes receiving a set of input images, generating a two-dimensional (2D) computer-generated representation of the faces depicted in the input images using at least one personalization encoder and a neural network, encoding style vectors as conditionally sampled vectors for input to a generative adversarial network (GAN) using a style encoder, the encoding being based on the 2D computer-generated representation and one or more user-provided responses to one or more prompts, and using the GAN to generate data representing a 3D mesh of a personalized avatar face based on the conditionally sampled vectors, and automatically fitting and rigging the 3D mesh onto the head of an avatar data model.
[0230] Method 1000 and related blocks 1002-1012 include features disclosed throughout this specification and its clauses, and can be modified in many ways to include features from different implementations. Furthermore, Method 1000 may be configured as computer-executable code stored on a computer-readable storage medium and / or to be executed as a sequence, set, or operation by a processing device or computer device. Other modifications are also applicable.
[0231] Blocks 1002-1012 may be executed (or repeated) in a different order than described above, and / or one or more blocks may be omitted. Method 1000 may be executed on a server (e.g., 102) and / or client devices (e.g., 110 or 116). Furthermore, parts of Method 1000 may be combined and executed sequentially or in parallel, according to any desired implementation.
[0232] The following describes in detail additional methods for automatic avatar generation (e.g., using a dual-GAN generator). Figure 11 is a flowchart of another exemplary method 1100 for the automatic generation of personalized avatars from 2D images in several implementation forms.
[0233] In some implementations, Method 1100 may be implemented, for example, on a server system, such as an online virtual experience platform 102 as shown in Figure 1, and / or using a generation pipeline configured as shown in Figure 2D. In some implementations, part or all of Method 1100 may be implemented on a system such as one or more client devices 110 and 116 as shown in Figure 1, and / or on both the server system and one or more client systems. In the examples described, the implementing system comprises one or more processors or processing circuits, and one or more storage devices such as a database or other accessible storage. In some implementations, different components of one or more servers and / or clients may execute different blocks or parts of Method 1100.
[0234] In this and other examples, to provide avatar generation, faces obtained from an input image may be detected and / or processed by one or more trained machine learning models. Before performing face detection or analysis, the user is provided with instructions indicating that such techniques will be used for avatar generation. If the user denies permission, automatic generation based on the input image is turned off (for example, a default character is used, or no avatar is generated at all). Images provided by the user are used specifically for avatar generation and are not stored. The user can turn off image analysis and automatic avatar generation at any time. Furthermore, face detection may be performed to encode facial features in the image, but face recognition is not performed. If the user permits the use of analysis for avatar generation, method 1100 begins with block 1102.
[0235] Block 1102 receives a set of input 2D images that capture a face and the user's response / prompt. For example, the images may be 2D images captured by an imaging device, such as a camera or one attached to a mobile phone. For example, the images may include a left perspective image, a front view image, and a right perspective image. For example, the input images and associated perspectives / viewpoints may be based on computer-generated training images provided for training the generation component 107. Block 1104 follows Block 1102.
[0236] In block 1104, a 2D representation of the face is generated by a trained neural network. The 2D representation is based on the input image, but in some implementations, the entire representation may be generated by the computer. In this way, the 2D representation is used not only as the base for the avatar's face, but may also be extended to include personalized features. Block 1106 follows block 1104.
[0237] In block 1106, at least two style vectors are encoded based on the user's responses / prompts to a 2D representation and / or one or more prompts (e.g., "What styles do you like?", "What styles do you want?"). The style vectors may be encoded by a trained style and content encoder, which is trained to output conditional density sampling based on the output of a trained dual-GAN generator.
[0238] In some implementations, the first style vector represents a bald head or a head without hair. Therefore, the first style vector cannot encode hair features. In one implementation, the second style vector represents hair features and is combined with encoded features from the first style vector. Therefore, the second style vector contains encoded features related to both the head and hair encoded from the 2D representation.
[0239] In some implementations, encoding a first style vector involves receiving a 2D representation with a style encoder, identifying the parameterization of a first GAN with a style encoder, and encoding the first style vector with a conditional density sampling of style features that are hair-independent, based on the 2D representation, and parameterized with a style encoder.
[0240] In some implementations, encoding the second style vector involves the style encoder receiving a 2D representation, the style encoder identifying the parameterization of the second GAN, and the style encoder encoding the second style vector using conditional density sampling of style features related to the hair drawn in the 2D representation and based on the parameterization. Block 1108 follows block 1106.
[0241] In block 1108, two data structures representing 3D meshes are generated by a trained dual-GAN generator based on input style vectors / conditional sampling vectors. For example, in some implementations, the first 3D mesh may depict a bald head, and the second 3D mesh may depict separated hair. In some implementations, both data structures may be three-plane or three-grid.
[0242] In some implementations, the first data representing the first 3D mesh is a first three-plane data structure, the second data representing the second 3D mesh is a second three-plane data structure, and the first and second three-plane data structures are unique. In addition, method 1100 may also include outputting a first decoded data from the first three-plane data structure and a second decoded data from the second three-plane data structure by an opacity decoder. In this and other examples, method 1100 may also include generating a first 3D mesh based on the first decoded data by a low-resolution rendering network and generating a second 3D mesh based on the second decoded data by a high-resolution rendering network.
[0243] In some implementations, the low-resolution rendering network includes a rendering algorithm configured to render a 3D mesh of a bald head, while the high-resolution rendering network includes a differential renderer coupled to a super-resolution neural network, and the high-resolution rendering network is configured to render a 3D mesh of hair that matches the spatial dimensions of the bald head. Block 1110 follows Block 1108.
[0244] In block 1110, the output mesh is assembled from two generated 3D meshes. For example, one mesh represents hair and the other represents a bald head, so the mesh representing hair is combined with the bald head to produce an output mesh representing a head with hair. Block 1112 follows block 1110.
[0245] In block 1112, the output 3D mesh is automatically fitted and rigged onto an avatar data model and / or avatar skeleton for deployment on a virtual experience platform or another platform. In some implementations, automatic fitting and rigging may include adjusting the output mesh to a specific head topology using a topology fitting algorithm, texturing the adjusted output mesh to create a textured output mesh using a multi-view texturing algorithm, and rigging the textured output mesh onto the avatar data model using an automatic rigging algorithm.
[0246] As described above, Method 1100 may be a computer-based method comprising: receiving a plurality of input images; generating a two-dimensional (2D) computer-generated representation of the faces depicted in the plurality of input images using at least one first encoder; encoding a first style vector for input to a first generative adversarial network (GAN) and a second GAN using a style encoder, wherein the encoding is based on the 2D computer-generated representation and encodes features unrelated to the hair depicted in the 2D computer-generated representation; and encoding a second style vector for input to the second GAN using a style encoder. The encoding is based on a 2D computer-generated representation and may include encoding features related to hair drawn in a 2D computer-generated representation; generating first data representing a first 3D mesh of a personalized avatar face based on a first style vector using a first GAN; generating second data representing a second 3D mesh of hair based on the first and second style vectors using a second GAN; concatenating the first and second 3D meshes to create an output mesh; and automatically fitting and rigging the output mesh onto the head of an avatar data model.
[0247] Method 1100 and related blocks 1102-1112 include features disclosed throughout this specification and its clauses, and can be modified in many ways to include features from different implementations. Furthermore, Method 1100 may be configured as computer-executable code stored on a computer-readable storage medium and / or to be executed as a sequence, set, or operation by a processing device or computer device. Other modifications are also applicable.
[0248] Blocks 1102-1112 may be executed (or repeated) in a different order than described above, and / or one or more blocks may be omitted. Method 1100 may be executed on a server (e.g., 102) and / or client devices (e.g., 110 or 116). Furthermore, parts of Method 1100 may be combined and executed sequentially or in parallel, according to any desired implementation form.
[0249] Below, a more detailed description of various computing devices that may be used to implement the different devices and components illustrated in Figures 1 to 11 is presented with reference to Figure 4.
[0250] Figure 12: Exemplary computing device Figure 12 is a block diagram of an exemplary computing device 1200 that may be used to implement one or more of the features described herein in several implementation configurations. In one example, device 400 may be used to implement a computer device (e.g., 102, 110, and / or 116 in Figure 1) and to perform an implementation of a suitable method described herein. The computing device 1200 may be any suitable computer system, server, or other electronic or hardware device. For example, the computing device 1200 may be a mainframe computer, desktop computer, workstation, portable computer, or electronic device (such as a portable device, mobile device, cell phone, smartphone, tablet computer, television, television set-top box, personal digital assistant (PDA), media player, game device, wearable device, etc.). In some implementations, device 1200 includes a processor 1202, memory 1204, input / output (I / O) interface 1206, and audio / video input / output devices 1214 (e.g., a display screen, touchscreen, display goggles or glasses, audio speaker, microphone, etc.).
[0251] The processor 1202 may be one or more processors and / or processing circuits for executing program code and controlling the basic operation of device 1200. “Processor” includes any suitable hardware and / or software system, mechanism, or component for processing data, signals, or other information. A processor may include a general-purpose central processing unit (CPU), multiple processing units, dedicated circuits for achieving a function, or other systems. Processing does not need to be limited to a specific geographical location or have temporal constraints. For example, a processor may perform its functions in “real-time,” “offline,” “batch mode,” etc. Parts of the processing may be performed by different (or the same) processing systems at different times and in different locations. A computer may be any processor communicating with memory.
[0252] Memory 1204 is typically provided within device 1200 for access by processor 1202 and is any suitable processor-readable storage medium suitable for storing instructions for execution by the processor, such as random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), flash memory, etc., and may be located separately from and / or integrated with processor 1202. Memory 1204 can store software that runs on server device 1200 by processor 1202, including the operating system 1208, application 1210, and associated data 1212. In some implementations, application 1210 may include instructions that enable processor 1202 to perform some or all of the functions described herein, for example, the methods shown in Figures 7 to 11. In some implementations, application 1210 may also include one or more trained models for automatically generating a stylized, customized, or personalized 3D avatar based on an input 2D image and / or responses / prompts to one or more prompts, as described herein.
[0253] For example, memory 1204 may contain software instructions for application 1210 that can provide an avatar automatically generated based on user preferences within an online virtual experience platform (e.g., 102). Any of the software in memory 1204 may, alternatively, be stored in any other suitable storage location or computer-readable medium. In addition, memory 1204 (and / or other connected storage devices) may store instructions and data used in the features described herein. Memory 1204 and any other type of storage (magnetic disks, optical disks, magnetic tapes, or other tangible media) may be considered “storage” or “storage devices”.
[0254] The I / O interface 1206 can provide functionality that enables the server device 1200 to interface with other systems and devices. For example, network communication devices, storage devices (e.g., memory and / or datastore 108), and input / output devices can communicate via interface 1206. In some implementations, the I / O interface can be connected to an interface device that includes input devices (such as keyboards, pointing devices, touchscreens, microphones, cameras, scanners, etc.) and / or output devices (such as display devices, speaker devices, printers, motors, etc.).
[0255] For illustrative purposes, Figure 12 shows one block for each of the following: processor 1202, memory 1204, I / O interface 1206, software blocks 1208 and 1210, and database 1212. These blocks may represent one or more processors or processing circuits, operating systems, memory, I / O interfaces, applications, and / or software modules. In other implementations, device 1200 may not have all of the components shown, and / or may have other elements, including other types of elements, instead of or in addition to the components shown herein. The online virtual experience platform 102 is described as performing the operations described in some implementations herein, but any preferred component or combination of components of the online virtual experience platform 102 or a similar system, or any preferred one or more processors associated with such a system, may perform the operations described.
[0256] A user device may also implement and / or be used with the features described herein. An exemplary user device may be a computer device including several components similar to device 1200, for example, a processor 1202, memory 1204, and an I / O interface 1206. An operating system, software, and applications suitable for the client device may be provided in memory and used by the processor. The I / O interface for the client device may be connected to network communication devices, as well as input and output devices, for example, a microphone for capturing sound, a camera for capturing images or videos, an audio speaker device for outputting sound, a display device for outputting images or videos, or other output devices. The display device in the audio / video input / output device 1214 may be connected to (or included in) device 1200 to display pre- and post-processed images as described herein, and such a display device may include any suitable display device, for example, an LCD, LED, or plasma display screen, a CRT, a television, a monitor, a touchscreen, a 3D display screen, a projector, or other visual display device. Some implementations can provide audio output devices, such as text-to-speech output or speech synthesis capabilities.
[0257] The methods, blocks, and / or operations described herein may be executed in an order different from that illustrated or described, and / or, where appropriate, concurrently with (partially or completely with) other blocks or operations. Some blocks or operations may be executed on a portion of the data and then later executed again on another portion of the data. In various implementations, not all of the described blocks and operations may necessarily be executed. In some implementations, blocks and operations may be executed multiple times in these methods, in different orders, and / or at different times.
[0258] In some implementations, some or all of the methods may be implemented on a system such as one or more client devices. In some implementations, one or more of the methods described herein may be implemented, for example, on a server system and / or on both a server system and a client system. In some implementations, one or more different components of a server and / or client may perform different blocks, operations, or other parts of the method.
[0259] One or more methods described herein (for example, methods 700, 800, 900, 1000, 1100, and 1200) may be implemented by computer program instructions or code that can be executed on a computer. For example, code may be executed by one or more digital processors (for example, microprocessors or other processing circuits) and stored on computer program products including magnetic, optical, electromagnetic, or semiconductor storage media, such as non-temporary computer-readable media (for example, storage media), such as semiconductor or solid-state memory, magnetic tape, removable computer diskettes, random-access memory (RAM), read-only memory (ROM), flash memory, rigid magnetic disks, optical disks, solid-state memory drives, etc. Program instructions may also be contained in and provided as electronic signals, for example, in the form of software as a service (SaaS) provided from a server (for example, a distributed system and / or a cloud computing system). Alternatively, one or more methods may be implemented in hardware (such as logic gates) or in a combination of hardware and software. Exemplary hardware may include programmable processors (e.g., field-programmable gate arrays (FPGAs), complex-programmable logic devices), general-purpose processors, graphics processors, application-specific integrated circuits (ASICs), and similar devices. One or more methods may run as part of or as a component of an application running on the system, or as an application or software running in conjunction with other applications and operating systems.
[0260] One or more methods described herein may be performed as standalone programs run on any type of computing device, programs run on a web browser, or mobile applications ("apps") run on mobile computing devices (e.g., mobile phones, smartphones, tablet computers, wearable devices (watches, armbands, jewelry, headwear, goggles, glasses, etc.), laptop computers, etc.). In one example, a client / server architecture may be used, for example, where the mobile computing device (as a client device) sends user input data to a server device and receives final output data for output (e.g., for display) from the server. In another example, all computations may be performed within a mobile app (and / or other app) on the mobile computing device. In yet another example, computations may be divided between the mobile computing device and one or more server devices.
[0261] Exemplary clause Clause 1. A computer-based method comprising: receiving a set of input images; generating a two-dimensional (2D) computer-generated representation of the faces depicted in the input images using at least one personalization encoder and a neural network; encoding style vectors as conditionally sampled vectors for input to a generative adversarial network (GAN) using a style encoder, the encoding being based on the 2D computer-generated representation and one or more user-provided responses to one or more prompts; generating data representing a 3D mesh of a personalized avatar face using the GAN based on the conditionally sampled vectors; and automatically fitting and rigging the 3D mesh onto the head of an avatar data model.
[0262] Clause 2. Generating a 2D computer-generated representation includes receiving a set of input images by a personalization encoder and outputting, by the personalization encoder, a feature vector of a face depicted in the set of input images, the feature vector being unique to the face and different from feature vectors of other faces, and receiving, by a neural network, the feature vector and outputting, by the neural network, a 2D representation based on the feature vector, the subject matter of Clause 1.
[0263] Clause 3. Encoding a style vector includes receiving, by a style encoder, a 2D representation, identifying, by the style encoder, parameterization of a GAN, and encoding, by the style encoder, a conditional sampling vector by conditional density sampling of one or more user-provided responses, the 2D representation, and style features based on the parameterization, the subject matter of any one of Clauses 1 to 2.
[0264] Clause 4. Generating a 3D mesh includes receiving, by a GAN, a conditional sampling vector and outputting, by the GAN, a three-plane data structure representing a 3D mesh in response to receiving the conditional sampling vector, the subject matter of any one of Clauses 1 to 3.
[0265] Clause 5. Further includes receiving, by a multi-plane renderer, data representing a 3D mesh from a GAN and outputting, by the multi-plane renderer, a scalar field representing the 3D mesh, the subject matter of any one of Clauses 1 to 4.
[0266] Clause 6. Further includes receiving, by a trained mesh generation algorithm, a scalar field and outputting, by the trained mesh generation algorithm, a 3D mesh as a polygon mesh of an isosurface represented by the scalar field, the subject matter of any one of Clauses 1 to 5.
[0267] Clause 7. The personalization encoder, neural network, style encoder, and GAN are deployed in a virtual experience platform, the avatar data model corresponds to an avatar participating in a virtual experience presented through the virtual experience platform, and the avatar is animatable, the subject matter according to any one of Clauses 1 to 6.
[0268] Clause 8. Automatically fitting and rigging includes adjusting a 3D mesh for a specific head topology by a topology fitting algorithm, texture mapping the adjusted mesh by a multi-view texture mapping algorithm to create a texture mapped mesh, and rigging the texture mapped mesh to an avatar data model by an automatic rigging algorithm, the subject matter according to any one of Clauses 1 to 7.
[0269] Clause 9. Further includes extracting body features of the body associated with the face from a set of input images and applying the extracted body features to a rigged avatar data model including a texture mapped mesh to create an animatable full-body avatar for the user represented in the set of input images, the subject matter according to any one of Clauses 1 to 8.
[0270] Clause 10. The set of input images includes two or more images of the user's face, and the method further includes capturing two or more images of the user's face by an image sensor operably communicating with the user device, receiving permission from the user to transmit the two or more images to the virtual experience platform, and transmitting the two or more images from the user device to the virtual experience platform, the subject matter according to any one of Clauses 1 to 9.
[0271] Clause 11. A system comprising a memory storing instructions, and a processing device coupled to the memory, the processing device being configured to access the memory and execute instructions, wherein the instructions cause the processing device to perform actions including receiving a set of input images, generating a two-dimensional (2D) computer-generated representation of a face depicted in the input images by at least one personalization encoder and a neural network, encoding a style vector by a style encoder as a conditionally sampled vector for input to a generative adversarial network (GAN), the encoding being based on the 2D computer-generated representation and one or more user-provided responses to one or more prompts, and using the GAN to generate data representing a 3D mesh of a personalized avatar face based on the conditionally sampled vector, and automatically fitting and rigging the 3D mesh onto the head of an avatar data model.
[0272] Clause 12. The subject matter described in any one of Clauses 1 to 11, wherein generating a 2D computer-generated representation includes: receiving a set of input images by a personalization encoder; outputting feature vectors of faces depicted in the set of input images by the personalization encoder, wherein the feature vectors are unique to each face and distinct from the feature vectors of other faces; receiving the feature vectors by a neural network; and outputting a 2D representation based on the feature vectors by the neural network.
[0273] Clause 13. Encoding a style vector is the subject matter described in any one of Clauses 1 to 12, which includes receiving a 2D representation by a style encoder, identifying the parameterization of a GAN by a style encoder, and encoding a conditional sampling vector by conditional density sampling of style features based on one or more user-provided responses, 2D representations, and parameterizations by a style encoder.
[0274] Clause 14. The subject matter described in any one of Clauses 1 to 13, wherein generating a 3D mesh includes, by a GAN, receiving a conditional sampling vector and by a GAN outputting a three-plane data structure representing the 3D mesh in response to receiving the conditional sampling vector.
[0275] Clause 15. The subject matter described in any one of Clauses 1 to 14, whose operation further includes receiving data representing a 3D mesh from a GAN by a multiplane renderer, and outputting a scalar field based on the data received from the GAN by a multiplane renderer.
[0276] Clause 16. The subject matter described in any one of Clauses 1 to 15, wherein the operation further includes receiving a scalar field by a trained mesh generation algorithm and outputting a 3D mesh as a polygon mesh of isosurfaces represented by the scalar field by the trained mesh generation algorithm.
[0277] Clause 17. Personalization encoders, neural networks, style encoders, and GANs are deployed on the virtual experience platform, and the avatar data model corresponds to the avatars participating in the virtual experience presented through the virtual experience platform, and the avatars are animable, subject matter as described in any one of Clauses 1 through 16.
[0278] Clause 18. Automatic fitting and rigging includes any one of Clauses 1 through 17, which includes adjusting a 3D mesh to a specific head topology by a topology fitting algorithm, texturing the adjusted mesh to create a textured mesh by a multi-view texturing algorithm, and rigging the textured mesh to an avatar data model by an automatic rigging algorithm.
[0279] Clause 19. The subject matter described in any one of Clauses 1 through 18, the operation further comprising extracting body features associated with faces from a set of input images, and applying the extracted body features to a rigged avatar data model including a textured mesh to create an animable full-body avatar for a user represented in the set of input images.
[0280] Clause 20. A non-temporary computer-readable medium on which instructions are stored, the instructions causing the processing device to perform operations including receiving a training set of input images, each image in the training set being a computer-generated image including a face and a plurality of features associated with each face; encoding a style vector for each image in the training set of input images using a style encoder as a conditional sampling vector for input to a generative adversarial network (GAN), the encoding being based on each image and a plurality of features associated therewith; the GAN generating data representing a 3D mesh of each face based on the conditional sampling vector; comparing the data generated by the GAN with each face in each image using a classifier that communicates operablely with the GAN, adjusting one or more classifier parameters of the classifier if the comparison indicates a positive comparison, adjusting one or more GAN parameters of the GAN if the comparison indicates a negative comparison; and deploying the GAN to a virtual experience platform in response to the convergence of the GAN based on each of the encoding, generation, and comparison.
[0281] Clause 21. A method implemented by a computer, comprising: receiving a plurality of input images; generating a two-dimensional (2D) computer-generated representation of the faces depicted in the plurality of input images using at least one first encoder; encoding a first style vector for input to a first generative adversarial network (GAN) and a second GAN using a style encoder, wherein the encoding is based on the 2D computer-generated representation and encodes features unrelated to the hair depicted in the 2D computer-generated representation; and encoding a second style vector for input to the second GAN using a style encoder, wherein the encoding The code is based on a 2D computer-generated representation, and the method is performed by a computer and includes encoding features related to hair drawn in a 2D computer-generated representation, generating first data representing a first 3D mesh of a personalized avatar face based on a first style vector using a first GAN, generating second data representing a second 3D mesh of hair based on the first and second style vectors using a second GAN, concatenating the first 3D mesh and the second 3D mesh to create an output mesh, and automatically fitting and rigging the output mesh onto the head of an avatar data model.
[0282] Clause 22. The subject described in any one of Clauses 1 through 21, wherein the first 3D mesh depicts a bald head and the second 3D mesh depicts separated hair.
[0283] Clause 23. Encoding a first style vector is the subject matter described in any one of Clauses 1 to 22, which includes: receiving a 2D representation by a style encoder; identifying the parameterization of a first GAN by a style encoder; and encoding the first style vector by a conditional density sampling of style features that are hair-independent, based on a 2D representation, and based on the parameterization by a style encoder.
[0284] Clause 24. Encoding a second style vector is the subject matter described in any one of Clauses 1 to 23, which includes: receiving a 2D representation by a style encoder; identifying the parameterization of a second GAN by a style encoder; and encoding the second style vector by a style encoder using conditional density sampling of style features relating to hair drawn in a 2D representation and based on the parameterization.
[0285] Clause 25. Subject matter as set forth in any one of Clauses 1 to 24, wherein the first data representing the first 3D mesh is a first three-plane data structure, the second data representing the second 3D mesh is a second three-plane data structure, the first three-plane data structure and the second three-plane data structure are unique, and the method performed by a computer includes outputting a first decoded data from the first three-plane data structure and a second decoded data from the second three-plane data structure by an opacity decoder.
[0286] Clause 26. The subject matter described in any one of Clauses 1 to 25, further comprising generating a first 3D mesh based on first decoded data by a low-resolution rendering network and generating a second 3D mesh based on second decoded data by a high-resolution rendering network.
[0287] Clause 27. Subject matter as described in any one of Clauses 1 to 26, wherein a low-resolution rendering network includes a rendering algorithm configured to render a 3D mesh of a bald head, and a high-resolution rendering network includes a differential renderer coupled to a super-resolution neural network, and the high-resolution rendering network is configured to render a 3D mesh of hair that matches the spatial dimensions of the bald head.
[0288] Clause 28. The first encoder, style encoder, first GAN, and second GAN are deployed in a virtual experience platform, and the plurality of input images are associated with and stored in the virtual experience platform, which is the subject matter described in any one of Clauses 1 to 27.
[0289] Clause 29. Automatically fitting and rigging includes adjusting an output mesh for a specific head topology by a topology fitting algorithm, creating a textured output mesh by texturing the adjusted output mesh by a multi-view texturing algorithm, and rigging the textured output mesh to an avatar data model by an automatic rigging algorithm, which is the subject matter described in any one of Clauses 1 to 28.
[0290] Clause 30. A system comprising memory storing instructions, and a processing device coupled to the memory, the processing device being configured to access the memory and execute instructions, the instructions to the processing device comprising: receiving a plurality of input images; generating a two-dimensional (2D) computer-generated representation of faces depicted in the plurality of input images by at least one first encoder; encoding a first style vector for inputs to a first generative adversarial network (GAN) and a second GAN by a style encoder, wherein the encoding is based on the 2D computer-generated representation and encodes features unrelated to hair depicted in the plurality of input images; and encoding for inputs to the second GAN by a style encoder. A system comprising a processing device that performs the following operations: encoding a second style vector, the encoding being based on a 2D computer-generated representation, encoding features related to hair depicted in multiple input images; generating first data representing a first 3D mesh of a personalized avatar face based on the first style vector using a first GAN; generating second data representing a second 3D mesh of hair based on the first and second style vectors using a second GAN; concatenating the first 3D mesh and the second 3D mesh to create an output mesh; and automatically fitting and rigging the output mesh onto the head of an avatar data model.
[0291] Clause 31. The subject described in any one of Clauses 1 through 30, wherein the first 3D mesh depicts a bald head and the second 3D mesh depicts separated hair.
[0292] Clause 32. Encoding a first style vector is the subject matter described in any one of Clauses 1 to 31, comprising: receiving a 2D representation by a style encoder; identifying the parameterization of a first GAN by a style encoder; and encoding the first style vector by a conditional density sampling of style features that are hair-independent, based on a 2D representation, and based on the parameterization by a style encoder.
[0293] Clause 33. Encoding a second style vector is the subject matter described in any one of Clauses 1 to 32, which includes: receiving a 2D representation by a style encoder; identifying the parameterization of a second GAN by a style encoder; and encoding the second style vector by a style encoder using conditional density sampling of style features relating to hair drawn in a 2D representation and based on the parameterization.
[0294] Clause 34. Subject matter as set forth in any one of Clauses 1 to 33, wherein the first data representing the first 3D mesh is a first three-plane data structure, the second data representing the second 3D mesh is a second three-plane data structure, the first three-plane data structure and the second three-plane data structure are unique, and the operation further comprises outputting a first decoded data from the first three-plane data structure and a second decoded data from the second three-plane data structure by an opacity decoder.
[0295] Clause 35. The operation of the subject matter described in any one of Clauses 1 to 34 further includes generating a first 3D mesh based on first decoded data by a low-resolution rendering network and generating a second 3D mesh based on second decoded data by a high-resolution rendering network.
[0296] Clause 36. Subject matter as described in any one of Clauses 1 to 35, wherein a low-resolution rendering network includes a rendering algorithm configured to render a 3D mesh of a bald head, and a high-resolution rendering network includes a differential renderer coupled to a super-resolution neural network, and the high-resolution rendering network is configured to render a 3D mesh of hair that matches the spatial dimensions of the bald head.
[0297] Clause 37. The subject matter described in any one of Clauses 1 through 36, wherein the first encoder, style encoder, first GAN, and second GAN are deployed on the virtual experience platform, and multiple input images are associated with and stored on the virtual experience platform.
[0298] Clause 38. Automatic fitting and rigging includes any one of Clauses 1 through 37, which includes adjusting the output mesh to a specific head topology by a topology fitting algorithm, texturing the adjusted output mesh to create a textured output mesh by a multiview texturing algorithm, and rigging the textured output mesh to an avatar data model by an automatic rigging algorithm.
[0299] Clause 39. A non-temporary computer-readable medium on which instructions are stored, wherein, in response to execution by a processing device, the processing device receives a plurality of input images, each of which depicts the same face, and each of which is a computer-generated two-dimensional image; and encodes a first style vector for input to a first generative adversarial network (GAN) and a second GAN using a style encoder, the encoding being based on the plurality of input images and encoding features unrelated to the hair depicted in the plurality of input images; and encodes a second style vector for input to the second GAN using a style encoder. A non-temporary computer-readable medium that performs the following operations: encoding, the encoding being based on multiple input images, encoding hair-related features depicted in the multiple input images; generating first data representing a first 3D mesh of a personalized avatar face based on a first style vector using a first GAN; generating second data representing a second 3D mesh of hair based on the first and second style vectors using a second GAN; concatenating the first and second 3D meshes to create an output mesh; and automatically fitting and rigging the output mesh onto the head of an avatar data model.
[0300] Clause 40. Subject matter as set forth in any one of Clauses 1 to 39, wherein encoding a first style vector includes the style encoder receiving 2D representations of multiple input images, the style encoder identifying the parameterization of a first GAN, and the style encoder encoding the first style vector using conditional density sampling of style features that are hair-independent, based on 2D representations, and based on the parameterization of a first GAN; and encoding a second style vector includes the style encoder identifying the parameterization of a second GAN, and the style encoder encoding the second style vector using conditional density sampling of style features that are hair-related in 2D representations and based on the parameterization of a second GAN.
[0301] conclusion This specification describes specific implementations, but these specific implementations are merely illustrative and not restrictive. Concepts illustrated in the examples may apply to other examples and implementations.
[0302] In situations where some of the implementations described herein may acquire or use user data (e.g., user images, user demographics, user behavior data on the platform, user search history, purchased and / or viewed products, user friendships on the platform, etc.), the user is provided with options to control whether, and how, such information is collected, stored, or used. That is, the implementations described herein will collect, store, and / or use user information in accordance with applicable regulations after obtaining explicit user authorization.
[0303] Users are provided with control over whether a program or feature collects user information about that particular user or other users associated with that program or feature. Each user whose information should be collected is presented with options (e.g., via the user interface) that allow the user to control the information collection related to that user and provide permission or authorization for whether information is collected and which parts of the information are collected. In addition, certain data may be modified in one or more ways before being stored or used, so that personally identifiable information is removed. For example, a user's identity may be modified so that no personally identifiable information can be determined (e.g., by substitution using a pseudonym, numbers, etc.). In another example, a user's geographical location may be generalized to a broader area (e.g., city, zip code, state, country, etc.).
[0304] It should be noted that the functional blocks, operations, features, methods, devices, and systems described herein may be integrated into or separated into different combinations of systems, devices, and functional blocks known to those skilled in the art. Any suitable programming language and programming technique may be used to implement routines in a particular implementation. Different programming techniques, such as procedural or object-oriented, may be employed. Routines may run on a single processing device or on multiple processors. Steps, operations, or calculations may be presented in a particular order, but this order may be modified in different particular implementations. In some implementations, multiple steps or operations described herein as sequential may be executed simultaneously. [Explanation of Symbols]
[0305] 100 Network Environment 102 Online Virtual Experience Platform 104 Virtual Experience (VE) Engine 105 Virtual Experiences 107 Generated Components 108 Datastores 110 First client device 112 Virtual Experience Applications 114, 120 users 116 Second client device 118 Virtual Experience Applications 122 Network 200 pipelines 202 2D Generation AI Stage 204 2D to 3D Generation AI Stage Stage 206: From 3D Mesh to Avatar 208 Morphing Stage 220 Personalization Encoders 220 Neural Networks 221 images 222 Neural Networks 223 Input Basic 2D Image 224 prompts 225 2D computer-generated images 240 Unconditional Model 241 Conditional Model 242 Styles Encoder 243 Style and / or Content Encoders 244 Generative Adversarial Network (GAN) Generator 245 Polygon Mesh 246 Multiplane Renderer 247 Marching Cube Algorithm 248 Trained Mesh Generation Algorithms 262 Topology Fitting Algorithms 264 Multi-view Texturing Algorithms 265 Avatars 266 Automatic Rigging Algorithms 271 dual-GAN generator 273 Opacity Decoder 275 Differential Renderer 277 Super-resolution neural networks 300 Training Architectures 302 training images 312 Identifiers 320 Training Architectures 321 GAN encoder 323 Renderer 325 GAN Decoder 330 Training Architectures 331 Content Latent Expression 333 Mapping Network 335 Classifier 342 The First GAN Generator 344 The second GAN generator 350 Architecture 400 Devices 401 Content Mapping Network 402 Content Latent Representation Vector 403 Content Mapping Decoder 404 Latent expression 405 Fourier Embedding 405, 407, 409, 411 layers 407 Fully connected layer 409 Gaussian Error Linear Unit 423 Upconvert Layer 425 Fully connected layer 501 Output 502 Output 503, 505 fully connected layer 507 Combine 509 Fully connected layer 511 color values 512, 513 fully connected layer 515 Maximum function 517 Maximum extraction density 601 Vision Transformer (ViT) Backbone 603 First fully coupled and transposed layer 605 Second fully coupled and transposed layer 607 Split layer 610 Style Vectors 611 Content Vectors 700 methods 800 ways 900 ways 1000 ways 1100 methods 1200 computing devices 1202 processors 1204 memory 1206 Input / Output (I / O) Interface 1208 Operating Systems 1210 Applications 1212 Related data 1214 Audio / Video Input / Output Devices
Claims
1. A method performed by a computer, The step of receiving a set of input images, The steps include generating a two-dimensional (2D) computer-generated representation of the face depicted in the input image using at least one personalization encoder and a neural network, A step of encoding a style vector as a conditionally sampled vector for input to a generative adversarial network (GAN) using a style encoder, wherein the encoding step is based on the 2D computer-generated representation and one or more user-provided responses to one or more prompts, The GAN generates data representing a 3D mesh of a personalized avatar face based on the conditional sampling vector, A computer-based method comprising the steps of automatically fitting and rigging the 3D mesh onto the head of an avatar data model.
2. The step of generating the aforementioned 2D computer-generated representation is: The steps include receiving the set of input images using the personalization encoder, A step of outputting a feature vector of the face depicted in the set of input images using the personalization encoder, wherein the feature vector is unique to the face and different from the feature vectors of other faces. The neural network receives the feature vector, A computer-based method according to claim 1, comprising the step of outputting the 2D representation based on the feature vectors using the neural network.
3. The step of encoding the style vector is: The steps include receiving the 2D representation using the style encoder, The steps include identifying the parameterization of the GAN using the style encoder, The computer-based method according to claim 1, comprising the steps of encoding the conditional sampling vector by conditional density sampling of the one or more user-provided responses, the 2D representations, and the style features based on the parameterization, using the style encoder.
4. The step of generating the aforementioned 3D mesh is: The GAN receives the conditional sampling vector, A computer-based method according to claim 1, comprising the step of outputting a three-plane data structure representing the 3D mesh in response to the GAN receiving the conditional sampling vector.
5. The steps include receiving the data representing the 3D mesh from the GAN using a multiplane renderer, The computer-based method according to claim 1, further comprising the step of outputting a scalar field representing the 3D mesh by the multiplane renderer.
6. The steps include receiving the scalar field using a trained mesh generation algorithm, The computer-based method according to claim 5, further comprising the step of outputting the 3D mesh as a polygon mesh of isosurfaces represented by the scalar field using the trained mesh generation algorithm.
7. The computer-implemented method according to claim 1, wherein the personalization encoder, neural network, style encoder, and GAN are deployed on a virtual experience platform, the avatar data model corresponds to an avatar participating in a virtual experience presented through the virtual experience platform, and the avatar is animable.
8. The steps of automatically fitting and rigging are: The process involves adjusting the 3D mesh for a specific head topology using a topology fitting algorithm, The steps include: creating a textured mesh by texturing the adjusted mesh using a multi-view texturing algorithm; The computer-based method according to claim 1, comprising the step of rigging the textured mesh onto the avatar data model using the automatic rigging algorithm.
9. The steps include extracting physical features of the body associated with the face from the aforementioned set of input images, The computer-based method according to claim 8, further comprising the steps of applying the extracted physical features to the rigged avatar data model including the texturized mesh to create an animable full-body avatar for a user represented in the set of input images.
10. The set of input images includes two or more images of the user's face, and the method is An image sensor that communicates with a user device in an operable manner includes the steps of capturing two or more images of the user's face, The steps include obtaining permission from the user to transmit the two or more images to the virtual experience platform, The method performed by a computer according to claim 1, further comprising the step of transmitting the two or more images from the user device to the virtual experience platform.
11. It is a system, The memory where the instructions are stored, A processing device coupled to the memory, wherein the processing device is configured to access the memory and execute the instruction, and the instruction is sent to the processing device, The step of receiving a set of input images, The steps include generating a two-dimensional (2D) computer-generated representation of the face depicted in the input image using at least one personalization encoder and a neural network, A step of encoding a style vector as a conditionally sampled vector for input to a generative adversarial network (GAN) using a style encoder, wherein the encoding step is based on the 2D computer-generated representation and one or more user-provided responses to one or more prompts, The GAN generates data representing a 3D mesh of a personalized avatar face based on the conditional sampling vector, The steps include automatically fitting and rigging the aforementioned 3D mesh onto the head of an avatar data model, A processing device that performs an operation including, A system that includes these features.
12. The step of generating the aforementioned 2D computer-generated representation is: The steps include receiving the set of input images using the personalization encoder, A step of outputting a feature vector of the face depicted in the set of input images using the personalization encoder, wherein the feature vector is unique to the face and different from the feature vectors of other faces. The neural network receives the feature vector, The system according to claim 11, comprising the step of outputting the 2D representation based on the feature vectors using the neural network.
13. The step of encoding the style vector is: The steps include receiving the 2D representation using the style encoder, The steps include identifying the parameterization of the GAN using the style encoder, The system according to claim 11, comprising the step of encoding the conditional sampling vector by conditional density sampling of the one or more user-provided responses, the 2D representation, and the style features based on the parameterization, using the style encoder.
14. The step of generating the aforementioned 3D mesh is: The GAN receives the conditional sampling vector, The system according to claim 11, further comprising the step of outputting a three-plane data structure representing the 3D mesh in response to the GAN receiving the conditional sampling vector.
15. The aforementioned operation is, The steps include receiving the data representing the 3D mesh from the GAN using a multiplane renderer, The system according to claim 11, further comprising the step of outputting a scalar field based on the data received from the GAN by the multiplane renderer.
16. The aforementioned operation is, The steps include receiving the scalar field using a trained mesh generation algorithm, The system according to claim 15, further comprising the step of outputting the 3D mesh as a polygon mesh of isosurfaces represented by the scalar field using the trained mesh generation algorithm.
17. The personalization encoder, neural network, style encoder, and GAN are deployed on a virtual experience platform, the avatar data model corresponds to an avatar participating in a virtual experience presented through the virtual experience platform, and the avatar is animable, according to claim 11.
18. The steps of automatically fitting and rigging are: The process involves adjusting the 3D mesh for a specific head topology using a topology fitting algorithm, The steps include: creating a textured mesh by texturing the adjusted mesh using a multi-view texturing algorithm; The system according to claim 11, further comprising the step of rigging the textured mesh onto the avatar data model using the automatic rigging algorithm.
19. The aforementioned operation is, The steps include extracting physical features of the body associated with the face from the aforementioned set of input images, The system according to claim 18, further comprising the step of applying the extracted physical features to the rigged avatar data model, which includes the texturized mesh, to create an animable full-body avatar for a user represented in the set of input images.
20. A non-temporary computer-readable medium in which instructions are stored, wherein the instructions are transmitted to the processing device in response to execution by the processing device. A step of receiving a training set of input images, wherein each image in the training set is a computer-generated image including a face and a plurality of features associated with each face, For each image in the aforementioned training set of input images, A step of encoding a style vector using a style encoder as a conditional sampling vector for input to a generative adversarial network (GAN), wherein the encoding step is based on each image and the plurality of features associated therewith, The GAN generates data representing the 3D mesh of each face based on the conditional sampling vector, A classifier that communicates with the GAN in an operable manner compares the data generated by the GAN with the respective faces in the respective images. If the comparison indicates a positive comparison, adjust one or more of the classifier parameters of the classifier. If the comparison indicates a negative comparison, the step is to adjust one or more GAN parameters of the GAN. The operation includes the step of deploying the GAN to a virtual experience platform in response to the convergence of the GAN based on each of the encoding, generation, and comparison steps, Non-temporary computer-readable media.
21. A method performed by a computer, Steps include receiving multiple input images, The steps include generating a two-dimensional (2D) computer-generated representation of the faces depicted in the plurality of input images using at least one first encoder, A step of encoding a first style vector for inputs to a first generative adversarial network (GAN) and a second GAN using a style encoder, wherein the encoding step is based on the 2D computer-generated representation and encodes features unrelated to the hair depicted in the 2D computer-generated representation, A step of encoding a second style vector for the input to the second GAN using the style encoder, wherein the encoding step is based on the 2D computer-generated representation, and the step of encoding features related to hair depicted in the 2D computer-generated representation, The first GAN generates first data representing a first 3D mesh of a personalized avatar face based on the first style vector, The steps include generating second data representing a second 3D mesh of the hair based on the first style vector and the second style vector using the second GAN, The steps include: creating an output mesh by concatenating the first 3D mesh and the second 3D mesh; The steps include automatically fitting and rigging the output mesh onto the head of the avatar data model, A method performed by a computer, including the above.
22. The computer-aided method according to claim 21, wherein the first 3D mesh depicts a bald head, and the second 3D mesh depicts separated hairs.
23. The step of encoding the first style vector is: The steps include receiving the 2D representation using the style encoder, The steps include identifying the parameterization of the first GAN using the style encoder, A computer-based method according to claim 21, comprising the step of encoding the first style vector by conditional density sampling of style features based on the parameterization, which is hair-independent and based on the 2D representation, using the style encoder.
24. The step of encoding the second style vector is: The steps include receiving the 2D representation using the style encoder, The steps include identifying the parameterization of the second GAN using the style encoder, The computer-based method according to claim 23, comprising the step of encoding the second style vector with the style encoder, relating to the hair depicted in the 2D representation, by conditional density sampling of style features based on the parameterization.
25. The first data representing the first 3D mesh is a first three-plane data structure, the second data representing the second 3D mesh is a second three-plane data structure, the first three-plane data structure and the second three-plane data structure are unique, and the method performed by the computer is The method performed by a computer according to claim 21, further comprising the step of outputting first decoded data from the first three-plane data structure and second decoded data from the second three-plane data structure using an opacity decoder.
26. A low-resolution rendering network generates the first 3D mesh based on the first decoded data, The computer-based method according to claim 25, further comprising the step of generating the second 3D mesh based on the second decoded data by a high-resolution rendering network.
27. The computer-based method according to claim 26, wherein the low-resolution rendering network includes a rendering algorithm configured to render a 3D mesh of a bald head, and the high-resolution rendering network includes a differential renderer coupled to a super-resolution neural network, and the high-resolution rendering network is configured to render a 3D mesh of hair that matches the spatial dimensions of the bald head.
28. The computer-based method according to claim 21, wherein the first encoder, style encoder, first GAN, and second GAN are deployed on a virtual experience platform, and the plurality of input images are associated with and stored on the virtual experience platform.
29. The steps of automatically fitting and rigging are: The process involves adjusting the output mesh for a specific head topology using a topology fitting algorithm, The steps include: creating a textured output mesh by texturing the adjusted output mesh using a multi-view texturing algorithm; The computer-based method according to claim 21, comprising the step of rigging the textured output mesh onto the avatar data model using an automated rigging algorithm.
30. It is a system, The memory where the instructions are stored, A processing device coupled to the memory, wherein the processing device is configured to access the memory and execute the instruction, and the instruction is sent to the processing device, Steps include receiving multiple input images, The steps include generating a two-dimensional (2D) computer-generated representation of the faces depicted in the plurality of input images using at least one first encoder, A step of encoding a first style vector for inputs to a first generative adversarial network (GAN) and a second GAN using a style encoder, wherein the encoding step is based on the 2D computer-generated representation and encodes features that are unrelated to hair depicted in the plurality of input images, A step of encoding a second style vector for the input to the second GAN using the style encoder, wherein the encoding step is based on the 2D computer-generated representation and includes a step of encoding features related to hair depicted in the plurality of input images. The first GAN generates first data representing a first 3D mesh of a personalized avatar face based on the first style vector, The steps include generating second data representing a second 3D mesh of the hair based on the first style vector and the second style vector using the second GAN, The steps include: creating an output mesh by concatenating the first 3D mesh and the second 3D mesh; The steps include automatically fitting and rigging the output mesh onto the head of the avatar data model, A processing device that performs an operation including, A system that includes these features.
31. The system according to claim 30, wherein the first 3D mesh depicts a bald head, and the second 3D mesh depicts separated hair.
32. The step of encoding the first style vector is: The steps include receiving the 2D representation using the style encoder, The steps include identifying the parameterization of the first GAN using the style encoder, The system according to claim 30, comprising the step of encoding the first style vector by conditional density sampling of style features based on the parameterization, which is hair-independent and based on the 2D representation, using the style encoder.
33. The step of encoding the second style vector is: The steps include receiving the 2D representation using the style encoder, The steps include identifying the parameterization of the second GAN using the style encoder, The system according to claim 32, comprising the step of encoding the second style vector with the style encoder using conditional density sampling of style features relating to the hair depicted in the 2D representation, based on the parameterization.
34. The first data representing the first 3D mesh is a first three-plane data structure, the second data representing the second 3D mesh is a second three-plane data structure, the first three-plane data structure and the second three-plane data structure are unique, and the operation is The system according to claim 30, further comprising the step of outputting first decoded data from the first three-plane data structure and second decoded data from the second three-plane data structure using an opacity decoder.
35. The aforementioned operation is, A low-resolution rendering network generates the first 3D mesh based on the first decoded data, The system according to claim 34, further comprising the step of generating the second 3D mesh based on the second decoded data by a high-resolution rendering network.
36. The system according to claim 35, wherein the low-resolution rendering network includes a rendering algorithm configured to render a 3D mesh of a bald head, and the high-resolution rendering network includes a differential renderer coupled to a super-resolution neural network, and the high-resolution rendering network is configured to render a 3D mesh of hair that matches the spatial dimensions of the bald head.
37. The system according to claim 30, wherein the first encoder, style encoder, first GAN, and second GAN are deployed on a virtual experience platform, and the plurality of input images are associated with and stored on the virtual experience platform.
38. The steps of automatically fitting and rigging are: The process involves adjusting the output mesh for a specific head topology using a topology fitting algorithm, The steps include: creating a textured output mesh by texturing the adjusted output mesh using a multi-view texturing algorithm; The system according to claim 30, further comprising the step of rigging the textured output mesh onto the avatar data model using an automated rigging algorithm.
39. A non-temporary computer-readable medium in which instructions are stored, wherein the instructions are transmitted to the processing device in response to execution by the processing device. A step of receiving multiple input images, wherein each of the multiple input images depicts the same face, and each of the multiple input images is a computer-generated two-dimensional image. A step of encoding a first style vector for inputs to a first generative adversarial network (GAN) and a second GAN using a style encoder, wherein the encoding step is based on the plurality of input images and encodes features that are unrelated to the hair depicted in the plurality of input images. A step of encoding a second style vector for the input to the second GAN using the style encoder, wherein the encoding step is based on the plurality of input images, and the step of encoding features related to hair depicted in the plurality of input images, The first GAN generates first data representing a first 3D mesh of a personalized avatar face based on the first style vector, The steps include generating second data representing a second 3D mesh of the hair based on the first style vector and the second style vector using the second GAN, The steps include: creating an output mesh by concatenating the first 3D mesh and the second 3D mesh; The steps include automatically fitting and rigging the output mesh onto the head of the avatar data model, A non-temporary computer-readable medium that enables the execution of actions including [specific actions].
40. The step of encoding the first style vector is: The steps include receiving a 2D representation of the plurality of input images using the style encoder, The steps include identifying the parameterization of the first GAN using the style encoder, The steps include encoding the first style vector using the style encoder, which is hair-independent and based on the 2D representation, by conditional density sampling of style features based on the parameterization of the first GAN, The step of encoding the second style vector is: The steps include identifying the parameterization of the second GAN using the style encoder, The non-temporal computer-readable medium according to claim 39, comprising the step of encoding the second style vector by conditional density sampling of style features relating to the hair depicted in the 2D representation, based on the parameterization of the second GAN, using the style encoder.