Avatar personalization using image generation
By receiving text prompts to generate and redirect avatar posture, combining predefined configurations and large-scale dataset learning, the scalability and accuracy of avatar personalization methods in the prior art are solved, and efficient personalized avatar generation is achieved.
Patent Information
- Application Number
- CN202380090707.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-08-30
- Filing Date
- 2023-12-20
- Publication Date
- 2025-08-08
AI Technical Summary
The methods of avatar personalization in the prior art are limited by predefined maps and avatars, cannot be expanded to millions of users, and have high computational costs, poor accuracy, and cannot be faithful to user identity and open world text understanding.
By receiving text prompts, a first pose is generated using the first model and redirecting it to the target avatar body, combining a predefined avatar configuration and a second model rendering avatar, learning the human body 3D representation using a large-scale image and video dataset to generate a personalized avatar image.
It achieves the expansion to millions of users without the need to fine-tune the user avatar, maintaining the loyalty and accuracy of the avatar appearance, providing efficient personalized avatar generation capabilities.
Smart Images

Figure CN120456961A_ABST
Abstract
Description
[0001] This application claims priority under 35 U.S.C. §119(e) to U.S. Provisional Application No. 63 / 482,288, filed on January 30, 2023. Background Art
[0002] field
[0003] The present disclosure generally relates to improving avatar personalization in virtual environments. More specifically, the present disclosure includes enabling zero-shot personalization of avatar appearance by leveraging large-scale image / video datasets to generate avatar poses based on textual prompts.
[0004] Related technologies
[0005] As social media and gaming extend social life into online media, users' virtual representations (e.g., avatars) increasingly focus on social presence and self-expression through visual features (e.g., stickers and avatars), leading to a demand for avatar personalization. Avatars are a suitable approach for personalization because they enable users to channel their identities into expressive virtual selves while mitigating deepfake and privacy concerns. However, in related art, expression is limited to a set of predefined stickers and avatars created by designers, limiting avatar personalization and user experience through social media. For example, text-to-image generation models can be personalized by fine-tuning on a few instances of a subject. However, due to the high computational / processing cost of fine-tuning for each new subject, this approach is not scalable to millions of users. This approach is also unreliable and may result in inaccurate representations of subject identities. Therefore, there is a need to provide users with improved virtual representation personalization capabilities that are applicable to many users, have improved accuracy, remain faithful to user identities, and include enhanced open-world text understanding. Summary of the Invention
[0006] The subject disclosure provides a system and method for personalized three-dimensional (3D) avatar generation. In one aspect of the disclosure, the method includes: receiving a text prompt; generating a first pose based on the text prompt using a first model; redirecting the first pose to a target avatar body; identifying a predefined avatar configuration corresponding to a user based on a user profile; transforming the target avatar body by applying the predefined avatar configuration to the target avatar body; and rendering an avatar based on the target avatar body having the predefined avatar configuration using a second model, wherein the avatar is in the first pose.
[0007] Optionally, the text prompt describes the scene and the avatar's interactions within the scene.
[0008] Optionally, the first pose is a 3D pose generated based on human joint and limb orientation and positioning parameters.
[0009] Optionally, the first pose is a human pose represented by Skinned Multi-Person Linear (SMPL) parameters.
[0010] Optionally, generating the first gesture further includes: extracting a first text embedding from the text prompt; mapping the first text embedding to a second text embedding retrieved from a dataset; selecting one or more second gestures corresponding to the second text embedding; and determining the first gesture based on the one or more second gestures.
[0011] Optionally, the computer-implemented method further comprises: training the first model on a dataset comprising body postures and corresponding text descriptions, the training comprising: identifying human bodies in images of the dataset, segmenting human bodies from the images, and extracting 3D Skinned Multi-PersonLinear (SMPL) model annotations from the human bodies segmented from the images.
[0012] Optionally, the target avatar body is a grayscale avatar-human representation.
[0013] Optionally, the redirecting includes matching corresponding joints from the first pose and the target avatar's body in position and orientation.
[0014] Optionally, the computer-implemented method further comprises: using the second model to generate an image of a scene in the virtual environment, the scene including the avatar interacting with objects in the scene.
[0015] Optionally, the second model performs conditional stable diffusion image inpainting to generate an image by image expansion from the avatar to fill the scene and objects in the scene, and the second model is conditioned on at least the avatar and the text prompt.
[0016] Another aspect of the present disclosure relates to a system configured for personalized avatar generation. The system includes: one or more processors; and a memory storing a plurality of instructions that, when executed by the one or more processors, cause the system to perform a plurality of operations. The operations include: receiving a text prompt describing a scene and an avatar's interactions within the scene; generating a first pose based on the text prompt using a first model; redirecting the first pose to a target avatar body; identifying a predefined avatar configuration corresponding to the user based on a user profile; transforming the target avatar body by applying the predefined avatar configuration to the target avatar body; and rendering an avatar based on the target avatar body having the predefined avatar configuration using a second model, wherein the avatar is in the first pose.
[0017] Optionally, the first pose is a 3D pose generated based on human joint and limb orientation and positioning parameters.
[0018] Optionally, the first pose is a human body pose represented by SMPL parameters.
[0019] Optionally, the one or more processors also execute multiple instructions for: extracting a first text embedding from the text prompt; mapping the first text embedding to a second text embedding retrieved from a dataset; selecting one or more second postures corresponding to the second text embedding; and determining the first posture based on the one or more second postures.
[0020] Optionally, the one or more processors also execute multiple instructions for performing the following operations: training the first model on a dataset comprising body poses and corresponding textual descriptions, the instructions causing the system to: identify human bodies in images of the dataset; segment human bodies from the images; and extract 3D skinned multi-person linear (SMPL) model annotations from the human bodies segmented from the images.
[0021] Optionally, the target avatar body is a grayscale avatar-human representation.
[0022] Optionally, the one or more processors further execute instructions for matching corresponding joints from the first pose and the target avatar body in position and orientation.
[0023] Optionally, the one or more processors further execute instructions for generating an image of a scene in a virtual environment using the second model, the scene including an avatar interacting with objects in the scene.
[0024] Optionally, the second model performs conditional stable diffusion image inpainting to generate an image by image expansion from the avatar to fill the scene and objects in the scene, and the second model is conditioned on at least the avatar and the text prompt.
[0025] Yet another aspect of the present disclosure relates to a non-transitory computer-readable storage medium having a plurality of instructions embodied thereon, the instructions being executable by one or more processors to perform one or more methods for personalized avatar generation and causing the one or more processors to: receive text input describing an avatar in a scene; generate a body pose based on the text prompt; redirect the body pose to a target avatar body; generate a personalized avatar based on a predefined avatar configuration applied to the target avatar body; and generate an image of the avatar in the scene based on the personalized avatar, the image including the avatar in the generated body pose having the predefined avatar configuration.
[0026] According to the content of this disclosure, these and other embodiments will be apparent.It should be understood that, according to the following detailed description, other configurations of the subject technology will be apparent to those skilled in the art, wherein, by way of example, the various configurations of the subject technology are shown and described.As will be appreciated, the subject technology can have other and different configurations, and the multiple details of these configurations can be modified in various other aspects, all of which do not depart from the scope of the subject technology.Therefore, the accompanying drawings and detailed description are essentially considered to be illustrative, rather than restrictive. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Certain embodiments of the present invention will now be described by way of example and with reference to the following drawings, in which:
[0028] Figure 1 is a schematic diagram of an environment 100 in which the methods, apparatuses, and systems described herein may be implemented according to some embodiments;
[0029] Figure 2 is a block diagram illustrating an overall framework of a system according to some embodiments;
[0030] Figure 3 is an exemplary model architecture of a text-to-3D pose generation model according to an embodiment;
[0031] Figure 4 Various examples of qualitative comparisons of text-to-3D pose generation models according to some embodiments are shown;
[0032] Figure 5 An example block diagram of a system for personalized avatar 3D pose generation according to some embodiments is shown;
[0033] Figure 6 is a flow chart of a method for personalized avatar 3D pose generation according to some embodiments; and
[0034] Figure 7 is a block diagram illustrating a computer system for performing, at least in part, one or more of the operations of the methods disclosed herein, according to some embodiments.
[0035] In the drawings, elements with the same or similar reference numerals are associated with the same or similar properties unless explicitly stated otherwise. DETAILED DESCRIPTION
[0036] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to one of ordinary skill in the art that the various embodiments of the present disclosure may be practiced without some of these specific details. In other instances, well-known structures and techniques are not shown in detail in order to avoid obscuring the present disclosure.
[0037] General Overview
[0038] Personalization and self-expression can be expressed through visual features in social media and virtual reality / augmented reality / mixed reality (VR / AR / MR) applications. Avatars are a suitable approach for personalization because they enable users to introduce their own identities into expressive virtual selves while mitigating deepfake and privacy concerns. The diversity of avatars (or other moving characters) represented using text-to-motion models trained on motion capture datasets is limited. However, with current technology, a person's expression is restricted to a set of pre-defined maps, features, and (pre-created) avatars. Different image personalization methods have poor fidelity to user identities, lack open-world text understanding, and / or are unable to scale to a large set of applications.
[0039] The various embodiments of the present disclosure describe a personalized avatar scene (PAS), which is a scalable method for zero-shot personalization of image generation that can generate images with a user's personalized avatar from a text prompt. According to the various embodiments, avatar image generation is independent of the avatar texture and style, and remains faithful to the avatar appearance without fine-tuning or training the user's avatar. Therefore, the methods and systems according to the various embodiments are easily scalable to millions of users. In order to present a 3D avatar with a pose that is faithful to a given input text (or text prompt), the various embodiments describe a pose generation model that utilizes large image and video datasets to learn a 3D representation of the human body, without relying on expensive human motion acquisition datasets to generate 3D poses (or images of 3D poses). 3D pose sequences are used to demonstrate the movement of an avatar, etc. Specifically, human motion can be represented as a sequence of 3D SMPL body parameters, where 21 body joints and root orientation are represented using 6D continuous SMPL.
[0040] PAS is trained in two stages: (1) a diffusion model for static 3D pose generation conditioned on text (i.e., a pose generation model) is trained on the Textual Pseudo-Pose (TPP) dataset; (2) the pre-trained diffusion model is extended to motion generation by adding temporal convolution and attention layers that model the new temporal dimension and are trained on motion acquisition data. In the first stage, PAS learns the distribution of human poses and their alignment with text. In the second stage, the model learns motion, i.e., how to connect poses in a temporally coherent manner.
[0041] According to various embodiments, a pose generation model is trained on a curated, large-scale dataset of in-the-wild human poses extracted from image-text datasets, resulting in significantly improved performance compared to other text-to-motion models. This model leverages a large-scale image dataset to learn 3D human pose parameters and overcomes the limitations of motion-capture datasets. The pose generation model learns to map different human poses to natural language descriptions using image-text datasets, rather than being limited to motion-capture data as traditional models do.
[0042] According to various embodiments, a human pose is generated from input text, and an avatar is rendered in the generated pose. The generated pose is redirected to the avatar body, enabling each user to render their avatar in the target generated pose. Finally, an image generation model conditioned on the input text (e.g., describing the avatar's actions and scene) and the rendered avatar is used to generate an image of the avatar contextualized in the scene. Redirecting the predicted pose to the avatar body and then rendering the user's avatar allows the avatar to remain strictly faithful to the context and appearance. The pose generation model provides zero-shot capabilities by learning to faithfully blend the rendered avatar into the scene, regardless of the avatar's style, photorealism, or context of the scene. Therefore, the pose generation model does not need to be retrained when a user's personal avatar style or appearance changes.
[0043] According to various embodiments, avatar body poses can be represented by 3D Skinned Multi-Person Linear (SMPL) model parameters. Embodiments learn human 3D pose parameters, perform text-to-3D pose generation using a transformer-based diffusion model, and then retarget the generated poses onto the avatar body to be rendered. SMPL is a realistic 3D human model based on skinning and blend shapes, learned from thousands of 3D body scans. The pose generation model is trained on a large-scale human pose dataset (e.g., text-pose pairs) constructed by extracting 3D pseudo-pose SMPL annotations from image-text datasets. According to embodiments, a large-scale Text Pseudo-Pose (TPP) dataset is extracted from an image-text dataset obtained by filtering images containing human bodies. The large-scale TPP dataset helps overcome the limited diversity of existing motion capture datasets in terms of both pose and text diversity, and has been shown to significantly simplify the dataset preparation required for human motion generation. By representing avatar body poses via 3D SMPL, the pose generation model can be easily adapted to various rendering engines to personalize the user's avatar appearance. After rendering the avatar in a pose aligned with the text prompt, the avatar is contextualized by outpainting the image from the avatar using a large-scale text-to-image generation model. In some embodiments, a stable diffusion model is fine-tuned on an image dataset with corresponding human masks and pseudo-pose annotations to generate images that are more faithful to the user's personal avatar appearance. Therefore, by introducing a human prior, PAS's zero-shot personalized image generation improves the fidelity of the user's virtual identity relative to the baseline.
[0044] Human motion can be represented as a sequence of 3D SMPL body parameters, where 21 body joints and root orientation are represented using 6D continuous SMPL. The global position of the body (e.g., avatar) in each frame can be represented by a 3D vector indicating the position in each of the x, y, and z dimensions. According to various embodiments, human motion generation utilizes a language model pre-trained on large-scale language data and naturally extends the static text-to-pose generation diffusion model to motion generation by adding temporal convolution and attention layers.
[0045] As disclosed herein, various embodiments provide a solution rooted in computer technology and originating in the field of computer networks for generating personalized avatar representations using a text-to-3D pose diffusion model. The disclosed subject matter facilitates more expressive body movements using any existing 3D rendering engine and enables high-quality, zero-shot personalization and motion generation of avatars. Consequently, the field of image generation technology and user experience are improved by producing high-quality, faithful images.
[0046] Example Architecture
[0047] Figure 1 is a schematic diagram of an environment 100 in which the methods, apparatuses, and systems described herein may be implemented according to some embodiments.
[0048] like Figure 1 As shown, environment 100 may include user device 110, platform 120, and network 130. The devices in environment 100 may be interconnected via wired connections, wireless connections, or a combination of wired and wireless connections.
[0049] User device 110 includes one or more devices capable of receiving, generating, storing, processing, and / or providing information associated with platform 120. For example, user device 110 may include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smartphone, a wireless phone, etc.), a head-mounted device (headset) or other wearable device (e.g., a virtual reality or augmented reality headset, smart glasses, a smart watch), or the like. In some embodiments, user device 110 may receive information from platform 120 and / or send information to platform 120 via network 130.
[0050] The platform 120 includes one or more devices as described elsewhere herein. In some embodiments, the platform 120 may include a cloud server or a group of cloud servers. In some embodiments, the platform 120 may be designed to be modular so that software components can be swapped in and out. Thus, the platform 120 may be easily and / or quickly reconfigured for different uses.
[0051] In some embodiments, as shown, the platform 120 can be hosted in a cloud computing environment 122. It is worth noting that although the embodiments described herein describe the platform 120 as being hosted in a cloud computing environment 122, in some embodiments, the platform 120 may not be cloud-based (i.e., may be implemented outside of a cloud computing environment) or may be partially cloud-based.
[0052] The cloud computing environment 122 includes an environment that hosts the platform 120. The cloud computing environment 122 can provide computing, software, data access, data storage (e.g., database), and other services without requiring an end user (e.g., user device 110) to be aware of the physical location and configuration of one or more systems and / or one or more devices hosting the platform 120. As shown, the cloud computing environment 122 can include a set of computing resources 124 (collectively, "computing resources 124" and individually, "computing resource 124").
[0053] The computing resources 124 include one or more personal computers, one or more workstation computers, one or more server devices, or one or more other types of computing and / or communication devices. In some embodiments, the computing resources 124 can host the platform 120. The computing resources 124 may include an application programming interface (API) layer that controls the various applications in the user device 110. The API layer can also provide tutorials about new features in the application to the user of the user device 110. Cloud resources may include computing instances executed in the computing resources 124, storage devices provided in the computing resources 124, data transmission devices provided by the computing resources 124, etc. In some embodiments, the computing resources 124 can communicate with other computing resources 124 via wired connections, wireless connections, or a combination of wired and wireless connections.
[0054] like Figure 1 As further shown, the computing resources 124 include a set of cloud resources, such as one or more applications ("APP") 124-1, one or more virtual machines ("VM") 124-2, one or more virtualized storage devices ("VS") 124-3, or one or more hypervisors ("HYP") 124-4, etc.
[0055] Applications 124-1 include one or more software applications that can be provided to or accessed by user device 110 and / or platform 120. Applications 124-1 can eliminate the need to install and execute software applications on user device 110. For example, applications 124-1 can include software associated with platform 120 and / or any other software that can be provided through cloud computing environment 122. In some embodiments, one application 124-1 can send information to or receive information from one or more other applications 124-1 through virtual machine 124-2. Applications 124-1 can include one or more modules configured to perform operations according to aspects of various embodiments. Such modules will be described in detail later.
[0056] The virtual machine 124-2 includes a software implementation of a machine (e.g., a computer) that executes programs, similar to a physical machine. The virtual machine 124-2 can be a system virtual machine or a process virtual machine, depending on the use and correspondence of the virtual machine 124-2 to any real machine. A system virtual machine can provide a complete system platform that supports executing a complete operating system ("OS"). A process virtual machine can execute a single program and can support a single process. In some embodiments, the virtual machine 124-2 can execute on behalf of a user (e.g., user device 110) and can manage the infrastructure of the cloud computing environment 122, such as data management, synchronization, or long-duration data transfer.
[0057] The virtualized storage device 124-3 includes one or more storage systems and / or one or more devices that use virtualization technology in the storage system or device of the computing resource 124. The virtualized storage device 124-3 can store multiple instructions that, when executed by the processor, cause the computing resource 124 to at least partially perform one or more operations in a method consistent with the present disclosure. In some embodiments, in the context of a storage system, the types of virtualization can include block virtualization and file virtualization. Block virtualization can refer to the extraction (or separation) of logical storage from physical storage, so that the storage system can be accessed without considering the physical storage or heterogeneous structure. This separation can allow administrators of the storage system to have flexibility in how the administrator manages the storage of end users. File virtualization can eliminate the dependency between data accessed at the file level and the physical storage location of the file. This can optimize storage usage, server consolidation and / or non-disruptive file migration performance.
[0058] Hypervisor 124-4 can provide hardware virtualization technologies that allow multiple operating systems (e.g., "guest operating systems") to execute simultaneously on a host computer (e.g., computing resource 124). Hypervisor 124-4 can provide a virtual operating platform to the guest operating systems and can manage the execution of the guest operating systems. Multiple instances of the various operating systems can share virtualized hardware resources.
[0059] The network 130 may include, for example, any one or more of a local area network (LAN), a wide area network (WAN), the Internet, etc. In addition, the network 130 may include, but is not limited to, any one or more of the following network topologies: these network topologies include bus networks, star networks, ring networks, mesh networks, star-bus networks, and tree or hierarchical networks, etc.
[0060] Figure 1The number and arrangement of devices and networks shown are provided as examples. In practice, there may be more than Figure 1 More devices and / or networks, fewer devices and / or networks, and Figure 1 The devices and / or networks shown in FIG. 1 are different devices and / or networks, or devices and / or networks arranged differently. In addition, Figure 1 Two or more of the devices shown may be implemented in a single device; or Figure 1 The single device shown may be implemented as multiple distributed devices. Additionally or alternatively, one set of devices (eg, one or more devices) of environment 100 may perform one or more functions described as being performed by another set of devices of environment 100.
[0061] Figure 2 is a block diagram illustrating the overall framework of a personalized avatar scene (PAS) system 200 according to one or more embodiments. The system 200 may include aspects including a training phase and an inference phase for one or more models. The system 200 may include one or more computing platforms configurable by machine-readable instructions. The machine-readable instructions may include one or more instruction modules. These instruction modules may include computer program modules. The instruction modules may include one or more of the following: a text encoder 210; a gesture generator 220; a gesture decoder 230; a redirection module 240; a rendering engine 250; a scene generator 260; and / or other instruction modules.
[0062] In some embodiments, one or more of the modules 210, 220, 230, 240, 250, and 260 may be included in the user device 110 and executed by one or more processors. In some embodiments, one or more of the modules 210, 220, 230, 240, 250, and 260 may be included in the cloud computing environment 122 and executed by the platform 120. In some embodiments, one or more of the modules 210, 220, 230, 240, 250, and 260 are included in and executed by a combination of the user device and the cloud computing environment.
[0063] Given an input text prompt, the text encoder 210 is a trained encoder configured to encode the input text prompt. The encoded text prompt is input to the gesture generator 220.
[0064] The pose generator 220 generates 3D body poses using a (diffusion-based) text-to-3D pose generation model. In some embodiments, the text-to-3D pose generation model maps a text embedding y from a pre-trained Contrastive Language-Image Pretraining (CLIP) model to a concatenation of a continuous body pose representation and a root orientation. According to various embodiments, the text-to-3D pose generation model can be based on a decoder-only transformer with a causal attention mask for a sequence of tokenized captions and their CLIP text embeddings, diffusion time-step embeddings, noisy body pose, and root orientation representations. and two final pose and orientation queries to predict the noise-free pose and root orientation (x p ,x r ). The text-to-3D pose generation model can be a pose generation model trained on one or more large-scale datasets (e.g., Image TPP (ITPP) dataset, HumanML3D test set, etc.) containing human poses and their textual descriptions (3D pose and text pairs). The large-scale dataset provides a variety of human poses and a large number of (text, 3D pose) sample pairs. During training, the text-to-3D pose generation model can process images to identify images with human bodies, and then extract 3D pseudo-pose SMPL annotations of human bodies. Reference Figure 3 A text-to-3D pose generation model according to one or more embodiments is described in further detail.
[0065] The gesture generated by the gesture generator 220 is input to the gesture decoder 230. The gesture decoder 230 decodes the gesture. In some embodiments, the decoded gesture is represented by 3D SMPL parameters.
[0066] Retargeting module 240 redirects or repositions the decoded gesture to a target avatar body. That is, the generated gesture is applied to the target avatar body. The target avatar body is a reference body representation that can be adjusted based on the user and the generated 3D gesture. In some embodiments, the target avatar body is a grayscale avatar-human representation. Given a generated 3D gesture, retargeting module 240 can retarget the gesture by converting the generated 3D gesture to a target avatar pose through an optimization process that matches the position and orientation of corresponding joints between the target avatar body and the generated gesture.
[0067] The rendering engine 250 uses a text-to-image generation model to render an image of a posed avatar based on a generated pose that conforms to the text prompt. For example, a latent stable diffusion text-to-image generation model (hereinafter referred to as a "personalized image generation model") can be used to generate the text-to-image prior. The posed avatar is a target avatar body in a generated pose (or target pose) with the user's personal avatar configuration. The appearance of the user's avatar is applied (or copied) to the target avatar body, allowing any user to render their own avatar in the generated pose. The rendered image can be an RGBA image. Based on the converted pose and the personalized avatar configuration (including the shape and texture of the avatar image), the avatar image is rendered in the target pose (e.g., the pose generated by the pose generator 220).
[0068] According to various embodiments, to perform zero-shot personalization given a generated 3D pose, an image of a posed avatar is rendered using a target avatar body as a reference input. To personalize image generation in system 200 with a high-quality user avatar, retargeting module 240 utilizes an internal avatar representation (i.e., the target avatar body) and rendering engine 250. Thus, system 200 can be used with a variety of avatar types, not limited to SMPL-based avatars. According to various embodiments, retargeting module 240 and rendering engine 250 are also "zero-shot" in that they do not require any training or prior knowledge of the user's avatar configuration to generate an image of the avatar based on an input text prompt.
[0069] The scene generator 260 generates a personalized image of a user's avatar in a scene corresponding to the input text. In some embodiments, the personalized image includes the user's avatar interacting with objects in the scene. The scene generator 260 can utilize a personalized image generation model to generate an image of the user's avatar and / or a scene including the avatar. Given a rendered image of a posed avatar and a text prompt describing the scene and interactions, the scene generator 260 performs conditional stable diffusion image inpainting to generate a personalized image by image expansion from the avatar to fill in the rest of the scene and objects. The personalized image generation model is conditioned on a rendered avatar in a target pose and an input text prompt describing the avatar's actions and scene. Using a grayscale human body representation as a reference for the target avatar's body improves the personalized image generation model's understanding of human body orientation and limb positioning, providing a more accurate image of the human body.
[0070] According to some embodiments, the rendered image of the posed avatar (from the rendering engine 250) may be pasted onto the generated image (from the scene generator 260) in a post-processing step.
[0071] According to various embodiments, hand and facial gesture parameters are added to the text-to-3D gesture generation model as additional targets for the redirection module 240. Thus, various embodiments can also control hand gestures and facial expressions through input text prompts, thereby achieving better expression through the rendered avatar. In this way, the personalized image of the user's avatar can include personalized gestures, facial expressions, hand movements, etc. generated based on the input text prompts.
[0072] In some embodiments, the personalized image generation model trains a convolutional neural network (UNet) on the learned latent space of an image autoencoder. The image autoencoder can be conditioned on a time step t and a CLIP text encoding (or embedding). The personalized image generation model is trained on a large-scale image-text dataset (e.g., an ITPP dataset of full-body human bodies), thereby having a broad open-world visual text understanding. Various embodiments may include, when training one or more models described herein, using segmentation (e.g., panoptic segmentation) to crop human bodies in the dataset to condition and train the UNet on the cropped human bodies. During inference (testing), the UNet is conditioned on images of avatars rendered in poses generated by the text-3D pose generation model. By operating the personalized image generation model in a compressed latent space rather than, for example, a pixel space, the personalized image generation model is more efficient than other large-scale text-to-image generation models. In addition, according to various embodiments, the personalized image generation model does not require a super-resolution step, etc. to generate high-resolution images.
[0073] According to various embodiments, the UNet can also be conditioned on a grayscale body rendered in a generated 3D pose. This introduces a human body prior into the personalized image generation model so that it can generate more realistic avatar-object interactions. Conditioning on a grayscale body rendering in a generated 3D pose also helps eliminate the gap between the human body (during training) and the avatar (during testing) because the grayscale body rendering is the same for everyone. In some embodiments, the UNet conditioning can be enhanced by downsampling the spatial dimensions and conditioning the UNet on the downsampling factor. This helps the personalized image generation model be less sensitive to the appearance of the avatar when stylizing the image because the downsampling removes texture details when conditioning, thereby improving the model's fidelity to the avatar's appearance.
[0074] According to various embodiments, the full set of personalized adjustments to the UNet can include conditioning on text describing the scene (i.e., text input) and avatar-object interactions (y), a rendered (RGBA) image of the avatar in the target pose (p), a personalized downsampling rate (w), and a time step (t). The full set of personalized adjustments to the UNet (of the personalized image generation model) can be achieved through the latent diffusion model loss To learn, the latent diffusion model loss can be expressed as:
[0075]
[0076] To condition UNet on an RGBA image p, ε(x) is used to encode the RGB channels into z p And z along the channel dimension p Connect to z t The α channel is individually downsampled to the spatial dimension of z (64×64) and concatenated along the channel dimension. To condition on the enhanced personalized downsampling rate w, the personalized downsampling rate w is encoded using a sinusoidal embedding, projected to the dimension of the CLIP text encoding, and concatenated with the text encoding for cross attention.
[0077] Various embodiments may include avatar motion generation through motion diffusion modeling using a UNet architecture. As described above, the rendered avatar body pose P can be represented as a body pose sequence of N frames [P1, P2, ... P] using 3D SMPL body parameters and avatar motion. N ]. A 6D continuous SMPL representation can be used for 21 body joints and root orientations, including the global position of each frame represented by a 3D vector indicating the position of the avatar in each of the x, y, and z dimensions, resulting in each According to various embodiments, avatar motion generation consists of three main components: (1) a pre-trained language model trained on large-scale language data, (2) a pose generation model (trained on the ITPP dataset), and (3) a set of temporal convolutional and attention layers that extend the pose generation model to the temporal dimension for motion generation and to capture dependencies between frames. t =α t ∈-σ t x) is used to train pose and motion diffusion models for numerical stability. In the UNet architecture, denoising is performed simultaneously on all motion frames, resulting in greater temporal coherence in the generated avatar motion, without any specific penalty on motion speed for smooth motion synthesis. The motion diffusion model also benefits from no classifier guidance at inference by conditioning it on empty text for a predetermined amount of time during training (e.g., 10% of the time).
[0078] Figure 33D pose generation model according to various embodiments. The model architecture 300 includes input text 310. The input text 310 can be a series of characters, a phrase, or a description of a desired pose, scene, or context surrounding an avatar. The input text 310 is encoded at a text encoder 320 and passed through a linear layer 330 to generate a text embedding. The linear layer 330 can be collectively referred to as at least one of linear 330-1, 330-2, 330-3, and 330-4. The text-to-3D pose generation model is a UNet-based diffusion model that is trained on a large-scale human pose dataset to generate 3D SMPL pose parameters conditioned on text embeddings. For example, the text embeddings can come from a large frozen language model.
[0079] The input sequence of the model architecture 300 includes the text embedding y, the token embedding 340, the diffusion time step t, and the noisy pose and root orientation representation All of these are projected to the transformer dimension to be decoded by the transformer decoder 350. The text embedding y can be a CLIP text embedding. The pose generation input can also include a batch size B and a channel dimension C (e.g., the channel dimension can be 135). The UNet of this model can be built from a sequence of: (1) a residual block with a 1×1 2D convolutional layer conditioned on the diffusion time step t embedding and the text embedding, and (2) an attention block that focuses on the text information and the diffusion time step t. During training (on the ITPP dataset), the translation parameters in all training examples can be set to zero (i.e., the person is set in the center of the scene).
[0080] In some embodiments, position embeddings are added to each token in the above sequence. During training, noiseless pose and root orientation representations (x p ,x r ).
[0081] The text-to-3D pose generation model can be trained using the mean squared error loss To directly predict the noise-free pose and root orientation, the mean squared error loss can be expressed as:
[0082]
[0083] where y is the avatar-object interaction.
[0084] According to various embodiments, the avatar body pose can be represented by 3D SMPL parameters. The system 200 can learn to predict the avatar body pose from text through a transformer (diffusion) based text-to-3D pose generation model, and then redirect the predicted pose to the avatar body that is subsequently rendered. In some embodiments, a pre-trained varying human pose prior (e.g., VPoser) is trained from a large dataset of human poses represented as SMPL bodies. The pre-trained varying human pose prior can be used as a varying autoencoder to effectively model the distribution of the pose prior, rather than learning to generate all 3D joint rotations. The pre-trained varying human pose prior can be configured to learn a latent representation of human pose and regularize the distribution of the latent representation to a normal distribution (e.g., regularize them in the natural distribution of the human body). The pre-trained varying human pose prior can be in a continuous embedding space that is well suited for progressive noise and denoising diffusion processes.
[0085] To improve sample quality during inference of the text-to-3D pose generation model, various embodiments include classifier-free bootstrapping by discarding a predetermined amount (e.g., 10% of the time) of text conditions during training. After training, the body pose encodings generated by the diffusion model are passed to a pre-trained VPoser decoder to generate 3D SMPL body rotation vectors.
[0086] In some embodiments, the pose generation model is extended to learn the temporal dimension of motion generation. The convolutional and attention layers of the UNet can be modified to include reshaping the input motion sequence into a tensor of shape B×C×N×1×1, where C is the length of the pose representation and N is the number of frames. A 1D temporal convolution layer is stacked (or added) after each 1×1 2D convolution. In some embodiments, a temporal attention layer is added after each cross-attention layer during the motion fine-tuning phase. This allows the model to train new 1D temporal layers from scratch while loading pre-trained convolution weights from the pose generation model. The kernel size of these temporal convolution layers can be set to three, as opposed to the unit convolution kernel size for 2D convolutions. A similar dimensionality decomposition strategy is applied to the attention layer, where a newly initialized temporal attention layer is stacked to each pre-trained attention block from the pose generation model (i.e., from the UNet). In some embodiments, the rotational position embedding is used for the temporal attention layer.
[0087] Table 1 evaluates text-to-3D pose generation based on PAS for text-to-human motion generation models by manual evaluation on 390 crowd-sourced prompts and automatic measurement on the HumanML3D test set. The PAS evaluated in Table 1 was trained on the ITPP dataset and the HumanML3D training set. Given each text prompt, five poses are generated and rendered according to the systems and methods of each embodiment. Multiple poses are ranked based on the CLIP similarity score between the rendered avatar as a 2D image and the input text prompt to select the best generated sample. To compare each benchmark, a motion sequence is generated given a text prompt, and the best representative frame is selected based on the CLIP similarity score with the frame encoded with the text. For example, benchmarks 1 to 4 can represent TEMOS (generating human motion based on text description), MotionCLIP (generating motion in CLIP space), AvatarCLIP (text-driven generation), and motion diffusion model (MDM), respectively.
[0088] Table 1: Comparison of PAS with other text-to-human motion generation models
[0089]
[0090] As shown in Table 1, PAS was compared with each method in terms of the correctness of each generated pose given a text prompt. The results are expressed as the percentage of user preferences for the text-to-3D pose generation model relative to each benchmark on 390 crowdsourced prompts. The prompt includes a description of the action as well as some context of the scene. The Fréchet Pose Distance (FPD), calculated as the Fréchet distance in the Vposer embedding space, measures the overall quality and diversity of the samples. In addition, the CLIP Similarity (CLIPSIM) score between the rendered avatar corresponding to each generated pose and its text description is calculated to measure the pose-to-text fidelity. As shown in Table 1, the pose generation model according to each embodiment (PAS) is superior to the other baseline methods in terms of fidelity to the input text.
[0091] Figure 4 Various examples of qualitative comparisons of the PAS model with other benchmarks are shown. Text 410 is an input text prompt. PAS 420 is an output pose rendered from a text-to-3D pose generation model according to various embodiments. Outputs 430 / 440 / 450 / 460 are output poses generated based on text 410 by other text-to-human motion generation models (i.e., benchmarks 1 to 4 in Table 1). Figure 4As shown, the generated gesture (PAS 420 ) shows increased fidelity to the input text (text 410 ).
[0092] Table 2 shows the results of an ablation study performed to understand the impact of different continuous pose representations and training datasets (e.g., 32-dimensional VPoser embeddings of body poses, 6D continuous vectors of all joints, and (according to the embodiments described) 6D continuous vectors of all joints regularized by a pre-trained VPoser). In this study, the HumanML3D dataset was modified for single-frame pose generation. Manually evaluated text fidelity is reported as the percentage of user preference for the last row setting of a selected set of 390 prompts (i.e., anything above 50% means that 6D+Vposer trained on ITPP data is favored). For CLIP similarity, according to various embodiments, SMPL avatars are rendered using 3D poses predicted by the text-to-3D pose generation model.
[0093] Table 2: Ablation study results
[0094]
[0095] As shown in Table 2, the results confirm the benefits of VPoser representation in the training of diffusion models or as a post-regularizer (indicated by the last two rows in Table 2). According to various embodiments, the study also compared the performance of the text-to-3D pose generation model when trained on ITPP data with that when trained on a modified HumanML3D dataset for pose generation. This study again confirms the impact of the ITPP dataset compared to the limited existing action collection datasets by providing a wide range of human actions and training more general action generation models.
[0096] The personalized image generation model of each embodiment provides fine-tuning and architectural modifications that improve the fidelity to the avatar appearance when compared to the image extension baseline on crowdsourced prompts. Table 3 evaluates the personalized image generation model according to PAS against the stable diffusion image inpainting / image extension benchmark. Table 3 reports the manual evaluation (avatar and text fidelity) and the automatic metric (CLIP similarity). The results of the manual evaluation are reported as a percentage of user preference for the output of each embodiment relative to the benchmark (i.e., any value above 50% indicates that the embodiment's method is favored). The "fine-tuned" method in Table 3 is the stable diffusion image inpainting / image extension model (baseline) fine-tuned on the ITPP dataset. This isolates the value of the architectural modifications made according to each embodiment.
[0097] Table 3: Comparison of PAS with other image generation models
[0098]
[0099] As shown in Table 3, image generation based on PAS demonstrates improved fidelity to avatar appearance compared to related methods / systems. The baseline model for stable diffusion image inpainting / image expansion is trained on random masks rather than human masks. This leads to the illusion of new limbs and clothing, disrupting the avatar's identity. Fine-tuning the human mask and architectural modifications performed according to various embodiments improve fidelity to avatar appearance. In particular, adjustments to the grayscale body rendering by the personalized image generation model, according to various embodiments, improve the model's understanding of human orientation and limb positioning, reducing the illusion of new body parts.
[0100] Figure 5 An example block diagram of a system 500 for generating personalized avatar 3D poses according to one or more embodiments is shown. The personalized avatar may be in a virtual environment configured for customization or user interaction. The system 500 may include one or more computing platforms configurable by machine-readable instructions. The machine-readable instructions may include one or more instruction modules. These instruction modules may include computer program modules. Figure 5 As shown, the instruction module may include one or more of the following: a receiving module 510; a generating module 520; an orientation module 530; a recognition module 540; a conversion module 550; a rendering module 560; and / or other instruction modules.
[0101] In some embodiments, one or more of the modules 510, 520, 530, 540, 550, and 560 may be included in the user device 110. In some embodiments, one or more of the modules 510, 520, 530, 540, 550, and 560 may be included in the cloud computing environment 122 and executed by the platform 120. In some embodiments, one or more of the modules 510, 520, 530, 540, 550, and 560 are included in and executed by a combination of the user device and the cloud computing environment.
[0102] The receiving module 510 is configured to receive a text prompt. The system 500 may also include extracting a text embedding from the text input. The text prompt may be encoded using a pre-trained text encoder to generate the text embedding. The text prompt may describe the scene and the avatar's interaction within the scene.
[0103] The generation module 520 is configured to generate a target pose based on the text prompt using a text-to-3D pose generation model. The target pose can be a 3D pose generated based on human joint and limb orientation and positioning parameters. In some embodiments, the target pose is a human pose represented by SMPL parameters. To generate the target pose, the text embedding from the text prompt can be mapped to the text embedding retrieved from the dataset. One or more body poses corresponding to the text embedding retrieved from the dataset can be selected and the target pose generated based on the selected body poses.
[0104] The orientation module 530 is configured to redirect the target pose to the target avatar body. The target avatar body can be, for example, a grayscale avatar-human representation that serves as a personalized reference for the avatar pose. The orientation module 530 can also be configured to match corresponding joints based on the position and orientation of the target pose and the target avatar body to redirect the target pose.
[0105] The identification module 540 is configured to identify an avatar configuration corresponding to the user who submitted the text prompt. The avatar configuration may be pre-defined on the user's profile or account.
[0106] The conversion module 550 is configured to apply a predefined avatar configuration to a target avatar body.
[0107] The rendering module 560 is configured to render a personalized avatar image using the personalized image generation model. The personalized avatar image is based on at least the target avatar body, the avatar configuration, and the target pose. Thus, the personalized avatar image can be rendered in a manner such that the avatar in the image mimics the avatar configuration and is in the generated target pose.
[0108] The system 500 may also include one or more modules configured to train a personalized image generation model on a dataset comprising body poses and corresponding text descriptions. The training may include identifying human bodies in images of the dataset, segmenting the human bodies, and extracting SMPL annotations from the segmented human bodies.
[0109] System 500 may also include one or more modules configured to use a personalized image generation model to generate an image of a scene in a virtual environment based on a personalized avatar image and textual prompts. The scene may include an avatar interacting with objects in the scene according to the textual prompts. In some embodiments, the rendered image of the avatar in a target pose may be overlaid or pasted onto the image of the scene. As described above, the personalized image generation model may perform conditional stable diffusion image inpainting to fill in an existing scene and objects in the existing scene by image expansion from the avatar to generate an image. The personalized image generation model may be conditioned on at least the avatar and the textual prompt. In some embodiments, the personalized image generation model is conditioned on the avatar and the textual prompt, as well as the avatar's interactions with objects in the scene, the downsampling value, and the time step.
[0110] although Figure 5 Example blocks of system 500 are shown, but in some implementations, system 500 may include more Figure 5 The blocks depicted in the system may be more blocks, fewer blocks, different blocks, or blocks arranged in a different manner. Additionally or alternatively, two or more of the multiple blocks of the system may be combined together.
[0111] Figure 6 6 is a flow chart of a method 600 for personalized avatar 3D pose generation according to one or more embodiments. The techniques described herein may be implemented as one or more methods performed by one or more physical computing devices; as one or more non-transitory computer-readable storage media storing a plurality of instructions that, when executed by one or more computing devices, cause the one or more methods to be performed; or as one or more physical computing devices specifically configured with a combination of hardware and software to perform the one or more methods.
[0112] In some embodiments, one or more of the steps of method 600 may be performed by one or more of modules 510, 520, 530, 540, 550, and 560. Figure 6 One or more operation blocks of method 600 may be performed by a processor circuit executing instructions stored in a storage circuit, a client device, a remote server, or a database, the client device, the remote server, or the database being communicatively coupled via a network. In some embodiments, methods consistent with the present disclosure may include performing at least one or more operations of method 600 in a different order, simultaneously, quasi-simultaneously, or overlapping in time.
[0113] like Figure 6As shown, at operation 610, method 600 includes receiving a text prompt describing a scene and optionally an interaction of an avatar within the scene. Method 600 may also include extracting a text embedding from the text input.
[0114] At operation 620, method 600 includes generating a first pose based on a text prompt using a first model (e.g., a text-to-3D pose generation model). In some embodiments, the first pose is a human pose represented by SMPL parameters. Method 600 may also include mapping a text embedding extracted from the text prompt to a text embedding retrieved from a dataset (e.g., a CLIP dataset); selecting one or more second poses corresponding to the text embeddings retrieved from the dataset; and generating the first pose based on the selected one or more second poses.
[0115] At operation 630, method 600 includes redirecting the first pose to a target avatar body. The target avatar body may be, for example, a grayscale avatar-human representation that serves as a reference for personalized avatars. Method 600 may also include matching corresponding joints based on positions and orientations from the first pose and the target avatar body to redirect the first pose.
[0116] At operation 640, method 600 includes identifying a personal avatar configuration corresponding to a user. A user may be associated with a text prompt (eg, the user who submitted the text prompt). The personal avatar configuration may be predefined on the user's profile or account.
[0117] At operation 650 , method 600 includes transforming the target avatar body by applying the personal avatar configuration to the target avatar body.
[0118] At operation 660 , method 600 includes rendering an avatar or an image of an avatar using the second model, the avatar being based on at least the target avatar body, the personal avatar configuration, and the first pose, such that the avatar is in the first pose.
[0119] The method 600 may also include: using the second model to generate an image of a scene in the virtual environment based on the rendered avatar (or avatar image) and the text prompt. The scene may include the avatar interacting with objects in the scene according to the text prompt.
[0120] Method 600 may further include training a second model on a dataset comprising body poses and corresponding textual descriptions. The training may include identifying human bodies in images of the dataset, segmenting the human bodies, and extracting SMPL annotations from the segmented human bodies. Method 600 may further include performing conditional stable diffusion image inpainting by the second model to fill in an existing scene and objects in the existing scene by image expansion from an avatar to generate an image. The second model may be conditioned on the avatar and the textual prompt. In some embodiments, the second model is conditioned on the avatar and the textual prompt, as well as the interaction of the avatar with the objects in the scene, the downsampling value, and the time step.
[0121] Although Figure 6 Several example blocks of method 600 are shown, but in some implementations, method 600 may include more Figure 6 The multiple blocks depicted in the method may be more blocks, fewer blocks, different blocks, or blocks arranged in a different manner. Additionally or alternatively, two or more of the multiple blocks of the method may be executed in parallel.
[0122] Hardware Overview
[0123] Figure 7 It shows Figure 1 A block diagram of an exemplary computer system 700 is provided, which may be used to support a user and platform, and one or more of the methods described herein may be implemented. In some aspects, the computer system 700 may be implemented using hardware, or a combination of software and hardware, either in a dedicated server, integrated into another entity, or distributed across multiple entities. The computer system 700 may include a desktop computer, a laptop computer, a tablet computer, a phablet, a smartphone, a feature phone, or a server computer. The server computer may be located remotely in a data center or stored locally.
[0124] The computer system 700 (e.g., user device 110 and computing environment 122) includes a bus 708 or other communication mechanism for transmitting information, and a processor 702 coupled with the bus 708 for processing information. By way of example, the computer system 700 may be implemented using one or more processors 702. The processor 702 may be a general-purpose microprocessor, a microcontroller, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable logic device (PLD), a controller, a state machine, gate logic, multiple discrete hardware components, or any other suitable entity that can perform computations on information or other manipulations of information.
[0125] In addition to hardware, the computer system 700 may include code that creates an execution environment for the computer programs discussed, for example, code constituting the following stored in the included memory 704: processor firmware; a protocol stack; a database management system; an operating system; or a combination of one or more thereof, such as random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable PROM (EPROM), registers, a hard disk, a removable disk, a compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), or any other suitable storage device, coupled to the bus 708 for storing information and instructions to be executed by the processor 702. The processor 702 and the memory 704 may be supplemented by, or incorporated in, special purpose logic circuitry.
[0126] Instructions may be stored in memory 704 and implemented in one or more computer program products, such as one or more modules of a plurality of computer program instructions encoded on a computer-readable medium for execution by or to control the operation of the computer system 700, and according to any method well known to those skilled in the art, including but not limited to computer languages such as data-oriented languages (e.g., SQL, dBase), system languages (e.g., C, object-oriented programming languages extending C (Objective-C), C++, assembly), structured languages (e.g., Java, .NET), and application languages (e.g., PHP, Ruby, Perl, Python). The instructions may also be implemented in computer languages such as array languages, aspect-oriented languages, assembly languages, authoring languages, command line interface languages, compiled languages, concurrent languages, curly-bracket languages, data flow languages, data structured languages, declarative languages, esoteric languages, extension languages, fourth generation languages, functional languages, interactive mode languages, interpreted languages, iterative languages, list-based languages, little languages, logic-based languages, machine languages, macro languages, metaprogramming languages, multiparadigm languages, numerical analysis, non-English-based languages, class-based object-oriented languages, prototype-based object-oriented languages, off-side rule languages, procedural languages, reflective languages, rule-based languages, scripting languages, stack-based languages, synchronous languages, syntax handling languages, language), visual language, wirth language, and XML-based language, etc. The memory 704 may also be used to store temporary variables or other intermediate information during execution of instructions to be executed by the processor 702.
[0127] As discussed herein, computer programs do not necessarily correspond to files in a file system. A program may be stored in a portion of a file (e.g., one or more scripts stored in a markup language document) that holds other programs or data, in a single file dedicated to the program in question, or in multiple collaborative files (e.g., files storing one or more modules, subroutines, or partial codes). A computer program may be deployed to execute on one or more computers that are located at a site or distributed across multiple sites and interconnected through a communication network. The process flow and logic flow described in this specification may be performed by one or more programmable processors that execute one or more computer programs to perform a function by operating on input data and generating output.
[0128] The computer system 700 also includes a data storage device 706, such as a magnetic disk or optical disk, coupled to the bus 708 for storing information and instructions. The computer system 700 can be coupled to various devices via an input / output module 710. The input / output module 710 can be any input / output module. An exemplary input / output module 710 includes a data port (e.g., a USB port). The input / output module 710 is configured to connect to a communication module 712. An exemplary communication module 712 (e.g., the communication module 218) includes a network interface card, such as an Ethernet card and a modem. In certain aspects, the input / output module 710 is configured to connect to multiple devices, such as an input device 714 (e.g., the user device 110) and / or an output device 716 (e.g., the user device 110). Exemplary input devices 714 include a keyboard and a pointing device (e.g., a mouse or a trackball), through which a user can provide input to the computer system 700. Other types of input devices 714 may also be used to provide interaction with the user, such as tactile input devices, visual input devices, audio input devices, or brain-computer interface devices. For example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form including sound input, voice input, tactile input, or brainwave input. Exemplary output devices 716 include display devices for displaying information to the user, such as liquid crystal display (LCD) monitors.
[0129] According to one aspect of the present disclosure, the user device 110 and the platform 120 can be implemented using a computer system 700 in response to the processor 702 executing one or more sequences of one or more instructions contained in the memory 704. Such instructions can be read into the memory 704 from another machine-readable medium (e.g., a data storage device 706). The execution of the sequence of instructions contained in the main memory 704 causes the processor 702 to perform the various process steps described herein. One or more processors in a multi-processing arrangement can also be used to execute the sequence of instructions contained in the memory 704. In alternative aspects, hard-wired circuitry can be used instead of software instructions, or hard-wired circuitry can be used in combination with software instructions to implement various aspects of the present disclosure. Therefore, the aspects of the present disclosure are not limited to any particular combination of hardware circuitry and software.
[0130] Various aspects of the subject matter described in this specification can be implemented in a computing system that includes a back-end component (e.g., a data server), a middleware component (e.g., an application server), or a front-end component (e.g., a client computer with a graphical user interface or web browser through which a user can interact with an implementation of the subject matter described in this specification); or various aspects of the subject matter described in this specification can be implemented in any combination of one or more such back-end components, one or more such middleware components, or one or more such front-end components. The components of the system can be interconnected via any form or medium of digital data communication (e.g., a communication network). The communication network (e.g., network 150) can include, for example, any one or more of the following: a LAN, a WAN, and the Internet. Furthermore, the communication network can include, but is not limited to, any one or more of the following tool topologies: a bus network, a star network, a ring network, a mesh network, a star-bus network, or a tree or hierarchical network. The communication module can be, for example, a modem or an Ethernet card.
[0131] Computer system 700 can comprise client and server.Client and server are usually far away from each other, and usually interact through communication network.The relationship between client and server is to generate by means of the computer program that runs on respective computer and has client-server relationship between each other.Computer system 700 can be but not limited to desktop computer, laptop computer or tablet computer.Computer system 700 can also be embedded in another device, this another device is such as but not limited to mobile phone, personal digital assistant (PDA), mobile audio player, global positioning system (GPS) receiver, video game console and / or TV set-top box.
[0132] As used herein, the term "machine-readable storage medium" or "computer-readable medium" refers to any one or more media that participate in providing instructions to processor 702 for execution. Such media can take a variety of forms, including, but not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical or magnetic disks, such as data storage device 706. Volatile media include dynamic memory, such as memory 704. Transmission media include coaxial cables, copper wire, and optical fiber, including the wires that form bus 708. Common forms of machine-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, DVDs, any other optical media, punch cards, paper tape, any other physical medium with a pattern of holes, RAM, PROM, EPROM, FLASH EPROM, any other memory chip or cartridge, or any other form of media that can be read by a computer. A machine-readable storage medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a combination of substances that effect a machine-readable propagated signal, or a combination of one or more thereof.
[0133] To illustrate the interchangeability of hardware and software, a number of items (e.g., various illustrative blocks, modules, components, methods, operations, instructions, and algorithms) have been generally described in terms of their functionality. Whether this functionality is implemented as hardware, software, or a combination of hardware and software depends on the specific application and the design constraints imposed on the overall system. A skilled artisan may implement the described functionality in varying ways for each specific application.
[0134] As used herein, the phrase "at least one of" following a list of items, together with the terms "and" or "or" used to separate any of those items, modifies the list as a whole, rather than modifying each member of the list (e.g., each item). The phrase "at least one of" does not require selection of at least one item; rather, the phrase is meant to include at least one of any of the items, and / or at least one of any combination of the items, and / or at least one of each of the items. As an example, the phrase "at least one of A, B, and C" or "at least one of A, B, or C" each refers to: only A, only B, or only C; any combination of A, B, and C; and / or at least one of each of A, B, and C.
[0135] To the extent the terms "including," "having," and the like are used in this specification or claims, such terms are intended to be open-ended in a manner similar to the term "comprising" when used as a transitional word in a claim. The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments.
[0136] Unless otherwise specified, elements referred to in the singular are not intended to mean "one and only one", but rather "one or more". All structural and functional equivalents to the elements of the various configurations described throughout this disclosure that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be encompassed by the subject technology. In addition, nothing disclosed herein is intended to be dedicated to the public, regardless of whether such disclosure is explicitly recorded in the description above. No claim element shall be interpreted under the provisions of 35 U.S.C. § 112, paragraph 6, unless the element is expressly described using the phrase "means for..." or, in the case of a method claim, using the phrase "step for..."
[0137] Although this specification includes many details, these details should not be interpreted as limiting the scope of the content that may be claimed, but should be interpreted as a description of the specific implementation of the subject matter. Certain features described in the context of different embodiments in this specification may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. In addition, although features may be described above as working in certain combinations and even initially claimed as such, in some cases, one or more features from the claimed combination may be removed from the combination, and the claimed combination may be directed to a variation of a sub-combination or a sub-combination.
[0138] The subject matter of this specification has been described in terms of particular aspects, but other aspects may be implemented and are within the scope of the following claims. For example, although the operations are depicted in a particular order in the accompanying drawings, this should not be interpreted as requiring that the operations be performed in the particular order shown or in a continuous order, or that all of the illustrated operations be performed to achieve the desired result. The actions recited in the claims can be performed in a different order and still achieve the desired result. As an example, the processes depicted in the accompanying drawings do not necessarily require the particular order shown or the continuous order to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of the various system components in the various aspects described above should not be understood as requiring such separation in all aspects, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products. Other variations are within the scope of the following claims.
Claims
1. A computer-implemented method for generating a personalized avatar, the computer-implemented method being executed by at least one processor, the method comprising: Receive text alerts; generating a first gesture based on the text prompt using a first model; redirecting the first gesture to a target avatar body; identifying a predefined avatar configuration corresponding to the user based on the user's profile; transforming the target avatar body by applying the predefined avatar configuration to the target avatar body; as well as An avatar is rendered based on the target avatar body having the predefined avatar configuration using a second model, wherein the avatar is in the first pose.
2. The computer-implemented method of claim 1 , wherein: The text prompts describe a scene and the avatar's interactions within the scene.
3. The computer-implemented method of claim 1 or 2, and any one or more of the following: a) wherein the first posture is a 3D posture generated based on human joint and limb orientation and positioning parameters; or b) Among them, The first pose is a human body pose represented by SMPL parameters.
4. The computer-implemented method of any preceding claim, generating the first posture further comprising: extracting a first text embedding from the text prompt; mapping the first text embedding to a second text embedding retrieved from a dataset; selecting one or more second gestures corresponding to the second text embedding; as well as The first posture is determined based on the one or more second postures.
5. The computer-implemented method according to any preceding claim, further comprising: The first model is trained on a dataset comprising body poses and corresponding text descriptions, the training comprising: Identify human bodies in images of said dataset, Segmenting the human body from the image, and 3D Skinned Multi-Person Linear (SMPL) model annotations are extracted from the human body segmented from the image.
6. A computer-implemented method according to any preceding claim, wherein: The target avatar body is a grayscale avatar-human representation.
7. A computer-implemented method according to any preceding claim, wherein: The redirection includes matching corresponding joints from the first pose and the target avatar body in position and orientation.
8. The computer-implemented method of any preceding claim, further comprising: An image of a scene in a virtual environment is generated using the second model, the scene including the avatar interacting with objects in the scene, in which case, optionally, the second model performs conditional stable diffusion image inpainting to generate an image by image expansion from the avatar to fill the scene and objects in the scene, and the second model is conditioned on at least the avatar and the text prompt.
9. A system for generating a personalized avatar, the system comprising: one or more processors; as well as a memory storing instructions that, when executed by the one or more processors, cause the system to: receiving a textual prompt describing a scene and an interaction of the avatar within the scene; generating a first gesture based on the text prompt using a first model; redirecting the first gesture to a target avatar body; identifying a predefined avatar configuration corresponding to the user based on the user's profile; transforming the target avatar body by applying the predefined avatar configuration to the target avatar body; as well as An avatar is rendered based on the target avatar body having the predefined avatar configuration using a second model, wherein the avatar is in the first pose.
10. The system of claim 9, and any one or more of the following: a) wherein the first posture is a 3D posture generated based on human joint and limb orientation and positioning parameters; or b) Among them, The first pose is a human body pose represented by SMPL parameters.
11. The system of claim 9 or 10, and one or more of the following: a) wherein the one or more processors further execute instructions for: extracting a first text embedding from the text prompt; mapping the first text embedding to a second text embedding retrieved from a dataset; selecting one or more second gestures corresponding to the second text embedding; and determining the first gesture based on the one or more second gestures; or b) Among them, The one or more processors further execute instructions for: training the first model on a dataset comprising body poses and corresponding textual descriptions, the instructions causing the system to: Identify human bodies in images of said dataset, Segmenting the human body from the image, and 3D Skinned Multi-Person Linear (SMPL) model annotations are extracted from the human body segmented from the image.
12. The system according to any one of claims 9 to 11, wherein: The target avatar body is a grayscale avatar-human representation.
13. The system according to any one of claims 9 to 12, and any one of the following: a) wherein the one or more processors further execute instructions for: matching corresponding joints from the first pose and the target avatar body in position and orientation; or b) Among them, The one or more processors also execute instructions for: An image of a scene in a virtual environment is generated using the second model, the scene including the avatar interacting with objects in the scene.
14. The system according to any one of claims 9 to 13, wherein: The second model performs conditional stable diffusion image inpainting to generate an image by image expansion from the avatar to fill the scene and objects in the scene, and the second model is conditioned on at least the avatar and the text prompt.
15. A non-transitory computer-readable storage medium having instructions embodied thereon, the instructions being executable by one or more processors to perform a method for personalized avatar generation and causing the one or more processors to: receiving text input describing an avatar in a scene; generating a body gesture based on the text prompt; redirecting the body posture to a target avatar body; generating a personalized avatar based on a predefined avatar configuration applied to the target avatar body; and An image of the avatar in the scene is generated based on the personalized avatar, the image including the avatar in the body pose with the predefined avatar configuration.