Generating animatable characters using 3D representations
By combining Gaussian bodies, posture-driven primitives and implicit neural fields, using implicit mesh learning and signed distance functions, the problem of generating accurate and efficient animated three-dimensional virtual images in the existing technology is solved, realizing virtual images with higher details and realism, and optimizing the use of computing resources.
Patent Information
- Application Number
- CN202411618896.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-01
- Filing Date
- 2024-11-13
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to generate accurate and efficient animated three-dimensional avatars, especially in capturing complex geometric details and appearance nuances, while high computing resources are consumed, limiting the speed and efficiency of real-time animation or interactive simulation.
By combining Gaussian bodies with pose-driven primitives and implicit neural fields, we use implicit mesh learning and signed distance functions to generate complex, high-fidelity 3D character dynamic representations.
Animateable virtual images with higher detail and realism are achieved, the appearance and geometric accuracy are optimized, and high-speed rendering in real-time or near-real-time applications are promoted.
Smart Images

Figure CN119991885A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of and priority to U.S. Provisional Application No. 63 / 548,261, filed on November 13, 2023, the entire contents of which are incorporated herein by reference in their entirety. Background Art
[0003] Machine learning models, such as neural networks, can be used to represent objects. For example, these models can capture the interplay of shape, texture, and lighting to build digital representations that resemble physical objects. However, creating models that accurately represent three-dimensional objects is challenging due to the complexity of capturing the range of geometric details and appearance nuances of three-dimensional objects. In addition, rendering such models into visual formats often requires significant computational resources, which can limit the speed and efficiency of generating real-time or near-real-time animations or interactive simulations. In addition, these models fail to meet performance criteria when used to generalize across a variety of instances without overfitting to specific examples. Summary of the invention
[0004] Implementations of the present disclosure relate to generating various animatable avatars. Compared to traditional systems (e.g., those that rely heavily on manual modeling and grid-based frameworks), the systems and methods described herein can combine Gaussians with pose-driven primitives and implicit neural fields. This hybrid implementation achieves dynamic representation of complex, high-fidelity 3D characters by manipulating Gaussian parameters (e.g., position, scale, orientation, opacity, color) informed by text descriptions. For example, the system and method can use implicit neural fields to predict Gaussian properties, allowing detailed and accurate textures and geometries to be generated. In addition, by utilizing implicit grid learning based on signed distance functions (SDFs), the present disclosure provides improvements in the stability and efficiency of learning many Gaussians, while also improving the extraction and rendering of complex avatars or other character type details. This enables the system and method to produce animatable avatars or other character types with higher detail and realism, and optimize for appearance and geometric accuracy, thereby facilitating high-speed rendering used in real-time or near-real-time applications.
[0005] At least one implementation involves one or more processors. One or more processors may include one or more circuits that can be used to assign multiple first elements of a three-dimensional (3D) model of a subject to multiple locations on the surface of the subject in an initial posture. One or more circuits can assign multiple second elements to the multiple first elements, each of the multiple second elements having an opacity corresponding to the distance between the second element and the surface of the subject. One or more circuits can update the multiple second elements based at least on the target posture of the subject and one or more attributes of the subject to determine multiple updated second elements. One or more circuits can render a representation of the subject based at least on the multiple updated second elements.
[0006] In some implementations, one or more circuits are used to update the plurality of updated second elements based at least on the evaluation of the one or more objective functions and the representation of the subject. In some implementations, one or more circuits are used to determine the opacity of each of the plurality of second elements using a signed distance function to represent the distance between the second element and the surface of the subject.
[0007] In some implementations, at least one first element of the plurality of first elements includes a position parameter corresponding to a position of the plurality of positions to which the at least one element is assigned, a scale parameter indicating a scale of the at least one first element in a 3D reference system in which the subject is located, and an orientation parameter indicating an orientation of the subject relative to the 3D reference system. In some implementations, at least one second element of the plurality of second elements includes a 3D Gaussian splash defined in a local reference system of the corresponding at least one first element of the plurality of first elements.
[0008] In some implementations, one or more circuits are used to receive an indication of one or more attributes of a subject as at least one of text data, speech data, audio data, image data, or video data. In some implementations, one or more processors are used to regularize the plurality of second elements based at least on position data for each of the plurality of second elements. In some implementations, one or more processors are used to update the plurality of updated second elements based at least on a mask determined from a representation of the subject and an alpha rendering determined from the plurality of updated second elements. In some implementations, one or more circuits are used to generate a representation to include a textured mesh of the subject.
[0009] At least one implementation relates to a system including one or more processing units to perform operations. The one or more processing units may perform operations to assign multiple first elements of a three-dimensional (3D) model of a subject to multiple locations on a surface of the subject at an initial pose. The one or more processing units may perform operations to assign multiple second elements to the multiple first elements, each of the multiple second elements having an opacity corresponding to a distance between the second element and the surface of the subject. The one or more processing units may perform operations to update the multiple second elements based on at least a target pose of the subject and one or more attributes of the subject to determine multiple updated second elements. The one or more processing units may perform operations to render a representation of the subject based at least on the multiple updated second elements.
[0010] In some implementations, one or more processing units are used to update the plurality of updated second elements based at least on an evaluation of one or more objective functions and a representation of the subject. In some implementations, one or more processing units are used to determine the opacity of each of the plurality of second elements using a signed distance function to represent the distance between the second element and the surface of the subject. In some implementations, at least one of the plurality of first elements includes a position parameter corresponding to a position in a plurality of positions to which the at least one element is assigned, a scale parameter indicating a scale of the at least one first element in a 3D reference system in which the subject is located, and an orientation parameter indicating an orientation of the subject relative to the 3D reference system.
[0011] In some implementations, at least one of the plurality of second elements comprises a 3D Gaussian splash defined in a local reference frame of a corresponding at least one of the plurality of first elements. In some implementations, the one or more processing units are used to receive an indication of one or more attributes of a subject as at least one of text data, speech data, audio data, image data, or video data. In some implementations, the one or more processing units are used to regularize the plurality of second elements based at least on position data of each of the plurality of second elements.
[0012] In some implementations, one or more processing units are used to update the plurality of updated second elements based at least on a mask determined from a representation of the subject and an alpha rendering determined from the plurality of updated second elements, and one or more of the processing units are used to generate the representation to include a textured mesh of the subject.
[0013] At least one implementation relates to a method. The method may include: assigning, by one or more processors, a plurality of first elements of a three-dimensional (3D) model of a subject to a plurality of locations on a surface of the subject in an initial pose. The method may include: assigning, by one or more processors, a plurality of second elements to the plurality of first elements, each of the plurality of second elements having an opacity corresponding to a distance between the second element and the surface of the subject. The method may include: updating, by one or more processors, the plurality of second elements based at least on a target pose of the subject and one or more attributes of the subject to determine a plurality of updated second elements. The method may include: rendering, by one or more processors, a representation of the subject based at least on the plurality of updated second elements.
[0014] In some implementations, updating the plurality of updated second elements is based at least on an evaluation of one or more objective functions and a representation of the subject, and wherein determining the opacity of each of the plurality of second elements includes representing a distance between the second element and a surface of the subject using a signed distance function.
[0015] The processors, systems and / or methods described herein may be implemented by or included in at least one of the following: a system for generating synthetic data; a system for performing simulation operations; a system for performing conversational AI operations; a system for performing collaborative content creation of 3D assets; a system including one or more language models (e.g., a large language model (LLM)); a system for generating or presenting virtual reality (VR) content, augmented reality (AR) content, and / or mixed reality (MR) content; a system for performing digital twin operations; a system for performing light transport simulations; a system for performing deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system associated with an autonomous or semi-autonomous machine (e.g., an in-vehicle infotainment system); a system including one or more virtual machines (VMs); a system implemented at least in part in a data center; or a system implemented at least in part using cloud computing resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The present system and method for a machine learning model for animatable object generation are described in detail below with reference to the accompanying drawings, wherein:
[0017] Figure 1 is a block diagram of an example system for generating animatable objects according to some embodiments of the present disclosure;
[0018] Figure 2 is a block diagram of an example system for generating animatable objects according to some embodiments of the present disclosure;
[0019] Figure 3is a flowchart of an example of a method for generating an animatable object according to some embodiments of the present disclosure;
[0020] Figure 4 is the use according to some embodiments of the present disclosure (e.g., Figure 1 to Figure 2 ) An example illustration of an object rendering of a 3D object generation system;
[0021] Figure 5 is used according to some embodiments of the present disclosure (e.g., Figure 1 to Figure 2 ) An example illustration of an object rendering of a 3D object generation system;
[0022] Figure 6 is used according to some embodiments of the present disclosure (e.g., Figure 1 to Figure 2 ) An example illustration of an object rendering of a 3D object generation system;
[0023] Figure 7 is an example illustration of a flawed method of generating an animatable avatar compared to an animatable 3D Gaussian avatar generated according to some embodiments of the present disclosure;
[0024] Figure 8 is a block diagram of an exemplary content streaming system suitable for implementing some embodiments of the present disclosure;
[0025] Fig. 9 is a block diagram of an exemplary computing device suitable for implementing some embodiments of the present disclosure; and
[0026] Fig.10 is a block diagram of an exemplary data center suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0027] The present disclosure relates to systems and methods for generating animatable characters using three-dimensional (3D) representations, such as primitive-based 3D Gaussian volumes (e.g., Gaussian splats) representations. For example, systems and methods according to the present disclosure may allow text and other input to inform the properties of a subject, such as an animatable avatar, which may be used to configure and / or optimize a 3D model of a subject using a 3D Gaussian volume.
[0028] Various 3D modeling techniques, such as mesh representations and neural radiance fields (NeRFs), can be used to generate a 3D representation of a subject, as well as allow the subject to deform (e.g., move). However, mesh representations can be rendered with low quality due to limitations in the underlying geometry of the mesh. NeRFs are computationally expensive, especially when rendering high-resolution images, and are therefore unlikely to successfully generate fine geometric details (e.g., loose clothing). Furthermore, various such techniques are unable to correctly represent poses that are beyond the underlying distribution of the representation, such as unseen body poses and complex body geometry.
[0029] Systems and methods according to the present disclosure can achieve more realistic and / or configurable subject animation by using a 3D model of a subject including 3D Gaussian bodies assigned to primitives (e.g., primitives defined using a skeleton-based parametric model). For example, multiple first elements (e.g., primitives) can be assigned to the surface of the subject. Multiple second elements (e.g., 3D Gaussian bodies) can be assigned to the first elements, such as assigning multiple second elements to each first element. The 3D Gaussian body can represent features of the subject and / or scene using color, opacity, scale, and rotation. Using primitives for avatars or other character types can achieve more natural animation of subject movement (which can be challenging for Gaussian bodies), and using Gaussian bodies can achieve efficient modeling, including fine details.
[0030] In some implementations, a field (e.g., a neural implicit field) is used to predict properties of a Gaussian. This can be performed for properties such as color, rotation, scaling, and / or opacity. This can allow for more stable Gaussian training, for example to mitigate noisy geometry and / or rendering. Properties can be predicted based on inputs such as text, speech, audio, image, and / or video data. In some implementations, the geometry of the Gaussian is determined based on the distance between the Gaussian and the surface of the subject. For example, the opacity of the Gaussian can be determined based on a signed distance field (SDF) function corresponding to the distance to the surface. This can address the transparent point cloud properties of 3D Gaussians, which otherwise may result in holes or other unrealistic features in the subject.
[0031] The 3D model (e.g., a 3D Gaussian volume) can be used to render an image of a subject in a variety of ways. For example, a textured mesh can be extracted from the 3D model and can be rendered quickly to meet performance criteria, such as for animation. Various objectives can be used to facilitate realistic generation of the 3D model, such as to optimize the 3D model. The objectives can include one or more fractional distillation sampling (SDS) objectives to update and / or optimize parameters of the 3D model, such as the shape, consistency, and / or color of the 3D model. The objectives can include a regularization objective to regularize the geometry of the avatar, and can include an alpha loss objective to match a mask rendered from an extracted mesh to an alpha rendering of the 3D model.
[0032] The systems and methods described herein may be used for a variety of purposes, such as, but not limited to, synthetic data generation, machine control, machine motion, machine driving, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or participant simulation and / or digital twins, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing, and / or any other suitable application.
[0033] The disclosed embodiments may be included in a variety of different systems, such as systems for performing synthetic data generation operations, automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, boating systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems including one or more virtual machines (VMs), systems implementing one or more language models (e.g., large language models (LLMs) and / or visual language models (VLMs)), systems implemented at least in part in a data center, systems for performing conversational AI operations, systems for performing light transport simulations, systems for performing collaborative content creation of 3D assets, systems implemented at least in part using cloud computing resources, and / or other types of systems.
[0034] refer to Figure 1 , Figure 11 is an example computing environment including system 100 according to some embodiments of the present disclosure. It should be understood that this arrangement and other arrangements described herein are only proposed as examples. As a supplement or alternative to the arrangements and elements shown, other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) can also be used, and some elements can be completely omitted. In addition, many elements described herein are functional entities that can be implemented as discrete or distributed components or in combination with other components, and can be implemented in any suitable combination and position. The various functions performed by the entities described herein can be performed by hardware, firmware and / or software. For example, various functions can be performed by a processor that executes instructions stored in a memory. System 100 may include any function, model (e.g., machine learning model), operation, routine, logic or instruction to perform functions such as transformer 112, object layer model 114, texture model 116 and / or renderer 118 described herein, such as configuring a machine learning model to generate an animated character using a 3D representation.
[0035] The system 100 may include or be coupled to one or more data sources 104. The data source 104 may include any of a variety of databases, data sets, or data repositories. The data source 104 may include data for configuring any of a variety of machine learning models (e.g., object layer modeler 114; texture model 116). The one or more data sources 104 may be maintained by one or more entities, which may be entities that maintain the system 100, or may be separate from the entities that maintain the system 100. In some implementations, the system 100 uses data from different data sets, such as by performing at least a first configuration (e.g., updating or training) of the models 114 and 116 using data from a first data source 104, and performing at least a second configuration of the models 114 and 116 using training data elements from a second data source 104. For example, the first data source 104 may include publicly available data, while the second data source 104 may include domain-specific data (access to which may be restricted compared to the data of the first data source 104). Image data 106 and video data 108 may include data from any suitable image or video dataset, including labeled and / or unlabeled image or video data. In some examples, data source 104 includes data from a large-scale image or video dataset (e.g., ImageNet) available from a variety of sources and services.
[0036] The data source 104 may include, but is not limited to, image data 106 and video data 108, such as any one or more of text, speech, audio, image, and / or video data. The system 100 may perform various pre-processing operations on the data, such as filtering, normalizing, compressing, decompressing, enlarging or reducing, cropping, and / or converting to grayscale (e.g., from image and / or video data). The images (e.g., including videos) of the image data 106 and the video data 108 may correspond to one or more views of a scene captured by an image or video capture device (e.g., a camera), or a computationally generated image, such as a simulated or virtual image or video (e.g., including by modifying an image from an image capture device). The images may each include a plurality of pixels, such as pixels arranged in rows and columns. The image may include image data assigned to one or more pixels of the image, such as color, brightness, contrast, intensity, depth (e.g., for a three-dimensional (3D) image), or various combinations thereof. Video data 108 may include video and / or video data structured into multiple frames (e.g., image frames, video frames), such as in the form of a sequence of frames, where each frame is assigned a time index (e.g., time step, time point) and has image data assigned to one or more pixels of the image.
[0037] In some implementations, the image data 106 and / or the video data 108 include camera pose information. The camera pose information may indicate a viewpoint by which the data is represented. For example, the camera pose information may indicate at least one of a position or an orientation of a camera (e.g., a real or virtual camera) by which the image data 106 and / or the video data 108 is captured or represented.
[0038] The system 100 can train, update, or configure one or more models (e.g., machine learning models) of the modeler system 110. The machine learning models (e.g., object layer model 114 and texture model 116) can include machine learning models or other models that can generate target outputs based on various types of inputs. The machine learning model can include one or more neural networks. The neural network can include an input layer, an output layer, and / or one or more intermediate layers (e.g., hidden layers), each of which can have respective nodes. The system 100 can train / update the neural network by modifying or updating one or more parameters (e.g., weights and / or biases) of respective nodes of the neural network in response to an evaluation of a candidate output of the neural network.
[0039] The machine learning models (e.g., the object layer model 114 and the texture model 116 of the modeler system 110) can be or include various neural network models, including models that can effectively operate or generate data (e.g., objects such as avatars, people, animals, characters, animations, etc.), including but not limited to image data, video data, text data, speech data, audio data, 3D model data, CAD data, or various combinations thereof. The machine learning model can include one or more transformers, recurrent neural networks (RNNs), long short-term memory (LSTM) models, other network types, or various combinations thereof. The machine learning model can include a generative model, such as a generative adversarial network (GAN), a Markov decision process, a variational autoencoder (VAE), a Bayesian network, an autoregressive model, an autoregressive encoder model (e.g., a model that includes an encoder to generate a latent representation (e.g., in an embedding space) of a model input (e.g., a representation of a different dimension than the input) and / or a decoder to generate an output representing the input from the latent representation), or various combinations thereof.
[0040] like Figure 1 As shown, the modeler system 110 can receive input 120 and can generate output 130 in response to the input 120. The input 120 can include any one or more text, voice, audio, image, other sensor modality data (e.g., LiDAR, RADAR, ultrasound, depth, etc. data), 3D asset data, CAD data, and / or video input data, and the modeler system 110 can generate output based on at least these data, such as generating 2D images, 3D images, and / or video output. For example, the input 120 can represent text information, such as "a person wearing a red hoodie and blue jeans", in response to which the modeler system 110 can generate an animatable object (e.g., a character or avatar) using a 3D representation.
[0041] The modeler system 110 can learn the positions of primitives. For example, the set of primitives can be geometric shapes that are located on the surface of the object in a configuration. The modeler system 110 can also learn the properties of Gaussians inside each primitive to represent the overall shape and color of the object. For example, the Gaussian can be a function applied within each primitive to model details such as contours and textures of object features such as shape, color, opacity, and rotation. The modeler system 110 can include a transformer 112, which can assign multiple primitives (e.g., first elements) of a three-dimensional (3D) model of a subject to multiple positions on the surface of the subject in an initial pose (or rest pose). For example, the transformer 112 can generate a set of basic geometric primitives, such as a cube, from a predefined rest pose (sometimes referred to as a "rest position" or "initial pose"). In some implementations, each primitive can have one or more attributes, such as position, rotation (e.g., along the X-axis, Y-axis, and Z-axis) and scale, thereby allowing resizing to fit the underlying template grid. For example, a template mesh can be used to mirror the outline and topology of a human body (or another object) so that primitives can fit closely to the body, thereby capturing the pose of the human body in a static or stationary state.
[0042] In some implementations, the transformer 112 may determine the placement of primitives so that each primitive is overlaid on the surface of the object. This may allow the modeler 112 to capture, at least with an initial level of accuracy, subtle geometric and visual features of the object, including but not limited to the shape, clothing, and hair of a person in a static pose. The accuracy with which the transformer 112 places primitives may affect the object layer model 114 and the texture model 116 to represent the object with greater accuracy, e.g., when the avatar's arms move, the corresponding primitives are configured by the transformer 112 to represent the motion.
[0043] In some implementations, the transformer 112 may assign multiple 3D Gaussian volumes (e.g., the second element) to multiple primitives (e.g., the first element). For example, within each geometric primitive generated by the transformer 112, a series of Gaussian distributions may be defined (e.g., Figure 2, which are shown as points within a cube in the figure). The Gaussian bodies can be characterized by their position, orientation (rotation), and scale within the primitive, as well as the Gaussian body fixed size, color, and shape. The multi-layered implementation allows the modeler system 110 to generate realistic and / or improved body configurations. The transformation of the Gaussian bodies from the local coordinate system within the primitive to the global environment (where the Gaussian bodies are aligned with the actual surface of the object) is performed by the transformer 112 using a local to world position transformation. This transformation can make each Gaussian body consistent with the surface contours of the underlying object, thereby capturing the appearance (e.g., of the virtual image) and any specific details, such as (e.g., clothing wrinkles or hair (of the virtual image). For example, the transformer 112 can generate a detailed representation of the geometry and surface features of the object. Reference will be made below to Figure 2 Additional information regarding the converter 112 is described in more detail.
[0044] The object layer model 114 of the modeler system 110 can be a first pre-trained neural network that receives as input the transformation of each Gaussian volume. The object layer model 114 can provide a geometric foundation by generating a signed distance field (SDF) that depicts the underlying geometry of the avatar and the contained 3D Gaussian volumes. The SDF can represent a scalar field where the value of each point represents its shortest distance to the surface of the avatar, with negative values representing points inside the geometry and positive values representing points outside. For example, the opacity of each Gaussian volume comes from the SDF, and distance affects the transparency to create a realistic rendering of the avatar. This relationship indicates that Gaussians that are aligned with the surface of the avatar contribute more to the visual output, while those that are farther away contribute less. In addition, the object layer model 114 can convert the SDF into a mesh representation of the avatar using differentiable marching tetrahedrons (DMTets). The mesh can form a visual structure on which textures and other surface details can be applied. Reference will be made below to Figure 2 Additional information regarding the object layer model 114 is described in more detail.
[0045] The texture model 116 of the modeler system 110 can be a second pre-trained neural network that receives the transformation of each Gaussian volume as input. The texture model 116 can characterize the visual aspects of the avatar, using the neural implicit field to assign color, rotation, scale, and opacity to each Gaussian volume. For example, the texture model 116 can use the transformed Gaussian volume position within the global coordinate system to assign visual attributes that enhance the fidelity or realism of the avatar. By using the neural implicit field and Query specification location The texture model 116 can ensure that the characteristics of the Gaussian body change consistently and smoothly across the surface of the avatar. Figure 2 Additional information regarding texture model 116 is described in more detail.
[0046] The renderer 118 of the modeler system 110 can apply the updated Gaussian positions and attributes via Gaussian splashing to produce a visual representation of the avatar (or a visual representation of an object). In some implementations, Gaussian splashing can include a process of projecting the color and opacity of a Gaussian volume onto an image plane, composing a composite image that captures the target pose with fine motion and surface detail. The renderer 118 aggregates or compiles the contributions of the individual Gaussian volumes into a unified visual field that accurately represents the avatar in the desired pose. For example, the renderer 118 creates an RGB image I and an alpha image I based on the updated positions and attributes of the 3D Gaussian volume. α .
[0047] In more detail, the renderer 118 projects the color information of the Gaussian volume onto the image plane to generate an RGB image I, thereby capturing the appearance of the avatar (e.g., as indicated by the text prompt). At the same time, the renderer 118 calculates the alpha image I α , encoding the transparency level of the Gaussian, which helps blend the avatar with various backgrounds and provide visual continuity in the rendered scene. The combination of the RGB image and the alpha image contributes to the realism of the avatar by allowing subtle visual effects (such as soft transitions between the edges of the avatar and overlapping Gaussians). In order to maintain the spatial integrity of the representation and / or prevent the Gaussian from deviating from its specified position relative to the primitive, the renderer 118 can apply a local position regularization loss, This can constrain the Gaussian volume to be within a certain radius of its associated primitive origin. This constraint verifies that the Gaussian volume helps form a coherent visual field when synthesizing a composite image via sputtering. The resulting visual output is then used to calculate the fractional distillation sampling (SDS) loss, as defined in Equation 2 (below). Through the sputtering process, the renderer 118 integrates the various attributes and positions of the Gaussian volume to form an interconnected and continuous visual representation of the virtual image. The following reference will be made to Figure 2 Additional information about the renderer 118 is described in more detail.
[0048] Further references Figure 1, the system 100 may receive one or more inputs 120. The inputs 120 may indicate one or more features of the 3D avatar representation for the system 100 to generate and / or animate. The inputs 120 may be received from one or more user input devices that may be coupled to the system 100. The inputs 120 may include any of a variety of data formats, including but not limited to text, voice, audio, images, CAD data, digital asset data, and / or video data, indicating instructions corresponding to features of the 3D avatar for the system 100 to generate and / or animate. The inputs 120 may indicate, for example but not limited to, information about features of the avatar and / or object to be represented by the 3D representation. In some implementations, the system 100 presents a prompt requesting one or more attributes or features via a user interface and receives the inputs 120 from the user interface. The inputs 120 may be received as semantic information (e.g., text, voice, speech, etc.) and / or image information (e.g., input representing pixels indicating an area in a scene).
[0049] Reference now Figure 2 , Figure 2 An example computing environment including a 3D object generation architecture 200 is depicted according to some embodiments of the present disclosure. The 3D object generation architecture 200 can be used to implement 3D Gaussian volume-based animatable virtual image generation with high-quality output. It should be understood that this arrangement and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used in addition to or in place of the arrangements and elements shown, and some elements may be omitted entirely. In addition, many of the elements described herein are functional entities that may be implemented as discrete or distributed components or combined with other components, and may be implemented in any suitable combination and location. The various functions performed by the entities described herein may be performed by hardware, firmware, and / or software. For example, the various functions may be performed by a processor executing instructions stored in a memory. The 3D object generation architecture 200 may include any functions, models (e.g., machine learning models), operations, routines, logic, or instructions to perform functions such as configuring, deploying, updating, and / or generating outputs of machine learning models, including Figure 1 The object layer model 114 and the texture model 116 are as described herein.
[0050] refer to Figure 1 ,refer to Figure 2, boxes 202-214 in, describe a GAvatar implementation that generates an animatable avatar based on a 3D Gaussian volume given a textual prompt. For example, the GAvatar implementation includes a primitive-based implicit 3D Gaussian volume representation that is used to allow animation of the avatar and uses a high-variance SDS loss to stabilize and amortize the learning of a large number of Gaussians. In addition, the GAvatar implementation uses an SDF to represent the underlying geometry of the 3D Gaussian volume, which allows the extraction of a high-quality textured mesh and regularizes the geometry of the avatar.
[0051] Referring to the primitive-based implicit 3D Gaussian volume representation, the GAvatar implementation can leverage this framework to construct the spatial distribution and orientation of Gaussian functions that are intrinsically associated with primitives attached to the avatar mesh. This approach ensures that each primitive adheres to the avatar's underlying geometry, determined by the rest pose and subsequent pose transformations. The use of 3D Gaussian volumes allows for more fine-grained control over the avatar's deformation, allowing for higher degrees of animation freedom without having to maintain continuity and smoothness of the avatar's motion. The SDS loss can be run in this context to optimize the avatar's parameters, accurately refining the model to account for subtle differences introduced by the text-to-image diffusion process.
[0052] With the incorporation of the reference SDF, the accuracy of defining the geometry of the avatar is significantly improved. The SDF can act as a scalar field, assigning a distance value to each point in space relative to the avatar's surface, with a sign indicating whether the point is inside or outside the geometry. This level of geometry definition can provide for the extraction of high-quality meshes. When integrated into the GAvatar implementation, the SDF can allow for high-resolution mesh generation and also aids in the regularization of the avatar's shape. By aligning a Gaussian distribution with the SDF, the model can achieve integration between abstract mathematical representations and tangible animation geometry, allowing the avatar's surface to be accurately depicted based on the desired text-driven animation.
[0053] In addition, the introduction of Gaussian sputtering as a tool for 3D scene reconstruction can improve efficiency and adaptability because it adopts a probabilistic approach to rendering, but its direct application to dynamic human (or animal) avatars or dynamic object generation introduces specific complexities (e.g., animation and training stability challenges). For example, the GAvatar implementation specifically addresses how to transform a Gaussian volume defined in the world coordinate system together with a deformable avatar and how to learn a Gaussian volume with consistent properties (e.g., color, rotation, scale, etc.) within a local neighborhood. In other words, Figure 2 The systems and methods described in the disclosure provide an improved architecture for dynamic deformation of primitives using methods that ensure attribute stability and spatial coherence, effectively facilitating realistic animation of virtual avatars in complex 3D scenes.
[0054] Gaussians that deform with avatars require a framework to ensure they remain consistent with changing poses, and simultaneously learning Gaussian properties that exhibit spatial consistency is critical to avoid unstable visual artifacts. The innovation of primitive-based implicit Gaussian representations proposes a twofold solution: it provides a consistent distribution of properties over the avatar surface and provides a stable reference frame for the Gaussians, thereby maintaining the structural coherence of the model through the spectrum of human motion and animation.
[0055] Still refer to Figure 1 to Figure 2 , the modeler system 110 can output a 3D model of a subject through a series of operations, first assigning a plurality of first elements (primitives) to the surface of a subject in an initial or static pose. These primitives can be positioned to correspond to the surface anatomical structure of the subject. The modeler system 110 can then assign a series of second elements (3D Gaussian volumes) to each primitive. Each Gaussian volume can include an opacity value related to its proximity to the surface of the subject. For example, a closer Gaussian has a lower transparency, while a farther Gaussian has a higher transparency. For example, a signed distance function (SDF) can be used to create a transparent effect that will / can reflect the real-world observation of the subject.
[0056] In response to determining the primitives and their respective Gaussian volumes of the subject in the static pose, the modeler system 110 can continue to update the pose. For example, the modeler system 110 can update the second element based on the subject's target pose and attributes (e.g., texture, color, other features represented by the neural field). In some implementations, the update is a targeted optimization that adjusts the position, rotation, scale, and opacity of the Gaussian volume to match the new target pose. In some implementations, these updated attributes can be derived from a text prompt.
[0057] In response to determining the primitives and Gaussian volumes of the subject in the target pose, the modeler system 110 can render the subject. In some implementations, the modeler system 110 can sample or test the fidelity of the model based on the target view or text prompt indicating the pose adjustment. For example, the fractional distillation sampling (SDS) framework can be used to optimize the arrangement and properties of the Gaussian volume. The modeler system 110 can use the rendered 3D subject in combination with the mask determined from the representation of the subject and the alpha rendering obtained from the Gaussian volume to further refine the model. Through an interactive process, the modeler system 110 can be enhanced to create accurate, high-quality renderings of the subject in various poses and appearances specified by the input prompt.
[0058] At block 202, Figure 1The modeler system 110 may perform Gaussian volume attribute calculations at a static pose of an animatable object (e.g., an animatable human avatar). First, at blocks 202 and 204, the modeler system 110 may perform primitive formulation, wherein an object (e.g., a human or animal body) is represented by a set of primitives attached to its surface. Figure 1 The transformer 112 can generate the primitive V k , each primitive can be composed of a position (P ) of an object (e.g., avatar, person, animal) at rest or in a target pose (θ) k ), Rotation (R k ) and ratio (S k ) is characterized and is shown as {P k , R k , S k Primitive-based 3D representations can represent 3D scenes by a set of primitives (such as cubes, points, or nerflets). Although reference Figure 2 Cube primitives are shown and described, but points or nerflets can be used to represent sets of primitives. For example, a set of K cube primitives {V1, ..., V k} can be attached to the SMPL-X grid where θ and β are SMPL-X pose and shape parameters, and LBS is the linear blending skinning function. Each primitive V k = {P k ,R k , S k} can be determined by its location Each axis ratio and orientation R k ∈SO(3)). The primitive parameters can be generated by (Formula 1):
[0059]
[0060]
[0061] The modeler system 110 may first determine (block 202) the grid-based primitive initialization Then a posture-dependent correction δP is applied (block 204) ω (θ), δR ω (θ),δS ω (θ), these corrections are represented by a neural network with parameters ω. The mesh-based initialization can then be determined by placing primitives on a 2D grid in the mesh’s uv-texture space and generating primitives at 3D locations on the mesh’s surface points corresponding to the uv coordinates.
[0062] In addition, the modeler system 110 can use fractional distillation sampling (SDS) to perform parameter η optimization of the 3D model g using a pre-trained text-to-image diffusion model. In some implementations, given a text cue y and a noise prediction of the diffusion model SDS works by minimizing the noise ∈ added to the rendered image I = g(η) and the noise predicted by the diffusion model The difference between them is used to optimize the model parameter η:
[0063]
[0064] where g(η) represents the differentiable rendering process of the 3D model, t is the noise level, and I t is a noisy image, is a weighting function. In some implementations, the weighting function can be designed to adjust the impact of different noise levels on the gradient calculation, making the optimization process sensitive to certain features in the image at various stages of rendering. In addition, the SDS optimization can refine the parameters of the 3D model through interactive back-propagation to keep the model quality and details during training consistent with the textual prompt y.
[0065] At block 202, a grid-based primitive initialization may be used as a reference configuration of the object without any pose-induced deformation. Represents, corresponding to the grid , where the pose θ and shape β parameters are at their default settings (e.g., providing a basis for applying further transformations). In addition, a set of cube primitives {V1, ..., V k Primitives may be initially aligned with the surface of a SMPL-X mesh by placing them on a 2D grid within the mesh's UV texture space. In some implementations, primitives are then generated at corresponding 2D locations on the mesh surface using the UV coordinates, thereby establishing their initial position at rest pose.
[0066] At block 204, the posture-dependent correction δP ω (θ), δR ω (θ),δS ω (θ)-represents the modification required to transform the primitive from a rest pose to a target pose. Corrections can account for changes that occur as a result of an object moving from a neutral rest position to a specific target pose. Primitive position, orientation, and scale can be altered to conform to the new pose. For example, primitives can be adjusted in real time or near real time based on pose parameters θ, using a neural network parameterized by ω to accurately capture the deformed shape of the object. After establishing the initial positions of the primitives, the neural network can apply pose-dependent corrections to these primitives to match the target pose, allowing the model to animate the rest position to a range of poses determined by the pose parameters θ.
[0067] After initialization and application of pose-dependent corrections, the modeler system 110 can further refine the representation of the animatable object by determining the properties of the Gaussian volumes contained within each primitive. In some implementations, determining the properties of the Gaussian volumes can include calculating Gaussian parameters that best represent the local surface characteristics of the mesh at the corresponding position of the primitive. The properties are modified to capture the texture normals and curvature details of the object. In some implementations, the properties are calculated using pose-dependent deformations applied to the primitives, using the underlying SMPL-X mesh as a reference to generate a high-fidelity, animatable 3D object.
[0068] For example, within each primitive, the modeler system 110 may define a set of N k 3D Gaussian volumes, each with a specific position established in the local coordinate system of the primitive Rotation and proportion Since primitives naturally deform according to the posture and shape of humans (or objects), the modeler system 110 can transform a set of 3D Gaussian Append to each primitive V k = {P k , R k , S k} and deform them along with the primitives. For example, each Gaussian (e.g., Gaussian 206) can be identified by its position in the local coordinates of the primitive Rotation and zoom And its color characteristics and opacity Definitions. As shown, Gaussians 206 are depicted in the form of cubes (primitives), where points of different colors, rotations, and scales represent individual 3D Gaussian volumes within the primitives, where each Gaussian may have unique properties that will / can help model the surface texture and shape of the object.
[0069] Additionally, in block 208, the local-to-world position transformation model may transform the Gaussian to its canonical position in world coordinates The and Can be defined as (Formula 3-5):
[0070]
[0071]
[0072]
[0073] In some implementations, this can be achieved by applying a global transformation corresponding to the primitives, thereby transforming the Gaussians from local position references within each primitive to a global environment consistent with the overall spatial orientation and scale of the object. This primitive-based Gaussian representation can naturally balance constraints and flexibility. This approach can provide an improvement over existing representation methods because it can provide greater flexibility than native primitive representations because it can allow primitives to deform outside of the cube by equipping the primitives with Gaussians. Therefore, by using Gaussian bodies, each primitive can adjust its shape more dynamically than if it were just a rigid cube, allowing more complex and subtle deformations. At the same time, the Gaussian bodies within each primitive share the motion of the primitive and are more constrained during animation. Therefore, when the Gaussian bodies are bound to their respective primitives (e.g., their movement is controlled and predictable during animation), it provides a balance between flexibility and constraints for the avatar animation system.
[0074] Referring to blocks 202, 204, and 208, the process includes transforming the avatar from a resting pose to a target pose based on manipulating primitives and their contained Gaussian volumes. The transformation is guided by a textual prompt that specifies a desired action or state of an object (e.g., an avatar) that affects the application of pose-related corrections and subsequent deformations. For example, if a textual prompt describes the avatar as "waving left hand," the modeler system 110 can interpret this to determine the necessary adjustments to the primitives and Gaussian volumes to achieve a left-hand wave from a resting left hand.
[0075] At block 202, a baseline configuration (e.g., position, orientation, scale, opacity) of primitives is initially established on the mesh of the avatar (or object) in a rest pose (or initial pose) so that the default pose and shape parameters of the mesh can be used as a reference during local-to-world rotation and scaling. Based on the text prompt, block 202 can ensure that the initial state of the avatar is neutral, thereby allowing a starting point to be provided for any pose transitions indicated by the prompt.
[0076] At block 204, the posture-dependent correction δP ω (θ), δR ω (θ),δS ω (θ)-is introduced to adjust the primitive from its initial resting pose to a target pose. The adjustment may be influenced by a textual prompt, where the modeler system 110 may attempt to mimic the described action or gesture by changing the geometry of the object accordingly. The correction may dynamically change the position, orientation, and / or scale of the primitive based on the pose parameter θ of the avatar to accommodate the deformation caused by the specific target pose.
[0077] At block 208, the local-to-world transformation model uses the outputs of blocks 202 (rest pose primitives) and 204 (target pose primitives) to perform deformations of primitives. For example, the transformation may align the pose of the avatar with the rules of the text prompt, attempting to accurately reflect the desired action or emotional state of the avatar (or object). For example, the modeler system 110 may apply a global transformation to the primitives, converting a Gaussian from local coordinates within each primitive to a global environment that reflects the overall spatial orientation of the avatar and can be scaled in the target pose. and Formulas 3-5 refine how the position, scale, and rotation of each Gaussian body adapts to the target pose, maintaining its representation as the avatar transitions from a rest pose to the target pose.
[0078] The local-to-world transformation at box 208 generates an output that includes a plurality of deformed primitives, no longer restricted to cubic form, and adjusted according to their corresponding Gaussians 210 (the points within the primitive that are shaded, rotated, and scaled). The contours of each adjusted primitive are matched to the dynamic pose structure of the avatar and displayed on the surface of the avatar. For example, the surface of the avatar can be deformed to capture a specified action, emotional state, clothing, and / or object derived from a text prompt.
[0079] Gaussian splatting at block 212 includes the modeler system 110 using the updated positions and properties of the Gaussian volumes from the object layer model 114 and the texture model 116 to render a visual representation of the avatar. The modeler system 110 may project the color and opacity of each Gaussian volume onto the image plane (RGB image I) to synthesize the final composite image. For example, splatting utilizes the transformed Gaussian parameters and and the opacity value derived from the SDF value calculated at block 208 The renderer 118 of the modeler system 110 may perform a splashing algorithm that aggregates the contributions of the individual Gaussian volumes to form an interconnected and continuous visual field, thereby producing an object 214 that embodies the target pose, with articulated motion and surface details. In some implementations, at block 216, the visual output from the Gaussian splashing may then be used to calculate the SDS loss L SDS , thereby allowing the modeler system 110 to refine the Gaussian volume properties to achieve consistency and alignment with the target appearance and pose.
[0080] Furthermore, after obtaining the position and properties of the 3D Gaussian volume, the renderer 118 may perform Gaussian sputtering to render the RGB image I and also the alpha image I α. For example, the RGB image captures color information projected from the Gaussian, while the alpha image represents transparency information, indicating how the visual elements of each Gaussian should blend with the background and with each other. In addition, the alpha image will / can directly affect the visual realism by allowing soft transitions and subtle visibility between the foreground avatar (or object) and its environment. The RGB image I can then be used to train the SDS loss defined in the objective formula 2. In order to prevent the Gaussian from deviating from the primitive, the renderer 118 also utilizes a local position regularization loss It constrains the Gaussian to be close to the origin of the relevant primitive.
[0081] The representation generated by Gaussian sputtering combines surface mesh details with a mix of volume primitives, laying a structural foundation that is technically capable of capturing a wide range of shapes, including shapes that differ from template meshes such as SMPL-X. This hybrid approach mitigates the difference between the coarse resolution provided by volume primitives and the high-fidelity surface details required for complex animations and poses. While the mesh provides detailed contours and the overall structure of the avatar, the volume primitives provide flexibility in representing a wider range of shape variations beyond the constraints of the predefined model. At the same time, Gaussian volumes can be used to detail more subtle distinctions that exceed the resolution of the primitives, such as subtle facial expressions, complex clothing textures, or dynamic hair movement. This layer of detail can ensure that the final rendered avatar (or object) accurately follows the desired pose and exhibits a level of detail and realism that is typically unattainable with traditional modeling techniques. Through this combined representation, the modeler system 110 can render avatars that present diverse and complex shapes and increase the level of detail, thereby improving the overall visual quality and realism of the animated character.
[0082] For more details, refer to Figure 1 The texture model 116 can generate outputs for predictions of color, rotation, and can scale the field of the object. For example, the position of the Gaussian Can be used to extract neural attribute fields Query the color of each Gaussian Rotation and zoom For example, to fully exploit the expressiveness of 3D Gaussians, the texture model 116 can be used to allow each Gaussian to have individual properties, such as color features, scale, rotation, and opacity. However, this can lead to unstable training where Gaussians within a local neighborhood have very different properties, resulting in noisy geometry and rendering. This is especially true when the gradient of the optimization objective has high variance, such as the SDS objective in Equation 2. To stabilize and amortize the training process, rather than directly optimizing the properties of the Gaussians, the texture model 116 can be implemented to predict these properties using a neural implicit field. As shown, for each Gaussian The modeler system 110 may first calculate the canonical position in the world coordinate system in Equation 3 in represents the static pose. Then, the texture model 116 can use the and The standard position of To query the color of each Gaussian Rotation Zoom and opacity This can be represented by a neural network with parameters φ and ψ:
[0083]
[0084] Among them, the texture model 116 can use a separate neural field to output the opacity of the Gaussian, while the other properties are determined by Separate neural fields are designed because the opacity of the Gaussian is closely related to the underlying geometry of the avatar and requires special handling. Querying the neural field, the texture model 116 can normalize the Gaussian body properties, which can then be shared across different poses and animations. The texture model 116 uses the neural implicit field to constrain nearby Gaussians to have consistent properties, thereby stabilizing and amortizing the training process and allowing high-quality avatar synthesis using high-variance losses. In some implementations, the implicit field can be regularized to promote smooth transitions of properties on the surface of the object, thereby mitigating abrupt changes and ensuring that adjacent Gaussians have gradually changing properties.
[0085] In extending the functionality of the texture model 116, the modeler system 110 may use an implicit Gaussian volume attribute field To provide texture extraction, thereby improving the fidelity of the final avatar rendering. For example, by querying the Gaussian color field, the modeler system 110 extracts a high-quality 3D texture that can be directly applied to the mesh. The intrinsic color properties of the Gaussian can provide an initial texture that captures the subtle appearance of the avatar. Subsequently, the textured mesh can be rendered in RGB. Using SDS loss The color field is iteratively fine-tuned to refine the texturing process, sharpening texture details and aligning the visual output with the geometric accuracy of the avatar form.
[0086] Utilizing the neural implicit fields in the texture model 116, improvements can be provided by implicitly enforcing spatial consistency between the properties of adjacent Gaussians. This provides a technical improvement because it addresses the dependency problem between Gaussians, ensuring that adjacent Gaussians exhibit similar properties. By not feeding each Gaussian individually (which would allow them to move without regard to their neighbors), the texture model 116 promotes a degree of interdependence, resulting in smoother transitions and more uniform properties across the surface of the avatar. This cohesive approach improves training stability because it mitigates the risks associated with high variance gradients that can arise during the optimization of complex models. Additionally, this approach facilitates improved and reliable synthesis of high-quality avatars because the consistency of properties between the resulting Gaussians provides more realistic and visually pleasing animations. Therefore, the properties predicted by the neural fields maintain the structural and visual integrity of the avatar across a variety of poses and movements.
[0087] For more details, refer to Figure 1 The object layer model 114 in the object layer model 114 can output a signed distance field (SDF) that represents the underlying geometry of the object (the outer layer of the object), where the SDF is the distance value from each point in the 3D space to the surface of the object. For example, the SDF value of each Gaussian volume can be obtained from the neural SDF S ψ Query, and can be done through the kernel function Converted to opacity The neural network of the object layer model 114 may be trained to obtain the value of the object layer model 114 by using a signed distance field (SDF) function S with parameter ψ. ψ To represent the underlying geometry of the 3D Gaussian volume. For example, the object layer model 114 may use a kernel function (Equation 8) to parameterize the opacity of each 3D Gaussian volume based on their signed distance to the surface:
[0088]
[0089] in is a bell-shaped kernel function with learnable parameters {γ,λ} that maps signed distances to opacity values. The opacity parameterization builds on the previous Gaussian that should be kept close to the surface for high opacity. The parameter λ controls the tightness of the high opacity neighborhood of the surface, and α controls the overall scale of the opacity. The SDF-based Gaussian opacity parameterization fits the primitive-based implicit Gaussian representation because the object layer model 114 can now take the above opacity field as Defined as the product of SDF and kernel function: where the neural network can be used to directly represent the SDF S ψ .
[0090] In addition, since the neural network uses SDF S ψ To represent the underlying geometry of the 3D Gaussian volume, the object layer model 114 can extract the mesh from the SDF via a differentiable marching tetrahedron (DMTet) (Formula 9):
[0091]
[0092] in is a learnable parameter representing the ψ For example, Can be used as a learnable offset for the SDF. The neural network may not use a level set of 0, since the Gaussian may not have the highest opacity value to achieve the desired rendering. In some implementations, the level set Adjusted during training to fine-tune the grid From SDF S ψ The threshold at which is extracted from , allowing manipulation of the resolution of the mesh and the details it captures from the underlying SDF representation. For example, The higher the value, the more fine details the mesh contains that are present in the SDF. Lower values result in a simpler mesh, focusing attention on the larger, more important geometric features of the avatar.
[0093] The DMTet method can synthesize high-resolution 3D shapes from simple inputs such as coarse voxels by adopting a hybrid 3D representation that combines implicit and explicit forms. Unlike traditional implicit methods that focus on regressing signed distance values, DMTet directly optimizes for reconstructing surfaces, enabling the synthesis of finer geometric details and reducing artifacts. In some implementations, the model uses a deformable tetrahedral grid to encode the discretized signed distance function and uses a differentiable marching tetrahedron layer to convert the implicit distance representation to an explicit surface mesh. This allows joint optimization of surface geometry and topology, as well as the generation of subdivision hierarchies through reconstruction and adversarial losses defined on the surface mesh.
[0094] Using the DMTet method, the neural network uses the SDF S of 3D Gaussian geometry ψ Allows the modeler system 110 to extract the avatar mesh by applying the DMTet process For example, by adjusting the level set parameters to control the extraction, thereby optimizing the balance between capturing detailed geometric features and maintaining computational efficiency. As shown in the figure, the DMTet process can create three layers of virtual images: (1) objects, (2) meshes, and (3) a set of Gaussian volumes defined relative to primitives, which appear as static poses of objects. For example, the object layer provides rough outlines, the mesh layer adds detailed surface geometry, and the Gaussian layer gives finer texture and shape details through individual attributes.
[0095] In addition, Gaussians are often used to create videos and simple visual effects, where speed and computational efficiency take precedence over the high fidelity (hi-fi) required for 3D asset generation. This application is largely due to the inherent limitations of having sufficient detail and accuracy when representing complex, dynamic 3D shapes and textures. However, the disclosed GAvatar implementation provides a significant technical advancement in this area. By combining Gaussian representations with modeling techniques such as signed distance fields (SDFs), differentiable marching tetrahedrons (DMTets), and neural implicit fields, the modeler system 110 expands the use cases of Gaussians. It allows the creation of high-fidelity 3D avatars that can be animated and transformed in a variety of poses and expressions with greater detail and realism. The GAvatar implementation improves the expressiveness and dynamic range of 3D models, solves the problem of inter-Gaussian dependencies, and provides a coherent and consistent visual output that meets the complex requirements of modern digital environments. Therefore, the GAvatar implementation provides an improved technical solution that expands the potential of Gaussian-based modeling and sets a new standard for generating detailed and realistic 3D assets.
[0096] Both the SDF and the extracted mesh can allow the object layer model 114 to regularize the geometry of the 3D Gaussian avatar (or another 3D Gaussian object) using various losses. For example, an Eikonal regularizer can be used to maintain an appropriate SDF, which is defined as (Formula 10):
[0097]
[0098] where p∈P contains the center points of all Gaussians in world coordinates and points sampled around the Gaussians using a normal distribution. In some implementations, an Eikonal regularizer helps ensure that the SDF maintains unit gradient outside the surface of the object, which is important for accurate representation of the distance field and subsequent geometry extraction. For example, during backpropagation, the regularizer adjusts the network parameter ψ to correct any deviation from the unit gradient condition.
[0099] In addition, the object layer model 114 may employ an alpha loss to transform the mask I rendered using the extracted mesh. M with the alpha image I from Gaussian sputteringα Matching (Formula 11):
[0100]
[0101] in The mask image I generated by the extracted grid is quantized M and the alpha image I generated by Gaussian sputtering α Since the transparency of the Gaussian allows for rendering of subtle visual features, the alpha loss helps align these renderings with the mesh outlines, confirming visual consistency. For example, the comparison ensures that the geometry captured by the Gaussian rendering is closely mirrored by the extracted mesh, providing a supervisory signal for the geometric fidelity within the model’s learned architecture.
[0102] In addition, an SDS loss for normals can be determined to supervise the normal rendering of the extracted mesh using differentiable rasterization. N The SDS gradient can be calculated as follows (Formula 12):
[0103]
[0104] Among them I N,t is a noise normal image. In some implementations, the noise normal image I N,t The object layer model 114 is used to train the object layer model 114 to resist potential perturbations, thereby enhancing the stability of the normal estimation. For example, the model may introduce synthetic noise during training to simulate real-world defects in the data. For example, the SDS normal loss The normal map can be used as input to the diffusion model by ensuring that it contributes to the supervision of the SDF neural network.
[0105] In some implementations, a normal consistency loss can be used It will mesh In some implementations, normal consistency enforces smoothness of the generated mesh by penalizing the differences between normals of adjacent vertices. For example, when reconstructing an avatar (e.g., an organic or new avatar), Can be used to enforce smooth transitions between surface elements to maintain a realistic appearance.
[0106] In addition, the reconstruction loss can be used For example, The object layer model 114 and the texture model 116 may be helped to refine the similarity of the virtual character to a given real-world image. The operation is performed by comparing the generated avatar image to a provided reference image (e.g., a photo of a person) and minimizing the appearance differences, especially in terms of pose and visual texture. To ensure that the output of the renderer 118 (including visual details such as shadows, highlights, and outlines) is aligned with those of the reference image. This allows the consistency between the reference image and the rendered image to be evaluated from multiple viewpoints, thus providing a multi-dimensional assessment of the accuracy of the model. Figure 1 For consistency, the modeler system 110 can verify whether the visual representation of the avatar is coherent when viewed from various angles.
[0107] Referring to block 202 (Gaussian volume property calculation for rest pose) in more detail, the modeler system 110 initiates the process of creating a 3D representation of an animatable object (e.g., a human avatar) starting from a rest pose, which is a baseline for the geometry of the object without any deformation caused by motion or action. Figure 1 The transformer 112 in the embodiment can generate a series of primitives (e.g., geometric shapes that approximate the shape of the object). These primitives are placed on the mesh surface of the object, aligned with its underlying structure. The primitives in the static pose can be generated by the transformer 112 using the mesh of the avatar in a neutral, undeformed state, using its position, rotation, and scale (denoted as {P k , R k , S k}) to define the parameters.
[0108] In response to the creation of these primitives and the generation of the Gaussian volume, the transformer 112 applies a local-to-world position transformation, such as defined by Equation 3. For example, the transformation can transform each Gaussian volume in the primitive coordinate system to The local position and parameters of the primitive {P k , R k , S k} as input and transform them into global positions in the world coordinate system
[0109] In some implementations, the transformed positions of the primitives become input to the object layer model 114 and the texture model 116. The object layer model 114 can use the transformed positions to create an SDF, which is a representation of the avatar's geometry. The SDF assigns a distance value to each point in space relative to the avatar's surface. The object layer model 114 can then apply a differentiable marching tetrahedron (DMTet) algorithm to the SDF to convert it into a mesh - the geometry of the avatar that can be visually rendered. During this mesh generation process, a normal consistency loss can be calculated. and SDS normal loss To ensure that the geometry of the mesh is smooth and accurately reflects the shape of the avatar. For example, a loss can be used to regularize the mesh generation process so that it is closely aligned with the surface details defined by the SDF.
[0110] At the same time, the texture model 116 can use the same transformed primitive position to assign color, rotation, scale, and opacity attributes to each Gaussian within the primitive. Opacity is specifically affected by the data of the SDF and is calculated using Equation 8: in is a kernel function that converts the signed distance values of the SDF into opacity values of the Gaussian volume. This relationship tightly ties the visual appearance of the avatar (its texture and shape) to the geometric representation generated by the SDF. After the texture model 116 defines the properties of each Gaussian volume (including its color and opacity), the Gaussian can perform Gaussian splatting, which is a process of using these properties to render the visual representation of the avatar (e.g., displayed against a contrasting background), thereby producing an image in which the avatar is highlighted (e.g., displayed as white against a black background).
[0111] Referring to block 204 (target pose generation) in more detail, the modeler system 110 may receive a target pose and determine a pose-related correction. For example, a correction is a modification to primitive parameters that will / may allow the avatar to move from its initial neutral rest pose to a desired target pose specified by a user or system. In some implementations, the modeler system 110 takes into account the avatar's current pose and shape parameters (θ, β) via a linear blend skinning (LBS) function LBS(θ, β). The pose-related correction -δP ω (θ), δR ω (θ),δS ω (θ) - is applied to the grid-based primitive initialization from block 202, and the updated primitive parameters obtained by Equation 1. These updated parameters represent the new position, orientation, and scale of the primitive that conforms to the target pose. In response to the primitive being adjusted for the target pose, the next step involves combining this information with the Gaussian volume attributes calculated at block 202. The primitive is represented by its position, rotation, scale, color characteristics, and opacity. Each Gaussian defined (206) By applying Equation 4 at block 208 and formula 5 The local-to-world transformation described in is transformed according to the new primitive configuration. For example, these formulas can adjust the Gaussian to the scale and rotation of the transformed primitive.
[0112] In some implementations, the object layer model 114 and the texture model 116 are used both for the transition from the rest pose to the target pose and for the final rendering of the avatar. After applying the local-to-world transform to the Gaussian volumes (at block 208), the object layer model 114 computes signed distance field (SDF) values for each Gaussian volume. The SDF gives a measure of how far a point is from the surface of the avatar, with the sign indicating whether the point is inside or outside the avatar. Using these SDF values, the object layer model can perform a differentiable marching tetrahedron (DMTet) to generate a mesh For example, a mesh is a 3D geometric representation of an avatar at a target pose. The SDF can also be used to calculate the Determines the opacity of the Gaussian. The closer it is to the surface, the higher the opacity. The farther it is from the surface, the lower the opacity.
[0113] In parallel (or sequentially), the texture model 116 can use the positions of the Gaussians in the world coordinate system to predict their visual properties, such as color and opacity. For example, this can be done using a neural implicit field, which is a function parameterized by a neural network that maps the position of each Gaussian to its visual properties. In this example, the neural implicit field will ensure that the properties vary smoothly and consistently across the surface of the avatar.
[0114] Thus, in target pose generation, both the object layer model 114 and the texture model 116 are used to produce an accurate and detailed avatar that can move realistically. The object layer model 114 provides the geometric integrity and motion of the avatar, while the texture model 116 provides visual realism by defining the appearance of the avatar's surface. The output of the model is then used for Gaussian sputtering at block 212, which renders the avatar in the target pose with the desired visual properties, resulting in object 214. In some implementations, the SDS loss can be calculated by The rendering process is iteratively refined (at block 216) to ensure that the visual output is aligned with a desired target pose specified by a user or system input.
[0115] In the optimization phase, the modeler system 110 can optimize the process of building the digital avatar by optimizing several model components and parameters simultaneously. It can be defined as (Formula 13):
[0116]
[0117] where the terms in the function are relevant to a particular purpose (note that weighting terms are omitted for brevity): Adjust fractional distillation; Regularize the positions of the Gaussians to stay close to their respective primitives; Keep the SDF with unit gradient; Ensure that the rendered image matches the intended transparency; Optimize normal images; Maintain the smoothness of the mesh normals. Using this goal, the modeler system 110 simultaneously optimizes the local position of the Gaussian volume Parameters of Gaussian attribute field SDF ψ and its associated opacity kernel parameters γ,λ, and primitive motion correction network δP ω , δR ω ,δS ω and SMPL-X shape parameter β. In some implementations, the object layer model 114 may refine the SDF S ψ , and the texture model 116 can adjust the Gaussian attribute field These models can be used together to update the Gaussian local position SDF and opacity kernel parameters γ, λ, correction network P for primitive motion ω , δR ω ,δS ω , and the SMPL-X shape parameter β.
[0118] For example, initialization can be a preparation phase where the avatar's UV map is split into a 64x64 grid, resulting in 4096 primitive regions. For each of these primitive regions, V k , a set of 512 Gaussians. The local position of the Gaussian A uniformly distributed 8x8x8 grid can be initialized within each primitive. The structure placement provides a starting configuration for subsequent refinement through the optimization process.
[0119] In some implementations, training involves iteratively adjusting the model using a procedure called Gaussian densification, which can occur every 100 iterations to accommodate the varying complexity of the avatar during animation. Rendering the RGB image I, the modeler system 110 can determine or identify the target pose θ from two sources: (1) the natural pose θ optimized with the above variables N ; (2) Random pose θ sampled from the animation database A To ensure realistic animation. Both the object layer model 114 and the texture model 116 can contribute to realistic avatar animation.
[0120] In some implementations, the total loss function Formula 13 can include the reconstruction loss, expressed as To improve the fidelity of virtual image generation. For example, An optimization cycle may be incorporated to minimize the difference between the generated avatar view and the actual image view. In addition, multi-view reconstruction may be implemented to allow the modeler system 110 to derive a more robust representation of the avatar. By using another model trained to synthesize multi-view images from a single input image or text, the modeler system 110 can use multiple view images across various viewing angles (e.g., front, back, left, right, top-down). Multi-view training can enhance the spatial and visual accuracy of the avatar, providing a detailed set of constraints that guide the avatar’s 3D model to more accurately and consistently reconstruct the user’s image from different angles.
[0121] Reference now Figure 3 , each block of the method 300 described herein includes a computing process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in a memory. The method can also be embodied as computer-usable instructions stored on a computer storage medium. The method can be provided by a stand-alone application, a service or a hosted service (standalone or in combination with another hosted service), or a plug-in to another product, to name a few. In addition, the method 300 is directed to Figure 1 and Figure 2 The systems and architectures are described by way of example. However, the method may additionally or alternatively be performed by any system or any combination of systems, including but not limited to the systems described herein.
[0122] Figure 3 is a flow chart illustrating a method 300 for generating a realistic animatable avatar from a text description according to some embodiments of the present disclosure. The various operations of method 300 may be implemented by the same or different devices or entities at different points in time. For example, one or more first devices may implement operations related to configuring a machine learning model, one or more second devices may implement operations related to rendering, and one or more third devices may implement operations related to receiving user input. One or more third devices may maintain a neural network model, or may access the neural network model using, for example, but not limited to, an API provided by one or more first devices and / or one or more second devices.
[0123] Method 300 includes, at box 310, assigning first elements of a three-dimensional (3D) model of a subject to locations on a surface. For example, one or more processing circuits may assign multiple first elements of the 3D model of the subject to multiple locations on the surface of the subject in an initial pose. The first element may be a primitive assigned to a subject or an object (e.g., a human body surface), and the initial pose may be a static pose. In some implementations, the primitive (e.g., the first element) may include (1) a position parameter corresponding to a position in the multiple locations to which at least one element is assigned, (2) a scale parameter indicating the scale of at least one first element in a 3D reference system in which the subject is located, and (3) an orientation parameter indicating the orientation of the subject relative to the 3D reference system.
[0124] In some implementations, primitives are assigned as first elements to the surface of a subject in an initial pose as a framework for subsequent transformations. These primitives are geometrically consistent with the natural contours of the subject, thereby providing an infrastructure from which detailed modeling and animation can be performed. For example, primitives can be cubes or other polyhedral elements whose size and orientation can be adjusted to be consistent with the anatomical features of the subject. In some implementations, input from a user or system can be used to design an avatar. For example, a processing circuit can receive an indication of one or more attributes of a subject as at least one of text data, voice data, audio data, image data, or video data. For example, a processing circuit can interpret a text input describing a desired posture or appearance and convert it into specific modeling parameters that define the posture and aesthetics of the avatar. In another example, an uploaded image or video can be used as a reference for the subject's attributes, where the processing circuit extracts key features and converts them into modeling parameters that guide the placement and configuration of primitives.
[0125] The method 300 includes, at box 320, assigning a second element to a first element, at least one second element having an opacity corresponding to a distance. For example, one or more processing circuits may assign a plurality of second elements to a plurality of first elements, each of the plurality of second elements having an opacity corresponding to a distance between the second element and a surface of a subject. The second element may be a 3D Gaussian body assigned to each primitive. In addition, the opacity may correspond to a signed distance function to account for Gaussian transparency. For example, the opacity of the 3D Gaussian body may reflect the proximity to the surface of the subject, which utilizes the characteristics of the SDF to dynamically adjust the visibility of each Gaussian body based on the spatial relationship of each Gaussian to the avatar.
[0126] In some implementations, determining the opacity of each second element in the plurality of second elements may include using a signed distance function (SDF) to represent the distance between the second element and the surface of the subject. The SDF may be used to represent the geometry of a 3D Gaussian volume by calculating the minimum distance from any point in space to the nearest surface point. In some implementations, at least one second element in the plurality of second elements includes a 3D Gaussian splash defined in the local reference frame of a corresponding at least one first element in the plurality of first elements. For example, the Gaussian splash within the local reference frame of the primitive provides control over the distribution and blending of these details, ensuring that each Gaussian makes the best contribution to the overall appearance of the avatar. For example, processing circuitry may exploit this to simulate complex textures such as fabrics, hair, or skin, where varying degrees of transparency and color are critical to realism.
[0127] The method 300 includes, at box 330, updating the second element based at least on the target pose of the subject and one or more attributes of the subject. For example, one or more processing circuits may update multiple second elements based at least on the target pose of the subject and one or more attributes of the subject to determine multiple updated second elements. In some implementations, the model may be updated and / or optimized relative to a target view (e.g., an animation from a static pose). In some implementations, one or more attributes may be represented by a neural field.
[0128] The updating of the second element can be based at least on the evaluation of one or more objective functions and the representation of the body. For example, the SDS loss is used to match the rendered image with the target appearance, thereby guiding the optimization of the Gaussian position and attributes to obtain a consistent visual output. For example, the processing circuit can iteratively refine the Gaussian parameters to ensure that the appearance of the virtual image in various postures is consistent with the input. In some implementations, the Eikonal rule can be used to regularize the multiple second elements based on at least the position data of each of the multiple second elements. For example, by preventing sudden changes in the SDF representing the surface of the virtual image, regularization can help maintain the geometric accuracy of the virtual image. For example, the SDF parameters can be adjusted based on the position data of the Gaussian body.
[0129] Additionally, in some implementations, an alpha loss may be used to optimize a 3D Gaussian volume, the method comprising updating a plurality of updated second elements based at least on a mask determined from a representation of the subject and an alpha rendering determined from the plurality of updated second elements. This includes updating a plurality of updated second elements based at least on a mask determined from a representation of the subject and an alpha rendering determined from the plurality of updated second elements. Additionally, the process may include comparing the alpha values of the generated image to a target alpha map to confirm that the opacity level of the Gaussian volume accurately reflects the visual depth and layering in the scene. For example, the opacity value may be fine-tuned to achieve a natural overlap between an avatar and its background or between different portions of the avatar itself.
[0130] The method 300 includes, at block 340, rendering a representation of the subject based on the updated second element. For example, one or more processing circuits may render a representation of the subject based at least on a plurality of updated second elements. In some implementations, the rendered image is generated from the 3D model in response to the primitives and Gaussians being configured and optimized. For example, the rendering of the representation may be a generated textured mesh of the subject. For example, the generated textured mesh may be a 3D model that combines geometric vertices, edges, and faces with surface textures in an attempt to represent the visual appearance and physical structure of an object or character. These textures include color maps, normal maps, and specular maps that simulate real-world surfaces. In some implementations, the rendering process may include shading techniques and light simulations to enhance the realism of the textured mesh. For example, ambient occlusion, shadow mapping, and reflection models may be applied to the 3D avatar to simulate real-world lighting conditions and interactions with the environment. In addition, in some implementations, rendering involves using ray tracing to achieve realistic lighting effects, in which light is simulated as it bounces off a surface, thereby producing natural shadows and reflections.
[0131] In some implementations, the processing circuitry includes at least one of: a system for generating synthetic data, a system for performing simulation operations, a system for performing collaborative content creation of 3D assets, a system for performing conversational AI operations, a system including one or more large language models (LLMs), a system including one or more visual language models (VLMs), a system for performing digital twin operations, a system for performing light transport simulations, a system for performing deep learning operations, a system implemented using an edge device, a system implemented using a robot, a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine, a system including one or more virtual machines (VMs), a system implemented at least in part in a data center, or a system implemented at least in part using cloud computing resources.
[0132] It should be appreciated that method 300 provides various improvements over existing systems. One improvement includes that method 300 achieves extremely fast rendering speeds due to the use of 3D Gaussian volumes, since the optimized post-processing circuit no longer needs to query Gaussian volume properties from implicit fields. For example, a generated avatar with 2.5 million Gaussian volumes can be rendered at a resolution of 1024x1024 at a speed of 100fps, which is much faster than NeRF-based methods. In addition, Gaussian rendering only takes about 3ms (300+fps), so it can be further accelerated by optimizing the speed of non-rendering operations (such as LBS and primitive transformations).
[0133] Another improvement includes generating realistic animatable avatars from text descriptions using Gaussian sputtering. As described herein, method 300 introduces a primitive-based 3D Gaussian volume representation that defines 3D Gaussian volumes within gesture-driven primitives. This representation naturally supports animation and allows flexible modeling of detailed avatar geometry and appearance by deforming Gaussians and primitives.
[0134] Another improvement includes using an implicit Gaussian attribute field to predict Gaussian attributes, which can stabilize and amortize the learning of a large number of Gaussians and allow the processing circuit to generate high-quality avatars using a high-variance optimization objective (such as SDS). In addition, after the avatar optimization, since the processing circuit can directly obtain the Gaussian attributes and can skip querying the attribute field, method 300 achieves extremely fast (100fps) rendering of the neural avatar at a resolution of 1024x1024. This is much faster than existing NeRF-based avatar models (which query the neural field for each new camera view and avatar pose).
[0135] Another improvement includes the implementation of a new implicit mesh learning method based on signed distance functions (SDFs) that links SDFs to Gaussian opacity. For example, it allows the processing circuit to regularize the underlying geometry of the Gaussian avatar and extract a high-quality textured mesh.
[0136] Reference now Figure 4 According to some embodiments of the present disclosure, using Figure 1 to Figure 23D object generation architecture described herein. As shown, an avatar generated using the 3D object generation architecture described herein and an avatar corresponding to a mesh normal and a textured mesh. For example, the renderings illustrate the detailed surface geometry and subtle texture output that can be achieved. Gaussian rendering reference avatars 410 and 430 of the avatars are shown. A mesh normal reference avatar 420 of the avatar is shown, and a textured mesh reference avatar 440 is shown. As a result, the visualization of the mesh normal and textured mesh has the accuracy of the geometry and the fidelity of the texture - creating a lifelike digital representation (e.g., as shown in a static pose).
[0137] Reference now Figure 5 According to some embodiments of the present disclosure, using Figure 1 to Figure 2 3D object generation architecture. Avatar 510 is shown without an implicit Gaussian volume attribute field, while avatar 520 is depicted with an implicit Gaussian volume attribute field. For example, avatar 510 is depicted based on disabling the implicit Gaussian volume attribute field and directly optimizing the Gaussian volume attributes. It can be observed that the generated avatar is significantly better than Figure 1 to Figure 2 The 3D object generation architecture is poor, with noticeable noise and color oversaturation. Therefore, it can be challenging when the architecture attempts to directly optimize millions of individual Gaussians with a high variance loss (such as SDS). In contrast, the implicit Gaussian volume attribute field allows for a more stable and robust optimization process (e.g., as shown in avatar 520). In addition, avatar 530 is shown without SDF-based grid learning, while avatar 540 is depicted with SDF-based grid learning. For example, avatar 510 is depicted based on disabling SDF-based grid learning and instead having the Gaussian volume attribute field additionally output Gaussian opacity. As shown in avatar 530, avatars generated without grid learning may be missing body parts and have distorted body shapes, while SDF-based grid learning handles these issues by regularizing the underlying geometry of the Gaussian avatar (as shown in avatar 540).
[0138] Reference now Figure 6 According to some embodiments of the present disclosure, using Figure 1 to Figure 2 An example illustration of an object rendering of a 3D object generation architecture. Figure 6 It should be understood that Figure 1 to Figure 2 The technical advantage of the 3D object generation architecture is that it allows the extraction of high-quality differentiable mesh representations of Gaussian virtual images. Figure 1 to Figure 2The mesh extraction method of the 3D object generation architecture of 60 (depicted in mesh 630) is compared with the Gaussian density-based method used in DreamGaussian (depicted in mesh 620), which is an implementation of extracting meshes from 3D Gaussian volumes. For example, mesh 630 is provided from the Gaussian volume attributes in GAvatar rendering 610 using a mesh extraction pipeline to obtain the final mesh, mesh 630. It should be appreciated that the mesh extracted by DreamGaussian (mesh 620) is noisy and lacks geometric details, while the mesh extracted by DreamGaussian (mesh 620) is noisy and lacks geometric details. Figure 1 to Figure 2 The method implemented by the 3D object generation architecture can obtain smoother meshes with fine-grained geometric details.
[0139] Reference now Figure 7 , a comparative example of a flawed method of generating an animatable avatar compared to generating an animatable 3D Gaussian avatar according to some embodiments of the present disclosure. For example, Figure 7 The animatable 3D Gaussian avatar method GAvatar is compared with other methods: DreamGaussian, AvatarCLIP, AvatarCraft and Fantasia3D. In addition, for the sake of completeness, Figure 7 Comparisons are also made with contemporary works DreamHumans and TADA. Figure 1 to Figure 2 The GAvatars 710 and 720 generated by the 3D object generation architecture can provide higher quality avatars in terms of geometry and appearance. DreamGaussian, AvatarCLIP, AvatarCraft, DreamHumans, TADA, and Fantasia3D are unable to model complex avatars. Therefore, the generated GAvatars 710 and 720 are significantly improved avatars compared to all methods.
[0140] Example content streaming system
[0141] See now Figure 8 , Figure 8 is an example system diagram for a content streaming system 800 according to some embodiments of the present disclosure. Figure 8 Includes one or more application servers 802 (which may include Figure 5 ), one or more client devices 804 (which may include components, features, and / or functionality similar to the example computing device 500 of Figure 5) and one or more networks 806 (which may be similar to one or more networks described herein). In some embodiments of the present disclosure, system 800 may be implemented to perform model training as well as runtime operations. Application sessions may correspond to game streaming applications (e.g., NVIDIA GeFORCE NOW), remote desktop applications, simulation applications (e.g., autonomous or semi-autonomous vehicle simulations), computer-aided design (CAD) applications, virtual reality (VR) and / or augmented reality (AR) streaming applications, deep learning applications, and / or other application types. For example, system 800 may be implemented to receive input indicating one or more features of an output to be generated using a neural network model, provide the input to the model so that the model generates an output, and use the output for various operations including display or simulation operations.
[0142] In the system 800, for an application session, one or more client devices 804 may receive only input data in response to input to one or more input devices, transmit the input data to one or more application servers 802, receive encoded display data from one or more application servers 802, and display the display data on a display 824. Thus, more computationally intensive calculations and processing are offloaded to one or more application servers 802 (e.g., rendering for graphical output of the application session—particularly ray or path tracing—performed by one or more GPUs of one or more game servers 802). In other words, the application session is streamed from one or more application servers 802 to one or more client devices 804, thereby reducing the requirements of one or more client devices 804 for graphics processing and rendering.
[0143] For example, with respect to instantiation of an application session, the client device 804 may display a frame of the application session on a display 824 based on receiving display data from one or more application servers 802. The client device 804 may receive input from one of the one or more input devices and generate input data in response, such as providing a prompt as input for generating a 3D avatar. The client device 804 may send the input data to the application server 802 via the communication interface 820 and via the network 806 (e.g., the Internet), and the application server 802 may receive the input data via the communication interface 818. The CPU may receive the input data, process the input data, and transmit the data to the GPU, which causes the GPU to generate a rendering of the application session. For example, the input data may represent movement or animation of a character of a user in a game session of a game application, firing a weapon, reloading, passing a ball, steering a vehicle, etc. The rendering component 812 may render the application session (e.g., representing the result of the input data), and the rendering capture component 814 may capture the rendering of the application session as display data (e.g., as image data capturing a rendered frame of the application session). The rendering of the application session may include lighting and / or shadow effects of ray or path tracing calculated using one or more parallel processing units (such as GPUs) of the application server 802, which may further use one or more dedicated hardware accelerators or processing cores to perform ray or path tracing techniques. In some embodiments, one or more virtual machines (VMs) - for example, including one or more virtual components such as vGPUs, vCPUs, etc. - may be used by the application server 802 to support the application session. The encoder 816 may then encode the display data to generate encoded display data, and the encoded display data may be sent to the client device 804 via the communication interface 818 over the network 806. The client device 804 may receive the encoded display data via the communication interface 820, and the decoder 822 may decode the encoded display data to generate the display data. The client device 804 may then display the display data via the display 824.
[0144] Example computing device
[0145] Fig. 9900 is a block diagram of an example computing device 900 suitable for implementing some embodiments of the present disclosure. The computing device 900 may include an interconnect system 902 that directly or indirectly couples the following devices: a memory 904, one or more central processing units (CPUs) 906, one or more graphics processing units (GPUs) 908, a communication interface 910, an input / output (I / O) port 912, an input / output component 914, a power supply 916, one or more presentation components 918 (e.g., one or more displays), and one or more logic units 920. In at least one embodiment, one or more computing devices 900 may include one or more virtual machines (VMs), and / or any of their components may include virtual components (e.g., virtual hardware components). For non-limiting examples, one or more of the GPUs 908 may include one or more vGPUs, one or more of the CPUs 906 may include one or more vCPUs, and / or one or more of the logic units 920 may include one or more virtual logic units. As such, one or more computing devices 900 may include discrete components (eg, a full GPU dedicated to computing device 900 ), virtual components (eg, a portion of a GPU dedicated to computing device 900 ), or a combination thereof.
[0146] although Fig. 9 The various blocks of are shown as being connected with wires via interconnect system 902, but this is not intended to be limiting and is merely for clarity. For example, in some embodiments, presentation component 918 (such as a display device) may be considered to be I / O component 914 (e.g., if the display is a touch screen). As another example, CPU 906 and / or GPU 908 may include memory (e.g., memory 904 may represent a storage device in addition to the memory of GPU 908, CPU 906, and / or other components). In other words, Fig. 9 The computing devices described herein are merely illustrative. No distinction is made between categories such as "workstations," "servers," "laptops," "desktop computers," "tablet computers," "client devices," "mobile devices," "handheld devices," "game consoles," "electronic control units (ECUs)," "virtual reality systems," and / or other device or system types, as all are contemplated herein. Fig. 9 within the range of computing devices.
[0147] Interconnection system 902 can represent one or more links or buses, such as address bus, data bus, control bus or its combination.Interconnection system 902 can be arranged in various topological structures, including but not limited to bus, star, ring, grid, tree or mixed topological structure.Interconnection system 902 may include one or more bus or link types, such as industrial standard architecture (ISA) bus, extended industrial standard architecture (EISA) bus, video electronics standard association (VESA) bus, peripheral component interconnect (PCI) bus, fast peripheral component interconnect (PCIe) bus and / or another type of bus or link. In some embodiments, there is a direct connection between components. For example, CPU 906 can be directly connected to memory 904. Further, CPU 906 can be directly connected to GPU 908. In the case of direct connection or point-to-point connection between components, interconnection system 902 may include PCIe link to perform the connection. In these examples, it is not necessary to include PCI bus in computing device 900.
[0148] The memory 904 may include any of a variety of computer-readable media. Computer-readable media may be any available media that can be accessed by the computing device 900. Computer-readable media may include volatile and non-volatile media, as well as removable and non-removable media. By way of example and not limitation, computer-readable media may include computer storage media and communication media.
[0149] Computer storage media may include volatile and nonvolatile media and / or removable and non-removable media implemented in any method or technology for storing information such as computer readable instructions, data structures, program modules, and / or other data types. For example, memory 904 may store computer readable instructions (e.g., representing programs and / or program elements, such as an operating system). Computer storage media may include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by computing device 900. As used herein, computer storage media does not include the signals themselves.
[0150] Computer storage media may embody computer readable instructions, data structures, program modules, and / or other data types in a modulated data signal (such as a carrier wave or other transmission mechanism), and include any information delivery media. The term "modulated data signal" may refer to a signal that has one or more of its characteristics set or changed in a manner that encodes the information in the signal. By way of example and not limitation, computer storage media may include wired media (such as a wired network or a direct wired connection) and wireless media (such as acoustic, RF, infrared, and other wireless media). Combinations of any of the above should also be included within the scope of computer readable media.
[0151] The CPU 906 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 900 to perform one or more of the methods and / or processes described herein. Each of the CPUs 906 may include one or more cores (e.g., 1, 2, 4, 8, 28, 72, etc.) capable of processing multiple software threads simultaneously. The CPU 906 may include any type of processor, and may include different types of processors (e.g., a processor with fewer cores for mobile devices and a processor with more cores for servers) depending on the type of computing device 900 implemented. For example, depending on the type of computing device 900, the processor may be an advanced RISC machine (ARM) processor implemented using reduced instruction set computing (RISC) or an x86 processor implemented using complex instruction set computing (CISC). In addition to one or more microprocessors or supplementary coprocessors such as mathematical coprocessors, the computing device 900 may also include one or more CPUs 906.
[0152] In addition to or in place of the CPU 906, one or more GPUs 908 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 900 to perform one or more of the methods and / or processes described herein. One or more of the GPUs 908 may be an integrated GPU (e.g., with one or more of the CPUs 906) and / or one or more of the GPUs 908 may be a discrete GPU. In an embodiment, one or more of the one or more GPUs 908 may be a coprocessor of one or more of the one or more CPUs 906. The GPU 908 may be used by the computing device 900 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, the GPU 908 may be used for general-purpose computing (GPGPU) on a GPU. The GPU 908 may include hundreds of thousands of cores capable of processing hundreds of thousands of software threads simultaneously. The GPU 908 may generate pixel data for outputting an image in response to a rendering command (e.g., a rendering command from the CPU 906 received via a host interface). GPU 908 may include graphics memory, such as display memory, for storing pixel data or any other suitable data, such as GPGPU data. Display memory may be included as part of memory 904. GPU 908 may include two or more GPUs operating in parallel (e.g., via a link). The link may directly connect the GPUs (e.g., using NVLINK) or may connect the GPUs through a switch (e.g., using NVSwitch). When combined together, each GPU 908 may generate pixel data or GPGPU data for different parts of the output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU may include its own memory, or may share memory with other GPUs.
[0153] In addition to or in lieu of the CPU 906 and / or GPU 908, one or more logic units 920 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 900 to perform one or more of the methods and / or processes described herein. In an embodiment, one or more CPUs 906, one or more GPUs 908, and / or one or more logic units 920 may perform any combination of methods, processes, and / or portions thereof, separately or jointly. One or more of the logic units 920 may be a part of and / or integrated into one or more of the CPUs 906 and / or GPUs 908 and / or one or more of the logic units 920 may be a discrete component or otherwise external to the CPUs 906 and / or GPUs 908. In an embodiment, one or more of the logic units 920 may be a coprocessor of one or more of the one or more CPUs 906 and / or one or more of the one or more GPUs 908.
[0154] Examples of logic unit 920 include one or more processing cores and / or components thereof, such as a data processing unit (DPU), a tensor core (TC), a tensor processing unit (TPU), a pixel vision core (PVC), a vision processing unit (VPU), an image processing unit (IPU), a graphics processing cluster (GPC), a texture processing cluster (TPC), a streaming multiprocessor (SM), a tree traversal unit (TTU), an artificial intelligence accelerator (AIA), a deep learning accelerator (DLA), an arithmetic logic unit (ALU), an application-specific integrated circuit (ASIC), a floating point unit (FPU), an input / output (I / O) element, a peripheral component interconnect (PCI) or a peripheral component interconnect express (PCIe) element, etc.
[0155] The communication interface 910 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 900 to communicate with other computing devices via an electronic communication network (including wired and / or wireless communications). The communication interface 910 may include components and functions that enable communication over any of a plurality of different networks (e.g., wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., via Ethernet or infinite bandwidth communications), low power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet). In one or more embodiments, the logic unit 920 and / or the communication interface 910 may include one or more data processing units (DPUs) to transmit data received over a network and / or over the interconnect system 902 directly to one or more GPUs 908 (e.g., a memory of one or more GPUs 908). In some embodiments, multiple computing devices 900 or components thereof, which may be similar or different in various aspects from one another, may be communicatively coupled to send and receive data for performing the various operations described herein, such as to facilitate latency reduction.
[0156] The I / O ports 912 may allow the computing device 900 to be logically coupled to other devices including I / O components 914, presentation components 918, and / or other components, some of which may be built into (e.g., integrated into) the computing device 900. Illustrative I / O components 914 include microphones, mice, keyboards, joysticks, gamepads, game controllers, satellite dishes, scanners, printers, wireless devices, and the like. The I / O components 914 may provide a natural user interface (NUI) that processes air gestures, voice, or other physiological input generated by a user, such as generating prompts, image data 106, and / or video data 108. In some instances, the input may be transmitted to an appropriate network element for further processing, such as modifying and registering an image. The NUI may implement any combination of voice recognition, stylus recognition, facial recognition, biometric recognition, gesture recognition on and near the screen, air gestures, head and eye tracking, and touch recognition associated with a display of the computing device 900 (as described in more detail below). The computing device 900 may include a depth camera, such as a stereo camera system, an infrared camera system, an RGB camera system, touch screen technology, and combinations thereof, for gesture detection and recognition. In addition, the computing device 900 may include an accelerometer or gyroscope (e.g., as part of an inertial measurement unit (IMU)) that allows detection of motion. In some examples, the computing device 900 may use the output of the accelerometer or gyroscope to render immersive augmented reality or virtual reality.
[0157] The power supply 916 may include a hardwired power supply, a battery power supply, or a combination thereof. The power supply 916 may provide power to the computing device 900 to enable the components of the computing device 900 to operate.
[0158] One or more presentation components 918 may include a display (e.g., a monitor, a touch screen, a television screen, a head-up display (HUD), other display types, or combinations thereof), speakers, and / or other presentation components. Presentation component 918 may receive data from other components (e.g., GPU 908, CPU 906, DPU, etc.) and output data (e.g., as images, video, sound, etc.).
[0159] Sample Data Center
[0160] Fig.10 An example data center 1000 is shown that can be used in at least one embodiment of the present disclosure, such as implementing system 100 and / or system 200 in one or more examples of data center 1000. Data center 1000 can include a data center infrastructure layer 1010, a framework layer 1020, a software layer 1030, and / or an application layer 1040.
[0161] like Fig.10 As shown, the data center infrastructure layer 1010 may include a resource coordinator 1012, grouped computing resources 1014, and node computing resources ("node CRs") 1016(1)-1016(N), where "N" represents any integer, positive integer. In at least one embodiment, the node CRs 1016(1)-1016(N) may include, but are not limited to, any number of central processing units ("CPUs") or other processors (including DPUs, accelerators, field programmable gate arrays (FPGAs), graphics processors or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memories), storage devices (e.g., solid-state or disk drives), network input / output ("NW I / O”) devices, network switches, virtual machines (“VMs”), power modules and / or cooling modules, etc. In some embodiments, one or more of the node CRs 1016(1)-1016(N) may correspond to a server having one or more of the above-mentioned computing resources. In addition, in some embodiments, the node CRs 1016(1)-1016(N) may include one or more virtual components, such as vGPUs, vCPUs, etc., and / or one or more of the node CRs 1016(1)-1016(N) may correspond to virtual machines (VMs).
[0162] In at least one embodiment, the grouped computing resources 1014 may include separate groups of nodes CR1016 contained in one or more racks (not shown) or in many racks contained in data centers (also not shown) at different geographical locations. The separate groups of nodes CR1016 in the grouped computing resources 1014 may include grouped computing, network, memory or storage resources, which may be configured or allocated to support one or more workloads. In at least one embodiment, several nodes CR1016 including CPUs, GPUs, DPUs and / or other processors may be grouped in one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any number of power modules, cooling modules and / or network switches in any combination.
[0163] Resource coordinator 1012 may configure or otherwise control one or more nodes CR 1016(1)-1016(N) and / or grouped computing resources 1014. In at least one embodiment, resource coordinator 1012 may comprise a software design infrastructure ("SDI") management entity for data center 1000. Resource coordinator 1012 may comprise hardware, software, or some combination thereof.
[0164] In at least one embodiment, Fig.10 As shown, the framework layer 1020 may include a job scheduler 1028, a configuration manager 1034, a resource manager 1036, and / or a distributed file system 1038. The framework layer 1020 may include a framework that supports software 1032 of the software layer 1030 and / or one or more applications 1042 of the application layer 1040. The software 1032 or the application 1042 may include network-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 1020 may be, but is not limited to, a type of free open source software web application framework that can utilize the distributed file system 1038 for large-scale data processing (e.g., "big data"), such as Apache Spark. TM(hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 1028 may include a Spark driver to facilitate the scheduling of workloads supported by the various layers of the data center 1000. The configuration manager 1034 may be able to configure different layers, such as the software layer 1030 and the framework layer 1020 including Spark and a distributed file system 1038 for supporting large-scale data processing. The resource manager 1036 may be able to manage clustered or grouped computing resources mapped to or allocated to support the distributed file system 1038 and the job scheduler 1028. In at least one embodiment, the clustered or grouped computing resources may include grouped computing resources 1014 at the data center infrastructure layer 1010. The resource manager 1036 may coordinate with the resource coordinator 1012 to manage these mapped or allocated computing resources.
[0165] In at least one embodiment, software 1032 included in software layer 1030 may include software used by at least portions of node CRs 1016(1)-1016(N), grouped computing resources 1014, and / or distributed file system 1038 of framework layer 1020. One or more types of software may include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.
[0166] In at least one embodiment, the applications 1042 included in the application layer 1040 may include one or more types of applications used by the nodes CRs 1016(1)-1016(N), the grouped computing resources 1114, and / or at least portions of the distributed file system 1038 of the framework layer 1020. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications (including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments), such as training, configuring, updating, and / or executing the machine learning models 104, 204.
[0167] In at least one embodiment, any of configuration manager 1034, resource manager 1036, and resource coordinator 1012 can implement any number and type of self-modification actions based on any number and type of data acquired in any technically feasible manner. The self-modification actions can save a data center operator of data center 1000 from making potentially poor configuration decisions and potentially avoiding underutilized and / or underperforming portions of a data center.
[0168] According to one or more embodiments described herein, data center 1000 may include tools, services, software, or other resources to train one or more machine learning models (e.g., implement machine learning model 112) or use one or more machine learning models to predict or infer information (e.g., generate scene representation 124, motion generator 128, and / or content model 204). For example, one or more machine learning models may be trained by computing weight parameters according to a neural network architecture using the software and / or computing resources described above with respect to data center 1000. In at least one embodiment, a trained or deployed machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 1000 using weight parameters computed by one or more training techniques (such as, but not limited to, those described herein).
[0169] In at least one embodiment, the data center 1000 may use a CPU, an application-specific integrated circuit (ASIC), a GPU, an FPGA, and / or other hardware (or virtual computing resources corresponding thereto) to perform training and / or reasoning using the above resources. In addition, the above one or more software and / or hardware resources may be configured to allow users to train or perform information reasoning services, such as image recognition, speech recognition, or other artificial intelligence services.
[0170] Example network environment
[0171] A network environment suitable for implementing embodiments of the present disclosure may include one or more client devices, servers, network attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be configured to communicate with the client device, server, or other device types. Fig. 9 The data center 1000 may be implemented on one or more instances of the computing device 900, for example, each device may include similar components, features and / or functions of the computing device 900. In addition, in the case of implementing a backend device (e.g., a server, NAS, etc.), the backend device may be included as part of the data center 1000, and the example of the data center 1000 is referred to in this article with respect to Fig.10 Describe in more detail.
[0172] The components of the network environment can communicate with each other via one or more networks that can be wired, wireless, or both. The network can include multiple networks or networks of networks. As an example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (such as the Internet and / or a public switched telephone network (PSTN)), and / or one or more private networks. In the case where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connections.
[0173] Compatible network environments may include one or more peer-to-peer network environments, in which case the network environment may not include a server, and one or more client-server network environments, in which case the network environment may include one or more servers. In a peer-to-peer network environment, the functionality described herein with respect to one or more servers may be implemented on any number of client devices.
[0174] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, and the like. A cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more of the servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework for one or more applications of the software supporting the software layer and / or the application layer. The software or application may include network-based service software or applications, respectively. In an embodiment, one or more of the client devices may use web-based service software or applications (e.g., by accessing the service software and / or application via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, the type of free and open source software web application framework, such as a distributed file system that may be used for large-scale data processing (e.g., "big data").
[0175] A cloud-based network environment can provide cloud computing and / or cloud storage that performs any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these different functions can be distributed in multiple locations from a central or core server (e.g., one or more data centers that can be distributed across states, regions, countries, the world, etc.). If the connection to the user (e.g., client device) is relatively close to the edge server, the core server can assign at least a portion of the function to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0176] One or more client devices may include the Fig. 9 At least some of the components, features, and functionality of one or more of the described example computing devices 900. By way of example and not limitation, the client device may be embodied as a personal computer (PC), a laptop computer, a mobile device, a smart phone, a tablet computer, a smart watch, a wearable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality headset, a global positioning system (GPS) or device, a video player, a camera, a surveillance device or system, a vehicle, a vessel, a spacecraft, a virtual machine, a drone, a robot, a handheld communication device, a hospital device, a gaming device or system, an entertainment system, a vehicle computer system, an embedded system controller, a remote control, an appliance, a consumer electronic device, a workstation, an edge device, any combination of these depicted devices, or any other suitable device.
[0177] The present disclosure may be described in the general context of computer code or machine-usable instructions (including computer-executable instructions, such as program modules) executed by a computer or other machine, such as a personal data assistant or other handheld device. In general, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs a particular task or implements a particular abstract data type. The present disclosure may be implemented in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The present disclosure may also be implemented in a distributed computing environment where tasks are performed by remote processing devices linked through a communications network.
[0178] As used herein, the description of "and / or" about two or more elements should be interpreted as meaning only one element, or a combination of elements. For example, "element A, element B, and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. In addition, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. In addition, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0179] The subject matter of the present disclosure is described in detail herein to meet statutory requirements. However, the description itself is not intended to limit the scope of the present disclosure. On the contrary, the inventors have anticipated that the claimed subject matter may also be embodied in other ways in conjunction with other current or future technologies to include different steps or combinations of steps similar to the steps described in this document. In addition, although the terms "step" and / or "box" may be used herein to refer to different elements of the method employed, unless and in addition to the order of the individual steps being explicitly described, the terms should not be interpreted as implying any particular order among or between the various steps disclosed herein.
Claims
1. One or more processors, including: One or more circuits for: assigning a plurality of first elements of the three-dimensional 3D model of the subject to a plurality of locations on the surface of the subject in the initial pose; assigning a plurality of second elements to the plurality of first elements, each second element of the plurality of second elements having an opacity corresponding to a distance between the second element and a surface of the subject; updating the plurality of second elements based at least on a target pose of the subject and one or more attributes of the subject to determine a plurality of updated second elements; as well as A representation of the body is rendered based at least on the plurality of updated second elements.
2. One or more processors according to claim 1, wherein: The one or more circuits are configured to update the plurality of updated second elements based at least on an evaluation of one or more objective functions and a representation of the subject.
3. One or more processors according to claim 1, wherein: The one or more circuits are configured to determine an opacity of each second element of the plurality of second elements using a signed distance function to represent a distance between the second element and a surface of the subject.
4. The one or more processors of claim 1, wherein: At least one first element of the plurality of first elements includes at least one of the following: a position parameter corresponding to a position in the plurality of positions to which the at least one element is assigned; a scale parameter indicating a scale of the at least one first element in a 3D reference system in which the subject is located; or An orientation parameter indicating the orientation of the body relative to the 3D reference system.
5. The one or more processors of claim 1, wherein: At least one second element of the plurality of second elements includes a 3D Gaussian sputtering defined in a local reference frame of at least one corresponding first element of the plurality of first elements.
6. The one or more processors of claim 1, wherein: The one or more circuits are configured to receive an indication of the one or more attributes of the subject as at least one of text data, voice data, audio data, image data, or video data.
7. The one or more processors of claim 1, wherein: The one or more processors are configured to regularize the plurality of second elements based at least on position data of each second element of the plurality of second elements.
8. The one or more processors of claim 1, wherein: The one or more processors are configured to update the plurality of updated second elements based at least on a mask determined from the representation of the subject and an alpha rendering determined from the plurality of updated second elements.
9. The one or more processors of claim 1, wherein: The one or more circuits are configured to generate the representation to include a textured mesh of the body.
10. The one or more processors of claim 1, wherein the one or more processors are included in at least one of: Systems for generating synthetic data; A system for performing simulation operations; A system for performing collaborative content creation of 3D assets; Systems for performing conversational AI operations; A system including one or more Large Language Models (LLMs); A system comprising one or more visual language models (VLMs); Systems for performing digital twin operations; A system for performing light transport simulations; Systems for performing deep learning operations; Systems implemented using edge devices; Systems implemented using robots; control systems for autonomous or semi-autonomous machines; Perception systems for autonomous or semi-autonomous machines; A system comprising one or more virtual machines VM; A system implemented at least in part in a data center; or A system implemented at least in part using cloud computing resources.
11. A system comprising: One or more processors configured to perform operations including: assigning a plurality of first elements of the three-dimensional 3D model of the subject to a plurality of locations on the surface of the subject in the initial pose; assigning a plurality of second elements to the plurality of first elements, each second element of the plurality of second elements having an opacity corresponding to a distance between the second element and a surface of the subject; updating the plurality of second elements based at least on a target pose of the subject and one or more attributes of the subject to determine a plurality of updated second elements; as well as A representation of the body is rendered based at least on the plurality of updated second elements.
12. The system according to claim 11, wherein: The one or more processors are configured to update the plurality of updated second elements based at least on an evaluation of one or more objective functions and a representation of the subject.
13. The system according to claim 11, wherein: The one or more processors are configured to determine an opacity of each second element of the plurality of second elements using a signed distance function to represent a distance between the second element and a surface of the subject.
14. The system according to claim 11, wherein: At least one first element of the plurality of first elements includes at least one of the following: a position parameter corresponding to a position in the plurality of positions to which the at least one element is assigned; a scale parameter indicating a scale of the at least one first element in a 3D reference system in which the subject is located; or An orientation parameter indicating the orientation of the body relative to the 3D reference system.
15. The system according to claim 11, wherein: At least one second element of the plurality of second elements includes a 3D Gaussian sputtering defined in a local reference frame of a corresponding at least one first element of the plurality of first elements.
16. The system according to claim 11, wherein: The one or more processors are configured to receive an indication of the one or more attributes of the subject as at least one of text data, voice data, audio data, image data, or video data.
17. The system of claim 11, wherein: The one or more processors are configured to regularize the plurality of second elements based at least on position data of each second element of the plurality of second elements.
18. The system of claim 11, wherein: The one or more processors are configured to update the plurality of updated second elements based at least on a mask determined from a representation of the subject and an alpha rendering determined from the plurality of updated second elements, and wherein the one or more processors are configured to generate the representation to include a textured mesh of the subject.
19. A method comprising: assigning, using one or more processors, a plurality of first elements of a three-dimensional (3D) model of a subject to a plurality of locations on a surface of the subject in an initial pose; assigning, using the one or more processors, a plurality of second elements to the plurality of first elements, each second element of the plurality of second elements having an opacity corresponding to a distance between the second element and a surface of the subject; updating, using the one or more processors, the plurality of second elements based at least on a target pose of the subject and one or more attributes of the subject to determine a plurality of updated second elements; as well as A representation of the subject is rendered using the one or more processors based at least on the plurality of updated second elements.
20. The method according to claim 19, wherein: The updating of the plurality of updated second elements is based at least on an evaluation of one or more objective functions and a representation of the subject, and wherein the determination of the opacity of each of the plurality of second elements comprises representing a distance between the second element and a surface of the subject using a signed distance function.
Citation Information
Cited By
Digital human model generation method and device
CN120931780A
Digital human model generation method and apparatus
CN120931780B