Synthetic human motion generation using neural networks for scene awareness
By combining the pre-trained motion diffusion model and scene perception components, more accurate scene-aware human motion is generated, which solves the problem of generating realistic motion in a variety of 3D scenes in the prior art, and achieves efficient motion generation in the absence of high-quality data.
Patent Information
- Application Number
- CN202510071684.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-17
- Filing Date
- 2025-01-16
- Publication Date
- 2025-07-18
AI Technical Summary
Existing human motion generation techniques are difficult to generate realistic motion in a variety of 3D scenarios, especially in the absence of high-quality training data, and the existing techniques often fail to capture the full scope and subtleties of human behavior.
By pre-training the motion diffusion model and combining scene perception components, neural networks are used to extract the representation of scene information to generate more accurate scene-aware human motion.
Generate more accurate and scene-aware human movements on limited motion-scene data, improving the authenticity and adaptability of motion generation.
Smart Images

Figure CN120340108A_ABST
Abstract
Description
Background Art
[0001] Human character motion (or simply, human motion) generation generally aims to create realistic and natural movements for virtual or animated characters that mimic the way humans move in the real world. For example, human motion generation may attempt to simulate the complex interactions of joints, muscles, and / or physical constraints to generate lifelike animations. Human motion generation typically plays a central role in computer graphics, animation, and / or virtual reality applications, as it can add a layer of authenticity and immersion to digital experiences in various industries and applications, such as video games, film and television production, simulation training, healthcare (e.g., for physical therapy simulation), and / or other scenarios. For example, in the entertainment industry, human motion generation can enable the creation of compelling and believable characters, thereby enhancing the overall viewing experience. In the context of training or design simulations, human motion generation can allow professionals to practice or design in a controlled environment without real-world risks. In the healthcare field, human motion generation can assist in rehabilitation and recovery by providing patients with interactive exercises tailored to their specific needs. These are just a few examples of how human motion generation can help bridge the gap between the digital and physical worlds.
[0002] Conventional synthetic human motion generation techniques have various drawbacks. For example, some techniques may attempt to generate human motion (e.g., character animation) in a specific three-dimensional (3D) scene based on an input text prompt that provides a certain type of instruction (e.g., "sit on the couch"). Generally, the goal is to generate physically realistic motion in terms of navigating the 3D scene (e.g., avoiding collisions when navigating around furniture) and interacting with objects in the scene (e.g., humans typically sit facing forward in a chair rather than sideways). However, it is difficult for conventional techniques to generate realistic motion for many 3D scenes. For example, conventional human motion generation techniques typically require high-quality training data that pairs captured human motion with the corresponding 3D scene and interactions in the 3D scene. The generation of this type of dataset can be very challenging and costly (e.g., requiring high-quality motion capture for specific characters, actions, objects, and / or 3D scenes of interest). Therefore, this type of training data is typically limited, so conventional models trained on certain characters, actions, objects, and / or 3D scenes generally do not generalize to other scenes, resulting in unrealistic and / or low-quality motion animations. Some techniques attempt to address this problem by placing high-quality motion capture sequences (captured without the environment) into a scanned scene environment. However, the synthetic motion obtained by these techniques often does not reflect reasonable human behavior in the real world. Finally, one conventional technique attempts to use reinforcement learning to address the lack of suitable training data. Reinforcement learning does not require any paired motion-scene data, but instead trains different strategies for each type of supported interaction. However, restricting the generated motion to specifically considered human interactions is unlikely to capture the full range and nuances of potential human motion. Therefore, improved human motion generation techniques are needed. SUMMARY OF THE INVENTION
[0003] Embodiments of the present disclosure relate to scene-aware human motion generation. Systems and methods are disclosed that pre-train a base motion diffusion model without scene information, connect a scene-aware component, and tune the resulting motion diffusion model on data with scene information.
[0004] Compared with the conventional systems described above, the motion diffusion model can be pre-trained on motion data, and a scene perception component (e.g., one or more layers of a neural network) can be connected and used to extract a representation of the scene information and inject it into the pre-trained motion diffusion model. For example, to predict the orientation of joint waypoints along a path in a specific 3D scene, a scene perception input channel that receives a representation of the 3D structure of the scene can be added to the pre-trained motion diffusion model. To predict the orientation of joint waypoints along a path that interacts with a 3D object in the 3D scene, a scene perception input channel that receives a representation of the 3D object and / or the surface of the 3D object can be added to the pre-trained motion diffusion model. Thus, the resulting scene perception motion diffusion model can be tuned on motion-scene data and used to generate human motion. Therefore, the techniques described herein can be used to generate scene-aware human motion of a character based on a representation of a 3D scene and / or a target 3D object with which the character is to interact. By combining the scene perception component with the underlying pre-trained motion diffusion model, the resulting scene perception motion diffusion model can be fine-tuned on a limited motion-scene data set, enabling the generation of more accurate and scene-aware human motion on much less motion-scene data than the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] The present system and method for scene-aware human motion generation will be described in detail below with reference to the accompanying drawings, in which:
[0006] Figure 1 is a block diagram of an example motion generation pipeline according to some embodiments of the present disclosure;
[0007] Figure 2 is a block diagram of an example scene perception motion diffusion model for a scene navigation component according to some embodiments of the present disclosure;
[0008] Figure 3 is a block diagram of an example scene perception motion diffusion model for a scene interaction component according to some embodiments of the present disclosure;
[0009] Figure 4 is a flowchart showing a method of generating a representation of scene-aware motion according to some embodiments of the present disclosure;
[0010] Figure 5 is a flowchart showing a method of generating a motion diffusion model according to some embodiments of the present disclosure;
[0011] Figure 6 is a block diagram of an example computing device suitable for implementing some embodiments of the present disclosure; and
[0012] Figure 7 It is a block diagram of an example data center suitable for implementing some embodiments of the present disclosure. Detailed implementation
[0013] Systems and methods related to scene-aware human motion generation are disclosed. In some embodiments, a diffusion model (e.g., a motion diffusion model) can be pre-trained on motion data and used as a base model, and a scene-aware component (e.g., one or more layers of a neural network) is used to extract a representation of scene information and inject the representation of the scene information into the pre-trained motion diffusion model. For example, to predict the orientation of joint (e.g., root joint) path points along a path in a specific 3D scene, a scene-aware input channel that receives a representation of the 3D structure of the scene (e.g., a two-dimensional (2D) occupancy grid or a 3D occupancy grid, a floor plan, a height map, semantic segmentation, etc.) can be added to the pre-trained motion diffusion model. In another example, to predict the orientation of joint (e.g., root joint) path points along a path that interacts with a 3D object in the 3D scene, a scene-aware input channel that receives a representation of the 3D object and / or the surface of the 3D object (e.g., a 3D point cloud) can be added to the pre-trained (e.g., motion) diffusion model. Thus, the resulting scene-aware motion diffusion model can be adjusted (e.g., fine-tuned) on motion-scene data and used to generate human motion. Compared with the prior art, the present technology can be used to generate more accurate and scene-aware human motion on much less motion-scene data than the prior art.
[0014] For example, given a representation of a 3D scene, a starting point, instructions (e.g., a text prompt such as an instruction for a character to "sit on a couch"), classification data (such as semantic segmentation of the 3D scene), and / or other inputs, any known path planning technique can be used to identify a target point to which a character in the 3D scene should move, identify a path to reach the target point (e.g., a chair) in the 3D scene and avoid collisions (e.g., with other furniture), identify a path to achieve an interaction with a target object in the 3D scene (e.g., sit on a chair), and / or identify one or more contact points between the character and the target object. Each path can be in the form of a 2D or 3D sequence of path points (or path points can be sampled along the path). In some embodiments, the sequence of path points can represent (or be used to generate corresponding) the successive 2D or 3D positions of one or more joints (e.g., the root joint) of the character being animated. These path points can be used as input to one or more diffusion models to predict the orientation of the corresponding joint at that path point. Additionally or alternatively, to identify a path to reach a target point and / or a path to achieve an interaction with a target object before predicting the orientation of a joint at a path point along the path, any known planning technique can also be used to identify the target point and / or one or more contact points between the character and the target object, and noisy intermediate path points can be used as input to one or more diffusion models to predict the position and orientation of the corresponding joint at that path point, thus effectively predicting the path and the pose along the path (e.g., the given positions of the starting point, the target point, and / or one or more contact points).
[0015] For example, to predict joint directions (and / or positions) for a motion sequence represented by a sequence of path points traversing a 2D or 3D scene, a scene-aware diffusion model can encode an instruction (e.g., a text instruction), the positions and noisy directions of the sequence of path points (and / or, if predicting the corresponding path point positions, the noisy positions of the path points), and a representation of the 2D or 3D structure of at least a portion of the 3D scene (e.g., a 2D or 3D occupancy grid, a floor plan, a height map, a patch of one of the foregoing (such as an egocentric patch), classification data (such as semantic segmentation representing objects or other parts of the scene in any number of classes), etc.). In some embodiments, the classification data can include layers for each object class in one or more object classes, such as furniture (e.g., different types), doors, windows, appliances, walls, lighting fixtures, power outlets and switches, electronic devices, personal items, other characters, audio sources, and / or other things in the scene. Thus, the scene-aware diffusion model can combine these encoded inputs to predict a denoised scene-aware motion sequence.
[0016] For example, a scene-aware diffusion model can iteratively predict and refine a denoised motion sequence based on the 2D structure or 3D structure of a 3D scene through a series of diffusion steps. In each diffusion step, the scene-aware diffusion model (e.g., a transformer-based model) can predict a denoised motion sequence based on the 2D structure or 3D structure of the 3D scene and diffuse the predicted motion sequence back to the previous diffusion step, thus effectively updating the state of the denoised motion sequence in the reverse order from the final diffusion step to the initial diffusion step based on the scene structure. By starting from the most refined representation of the scene-aware motion and diffusing it back to the previous step, the denoised scene-aware motion sequence predicted in each diffusion step benefits from the improvements accumulated in subsequent steps, improving the scene-aware temporal dependencies, where later states can be affected by earlier states, and providing an opportunity to correct any errors or inaccuracies introduced in the earlier steps, thereby generating a more accurate and realistic scene-aware motion sequence.
[0017] In some embodiments, to predict the joint orientations (and / or positions) of a motion sequence represented by a sequence of path points along a path of interacting with a target object in a 3D scene (e.g., sitting on a chair), the scene-aware diffusion model can encode an instruction (e.g., a text instruction), the positions and noisy directions of the sequence of path points (and / or the noisy positions of the path points if predicting the corresponding path point positions), the 3D structure (e.g., a 3D point cloud) of the 3D object or its surface, and the representation of one or more contact positions on the 3D object (e.g., the positions where the arm or pelvis contacts the chair, whether pre-determined during path planning using any known technique or noisy contact positions to be predicted by the scene-aware motion diffusion model). Thus, the scene-aware diffusion model can combine these encoded inputs to predict a denoised scene-aware motion sequence. In some embodiments, the scene-aware diffusion model can iteratively predict and refine a denoised motion sequence based on the 3D structure of the 3D object it interacts with through a series of diffusion steps. In each diffusion step, the scene-aware diffusion model (e.g., a transformer) can predict a denoised motion sequence based on the 3D structure of the 3D object and diffuse the predicted motion sequence back to the previous step, thus effectively updating the state of the scene-aware motion sequence in the reverse order based on the structure of the object in the scene, effectively incorporating the accumulated improvements, improving the scene-aware (e.g., object-aware) temporal dependencies, and providing an opportunity to correct any errors or inaccuracies introduced in the earlier steps, thereby generating a more accurate and realistic scene-aware motion sequence.
[0018] In some embodiments, data augmentation can be used to generate training data that pairs (e.g., captured) motion data with the corresponding 3D object being interacted with, to retarget the initially captured motion about one object onto another object (e.g., retarget the captured motion data of sitting on a particular chair onto a different chair). Compared with the prior art that retargets the contact positions (e.g., the positions where the arm or pelvis touches the chair) to the target positions where the corresponding joints of the skeletal structure touch the target object, in some embodiments, a 3D model (e.g., 3D mesh) can be used to model the body surface structure of the character, so that the contact positions can be retargeted to the target positions where the corresponding positions on the body surface touch the target object. Therefore, the resulting retargeted motion-object interaction data is more accurate than the prior art, and training a scene-aware diffusion model (e.g., fine-tuning a pre-trained diffusion model) using this training data can improve the accuracy of the resulting generated motion.
[0019] In an example training embodiment, a pre-trained base diffusion model can be adjusted or otherwise adapted to the scene content using fine-tuning (e.g., freezing one or more layers of the pre-trained model), parameter-efficient fine-tuning (PEFT) (e.g., low-rank adaptation (LoRA), prefix tuning, prompt tuning, p-tuning), some other techniques for updating one or more trainable parameters (e.g., network weights, rank decomposition matrices, hard prompts, soft prompts), and / or other means. For example, the adjustment may involve adding one or more scene-aware layers that extract representations of scene information and inject them into the pre-trained base diffusion model, and using the (e.g., retargeted) motion-object data to train the resulting model (e.g., fixing one or more pre-trained layers of the pre-trained base diffusion model) to learn the corresponding weights of the added layers.
[0020] Thus, the techniques described herein can be used to generate scene-aware human motions for a character based on a representation of a 3D scene and / or a representation of the target 3D object with which the character is interacting. By combining the scene-aware component with the base pre-trained diffusion model, the resulting scene-aware diffusion model can be fine-tuned on much more limited motion-scene data, enabling the generation of more accurate and scene-aware human motions on much less motion-scene data than the prior art.
[0021] Refer to Figure 1 , Figure 1An example motion generation pipeline 100 in accordance with some embodiments of the present disclosure. It should be understood that such and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, orders, function groupings, etc.) may be used in addition to or in place of the shown arrangements and elements, and some elements may be omitted altogether. Moreover, many of the elements described herein are functional entities that may be implemented as discrete or distributed components or in combination with other components and implemented in any suitable combination and location. The various functions described herein as being performed by entities may be performed by hardware, firmware, and / or software. For example, the various functions may be implemented by a processor executing instructions stored in a memory.
[0022] In Figure 1 In the example shown, the motion generation pipeline 100 includes a planning component 120 that receives a text cue 105, a representation of a 3D scene 110, and classification data 115 representing one or more aspects of the 3D scene 110. In some embodiments, the planning component 120 uses one or more of these inputs to identify target points in the 3D scene, where a character in the 3D scene 110 should move to one or more positions of one or more path points along a path to the target point. Thus, the scene navigation component 132 can use the known positions of the starting point, target point, and / or path points to predict the orientation of the corresponding joints at the path points. In some embodiments, the planning component 120 predicts the path (e.g., the positions of one or more path points), and the scene navigation component 132 predicts the pose (e.g., the orientation of the corresponding joints) at the path points along the path. In some embodiments, the planning component 120 predicts the target point, and the scene navigation component 132 predicts the path (e.g., the positions of one or more path points) and the pose (e.g., the orientation of the corresponding joints) at the path points along the path.
[0023] In some embodiments, the text cue 105 may represent instructions for a character initially located at a starting point 126 to interact with a target object (e.g., a chair 124) in the 3D scene 110. Thus, the planning component 120 can identify the target object from the text cue 105, identify a first target point 128 to which the character in the 3D scene 110 can move before interacting with the target object, identify the corresponding interaction (e.g., sitting on the chair 124), identify a second target point 130 to which the character can move via the interaction, and / or identify one or more contact points between the character and the target object. As an addition to or alternative to the scene navigation component 132 predicting the path 127 between the starting point 126 and the first target point 128 and / or the pose (e.g., the motion sequence 135) along the path 127, the scene interaction component 140 can predict the path 129 between the first target point 128 and the second target point 130 and / or the pose (e.g., the motion sequence 145) along the path 129.
[0024] Typically, according to embodiments, there can be multiple inputs. For example, the motion generation pipeline 100 can be incorporated into or triggered by the user interface of a character animation, robotics, and / or other types of applications that generate representations of motion and / or animate motion, and the user interface can accept one or more user inputs that represent instructions for a character, robot, or other entity to move and / or interact with the 3D scene 110. In Figure 1 the illustrated embodiment, the instructions are embodied in the text prompt 105 in natural language, but this is not necessary. Additionally or alternatively, the user interface can accept and / or encode instructions represented by voice commands, detected gestures, joystick or gamepad inputs, virtual or augmented reality controllers, spatial coordinates identified via the user interface, and / or other types of inputs.
[0025] In embodiments that include the text prompt 105, the planning component 120 can use any known technique to evaluate the text prompt 105, a representation of the current state of the 3D scene 110 (e.g., a 2D occupancy grid or a 3D occupancy grid, a floor plan, a height map), and / or the corresponding classification data 115 (e.g., semantic segmentation of the 3D scene 110) to identify one or more target points for the character to move in the 3D scene 110 and / or the corresponding target orientations at each target point. For example, the planning component 120 can use natural language processing (e.g., named entity recognition, keyword extraction) to identify and extract relevant spatial information, target objects in the 3D scene 110, and / or orientation cues referenced in the text prompt 105. Additionally or alternatively, the planning component 120 can use one or more machine learning models (e.g., one or more language models) to interpret the text prompt 105 and infer the expected target locations and / or target orientations. In some embodiments, the planning component 120 can use any known technique to evaluate the text prompt 105, the 3D scene 110, and / or the classification data 115 to generate a path through the 3D scene 110 (e.g., path 127) and / or a path to achieve the interaction specified in the text prompt 105 (e.g., path 129). For example, the planning component 120 can use a pathfinding algorithm (e.g., A*) to generate a path that avoids obstacles in the 3D scene 110. Note that Figure 1 an embodiment where the planning component 120 generates the path 127 and the path 129 is shown, but this is not necessary.
[0026] Figure 2 is a block diagram of an example scene-aware diffusion model 200 for a scene navigation component according to some embodiments of the present disclosure. For example, Figure 1The scene navigation component 132 can use the scene-aware diffusion model 200 to predict a motion sequence that includes joint orientations (and / or positions) of a sequence of path points along a path from a starting point 126 to a first target point 128 through a 3D scene 110.
[0027] In some embodiments, the scene-aware diffusion model 200 can be implemented using a neural network. Although the scene-aware diffusion model 200 and other models and functions described herein can be implemented using a neural network (or a part thereof), this is not intended to be limiting. Generally, the models and / or functions described herein can be implemented using any of a variety of different networks or machine learning models, such as machine learning models using linear regression, logistic regression, decision trees, support vector machines (SVMs), naive Bayes, k-nearest neighbors (Knn), k-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoders, convolutional, transformers, recurrent, perceptrons, long / short-term memory (LSTM), Hopfield, Boltzmann, deep belief, deconvolution, generative adversarial, liquid machines, etc.) and / or other types of machine learning models.
[0028] Generally, a path can be represented as a sequence of 2D or 3D path points x 1 …x N and the scene-aware diffusion model 200 can iteratively predict and refine the denoised motion sequence over a series of t diffusion steps. For example, the scene navigation component 132 can initially construct a representation of the sequence using one or more data structures representing the position (e.g., 3D position, 2D ground projection), orientation, and / or other features of one or more joints (e.g., a root joint such as the pelvis) of a character at each path point, filling known parameters (e.g., the position and orientation of the starting point x 1 and the target point x N ) in the corresponding elements of the one or more data structures and filling the remaining elements with random noise (e.g., unknowns to be predicted). Thus, the scene navigation component 132 can generate a representation of the motion sequence at a particular diffusion step t and apply it to the scene-aware motion diffusion model 200 to predict the denoised motion sequence Figure 1 FIG. shows an embodiment in which the scene navigation component 132 starts from the last diffusion step t = T and uses the scene-aware diffusion model 200 to predict the denoised motion sequence at the first diffusion step t = 0. After this first iteration, the scene navigation component 132 can diffuse the denoised motion sequence to a state prior to the last diffusion step t = T - 1 (at Figure 2a state corresponding to the result of the previous step prediction (shown in the figure) and use it as input for the scene-aware motion diffusion model 200 to predict again the denoised motion sequence at the first diffusion step t = 0 The scene navigation component 132 can repeat this process, thereby effectively updating the state of the denoised motion sequence in the reverse order from the final diffusion step to the initial diffusion step. This reverse diffusion process is only intended as an example, and other diffusion techniques with any number and order of diffusion steps can be implemented within the scope of the present disclosure.
[0029] In some embodiments, the scene-aware diffusion model 200 can include a base diffusion model (e.g., including layer 210, transformer encoder 230, and layer 240), which can be pre-trained on motion data (e.g., input motion sequence and output motion sequence) without scene data using any known technique. Figure 2 In the illustrated embodiment, the scene-aware diffusion model 200 further includes a scene-aware component, which includes layer 220 and transformer encoder 235. The transformer encoder can be integrated with or connected to the pre-trained base diffusion model, thereby forming an input channel for scene data 215.
[0030] In some embodiments, the scene navigation component 132 can generate scene data 215, which can represent one or more features of the 3D scene 110 and can take any suitable form. For example, the scene data 215 can represent the 2D or 3D structure of at least a part of the 3D scene 110 (e.g., 2D or 3D occupancy grid, floor plan, height map, patches of the foregoing, such as egocentric patches, classification data, such as semantic segmentation representing any number of classes of objects or other parts of the scene, etc.). In some embodiments, the classification data can include layers for each of one or more classes of objects, such as (for example, different types of) furniture, doors, windows, appliances, walls, lighting fixtures, power outlets and switches, electronic devices, personal items, other characters, audio sources, and / or other things that may be present in the scene. Thus, the scene navigation component 132 can generate a representation of the scene data 215 and apply the scene data 215 to layer 220 to extract a set of features (e.g., encoded 2D feature map) representing the scene data 215.
[0031] To inject the scene data 215 into the pre-trained base diffusion model, the scene-aware diffusion model 200 can generate sampled features by sampling the encoded feature map at elements corresponding to the positions of the (e.g., 2D) path points of the motion sequence at a specific diffusion step t and use it as input for the scene-aware motion diffusion model 200 to predict again the denoised motion sequence at the first diffusion step t = 0 (For example, these features can be sampled at the location being denoised, which can change at each diffusion step t based on the predicted denoised motion sequence at the previous step t+1).
[0032] Thus, the scene-aware diffusion model 200 can use the motion sequence at a specific diffusion step t and the corresponding sampled features of the scene data 215 to predict the denoised motion sequence For example, the scene-aware diffusion model 200 can encode Figure 1 a representation of the text prompt 105 (the diffusion step is iterating its representation) and / or some other input, and use the encoded representation as a conditioning input 205. The layer 210 can be used to resize the input representation of the motion sequence to a certain specified dimension of the transformer encoder 230. For example, each waypoint can be represented in 3D (e.g., x, y, heading angle), and the layer 210 can be used to resize each waypoint to a larger dimension (e.g., 64 or 128), such as a dimension corresponding to the dimension of the sampled features of the encoded feature map of the scene data 215. The positional encoding 225 can be used to encode the representation of the time step and merge it (e.g., add) into each corresponding waypoint or sampled scene feature (e.g., merge it into the corresponding token representing each corresponding waypoint or sampled scene feature). Each of the transformer encoder 230 and the transformer encoder 235 can include any number of layers, and each layer can include any number of attention heads that effectively weight different elements in the input based on importance.
[0033] The transformer encoder 235 can be connected to the transformer encoder 230 in any suitable way. In an example implementation, the transformer encoder 235 can be connected to the transformer encoder 230 via one or more (e.g., linear) layers followed by an addition layer (e.g., residual or skip connection). For example, the transformer encoder 235 can output an encoded representation of the scene data 215, the encoded representation of the scene data 215 can be processed using one or more linear layers (shown as arrows between the transformer encoder 235 and the transformer encoder 230 in Figure 2 ), and the result can be added to the intermediate features of the transformer encoder 230 (e.g., the pre-trained base diffusion transformer).
[0034] In some embodiments, a base diffusion model (e.g., layer 210, transformer encoder 230, and layer 240) can be pre-trained on motion data without scene data, and the scene-aware diffusion model 200 can be adjusted (e.g., fine-tuned) on motion-scene data (e.g., input training data including a noisy motion sequence and scene data 215, and a corresponding ground-truth denoised motion sequence). For example, the scene-aware components of the scene-aware motion diffusion model 200 (e.g., layer 220, transformer encoder 235, one or more layers connecting transformer encoder 235 to transformer encoder 230) can be initialized (e.g., with initial values of trainable parameters, such as weights set to zero), the base diffusion model can be frozen or locked ( Figure 2 shown by a padlock in
[0035] ), and paired motion-scene training data can be used to train the scene-aware diffusion model 200 to learn the values of the trainable parameters of the scene-aware components. Generally, any known motion-scene dataset can be used and / or any technique can be used to generate a motion-scene dataset with input and ground-truth training data for the scene-aware diffusion model 200. However, by combining the scene-aware components with a base pre-trained motion diffusion model, the resulting scene-aware diffusion model 200 can be fine-tuned on more limited motion-scene data, enabling more accurate and scene-aware human motion to be generated on much less motion-scene data than in the prior art. Figure 1 Thus, Figure 2 the scene navigation component 132 of
[0036] Figure 3 can use the scene-aware diffusion model 200 of to generate and / or denoise a motion sequence 135 representing the position and / or orientation of one or more joints at each of a plurality of 2D or 3D path points along a path through the 3D scene 110. In some embodiments, the motion sequence 135 can represent the position and orientation of a representative joint such as a root joint. The root joint can be a designated pivot point in the skeletal structure of the character being animated (e.g., located at the base of the spine or pelvis). The root joint can be used as a basic reference point for the entire skeleton, affecting the position and orientation of the entire body. Thus, manipulating the root joint in the animation can be effectively used to reposition the entire character in the 3D scene 110. In some embodiments, the scene navigation component 132 can use any known motion retargeting technique to generate an animation of the entire body of the character from the position and orientation of the root joint as the character progresses through the path points of the motion sequence 135. Thus, the scene navigation component 132 can animate the movement of the character from the starting point 126 to the first target point 128 in the 3D scene 110.
[0036] Figure 3A block diagram of an example scene-aware diffusion model 300 for a scene interaction component according to some embodiments of the present disclosure. For example, Figure 1 The scene interaction component 140 of can use the scene-aware diffusion model 300 to predict a motion sequence that includes joint orientations (and / or positions) of a sequence of path points along a path from a first target point 128 to a second target point 130 in the 3D scene 110.
[0037] Generally, Figure 2 The components of the scene-aware diffusion model 200 of and Figure 3 The components of the scene-aware diffusion model 300 of (which share reference numerals with the corresponding components of the scene-aware motion diffusion model 200) can perform similar functions, but their architectures can be different. As a non-limiting example, Figure 1 The scene navigation component 132 of can use Figure 2 The scene-aware diffusion model 200 of to predict a motion sequence that represents some quantity of features (e.g., the 2D position and / or heading angle of the ground projection of the root joint of a character) for each of a plurality of path points that form a path through the 3D scene. In contrast, Figure 1 The scene interaction component 140 of can use Figure 3 The scene-aware diffusion model 300 of to predict a motion sequence that represents some other quantity of features (e.g., the 3D positions and / or 3D orientations of all joints in a character's body) for each of a plurality of path points that form a path for interacting with an object in the 3D scene (e.g., an animation representing a character sitting on a chair). Thus, the sizes of the instances of layer 210 in the scene-aware diffusion model 200 and the scene-aware diffusion model 300 can be different and can accept input data with different dimensions. Additionally or alternatively, the number of path points represented in the motion sequences accepted and / or generated by the scene-aware diffusion models 200 and 300 can be different, the number of layers in one or more components of the scene-aware diffusion models 200 and 300 can be different, different conditional inputs 205 can be injected into the scene-aware diffusion models 200 and 300, and / or some other aspects of the architecture can be different.
[0038] In Figure 3 The embodiment shown, the scene-aware diffusion model 300 includes a scene-aware component that includes layer 320 and a transformer encoder 235, and the scene-aware component can be integrated with or connected to a pre-trained base diffusion model (e.g., layer 210, transformer encoder 230, and / or layer 240) to form an input channel (e.g., object interaction data 315) for a representation of a target object and / or a representation of an interaction with the target object.
[0039] For example, Figure 1The scene interaction component 140 can generate object interaction data 315, which can encode a representation of a target object, a surface, or other parts of the 3D scene or the 3D structure of an object in the 3D scene (e.g., a 3D point cloud, a signed distance field representing the distance and direction from any point to the 3D object), and / or a representation of the interaction between the character and the target object, surface, or other parts of the 3D scene (e.g., a base point set (BPS) representation of contact and / or proximity). For example, the scene interaction component 140 can use one or more BPS representations to encode the interaction at any given point in time. Generally, the BPS representation can be used to encode the geometry and appearance (e.g., a 3D point cloud, a 3D mesh, or other 3D model) of a 3D object based on the minimum distance between a set of (e.g., randomly selected) base points and the corresponding nearest points on the 3D object.
[0040] As a non-limiting example, the scene interaction component 140 can define a set of 3D base points (e.g., randomly selected, aligned with the voxels in the 3D mesh), which can be used as a reference frame. To encode the corresponding parts (e.g., 3D crops or samples) of the target object (e.g., a 3D model of a chair) and / or the 3D scene (e.g., around the character and / or the target object), the scene interaction component 140 can identify and select the nearest vertex of the target object for each base point, determine the distance between each base point and the corresponding nearest vertex, and generate a (e.g., concatenated) representation of these distances. To encode the representation of the interaction (e.g., contact and / or proximity) between the character and the target object, the scene interaction component 140 can use the selected vertices of the target object as a set of 3D base points, and for each such point, identify and select the nearest vertex of the 3D representation of the character, determine the distance between each such point and the corresponding nearest vertex, and generate a (e.g., concatenated) representation of these distances. Thus, the scene interaction component 140 can generate and use one or more of the above representations as the object interaction data 315, which can associate the encoded representation of proximity and / or contact with each of the multiple 3D points, and the scene interaction component 140 can apply the object interaction data 315 to the layer 320 (e.g., which can form a 3D convolutional neural network) to encode the object interaction data 315 into the 3D mesh and extract a set of 3D features (e.g., an encoded 3D feature map) representing the object interaction data 315. This is merely an example, and any known way of encoding a representation of a target object, surface, or other parts of the 3D scene or the 3D structure of an object in the 3D scene and / or a representation of the interaction between the character and the target object, surface, or other parts of the 3D scene can be implemented within the scope of the present disclosure.
[0041] Continuing with the above example, to inject the object interaction data 315 into the pre-trained base diffusion model, the scene-aware diffusion model 300 can generate the sampled features by sampling the encoded 3D feature map at elements of the 3D positions of the path points corresponding to the motion sequence at a specific diffusion step t at the joint positions (e.g., the features can be sampled at the 3D joint positions being denoised, which can change at each diffusion step t based on the predicted denoised motion sequence at the previous step t+1). Thus, the scene-aware diffusion model 300 can use the state of the motion sequence at a specific diffusion step t and the corresponding sampled features of the object interaction data 315 to predict the denoised motion sequence
[0042] Figure 3 In some embodiments, Figure 3 the base diffusion model (e.g., layer 210, transformer encoder 230, and layer 240) can be pre-trained on motion data without scene data, and the scene-aware diffusion model 300 can be adjusted (e.g., fine-tuned) on motion-scene data (e.g., input training data including noisy motion sequences and object interaction data 315, and the corresponding ground truth denoised motion sequences). For example, the scene-aware components of the scene-aware diffusion model 300 (e.g., layer 320, transformer encoder 235, one or more layers connecting the transformer encoder 235 to the transformer encoder 230) can be initialized (e.g., with initial values of trainable parameters, such as weights set to zero), the base diffusion model can be frozen or locked Figure 3 (shown with a padlock in
[0043] Figure 3 ), and the scene-aware diffusion model 300 can be trained using paired motion-scene training data to learn the values of the trainable parameters of the scene-aware components. Generally, any known motion-scene dataset can be used and / or any technique can be used to generate a motion-scene dataset with input and ground truth training data corresponding to the scene-aware diffusion model 300. However, by combining the scene-aware components with the base pre-trained motion diffusion model, the resulting scene-aware diffusion model 300 can be fine-tuned on a more limited amount of motion-scene data, enabling the generation of more accurate and scene-aware human motions on much less motion-scene data than in the prior art.In some embodiments, data augmentation can be used to generate training data that pairs (e.g., captured) motion data with corresponding 3D objects being interacted with, to retarget initially captured motion about one object to another object (e.g., retarget motion data captured while sitting in a particular chair to a different chair). Compared to prior art that retargets contact positions (e.g., where the arm or pelvis contacts the chair) to target positions where corresponding joints of the skeletal structure contact the target object, in some embodiments, a 3D model (e.g., 3D mesh) can be used to model the surface structure of the character's body, so that contact positions can be retargeted to target positions where corresponding positions on the body surface contact the target object. Thus, the resulting retargeted motion-object interaction data is more accurate than prior art. Accordingly, the training data can be used to fine-tune the scene-aware diffusion model 300, which should improve the accuracy of the resulting generated motion.
[0044] Accordingly, the returned Figure 1 , Figure 1 scene interaction component 140 can use Figure 3 the scene-aware diffusion model 300 to generate and / or denoise a motion sequence 145 representing the position and / or orientation of one or more joints at each of a plurality of 2D or 3D path points along a path of interacting with a target object (e.g., chair 124) in the 3D scene 110. In some embodiments, the motion sequence 145 can represent the position and orientation of multiple joints in the skeletal structure of an animated character for each path point. Thus, in this example, the scene interaction component 140 can use these positions and orientations to animate the body of the character as the character progresses through the path points of the motion sequence 145 without motion retargeting. Accordingly, the scene interaction component 140 can animate the character moving from a first target point 128 to a second target point 130 in the 3D scene 110.
[0045] Now referring to Figure 4 and Figure 5 , each block of the methods 400 and 500 described herein includes a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be implemented by a processor executing instructions stored in a memory. The methods 400 and 500 can also be embodied as computer-usable instructions stored on a computer storage medium. The methods 400 and 500 can be provided by a stand-alone application, service, or hosted service (independently or in combination with other hosted services) or a plug-in of another product, to name just a few. Additionally, by way of example, with respect to Figure 1The motion generation pipeline 100 describes methods 400 and 500. However, these methods may be performed additionally or alternatively by any one system or any combination of systems, including but not limited to those described herein.
[0046] Figure 4 FIG. 4 is a flowchart showing a method 400 for generating a representation of scene-aware motion according to some embodiments of the present disclosure. At block B402, method 400 includes: generating a representation of scene-aware motion based at least on processing a representation of at least a portion of a three-dimensional (3D) scene using a diffusion model that includes a scene-aware component and a pre-trained motion diffusion model, the representation of scene-aware motion including one or more orientations of one or more joint path points along one or more paths of a character in the 3D scene.
[0047] For example, with respect to Figure 1 the motion generation pipeline 100, the scene navigation component 132 may use Figure 2 the scene-aware diffusion model 200 to predict a motion sequence that represents the position and / or orientation of one or more joints at each of one or more path points along a path 127 from a starting point 126 to a first target point 128. More specifically, Figure 2 the scene-aware diffusion model 200 may include a scene-aware component (e.g., layer 220, transformer encoder 235, one or more layers connecting the transformer encoder 235 to the transformer encoder 230) and a pre-trained diffusion model (e.g., layer 210, transformer encoder 230, and layer 240), and the scene-aware component may form an input channel for scene data 215, which may represent the 2D or 3D structure of at least a portion of the 3D scene 110 at a particular time (e.g., a 2D or 3D occupancy grid, floor plan, height map, a patch of one of the foregoing (such as an egocentric patch), classification data (such as semantic segmentation representing objects or other parts of any number of classes of the scene, etc.)). Thus, the scene-aware diffusion model 200 may use the state of the motion sequence at a particular diffusion step t and the corresponding sampled features of the scene data 215 to predict a denoised motion sequence
[0048] In another example, the scene interaction component 140 may use Figure 3 the scene-aware diffusion model 300 to predict a motion sequence that represents the position and / or orientation of one or more joints at each of one or more path points along a path 129 from the first target point 128 to a second target point 130. More specifically, Figure 3The scene-aware diffusion model 300 can include a scene-aware component (e.g., layer 320, transformer encoder 235, one or more layers connecting transformer encoder 235 to transformer encoder 230) and a pre-trained diffusion model (e.g., layer 210, transformer encoder 230, and layer 240), and the scene-aware component can form an input channel for object interaction data 315, which can encode a representation of the target object, surface, or other parts in the 3D scene or the 3D structure of an object in the 3D scene, and / or a representation of the interaction between the role and the target object, surface, or other parts in the 3D scene (e.g., the BPS representation of contact and / or proximity). Thus, the scene-aware diffusion model 300 can use the state of the motion sequence at a specific diffusion step t and the corresponding sampled features of the object interaction data 315 to predict the denoised motion sequence
[0049] Figure 5 is a flowchart showing a method 500 for generating a diffusion model for motion according to some embodiments of the present disclosure. At block B502, method 500 includes: pre-training a diffusion model using motion data. For example, regarding Figure 2 , the scene-aware diffusion model 200 can include a base diffusion model (e.g., layer 210, transformer encoder 230, and layer 240), which can be pre-trained on motion data without scene data using any known technique (e.g., by Figure 6 the computing device 600).
[0050] At block B504, method 500 includes: connecting a scene-aware component to the pre-trained diffusion model. For example, regarding Figure 2 , the scene-aware component including layer 220 and transformer encoder 235 can be integrated with or (e.g., using Figure 6 the computing device 600) connected to the pre-trained base diffusion model, thereby forming an input channel for scene data 215. In some embodiments, transformer encoder 235 can be connected to transformer encoder 230 via one or more (e.g., linear) layers followed by an addition layer (e.g., residual or skip connection). Layer 220, transformer encoder 235, and / or one or more layers connecting transformer encoder 235 to transformer encoder 230 can be initialized (e.g., with initial values of trainable parameters, such as weights being set to zero).
[0051] At block B506, method 500 includes: adjusting the resulting diffusion model using motion-scene training data. For example, regarding Figure 2 , the scene-aware diffusion model 200, the base diffusion model can be frozen or locked (e.g., byFigure 6 computing device 600), and the scene-aware diffusion model 200 can be trained using paired motion-scene training data (e.g., by Figure 6 computing device 600) to learn the values of the trainable parameters of the scene-aware component. Generally, any known motion-scene dataset can be used and / or any technique can be used to generate a motion-scene dataset having input and ground truth training data corresponding to the scene-aware diffusion model 200. However, by combining the scene-aware component with a base pre-trained diffusion model, the resulting scene-aware diffusion model 200 can be fine-tuned on a more limited amount of motion-scene data, enabling the generation of more accurate and scene-aware human motions on much less motion-scene data than in the prior art.
[0052] The systems and methods described herein can be used for various purposes, by way of example and not limitation, for machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twins, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing, generative AI, and / or any other suitable application.
[0053] The disclosed embodiments can be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, boating systems, intelligent area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems including one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems implementing one or more language models (such as one or more large language models (LLMs)), systems for performing light transport simulation, systems for performing collaborative content creation of 3D assets, systems implemented at least partially using cloud computing resources, and / or other types of systems.
[0054] Example computing device
[0055] Figure 6FIG. 600 is a block diagram of an example computing device 600 suitable for implementing some embodiments of the present disclosure. The computing device 600 may include an interconnect system 602 that directly or indirectly couples the following devices: a memory 604, one or more central processing units (CPUs) 606, one or more graphics processing units (GPUs) 608, a communication interface 610, input / output (I / O) ports 612, input / output components 614, a power supply 616, one or more presentation components 618 (such as a display), and one or more logic units 620. In at least one embodiment, the computing device 600 may include one or more virtual machines (VMs), and / or any of its components may include virtual components (e.g., virtual hardware components). For non-limiting examples, one or more GPUs 608 may include one or more vGPUs, one or more CPUs 606 may include one or more vCPUs, and / or one or more logic units 620 may include one or more virtual logic units. Thus, the computing device 600 may include discrete components (e.g., a complete GPU dedicated to the computing device 600), virtual components (e.g., a portion of a GPU dedicated to the computing device 600), or a combination thereof.
[0056] Although Figure 6 each of the boxes is shown as being connected via the interconnect system 602 with lines, this is not intended to be limiting and is merely for clarity. For example, in some embodiments, a presentation component 618 such as a display device may be considered an I / O component 614 (e.g., if the display is a touchscreen). As another example, the CPU 606 and / or GPU 608 may include memory (e.g., the memory 604 may represent a storage device in addition to the memory of the GPU 608, CPU 606, and / or other components). In other words, Figure 6 the computing device is merely illustrative. No distinction is made among categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "game console", "electronic control unit (ECU)", "virtual reality system", and / or other device or system types, because all of these are considered within the scope of Figure 6 the computing device.
[0057] The interconnect system 602 may represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnect system 602 may include one or more types of links or buses, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or link. In some embodiments, there are direct connections between components. As an example, the CPU 606 may be directly connected to the memory 604. Additionally, the CPU 606 may be directly connected to the GPU 608. Where there are direct or point-to-point connections between components, the interconnect system 602 may include a PCIe link to perform the connection. In these examples, the computing device 600 need not include a PCI bus.
[0058] The memory 604 may include any of a variety of computer-readable media. Computer-readable media can be any available media that can be accessed by the computing device 600. Computer-readable media can include both volatile and non-volatile media and removable and non-removable media. By way of example and not limitation, computer-readable media may include computer storage media and communication media.
[0059] Computer storage media can include volatile and non-volatile media and / or removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, the memory 604 may store computer-readable instructions (e.g., which represent programs and / or program elements, such as an operating system). Computer storage media may include, but are not limited to, RAM, ROM, EEPROM, flash memory or other storage technology, CD-ROM, digital versatile disks (DVDs) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by the computing device 600. As used herein, computer storage media does not include signals per se.
[0060] A computer storage medium can include computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave, or other transmission mechanism, and includes any information conveyance medium. The term "modulated data signal" can refer to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, a computer storage medium can include wired media such as a wired network or direct wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media. Any combination of the foregoing should also be included within the scope of computer-readable media.
[0061] The CPU 606 can be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 600 to perform one or more of the methods and / or processes described herein. Each of the CPUs 606 can include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of simultaneously processing a large number of software threads. The CPU 606 can include any type of processor and can include different types of processors depending on the type of computing device 600 being implemented (e.g., a processor with fewer cores for a mobile device and a processor with more cores for a server). For example, depending on the type of computing device 600, the processor can be an Advanced RISC Machine (ARM) processor implemented using reduced instruction set computing (RISC) or an x86 processor implemented using complex instruction set computing (CISC). In addition to one or more microprocessors or complementary coprocessors such as a math coprocessor, the computing device 600 can also include one or more CPUs 606.
[0062] In addition to or instead of the CPU 606, the GPU 608 can also be configured to execute at least some computer-readable instructions to control one or more components of the computing device 600 to execute one or more methods and / or processes described herein. One or more GPUs 608 can be an integrated GPU (e.g., having one or more CPUs 606) and / or one or more GPUs 608 can be a discrete GPU. In an embodiment, one or more GPUs 608 can be a coprocessor of one or more CPUs 606. The computing device 600 can use the GPU 608 to render graphics (e.g., 3D graphics) or perform general computing. For example, the GPU 608 can be used for general-purpose computing on the GPU (GPGPU). The GPU 608 can include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU 608 can generate pixel data for an output image in response to a rendering command (e.g., a rendering command received from the CPU 606 via the host interface). The GPU 608 can include graphics memory such as display memory for storing pixel data or any other suitable data (e.g., GPGPU data). The display memory can be included as part of the memory 604. The GPU 608 can include two or more GPUs operating in parallel (e.g., via a link). The link can directly connect the GPUs (e.g., using NVLINK) or can connect the GPUs through a switch (e.g., using NVSwitch). When combined, each GPU 608 can generate pixel data or GPGPU data for a different part of the output or for a different output (e.g., the first GPU for the first image and the second GPU for the second image). Each GPU can include its own memory or can share memory with other GPUs.
[0063] In addition to or instead of the CPU 606 and / or the GPU 608, the logic unit 620 can be configured to execute at least some computer-readable instructions to control one or more components of the computing device 600 to execute one or more methods and / or processes described herein. In an embodiment, the CPU 606, the GPU 608, and / or the logic unit 620 can execute any combination of methods, processes, and / or portions thereof discretely or jointly. One or more logic units 620 can be part of and / or integrated in one or more CPUs 606 and / or one or more GPUs 608, and / or one or more logic units 620 can be discrete components of the CPU 606 and / or the GPU 608 or otherwise external to them. In an embodiment, one or more logic units 620 can be a processor of one or more CPUs 606 and / or one or more GPUs 608.
[0064] Examples of the logic unit 620 include one or more processing cores and / or their components, such as data processing units (DPUs), tensor cores (TCs), tensor processing units (TPUs), pixel vision cores (PVCs), vision processing units (VPUs), graphics processing clusters (GPCs), texture processing clusters (TPCs), streaming multiprocessors (SMs), tree traversal units (TTUs), artificial intelligence accelerators (AIAs), deep learning accelerators (DLAs), arithmetic logic units (ALUs)), application specific integrated circuits (ASICs), floating point units (FPUs), input / output (I / O) components, peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) components, etc.
[0065] The communication interface 610 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 600 to communicate with other computing devices via an electronic communication network, including wired and / or wireless communication. The communication interface 610 may include components and functionality that enable communication over any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication via Ethernet or InfiniBand), low power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, one or more of the logic units 620 and / or the communication interface 610 may include one or more data processing units (DPUs) for directly transferring data received over the network and / or via the interconnect system 602 to one or more GPUs 608 (e.g., their memory).
[0066] The I / O port 612 enables the computing device 600 to be logically coupled to other devices including, among other things, I / O components 614, presentation components 618, and / or other components, some of which may be built into (e.g., integrated into) the computing device 600. Illustrative I / O components 614 include microphones, mice, keyboards, joysticks, gamepads, game controllers, dish satellite antennas, scanners, printers, wireless devices, and the like. The I / O components 614 may provide a natural user interface (NUI) that processes user-generated air gestures, voice, or other physiological inputs. In some instances, the input may be transmitted to an appropriate network element for further processing. The NUI may implement any combination of speech recognition, stylus recognition, face recognition, biometric recognition, on-screen and near-screen gesture recognition, air gestures, head and eye tracking, and touch recognition associated with the display of the computing device 600 (described in more detail below). The computing device 600 may include depth cameras such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technologies, and combinations thereof for gesture detection and recognition. Additionally, the computing device 600 may include an accelerometer or gyroscope (e.g., as part of an inertial measurement unit (IMU)) that implements motion detection. In some examples, the output of the accelerometer or gyroscope may be used by the computing device 600 to render immersive augmented reality or virtual reality.
[0067] The power supply 616 may include a hard-wired power supply, a battery power supply, or a combination thereof. The power supply 616 may power the computing device 600 to enable the components of the computing device 600 to operate.
[0068] The presentation component 618 may include a display (e.g., a monitor, touchscreen, television screen, head-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The presentation component 618 may receive data from other components (e.g., GPU 608, CPU 606, DPU, etc.) and output the data (e.g., as images, videos, sounds, etc.).
[0069] Example data center
[0070] Figure 7 An example data center 700 is shown, which may be used in at least one embodiment of the present disclosure. The data center 700 may include a data center infrastructure layer 710, a framework layer 720, a software layer 730, and an application layer 740.
[0071] As Figure 7As shown, the data center infrastructure layer 710 may include a resource coordinator 712, grouped computing resources 714, and node computing resources (“node C.R.”) 716(1)-716(N), where “N” represents any whole positive integer. In at least one embodiment, the node C.R. 716(1)-716(N) may include, but is not limited to, any number of central processing units (“CPU”) or other processors (including DPU, accelerator, field programmable gate array (FPGA), graphics processor or graphics processing unit (GPU), etc.), memory devices (such as dynamic read-only memory), storage devices (such as solid state drive or disk drive), network input / output (NW I / O) devices, network switches, virtual machines (VM), power modules, and cooling modules, etc. In some embodiments, one or more of the node C.R. 716(1)-716(N) may correspond to a server having one or more of the above computing resources. Additionally, in some embodiments, the node C.R. 716(1)-716(N) may include one or more virtual components, such as vGPU, vCPU, etc., and / or one or more of the node C.R. 716(1)-716(N) may correspond to a virtual machine (VM).
[0072] In at least one embodiment, the grouped computing resources 714 may include separate groupings (not shown) of node C.R. 716 housed within one or more racks, or numerous racks (also not shown) within data centers located in various geographical locations. The separate groupings of node C.R. 716 within the grouped computing resources 714 may include grouped computing, network, memory, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, several node C.R. 716 including CPU, GPU, DPU, and / or other processors may be grouped within one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any number of power modules, cooling modules, and / or network switches in any combination.
[0073] The resource coordinator 712 may configure or otherwise control one or more of the node C.R. 716(1)-716(N) and / or the grouped computing resources 714. In at least one embodiment, the resource coordinator 712 may include a software design infrastructure (“SDI”) management entity for the data center 700. The resource coordinator 712 may include hardware, software, or some combination thereof.
[0074] In at least one embodiment, as Figure 7As shown, the framework layer 720 may include a job scheduler 728, a configuration manager 734, a resource manager 736, and a distributed file system 738. The framework layer 720 may include a framework for software 732 that supports the software layer 730 and / or one or more applications 742 of the application layer 740. The software 732 or the application 742 may respectively include web-based service software or applications, such as service software or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 720 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark that can use the distributed file system 738 for large-scale data processing (e.g., "big data"). TM (hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 728 may include a Spark driver for facilitating the scheduling of workloads supported by the various layers of the data center 700. In at least one embodiment, the configuration manager 734 may be capable of configuring different layers, such as the software layer 730 and the framework layer 720 including Spark and the distributed file system 738 for supporting large-scale data processing. The resource manager 736 is capable of managing the cluster or grouped computing resources mapped to or allocated for supporting the distributed file system 738 and the job scheduler 728. In at least one embodiment, the cluster or grouped computing resources may include grouped computing resources 714 at the data center infrastructure layer 710. The resource manager 736 may coordinate with the resource coordinator 712 to manage these mapped or allocated computing resources.
[0075] In at least one embodiment, the software 732 included in the software layer 730 may include software used by at least a portion of the nodes C.R. 716(1)-716(N), the grouped computing resources 714, and / or the distributed file system 738 of the framework layer 720. One or more types of software may include, but are not limited to, Internet web search software, email virus browsing software, database software, and streaming video content software.
[0076] In at least one embodiment, one or more applications 742 included in the application layer 740 may include one or more types of applications used by at least portions of nodes C.R. 716(1)-716(N), the grouped computing resources 714, and / or the distributed file system 738 of the framework layer 720. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing and machine learning applications, including training or inference software, machine learning framework software (such as PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.
[0077] In at least one embodiment, any one of the configuration manager 734, the resource manager 736, and the resource coordinator 712 may implement any number and type of self-modifying actions based on any amount and type of data obtained in any technically feasible manner. The self-modifying actions may relieve the data center operator of the data center 700 from making potentially bad configuration decisions and may avoid underutilized and / or poorly performing portions of the data center.
[0078] The data center 700 may include tools, services, software, or other resources for training one or more machine learning models or using one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, a machine learning model may be trained by calculating weight parameters according to a neural network architecture by using the software and computing resources described above with respect to the data center 700. In at least one embodiment, by using the weight parameters calculated by one or more training techniques, the resources described above with respect to the data center 700 may be used to infer or predict information using the trained machine learning model corresponding to one or more neural networks, such as, but not limited to, those described herein.
[0079] In at least one embodiment, the data center 700 may use a CPU, an application specific integrated circuit (ASIC), a GPU, an FPGA, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inference using the above resources. In addition, one or more of the above software and / or hardware resources may be configured as a service to allow a user to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.
[0080] Example Network Environment
[0081] The network environment suitable for implementing the embodiments of the present disclosure may include one or more client devices, servers, network attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be implemented on one or more instances of the Figure 6 computing device 600 - for example, each device may include similar components, features, and / or functions of the computing device 600. Additionally, in the case of implementing backend devices (e.g., servers, NAS, etc.), the backend devices may be included as part of a data center 700, an example of which is described herein with respect to Figure 7 more detail.
[0082] The components of the network environment may communicate with each other via a network, which may be wired, wireless, or both. The network may include multiple networks, or networks within multiple networks. By way of example, the network may include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (e.g., the Internet and / or the public switched telephone network (PSTN)), and / or one or more private networks. In the case where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) may provide a wireless connection.
[0083] A compatible network environment may include one or more peer - to - peer network environments (in which case servers may not be included in the network environment), and one or more client - server network environments (in which case one or more servers may be included in the network environment). In a peer - to - peer network environment, the functions described herein with respect to servers may be implemented on any number of client devices.
[0084] In at least one embodiment, the network environment may include one or more cloud - based network environments, distributed computing environments, combinations thereof, etc. A cloud - based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework for supporting one or more applications of a software layer and / or an application layer. The software or application may respectively include network - based service software or applications. In an embodiment, one or more client devices may use network - based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a type of free and open - source software web application framework, such as may be used for large - scale data processing (e.g., “big data”) using a distributed file system.
[0085] A cloud-based network environment can provide cloud computing and / or cloud storage that perform any combination of the computing and / or data storage functions (or one or more parts thereof) described herein. Any of these various functions can be distributed across multiple locations from a central or core server (e.g., one or more data centers that can be distributed across states, regions, countries, globally, etc.). If the connection to a user (e.g., a client device) is relatively close to an edge server, the core server can assign at least a portion of the function to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0086] A client device can include at least some of the components, features, and functions of the example computing device 600 described herein. By way of example and not limitation, a client device can be embodied as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, camera, surveillance device or system, vehicle, boat, aircraft, virtual machine, drone, robot, handheld communication device, hospital device, gaming device or system, entertainment system, in-vehicle computer system, embedded system controller, remote control, appliance, consumer electronic device, workstation, edge device, any combination of these described devices, or any other suitable device. Figure 6 The present disclosure can be described in the general context of machine-usable instructions or computer code, including computer-executable instructions such as program modules, that are executed by a computer or other machine such as a personal digital assistant or other handheld device. Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs particular tasks or implements particular abstract data types. The present disclosure can be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The present disclosure can also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communication network.
[0087]
[0088] As used herein, the recitation of "and / or" with respect to two or more elements shall be construed to mean only one element or a combination of elements. For example, "element A, element B, and / or element C" can include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Additionally, "at least one of element A or element B" can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, "at least one of element A and element B" can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0089] The subject matter of the present disclosure is described in detail herein to meet statutory requirements. However, the description itself is not intended to limit the scope of the present disclosure. On the contrary, the inventors have contemplated that the claimed subject matter may also be embodied in other ways, including steps different from the steps described herein or combinations of steps similar to those described in connection with other current or future technologies. Moreover, although the terms "step" and / or "block" may be used herein to imply different elements of a method employed, these terms should not be construed to imply any particular order among or between the various steps disclosed herein unless the order of the steps is expressly recited.
Claims
1. A processor, comprising: One or more processing units, the one or more processing units configured to: generate a representation of scene-aware motion based at least on processing a representation of at least a portion of a three-dimensional (3D) scene using a diffusion model that includes a scene perception component and a pre-trained motion diffusion model, the representation of scene-aware motion including one or more orientations of one or more joint path points along one or more paths of a character depicted at least partially within the 3D scene.
2. The processor of claim 1, wherein the one or more processing units are further configured to: generate the diffusion model based at least on adding the scene perception component to the pre-trained motion diffusion model and tuning the diffusion model using motion-scene training data.
3. The processor according to claim 1, wherein the processing using the diffusion model includes: Inject a top-down height map of the 3D scene into the pre-trained motion diffusion model.
4. The processor according to claim 1, wherein the processing using the diffusion model includes: Inject a 3D point cloud representing at least a portion of a 3D object within the 3D scene into the pre-trained motion diffusion model.
5. The processor according to claim 1, wherein the processing using the diffusion model comprises: Inject classification data representing one or more classified positions of one or more classified objects within the 3D scene into the pre-trained motion diffusion model.
6. The processor according to claim 1, wherein the processing using the diffusion model includes: Inject classification data representing one or more classified positions of one or more other characters or one or more audio sources within the 3D scene into the pre-trained motion diffusion model.
7. The processor of claim 1, wherein the one or more processing units are further configured to: update the diffusion model using training data generated at least based on retargeting motion data including one or more contact positions with a first object to one or more corresponding positions where the modeled body surface contacts a target object.
8. The processor of claim 1, wherein the processor is included in at least one of the following: A system for performing simulation operations; A system for performing digital twin operations; A system for performing optical transmission simulation; A system for performing collaborative content creation of 3D assets; A system for performing deep learning operations; A system for performing remote operations; A system for performing real-time streaming; A system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; A system implemented using an edge device; A system implemented using a robot; A system for generating synthetic data; A system for generating synthetic data using AI; A system including one or more virtual machines (VMs); A system implemented at least partially within a data center; or A system implemented at least partially using cloud computing resources.
9. A system, comprising: One or more processing units, the one or more processing units configured to generate a representation of scene-aware motion corresponding to a character depicted at least partially within the 3D scene based at least on processing a representation of at least a portion of the 3D scene using a diffusion model.
10. The system according to claim 9, wherein the one or more processing units are further configured to: generate the diffusion model based at least on adding a scene perception component to a pre-trained motion diffusion model and adjusting the diffusion model using motion-scene training data.
11. The system according to claim 9, wherein the processing using the diffusion model comprises: Inject a top-down height map of the 3D scene into the pre-trained motion diffusion model of the diffusion model.
12. The system according to claim 9, wherein the processing using the diffusion model comprises: Inject a 3D point cloud representing at least a portion of the 3D objects in the 3D scene into the pre-trained motion diffusion model of the diffusion model.
13. The system according to claim 9, wherein the processing using the diffusion model comprises: Inject classification data representing one or more classified positions of one or more classified objects in the 3D scene into the pre-trained motion diffusion model of the diffusion model.
14. The system according to claim 9, wherein the processing using the diffusion model comprises: Inject classification data representing one or more classified positions of one or more other characters or one or more audio sources in the 3D scene into the pre-trained motion diffusion model of the diffusion model.
15. The system according to claim 9, wherein the one or more processing units are further configured to: update the diffusion model using training data generated at least based on retargeting motion data including one or more contact positions with a first object to one or more corresponding positions where a modeled body surface contacts a target object.
16. The system according to claim 9, wherein the system is included in at least one of the following: A system for performing simulation operations; A system for performing digital twin operations; A system for performing optical transmission simulation; A system for performing collaborative content creation of 3D assets; A system for performing deep learning operations; A system for performing remote operations; A system for performing real-time streaming; A system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; A system implemented using an edge device; A system implemented using a robot; A system for generating synthetic data; A system for generating synthetic data using AI; A system including one or more virtual machines VM; A system implemented at least partially in a data center; or A system implemented at least partially using cloud computing resources.
17. A method, comprising: Generating a representation of one or more orientations of one or more path points along a path of a character in the 3D scene based at least on injecting a representation of at least a portion of a three-dimensional (3D) scene into a pre-trained diffusion model.
18. The method according to claim 17, further comprising: Generating the diffusion model based at least on adding a scene perception component to the pre-trained diffusion model and adjusting the diffusion model using motion-scene training data.
19. The method according to claim 17, further comprising: Updating a diffusion model including the pre-trained diffusion model using training data generated at least based on retargeting motion data including one or more contact positions with a first object to one or more corresponding positions where a modeled body surface contacts a target object.
20. The method according to claim 17, wherein the method is performed by at least one of the following: A system for performing simulation operations; A system for performing digital twin operations; Systems for performing optical transmission simulations; Systems for performing collaborative content creation of 3D assets; Systems for performing deep learning operations; Systems for performing remote operations; Systems for performing real-time streaming; Systems for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; Systems implemented using edge devices; Systems implemented using robots; Systems for generating synthetic data; Systems for generating synthetic data using AI; Systems comprising one or more virtual machines (VMs); Systems implemented at least partially in a data center; or Systems implemented at least partially using cloud computing resources.
Citation Information
Cited By
Multi-graph reference digital life body generation method and device, equipment and storage medium
CN121747158A
Multi-image reference number life body generation method, device, equipment and storage medium
CN121747158B
Unmanned aerial vehicle control method based on multi-modal fusion and related equipment
CN121765632A