Generating Scene-Aware Synthetic Human Motion Using Neural Networks
By integrating a scene-aware component into a pre-trained motion diffusion model, the generation of realistic human motion in 3D scenes is enhanced, addressing the limitations of conventional methods by improving accuracy and reducing data requirements.
Patent Information
- Application Number
- DE102025100765
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-17
- Filing Date
- 2025-01-10
- Publication Date
- 2025-07-17
AI Technical Summary
Conventional techniques for generating human motion in 3D scenes face challenges due to the need for high-quality training data, which is costly and limited, leading to non-realistic and low-quality animations, and existing methods that do not require paired motion scene data often fail to capture the full range of human motion accurately.
A motion diffusion model is pre-trained on motion data and enhanced with a scene-aware component to extract and inject scene information, allowing for the generation of more accurate and realistic human motion by predicting joint orientations and positions based on 3D scene structures or object interactions.
The approach enables the generation of scene-aware human motion with improved accuracy using less motion scene data, allowing for realistic animations in various 3D environments without the need for extensive high-quality training data.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND
[0001] Human-type character motion generation (or simply human motion) typically strives to create realistic and natural movements for virtual or animated characters, mimicking the way humans move in the real world. Human motion generation, for example, may attempt to simulate the complex interplay of joints, muscles, and / or physical limitations to produce lifelike animations. Human motion generation often plays a central role in computer graphics, animation, and / or virtual reality applications, as it can add a layer of authenticity and immersion to digital experiences in various industries and applications, such as video games, film and television production, simulation training, healthcare (e.g., for physical therapy simulations), and / or other scenarios.In the entertainment industry, human motion generation can enable the creation of convincing and believable characters, enhancing the overall viewing experience. In the context of training or design simulations, human motion generation can allow professionals to practice or design in a controlled environment without the risks of the real world. In healthcare, human motion generation can aid rehabilitation and recovery by providing patients with interactive exercises tailored to their specific needs. These are just a few examples where human motion generation can help bridge the gap between the digital and physical worlds.
[0002] Conventional techniques for generating synthetic human motion suffer from a variety of drawbacks. For example, some techniques may seek to generate human motion (e.g., a character animation) within a particular three-dimensional (3D) scene based on an input text prompt that provides some kind of instruction (e.g., "Sit on the couch"). Typically, the goal is to generate motion that is physically realistic both in terms of navigating the 3D scene (e.g., avoiding collisions while navigating around furniture) and interacting with objects in the scene (e.g., people usually sit facing forward in a chair, not sideways). However, conventional techniques struggle to generate realistic motion for many 3D scenes.For example, conventional techniques for generating human motion typically require high-quality training data that pairs captured human motion with corresponding 3D scenes and interactions within the 3D scene. This type of dataset can be challenging and costly to generate (e.g., requiring high-quality motion capture with the particular characters, actions, objects, and / or 3D scenes of interest). Therefore, this type of training data is usually limited, so that conventional models trained on specific characters, actions, objects, and / or 3D scenes typically cannot generalize to others, resulting in unrealistic and / or low-quality motion animations. Some techniques attempt to address this concern by placing high-quality motion capture sequences (captured without the environment) into scanned scene environments.However, the synthetic motion resulting from these techniques often does not adequately reflect real-world human behavior. Finally, one conventional technique attempts to address the lack of suitable training data using reinforcement learning, which does not require paired motion scene data but instead trains a different strategy for each type of supported interaction. However, limiting generated motion to specifically considered human interactions is unlikely to capture the full range and subtleties of potential human motion. Therefore, there is a need for improved human motion generation techniques. SUMMARY
[0003] The invention is defined by the claims. To illustrate the invention, aspects and embodiments that may or may not fall within the scope of the claims are described herein.
[0004] A motion diffusion model is disclosed that can be pre-trained on motion data, and a scene-aware component (e.g., one or more layers of a neural network) can be connected and used to extract a representation of scene information and inject it into the pre-trained motion diffusion model. For example, to predict orientations of joint waypoints along a path through a particular 3D scene, a scene-aware input channel that accepts a representation of the 3D structure of the scene can be added to a pre-trained motion diffusion model. For example, to predict orientations of joint waypoints along a path that interacts with a 3D object in the 3D scene, a scene-aware input channel that accepts a representation of the 3D object and / or a surface thereof can be added to a pre-trained motion diffusion model.The resulting scene-aware motion diffusion model(s) can therefore be tailored to motion scene data and used to generate human motion.
[0005] Embodiments of the present disclosure relate to the generation of scene-aware human motion. Systems and methods are disclosed that pretrain a base motion diffusion model without scene information, connect a scene-aware component, and tune the resulting motion diffusion model to data with scene information.
[0006] Unlike conventional systems such as those described above, a motion diffusion model can be pre-trained on motion data, and a scene-aware component (e.g., one or more layers of a neural network) can be connected and used to extract a representation of scene information and inject it into the pre-trained motion diffusion model. For example, to predict joint waypoint orientations along a path through a given 3D scene, a scene-aware input channel that accepts a representation of the scene's 3D structure can be added to a pre-trained motion diffusion model.For example, to predict joint waypoint orientations along a path that interacts with a 3D object in the 3D scene, a scene-aware input channel that accepts a representation of the 3D object and / or a surface thereof can be added to a pre-trained motion diffusion model. The resulting scene-aware motion diffusion model(s) can therefore be tuned to motion scene data and used to generate human motion. Accordingly, the techniques described herein can be employed to generate scene-aware human motion for a character based on a representation of a 3D scene and / or a target 3D object with which the character may interact.By incorporating a scene-aware component with a pre-trained baseline motion diffusion model, the resulting scene-aware motion diffusion model can be fine-tuned to a limited set of motion scene data, enabling the generation of more accurate and scene-aware human motion on far less motion scene data than previous techniques.
[0007] The disclosure extends to any new aspect or feature described and / or illustrated herein.
[0008] Further features of the disclosure are characterized by the independent and dependent claims.
[0009] Any feature of one aspect of the disclosure may be applied to other aspects of the disclosure in any suitable combination. In particular, method aspects may be applied to device or system aspects, and vice versa.
[0010] Furthermore, features implemented in hardware may be implemented in software and vice versa. Any reference to software and hardware features herein should be construed accordingly.
[0011] Any system or device feature described herein may also be provided as a method feature, and vice versa. System and / or device aspects described functionally (including means plus functional features) may alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and associated memory.
[0012] It is also understood that particular combinations of the various features described and defined in aspects of the disclosure may be independently implemented and / or provided and / or used.
[0013] The disclosure also provides computer programs and computer program products comprising software code that, when executed on a data processing device, is adapted to perform any of the methods and / or embody any of the device and system features described herein, including all component steps of any method.
[0014] The disclosure also provides a computer or computing system (including networked or distributed systems) having an operating system that provides a computer program for performing methods described herein and / or for embodying any device or system features described herein.
[0015] The disclosure also provides a computer-readable medium having stored thereon one or more of the above-mentioned computer programs.
[0016] The disclosure also provides a signal carrying one or more of the above-mentioned computer programs.
[0017] The disclosure extends to methods and / or devices and / or systems as described herein with reference to the accompanying drawings.
[0018] Aspects and embodiments of the present disclosure will now be described, by way of example only, with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The present systems and methods for generating scene-aware human motion are described below with reference to the accompanying drawing figures, wherein: Fig. 1 is a block diagram of an exemplary motion generation pipeline according to some embodiments of the present disclosure; Fig. 2 is a block diagram of an exemplary scene-aware motion diffusion model for a scene navigation component according to some embodiments of the present disclosure; Fig. 3 is a block diagram of an exemplary scene-aware motion diffusion model for a scene interaction component according to some embodiments of the present disclosure; Fig. 4 is a flowchart illustrating a method for generating a scene-aware motion representation according to some embodiments of the present disclosure. FIG. Fig. 5 is a flowchart illustrating a method for generating a motion diffusion model according to some embodiments of the present disclosure. Fig. 6 is a block diagram of an exemplary computing device suitable for use in implementing some embodiments of the present disclosure; and Fig. 7 is a block diagram of an exemplary data center suitable for use in implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0020] Systems and methods relating to the generation of scene-aware human motion are disclosed. In some embodiments, a diffusion model (e.g., a motion diffusion model) may be pre-trained on motion data and used as a base model, and a scene-aware component (e.g., one or more layers of a neural network) may be used to extract a representation of scene information and inject it into the pre-trained motion diffusion model. For example, to predict orientations of joint waypoints (e.g., root joint waypoints) along a path through a given 3D scene, a scene-aware input channel that accepts a representation of the scene's 3D structure (e.g., a two-dimensional (2D) or 3D occupancy grid, a ground map, height map, semantic segmentation, etc.) may be added to a pre-trained motion diffusion model.In another example, to predict the orientations of joint waypoints (e.g., root joint waypoints) along a path that interacts with a 3D object in the 3D scene, a scene-aware input channel that accepts a representation of the 3D object and / or a surface thereof (e.g., a 3D point cloud) can be added to a pre-trained (e.g., motion) diffusion model. The resulting scene-aware motion diffusion model(s) can therefore be (e.g., fine-tuned) on motion scene data and used to generate human motion. The present techniques can be used to generate more accurate and scene-aware human motion on far less motion scene data than previous techniques.
[0021] For example, given a representation of a 3D scene, a starting point, an instruction (e.g., a textual prompt such as an instruction for a character to "sit on the couch"), classification data such as semantic segmentation of the 3D scene, and / or other input(s), any known path planning technique may be used to identify a target point in the 3D scene to which the character should move, identify a path to the target point through the 3D scene (e.g., to a chair) that avoids collisions (e.g., with other furniture), identify a path that implements interaction with a target object in the 3D scene (e.g., sitting down on the chair), and / or identify one or more points of contact between the character and the target object. Each path may take the form of a sequence of 2D or 3D waypoints (or the waypoints may be sampled along the path).In some embodiments, the sequence of waypoints may represent consecutive 2D or 3D positions of one or more joints (e.g., a root joint) of the character being animated. These waypoints may be used as inputs to one or more diffusion models to predict orientations of the corresponding joint(s) at the waypoints.Additionally or alternatively, to identify a path to the destination point and / or a path that implements interaction with a target object prior to predicting orientations of joint(s) at waypoints along the path(s), any known planning technique may be used to identify the destination point and / or one or more points of contact between the character and the target object, and noisy intermediate waypoints may be used as inputs to one or more diffusion models to predict positions and orientations of the corresponding joint(s) at the waypoints, effectively predicting the path(s) and poses along the path(s) (e.g., given positions of an origin point, destination point, and / or one or more points of contact).
[0022] For example, to predict joint orientations (and / or positions) for a motion sequence represented by a sequence of waypoints along a path through a 2D or 3D scene, a scene-aware diffusion model may encode representations of an instruction (e.g., a text instruction), positions and noisy orientations of the sequence of waypoints (and / or noisy positions of waypoints when predicting corresponding waypoint positions), and a 2D or 3D structure of at least a portion of the 3D scene (e.g., a 2D or 3D occupancy grid, a ground map, a height map, a patch of any of the foregoing, such as an egocentric patch, classification data, such as a semantic segmentation that may represent any number of classes of objects or other parts of the scene, etc.). In some embodiments, the classification data may include a layer for each of one or more object classes, such as (e.g.,Different types of furniture, doors, windows, appliances, walls, lighting fixtures, outlets and switches, electronic devices, personal items, one or more other characters, audio sources, and / or other things in the scene. The scene-aware diffusion model can therefore combine these encoded inputs to predict a denoised scene-aware motion sequence.
[0023] For example, the scene-aware diffusion model can iteratively predict and refine a denoised motion sequence over a series of diffusion steps based on the 2D or 3D structure of the 3D scene. In each diffusion step, the scene-aware diffusion model (e.g., a transformer-based model) can predict a denoised motion sequence based on the 2D or 3D structure of the 3D scene and diffuse the predicted motion sequence back to the previous diffusion step, effectively updating the state of the denoised motion sequence based on scene structure in reverse order from the final diffusion step to the initial one.By starting with the most refined representation of the scene-aware motion and diffusing it back to the previous step, the denoised scene-aware motion sequence predicted in each diffusion step benefits from the accumulated improvements made at later steps, improving scene-aware temporal dependence to correct errors or inaccuracies introduced in earlier steps, resulting in a more accurate and realistic scene-aware motion sequence.
[0024] In some embodiments, to predict joint orientations (and / or positions) for a movement sequence represented by a sequence of waypoints along a path that interacts with a target object in a 3D scene (e.g., sitting down on a chair), a scene-aware diffusion model may use representations of an instruction (e.g., a text instruction), positions and noisy orientations of the sequence of waypoints and / or noisy positions of waypoints when predicting corresponding waypoint positions), a 3D structure of the 3D object on a surface thereof (e.g., a 3D point cloud), one or more touch locations on the 3D object (e.g.,The scene-aware diffusion model may therefore combine these encoded inputs to predict a denoised, scene-aware motion sequence. In some embodiments, the scene-aware diffusion model may iteratively predict and refine a denoised motion sequence over a series of diffusion steps based on the 3D structure of the 3D object being interacted with.In each diffusion step, the scene-aware diffusion model (e.g., a transformer) can predict a denoised motion sequence based on the 3D structure of the 3D object and diffuse the predicted motion sequence back to the previous step, effectively updating the state of the scene-aware motion sequence based on the structure of the object in the scene in reverse order, effectively incorporating accumulated improvements, improving scene-aware (e.g., object-aware) temporal dependencies, and providing an opportunity to correct errors or inaccuracies introduced in previous steps, resulting in a more accurate and realistic scene-aware motion sequence.
[0025] In some embodiments, training data that pairs (e.g., captured) motion data with a corresponding 3D object being interacted with may be generated using data augmentation to retarget motion originally captured with respect to one object to a different object (e.g., retargeting captured motion data when sitting down from a particular chair to a different chair). In contrast to prior techniques that retarget touch locations (e.g., the locations where arms or pelvis touch the chair) to target locations where corresponding joints of a skeletal structure touch the target object, in some embodiments, the surface structure of the character's body may be retargeted using a 3D model (e.g.,A 3D mesh can be modeled to touch locations that can be realigned to target locations where corresponding locations on the body surface touch the target object. Therefore, the resulting realigned motion-object interaction data is more accurate than previous techniques, and training a scene-aware diffusion model (e.g., fine-tuning a pre-trained diffusion model) using this training data increases the accuracy of the resulting generated motion.
[0026] In an example training embodiment, a pre-trained base diffusion model may be tuned or adapted to scene content using fine-tuning (e.g., freezing one or more layers of the pre-trained model), Parameter-Efficient Fine-Tuning (PEFT) (e.g., Low-Rank Adaptation (LoRA), Prefix Tuning, Prompt Tuning, p-Tuning), another technique that updates one or more trainable parameters (e.g., network weights, rank decomposition matrices, hard prompts, soft prompts)), and / or otherwise. Fine-tuning may, for example, involve adding one or more scene-aware layers that extract a representation of scene information and injecting it into the pre-trained base diffusion model and training the resulting model (e.g., fixing one or more pre-trained layers of the pre-trained base diffusion model) using (e.g.,realigned) motion object data to learn appropriate weights for the added layer(s).
[0027] Therefore, the techniques described herein can be used to generate scene-aware human motion for a character based on a representation of a 3D scene and / or a 3D target object with which the character can interact. By incorporating a scene-aware component with a pre-trained base diffusion model, the resulting scene-aware diffusion model can be fine-tuned to more constrained motion scene data, enabling the generation of more accurate and scene-aware human motion on far less motion scene data than previous techniques.
[0028] With reference to 1, Fig. 1 illustrates an exemplary motion generation pipeline 100 according to some embodiments of the present disclosure. It should be understood that these and other arrangements described herein are set forth as examples only. Other arrangements and elements (e.g., engines, interfaces, functions, sequences, groupings of functions, etc.) may be used in addition to or in place of those shown, and some elements may be omitted entirely. Further, many of the elements described herein are functional units that may be implemented as separate or distributed components or in conjunction with other components, and in any suitable combination and location. Various functions described herein as being performed by units may be performed by hardware, firmware, and / or software.For example, various functions can be performed by a processor executing instructions stored in memory.
[0029] In the Fig. 1, the motion generation pipeline 100 includes a planning component 120 that accepts a text prompt 105, a representation of a 3D scene 110, and classification data 115 representing one or more aspects of the 3D scene 110. In some embodiments, the planning component 120 uses one or more of these inputs to identify a destination point in the 3D scene to which a character in the 3D scene 110 should move and / or the position(s) of one or more waypoints along a path to the destination point. Therefore, a scene navigation component 132 may use positions of a known origin point, the destination point, and / or the waypoint(s) to predict orientations of the corresponding joint(s) at the waypoints. In some embodiments, the planning component 120 predicts the path (e.g.,Positions of one or more waypoints), and the scene navigation component 132 predicts poses (e.g., orientations of the corresponding joint(s)) at the waypoints along the path. In some embodiments, the planning component 120 predicts the destination point, and the scene navigation component 132 predicts the path (e.g., positions of one or more waypoints) and poses (e.g., orientations of the corresponding joint(s)) at the waypoints along the path.
[0030] In some embodiments, the textual prompt 105 may represent an instruction for a character initially located at an exit point 126 to interact with a target object in the 3D scene 110 (e.g., a chair 124). The planning component 120 may therefore identify the target object from the textual prompt 105, identify a first target point 128 in the 3D scene 110 to which the character may move prior to interacting with the target object, identify a corresponding interaction (e.g., sitting down on the chair 124), identify a second target point 130 to which the character may move via the interaction, and / or identify one or more points of contact between the character and the target object. In addition or alternatively to the scene navigation component 132 predicting a path 127 between the exit point 126 and the first target point 128 and / or poses along the path 127 (e.g.,Motion sequence 135), a scene interaction component 140 may predict a path 129 between the first target point 128 and the second target point 130 and / or poses along the path 129 (e.g., motion sequence 145).
[0031] In general, a variety of inputs are possible depending on the implementation. For example, the motion generation pipeline 100 may be integrated into or triggered by a user interface for character animation, robotics, and / or another type of application that generates a representation of motion and / or animates motion, and the user interface may accept one or more user inputs representing an instruction for a character, robot, or other identity to move within and / or interact with the 3D scene 110. In the embodiment described in Fig. 1, a natural language instruction is embodied in text prompt 105, but this need not be the case. Additionally or alternatively, the user interface may accept and / or encode an instruction represented by a voice command, a detected gesture, joystick or gamepad input, one or more virtual and / or augmented reality controllers, spatial coordinates identified via the user interface, and / or other types of input.
[0032] In embodiments that include the textual prompt 105, the planning component 120 may use any known technique to evaluate the textual prompt 105, a representation of a current state of the 3D scene 110 (e.g., a 2D or 3D occupancy grid, a floor map, a height map), and / or corresponding classification data 115 (e.g., a semantic segmentation of the 3D scene 110) to identify one or more target points for the character's movement in the 3D scene 110 and / or a corresponding target orientation at each of the target points. For example, the planning component 120 may use natural language processing (e.g., named entity recognition, keyword extraction) to identify and extract relevant spatial information, a target object in the 3D scene 110, and / or directional cues referenced in the textual prompt 105.Additionally or alternatively, the planning component 120 may use one or more machine learning models (e.g., one or more language models) to analyze the textual prompt 105 and derive an intended target position and / or target orientation. In some embodiments, the planning component 120 may use any known technique for evaluating the textual prompt 105, the 3D scene 110, and / or the classification data 115 to generate a path through the 3D scene 110 (e.g., path 127) and / or a path that implements an interaction specified in the textual prompt 105 (e.g., path 129). For example, the planning component 120 may use a pathfinding algorithm (e.g., A*) to generate a path that avoids obstacles in the 3D scene 110. Note that . Fig. 1 illustrates an embodiment in which the planning component 120 generates paths 127 and 129, but this need not be the case.
[0033] Fig. 2 is a block diagram of an exemplary scene-aware diffusion model 200 for a scene navigation component according to some embodiments of the present disclosure. The scene-aware diffusion model 200 may be used by the scene navigation component 132 of the Fig. 1 can be used to predict a motion sequence comprising joint orientations (and / or positions) of a sequence of waypoints along a path from the starting point 126 to the first destination point 128 through the 3D scene 110.
[0034] In some embodiments, the scene-aware diffusion model 200 may be implemented using one or more neural networks. Although the scene-aware diffusion model 200 and other models and functionality described herein may be implemented using one or more neural networks (or a portion thereof), this is not intended to be limiting. In general, the models and / or functionality described herein may be implemented using any of a number of different networks or machine learning models, such as one or more machine learning models using linear regression, logistic regression, decision trees, support vector machines (SVMs), Naive Bayes, k-nearest neighbor (Knn), k-means clustering, random forest, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g.,, autoencoders, convolutional algorithms, transformers, recurrent networks, perceptrons, long / short term memory (LSTM), Hopfield, Boltzmann, deep belief, deconvolutional, generative adversarial, liquid state machine, etc.) and / or other types of machine learning models.
[0035] In general, a path can be defined as a sequence of 2D or 3D waypoints x 1 ... x N and the scene-aware diffusion model 200 can iteratively generate a denoised motion sequence x̂ 1 ... x Nover a series of t diffusion steps. The scene navigation component 132 may initially build a representation of the sequence using one or more data structures representing position (e.g., 3D position, a 2D ground projection), orientation, and / or other features of one or more joints (e.g., a root joint, such as a pelvis) of the character at each of the waypoints that contain known parameters (e.g., position and orientation of the origin x 1 and the target point x N ) in corresponding elements of the one or more data structures, and fill the remaining elements (e.g., the unknowns to be predicted) with random noise. Therefore, the scene navigation component 132 can generate a representation of the motion sequence xt1...xtN at a particular diffusion step t on the scene-aware motion diffusion model 200 and apply it to generate a denoised motion sequence x̂ 1 ... x N to predict. Fig. 1 illustrates an embodiment in which the scene navigation component 132 starts at the last diffusion step t=T using the scene-aware diffusion model 200 to generate a denoised motion sequence x^01...x^0N at the first diffusion step t=0. After this first iteration, the scene navigation component 132 may diffuse the denoised motion sequence to a state corresponding to the one preceding the last diffusion step t=T-1 (illustrated in Fig. 2 as the predicted results of the previous step), and use this as an input to the scene-aware motion diffusion model 200 to again produce a denoised motion sequence x^01...x^0N at the first diffusion step t=0. The scene navigation component 132 may repeat this process, effectively updating the state of the denoised motion sequence in reverse order from the final diffusion step to the initial one. This reverse diffusion process is intended to serve as an example, and other diffusion techniques with any number and order of diffusion steps may be implemented within the scope of the present disclosure.
[0036] In some embodiments, the scene-aware diffusion model 200 may include a base diffusion model (e.g., comprising one or more layers 210, a transformer encoder 230, and one or more layers 240) that may be pre-trained on motion data (e.g., input and output motion sequences) without scene data using a known technique. In the embodiment described in Fig. 2, the scene-aware diffusion model 200 also includes a scene-aware component that includes one or more layers 220 and a transformer encoder 235 that may be integrated with or connected to the pre-trained base diffusion model and form an input channel for scene data 215.
[0037] In some embodiments, the scene navigation component 132 may generate the scene data 215, which may represent one or more features of the 3D scene 110 and may take any suitable form. For example, the scene data 215 may represent a 2D or 3D structure of at least a portion of the 3D scene 110 (e.g., a 2D or 3D occupancy grid, a ground map, a height map, a patch of one or more of the foregoing, such as an egocentric patch, classification data, such as a semantic segmentation representing any number of classes of objects or other parts of the scene, etc.). In some embodiments, the classification data may include a layer for each of one or more object classes, such as (e.g.,different types of) furniture, doors, windows, appliances, walls, lighting fixtures, outlets and switches, electronic devices, personal items, one or more other characters, audio sources, and / or other things that may be located in the scene. The scene navigation component 132 may therefore generate a representation of the scene data 215 and apply the scene data 215 to the layer(s) 220 to extract a set of features (e.g., an encoded 2D feature map) representing the scene data 215.
[0038] To inject the scene data 215 into the pre-trained base diffusion model, the scene-aware diffusion model 200 may sample features ft1...xtN by scanning the coded feature map at elements corresponding to the positions of the (e.g. 2D) waypoints of the movement sequence xt1...xtN at a particular diffusion step t (e.g., locations can be sampled that are denoised, which changes at each diffusion step t based on the denoised motion sequence predicted at the previous step t+1).
[0039] Therefore, the scene-aware diffusion model 200 can determine the state of the motion sequence xt1...xtN at a given diffusion step t and corresponding sampled features ft1...ftN of the scene data 215 to create a denoised motion sequence x^01...x^0N For example, the scene-aware diffusion model 200 may represent the text prompt 105) of the Fig. 1, encode a representation of the repetitive diffusion step and / or another input and use the encoded representation as a conditioning input 205. One or more layers 210 may serve to process the input representation of the movement sequence xt1...xtN to a particular dimensionality for the transformer encoder 230. For example, each waypoint may be represented in 3D (e.g., x, y, heading angle), and the one or more layers 210 may scale each waypoint to a larger dimensionality (e.g., 64 or 128), such as one corresponding to the dimensionality of the sampled features of the encoded feature map of the scene data 215. Position encodes 225 may be used to encode and integrate (e.g., add) a representation of the time step into each corresponding waypoint or sampled scene feature (e.g., a corresponding token). The transformer encoders 230 and 235 may each include any number of layers, and each layer may include a number of attention heads that effectively weight different elements in the input based on importance.
[0040] The transformer encoder 235 may be connected to the transformer encoder 230 in any suitable manner. In an exemplary implementation, the transformer encoder 235 may be connected to the transformer encoder 230 via one or more (e.g., linear) layers followed by an addition (e.g., a remainder or jump connection). For example, the transformer encoder 235 may output an encoded representation of the scene data 215, one or more linear layers (in Fig. 2 as arrows between the transformer encoder 235 and the transformer encoder 230) may be used to process the encoded representation of the scene data 215, and the result may be added to the intermediate features of the transformer encoder 230 (e.g., a pre-trained basic diffusion transformer).
[0041] In some embodiments, the base diffusion model (e.g., layer or layers 210, transformer encoder 230, and layer or layers 240) may be pre-trained on motion data without scene data, and the scene-aware diffusion model 200 may be tuned (e.g., fine-tuned) on motion scene data (e.g., input training data comprising noisy motion sequences and scene data 215 and corresponding denoised ground truth motion sequences). For example, the scene-aware component of the scene-aware motion diffusion model 200 (e.g., the layer(s) 220, the transformer encoder 235, one or more layers connecting the transformer encoder 235 to the transformer encoder 230) may also be initialized (e.g., with initial values for trainable parameters, such as weights, set to zero), the base diffusion model may be frozen or locked (with padlocks in Fig. 2), and the scene-aware diffusion model 200 may be trained using paired motion scene training data to learn values for the trainable parameters of the scene-aware component. In general, any known motion scene dataset may be used and / or any technique may be used to generate a motion scene dataset with input and ground truth training data for the scene-aware diffusion model 200. However, by incorporating a scene-aware component with a pre-trained diffusion model, the resulting scene-aware diffusion model 200 may be fine-tuned to more constrained motion scene data, enabling the generation of more accurate and scene-aware human motion on far less motion scene data than previous techniques.
[0042] The scene navigation component 132 of the Fig. 1, the scene-aware diffusion model 200 of the Fig. 2 to generate and / or denoise a motion sequence 135 representing one or more positions and / or orientations of one or more joints at each of the plurality of 2D or 3D waypoints along the path through the 3D scene 110. In some embodiments, the motion sequence 135 may represent position and orientation for a representative joint, such as a root joint. The root joint may be a particular pivot point in a skeletal structure of the character to be animated (e.g., located at the base of the spine or pelvis). The root joint may serve as a fundamental reference point for the entire skeleton, affecting the position and orientation of the entire body. Therefore, manipulating the root joint during animation may effectively serve to relocate the entire character within the 3D scene 110.In some embodiments, the scene navigation component 132 may use known motion inking techniques to generate a full-body animation of the character from the position and orientation of the root joint as it moves through the waypoints of the motion sequence 135. The scene navigation component 132 may therefore animate the character moving from the starting point 126 to the first destination point 128 in the 3D scene 110.
[0043] Fig. 3 is a block diagram of an exemplary scene-aware diffusion model 300 for a scene interaction component according to some embodiments of the present disclosure. The scene-aware diffusion model 300 may, for example, be used by the scene interaction component 140 of the Fig. 1 can be used to predict a motion sequence comprising joint orientations (and / or positions) of a sequence of waypoints along a path from the first target point 128 to the second target point 130 in the 3D scene 110.
[0044] In general, the components of the scene-aware diffusion model 200 of the Fig. 2 and the scene-aware diffusion model 300 of the Fig. 3, which share a reference number with a corresponding component of the scene-aware motion diffusion model 200, implement similar functionality, although their architectures may differ. As a non-limiting example, the scene navigation component 132 of Fig. 1, the scene-aware diffusion model 200 of the Fig. 2 to predict a motion sequence representing a number of features (e.g., 2D position and / or heading angle for a ground projection of the character's root joint) for each of a plurality of waypoints forming a path through a 3D scene. In contrast, the scene interaction component 140 of Fig. 1 the scene-aware diffusion model 300 of the Fig. 3 to predict a motion sequence representing a different number of features (e.g., 3D position and / or 3D orientation for all joints in the character's body) for each of a plurality of waypoints that form a path that interacts with an object in the 3D scene (e.g., representing an animation of the character sitting down in a chair). The instances of layer(s) 210 in the scene-aware diffusion models 200 and 300 can therefore be scaled differently and can accept input data with different dimensionalities.Additionally or alternatively, the number of waypoints represented in the motion sequence accepted and / or generated by the scene-aware diffusion models 200 and 300 may differ, the number of layers in one or more components of the scene-aware diffusion models 200 and 300 may differ, different conditioning input 205 may be injected into the scene-aware diffusion models 200 and 300, and / or another aspect of the architecture may differ.
[0045] In the embodiment shown in Fig. 3, the scene-aware diffusion model 300 includes a scene-aware component comprising layer(s) 320 and the transformer encoder 235, which may be integrated with or connected to the pre-trained base diffusion model (e.g., layer(s) 210, transformer encoder 230, and / or layer(s) 240) by forming an input channel for a representation of the target object and / or an interaction with the target object (e.g., object interaction data 315).
[0046] The scene interaction component 140 of the Fig. 1 may, for example, generate the object interaction data 315, which may encode a representation of a 3D structure of a target object, surface, or other portion of a 3D scene or an object in a 3D scene (e.g., a 3D point cloud, a signed distance field representing the distance and direction from any point to the 3D object), and / or a representation of an interaction (e.g., a base point set (BPS) representation of touch and / or proximity) between a character and the target object, surface, or other portion of the 3D scene. The scene interaction component 140 may, for example, encode the interaction at a given point in time using one or more BPS representations. In general, a BPS representation may be used to represent the geometry and appearance of a 3D object (e.g.,a 3D point cloud, a 3D mesh or a 3D model) in terms of the minimum distances between a set of (e.g. randomly selected) base points and corresponding nearest points on the 3D object.
[0047] As a non-limiting example, the scene interaction component 140 may define a set of 3D base points (e.g., randomly selected, aligned to voxels of a 3D grid) that may be used as a frame of reference. To encode a target object (e.g., a 3D model of a chair) and / or a corresponding portion (e.g., a 3D cutout or pattern) of the 3D scene (e.g., the environment of the character and / or the target object), the scene interaction component 140 may identify and select the nearest vertex of the target object for each base point, determine the distance between each base point and the corresponding nearest vertex, and generate a (e.g., concatenated) representation of these distances. To encode a representation of the interaction (e.g.,To determine proximity (e.g., touch and / or proximity) between the character and the target object, the scene interaction component 140 may use the selected vertices of the target object as a set of 3D base points, and for each such point, identify and select the nearest vertex of a 3D representation of the character, determine the distance between each such point and the corresponding nearest vertex, and generate a (e.g., concatenated) representation of these distances. The scene interaction component 140 may therefore generate and use a representation of one or more of the above as the object interaction data 315, which may associate a corrected representation of proximity and / or touch with each of a plurality of 3D points, and the scene interaction component 140 may apply the object interaction data 315 to the layer(s) 320 (e.g.,a 3D convolutional neural network) to encode the object interaction data 315 into a 3D grid and extract a set of 3D features (e.g., an encoded 3D feature map) representing the object interaction data 315. This is simply intended to be an example, and known ways of encoding a representation of a 3D structure of a target object, surface, or other portion of a 3D scene or an object in a 3D scene and / or a representation of an interaction between a character and the target object, surface, or other portion of the 3D scene may be implemented within the scope of the present disclosure.
[0048] Continuing the above example, to inject the interaction scene data 315 into the pre-trained base diffusion model, the scene-aware diffusion model 300 may sample features st1...stN by scanning the coded feature map at elements that correspond to the 3D positions of the joints at the waypoints of the movement sequence xt1...xtN at a particular diffusion step t (e.g., features at the 3D joint locations being denoised may be sampled, which may change at each diffusion step t based on the denoised motion sequence predicted at the previous step t+1). Therefore, the scene-aware diffusion model 300 may estimate the state of the motion sequence xt1...xtN at a given diffusion step t and corresponding sampled features st1...stN the object interaction data 315 to create a denoised motion sequence x^01...x^0N to predict.
[0049] In some embodiments, the base diffusion model (e.g., layer(s) 210, the transformer encoder 230, and layer(s) 240) of the Fig. 3 may be pre-trained on motion data without scene data, and the scene-aware diffusion model 300 may be tuned (e.g., fine-tuned) on motion scene data (e.g., input training data comprising noisy motion sequences and object interaction data 315 and corresponding denoised ground truth motion sequences). For example, the scene-aware component of the scene-aware diffusion model 300 (e.g., layer(s) 320, the transformer encoder 235, one or more layers connecting the transformer encoder 235 to the transformer encoder 230) may be initialized (e.g., with initial values for trainable parameters, such as zeroed weights), the base diffusion model may be frozen or locked (with padlocks in Fig. 3), and the scene-aware diffusion model 300 may be trained using paired motion scene training data to learn values for the trainable parameters of the scene-aware component. In general, any known motion scene dataset may be used and / or any technique may be used to generate a motion scene dataset with input and ground truth training data corresponding to the scene-aware diffusion model 300. However, by incorporating a scene-aware component with a pre-trained diffusion model, the resulting scene-aware diffusion model 300 may be fine-tuned to more constrained motion scene data, enabling the generation of more accurate and scene-aware human motion on far less motion scene data than previous techniques.
[0050] In some embodiments, training data that pairs (e.g., captured) motion data with a corresponding 3D object being interacted with may be generated using data augmentation to retarget motion originally captured with respect to one object to a different object (e.g., retargeting captured motion data when sitting down from a particular chair to a different chair). In contrast to prior techniques that retarget touch locations (e.g., the locations where arms or pelvis touch the chair) to target locations where corresponding joints of a skeletal structure touch the target object, in some embodiments, the surface structure of the character's body may be retargeted using a 3D model (e.g.,a 3D mesh) to touch locations that can be realigned to target locations where corresponding locations on the body surface touch the target object. Therefore, the resulting realigned moving object interaction data is more accurate than previous techniques. As a result, the scene-aware diffusion model 300 can be fine-tuned using this training data, which should increase the accuracy of the resulting generated motion.
[0051] Therefore, and returning to Fig. 1, the scene interaction component 140 of the Fig. 1 the scene-aware diffusion model 300 of the Fig. 3 to generate and / or denoise a motion sequence 145 representing one or more positions and / or orientations of one or more joints at each of a plurality of 2D or 3D waypoints along the path that interacts with the target object (e.g., the chair 124) in the 3D scene 110. In some embodiments, the motion sequence 145 may represent positions and orientations for a plurality of joints in a skeletal structure of the animated character for each of the waypoints. In this example, the scene interaction component 140 may therefore use these positions and orientations to generate an animation of the character's body as it progresses through the waypoints of the motion sequence 145 without requiring motion inking. The scene interaction component 140 may therefore animate the character moving from the first target point 128 to the second target point 130 in the 3D scene 110.
[0052] With reference to the Fig. 4 and Fig. 5, each block of the methods 400 and 500 described herein comprises a computational process that may be performed using any combination of hardware, firmware, and / or software. For example, various functions may be performed by a processor executing instructions stored in memory. The methods 400 and 500 may also be embodied as computer-usable instructions stored on computer storage media. The methods 400 and 500 may be provided by a standalone application, a service, or a hosted service (standalone or in combination with another hosted service), or a plug-in to another product, to name a few. Additionally, the methods 400 and 500 are described, by way of example, with reference to the motion generation pipeline 100 of the Fig. 1. However, these methods may additionally or alternatively be performed by any system or combination of systems, including, but not limited to, the systems described herein.
[0053] Fig. 4 is a flowchart illustrating a method 400 for generating a scene-aware motion representation according to some embodiments of the present disclosure. The method 400 includes, at block B402, generating, based on at least processing a representation of at least a portion of a three-dimensional (3D) scene using a diffusion model comprising a scene-aware component and a pre-trained motion diffusion model, a scene-aware motion representation comprising one or more orientations of one or more joint waypoints along one or more paths of a character in the 3D scene.
[0054] Referring, for example, to the motion generation pipeline 100 of Fig. 1, the scene navigation component 132 may use the scene-aware diffusion model 200 of the Fig. 2 to predict a motion sequence representing positions and / or orientations of one or more joints at one or more waypoints along the path 127 from the starting point 126 to the first destination point 128. More specifically, the scene-aware diffusion model 200 of the Fig. 2, a scene-aware component (e.g., layer(s) 220, the transformer encoder 235, one or more layers connecting the transformer encoder 235 to the transformer encoder 230), and a pre-trained diffusion model (e.g., layer(s) 210, the transformer encoder 230, and layer(s) 240), and the scene-aware component may form an input channel for the scene data 215, which may represent a 2D or 3D structure of at least a portion of the 3D scene 110 at a particular time (e.g., a 2D or 3D occupancy grid, a ground map, a height map, a patch of any of the foregoing, such as an egocentric patch, classification data, such as a semantic segmentation representing any number of classes of objects or other parts of the scene, etc.). Therefore, the scene-aware diffusion model 200 can determine the state of the motion sequence xt1...xtN at a given diffusion step t and corresponding sampled features ft1...ftN of the scene data 215 to create a denoised motion sequence x^01...x^0N to predict.
[0055] In another example, the scene interaction component 140 may use the scene-aware diffusion model 300 of the Fig. 3 to predict a motion sequence representing positions and / or orientations of one or more joints at one or more waypoints along the path 129 from the first target point 128 to the second target point 130. More specifically, the scene-aware diffusion model 300 of the Fig. 3 may comprise a scene-aware component (e.g., layer(s) 320, the transformer encoder 235, one or more layers connecting the transformer encoder 235 to the transformer encoder 230) and a pre-trained diffusion model (e.g., layer(s) 210, the transformer encoder 230, and layer(s) 240). The scene-aware component may form an input channel for the object interaction data 315 encoding a representation of a 3D structure of a target object, surface, or other portion of a 3D scene or an object in a 3D scene and / or a representation of an interaction (e.g., a BPS representation of touch and / or proximity) between a character and the target object, surface, or other portion of the 3D scene. Therefore, the scene-aware diffusion model 300 may determine the state of the motion sequence xt1...xtN at a given diffusion step t and corresponding sampled features st1...stN the object interaction data 315 to create a denoised motion sequence x^01...x^0N to predict.
[0056] Fig. 5 is a flowchart illustrating a method 500 for generating a diffusion model for motion according to some embodiments of the present disclosure. The method 500 includes, at block B502, pretraining a diffusion model using motion data. Referring to Fig. 2, the scene-aware diffusion model 200 may, for example, comprise a base diffusion model (e.g., layer(s) 210, the transformer encoder 230, and layer(s) 240), which may be pre-trained (e.g., by the computing device 600 of the Fig. 6) to motion data without scene data using any known technique.
[0057] The method 500 includes, at block B504, connecting a scene-aware component to the pre-trained diffusion model. With reference to Fig. 2, a scene-aware component comprising layer(s) 220 and a transformer encoder 235 may be integrated into or connected to the pre-trained base diffusion model (e.g., using the computing device 600 of the Fig. 6) by forming an input channel for scene data 215. In some embodiments, transformer encoder 235 may be connected to transformer encoder 230 via one or more (e.g., linear) layers followed by an addition (e.g., a residual or skip connection). Layer(s) 220, transformer encoder 235, and / or the one or more layers connecting transformer encoder 235 to transformer encoder 230 may be initialized (e.g., with initial values for trainable parameters, such as weights set to zero).
[0058] The method 500 includes, at block B506, tuning the resulting diffusion model using motion scene training data. Referring to Fig. 2, for example, the scene-aware diffusion model 200, the base diffusion model, may be frozen or locked (e.g., by the computing device 600 of the Fig. 6), and the scene-aware diffusion model 200 may (e.g., by the computing device 600 of the Fig. 6) be trained using paired motion scene training data to learn values for the trainable parameters of the scene-aware component. In general, any known motion scene dataset can be used and / or any technique can be used to generate a motion scene dataset with input and ground truth training data corresponding to the scene-aware diffusion model 200. However, by incorporating a scene-aware component with a pre-trained diffusion model, the resulting scene-aware diffusion model 200 can be fine-tuned to more constrained motion scene data, enabling the generation of more accurate and scene-aware human motion on far less motion scene data than previous techniques.
[0059] The systems and methods described herein may be used for a variety of purposes, including, but not limited to, machine control, machine locomotion, machine propulsion, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twinning, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actuator simulation and / or digital twinning, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets, cloud computing, and / or other suitable applications.
[0060] The disclosed embodiments may be used in a variety of different systems, such as automotive systems (e.g., a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine), systems implemented using a robot, aviation systems, media systems, boat systems, intelligent area surveillance systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using an edge device, systems including one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems including one or more language models,System implementations such as one or more large language models (LLMs) for performing light transport simulation, systems for performing collaborative content creation for 3D assets, systems implemented at least partially using cloud computing resources, and / or other types of systems. EXAMPLE CALCULATION DEVICE
[0061] Fig. 6 is a block diagram of exemplary computing device(s) 600 suitable for use in implementing some embodiments of the present disclosure. Computing device 600 may include an interconnect system 602 that directly or indirectly couples the following devices: memory 604, one or more central processing units (CPUs) 606, one or more graphics processing units (GPUs) 608, a communications interface 610, input / output (I / O) ports 612, input / output components 614, a power supply 616, one or more presentation components 618 (e.g., display(s)), and one or more logic units 620. In at least one embodiment, the one or more computing devices 600 may include one or more virtual machines (VMs), and / or each of the components thereof may include virtual components (e.g., virtual hardware components).As non-limiting examples, one or more of the GPUs 608 may include one or more vGPUs, one or more of the CPUs 606 may include one or more vCPUs, and / or one or more of the logic units 620 may include one or more virtual logic units. Thus, a computing device 600 may include discrete components (e.g., a full GPU associated with the computing device 600), virtual components (e.g., a portion of a GPU associated with the computing device 600), or a combination thereof.
[0062] Although the different blocks of the Fig. 6 are shown as being connected with lines via the interconnect system 602, this is not intended to be limiting and is for clarity only. For example, in some embodiments, a presentation component 618, such as a display device, may be considered an I / O component 614 (e.g., if the display is a touchscreen). As another example, the CPUs 606 and / or GPUs 608 may include memory (e.g., memory 604 may be representative of a storage device in addition to the memory of the GPUs 608, the CPUs 606, and / or other components). In other words, the computing device of the Fig. 6 is merely illustrative. No distinction is made between categories such as ‘workstation’, ‘server’, ‘laptop’, ‘desktop’, ‘tablet’, ‘client device’, ‘mobile device’, ‘handheld device’, ‘game console’, ‘electronic control unit (ECU)’, ‘virtual reality system’ and / or other types of device or system, as all are considered to be within the scope of the computing device of the Fig. 6 can be considered.
[0063] The interconnect system 602 may represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnect system 602 may include one or more types of buses or links, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or any other type of bus or link. In some embodiments, there are direct connections between components. As one example, the CPU 606 may be directly connected to the memory 604. Further, the CPU 606 may be directly connected to the GPU 608. For a direct or point-to-point connection between components, the interconnect system 602 may include a PCIe link to establish the connection.In these examples, no PCI bus needs to be included in the computing device 600.
[0064] Memory 604 may include a variety of computer-readable media. The computer-readable media may be any available media accessible by computing device 600. The computer-readable media may include both volatile and non-volatile media, and removable and non-removable media. By way of example and without limitation, the computer-readable media may include computer storage media and communication media.
[0065] The computer storage media may include both volatile and non-volatile media and / or removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, and / or other data types. For example, the memory 604 may store computer-readable instructions (e.g., representing one or more programs) and / or one or more program elements, such as an operating system. Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, Digital Versatile Disks (DVD) or other optical disk storage, magnetic cartridges, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by the computing device 600.As used herein, computer storage media do not include signals per se.
[0066] The computer storage media may embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave, or other transport mechanism, and includes any media for conveying information. The term "modulated data signal" may refer to a signal, one or more characteristics of which are adjusted or changed such that information is encoded in the signal. The computer storage media may include, by way of example and not limitation, wired media, such as a wired network or a direct-wired connection, and wireless media, such as acoustic, RF, infrared, or other wireless media. Combinations of the above should also be included within the scope of computer-readable media.
[0067] The CPU(s) 606 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 600 to perform one or more of the methods and / or processes described herein. The CPU(s) 606 may each include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of handling a variety of software threads concurrently. The CPU(s) 606 may include any type of processor and may include different types of processors depending on the type of computing device 600 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers).Depending on the type of computing device 600, the processor may be, for example, an Advanced RISC Machines (ARM) processor implemented using a Reduced Instruction Set Computing (RISC) processor or an x86 processor implemented using Complex Instruction Set Computing (CISC). Computing device 600 may include one or more CPUs 606 in addition to one or more microprocessors or supplemental coprocessors, such as math coprocessors.
[0068] In addition or alternatively to the one or more CPUs 606, the one or more GPUs 608 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 600 to perform one or more of the methods and / or processes described herein. One or more of the GPUs 608 may be an integrated GPU (e.g., with one or more of the CPUs 606) and / or one or more of the GPUs 608 may be a discrete GPU. In embodiments, one or more of the GPU(s) 608 may be a coprocessor of one or more of the CPU(s) 606. The GPU(s) 608 may be used by the computing device 600 to render graphics (e.g., 3D graphics) or to perform general purpose computations. For example, the GPU(s) 608 may be used for general purpose computations on GPUs (GPGPU).The GPU(s) 608 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU(s) 608 may generate pixel data for output images in response to rendering commands (e.g., rendering commands from the CPU(s) 606 received via a host interface). The GPU(s) 608 may include graphics memory, such as display memory for storing pixel data or other useful data, such as GPGPU data. The display memory may be included as part of the memory 604. The GPU(s) 608 may include two or more GPUs operating in parallel (e.g., via a link). The link may connect the GPUs directly (e.g., using NVLINK) or connect the GPUs via a switch (e.g., using NVSwitch).When combined, each GPU can generate 608 pixel data or GPGPU data for different portions of an output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can include its own memory or can share memory with other GPUs.
[0069] In addition to or alternatively to the one or more CPUs 606 and / or the one or more GPUs 608, the one or more logic units 620 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 600 to perform one or more of the methods and / or processes described herein. In embodiments, the CPU(s) 606, the GPU(s) 608, and / or the logic unit(s) 620 may separately or jointly perform any combination of methods, processes, and / or portions thereof. One or more of the logic units 620 may be part of and / or integrated with one or more of the CPUs 606 and / or one or more of the GPUs 608, and / or one or more of the logic units 620 may be discrete components or otherwise external to the CPUs 606 and / or the GPUs 608.In embodiments, one or more of the logic units 620 may be a co-processor of one or more of the CPUs 606 and / or one or more of the GPUs 608.
[0070] Examples of the logic unit(s) 620 include one or more processing cores and / or components thereof, such as data processing units (DPUs), tensor cores (TCs), tensor processing units (TPUs), pixel visual cores (PVCs), vision processing units (VPUs), graphics processing clusters (GPCs), texture processing clusters (TPCs), streaming multiprocessors (SMs), tree traversal units (TTUs), artificial intelligence accelerators (AIAs), deep learning accelerators (DLAs), arithmetic logic units (ALUs), application-specific integrated circuits (Application-Specific Integrated Circuits, ASICs), Floating Point Units (FPUs),Input / output (I / O) elements, peripheral component interconnect (PCI) or PCI Express (PCIe) elements, and / or the like.
[0071] The communication interface 610 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 600 to communicate with other computers over an electronic network, including wired and / or wireless communication. The communication interface 610 may include components and functionality to enable communication over a number of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communicating over Ethernet or InfiniBand), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.In one or more embodiments, the one or more logic units 620 and / or the communication interface 610 may include one or more data processing units (DPUs) to transfer data received over a network and / or via the interconnect system 602 directly to one or more GPUs 608 (e.g., a memory thereof).
[0072] The I / O ports 612 may enable the computing device 600 to be logically coupled to other devices, including the I / O components 614, the one or more presentation components 618, and / or other components, some of which may be built into (e.g., integrated) the computing device 600. Illustrative I / O components 614 include a microphone, a mouse, a keyboard, a joystick, a game pad, a game controller, a satellite dish, a scanner, a printer, a wireless device, etc. The I / O components 614 may provide a natural user interface (NUI) that processes air gestures, speech, or other physiological inputs generated by a user. In some cases, the inputs may be transmitted to a suitable network element for further processing.An NUI may implement any combination of speech capture, stylus capture, facial recognition, biometric recognition, both on-screen and off-screen gesture recognition, air gestures, head and eye tracking, and touch sensing (as further described below) associated with a display of the computing device 600. The computing device 600 may include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof, for gesture capture and recognition. Additionally, the computing device 600 may include accelerometers or gyroscopes (e.g., as part of an inertial measurement unit (IMU)) that enable motion capture. In some examples, the outputs of the accelerometers and gyroscopes from the computing device 600 may be used to render immersive augmented reality or virtual reality.
[0073] The power supply 616 may include a wired power supply, a battery power supply, or a combination thereof. The power supply 616 may provide power to the computing device 600 to enable the functioning of the components of the computing device 600.
[0074] The presentation component(s) 618 may include a display (e.g., a monitor, a touchscreen, a television screen, a heads-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The one or more presentation components 618 may receive data from other components (e.g., the one or more GPUs 608, the one or more CPUs 606, DPUs, etc.) and output the data (e.g., as an image, video, audio, etc.). EXEMPLARY DATA CENTER
[0075] Fig. Figure 7 illustrates an example data center 700 that may be used in at least one embodiment of the present disclosure. Data center 700 may include a data center infrastructure layer 710, a framework layer 720, a software layer 730, and / or an application layer 740.
[0076] As in Fig. 7, the data center infrastructure layer 710 may include a resource orchestrator 712, clustered compute resources 714, and node compute resources ("node CRs") 716(1)-716(N), where "N" represents any positive integer. In at least one embodiment, the node CRs 716(1)-716(N) may include, but are not limited to, any number of central processing units (CPUs) or other processors (including DPUs, accelerators, field-programmable gate arrays (FPGAs), graphics processing units or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memories), storage devices (e.g., solid-state or disk drives), network input / output (NW I / O) devices, network switches, virtual machines (VMs), power modules, and / or cooling modules, etc. In some embodiments, one or more Node CRs of the Node CRs 716(1)-716(N) may correspond to a server having one or more of the computing resources mentioned above. Furthermore, in some embodiments, the node CRs 716(1)-716(N) may include one or more virtual components, such as vGPUs, vCPUs, and / or the like, and / or one or more of the node CRs 716(1)-716(N) may correspond to a virtual machine (VM).
[0077] In at least one embodiment, the grouped computing resources 714 may include separate groupings of node CRs 716 housed in one or more racks (not shown) or in many racks in data centers in different geographic locations (also not shown). Separate groupings of node CRs 716 within grouped computing resources 714 may include grouped computing, networking, memory, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, multiple node CRs 716, including CPUs, GPUs, DPUs, and / or other processors, may be grouped in one or more racks to provide computing resources to support one or more workloads.The one or more racks may also include any number of power modules, cooling modules, and / or network switches in any combination.
[0078] The resource orchestrator 712 may configure or otherwise control one or more node CRs 716(1)-716(N) and / or clustered computing resources 714. In at least one embodiment, the resource orchestrator 712 may include a software design infrastructure (SDI) management entity for the data center 700. The resource orchestrator 712 may include hardware or software, or a combination thereof.
[0079] In at least one embodiment, as in Fig. 7, the framework layer 720 may include a job scheduler 728, a configuration manager 734, a resource manager 736, and / or a distributed file system 738. The framework layer 720 may include a framework for supporting software 732 of the software layer 730 and / or one or more applications 742 of the application layer 740. The software 732 or the application(s) 742 may each include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 720 may be some type of free and open source software web application framework, such as, but not limited to, Apache Spark™ (hereinafter "Spark"), which may utilize a distributed file system 738 for processing large amounts of data (e.g., "Big Data").In at least one embodiment, the job scheduler 728 may include a Spark driver to facilitate scheduling workloads supported by different layers of the data center 700. The configuration manager 734 may be capable of configuring different layers, such as the software layer 730 and the framework layer 720, including Spark and the distributed file system 738 to support large-scale data processing. The resource manager 736 may be capable of managing clustered or grouped compute resources mapped to or associated with the distributed file system 738 and the job scheduler 728 for support. In at least one embodiment, the clustered or grouped compute resources may include the grouped compute resources 714 at the infrastructure layer 710 of the data center.The resource manager 736 may coordinate with the resource orchestrator 712 to manage these allocated or assigned computing resources.
[0080] In at least one embodiment, the software 732 included in the software layer 730 may include software used by at least portions of the node CRs 716(1)-716(N), clustered computing resources 714, and / or the distributed file system 738 of the framework layer 720. One or more types of software may include, but are not limited to, Internet web search software, email virus scanning software, database software, and streaming video content software.
[0081] In at least one embodiment, the application(s) 742 included in the application layer 740 may include one or more types of applications used by at least portions of the node CRs 716(1)-716(N), clustered compute resources 714, and / or the distributed file system 738 of the framework layer 720. One or more types of applications may include, but are not limited to, any number of genomic applications, cognitive computation, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in connection with one or more embodiments.
[0082] In at least one embodiment, the configuration manager 734, the resource manager 736, and / or the resource orchestrator 712 may implement any number and type of self-modifying actions based on any amount and type of data collected in any technically feasible manner. Self-modifying actions may relieve a data center operator of the data center 700 from potentially making poor configuration decisions and potentially avoid underutilized and / or poorly performing sections of a data center.
[0083] Data center 700 may include tools, services, software, or other resources to train one or more machine learning models or to predict or infer information using one or more machine learning models according to one or more embodiments described herein. For example, one or more machine learning models may be trained by calculating weighting parameters according to a neural network architecture using software and / or computational resources described above with reference to data center 700.In at least one embodiment, trained or deployed machine learning models corresponding to one or more neural networks may be used to infer or predict information using resources described above with reference to data center 700 using weighting parameters calculated by one or more training techniques such as those described herein.
[0084] In at least one embodiment, the data center 700 may utilize CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, and / or other hardware (or virtual computing resources equivalent thereto) to perform training and / or inference using the resources described above. Furthermore, one or more of the software and / or hardware resources described above may be configured as a service to allow users to train or perform information inference, such as image recognition, speech recognition, or other useful intelligent services. EXAMPLE NETWORK ENVIRONMENTS
[0085] Network environments suitable for use in implementing embodiments of the disclosure may include one or more client devices, servers, network attached storage (NAS), other backend devices, and / or device types. The client devices, servers, and / or other device types (e.g., each device) may be connected to one or more instances of the computing device(s) 600 of the Fig. 6, e.g., each device may include similar components, features, and / or functionality of the computing device(s) 600. If backend devices (e.g., servers, NAS, etc.) are implemented, the backend devices may additionally be included as part of a data center 700, an example of which is described in more detail herein with reference to Fig. 7 is described.
[0086] The components of a network environment can communicate with each other over one or more networks, which can be wired, wireless, or both. The network can include multiple networks or a network of networks. By way of example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks such as the Internet and / or public networks such as a public switched telephone network (PSTN), and / or one or more private networks. If the network includes a wireless telecommunications network, components such as a base station, a communications tower, or even access points (as well as other components) can provide wireless connectivity.
[0087] Compatible network environments may include one or more peer-to-peer network environments, in which case a server may not be included in a network environment, and one or more client-server network environments, in which case one or more servers may be included in a network environment. In peer-to-peer network environments, the functionality described herein may be implemented with reference to one or more servers on any number of client devices.
[0088] In at least one embodiment, a network environment may include one or more cloud-based network environments, a distributed computing environment, a combination thereof, etc. A cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. A framework layer may include a framework for supporting software of a software layer and / or one or more applications of an application layer. The software or application(s) may each include web-based service software or applications. In embodiments, one or more of the client devices may use the web-based service software or applications (e.g.,by accessing the service software and / or applications through one or more application programming interfaces (APIs). The framework layer may be some type of free and open-source software web application framework, e.g., using, but not limited to, a distributed file system for processing large amounts of data (e.g., "big data").
[0089] A cloud-based network environment may provide cloud computing and / or cloud storage that performs any combination of the computing and / or data storage functions described herein (or one or more portions thereof). Any of these various functions may be distributed across multiple locations from central or core servers (e.g., from one or more data centers distributed across a state, region, country, the world, etc.). When a connection to a user (e.g., a client device) is relatively close to one or more edge servers, one or more core servers may allocate at least a portion of the functionality of the edge server(s). A cloud-based network environment may be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0090] The client device(s) may include at least some of the components, features, and functionality of the example computing device(s) 600 described herein with reference to Fig.6. By way of example and not limitation, a client device may be a personal computer (PC), a laptop computer, a mobile device, a smartphone, a tablet computer, a wearable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality headset, a global positioning system (GPS) or device, a video player, a video camera, a surveillance device or system, a vehicle, a boat, an air vehicle, a virtual machine, a drone, a robot, a handheld communication device, a hospital device, a gaming device or system, an entertainment device, a vehicle computing system, an embedded system controller, a household appliance, a consumer electronics system, a workstation, an edge device, any combination of these devices, or any other suitable device.
[0091] The disclosure of this application also includes the following numbered clauses: Clause 1. A processor comprising: one or more processing units for generating, based at least on processing a representation of at least a portion of a three-dimensional scene (3D scene) using a diffusion model comprising a scene-aware component and a pre-trained motion diffusion model, a representation of the scene-aware motion comprising one or more orientations of one or more joint waypoints along one or more paths of a character at least partially mapped in the 3D scene. Clause 2. The processor of clause 1, wherein the one or more processing units are further to generate the diffusion model based at least on adding the scene-aware component to the pre-trained motion diffusion model and fine-tuning the diffusion model using motion scene training data. Clause 3. The processor of any preceding clause, wherein processing using the diffusion model comprises injecting a top-down height map of the 3D scene into the pre-trained motion diffusion model. Clause 4. The processor of any preceding clause, wherein processing using the diffusion model comprises injecting a 3D point cloud representing at least a portion of a 3D object in the 3D scene into the pre-trained motion diffusion model. Clause 5. The processor of any preceding clause, wherein processing using the diffusion model comprises injecting classification data representing one or more classified locations of one or more classified objects in the 3D scene into the pre-trained motion diffusion model. Clause 6. The processor of any preceding clause, wherein processing using the diffusion model comprises injecting classification data representing one or more classified locations of one or more other characters or one or more audio sources in the 3D scene into the pre-trained motion diffusion model. Clause 7. The processor of any preceding clause, wherein the one or more processing units are further to update the diffusion model using training data generated based at least on reorienting motion data comprising one or more contact locations with a first object to one or more corresponding locations at which a modeled body surface comes into contact with a target object. Clause 8. The processor according to any preceding clause, wherein the processor consists of at least one of the following: a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for generating synthetic data; a system for generating synthetic data using AI; a system that contains one or more virtual machines (VMs); a system that is implemented at least partially in a data center; or a system that is implemented at least in part using cloud computing resources. Clause 9 A system comprising one or more processing units for generating, based at least on processing a representation of at least a portion of a three-dimensional (3D) scene using a diffusion model, a representation of scene-aware motion corresponding to a character at least partially depicted in the 3D scene. Clause 10. The system of clause 9, wherein the one or more processing units are further to generate the diffusion model based at least on adding a scene-aware component to a pre-trained motion diffusion model and fine-tuning the diffusion model using motion scene training data. Clause 11. The system of Clause 9 or Clause 10, wherein processing using the diffusion model comprises injecting the 3D scene of a height map from top to bottom into the pre-trained diffusion model of the diffusion model. Clause 12. The system of any of clauses 9-11, wherein processing using the diffusion model comprises injecting a 3D point cloud representing at least a portion of a 3D object in the 3D scene into the pre-trained diffusion model of the diffusion model. Clause 13. The system of any of clauses 9-12, wherein processing using the diffusion model comprises injecting classification data representing one or more classified locations of one or more classified objects in the 3D scene into the pre-trained diffusion model of the diffusion model. Clause 14. The system of any of clauses 9-13, wherein processing using the diffusion model comprises injecting classification data representing one or more classified locations of one or more other characters or one or more audio sources in the 3D scene into the pre-trained diffusion model of the diffusion model. Clause 15. The system of any of clauses 9-14, wherein the one or more processing units are further to update the diffusion model using training data generated based at least on reorienting motion data comprising one or more contact locations with a first object to one or more corresponding locations at which a modeled body surface comes into contact with a target object. Clause 16. The system under any of Clauses 9-15, where the system consists of at least one of the following: a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote surgery; a system for performing real-time streaming; a system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for generating synthetic data; a system for generating synthetic data using AI; a system that contains one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system implemented at least in part using cloud computing resources. Clause 17 A procedure comprising: Generating, based on at least injecting a representation of at least a portion of a three-dimensional scene (3D scene) into a pre-trained diffusion model, a representation of one or more orientations of one or more waypoints along one or more paths of a character in the 3D scene. Clause 18. The method of clause 17, further comprising generating a diffusion model based at least on adding a scene-aware component to the pre-trained diffusion model and fine-tuning the diffusion model using motion scene training data. Clause 19. The method of clause 17 or clause 18, further comprising updating a diffusion model comprising the pre-trained diffusion model using training data generated based at least on reorienting motion data comprising one or more contact locations with a first object to one or more corresponding locations at which a modeled body surface comes into contact with the target object. Clause 20. The procedure of any of Clauses 17 to 19, where the procedure is carried out by at least one of the following: a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote surgery; a system for performing real-time streaming; a system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for generating synthetic data; a system for generating synthetic data using AI; a system that contains one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system implemented at least in part using cloud computing resources.
[0092] The disclosure may be described in the general context of computer code or machine-usable instructions, including computer-executable instructions, such as program modules, executed by a computer or other machine, such as a personal data assistant or other handheld device. In general, program modules, which include routines, programs, objects, components, data structures, etc., refer to code that performs particular tasks or implements particular abstract data types. The disclosure may be implemented in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computing, more special-purpose computing devices, etc. The disclosure may also be implemented in distributed computing environments in which tasks are performed by remote processing devices linked by a communications network.
[0093] As used herein, any reference to "and / or" in reference to two or more elements should be construed to mean only one element or a combination of elements. For example, "element A, element B, and / or element C" may mean element A only, element B only, element C only, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Additionally, "at least one of element A or element B" may include at least one element A, at least one element B, or at least one element A and at least one element B. Further, "at least one element A and element B" may include at least one element A, at least one element B, or at least one element A and at least one element B.
[0094] The subject matter of the present disclosure is described herein with a certain degree of specificity to comply with legal requirements. However, the description itself is not intended to limit the scope of this disclosure. Instead, the inventors have contemplated that the claimed subject matter could be embodied in other ways to include different steps or combinations of steps similar to those described in this specification in connection with other present or future technologies. Furthermore, and although the terms "step" and / or "block" may be used herein to refer to different elements of employed methods, the terms should not be construed as implying a particular order among or between various steps disclosed herein, except and only where the order of individual steps is explicitly described.
[0095] It is to be understood that aspects and embodiments described above are purely exemplary and that modifications of details may be made within the scope of the claims.
[0096] Each device, method, and feature disclosed in the specification, and (where appropriate) the claims and drawings, may be provided independently or in any suitable combination.
[0097] Reference signs appearing in the claims are for illustrative purposes only and do not limit the scope of the claims.
Claims
[1] Processor comprising: one or more processing units for generating, based at least on processing a representation of at least a portion of a three-dimensional scene (3D scene) using a diffusion model comprising a scene-aware component and a pre-trained motion diffusion model, a representation of the scene-aware motion comprising one or more orientations of one or more joint waypoints along one or more paths of a character at least partially mapped in the 3D scene. [2] The processor of claim 1, wherein the one or more processing units are further to generate the diffusion model based at least on adding the scene-aware component to the pre-trained motion diffusion model and fine-tuning the diffusion model using motion scene training data. [3] A processor according to any preceding claim, wherein processing using the diffusion model comprises injecting a top-down height map of the 3D scene into the pre-trained motion diffusion model. [4] A processor according to any preceding claim, wherein processing using the diffusion model comprises injecting a 3D point cloud representing at least a portion of a 3D object in the 3D scene into the pre-trained motion diffusion model. [5] A processor according to any preceding claim, wherein processing using the diffusion model comprises injecting classification data representing one or more classified locations of one or more classified objects in the 3D scene into the pre-trained motion diffusion model. [6] A processor according to any preceding claim, wherein processing using the diffusion model comprises injecting classification data representing one or more classified locations of one or more other characters or one or more audio sources in the 3D scene into the pre-trained motion diffusion model. [7] A processor according to any preceding claim, wherein the one or more processing units are further to update the diffusion model using training data generated based at least on reorienting motion data comprising one or more contact locations with a first object to one or more corresponding locations at which a modeled body surface comes into contact with a target object. [8] A processor according to any preceding claim, wherein the processor comprises at least one of the following: a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for generating synthetic data; a system for generating synthetic data using AI; a system that contains one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system implemented at least in part using cloud computing resources. [9] System comprising one or more processing units for generating, based at least on processing a representation of at least a portion of a three-dimensional scene (3D scene) using a diffusion model, a representation of scene-aware motion corresponding to a character at least partially depicted in the 3D scene. [10] The system of claim 9, wherein the one or more processing units are further to generate the diffusion model based at least on adding a scene-aware component to a pre-trained motion diffusion model and fine-tuning the diffusion model using motion scene training data. [11] The system of claim 9 or claim 10, wherein processing using the diffusion model comprises injecting a top-down height map of the 3D scene into the pre-trained diffusion model of the diffusion model. [12] The system of any of claims 9-11, wherein processing using the diffusion model comprises injecting a 3D point cloud representing at least a portion of a 3D object in the 3D scene into the pre-trained diffusion model of the diffusion model. [13] The system of any of claims 9-12, wherein processing using the diffusion model comprises injecting classification data representing one or more classified locations of one or more classified objects in the 3D scene into the pre-trained diffusion model of the diffusion model. [14] The system of any of claims 9-13, wherein processing using the diffusion model comprises injecting classification data representing one or more classified locations of one or more other characters or one or more audio sources in the 3D scene into the pre-trained diffusion model of the diffusion model. [15] The system of any of claims 9-14, wherein the one or more processing units are further to update the diffusion model using training data generated based at least on reorienting motion data comprising one or more contact locations with a first object to one or more corresponding locations at which a modeled body surface comes into contact with a target object. [16] System according to one of claims 9 to 15, wherein the system comprises at least one of the following: a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote surgery; a system for performing real-time streaming; a system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for generating synthetic data; a system for generating synthetic data using AI; a system that contains one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system implemented at least in part using cloud computing resources. [17] Method comprising: Generating, based on at least injecting a representation of at least a portion of a three-dimensional scene (3D scene) into a pre-trained diffusion model, a representation of one or more orientations of one or more waypoints along one or more paths of a character in the 3D scene. [18] The method of claim 17, further comprising generating a diffusion model based at least on adding a scene-aware component to the pre-trained diffusion model and fine-tuning the diffusion model using motion scene training data. [19] The method of claim 17 or claim 18, further comprising updating a diffusion model comprising the pre-trained diffusion model using training data generated based at least on reorienting motion data comprising one or more contact locations with a first object to one or more corresponding locations at which a modeled body surface comes into contact with the target object. [20] A method according to any one of claims 17 to 19, wherein the method is carried out by at least one of the following: a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote surgery; a system for performing real-time streaming; a system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for generating synthetic data; a system for generating synthetic data using AI; a system that contains one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system implemented at least in part using cloud computing resources.