METHOD FOR PHYSICALLY BASED ANIMATIONS STARTING FROM JOINTS WITH PARTIAL CONDITIONS

The method addresses the challenges of labor-intensive and non-realistic character animations by using a machine learning model to generate animations based on sparse joint constraints and force considerations, resulting in more efficient and realistic animations.

DE102024132811A1Pending Publication Date: 2025-05-15NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102024132811
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-09-10
Filing Date
2024-11-11
Publication Date
2025-05-15

AI Technical Summary

Technical Problem

Existing methods for creating computer-generated character animations are labor-intensive, time-consuming, and often result in animations that are not physically realistic, as they do not account for the forces acting on the character's joints.

Method used

A computer-implemented method that uses a trained machine learning model to generate animations by specifying the positions and orientations of a subset of a character's joints, rather than all joints, and takes into account the forces acting on the joints to produce more physically realistic animations.

Benefits of technology

The method enables the creation of animations that are more physically realistic and reduces the labor and time required for animation creation, as it can animate a character by indicating the positions and orientations of a subset of joints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

One embodiment of a method for animating characters comprises obtaining a first state of a character and one or more constraints for one or more movements associated with a subset of joints belonging to the character, generating a first action to be performed for the character via a trained machine learning model and based on the first state and the one or more constraints, and causing the character to perform the first action within a computer-based or physical environment.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUNDTechnical area

[0001] Embodiments of the present disclosure relate generally to robotics, virtual character control, artificial intelligence, and machine learning, and more particularly to methods for physics-based animations from joints with partial prerequisites. Description of the related art

[0002] Character animation is the process of creating a series of different poses, expressions, and / or actions for a character that can be performed sequentially. Character animations can be created in a variety of ways, including hand-drawn animations, stop-motion animations, and computer-generated animations.

[0003] Computer-generated character animations are typically created through a largely manual process, with animators using software to design and move three-dimensional (3D) virtual models of characters so that the characters can move within specific animation sequences. For example, an animator might use software to specify the positions and orientations of the joints of a character's head, torso, arms, etc., within a series of keyframes of a particular animation. To create a complete animation, the software might use kinematic modeling to calculate the positions and orientations of the same joints within frames that occur between the keyframes. The character can then be animated to move in a manner that follows the positions and orientations of the joints within the keyframes and the frames in between.

[0004] A disadvantage of the approach described above for creating computer-generated character animations is that the animator generally must specify the positions and orientations of all the character's joints within keyframes to create that character's animation. There are few, if any, conventional software programs that can automatically determine physically plausible positions and orientations for a character's joints that have not been specified by an animator in one of the keyframes. Because an animator must specify the positions and orientations of all a character's joints within the keyframes of an animation, manually creating character animations is typically very labor-intensive and time-consuming, and sometimes inaccurate.

[0005] Another disadvantage of the above approach to creating computer-generated character animations is that the kinematic modeling used to calculate the positions and orientations of joints within intermediate frames does not account for the forces that move those joints. Instead, kinematic modeling only calculates the joint movement required to transition between the positions and orientations of the joints within keyframes. Because forces are not taken into account, the resulting animations are often not physically realistic, which negatively impacts the overall visual quality.

[0006] As the above shows, more effective methods for generating computer-aided character animations are needed in the current state of the art. SUMMARY

[0007] An embodiment of the present disclosure sets forth a computer-implemented method for animating a character. The method includes obtaining a first state of a character and one or more constraints for one or more movements associated with a subset of joints associated with the character. The method further includes generating a first action for the character to perform using a trained machine learning model and based on the first state and the one or more constraints. The method further includes causing the character to perform the first action in a computer-based or physical environment.

[0008] Other embodiments of the present disclosure include, without limitation, one or more computer-readable media containing instructions for performing one or more aspects of the disclosed methods, and one or more computer systems for performing one or more aspects of the disclosed methods.

[0009] A technical advantage of the disclosed methods over the prior art is that the disclosed methods can animate a physical or virtual character by specifying the positions and orientations of a subset of a character's joints, rather than all of them, in any number of frames of an animation. Furthermore, the disclosed methods can produce animations that are more physically realistic than animations produced with kinematic models that do not account for the forces moving a character's joints. These technical advantages represent one or more technological improvements over the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order that the above-mentioned features of the various embodiments may be understood in detail, a more particular description of the inventive concepts briefly summarized above may be made by reference to various embodiments, some of which are illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings illustrate only typical embodiments of the inventive concepts and are therefore in no way limiting in scope, and that other equally effective embodiments may exist. Fig. 1 illustrates a computerized system configured to implement one or more aspects of the various embodiments; Fig. 2 is a more detailed representation of the machine learning server of Fig. 1 according to various embodiments; Fig. 3 is a more detailed representation of the computing device of Fig. 1 according to various embodiments; Fig. 4 is a more detailed representation of the model trainer of Fig. 1 according to various embodiments; Fig. 5 shows how the motion tracking model of Fig. 1 is trained according to various embodiments; Fig. 6 is a more detailed representation of the prior module of Fig. 5 according to various embodiments; Fig. 7 is a more detailed representation of the control application of Fig. 1 according to various embodiments; Fig. 8 is a more detailed illustration of how the motion tracking model of Fig. 7 is used to control a character according to various embodiments; Fig. 9 shows a flowchart of the method steps for training a motion tracking model according to various embodiments; and Fig. 10 is a flowchart of method steps for generating an animation of a character based on sparse motion constraints, according to various embodiments. DETAILED DESCRIPTION

[0011] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of the various embodiments. However, it will be apparent to one skilled in the art that the concepts may be practiced without one or more of these specific details. General overview

[0012] Embodiments of the present disclosure provide methods for animating characters using sparse motion constraints. In some embodiments, the sparse motion constraints may specify the positions and / or orientations of any number of joints of a character in any number of frames of an animation. Given the sparse motion constraints, a control application samples a latent prior distribution generated via a trained motion tracking model, taking into account the sparse motion constraints and a current state of a character, to obtain a sampled latent vector.The character may be a virtual character in a computer-based environment or a physical robot in a real-world environment, and the character's current state may be obtained from the computer-based environment or sensed using sensors in the real-world environment. The control application inputs the sampled latent vector and the character's current state to a controller of the motion tracking model to generate an action. The control application may then control the character within the computer-based environment or the real-world environment using the action. Controlling the character may result in an updated state of the character, and the above process may be repeated to generate another action for controlling the character using the updated state of the character, the sparse motion constraints, and the motion tracking model.

[0013] A model trainer trains the motion tracking model, which in some embodiments may be a variational autoencoder (VAE). In some embodiments, the model trainer first samples a motion and a time step within the motion from a set of motion recordings. The model trainer then samples a mask, which is used to mask out random joints and / or frames from the sampled motion. The model trainer computes an action for each frame using the motion tracking model, the masked motion, and a current state of the character. The model trainer simulates the character performing the action in the environment and obtains a motion of the character from the environment.Then, the model trainer calculates a reward based on a comparison of the obtained motion with the sampled motion, and the model trainer updates the parameters of the motion tracking model based on the reward.

[0014] Character animation techniques have many real-world applications. For example, these techniques can be used to animate a character in a virtual or extended reality (XR) environment, such as a gaming environment. As another example, these techniques can be used to control a physical robot in a real-world environment.

[0015] The above examples are not intended to be limiting in any way. As those skilled in the art will appreciate, the character animation techniques described herein can be used in any suitable application. System overview

[0016] Fig. Figure 1 shows a block diagram of a computer-based system 100 configured to implement one or more aspects of various embodiments. As illustrated, the system 100 includes a machine learning server 110, a data store 120, and a computing system 140 communicating over a network 130, which may be a wide area network (WAN) such as the Internet, a local area network (LAN), a cellular network, and / or any other suitable network.

[0017] As shown, a model trainer 116 executes on one or more processors 112 of the machine learning server 110 and is stored in a system memory 114 of the machine learning server 110. The processor(s) 112 receive user input from input devices, such as a keyboard or a mouse. In operation, the processor(s) 112 may include one or more primary processors of the machine learning server 110 that control and coordinate the operations of other system components. In particular, the processor(s) 112 may issue instructions that control the operation of one or more graphics processing units (GPUs) (not shown) and / or other parallel processing circuitry (e.g., parallel processing units, deep learning accelerators, etc.) that includes circuitry optimized for graphics and video processing, including, for example, video output circuitry.The GPU(s) may provide pixels to a display device, which may be any conventional cathode ray tube, liquid crystal display, light-emitting diode display, and / or the like.

[0018] The system memory 114 of the machine learning server 110 stores content, such as software applications and data, for use by the processor(s) 112 and GPU(s) and / or other processing units. The system memory 114 may be any type of memory capable of storing data and software applications, such as random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash ROM), or any suitable combination of the foregoing. In some embodiments, memory (not shown) may supplement or replace the system memory 114. The memory may include any number and type of external memory accessible by the processor 112 and / or the GPU.For example, the memory may include a secure digital card, an external flash memory, a portable CD read-only memory (CD-ROM), an optical storage device, a magnetic storage device, and / or any suitable combination of the foregoing devices.

[0019] The machine learning server 110 shown herein is for illustrative purposes only, and variations and modifications are possible without departing from the scope of the present disclosure. For example, the number of processors 112, the number of GPUs and / or other types of processing units, the number of system memories 114, and / or the number of applications contained in the system memory 114 may be changed as desired. Furthermore, the connection topology between the various units in Fig. 1 may be modified as desired. In some embodiments, any combination of the processor(s) 112, the system memory 114, and / or the GPU(s) may be included in and / or replaced by any type of virtual computing system, distributed computing system, and / or cloud computing environment, such as a public, private, or hybrid cloud system.

[0020] In some embodiments, the model trainer 116 is configured to train one or more machine learning models, including a motion tracking model 150 trained to generate actions for animating a character given a sparse set of joint constraints. Methods that the model trainer 116 may use to train the motion tracking model 150 are described below in connection with the Fig. 4-5 and 9. Training data and / or trained (or deployed) machine learning models, including motion tracking model 150, may be stored in data storage 120. In some embodiments, data storage 120 may include any device or devices for storage, such as one or more hard disk drives, one or more flash drives, optical storage, a network-attached storage (NAS), and / or a storage area network (SAN). Although illustrated as being accessible via network 130, in at least one embodiment, machine learning server 110 may include data storage 120.

[0021] To illustrate, the data store 120 also stores motion recordings 154. The motion recordings 154 are used to train the motion tracking model 152. In some embodiments, the motion recordings 154 include recorded human movements used to evaluate the generated movements of the motion tracking model 152. In various examples, the motion recordings 154 are compiled from various human activities captured, for example, by motion capture techniques.

[0022] As shown, a control application 146 using the trained motion tracking model 152 is stored in memory 144 and executed on the processor(s) 142 of the computing device 140. The control application 146 is described below in connection with the Fig. 7-8 and 10. To illustrate, the control application 146 uses the motion tracking model 152 to control a character 160 moving in an environment 170.

[0023] The environment 170 in which the character 160 performs actions can be either a computer-based environment or a physical environment. A computer-based environment can, in some embodiments, be simulated in any technically feasible way, for example, using a 3D engine, a generative model (e.g., a neural network) that predicts the next state after an action, etc. In a computer-based virtual 3D environment, the character 160 can, for example, navigate through a digital landscape, such as a simulation of a cityscape with flowing traffic and pedestrians, a fantasy world with dynamic terrain and interactive elements, and / or the like. Computer-based environments can be used in the development of video games, virtual reality (VR) applications, advanced AI training simulations, and / or the like.In a physical environment, the character 160, for example, a humanoid robot, may navigate real-world scenarios, for example, a robot may move through a warehouse to perform logistical operations, maneuver in a hospital to deliver supplies, operate in hazardous environments such as nuclear power plants where the presence of humans is risky, and / or the like.

[0024] Fig. 2 is a more detailed illustration of the machine learning server 110 of Fig. 1 according to various embodiments. In some embodiments, the machine learning server 110 may comprise any type of computing system, including, but not limited to, a server, a server platform, a desktop computer, a laptop, a handheld / mobile device, a digital kiosk, an in-vehicle infotainment system, and / or a portable device. In some embodiments, the machine learning server 110 is a server machine operating in a data center or cloud computing environment that provides scalable computing resources as a service over a network.

[0025] In some embodiments, the machine learning server 110 includes, among other things, the processor(s) 112 and the memory(s) 114 coupled to a parallel processing subsystem 212 via a memory bridge 205 and a communication path 206. The memory bridge 205 is further coupled to an I / O (input / output) bridge 207 via a communication path 206, and the I / O bridge 207 is in turn coupled to a switch 216.

[0026] In some embodiments, the I / O bridge 207 is configured to receive user input from optional input devices 208, such as a keyboard, a mouse, a touchscreen, sensor data analysis (e.g., evaluation of gestures, speech, or other information about one or more uses in a field of view or a sensory field of one or more sensors), and / or the like, and forward the input information to the processor(s) 112 for processing. In some embodiments, the machine learning server 110 may be a server machine in a cloud computing environment. In such embodiments, the machine learning server 110 may not include input devices 208, but may receive equivalent input data by receiving commands (e.g.,in response to one or more inputs from a remote computing device) in the form of messages transmitted over a network and received via a network adapter 218. In some embodiments, the switch 216 is configured to provide connections between the I / O bridge 207 and other components of the machine learning server 110, such as a network adapter 218 and various add-in cards 220 and 221.

[0027] In some embodiments, the I / O bridge 207 is coupled to a system disk 214, which may be configured to store content, applications, and data for use by the processor(s) 112 and the parallel processing subsystem 212. In some embodiments, the system disk 214 provides non-volatile storage for applications and data and may include fixed or removable hard disk drives, flash memory devices, and CD-ROM (Compact Disc Read-Only Memory), DVD-ROM (Digital Versatile Disc-ROM), Blu-ray, HD-DVD (High-Definition DVD), or other magnetic, optical, or solid-state storage devices. In some embodiments, other components, such as Universal Serial Bus or other connectors, CD drives, DVD drives, movie recorders, and the like, may also be connected to the I / O bridge 207.

[0028] In some embodiments, memory bridge 205 may be a northbridge chip and I / O bridge 207 may be a southbridge chip. Additionally, communication paths 206 and 213, as well as other communication paths within machine learning server 110, may be implemented using any technically suitable protocols, including, but not limited to, AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol known in the art.

[0029] In some embodiments, the parallel processing subsystem 212 includes a graphics subsystem that provides pixels to an optional display 210, which may be a conventional cathode ray tube, a liquid crystal display, a light-emitting diode display, and / or the like. In such embodiments, the parallel processing subsystem 212 may include circuitry optimized for graphics and video processing, including, for example, video output circuitry. Such circuitry may be integrated into one or more parallel processing units (PPUs), also referred to as parallel processors, included in the parallel processing subsystem 212.

[0030] In some embodiments, the parallel processing subsystem 212 includes circuitry optimized for general purpose and / or computational processing (e.g., that is undergoing optimization). Again, such circuitry may be included in one or more PPUs included in the parallel processing subsystem 212 and configured to perform such general purpose and / or computational operations. In other embodiments, the one or more PPUs included in the parallel processing subsystem 212 may be configured to perform graphics processing, general purpose processing, and / or computational processing operations. The system memory 114 includes at least one device driver configured to manage the processing operations of the one or more PPUs within the parallel processing subsystem 212.In addition, the system memory 114 contains the model trainer 116, which is described below in connection with the . Fig. 4-5 and 9. Although described herein primarily with respect to the model trainer 116, the methods disclosed herein may also be implemented, in whole or in part, in other software and / or hardware, such as the parallel processing subsystem 212.

[0031] In some embodiments, the parallel processing subsystem 212 may be integrated with one or more of the other elements of Fig. 2 to form a single system. For example, the parallel processing subsystem 212 may be integrated with the processor(s) 112 and other interconnect circuitry on a single chip to form a system on a chip (SoC).

[0032] In some embodiments, processor(s) 112 comprise the primary processor of machine learning server 110, which controls and coordinates the operations of other system components. In some embodiments, processor(s) 112 issue instructions that control the operation of PPUs. In some embodiments, communication path 213 is a PCI Express connection with dedicated lanes assigned to each PPU. Other communication paths may also be used. The PPU advantageously implements a highly parallel processing architecture, and the PPU may be provided with any amount of local parallel processing (PP) memory.

[0033] It should be understood that the system shown here is for illustrative purposes only and that variations and modifications are possible. The interconnect topology, including the number and arrangement of bridges, the number of processors 112, and the number of parallel processing subsystems 212, can be changed as desired. For example, in some embodiments, system memory 114 may be connected directly to processor(s) 112 rather than through memory bridge 205, and other devices may communicate with system memory 114 via memory bridge 205 and processor(s) 112. In other implementations, parallel processing subsystem 212 may be connected to I / O bridge 207 or directly to processor(s) 112 rather than to memory bridge 205.In still other embodiments, the I / O bridge 207 and the memory bridge 205 may be integrated into a single chip rather than existing as one or more discrete devices. In certain embodiments, one or more of the devices shown in FIG. Fig. 2 may be missing. For example, the switch 216 may be omitted and the network adapter 218 and the add-in cards 220, 221 may be connected directly to the I / O bridge 207. Finally, in certain embodiments, one or more of the components shown in Fig. 2 may be implemented as virtualized resources in a virtual computing environment, such as a cloud computing environment. In particular, in some embodiments, the parallel processing subsystem 212 may be implemented as a virtualized parallel processing subsystem. For example, the parallel processing subsystem 212 may be implemented as virtual graphics processing unit(s) (vGPU(s)) that render graphics on a virtual machine (VM) running on a server machine whose GPU and other physical resources are shared among one or more VMs.

[0034] Fig. 3 is a more detailed illustration of the computer system 140 of Fig. 1 according to various embodiments. In some embodiments, computing system 140 may comprise any type of computing system, including, but not limited to, a server, a server platform, a desktop computer, a laptop, a handheld / mobile device, a digital kiosk, an in-vehicle infotainment system, and / or a portable device. In some embodiments, computing system 140 is a server operating in a data center or cloud computing environment that provides scalable computing resources as a service over a network.

[0035] In some embodiments, computer system 140 includes, among other things, processor(s) 142 and memory(s) 144 coupled to a parallel processing subsystem 312 via a memory bridge 305 and a communication path 306. Memory bridge 305 is further coupled to an I / O (input / output) bridge 307 via a communication path 306, and I / O bridge 307 is in turn coupled to a switch 316.

[0036] In some embodiments, the I / O bridge 307 is configured to receive user input from optional input devices 308, such as a keyboard, a mouse, a touchscreen, sensor data analysis (e.g., evaluation of gestures, speech, or other information about one or more uses in a field of view or a sensory field of one or more sensors), and / or the like, and forward the input information to the processor(s) 142 for processing. In some embodiments, the computer system 140 may be a server machine in a cloud computing environment. In such embodiments, the computer system 140 may not include the input devices 308, but may provide equivalent input information by receiving commands (e.g.,in response to one or more inputs from a remote computing device) in the form of messages transmitted over a network and received via a network adapter 318. In some embodiments, switch 316 is configured to provide connections between I / O bridge 307 and other components of computer system 140, such as a network adapter 318 and various add-in cards 320 and 321.

[0037] In some embodiments, the I / O bridge 307 is coupled to a system disk 314, which may be configured to store content, applications, and data for use by the processor(s) 312 and the parallel processing subsystem 312. In some embodiments, the system disk 314 provides non-volatile storage for applications and data and may include fixed or removable hard disk drives, flash memory devices, and CD-ROM (Compact Disc Read-Only Memory), DVD-ROM (Digital Versatile Disc-ROM), Blu-ray, HD-DVD (High-Definition DVD), or other magnetic, optical, or solid-state storage devices. In some embodiments, other components, such as Universal Serial Bus or other connectors, CD drives, DVD drives, movie recorders, and the like, may also be connected to the I / O bridge 307.

[0038] In some embodiments, memory bridge 305 may be a northbridge chip and I / O bridge 307 may be a southbridge chip. Additionally, communication paths 306 and 313, as well as other communication paths within computer system 140, may be implemented using any technically suitable protocols, including, but not limited to, AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol known in the art.

[0039] In some embodiments, parallel processing subsystem 312 includes a graphics subsystem that provides pixels to an optional display 310, which may be a conventional cathode ray tube, a liquid crystal display, a light-emitting diode display, and / or the like. In such embodiments, parallel processing subsystem 312 may include circuitry optimized for graphics and video processing, including, for example, video output circuitry. Such circuitry may be integrated into one or more parallel processing units (PPUs), also referred to as parallel processors, included in parallel processing subsystem 312.

[0040] In some embodiments, the parallel processing subsystem 312 includes circuitry optimized for general purpose and / or computational processing (e.g., that has undergone optimization). Again, such circuitry may be integrated via one or more PPUs included within the parallel processing subsystem 312 and configured to perform such general purpose and / or computational operations. In other embodiments, the one or more PPUs included within the parallel processing subsystem 312 may be configured to perform graphics processing, general purpose processing, and / or computational processing operations. The system memory 144 includes at least one device driver configured to manage the processing operations of the one or more PPUs within the parallel processing subsystem 312.In addition, the system memory 144 contains the control application 146 which, in conjunction with the . Fig. 7-8 and 10. Although described herein primarily with respect to the control application 146, the methods disclosed herein may also be implemented, either in whole or in part, in other software and / or hardware, such as the parallel processing subsystem 312.

[0041] In some embodiments, the parallel processing subsystem 312 may be integrated with one or more of the other elements of Fig. 3 to form a single system. For example, the parallel processing subsystem 312 may be integrated with the processor(s) 142 and other interconnect circuitry on a single chip to form a system on a chip (SoC).

[0042] In some embodiments, processor(s) 142 comprise(s) the primary processor of computing system 140, which controls and coordinates the operations of other system components. In some embodiments, processor(s) 142 issue(s) instructions that control the operation of PPUs. In some embodiments, communication path 313 is a PCI Express interconnect with dedicated connections or lanes assigned to each PPU. Other communication paths may also be used. The PPU advantageously implements a highly parallel processing architecture, and the PPU may be equipped with any amount of local parallel processing (PP) memory.

[0043] It should be noted that the system shown here is for illustrative purposes only, and variations and modifications are possible. The interconnect topology, including the number and arrangement of bridges, the number of processors 312, and the number of parallel processing subsystems 312, can be changed as desired. For example, in some embodiments, system memory 144 may be connected directly to processor(s) 142 rather than through memory bridge 305, and other devices may communicate with system memory 144 via memory bridge 305 and processor(s) 142. In other embodiments, parallel processing subsystem 312 may be connected to I / O bridge 307 or directly to processor(s) 142 rather than to memory bridge 305.In still other embodiments, the I / O bridge 307 and the memory bridge 305 may be integrated into a single chip rather than existing as one or more discrete devices. In certain embodiments, one or more of the devices shown in FIG. Fig. 3 are missing. For example, the switch 316 may be omitted and the network adapter 318 and the add-in cards 320, 321 would be connected directly to the I / O bridge 307. Finally, in certain embodiments, one or more of the components shown in Fig. 3 may be implemented as virtualized resources in a virtual computing environment, such as a cloud computing environment. In particular, in some embodiments, the parallel processing subsystem 312 may be implemented as a virtualized parallel processing subsystem. For example, the parallel processing subsystem 312 may be implemented as virtual graphics processing unit(s) (vGPU(s)) that renders graphics on a virtual machine(s) (VM(s)) running on a server machine(s) whose GPU(s) and other physical resources are shared by one or more VMs. Physics-based animation of joints with partial prerequisites

[0044] Fig. 4 is a more detailed representation of the model trainer 116 from Fig. 1 according to various embodiments. As shown, the model trainer 116 includes an initialization module 402, the motion tracking model 151, and a reinforcement learning module 404.

[0045] During training, a character 160 (contained in the environment 170) interacts with the environment 170 according to actions 305 generated using the motion tracking model 151 being trained. The motion tracking model 151 is a partially constrained physical controller that receives as input a current state 301 of the character 160 and a sparse number of body parts, as well as future time steps. With these inputs, the motion tracking model 151 generates the actions 305, which in some embodiments may be motor actuations, such that the character 160 tracks the requested body parts in the requested time periods. For example, the actions 305 may be one-step primitive motor controls.

[0046] The initialization module 402 generates the sparse set of body parts and future time steps by sampling motion recordings 154 and sampling a random mask for future poses. The motion recordings 154 can be a motion capture dataset that contains captured motions of people performing various motions. In some embodiments, the motion recordings 154 can contain the positions and rotations for each joint in each frame of the captured motions. The initialization module 402 samples a motion from the motion recordings 154 and a time step within the motion to start with. The initialization module 402 also samples a random mask that randomly selects joints from randomly selected frames of the sampled motion to occlude, in order to have few orto generate sparse future poses missing the selected joints and / or images, which may not involve masking in some cases where no joints or images are randomly selected to be hidden.

[0047] After simulating character 160 while performing an action generated by motion tracking model 151, taking into account state 301 and the sparse set of body parts and future time steps output by initialization module 402, reinforcement learning module 404 learns to recreate the originally recorded motion based on a comparison of a next state 302 of character 170 from the simulation (which may also be input to motion tracking model 151 at a later time) with a ground truth state of character 170 in a corresponding frame of motion sampled by initialization module 402. Specifically, based on the comparison, reinforcement learning module 404 calculates a reward for tracking the sparse set of body parts for the current frame.The reinforcement learning module 404 then updates the parameters of the motion tracking model 151 using the reward and a backpropagation method. In some embodiments, a proximal policy optimization (PPO) technique may be employed during training.

[0048] More formally, a reinforcement learning agent can interact with an environment (e.g., environment 170) according to a policy π to train the motion tracking model 151, which is a controller with few constraints. At each step t, the agent observes a state s t and samples or determines an action a t from Directive a t ~ π(a t |s t ). The environment then moves according to the environmental dynamics p(s t+1 |s t , a t ) to the next state. The agent's goal is to learn a policy that maximizes the discounted cumulative reward: J=Ep(τ|π)[∑t=0Tγtrt|s0=s], where p(τ|π)=p(s0)∏t=0T−1p(st+1|st,at)π(at|st) the probability of a trajectory τ=(s 0 ,a 0, r 0 ,...,s T-1 ,a T-1 ,r T-1 ,s T ) and γ ∈ [0,1) is a discount factor that determines the effective horizon of the policy.

[0049] In physics-based motion tracking, the goal is to generate controls (e.g., motor actuations) that allow a simulated character to perform a sequence of simulated poses. t = (p t ,θ t ) corresponding to a kinematic target movement q̂ = (p̂,θ ̂ ) are very similar. A movement can be described as a sequence of poses over time q t:t+K where each kinematic pose q t = (p t , θ t ) by the Cartesian 3D positions of the joints pt=(pt0,pt1,…,ptJ) of a character J and their local rotations θt=(θt0,θt1,…,θtJ) is shown.

[0050] The goal of training the motion tracking model 151 is to go beyond the usual case of tracking whole-body reference movements and to track sparse or few reference movements, which here are q^tsparse is referred to. A sparse pose represents an incomplete representation of joint positions and rotations, where only the features of some joints can be observed. The set of observed joints can also vary from image to image within a reference motion. For example, some images may contain a complete description of the pose, while joints in other images may not be observed at all. The initialization module 402 can obtain a sparse motion by capturing a full-body motion and blanking out different elements in space and time. Subsequently, the motion-tracking model 151 can be trained using reinforcement learning based on the sparse motion.

[0051] Fig. 5 illustrates how the motion-tracking model 151 is of Fig. 1 according to various embodiments. As shown, in some embodiments, the motion tracking model 151 may be a variational autoencoder (VAE) comprising an encoder 508, a prior module 510, and a controller 516 that is a decoder. Although described primarily in relation to a VAE, in some embodiments, any technically feasible machine learning model, including other generative models such as diffusion models, may be used as the motion tracking model 151. For example, in some embodiments that do not use a VAE, the motion tracking model 151 may directly map a character's current state and masked targets to an action. In such cases, the motion tracking model 151 may have any suitable architecture, e.g.,a single model that is trained and then reused at inference time, rather than multiple models such as an encoder, a prior module, and a controller as in the case of the VAE. To train the motion tracking model 151, the model trainer 116 first samples a motion from the motion recordings 154 and a time step within the motion, which is represented as sampled motion 502. In some embodiments, the sampling may prioritize training on the most difficult-to-imitate motions by using a motion sampling rate per segment that is proportional to the number of times the motion tracking model 151 failed to track an image belonging to that segment relative to the total number of images present in that segment and smoothed over time using standard discounted accumulation.

[0052] The model trainer 116 also samples a mask for the sampled motion 502. In some embodiments, the mask may randomly select joints from randomly selected images of the sampled motion 502 (which, as described above in connection with Fig. 4, in some cases may not include masking) to generate sparse future poses 504, missing the selected joints and / or images. The model trainer 116 inputs the sparse future poses 504 and a current state 522 of the character to a prior module 510, which is trained to learn which part of a latent space to use given the inputs, represented as a latent prior distribution 512. The latent prior distribution 512 is a (prior) distribution in the latent space representing the solutions, so the prior distribution can be viewed as a distribution over possible solutions.The model trainer 116 also inputs the sparse future poses 504, the complete future reference poses 506 from the sampled motion 502, and the current state 522 into the encoder 508, which observes the complete reference motion that must be generated and is also used to learn the latent prior distribution 512 to be used.

[0053] To control the character in the environment 170, the model trainer 116 samples the latent distribution 512 using the outputs of the prior module 510 and the encoder 508 to obtain a sampled latent vector 514. The model trainer 116 then inputs the sampled latent vector 514 and the character's current state 522 to the controller 516, which outputs an action 518. The model trainer 116 then simulates the character performing the action and receives character movement from the environment 170. In some embodiments, the model trainer 116 transmits the action 518 to a controller of the character, such as a proportional-derivative controller (PD controller), which controls the character's joints to move in the environment 170 according to the action 518.

[0054] The simulation 520 of the character performing the action 518 results in an updated state 524 of the character, which the model trainer 116 uses to calculate a reward for updating parameters of the motion tracking model 151. In some embodiments, the model trainer 116 calculates a reward based on a comparison of the updated state 524 after the simulation 520 with the state in a corresponding frame of the sampled motion 502.In some embodiments, the reward may include: (1) a term that rewards smaller distances between obtained joint positions and joint positions in the sampled motion, (2) a term that rewards smaller angles between obtained joint rotations and joint rotations in the sampled motion, (3) a term that rewards smaller differences between an obtained root joint height and a root joint height in the sampled motion, (4) a term that rewards smaller differences between obtained joint velocities and joint velocities in the sampled motion, and (5) a term that penalizes the energy consumption in the obtained motion. More generally expressed, in some embodiments the reward may be any comparison metric (i.e.be the similarity metric) between the generated motion and the kinematic target motion from the corresponding image of the sampled motion 502. The model trainer 116 can update the parameters of the encoder 508, the prior module 510, and the controller 516 in the motion tracking model 151 using the reward and a backpropagation method. In some embodiments, a PPO method can be used during training.

[0055] In some embodiments, the model trainer 116 may also determine whether to prematurely terminate training using the sampled motion 502. In such cases, the model trainer 116 may determine premature termination if the obtained motion of the character in the updated state 524 deviates too much in position and / or orientation from the sampled motion 502 used as a reference. More generally, in some embodiments, the model trainer 116 may determine premature termination based on an error defined as a mismatch in a technically feasible similarity metric.

[0056] In some embodiments, the prior module 510 more formally takes as input a set of K future sparse constraints and returns the distribution 512 over the latent space p(z|st,qt+1:t+Ksparse) from which the motion tracking model 151 can later sample. The encoder 508, used only during training, also observes unmasked whole-body poses. The task of the encoder 508 is to overcome the ambiguity between the acceptable solutions by providing a residual for the prior distribution 512. The controller 516, also referred to as the decoder, observes the current state s t and the sampled latent or the sampled vector z t and generates an action distribution π(a|s t , e.g. t ).

[0057] In some embodiments, the motion tracking model 151 may be trained using reinforcement learning with a motion tracking objective whose goal is to produce actions that reproduce a desired reference motion when only sparse observations are available for future frames. Sparsity can occur both temporally and spatially (at the joints). Compared to mimicking whole-body motions, sparsity introduces a key issue of ambiguity because the problem becomes underspecified. For example, if the reference motion only specifies the position of the pelvis, there are multiple possible whole-body motions (solutions) that satisfy the specified constraints (e.g., the hands may remain static next to the body or swing at the side of the character).To address the above-mentioned challenges, the motion-tracking model 151 can be implemented as a VAE with a learned prior, a framework that is capable of modeling these types of multimodal solutions. As described, the motion-tracking model 151 has three learned models: the prior module 510. p(ztp|st,qt+1:t+Ksparse) the encoder 508 q(ztq|st,qt+1:t+Ksparse,qt+1:t+K,) that shifts the latent distribution during training, and a control 516 (decoder policy) π(a t |s t , z t). The learnable prior module 510 allows the system to distinguish between multiple valid solutions for different constraints, so that the action distribution for a head-only constraint can capture a wider variety of movements than for a VR (head and hands) constraint. In some embodiments, both the encoder 508 and a critic (not shown), which sees the full pose and provides a prediction of the cumulative discounted reward that the controller 516 is expected to receive based on the movement to which the controller 516 is conditioned, are modeled as fully connected networks. In such cases, both the encoder 508 and the critic observe the current pose s t , the complete unmasked future poses ŝ t+1:t+Kand the binary mask for future poses. In some embodiments, the controller 516 is also a fully connected network that stores the current state of t and the sampled latent z t observed to a t .In some embodiments, the prior module 510 may be used as described below in conjunction with Fig. 6 be described.

[0058] In some embodiments, the latent distribution may be determined during training using a residual encoder zt~N(μtp+μtq,σtp) with a linearly increasing KL divergence coefficient. Then, during inference, the latent is calculated directly from the prior or the prior distribution zt~N(μtp,σtp). sampled. To train the motion tracking model 151, the model trainer 116 may optimize a proxy reconstruction loss through direct reinforcement learning optimization. The reward during training may be formulated as a whole-body facial expression and viewed through the lens of the goal-conditioned reinforcement learning framework. In some embodiments, the policy's action distribution is calculated using a multidimensional Gaussian distribution with a fixed diagonal covariance matrix σ π = exp (-2.9) and the reward r t according to: rt=r(q^t,qt)=wgtrtgr+wgrrtgr+wrhrtrh+wjavrtjav+wenrgrtenrg, where the individual components are defined as: global translation reward of the root height rtrh=e−crh‖p^troot−hright−ptroot−height‖ and joint angular velocity rtjav=e−cjvv‖v^t−vt‖. To mitigate jitter and promote more stable behaviors, an energy reduction reward rtenrg=−∑j|τjωj|2 contain, where τ j and ω j correspond to the torque and angular velocity of joint j. w and c can be manually specified coefficients for combining the different reward conditions. In some embodiments, at each step during training, a full-body motion tracker observes the character's current state, which consists of the current 3D body pose and velocity, canonicalized to the character's local image: st=(θt⊖θtroot,pt−ptroot,wt⊖θtroot), where ⊖ denotes the difference between two quaternions. The target poses are represented, for example, by K = 10 future poses. Where each joint q^t+kj relative to the current pose q^j=(θ^j⊖tj,θ^j⊖θtroot,p^j−ptj,p^j−ptroot). is canonized.

[0059] Furthermore, a pose with missing information is referred to here as q̂ sparse A pose with missing information can be the result of masking ground-truth motion capture data or, for example, sparse input from an animator.

[0060] To support "any-joint any-time," or any joint at any time within the provided context length at the time of inference, the motion tracking model 151 can be represented with any sparseness pattern based on an animator's requirements. To handle such scenarios, the motion tracking model 151 can be trained with different sparseness patterns by randomly sampling joint masks and time gaps, sequences of images in which all joints are hidden. This results in a mask ∈ ℝ K·(J·2), K future steps with J joints supporting both position and rotation constraints.

[0061] As described, in some embodiments, the model trainer 116 may also implement early termination when the received motion of the character in the updated state 524 deviates too far in position and / or orientation (or based on another similarity metric) from the sampled motion 502 used as a reference, and adaptive state initialization that complements early termination by prioritizing training on motions that are most difficult to mimic. Early termination has two goals: (1) dynamic motions are typically more difficult to track, and early termination ensures that the training procedure focuses on motions that are "close" to the target motion; and (2) because rewards are non-negative, early termination of the episode serves as a strong gain-optimizing signal for the agent to stay within the tracking boundaries.In particular, in some embodiments, an episode may be terminated in any state if one of the reward components. rtgt,rtgr,rtrh,rtjav,rtenrg falls below a certain threshold. Experience has shown that a threshold of 0.2 works well. In adaptive state initialization, in some embodiments, each recorded motion may be divided into segments (e.g., 0.5-second segments). During training, the initialization module 402 of the model trainer 116 may maintain a sampling rate per segment proportional to the number of times the policy failed to track an image belonging to that segment. The sampling rate per segment may also be smoothed over time using standard discounted accumulation, where the sampling rate of a segment i is updated at each epoch e according to the following formula: wei=num failuresitotal framesi+0.7⋅we−1i.

[0062] When a new episode begins, the target movement and the initial time within that movement can be sampled proportionally to the weights of equation (5).

[0063] After training, the prior module 510 and the controller 516 can be used during the inference phase to generate a physically animated full-body movement of the character based on sparse constraints provided by a user. More specifically, the user can provide the initial pose for the character 160 and a series of constraints (joint positions and rotations) for subsequent time steps. For example, the user can draw a curve that a specific joint (e.g., a pelvic joint) should follow, or different curves for different joints (e.g., a head joint and wrists). At each time step, a latent z tfrom the prior distribution 512 zt∼N(μtp,σtp) based on the character's current state and sparse future constraints. The current state, along with the sampled latent, is then communicated to controller 516, which generates realistic full-body movements that meet user specifications. In general, user-provided sparse motion constraints can specify the positions and / or orientations of any number of a character's joints in any number of frames of an animation, including fully observed motion where no masking is applied. That is, controller 516 supports "any-joint-any-time" tracking, from all joints in all frames to no visible joints at all.

[0064] Fig. 6 is a more detailed representation of the prior module 510 of Fig. 5 according to various embodiments. As shown, the prior module 510 takes as input a current pose 620 and a number K of future target joint positions 602. A representation 604 of the future joint target positions 602 per joint and a representation 622 of the current pose per joint represent each pose by individual joints of the pose. A joint embedding JE j, which is a position joint embedding, is appended to each future constraint of joint j. For example, the joint embedding 608 is appended to the future joint constraint 606 to produce a combined representation 610. The prior module 510 encodes 609 the combined representations using an encoder (not shown) shared by all joint representations to produce a position encoding 612, which is masked at the joints and time to produce a masked joint position encoding 616. The prior module 510 encodes 611 the representation 622 of the current whole-body pose per joint using a separate encoder (not shown). Time-domain capture can be achieved by applying a standard position encoding. The prior module 510 also appends a type embedding TE to the masked joint position encoding 616 typefor each input type 618. The prior module 510 then feeds the masked joint position encoding 616 with the appended type embeddings 618 into a transformer encoder 170, which outputs the prior distribution 512. The transformer encoder 170 learns to care about the current pose and each future joint in each time frame up to K future time frames based on what is important.

[0065] More formally, in some embodiments, the input to the prior module 510 having a transformer architecture includes the current state s t , the K future poses [q̂ t+1 , ..., q t+K ] and a text encoding for the current movement text t The inputs are preprocessed as follows. For future poses, each joint q^Tj in any future pose with a joint embedding JE jAn encoder, shared by all future joint constraints, encodes the combined representation (e^τj), followed by a position coding over the time domain (e˜rj). This leads to a future pose-encoding tensor [ẽ t+1 , ..., ẽ t+K ] ∈ ℝ K·J·2×dim , J joints, K future poses and two possible types of constraints (position, rotation). For a current pose, the pose q t to e t encoded 611, using an encoder different from the encoder used for the future joint constraints. The resulting representation (K + 1, dim) is then fed to the transformer encoder 170, followed by two output heads to generate the prior distribution and prior 512, respectively.

[0066] Fig. 7 is a more detailed representation of the control application 146 of Fig. 1 according to various embodiments. As shown, the control application 146 includes the motion tracking model 152. In operation, the control application 146 receives sparse motion constraints 702. The sparse motion constraints 702 may specify the positions and / or orientations of any number of a character's joints in any number of frames of an animation, including fully observed motion where no masking is applied. The sparse motion constraints 702 may be specified by a user in any technically feasible manner, for example, in some embodiments, via a graphical user interface (GUI). The control application 146 inputs the sparse motion constraints 702 and a current state of the character to the motion tracking model 152 to generate an action 704.The control application 146 then controls a character within the environment 170 using the generated action. For example, in some embodiments, the control application 146 may transmit the action 704 to a controller of the character, such as a PD controller, which controls the character's joints to move according to the action within the environment 170. The control application 146 receives an updated state 706 of the character from the environment 170, which, along with the sparse motion constraints 702, may be used to generate further actions to control the character.

[0067] Fig. 8 is a more detailed illustration of how the motion tracking model 152 of Fig. 7 according to various embodiments to control a character. As shown, upon input of sparse motion constraints 802 (which are the same as those described above in connection with Fig. 7), the prior module 510 generates the prior distribution 512 based on the sparse motion constraints and a current state of a character, and the control application 146 then samples the prior distribution 512 to obtain a latent vector 804. The control application 146 then inputs the sampled latent vector 804 and a current state 808 of the character to the controller 516, which outputs an action 806. The control application 146 controls a character in the environment 170 using the action 806. As described, in some embodiments, the control application 146 may transmit the action 806 to a controller of the character, for example, a PD controller, which controls the character's joints to move according to the action in the environment 170.Thereafter, the control application 146 receives an updated status of the character from the environment 107, and the above process can be repeated to generate another action to control the character in a subsequent time step, and so on. Note that the motion tracking model 152 does not include the encoder 508 of the motion tracking model 151, since the encoder 508 can be discarded after training the motion tracking model 151.

[0068] Fig. 9 shows a flowchart of the method steps for training a motion tracking model according to various embodiments. Although the method steps in connection with the Fig. 1 to 8, those skilled in the art will understand that any system configured to perform the method steps in any order falls within the scope of the present invention.

[0069] As illustrated, a method 900 begins at step 902, where the model trainer 116 samples a motion from the motion recordings 154 and a time step within the motion. In some embodiments, the sampling may prioritize training on the most difficult-to-imitate motions by using a motion sampling rate per segment that is proportional to the number of times the motion tracking model 151 failed to track an image belonging to that segment relative to the total number of images output in that segment and smoothed over time using standard discounted accumulation, as described above in connection with Fig. 5 is described.

[0070] At step 904, the model trainer 116 samples a mask for the sampled motion. In some embodiments, the mask may filter randomly selected joints from randomly selected images of the sampled motion, which may not involve masking if no joints or images are randomly selected.

[0071] In step 906, the model trainer 116 calculates an action for an image using the motion tracking model 151 for the sampled motion masked by the sampled mask. In some embodiments, the prior module 510 of the model trainer 116 generates the latent distribution 512 considering the sparse motion constraints and a character's current state as inputs, and the model trainer 116 samples the latent distribution 512 to obtain a latent vector. Then, the model trainer 116 inputs the latent vector and the character's current state to the controller 516, which outputs the action.

[0072] At step 908, the model trainer 116 simulates the character performing the action and obtains a character movement from the environment 170. In some embodiments, the model trainer 116 transmits the action calculated at step 906 to a controller of the character, such as a PD controller, which controls the character's joints to move in the environment 170 according to the action.

[0073] In step 910, the model trainer 116 determines whether to prematurely terminate training using the sampled motion and the sampled mask. In some embodiments, the model trainer 116 determines to prematurely terminate training if the obtained motion of the character has deviated too much in position and / or orientation from the sampled motion used as a reference, as described above in connection with Fig. 5. More generally, in some embodiments, the model trainer 116 may determine early termination based on an error defined as a mismatch in a technically feasible similarity metric.

[0074] If the model trainer 116 determines not to terminate early, the model trainer 116 calculates a reward based on a comparison of the obtained motion with the sampled motion in step 912. In some embodiments, the reward may include: (1) a term that rewards smaller distances between obtained joint positions and joint positions in the sampled motion, (2) a term that rewards smaller angles between obtained joint rotations and joint rotations in the sampled motion, (3) a term that rewards smaller differences between a obtained root joint height and a root joint height in the sampled motion, (4) a term that rewards smaller differences between obtained joint velocities and joint velocities in the sampled motion, and (5) a term that penalizes energy consumption in the obtained motion.In some embodiments, the reward of equation (2) may be used. More generally, in some embodiments, the reward may be any comparison metric between the generated motion and the received motion used as the target kinematic motion.

[0075] In step 914, the model trainer 116 updates the parameters of the motion tracking model 151 based on the reward calculated in step 912. In some embodiments, the model trainer 116 may update the parameters of the encoder 508, the prior module 510, and the controller 516 in the motion tracking model 151 using the reward and a backpropagation method. In some embodiments, a PPO technique may be employed. The method 900 then returns to step 906, where the model trainer 116 calculates an action for another frame using the motion tracking model 151 for the sampled motion masked by the sampled mask.

[0076] On the other hand, if the model trainer 116 determines early termination in step 910, the model trainer 116 determines whether to continue training in step 916. For example, training may be terminated after a certain number of training iterations or if the reward does not improve over a number of training iterations. If the model trainer 116 decides to continue training, the method 900 returns to step 902, where the model trainer 116 samples another motion from the motion recordings 154 and a time step within the motion. If, on the other hand, the model trainer 116 decides to terminate training, the method 900 ends.

[0077] Fig. 10 is a flowchart of the method steps for generating an animation of a character having sparse motion constraints, according to various embodiments. Although the method steps in connection with the Fig.1 to 8, those skilled in the art will understand that any system configured to perform the method steps in any order falls within the scope of the present invention.

[0078] As illustrated, a method 1000 begins at step 1002, where the control application 146 receives sparse motion constraints. The sparse motion constraints may be specified by a user in any technically feasible manner, such as via a GUI. The sparse motion constraints may include positions and / or orientations of any number of a character's joints in any number of frames of an animation, including fully observed motion with no masking applied, thereby supporting "any-joint-any-time" tracking, ranging from all joints in all frames to no visible joints at all.

[0079] At step 1004, the control application 146 samples the prior distribution 512 based on the sparse motion constraints and a character's state to obtain a latent vector. In some embodiments, the control application 146 inputs the sparse motion constraints and the character's state to the prior module 510, which generates the prior distribution 512, and the control application 146 then samples the prior distribution 512 to obtain the latent vector.

[0080] At step 1006, the control application 146 generates an action based on the latent vector and the character's state. In some embodiments, the control application 146 inputs the latent vector and a current state of the character to the controller 516, which outputs the action.

[0081] At step 1008, the control application 146 controls the character in the environment 170 using the generated action. In some embodiments, the control application 146 transmits the action to a controller of the character, such as a PD controller, which controls the character's joints to move in the environment 170 according to the action.

[0082] In step 1010, the control application 146 obtains a state of the character from the environment 107. The state may include updated joint positions of the character after the action is performed.

[0083] If the control application 146 determines in step 1012 that control of the character is to continue, the method 1000 returns to step 1004, where the control application 146 again samples the prior distribution 512 based on the sparse motion constraints received in step 1002 and the state of the character received in step 1010 to obtain another latent vector.

[0084] In summary, methods for animating characters using sparse motion constraints are disclosed. In some embodiments, the sparse motion constraints may specify the positions and / or orientations of any number of joints of a character in any number of frames of an animation. Based on sparse motion constraints, a control application samples a latent prior distribution generated via a trained motion tracking model considering the sparse motion constraints and a current state of a character to obtain a sampled latent vector. The character may be a virtual character in a computer-based environment or a physical robot in a real-world environment, and the character's current state may be obtained from the computer-based environment or detected using sensors in the real-world environment.The control application inputs the sampled latent vector and the current state of the character to a controller of the motion tracking model to generate an action. The control application can then control the character within the computer-based environment or the real-world environment using the action. Controlling the character can result in an updated state of the character, and the above process can be repeated to generate another action for controlling the character using the updated state of the character, the sparse motion constraints, and the motion tracking model.

[0085] A model trainer trains the motion tracking model, which in some embodiments may be a variational autoencoder (VAE). In some embodiments, the model trainer first samples a motion and a time step within the motion from a set of motion recordings. The model trainer then samples a mask, which is used to mask out random joints and / or frames from the sampled motion. The model trainer computes an action for each frame using the motion tracking model, the masked motion, and a current state of the character. The model trainer simulates the character performing the action in the environment and obtains a motion of the character from the environment.Then, the model trainer calculates a reward based on a comparison of the obtained motion with the sampled motion, and the model trainer updates the parameters of the motion tracking model based on the reward.

[0086] A technical advantage of the disclosed methods over the prior art is that the disclosed methods can animate a physical or virtual character by specifying the positions and orientations of a subset of a character's joints, rather than all of them, in any number of frames of an animation. Furthermore, the disclosed methods can produce animations that are more physically realistic than animations produced using kinematic models that do not account for the forces moving a character's joints. These technical advantages represent one or more technological improvements over the prior art. 1. In some embodiments, a computer-implemented method for animating characters includes obtaining a first state of a character and one or more constraints for one or more movements associated with a subset of joints associated with the character, generating a first action for the character to perform via a trained machine learning model and based on the first state and the one or more constraints, and causing the character to perform the first action within a computer-based or physical environment. 2. The computer-implemented method of sentence 1, wherein generating the first action comprises sampling a prior distribution based on the first state and the one or more constraints to generate a latent vector, and processing the latent vector and the first state using a controller included in the machine learning model to generate the first action. 3. The computer-implemented method of sentence 1 or 2, further comprising processing the first state and the one or more constraints using a transformer encoder to generate the prior distribution. 4. The computer-implemented method of any one of sentences 1-3, further comprising training a first machine learning model to generate the trained machine learning model, wherein the first machine learning model comprises an encoder. 5. The computer-implemented method of any one of sentences 1-4, wherein the trained machine learning model comprises at least one trained variational autoencoder, VAE, or a trained generative model. 6. The computer-implemented method of any one of sentences 1-5, further comprising generating a second action for the character to perform via the trained machine learning model and based on the one or more constraints and a second state of the character after performing the first action, and causing the character to perform the second action within the computer-based or physical environment. 7. The computer-implemented method according to any one of sentences 1-6, further comprising training a first machine learning model to generate the trained machine learning model by sampling a first motion from a set of motion recordings and a time step within the first motion to generate a sampled motion, removing at least one joint or at least one image within the sampled motion to generate a masked motion, generating, via the first machine learning model and based on a second state of the character and the masked motion, a second action for the character to perform, causing the character to perform the second action within the computer-based environment to reach a third state of the character,and updating one or more parameters of the first machine learning model based on a comparison between the third state and a fourth state of the character included in the sampled motion. 8. The computer-implemented method of any one of sentences 1-7, further comprising training a first machine learning model to generate the trained machine learning model based on a reward that is a metric of a comparison between movements generated by the first machine learning model and movements sampled from a set of movement recordings. 9. The computer-implemented method of any one of clauses 1-8, wherein a controller controlling one or more joints of the character causes the character to move according to the first action. 10. The computer-implemented method of any one of clauses 1-9, wherein the character comprises either a virtual character or a physical robot. 11. In some embodiments, one or more non-transitory computer-readable media store instructions that, when executed by at least one processor, cause the at least one processor to perform the steps of obtaining a first state of a character and one or more constraints for one or more movements associated with a subset of joints associated with the character, generating a first action for the character to perform via a trained machine learning model and based on the first state and the one or more constraints, and causing the character to perform the first action within a computer-based or physical environment. 12. One or more non-transitory computer-readable media according to sentence 11, wherein generating the first action comprises sampling a prior distribution based on the first state and the one or more constraints to generate a latent vector, and processing the latent vector and the first state using a controller included in the trained machine learning model to generate the first action. 13. One or more non-transitory computer-readable media according to sentence 11 or 12, wherein the instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of processing the first state and the one or more constraints using a transformer encoder to generate the prior distribution. 14. One or more non-transitory computer-readable media according to any one of clauses 11-13, wherein the instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of training a first machine learning model to generate the trained machine learning model, wherein the first machine learning model includes an encoder. 15. One or more non-transitory computer-readable media according to any one of sentences 11-14, wherein the instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of training a first machine learning model to generate the trained machine learning model by sampling a first motion from a set of motion recordings and a time step within the first motion to generate a sampled motion, removing at least one joint or at least one image within the sampled motion to generate a masked motion, generating, via the first machine learning model and based on a second state of the character and the masked motion, a second action for the character to perform, causing the character to perform the second action within the computer-based environment,to achieve a third state of the character, and updating one or more parameters of the first machine learning model based on a comparison between the third state and a fourth state of the character contained in the sampled motion. 16. One or more non-transitory computer-readable media according to any one of sentences 11-15, wherein the instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of terminating training the first machine learning model using the sampled motion based on a similarity between the third state and the fourth state being less than a predefined threshold. 17. One or more non-transitory computer-readable media according to any one of sentences 11-16, wherein the instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of training a first machine learning model to generate the trained machine learning model based on a reward that is a metric of a comparison between movements generated by the first machine learning model and movements sampled from a set of motion recordings. 18. One or more non-transitory computer-readable media according to any one of clauses 11-17, wherein a controller controlling one or more joints of the character causes the character to move according to the first action. 19. One or more non-transitory computer-readable media according to any one of sentences 11-18, wherein the environment is at least one of a simulation environment, an augmented reality environment, XR, a game environment, and a physical environment. 20. In some embodiments, a system includes one or more memories storing instructions and one or more processors coupled to the one or more memories and configured, upon execution of the instructions, to obtain a first state of a character and one or more constraints on one or more movements associated with a subset of joints associated with the character, to generate, via a trained machine learning model and based on the first state and the one or more constraints, a first action for the character to perform, and to cause the character to perform the first action within a computer-based or physical environment.

[0087] All combinations of claim elements recited in the claims and / or elements described in this application fall in any way within the intended scope of the present disclosure and protection.

[0088] The descriptions of the various embodiments are provided for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments.

[0089] Aspects of the present embodiments may be embodied as a system, method, or computer program product. Accordingly, aspects of the present disclosure may take the form of a pure hardware embodiment, a pure software embodiment (including firmware, resident software, microcode, etc.), or an embodiment that combines software and hardware aspects, which may be generally referred to herein as a "module" or "system." Furthermore, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code embodied thereon.

[0090] Any combination of one or more computer-readable media may be used. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing.More specific examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer diskette, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium can be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0091] Aspects of the present disclosure are described above with reference to flowchart illustrations and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, may be implemented by computer program instructions. These computer program instructions may be supplied to a processor of a general-purpose computer, a special-purpose computer, or other programmable computing device to produce a machine. The instructions, when executed by the processor of the computer or other programmable computing device, enable the implementation of the functions / acts specified in the flowchart and / or block diagram.Such processors may be, without limitation, general-purpose processors, special-purpose processors, application-specific processors, or field-programmable gate arrays.

[0092] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of code comprising one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions specified in the blocks may occur out of the order shown in the figures. For example, two blocks shown in sequence may actually execute substantially concurrently, or the blocks may sometimes execute in reverse order, depending on the functionality involved.It is also understood that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, may be implemented by special purpose hardware-based systems that perform the specified functions or acts, or by combinations of special purpose hardware and computer instructions.

[0093] While the foregoing refers to embodiments of the present disclosure, other and further embodiments of the disclosure may be devised without departing from the basic scope of the disclosure, and the scope of the disclosure is determined by the following claims.

Claims

[1] A computer-implemented method for animating characters, the method comprising: Maintaining a first state of a character and one or more constraints on one or more movements associated with a subset of joints belonging to the character; Generating a first action for the character to perform via a trained machine learning model and based on the first state and the one or more constraints; and Causing the character to perform the first action within a computer-based or physical environment. [2] The computer-implemented method of claim 1, wherein generating the first action comprises: sampling a prior distribution based on the first state and the one or more constraints to generate a latent vector, and Processing the latent vector and the first state using a controller included in the machine learning model to generate the first action. [3] The computer-implemented method of claim 1 or 2, further comprising processing the first state and the one or more constraints using a transformer encoder to generate the prior distribution. [4] The computer-implemented method of any preceding claim, further comprising training a first machine learning model to generate the trained machine learning model, wherein the first machine learning model comprises an encoder. [5] A computer-implemented method according to any one of the preceding claims, wherein the trained machine learning model comprises at least one trained variational autoencoder, VAE, or a trained generative model. [6] A computer-implemented method according to any one of the preceding claims, further comprising: Generating a second action for the character to perform via the trained machine learning model and based on one or the multiple restrictions and a second state of the character after performing the first action; and Causing the character to perform the second action within the computer-based or physical environment. [7] A computer-implemented method according to any one of sentences 1-6, further comprising training a first machine learning model to generate the trained machine learning model by: sampling a first motion from a set of motion recordings and a time step within the first motion to generate a sampled motion; removing at least one joint or at least one image within the sampled motion to create a masked motion; Generating, via the first machine learning model and based on a second state of the character and the masked movement, a second action for the character to perform; Causing the character to perform the second action within the computer-based environment to achieve a third state of the character; and Updating one or more parameters of the first machine learning model based on a comparison between the third state and a fourth state of the character included in the sampled motion. [8] A computer-implemented method according to any preceding claim, further comprising training a first machine learning model to generate the trained machine learning model based on a reward that is a metric of a comparison between movements generated by the first machine learning model and movements sampled from a set of movement recordings. [9] A computer-implemented method according to any one of the preceding claims, wherein a controller controlling one or more joints of the character causes the character to move according to the first action. [10] A computer-implemented method according to any one of the preceding claims, wherein the character comprises either a virtual character or a physical robot. [11] One or more non-transitory computer-readable media storing instructions that, when executed by at least one processor, cause the at least one processor to perform the steps of: Maintaining a first state of a character and one or more constraints on one or more movements associated with a subset of joints belonging to the character; Generating a first action for the character to perform via a trained machine learning model and based on the first state and the one or more constraints; and Causing the character to perform the first action within a computer-based or physical environment. [12] One or more non-transitory computer-readable media according to claim 11, wherein generating the first action comprises: sampling a prior distribution based on the first state and the one or more constraints to generate a latent vector; and Processing the latent vector and the first state using a controller included in the trained machine learning model to generate the first action. [13] One or more non-transitory computer-readable media according to claim 12, wherein the instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of processing the first state and the one or more constraints using a transformer encoder to generate the prior distribution. [14] One or more non-transitory computer-readable media according to claim 13, wherein the instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of training a first machine learning model to generate the trained machine learning model, the first machine learning model comprising an encoder. [15] One or more non-transitory computer-readable media according to any one of claims 11 to 14, wherein the instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of training a first machine learning model to generate the trained machine learning model by: sampling a first motion from a set of motion recordings and a time step within the first motion to generate a sampled motion; removing at least one joint or at least one image within the sampled motion to create a masked motion; Generating, via the first machine learning model and based on a second state of the character and the masked movement, a second action for the character to perform; Causing the character to perform the second action within the computer-based environment to achieve a third state of the character; and Updating one or more parameters of the first machine learning model based on a comparison between the third state and a fourth state of the character included in the sampled motion. [16] One or more non-transitory computer-readable media according to claim 15, wherein the instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of terminating training the first machine learning model using the sampled motion based on a similarity between the third state and the fourth state being less than a predefined threshold. [17] One or more non-transitory computer-readable media according to any one of claims 11 to 16, wherein the instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of training a first machine learning model to generate the trained machine learning model based on a reward that is a metric of a comparison between movements generated by the first machine learning model and movements sampled from a set of motion recordings. [18] One or more non-transitory computer-readable media according to one of claims 11 to 17, wherein a controller controlling one or more joints of the character causes the character to move according to the first action. [19] One or more non-transitory computer-readable media according to any one of claims 11 to 18, wherein the environment is at least one of a simulation environment, an augmented reality environment, XR, a game environment, and a physical environment. [20] System comprising: one or more memories that store instructions; and one or more processors coupled to the one or more memories and configured to, in executing the instructions: to obtain a first state of a character and one or more constraints on one or more movements associated with a subset of joints belonging to the character; to generate, via a trained machine learning model and based on the first state and the one or more constraints, a first action for the character to perform; and to cause the character to perform the first action within a computer-based or physical environment.