Whole body motion tracking for use in virtual environments

Through the diffusion model of a multi-layer perception (MLP) network, the whole-body posture is predicted based on sparse upper body tracking signals, and the accuracy and comfort problems of whole-body motion tracking in the prior art are solved, achieving efficient whole-body posture prediction.

CN120476369APending Publication Date: 2025-08-12CTRL-LABS CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480006109.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-12
Filing Date
2024-01-16
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The prior art is difficult to accurately track whole-body movement, especially lower body posture, and existing methods usually rely on more than 3 inputs, resulting in high cost, low accuracy and poor user comfort.

Method used

A diffusion model using a multi-layer perception (MLP) network is used to predict the whole-body posture based on sparse upper body tracking signals. The upper body and lower body posture are generated by training the diffusion model, and the time step embedding is used to alleviate the jitter problem.

Benefits of technology

Highly accurate full-body motion tracking, especially lower body posture prediction, reduces costs and improves user experience comfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120476369A_ABST
    Figure CN120476369A_ABST
Patent Text Reader

Abstract

A method for whole body motion tracking includes receiving tracking signals from a plurality of sensors associated with an upper body of a person; and determining a motion feature and a joint feature based on the tracking signal. The method further includes training a diffusion model, the diffusion model including a multi-layer perception (MLP) network; and generating a plurality of inputs to the trained diffusion model, the plurality of inputs including motion features and joint features. The method includes providing the plurality of inputs to a trained diffusion model to generate a plurality of outputs. The plurality of outputs includes a sequence of whole body poses, and the sequence of whole body poses includes an upper body pose and a lower body pose.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 479,924, filed January 13, 2023, the entire contents of which are incorporated herein. Technical Field

[0003] The present disclosure relates generally to accurately tracking full-body motion for realistically controlling a full-body avatar in a virtual environment, and more particularly to tracking the full body from sparse upper-body tracking signals. Background Art

[0004] The movement of an avatar in augmented reality (AR) and virtual reality (VR) applications is typically determined by the user's real-world motion tracked via inertial measurement unit (IMU) sensors in a head-mounted device (HMD) and / or handheld device. Full-body motion tracking is desirable because it provides an engaging experience in which the user can interact with the virtual environment with an enhanced sense of presence.

[0005] A particular challenge with full-body motion tracking is that the available tracking signals from standalone HMDs are typically limited to tracking the user's head and wrists. While this signal may be helpful for reconstructing upper-body motion, lower-body motion is not directly tracked. Any lower-body tracking solution that uses additional tracking signals from lower-body joints (e.g., additional IMUs) and is applied to avatar motion in AR / VR applications will incur higher costs, lower accuracy, and sacrifice user comfort. Existing methods for full-body motion tracking rely on more than three inputs and / or have difficulty predicting the whole-body pose, especially the lower-body pose.

[0006] Since user motion tracking is the primary source of avatar manipulation in AR / VR applications, there is a need to improve the accuracy of motion tracking, and more specifically, to provide full-body motion tracking. Summary of the Invention

[0007] According to some embodiments, a method for whole-body motion tracking includes: receiving tracking signals from a set of sensors associated with a person's upper body; and determining motion features and joint features based on the tracking signals. The method also includes: training a diffusion model, the diffusion model including a multi-layer perceptron (MLP) network; and generating a set of inputs to the trained diffusion model, the set of inputs including motion features and joint features. The method also includes: providing the set of inputs to the trained diffusion model to generate a set of outputs. The set of outputs includes a whole-body pose sequence, the whole-body pose sequence including an upper-body pose and a lower-body pose.

[0008] According to some embodiments, a non-transitory computer-readable medium stores a program for whole-body motion tracking, which, when executed by a computer, configures the computer to: receive tracking signals from a set of sensors associated with a person's upper body; and determine motion features and joint features based on the tracking signals. The program, when executed, further configures the computer to: train a diffusion model comprising a multi-layer perceptron (MLP) network; generate a set of inputs to the trained diffusion model, the set of inputs comprising motion features and joint features; and provide the set of inputs to the trained diffusion model to generate a set of outputs. The set of outputs comprises a whole-body pose sequence, the whole-body pose sequence comprising upper-body poses and lower-body poses.

[0009] According to some embodiments, a system for whole-body motion tracking includes: a processor; and a non-transitory computer-readable medium storing a set of instructions that, when executed by the processor, configure the processor to: receive tracking signals from a set of sensors associated with a person's upper body; and determine motion features and joint features based on the tracking signals. The instructions, when executed, configure the processor to: train a diffusion model comprising a multi-layer perceptron (MLP) network; generate a set of inputs to the trained diffusion model, the set of inputs comprising motion features and joint features; and provide the set of inputs to the trained diffusion model to generate a set of outputs. The set of outputs comprises a whole-body pose sequence, the whole-body pose sequence comprising an upper-body pose and a lower-body pose. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings, which are included to provide a further understanding and are incorporated in and constitute a part of this specification, illustrate the disclosed embodiments and together with the description serve to explain the principles of the disclosed embodiments.

[0011] Figure 1 A network architecture for implementing full-body motion tracking is shown in accordance with some embodiments.

[0012] Figure 2is a diagram showing the Figure 1 A block diagram with details of the devices used in the architecture.

[0013] Figure 3 is a flow chart illustrating a process for whole-body motion tracking according to some embodiments.

[0014] Figure 4 An example of whole-body motion synthesis using a diffusion model is shown, in accordance with some embodiments.

[0015] Figure 5 is a block diagram illustrating an MLP-based network according to some embodiments.

[0016] Figure 6 is a block diagram illustrating an MLP-based diffusion model according to some embodiments.

[0017] Figure 7 A qualitative comparison between the diffusion model 600 and AvatarPoser of some embodiments is shown.

[0018] Figure 8 A visualization of the motion trajectory between the diffusion model 600 and the AvatarPoser of some embodiments is shown.

[0019] Figure 9 is a block diagram illustrating an exemplary computer system with which aspects of the subject technology may be implemented, according to some embodiments.

[0020] In one or more embodiments, not all components depicted in each figure are required, and one or more embodiments may include additional components not shown in the figures. Various modifications may be made to the arrangement and types of these components without departing from the scope of the present disclosure. Additional components, different components, or fewer components may be used within the scope of the present disclosure. DETAILED DESCRIPTION

[0021] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the various embodiments of the present disclosure may be practiced without some of these specific details. In other instances, well-known structures and techniques are not shown in detail in order to avoid obscuring the present disclosure.

[0022] According to some embodiments, the term "virtual reality" as used herein refers to a computer-generated simulation of an immersive, two-dimensional or (more typically) three-dimensional environment that people can explore and interact with through sensory stimulation (typically using a headset or other dedicated equipment). Virtual reality (abbreviated as "VR") can provide a sensory-rich experience that simulates and / or replicates the real world or imagined scenes, enabling users to participate in activities, manipulate objects and perceive the simulated environment as if it were real. VR systems typically include visual feedback, auditory feedback and / or tactile feedback to enhance the sense of presence and immersion, bringing users into a digitally generated world where users can interact, learn or experience various scenes in a highly immersive and interactive manner. As used herein, the term "virtual reality" is understood to include the Internet and "augmented reality".

[0023] According to some embodiments, the term "diffusion model" as used herein refers to a likelihood-based generative model that learns to invert random Gaussian noise introduced by a Markov chain in order to recover the desired data samples from the noise. Diffusion models may require injecting step-size embeddings into the network during both the training and inference phases of the model.

[0024] According to some embodiments, the term "full body motion tracking" as used herein refers to technology that captures and interprets the motion and position of a person's entire body in real time. The technology utilizes a combination of sensors, cameras, and / or other tracking devices to collect data about the user's motion, including joint angles, limb positions, and gestures. After comprehensively capturing these motions, the motions are reconstructed and converted into digital representations to allow the user's body motion to be accurately mapped to an avatar or virtual character within a digital environment. This technology enables immersive experiences in virtual reality (VR), augmented reality (AR), gaming, sports analysis, healthcare, and various other applications by accurately tracking the user's body movements and replicating these movements into digital space. Full body motion tracking may be equivalently referred to herein as "full body motion synthesis" or "full body motion prediction."

[0025] As a technical solution to the above-mentioned technical problems, some embodiments of the present disclosure provide a model that tracks full-body motion (including upper and lower-body motion) based solely on sparse upper-body tracking signals and provides accurate posture prediction, particularly for the lower body. For example, some embodiments can achieve high-fidelity full-body tracking using only the three standard inputs (head and hands) provided by most HMDs. Compared to existing methods for motion prediction, this model can be more robust to tracking signal loss.

[0026] For example, the model can use a multi-layer perceptron (MLP) architecture and a conditioning scheme for motion data to predict accurate and smooth full-body motion, particularly lower-body motion. The model can include a compact architecture for generating realistic and smooth motion while achieving real-time inference speed, making it useful for online body tracking applications (e.g., AR / VR applications).

[0027] According to some embodiments, the model can be a conditional diffusion model. Time step embeddings can be injected during the diffusion process to alleviate jitter issues and improve model performance and robustness to tracking signal loss. Time step embeddings can implement a block-wise injection scheme that adds a diffusion time step embedding before each intermediate block of the neural network (NN). This injection scheme achieves gradual denoising and produces smooth motion sequences.

[0028] Given a sequence of N observed joint features, some embodiments include: predicting full body pose based on input joint features / output joint features from the N observed joint frames. The diffusion model can be a conditional model that is configured to generate a full body pose sequence conditioned on the sparse tracking of the observed joint features. In other words, the diffusion model of some embodiments can predict body pose. The diffusion model can utilize an MLP-based network to perform full body motion synthesis based on sparse tracking signals, such that each block M of the MLP network includes both a convolutional layer and a fully connected layer, and the convolutional layer and the fully connected layer are responsible for merging temporal information and merging spatial information, respectively.

[0029] In some embodiments, the motion features and the observed joint features at time t can be passed through fully connected layers separately to obtain intermediate features. The intermediate features of N frames can be connected and fed to the MLP network. The embedding of time step t can be repeatedly injected before each block M of the MLP network, rather than taking the embedding of time step t as an additional input to the network (for example, by connecting the embedding to the joint features). The time step embedding can be projected to match the dimension of the input joint features and passed through the fully connected layers and sigmoid linear unit (SiLU) activation layers of the MLP network. Subsequently, the obtained intermediate features can be added directly to the input intermediate activations. In this way, the diffusion model according to each embodiment can greatly alleviate the jitter problem and achieve the synthesis of smooth motion for application to online avatars.

[0030] In some embodiments, the N observed joint frames can be set to 196 joint frames (i.e., N=196), and the joint rotation can be represented by a 6D reparameterization. The MLP network may include 12 blocks (i.e., M=12). The diffusion model can be trained using two settings to predict the global orientation of the root joint and the relative rotations of other joints. During inference, the diffusion model can be applied autoregressively to longer sequences.

[0031] Some embodiments provide a computer-implemented method comprising: receiving motion data from a sensor regarding an augmented reality (AR) device / virtual reality (VR) device, the motion data representing a sparse tracking signal; generating a model for predicting full-body posture, the model comprising a multi-layer network; determining motion features and joint features based on the motion data; obtaining intermediate features using the model based on the motion features and joint features; generating a full-body posture sequence based on the intermediate features; and estimating the position of the user's legs based on the full-body posture sequence.

[0032] Some embodiments provide a non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform a method comprising: receiving motion data from a sensor regarding an augmented reality (AR) device / virtual reality (VR) device, the motion data representing a sparse tracking signal; generating a model for predicting whole-body posture, the model comprising a multi-layer network; determining motion features and joint features based on the motion data; obtaining intermediate features using the model based on the motion features and joint features; generating a whole-body posture sequence based on the intermediate features; and estimating the position of the user's legs based on the whole-body posture sequence.

[0033] Figure 1 A network architecture 100 for implementing full-body motion tracking according to some embodiments is shown. The architecture 100 may include a server 130 and a database 152, which are communicatively coupled to a plurality of client devices 110 via a network 150. The client devices 110 may include any of the following: a laptop computer; a desktop computer; or a mobile device such as a head-mounted device (HMD), a smartphone, a handheld device, a video player, or a tablet device. The database 152 may store, for example, backup files from a matrix, video, and processed data.

[0034] The network 150 may include, for example, any one or more of the following: a local area network (LAN); a wide area network (WAN); and the Internet. Furthermore, the network 150 may include, but is not limited to, any one or more of the following network topologies, including a bus network, a star network, a ring network, a mesh network, a star-bus network, and a tree or hierarchical network.

[0035] Figure 2 1 is a block diagram illustrating details of a system 200 for use in the network architecture disclosed herein (e.g., architecture 100), according to some embodiments, having at least one client device 110 and at least one server 130. The client device 110 and the server 130 are communicatively coupled via a network 150 via respective communication modules 218-1 and 218-2 (hereinafter collectively referred to as "communication modules 218"). The communication modules 218 are configured to interface with the network 150 to send and receive information, such as requests, uploaded content, messages, and commands, to other devices on the network 150. The communication modules 218 may be, for example, a modem or an Ethernet card and may include radio hardware and software for wireless communication (e.g., via electromagnetic radiation, such as radio frequency (RF), near field communication (NFC), Wi-Fi, and Bluetooth radio technologies). The client device 110 may be coupled to an input device 214 and an output device 216. A user may interact with the client device 110 via the input device 214 and the output device 216. The input device 214 may include a mouse, keyboard, pointer, touch screen, microphone, joystick, virtual joystick, or touch screen display that a user can use to interact with the client device 110. In some embodiments, the input device 214 may include a camera, microphone, and sensors such as touch sensors, acoustic sensors, inertial motion units (IMUs), and other sensors configured to provide input data to the VR / AR head-mounted viewer. The output device 216 may be a screen display, touch screen, and speaker.

[0036] The client device 110 may further include a processor 212-1 configured to execute instructions stored in a memory 220-1 and cause the client device 110 to perform at least some of the operations in the method consistent with the present disclosure. The memory 220-1 may further include an application 222 configured to run in the client device 110 and coupled to the input device 214 and the output device 216. The application 222 may be downloaded by a user from the server 130 and may be hosted by the server 130. The application 222 may include specific instructions that, when executed by the processor 212-1, cause operations to be performed according to the method described herein. In some embodiments, the application 222 runs on an operating system (OS) installed in the client device 110. In some embodiments, the application 222 may run in a web browser. In some embodiments, the processor is configured to control a graphical user interface (GUI) for a user of one of the client devices 110 to access the server 130.

[0037] Database 252 can store data and files associated with server 130 from application 222. In some embodiments, client device 110 is a mobile phone that is used to collect videos or pictures and upload them to server 130 using video or image collection application 222 for storage in database 252.

[0038] The server 130 includes a memory 220-2, a processor 212-2, and a communication module 218-2. Hereinafter, the processors 212-1 and 212-2, and the memories 220-1 and 220-2 will be collectively referred to as "processor 212" and "memory 220," respectively. The processor 212 is configured to execute instructions stored in the memory 220. In some embodiments, the memory 220-2 includes an application engine 232. The application engine 232 can be configured to perform operations and methods according to aspects of the embodiments. The application engine 232 can share features and resources with the client device, or provide features and resources, including: multiple tools associated with the collection and acquisition of data, images, or videos, or applications (e.g., application 222) that use data, images, or videos obtained using the application engine 232. A user can access the application engine 232 through the application 222 installed in the memory 220-1 of the client device 110. Thus, application 222 may be installed by server 130 and, through any of a number of tools, execute scripts and other routines provided by server 130. Execution of application 222 may be controlled by processor 212-1.

[0039] Figure 33 is a flow chart illustrating a process 300 for whole-body motion tracking performed by a server (e.g., server 130, etc.) or a device (e.g., client device 110, etc.) in accordance with some embodiments. In some embodiments, one or more operations in process 300 may be performed by a processor circuit (e.g., processor 212, etc.) executing instructions stored in a memory circuit (e.g., memory 220, etc.) of a system disclosed herein (e.g., system 200). Furthermore, in some embodiments, a process consistent with the present disclosure may include at least the operations in process 300 performed in a different order, simultaneously, quasi-simultaneously, or in a temporally overlapping manner.

[0040] At 310, process 300 receives tracking signals from a plurality of sensors associated with an upper torso of a person. In some embodiments, the plurality of sensors include, but are not limited to, an inertial measurement unit (IMU). The sensors may be mounted to an upper torso device, such as a head mounted device (HMD) or a handheld device. For example, the sensors may be located in an HMD and two handheld devices, one of which is located in each hand of the person. The tracking signals may include, but are not limited to, the orientation and translation of each of the devices.

[0041] At 320 , process 300 determines motion characteristics and joint characteristics based on the tracking signals.

[0042] At 330, process 300 trains a diffusion model comprising a multi-layer perceptron (MLP) network. In some embodiments, the MLP network comprises a plurality of blocks. Each of the plurality of blocks may comprise a convolutional layer and a fully connected layer. Each block may also comprise a sigmoid linear unit activation layer and layer normalization.

[0043] At 340, process 300 generates multiple inputs to the trained diffusion model. The multiple inputs may include motion features and joint features. In some embodiments, the inputs to the diffusion model also include intermediate features generated based on the motion features and joint features.

[0044] In some embodiments, the MLP network includes multiple blocks, and the process 300 provides a time-step embedding to each block. For example, the time-step embedding can be provided to each block through a fully connected layer and a sigmoid linear unit activation layer.

[0045] At 350, process 300 provides the plurality of inputs to the trained diffusion model to generate a plurality of outputs. In some embodiments, the outputs include a full-body pose sequence. The full-body pose sequence includes an upper-body pose and a lower-body pose.

[0046] In some embodiments, the output may also include the position of the person's lower body.Process 300 may also include estimating the position of the lower body based on the full body pose sequence.

[0047] In some embodiments, output is generated from the MLP network based on the intermediate features. For example, a full body pose sequence can be generated based on the intermediate features.

[0048] Experimental methods

[0049] Problem Statement

[0050] Some embodiments predict full body motion from sparse tracking signals (i.e., the orientation and translation of the head-mounted view and two hand controllers). Given a sequence of N observed joint features In the case of N frames, it is expected to predict the whole body posture Where C and S represent the dimensions of input joint features / output joint features. In this example, a skinned multi-person linear model (SMPL) is used to represent human poses. Only the first 22 joints of the SMPL model are used, ignoring the hand joints. Therefore, Represents the global orientation of the pelvis and the relative rotation of each joint.

[0051] Some of the examples below use a simple MLP-based network to perform full-body motion synthesis based on sparse tracking signals. The performance can be further improved by leveraging the proposed MLP-based architecture to support a conditional generative diffusion model, called "Avatars GrowLegs" (AGRoL).

[0052] Figure 4 An example of full body motion synthesis using a diffusion model based on HMD input and hand controller input according to some embodiments is shown. The RGB axes describe the orientation of the head and hands, which are used as input to the model.

[0053] MLP-based networks

[0054] Figure 5 is a block diagram illustrating an MLP-based network 500 according to some embodiments. In this example, the MLP-based network 500 includes only four types of components: fully connected layers, SiLU activation layers, 1D convolutional layers with a kernel size of 1, and layer normalization. FC, LN, and SiLU represent fully connected (FC) layers 510, layer normalization 520, and SiLU activation layers 530, respectively. 1×1 convolution (1×1Conv) represents a 1D convolutional layer with a kernel size of 1 540. Note that 1×1Conv here is equivalent to adding a kernel size of 1 to the input tensor. The fully connected layer operates on the first dimension of N, while the FC layer operates on the last dimension. N represents the time dimension, represents the dimensionality of the latent space.

[0055] Each intermediate block 550 of the MLP-based network 500 includes a convolutional layer 540 and a fully connected (FC) layer 510, which are responsible for temporal information merging and spatial information merging, respectively. In this example, the intermediate block 550 is repeated M times.

[0056] Some embodiments use skip connections as pre-normalization for layers. The first layer (FC layer 510) of the MLP-based network 500 converts the input data 560 (denoted as p 1:N ) projected into the latent space The last layer (FC layer 510) transforms the latent space into the output space of full body pose (For example, output data 570, represented as ).

[0057] Diffusion Model

[0058] The diffusion model is a generative model that learns to invert the random Gaussian noise introduced by a Markov chain in order to recover the expected data samples from the noise.

[0059] Figure 6 is a block diagram illustrating an MLP-based diffusion model 600 according to some embodiments. The diffusion model 600 includes an MLP-based network (e.g., MLP-based network 500) and FC layers 510 at the input and output. In this example, t is the noise step, and the input 610 (denoted as ) represents a motion sequence of length N at step length t, which can be pure Gaussian noise at t=0. Input 620 (denoted as p 1:N ) represents the sparse upper body signal of length N. Output 630 (represented as ) represents the denoised motion sequence at step size t.

[0060] In the forward diffusion process, the sample motion sequence in a given data distribution In the case of , the Markov noise process can be written as:

[0061]

[0062] where α t ∈(0,1)∈(0,1) is a constant hyperparameter, and I is the identity matrix. When t→∞, tends to an isotropic Gaussian distribution. Then, in the back-diffusion process, a model p with parameters θ is trained θ , with the variance fixed as Gaussian noise input x TGenerate samples from ~N(0,1). Formally,

[0063]

[0064] where μ θ can be rewritten as,

[0065]

[0066] in Therefore, the diffusion model 600 must learn the t and time step t prediction noise ∈ θ (x t ,t).

[0067] The desired situation is to use the diffusion model 600 to generate the joint feature p 1:N The sparse tracking of is conditioned on the full body pose sequence. Therefore, the back diffusion process becomes conditional: In addition, clean body posture is predicted directly without predicting the residual noise ∈ θ (x t ,t). Model f θ The output is represented as The objective function can be expressed as

[0068]

[0069] like Figure 6 As shown in the example in , some embodiments use the above-mentioned MLP-based network 500 as a model for predicting full body posture. θ At time step t, the motion feature and the observed joint features p 1:N First, they are passed through the fully connected (FC) layer 510 to obtain the intermediate features 640, which are represented as and and is defined as:

[0070]

[0071]

[0072] These intermediate features 640 may be concatenated together and fed into the MLP-based network 500 .

[0073]

[0074] Blocked time step embedding. In some diffusion models, the embedding of time step t is fed to the network as an additional input. However, because some embodiments use an MLP, the diffusion model 600 may be insensitive to the value of the time step embedding, which can hinder the learning denoising process and lead to predicted motion with severe jitter (as shown below (see "Experiments")).

[0075] To solve this problem, Figure 6 As shown in the example of [ ], some embodiments repeatedly inject the time step embedding 650 before each block of the MLP network. The time step embedding is projected to match the input feature dimension through the fully connected (FC) layer 510 and the SiLU activation layer 530, and the resulting feature is directly added to the input intermediate activation. As shown below (see "Experiments"), the proposed strategy can greatly alleviate the jitter problem and achieve smooth motion synthesis.

[0076] experiment

[0077] An embodiment of the diffusion model 600 is trained and evaluated on the AMASS dataset. Two settings are used for training and testing to compare with prior art. For the first setting, three subsets are used: CMU, BMLr, and HDM05. For the second setting, a data split is used with CMU, MPI Limits, Total Capture, Eyes Japn, KIT, BioMotion-Lab, BMLMovi, EKUT, ACCAD, MPI Mosh, SFU, and HDM05 as training data and HumanEval and Transition as test data. In both settings, the SMPL human body model is used as the human pose representation and the diffusion model 600 is trained to predict the global orientation of the root joint and the relative rotations of the other joints.

[0078] Implementation details

[0079] For simplicity and continuity, joint rotations are represented by 6D reparameterization. Therefore, for the body pose sequence If not stated otherwise, the number of frames is set to N=196.

[0080] MLP network. In this example, the MLP-based network 500 is built using 12 blocks (M=12). All latent features in the MLP-based network 500 have the same shape N×512. The network is trained using a batch size of 256 and the Adam optimizer. The learning rate is initially set to 3e-4 and is reduced to 1e-5 after 200,000 iterations. Weight decay is set to 1e-4 throughout training. During inference, the diffusion model 600 is applied to longer sequences in an autoregressive manner.

[0081] MLP-based diffusion model (AGRoL). The architecture of the MLP-based network 500 remains unchanged in the diffusion model 600. To inject the time-step embeddings used in the diffusion process into the network, in each MLP block, the time-step embeddings 650 are passed to the fully connected (FC) layer 510 and the SiLU activation layer 530 and the input features are added. The network is trained with exactly the same hyperparameters as the MLP-based network 500, but using the AdamW optimizer. During training, the sampling steps are set to 1000 with a cosine noise schedule. The denoising diffusion implicit model (DDIM) technique is used to accelerate sampling, and only 5 steps are sampled during inference.

[0082] All experiments were performed using the Pytorch framework on a single NVIDIA V100 graphics card.

[0083]

[0084]

[0085] Table 1

[0086] Table 1 provides a comparison of the methods of some embodiments with the current state-of-the-art techniques on a subset of the AMASS dataset. Table 1 reports the MPJPE [cm], MPJRE [deg], MPJVE [cm / s], and Jitter [10 2 m / s 3 ] metrics. In this example, AGRoL achieves the best performance on MPJPE, MPJRE, and MPJVE, and outperforms the other models, especially on the lower body position error (Lower PE) metric and the jitter metric, indicating that the diffusion model 600 generates accurate lower body motion and smooth motion in this example. The best results are in bold, and the second-best results are underlined.

[0087]

[0088] Table 2

[0089] Table 2 provides a comparison of the methods of some embodiments with the current state-of-the-art techniques on the AMASS dataset. Table 2 reports the MPJPE [cm] index, MPJRE [deg] index, MPJVE [cm / s] index and jitter [10 2 m / s 3 ] metrics. * indicates Avatar-Poser was retrained using public code. Represents techniques that use pelvic position and pelvic rotation during inference, which may not be directly comparable because some embodiments assume pelvic information is not available during training and testing. The best results are in bold, and the second best results are underlined.

[0090] Evaluation Metrics

[0091] In this example, 10 metrics are used to evaluate the diffusion model 600. These metrics can be divided into three categories. The first category is rotation-related metrics, including mean per-joint rotation error [degrees] (MPJRE) and root rotation error [degrees] (root RE). These metrics measure the average relative rotation error of all joints and the global rotation error of the root joint. The second category is velocity-related metrics, including mean per-joint velocity error [cm / s] (MPJVE) and jitter. Mean per-joint velocity error [cm / s] (MPJVE) measures the average velocity error of all joints. Jitter is expressed as 10 2 Measuring the average jerk (the time derivative of acceleration) of all joints in global space in m / s, jitter reflects the smoothness of motion. The third category is position-related metrics, which include all other metrics. Specifically, the mean per-joint position error [cm] (MPJPE) measures the average position error of all joints. Root PE evaluates the root position error. Hand PE measures the average position error of both hands. Upper body PE and lower body PE evaluate the average position error of upper and lower body joints, respectively.

[0092] Evaluation results

[0093] In this example, the diffusion model 600 was evaluated on the AMASS dataset using two different protocols. As shown in Tables 1 and 2, the MLP-based network 500 outperformed most previous techniques and achieved comparable results, demonstrating the effectiveness of our proposed simple network. With the help of the diffusion process, the AGRoL diffusion model 600 further improved the performance of the MLP-based network 500 and surpassed all previous techniques. Furthermore, the AGRoL diffusion model 600 significantly reduced the jitter error, meaning that the generated motion was much smoother than that of the other models. Figure 7 and Figure 8 Some examples are visualized in . Figure 7Shows the reconstruction error comparison between the AGROL diffusion model 600 and AvatarPoser. Figure 8 In

[15] , a smoothness comparison between the AGRoL diffusion model 600 and AvatarPoser is shown by visualizing the pose trajectories.

[0094] Figure 7 A qualitative comparison of the AGRoL diffusion model 600 (bottom half) and AvatarPoser (top half) of some embodiments on a test sequence from the AMASS dataset is shown. The predicted skeleton and body mesh are visualized in the figure. The green skeleton represents the motion predicted using the AGRoL diffusion model 600. The red skeleton represents the motion predicted using AvatarPoser. The blue skeleton represents the ground truth motion. As shown in the figure, the predicted motion of the AGRoL diffusion model 600 is more accurate than the predicted motion of AvatarPoser.

[0095] Figure 8 A visualization of the motion trajectories between the AGRoL diffusion model 600 and AvatarPoser of some embodiments is shown. The trajectories of the predicted motion are visualized in the figure. The image on the left shows the ground truth motion with a blue skeleton. The image in the middle shows the predicted motion of the AGRoL diffusion model 600 with a green skeleton. The image on the right shows the predicted motion of AvatarPoser with a red skeleton. The light purple vectors in the figure represent the velocity vector of each joint. By visualizing the motion trajectories, the jitter problem and the foot sliding problem can be better seen from the figure. Smooth motion tends to have regular posture trajectories with stable changes in the velocity vector of each joint. The density of the posture trajectories changes with walking speed, and the trajectories become denser when the person slows down. Therefore, if there is no foot sliding, you should occasionally see changes in the density of the posture trajectories.

[0096] Ablation experiments

[0097] The MLP-based network 500 of some embodiments is compared with other networks using the diffusion model 600 described above to show the effectiveness of the MLP-based network 500. For the diffusion model 600, the time-step embedding is also ablated, and different strategies are used to add the time-step embedding. The impact of the additional loss and the number of sampling steps used during inference are also studied.

[0098] Architecture

[0099] To verify the effectiveness of the MLP-based network 500 in the diffusion model setting, the MLP-based network 500 is replaced with other types of networks and the results are compared. Two architectures are considered: the network from AvatarPoser, and the transformer network. In the case of the transformer network, instead of repeatedly injecting the temporal position embedding into each block, the temporal position embedding is combined with the input features. and concatenate them before feeding them into the network. The same strategy was applied to the AvatarPoser network, as this model also used transformer blocks in the early stages. To establish a fair comparison with the AvatarPoser architecture, two versions of the model were trained: one using the original settings and the other using more transformer layers to achieve a size comparable to the proposed diffusion model 600. The same experiments were performed on the transformer network. As shown in Table 3, the MLP-based network 500 proposed in some embodiments achieved superior results when trained in a diffusion manner compared to the other networks.

[0100]

[0101] Table 3

[0102] Table 3 summarizes ablation experiments on the network architectures used in diffusion model 600 in some embodiments. MLP-based network 500 was replaced with other networks and trained using the same hyperparameters as the diffusion model. In this example, MLP-based network 500 outperformed all other networks on most metrics. AvatarPoser-Large represents a network with the same architecture as AvatarPoser but with additional transformer layers. The best results are in bold, and the second-best results are underlined.

[0103] Diffusion time step embedding

[0104] Some embodiments of the diffusion model 600 use a time-step embedding to indicate the noise addition step t during the diffusion process. In this example, a sinusoidal position embedding is used as the time-step embedding. Table 4 shows the results of the AGRoL diffusion model 600 without the time-step embedding. The diffusion model 600 can still achieve good performance on metrics related to position error and rotation error, while performance on metrics related to velocity error (MPJVE and jitter) may degrade significantly. Due to the lack of a time-step embedding, the diffusion model 600 may not know which step it is at and may therefore not be able to denoise correctly.

[0105] In this example, to apply time-step embeddings to the diffusion model 600, three strategies are ablated: Add, Concat, and RepIn. Compared to RepIn, which repeatedly passes the time-step embeddings through the linear layer and injects the time-step embeddings into each block of the MLP network, in Add and Concat, the time-step embeddings are used only once at the beginning of the network. Here, the time-step embeddings are first passed through the fully connected layer and the SiLU activation layer to obtain the latent features. Then it is fed into the network. Specifically, Add u and the input features and Sum, so the output of the network is Concat u with the input features and Connected together, the output of the network is RepIn represents the strategy of adding time-step embeddings. Specifically, for each block of the MLP network, the time-step embeddings are projected to pass through the fully connected layer and the SiLU activation layer respectively, and then the obtained features u i ,i∈[0,..M] is added to the input features of its corresponding block. As shown in Table 4, the proposed strategy can greatly improve the speed-related metrics and alleviate the jitter problem to generate smooth motion.

[0106]

[0107] Table 4

[0108] Table 4 summarizes the ablation of time step embedding according to some embodiments. w / o time represents the results of AGRoL without time step embedding. Add sums the features from the time step embedding with the input features. Concat concatenates the features from the time step embedding with the input features. In Add and Concat, the time step embedding is fed only once at the top of the network. Repeated injection (RepIn) represents the strategy of injecting the time step embedding into each block of the network. As shown in the table, time step embedding mainly affects the MPJVE metric and the Jitter metric. In the absence of time step embedding, or when the time step embedding is added incorrectly, large errors in speed-related metrics may occur, causing serious jitter problems.

[0109] Additional losses

[0110] Apart from In addition, three other geometric losses during training are studied:

[0111]

[0112] where FK() is the forward kinematics function that takes as input the local body joint rotations and outputs the positions of these joints in the global coordinate space. represents the position loss of the joint, represents the velocity loss of the joint in 3D space, and Indicates foot contact loss, Forces the foot to be still when there is no foot movement. i ∈{0,1} represents a binary mask and is equal to 0 when the velocity of the foot joint is zero.

[0113] In this example, the diffusion model 600 is trained with different combinations of additional losses, with the weights of these combinations set to 1. The additional geometric loss does not bring any additional performance to the diffusion model 600. The diffusion model 600 achieves good results when trained only with the denoising objective function formula (4). One reason why the additional loss does not improve the performance of the AGRoL diffusion model 600 may be due to the internal workings of the back-diffusion process, which does not interact with the additional geometric loss without proper tuning.

[0114]

[0115]

[0116] Table 5

[0117]

[0118] Table 6

[0119]

[0120] Table 7

[0121] Table 5 summarizes the ablation of the additional loss used during the training of the diffusion model 600 according to some embodiments.

[0122] Table 6 summarizes the ablation of the number of sampling steps during inference according to some embodiments. The input and output lengths are fixed to N=196.

[0123] Table 7 summarizes the robustness of the diffusion model 600 to joint tracking loss, according to some embodiments. Different techniques were evaluated by randomly masking a portion (10%) of the input frames during inference on the AMASS dataset. Each technique was tested five times and the results were averaged. In this example, the AGRoL diffusion model 600 achieved the best performance among all techniques, demonstrating its robustness to joint tracking loss.

[0124] Number of sampling steps during inference

[0125] In this example, the number of sampling steps used during inference was varied. An embodiment of the diffusion model 600 was used, which was trained with 1000 sampling steps and tested with a subset of the steps in the diffusion process. Five DDIM sampling steps were used, allowing the diffusion model 600 to achieve excellent performance on most metrics while being fast.

[0126] Robustness to tracking loss

[0127] In this example, the robustness of some embodiments of the diffusion model 600 to tracking loss of the input joints is studied. In practice, a common problem in VR applications is that the input becomes sparse in time due to the hands or controllers being out of the field of view, resulting in a loss of joint tracking signals on some frames. All available techniques are evaluated for their performance with respect to tracking loss by randomly masking a portion of the input frames during inference. Table 7 shows the results. The performance of the other techniques degrades significantly, indicating that these techniques are not robust to the tracking loss problem. In contrast, the accuracy of the diffusion model 600 decreases less, indicating that the diffusion model 600 can accurately model motion based on highly sparse tracking input.

[0128] Inference speed

[0129] The AGRoL diffusion model 600 of some embodiments achieves real-time inference speed due to its lightweight architecture combined with DDIM sampling. On a single NVIDIA V100 GPU, a single run of the AGRoL generation process with five DDIM sampling steps produced 196 output frames in 35 milliseconds (ms). The predicted MLP-based diffusion model 600 takes 196 frames as input and predicts the final result for 196 frames in a single forward pass. It is even faster, taking only 6ms on a single NVIDIA V100 GPU.

[0130] Conclusions and limitations

[0131] Some embodiments provide an MLP-based architecture with building blocks for achieving competitive performance in whole-body motion synthesis tasks. Some embodiments provide AGRoL, a conditional diffusion model 600 for whole-body motion synthesis based on sparse tracking signals. The AGRoL diffusion model 600 utilizes a simple yet effective conditioning scheme for structured human motion data. This paper demonstrates that this lightweight diffusion-based model generates realistic and smooth human motion while achieving real-time inference speed, making it suitable for online AR / VR applications.

[0132] Figure 99 is a block diagram illustrating an exemplary computer system 900 with which aspects of the subject technology may be implemented. In certain aspects, computer system 900 may be implemented using hardware or a combination of software and hardware in a dedicated server, integrated into another entity, or distributed across multiple entities.

[0133] The computer system 900 (e.g., a server and / or client) includes a bus 908 or other communication mechanism for communicating information, and a processor 902 coupled with the bus 908 for processing information. As an example, the computer system 900 can be implemented using one or more processors 902. The processor 902 can be a general-purpose microprocessor, a microcontroller, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable logic device (PLD), a controller, a state machine, gate logic, discrete hardware components, or any other suitable entity that can perform calculations or other information manipulations.

[0134] In addition to the hardware, the computer system 900 may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof, which is stored in an included memory 904, such as random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable PROM (EPROM), registers, a hard disk, a removable disk, a compact disc read-only memory (CD-ROM), a digital video disc (DVD), or any other suitable storage device, coupled to the bus 908 for storing information and instructions to be executed by the processor 902. The processor 902 and the memory 904 may be supplemented by, or incorporated in, special purpose logic circuitry.

[0135] Instructions may be stored in memory 904 and may be implemented in one or more computer program products (i.e., one or more modules of computer program instructions) encoded on a computer-readable medium for execution by the computer system 900 or for controlling the operation of the computer system 900, and according to any method known to those skilled in the art, including but not limited to computer languages such as data-oriented languages (e.g., SQL, dBase), system languages (e.g., C, Objective-C, C++, Assembly), structured languages (e.g., Java, .NET), and application languages (e.g., PHP, Ruby, Perl, Python). The instructions may also be implemented in computer languages such as array languages, aspect-oriented languages, assembly languages, authoring languages, command-line interface languages, compiled languages, concurrent languages, curly-bracket languages, dataflow languages, data structured languages, declarative languages, esoteric languages, extension languages, fourth generation languages, functional languages, interactive pattern languages, interpreted languages, iterative languages, list-based languages, little languages, logic-based languages, machine languages, macro languages, metaprogramming languages, multiparadigm languages, numerical analysis, non-English-based languages, class-based object-oriented languages, prototype-based object-oriented languages, off-side rule languages, procedural languages, reflective languages, rule-based languages, scripting languages, stack-based languages, synchronous languages, syntax handling languages, visual languages, Wirth languages, and XML-based languages. Memory 904 may also be used for storing temporary variables or other intermediate information during execution of instructions to be executed by processor 902 .

[0136] As discussed herein, computer programs do not necessarily correspond to files in a file system. A program may be stored in a portion of a file (e.g., one or more scripts stored in a markup language document) that holds other programs or data, in a single file dedicated to the program in question, or in multiple collaborative files (e.g., files storing one or more modules, subroutines, or partial codes). A computer program may be deployed to execute on one or more computers that are located at a site or distributed across multiple sites and interconnected through a communication network. The processes and logic flows described in this specification may be performed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data and generating output.

[0137] The computer system 900 also includes a data storage device 906, such as a magnetic disk or optical disk, coupled to the bus 908 for storing information and instructions. The computer system 900 can be coupled to various devices via an input / output module 910. The input / output module 910 can be any input / output module. Exemplary input / output modules 910 include data ports such as universal serial bus (USB) ports. The input / output module 910 is configured to connect to a communication module 912. Exemplary communication modules 912 include network interface cards, such as Ethernet cards and modems. In certain aspects, the input / output module 910 is configured to connect to multiple devices, such as input devices 914 and / or output devices 916. Exemplary input devices 914 include keyboards and pointing devices, such as mice or trackballs, through which a user can provide input to the computer system 900. Other types of input devices 914 can also be used to provide interaction with the user, such as tactile input devices, visual input devices, audio input devices, or brain-computer interface devices. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and any form of input from the user can be received, including acoustic input, voice input, tactile input, or brainwave input. Exemplary output device 916 includes a display device for displaying information to the user, such as a liquid crystal display monitor.

[0138] According to one aspect of the present disclosure, a system for whole-body motion tracking (e.g., the system 200 described above) can be implemented using the computer system 900 in response to the processor 902 executing one or more sequences of one or more instructions contained in the memory 904. Such instructions can be read into the memory 904 from another machine-readable medium (e.g., a data storage device 906). Execution of the sequence of instructions contained in the main memory 904 causes the processor 902 to perform the process steps described herein. One or more processors in a multi-processing arrangement can also be employed to execute the sequence of instructions contained in the memory 904. In alternative aspects, hard-wired circuitry can be used in place of software instructions, or hard-wired circuitry can be used in combination with software instructions to implement various aspects of the present disclosure. Accordingly, aspects of the present disclosure are not limited to any specific combination of hardware circuitry and software.

[0139] Aspects of the subject matter described in this specification can be implemented in a computing system that includes a back-end component (e.g., a data server), or includes a middleware component (e.g., an application server), or includes a front-end component (e.g., a client computer with a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described in this specification); or aspects of the subject matter described in this specification can be implemented in any combination of one or more such back-end components, middleware components, or front-end components. The components of the system can be interconnected through any form or medium of digital data communication (e.g., a communication network). The communication network can include, for example, any one or more of the following: a LAN; a WAN, and the Internet. In addition, the communication network can include, but is not limited to, any one or more of the following network topologies, including, for example, a bus network, a star network, a ring network, a mesh network, a star bus network, or a tree or hierarchical network. The communication module can be, for example, a modem or an Ethernet card.

[0140] Computer system 900 can include client and server.Client and server are usually far away from each other, and usually interact through communication network.The relationship between client and server is to produce by means of computer program that runs on respective computer and has client-server relationship between each other.Computer system 900 can be, for example, but not limited to: desktop computer, laptop computer or tablet computer.Computer system 900 can also be embedded in another device, and this another device is, for example, but not limited to: mobile phone, personal digital assistant (PDA), mobile audio player, global positioning system (GPS) receiver, video game console, and / or TV set-top box.

[0141] As used herein, the term "machine-readable storage medium" or "computer-readable medium" refers to any medium or media that participates in providing instructions to processor 902 for execution. Such media can take many forms, including, but not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical or magnetic disks, such as data storage device 906. Volatile media include dynamic memory, such as memory 904. Transmission media include coaxial cables, copper wire, and optical fiber, including the wires that make up bus 908. Common forms of machine-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic medium, CD-ROMs, DVDs, any other optical medium, punch cards, paper tape, any other physical medium with a pattern of holes, RAM, PROMs, EPROMs, FLASH EPROMs, any other memory chip or cartridge, or any other medium that can be read by a computer. The machine-readable storage medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or a combination of one or more of them.

[0142] When the computing system 900 reads application data and provides an application, information may be read from the application data and stored in a memory device (e.g., memory 904). Additionally, data from a memory 904 server accessed via a network, bus 908, or data storage device 906 may be read and loaded into memory 904. Although data is described as being found in memory 904, it will be understood that the data need not be stored in memory 904 and may be stored in other memory (e.g., data storage device 906) accessible to the processor 902 or distributed across several media.

[0143] Although this specification contains many details, these details should not be interpreted as limitations on the scope of what may be claimed, but rather as descriptions of specific implementations of the subject matter. Certain features described in this specification in the context of different embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as working in certain combinations, and even initially claimed as such, in some cases, one or more features from a claimed combination may be removed from the combination, and a claimed combination may be directed to a sub-combination or a variation of a sub-combination.

[0144] Many of the features and applications described above can be implemented as software processes that are designated as a set of instructions recorded on a computer-readable storage medium (alternatively referred to as a computer-readable medium, a machine-readable medium, or a machine-readable storage medium). When these instructions are executed by one or more processing units (e.g., one or more processors, processor cores, or other processing units), the instructions cause the one or more processing units to perform the actions indicated in the instructions. Examples of computer-readable media include, but are not limited to, RAM, ROM, read-only compact discs (CD-ROM), recordable compact discs (CD-R), rewritable compact discs (CD-RW), read-only digital versatile discs (e.g., DVD-ROM, dual-layer DVD-ROM), various recordable / rewritable DVDs (e.g., DVD-RAM, DVD-RW, DVD+RW, etc.), flash memory (e.g., SD card, mini-SD card, micro SD card, etc.), magnetic and / or solid-state hard drives, ultra-density optical discs, any other optical or magnetic media, and floppy disks. In one or more embodiments, computer-readable media does not include carrier waves and electronic signals transmitted wirelessly or via a wired connection, or any other transient signals. For example, computer-readable media can be entirely limited to tangible physical objects that store information in a computer-readable form. In one or more embodiments, computer-readable media is a non-transitory computer-readable medium, a computer-readable storage medium, or a non-transitory computer-readable storage medium.

[0145] In one or more embodiments, a computer program product (also referred to as a program, software, software application, script, or code) may be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and the computer program product may be deployed in any form, including as a standalone program, or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple collaborative files (e.g., multiple files that store one or more modules, one or more subroutines, or portions of code). A computer program may be deployed to execute on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.

[0146] Although the above discussion primarily refers to microprocessors or multi-core processors executing software, one or more embodiments are implemented by one or more integrated circuits, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). In one or more embodiments, such integrated circuits execute instructions stored on the circuits themselves.

[0147] It will be appreciated by those skilled in the art that the various illustrative blocks, modules, elements, parts, methods and algorithms described herein can be implemented as electronic hardware, computer software or a combination thereof. In order to illustrate the interchangeability of hardware and software, various illustrative blocks, modules, elements, parts, methods and algorithms have been generally described in terms of their functionality. Whether this function is implemented as hardware or software depends on the specific application and the design constraints imposed on the entire system. Those skilled in the art can implement the described functions in different ways for each specific application. Various components and blocks can be arranged in different ways (e.g., arranged in different orders, or divided in different ways), all of which do not depart from the scope of this subject technology.

[0148] It should be understood that any particular order or the hierarchy of the blocks in the disclosed process are illustrations of example methods. Based on embodiment preference, it should be understood that the particular order or the hierarchy of the blocks in these processes can be rearranged, or not all of the blocks shown are executed. Any block in these blocks can be executed simultaneously. In one or more embodiments, multitasking and parallel processing may be advantageous. In addition, the separation of the various system components in the above-described embodiments should not be understood as requiring such separation in all embodiments, but it should be understood that the described program components and system can be integrated into a single software product or encapsulated in a plurality of software products usually.

[0149] For example, the subject technology has been described with respect to the various aspects described above. This disclosure is provided to enable anyone skilled in the art to practice the various aspects described herein. This disclosure provides various examples of the subject technology, and the subject technology is not limited to these examples. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects.

[0150] Unless otherwise specified, reference to an element in the singular is not intended to mean "one and only one," but rather "one or more." Unless otherwise specified, the term "some" refers to one or more. Masculine pronouns (e.g., his) include feminine and neuter pronouns (e.g., her and its), and vice versa. Headings and subheadings, if any, are used for convenience only and do not limit the disclosure.

[0151] To the extent that the terms "including," "having," etc. are used in this specification or the claims, such terms are intended to be open ended in a manner similar to how the term "comprising" is interpreted as "comprising" when used as a transitional word in a claim.

[0152] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments. In one aspect, the various alternative configurations and operations described herein may be considered to be at least equivalent.

[0153] As used herein, the phrase "at least one of" following a list of items, together with the terms "and" or "or" used to separate any of those items, modifies the list as a whole, rather than modifying each element of the list (i.e., each item). The phrase "at least one of" does not require selection of at least one item; rather, the phrase is meant to include at least one of any of the items, and / or at least one of any combination of the items, and / or at least one of each of the items. As an example, the phrase "at least one of A, B, and C" or "at least one of A, B, or C" each refers to: only A, only B, or only C; any combination of A, B, and C; and / or at least one of each of A, B, and C.

[0154] Phrases such as “aspect” do not imply that the aspect is essential to the subject technology or that the aspect applies to all configurations of the subject technology. Disclosure relating to an aspect may apply to all configurations, or to one or more configurations. An aspect may provide one or more examples. Phrases such as “an aspect” may refer to one or more aspects, and vice versa. Phrases such as “an embodiment” do not imply that the embodiment is essential to the subject technology or that the embodiment applies to all configurations of the subject technology. Disclosure relating to an embodiment may apply to all configurations, or to one or more configurations. An embodiment may provide one or more examples. Phrases such as “an embodiment” may refer to one or more embodiments, and vice versa. Phrases such as “configuration” do not imply that the configuration is essential to the subject technology or that the configuration applies to all configurations of the subject technology. Disclosure relating to a configuration may apply to all configurations, or to one or more configurations. A configuration may provide one or more examples. Phrases such as “a configuration” may refer to one or more configurations, and vice versa.

[0155] In one aspect, unless otherwise indicated, all measurements, values, levels, positions, amplitudes, dimensions, and other specifications set forth in this specification (including in the claims) are approximate and not exact. In one aspect, these measurements, values, levels, positions, amplitudes, dimensions, and other specifications are intended to have reasonable ranges that are consistent with the functions to which these measurements, values, levels, positions, amplitudes, dimensions, and other specifications relate and with customary practices in the art. It should be understood that some or all steps, operations, or processes may be performed automatically without user intervention.

[0156] Method claims may be provided to present elements of the various steps, operations or processes in a sample order, and are not meant to be limited to the specific order or hierarchy presented.

[0157] In one aspect, a method may be an operation, instruction, or function, and vice versa. In one aspect, a claim may be amended to include some or all of the words (e.g., instructions, operations, functions, or components), one or more words, one or more sentences, one or more phrases, one or more paragraphs, and / or one or more claims recited in one or more other claims.

[0158] All structural and functional equivalents of the elements of the various configurations described throughout this disclosure that are known or later become known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be encompassed by the subject technology. In addition, nothing disclosed herein is intended to be dedicated to the public, regardless of whether such disclosure is explicitly stated in the above description. No claim element shall be construed under 35 U.S.C. § 112, sixth paragraph, unless the element is expressly described using the phrase "means for..." or, in the case of a method claim, using the phrase "step for..."

[0159] The invention title, background technology and description of the drawings of the present disclosure are incorporated into this disclosure and are provided as illustrative examples of the present disclosure rather than as limiting descriptions. It is submitted with the understanding that they will not be used to limit the scope or meaning of the claims. In addition, in the detailed description, it can be seen that this specification provides illustrative examples and that various features are grouped together in various embodiments for the purpose of simplifying the present disclosure. This approach of the present disclosure should not be interpreted as reflecting the intention that the included subject matter requires more features than those expressly recited in any claim. Rather, as reflected in the claims, the inventive subject matter lies in less than all the features of a single disclosed configuration or operation. The claims are incorporated into the detailed description herein, with each claim independently and individually representing patentable subject matter.

[0160] The claims are not intended to be limited to the aspects described herein, but rather should be accorded the full scope consistent with the language of the claims and including all legal equivalents. Nevertheless, no claim is intended to encompass subject matter that fails to satisfy the requirements of 35 U.S.C. §§ 101, 102, or 103, nor should the claims be construed in such a manner.

[0161] Embodiments consistent with the present disclosure may be combined with any combination of features or aspects of the embodiments described herein.

Claims

1. A method for whole-body motion tracking, comprising: receiving tracking signals from a plurality of sensors associated with an upper body of a person; determining motion characteristics and joint characteristics based on the tracking signal; training a diffusion model, the diffusion model comprising a multi-layer perceptron (MLP) network; generating a plurality of inputs to the trained diffusion model, the plurality of inputs comprising the motion features and the joint features; as well as providing the plurality of inputs to the trained diffusion model to generate a plurality of outputs, The multiple outputs include a full-body posture sequence, and the full-body posture sequence includes an upper-body posture and a lower-body posture.

2. The method according to claim 1, further comprising: An intermediate feature is generated according to the motion feature and the joint feature, wherein the plurality of inputs to the diffusion model include the intermediate feature, and the whole body pose sequence is generated based on the intermediate feature.

3. The method according to claim 2, wherein: Generating the plurality of outputs includes generating the plurality of outputs from the MLP network based on the intermediate features.

4. The method according to claim 1, wherein The plurality of outputs includes a position of a lower body of the person, and the method further includes estimating the position of the lower body based on the full-body pose sequence.

5. The method according to claim 1, wherein The MLP network includes a plurality of blocks, and the method further includes providing a time step embedding to each of the plurality of blocks.

6. The method according to claim 5, wherein: The time-step embedding is provided to each of the plurality of blocks through a fully connected layer and a sigmoid linear unit activation layer.

7. The method according to claim 5, wherein: Each of the plurality of blocks includes a convolutional layer and a fully connected layer.

8. The method according to claim 7, wherein: Each of the plurality of blocks further includes a sigmoid linear unit activation layer and layer normalization.

9. The method according to claim 1, wherein The plurality of sensors are inertial measurement units (IMUs).

10. The method according to claim 1, wherein The multiple sensors include a first sensor installed in a first handheld device, a second sensor installed in a second handheld device, and a third sensor installed in a head-mounted device (HMD), and the tracking signal includes a first orientation and a first translation of the first handheld device, a second orientation and a second translation of the second handheld device, and a third orientation and a third translation of the head-mounted device.

11. A non-transitory computer-readable medium storing a program for whole-body motion tracking, wherein when the program is executed by a computer, the computer is configured to: receiving tracking signals from a plurality of sensors associated with an upper body of a person; determining motion characteristics and joint characteristics based on the tracking signal; training a diffusion model, the diffusion model comprising a multi-layer perceptron (MLP) network; generating a plurality of inputs to the trained diffusion model, the plurality of inputs comprising the motion features and the joint features; as well as providing the plurality of inputs to the trained diffusion model to generate a plurality of outputs, The multiple outputs include a full-body posture sequence, and the full-body posture sequence includes an upper-body posture and a lower-body posture.

12. The non-transitory computer-readable storage medium of claim 11, wherein: When executed by the computer, the program further configures the computer to: Generate an intermediate feature according to the motion feature and the joint feature, wherein the plurality of inputs to the diffusion model include the intermediate features, and the whole body pose sequence is generated based on the intermediate features, and The generating the plurality of outputs comprises generating the plurality of outputs from the MLP network based on the intermediate features.

13. The non-transitory computer readable medium of claim 11, wherein: The plurality of outputs includes a position of the person's lower body, the MLP network includes a plurality of blocks, and the program, when executed by the computer, further configures the computer to: estimating a position of the lower body based on the whole body pose sequence; A time step embedding is provided to each of the plurality of blocks, wherein the time step embedding is provided to each of the plurality of blocks through a fully connected layer and a sigmoid linear unit activation layer.

14. The non-transitory computer readable medium of claim 13, wherein: Each of the plurality of blocks includes a convolutional layer, a fully connected layer, a sigmoid linear unit activation layer, and layer normalization.

15. The non-transitory computer readable medium of claim 11, wherein: The plurality of sensors are inertial measurement units (IMUs).

16. The non-transitory computer readable medium of claim 11, wherein: The multiple sensors include a first sensor installed in a first handheld device, a second sensor installed in a second handheld device, and a third sensor installed in a head-mounted device (HMD), and the tracking signal includes a first orientation and a first translation of the first handheld device, a second orientation and a second translation of the second handheld device, and a third orientation and a third translation of the head-mounted device.

17. A system for whole-body motion tracking, comprising: processor; as well as A non-transitory computer-readable medium storing a set of instructions that, when executed by the processor, configure the processor to: receiving tracking signals from a plurality of sensors associated with an upper body of a person; determining motion characteristics and joint characteristics based on the tracking signal; training a diffusion model, the diffusion model comprising a multi-layer perceptron (MLP) network; generating a plurality of inputs to the trained diffusion model, the plurality of inputs comprising the motion features and the joint features; as well as providing the plurality of inputs to the trained diffusion model to generate a plurality of outputs, The multiple outputs include a full-body posture sequence, and the full-body posture sequence includes an upper-body posture and a lower-body posture.

18. The system according to claim 17, wherein: When executed by the processor, the instructions further configure the processor to: Generate an intermediate feature according to the motion feature and the joint feature, wherein the plurality of inputs to the diffusion model include the intermediate features, and the whole body pose sequence is generated based on the intermediate features, and The generating the plurality of outputs comprises generating the plurality of outputs from the MLP network based on the intermediate features.

19. The system according to claim 17, wherein: The plurality of outputs includes a position of a lower body of the person, the MLP network includes a plurality of blocks, and the instructions, when executed by the processor, further configure the processor to: estimating a position of the lower body based on the whole body pose sequence; providing a time step embedding to each of the plurality of blocks, wherein the time step embedding is provided to each of the plurality of blocks through a fully connected layer and a sigmoid linear unit activation layer, Each of the multiple blocks includes a convolutional layer, a fully connected layer, a S-type linear unit activation layer and layer normalization.

20. The system of claim 17, wherein: The multiple sensors are inertial measurement units (IMUs), the multiple sensors include a first sensor installed in a first handheld device, a second sensor installed in a second handheld device, and a third sensor installed in a head-mounted device (HMD), and the tracking signal includes a first orientation and a first translation of the first handheld device, a second orientation and a second translation of the second handheld device, and a third orientation and a third translation of the head-mounted device.