Whole-body motion generation system and method based on autoregressive diffusion converter

Through the combination of autoregressive diffusion transformer and multimodal encoder, reference motion and progressive training strategies are introduced, which solves the problems of insufficient generated motion quality and inaccurate multimodal control in the prior art, and achieves long-term and consistent style and space-time dynamics of three-dimensional human motion generation.

CN120510256AActive Publication Date: 2025-08-19UESTC (SHENZHEN) ADVANCED RES INST
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510426914.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-08-19
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The existing whole-body movement generation technology is difficult to generate three-dimensional human movements that are consistent in style and space-time dynamics. The multimodal condition control is inaccurate, the quality of data integration is insufficient, and the scarcity of text labeling leads to poor generalization.

Method used

The autoregressive diffusion transformer is used to combine with a multimodal encoder to project multimodal features into a unified space, introduce reference motion as a condition, and train the model with a progressive weak to strong condition strategy, gradually from abstract semantics to dense space-time alignment to generate full-body movement.

Benefits of technology

It significantly improves the authenticity and controllability of generated movements, supports diverse motion generation tasks, and realizes long-term, style, space-time dynamics and three-dimensional human movements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510256A_ABST
    Figure CN120510256A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a whole body motion generation system and method based on an autoregressive diffusion transverter, the system comprises the autoregressive diffusion transverter, a unified human body motion module and a multi-modal encoder, the multi-modal encoder respectively extracts each modal feature in input data, and the modal features in the input data are extracted by the multi-modal encoder; each modal feature is projected to a unified multi-modal feature space; the unified human body motion module is used for splicing the projected multi-modal features to form a fixed prefix context and inputting the fixed prefix context into a diffusion converter; the diffusion converter generates and outputs whole-body movement. The problems that the generated motion quality is insufficient, scene generalization is limited, multi-modal condition control is inaccurate, and motion with time-space consistency is difficult to generate are solved. According to the method, the authenticity and controllability of motion generation are remarkably improved, and meanwhile diversified motion generation tasks are supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of motion generation, and in particular to a whole-body human motion generation system and method based on an autoregressive diffusion transformer. Background Art

[0002] Full-body human motion generation technology, with its highly realistic and controllable characteristics, has shown broad application potential in multiple fields. However, existing technologies in the field of full-body human motion generation have significant shortcomings. Most are single-task models that cannot accommodate diverse tasks. Some models that focus on multimodality can be divided into three types: the first is to train a separate model for each modal condition, which prevents interaction between different modalities; the second is to use independently trained branches to process different inputs, which is prone to modal conflicts between branches; and the third is to directly mix multiple modal conditions for training. However, due to differences in conditional constraints, it is difficult for the model to accurately learn the connection between conditions and motion. In terms of data, when integrating multiple data sources, quantity is emphasized over quality, and visual language annotation is not combined. This leads to scarce text annotations or serious hallucination problems, which in turn leads to poor generalization from text to motion. In addition, existing technologies are generally only able to generate short movements in a single shot, and it is difficult to generate long-term three-dimensional human motion with consistent style and spatiotemporal dynamics. Summary of the Invention

[0003] The technical problem to be solved by the embodiments of the present invention is to provide a whole-body human motion generation system and method based on an autoregressive diffusion transformer, so as to generate long-term three-dimensional human motion with consistent style and spatiotemporal dynamics.

[0004] In order to solve the above technical problems, the embodiment of the present invention proposes a whole-body human motion generation system based on an autoregressive diffusion transformer, including an autoregressive diffusion transformer, a unified human motion module and a multimodal encoder, wherein: The multimodal encoder extracts each modal feature from input data and projects each modal feature into a unified multimodal feature space; the input data includes text, global motion, speech, music, and reference motion; The unified human motion module concatenates the projected multimodal features with the noise motion features to form a fixed prefix context which is input into the diffusion transformer. The diffusion transformer generates and outputs whole-body human motion.

[0005] Accordingly, an embodiment of the present invention further provides a method for generating whole-body human motion based on an autoregressive diffusion transformer, comprising: Model construction step: constructing a whole-body human motion generation model based on an autoregressive diffusion transformer, wherein the model includes an autoregressive diffusion transformer, a unified human motion module, and a multimodal encoder; The model extracts features corresponding to each modality of text, global motion, speech, music, reference motion, and noise motion features through a multimodal encoder; The unified human motion module concatenates the projected multimodal features with the noise motion features to form a fixed prefix context that is input into the diffusion transformer. The diffusion transformer generates and outputs whole-body human motion; Model training steps: The model is trained using a progressive weak-to-strong conditioning strategy, gradually moving from abstract semantic constraints to dense spatiotemporal alignment constrained motion generation; Motion generation step: Gaussian noise sampled from the Gaussian distribution is input into the trained full-body human motion generation model to generate the corresponding full-body human motion.

[0006] The present invention addresses the issues of insufficient motion quality, limited scene generalization, imprecise control of multimodal conditions, and difficulty generating long-duration, spatiotemporally consistent motion. By introducing a reference motion as a special condition, combined with a diffusion transformer and a progressive training strategy, the present invention significantly improves the realism and controllability of generated motion, while supporting a variety of motion generation tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 1 is a flow chart of a whole-body motion generation system based on an autoregressive diffusion transformer according to an embodiment of the present invention. DETAILED DESCRIPTION

[0008] It should be noted that, unless there is a conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The present invention is further described in detail below with reference to the drawings and specific embodiments.

[0009] In the embodiments of the present invention, if there are directional indications (such as up, down, left, right, front, back, etc.), they are only used to explain the relative position relationship and movement status of the various components under a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0010] In addition, the terms "first," "second," and so on, used in this disclosure are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Therefore, features specified as "first" or "second" may explicitly or implicitly include at least one of these features.

[0011] Please refer to Figure 1 The whole-body motion generation system based on the autoregressive diffusion transformer of an embodiment of the present invention includes an autoregressive diffusion transformer, a unified human motion module and a multimodal encoder.

[0012] The proposed system framework uses Diffusion Transformers (DiTs) as its core model. By treating multimodal conditions as fixed prefix context sequences and fusing them through a self-attention mechanism, the proposed system gradually transitions from high-level semantic constraints to dense spatiotemporal alignment constraints for motion generation, employing a progressive weak-to-strong conditioning strategy. Furthermore, by introducing reference motion as a special condition, the proposed system significantly improves the quality of generated motion and ensures consistency between the reference and generated motions, thus supporting autoregressive interactive motion generation and easily extending to long-term motion generation.

[0013] The multimodal encoder extracts each modal feature from the input data respectively and projects each modal feature into a unified multimodal feature space; the input data includes text, global motion, speech, music, and reference motion.

[0014] To preserve rich cross-modal conditional information, the present invention uses modality-specific encoders (such as the pre-trained T5-XXL for text, the WAV encoder for speech, and Librosa for music) to extract features from each modality. The multimodal encoder includes a text encoder, a global motion encoder, a speech encoder, a music encoder, a whole-body noise motion encoder, and a reference motion encoder. The text encoder uses the T5 encoder. The reference motion encoder, the whole-body noise motion encoder, and the global motion encoder all use multiple linear layers, each encoding a body part (identified by index). The input format clearly identifies which index corresponds to which body part. All three encoders encode a single frame of the human body and do not perform temporal encoding. The music encoder uses Librosa to extract intensity (OnsetEnvelope), MFCC (Mel-Frequency Cepstral Coefficients), chroma features (Chroma CENS), onset, and tempo. The speech encoder uses learnable convolution to extract features. The reference motion encoder and the whole-body noise motion encoder share the same architecture, but differ in parameters.

[0015] The unified human motion module concatenates the projected multimodal features with the noise motion features to form a fixed prefix context, which is then input into the diffusion transformer. To project the encoded features of different modal conditions into a unified multimodal feature space with a common dimension, the present invention uses a learnable linear projection layer to align the features of each modality to the motion embedding dimension and project them into the unified multimodal feature space.

[0016] The fusion formula of the unified human motion module is: , Where f represents the modality-specific encoder, h represents the projection layer, and c is the concatenated multimodal conditional representation (with fixed arrangement).

[0017] To constrain the physical properties of motion, the present invention directly predicts motion Rather than noise. Therefore, the diffusion target of the present invention is defined as follows: ; in, represents the data distribution, and T is the maximum number of diffusion steps. G represents the learned denoising function, and x t Indicates steps t The noise motion is expressed as , where each The i-th pose in the motion corresponding to step t.

[0018] The diffusion transformer generates and outputs whole-body human motion.

[0019] The present invention generates corresponding spatiotemporal masks according to the task types of global spatiotemporal controllable tasks, and realizes multiple global spatiotemporal controllable tasks through a global motion condition and the spatiotemporal masks corresponding to the tasks.

[0020] The present invention divides different global spatiotemporal controllable tasks into: spatially dense tasks covering multiple joints and spatially sparse tasks focusing only on sparse joints; for spatially intensive tasks, the global motion conditions are decomposed into sparse global motion representation and dense relative representation (spatial sparsity does not require decomposition, spatially sparse tasks generally do not cause major problems, while spatially dense tasks cause greater problems, so the dense ones need to be decomposed into sparse ones).

[0021] The present invention realizes global spatiotemporal controllable generation through the following steps: 1. Global motion conditions and task-dependent mask strategies: The conditions for global spatiotemporal controllable generation tasks can all be derived from a global motion condition (using the x, y, and z coordinates of the 3D global space to accurately describe the spatial positions of human joints). This global motion condition is then used to design different spatiotemporal masks for each task (e.g., motion prediction, interpolation, completion, or joint / trajectory-guided synthesis). For motion prediction, the initial motion (e.g., the first 25% of frames) is visible, while the subsequent motion (e.g., the last 75%) is invisible. For motion interpolation, the initial motion (e.g., the first 25% of frames) and the subsequent motion (e.g., the last 25%) are visible, while the intermediate motion (e.g., the middle 50%) is invisible. For motion completion, some joints in some frames are invisible, while the majority of the rest is visible. For joint / trajectory-guided synthesis, only a few joints are visible in certain frames, while the majority of the rest is invisible. These tasks can be combined to achieve multi-task training by applying different types of spatiotemporal masks to the global motion condition.

[0022] 2. Global position gradient optimization and global motion representation decomposition: In order to ensure that the local motion representation output by the model (local motion representation: as the output of model building, it is a form of representation specifically designed to reduce the difficulty of the model learning motion features. It specifically covers the root (pelvic) joint representation, including the height, linear velocity, and angular velocity of the root joint; the remaining joints are represented as information such as position and rotation relative to the root joint. In this way, structured and hierarchical data is provided to the model, helping the model to learn motion features more efficiently) meets the constraints of the global motion conditions, in addition to directly inputting the global motion conditions, the local motion representation output by the model is generally integrated to obtain a global motion representation, and then the difference is calculated with the input global motion conditions. The output of the model is guided by gradient optimization during the diffusion model denoising process. However, this gradient optimization method is often prone to generating unnatural motion when multiple joints are involved.

[0023] To this end, the present invention divides these different global spatiotemporal controllable tasks into: spatially dense tasks (such as motion prediction, interpolation and completion) covering multiple joints, and spatially sparse tasks (such as joint / trajectory-guided synthesis) only focusing on sparse joints.

[0024] For spatially intensive tasks, this paper decomposes the global motion condition into a sparse global motion representation (retaining only the sparse global representation of joints) and a dense relative representation (converting the dense joints from the global representation to the local representation). This is achieved by optimizing the global position of the constrained sparse joints and injecting noise motion into the dense relative representation during inference. The injection takes the form of a splicing operation using the mean of the diffusion denoising model. The formulation of global spatiotemporal controllable generation is: ; ; Where c_g is the global motion condition, M_t is the task spatiotemporal mask, M_d is the global motion decomposition mask, S is the noise addition process, M' represents the form of mask conversion to local representation, G is the distance between the calculation condition and the output, which can be referred to the G function definition in the OmniControl paper; μ t It refers to the mean of the Gaussian distribution (related to the theory of diffusion), and τ is a hyperparameter that controls the strength / amplitude of the gradient.

[0025] As an implementation method, the present invention adopts a progressive training strategy to gradually introduce fine-grained conditions. The specific steps are as follows: Phase 1: Training using only textual conditions to establish alignment between motion and abstract semantics. At this time, the conditions only constrain motion at the abstract level.

[0026] Phase 2: Based on the text condition, a reference motion condition is introduced to strengthen the spatiotemporal constraints at the start of the generated motion. At this point, it is only necessary to ensure that the generated motion maintains a connection between the start and the end of the reference motion.

[0027] Phase 3: Next, global spatiotemporal conditions are introduced to enhance the spatiotemporal constraints for generating motion sparseness. At this point, constraints are only applied to certain frames and joints.

[0028] Phase 4: Finally, music and speech are introduced to enhance the dense spatiotemporal alignment of generated motion across all timeframes. This ensures that constraints exist across all timeframes, and all modal conditions can coexist. Training under full multimodal conditions ensures high-quality generated motion and precise control of multimodal conditions.

[0029] The present invention achieves the ability to arbitrarily combine conditions during reasoning by assigning probabilities to different conditions during the training process.

[0030] As an embodiment, the present invention uses data in a preset data set for training, and the preset data set is constructed according to the following steps: Collect motion capture data covering a variety of motion generation tasks; Unifying the data into SMPL-X format and transforming into a unified coordinate system; Unify to a preset fps (e.g. 30 fps); Rendering motion as video and generating hierarchical text annotations using visual language models. This method automatically generates hierarchical text annotations by rendering motion as video and combining it with contextual data such as the original text, category, and dataset description using visual language models (VLMs). This integrates visual and textual information to accurately and automatically generate hierarchical text annotations. Music or speech data is directly included in the pre-set dataset.

[0031] This invention is used for full-body human motion generation. The system framework utilizes an autoregressive diffusion transformer in a unified sequence-to-sequence approach to support a variety of multimodal motion generation tasks, including text-to-motion, music-to-dance, speech-to-gesture, human-in-scene interaction generation, object-to-human interaction generation, and human-to-human interaction generation, as well as global spatiotemporal control tasks (such as motion prediction, intermediate frame generation, motion completion, and joint / trajectory-guided synthesis). Flexible combinations of these tasks are also supported.

[0032] The present invention innovatively uses reference motion as a new conditional signal, significantly improving the consistency, style, and temporal dynamics of the generated content, which is extremely critical for generating realistic animations. To resolve the conflict between multimodal conditions of different granularities, the present invention introduces a progressive weak-to-strong conditional strategy, gradually adding conditions from text, to reference motion, to global motion conditions, to music and speech, to achieve a learning process from abstract level constraints to strict spatiotemporal alignment constraints, in order to achieve unified multi-task modeling. The global motion condition is further decomposed into a sparse global representation and a dense relative representation to maintain the naturalness of the motion.

[0033] The present invention can effectively integrate multiple motion capture data sources (motion capture data: refers to the data obtained by accurately measuring and recording human motion using advanced high-precision equipment, such as wearable sensors, inertial measurement units (IMUs), etc. These devices can capture the motion information of various parts of the human body in real time, including position, angle, speed, etc. The data collected by such equipment is usually of high quality and can accurately reflect the real motion state of the human body, providing a reliable data foundation for subsequent human motion analysis, modeling and generation), covering a variety of tasks, including text-to-motion generation, audio-to-dance generation, speech-to-gesture generation, global control signal-to-controllable motion generation, as well as the generation of human interaction in scenes, object-to-human interaction, and human-to-human interaction. To ensure detailed and consistent text annotation, the present invention renders the motion sequence into video and uses advanced natural language processing models to automatically generate structured and hierarchical descriptions to capture low-level actions and high-level semantics.

[0034] The method for generating whole-body motion based on an autoregressive diffusion transformer according to an embodiment of the present invention includes: Model construction step: constructing a whole-body human motion generation model based on an autoregressive diffusion transformer, wherein the model includes an autoregressive diffusion transformer, a unified human motion module, and a multimodal encoder; The model extracts features corresponding to each modality of text, global motion, speech, music, reference motion, and noise motion features through a multimodal encoder, and projects the features of each modality into a unified multimodal feature space; The unified human motion module concatenates the projected multimodal features with the noise motion features to form a fixed prefix context that is input into the diffusion transformer. The diffusion transformer generates and outputs whole-body human motion; Model training steps: The model is trained using a progressive weak-to-strong conditioning strategy, gradually moving from abstract semantic constraints to dense spatiotemporal alignment constrained motion generation; Motion generation step: Gaussian noise sampled from the Gaussian distribution is input into the trained full-body human motion generation model to generate the corresponding full-body human motion.

[0035] As an embodiment, the multimodal encoder includes a text encoder, a global motion encoder, a speech encoder, a music encoder, a whole body noise motion encoder, a reference motion encoder, etc.

[0036] As an implementation method, a progressive weak-to-strong condition strategy is adopted in the model training step: Phase 1: Training using only text conditions to establish alignment between motion and abstract semantics; The second stage: based on the text condition, the reference motion condition is introduced to strengthen the spatiotemporal constraints of the generated motion at the start; The third stage: Then, global spatiotemporal conditions are introduced to enhance the spatiotemporal constraints of motion sparsity. Phase 4: Finally, music and speech are introduced to enhance the dense spatiotemporal alignment of generated motion across all time frames.

[0037] As an implementation method, the model generates a corresponding spatiotemporal mask according to the task type of the global spatiotemporal controllable task, and realizes multiple global spatiotemporal controllable tasks through a global motion condition and the spatiotemporal mask corresponding to the task; Different global spatiotemporal controllable tasks are divided into spatially dense tasks covering multiple joints and spatially sparse tasks focusing only on sparse joints; For spatially intensive tasks, the global motion condition is decomposed into a sparse global motion representation and a dense relative representation; The formulation of global spatiotemporal controllable generation is: ; ; Among them, c g is the global motion condition, M t is the task spatiotemporal mask, M d is the global motion decomposition mask, S is the noise addition process, M' represents the form of mask conversion to local representation, G is the distance between the calculation condition and the output; μ t is the mean of the Gaussian distribution, and τ is a hyperparameter.

[0038] As an implementation method, the model training step also includes a training dataset construction step before the following steps: Collect motion capture data covering a variety of motion generation tasks; Unifying the data into SMPL-X format and transforming into a unified coordinate system; Unify to the preset fps; Render motion as video and generate hierarchical text annotations using a vision-language model.

[0039] The present invention supports multiple input conditions, such as text, music, and voice, and generates a corresponding 3D human motion process or model based on these constraints. This process integrates information from different modalities to achieve multi-source information-driven human motion synthesis, meeting the needs of human motion in various scenarios.

[0040] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A whole-body human motion generation system based on an autoregressive diffusion transformer, characterized in that: It includes an autoregressive diffusion transformer, a unified human motion module and a multimodal encoder, among which, The multimodal encoder extracts each modal feature from input data and projects each modal feature into a unified multimodal feature space; the input data includes text, global motion, speech, music, and reference motion; The unified human motion module concatenates the projected multimodal features with the noise motion features to form a fixed prefix context which is input into the diffusion transformer. The diffusion transformer generates and outputs whole-body human motion.

2. The whole body motion generation system based on the autoregressive diffusion transformer according to claim 1, characterized in that: The multimodal encoder includes a text encoder, a global motion encoder, a speech encoder, a music encoder, a whole body noise motion encoder, and a reference motion encoder.

3. The whole body motion generation system based on autoregressive diffusion transformer according to claim 1, characterized in that: Adopt a progressive weak to strong condition strategy: Phase 1: Training using only text conditions to establish alignment between motion and abstract semantics; The second stage: based on the text condition, the reference motion condition is introduced to strengthen the spatiotemporal constraints of the generated motion at the start; The third stage: Then, global spatiotemporal conditions are introduced to enhance the spatiotemporal constraints of motion sparsity. Phase 4: Finally, music and speech are introduced to enhance the dense spatiotemporal alignment of generated motion across all time frames.

4. The whole body motion generation system based on autoregressive diffusion transformer according to claim 3, characterized in that: The system generates a corresponding spatiotemporal mask according to the task type of the global spatiotemporal controllable task, and realizes multiple global spatiotemporal controllable tasks through a global motion condition and the spatiotemporal mask corresponding to the task; Different global spatiotemporal controllable tasks are divided into spatially dense tasks covering multiple joints and spatially sparse tasks focusing only on sparse joints; For spatially intensive tasks, the global motion condition is decomposed into a sparse global motion representation and a dense relative representation; The formulation of global spatiotemporal controllable generation is: ; ; Among them, c g is the global motion condition, M t is the task spatiotemporal mask, M d is the global motion decomposition mask, S is the noise addition process, M' represents the form of mask conversion to local representation, G is the distance between the calculation condition and the output; μ t is the mean of the Gaussian distribution, and τ is a hyperparameter.

5. The whole body motion generation system based on autoregressive diffusion transformer according to claim 1, characterized in that: The system is trained using data from a preset dataset, which is constructed according to the following steps: Collect motion capture data covering a variety of motion generation tasks; Unifying the data into SMPL-X format and transforming into a unified coordinate system; Unify to the preset fps; Render motion as video and generate hierarchical text annotations using a vision-language model.

6. A method for generating whole-body human motion based on an autoregressive diffusion transformer, characterized in that: include: Model construction step: constructing a whole-body human motion generation model based on an autoregressive diffusion transformer, wherein the model includes an autoregressive diffusion transformer, a unified human motion module, and a multimodal encoder; The model extracts features corresponding to each modality of text, global motion, speech, music, reference motion, and noise motion features through a multimodal encoder; The unified human motion module concatenates the projected multimodal features with the noise motion features to form a fixed prefix context that is input into the diffusion transformer. The diffusion transformer generates and outputs whole-body human motion; Model training steps: The model is trained using a progressive weak-to-strong conditioning strategy, gradually moving from abstract semantic constraints to dense spatiotemporal alignment constrained motion generation; Motion generation step: Gaussian noise sampled from the Gaussian distribution is input into the trained full-body human motion generation model to generate the corresponding full-body human motion.

7. The method for generating whole-body motion based on an autoregressive diffusion transformer according to claim 6, wherein: The multimodal encoder includes a text encoder, a global motion encoder, a speech encoder, a music encoder, a whole body noise motion encoder, and a reference motion encoder.

8. The method for generating whole-body motion based on an autoregressive diffusion transformer according to claim 6, wherein: In the model training step, a progressive weak-to-strong condition strategy is adopted: Phase 1: Training using only text conditions to establish alignment between motion and abstract semantics; The second stage: based on the text condition, the reference motion condition is introduced to strengthen the spatiotemporal constraints of the generated motion at the start; The third stage: Then, global spatiotemporal conditions are introduced to enhance the spatiotemporal constraints of motion sparsity. Phase 4: Finally, music and speech are introduced to enhance the dense spatiotemporal alignment of generated motion across all time frames.

9. The method for generating whole-body motion based on an autoregressive diffusion transformer according to claim 6, wherein: The model generates corresponding spatiotemporal masks according to the task type of the global spatiotemporal controllable task, and realizes multiple global spatiotemporal controllable tasks through a global motion condition and the spatiotemporal mask corresponding to the task; different global spatiotemporal controllable tasks are divided into: spatially dense tasks covering multiple joints and spatially sparse tasks focusing only on sparse joints; For spatially intensive tasks, the global motion condition is decomposed into a sparse global motion representation and a dense relative representation; The formulation of global spatiotemporal controllable generation is: ; ; Among them, c g is the global motion condition, M t is the task spatiotemporal mask, M d is the global motion decomposition mask, S is the noise addition process, M' represents the form of mask conversion to local representation, G is the distance between the calculation condition and the output; μ t is the mean of the Gaussian distribution, and τ is a hyperparameter.

10. The method for generating whole-body motion based on an autoregressive diffusion transformer according to claim 6, wherein: The model training step also includes the training dataset construction step: Collect motion capture data covering a variety of motion generation tasks; Unifying the data into SMPL-X format and transforming into a unified coordinate system; Unify to the preset fps; Render motion as video and generate hierarchical text annotations using a vision-language model.

Citation Information

Patent Citations

  • Virtual object action generation method and device, computer equipment and storage medium

    CN116977509A

  • Self-supervised skeleton behavior recognition method, system and equipment based on action semantic guidance and medium

    CN119131880A

  • Multimodal coding alignment method and device based on diffusion model

    CN119599027A

  • Human motion generation method and system

    US20240193797A1

  • Systems and methods for multi-modal language models

    US20240370718A1