Multi-modal driven human body action generation method based on large language model
By constructing a multimodal alignment fine-grained action data set and semantic-aware decoupled action discretization technology, the generalization ability and multimodal fusion problems in the existing technology are solved, and a high-precision and controllable three-dimensional human body action generation is achieved.
Patent Information
- Application Number
- CN202510674138.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-02
AI Technical Summary
The existing three-dimensional human body movement generation technology has weak generalization ability when facing novel descriptions or multimodal fusion inputs, the text-modal alignment is not fine, the semantic gap between modals is obvious, and the lack of multimodal joint driving and detail control capabilities, which limits the adaptability in complex scenarios and multimodal fusion tasks.
By constructing a fine-grained action dataset with multimodal alignment, semantic-aware decoupling action discretization technology is used to decouple the whole body movements according to the body parts and generate semantically aligned atomic action tokens through independent vector quantization encoders, construct mixed action sentences, and multimodal action generation based on large language models, supporting zero-sample generation and part-level control.
It significantly improves the semantic accuracy and controllability of action generation, enhances the generalization ability and multimodal adaptability of the model, and can independently infer and generate coordinated actions that meet multimodal conditions in complex scenarios.
Smart Images

Figure CN120580356A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a three-dimensional human motion generation technology, and in particular to a multi-modal driven human motion generation method based on a large language model. Background Art
[0002] With the rapid development of three-dimensional human motion generation technology, condition-driven human motion generation (such as motion generation based on text, music or voice) has broad application prospects in animation production, virtual reality, games, and human-computer interaction. In recent years, methods based on diffusion models and autoregressive language models have made significant progress. Existing methods such as "A text-driven digital human motion generation method" (CN202510137065.9) and "Controllable music-driven three-dimensional dance motion generation model training method, generation method and device" (CN202411486370.0) can better generate human motion sequences that conform to text or audio descriptions. However, such methods generally face the following key technical bottlenecks:
[0003] 1. Weak generalization ability: Most existing methods rely on large amounts of supervised data for training on specific tasks (such as text-to-action), which makes it difficult to generalize when faced with novel descriptions or multimodal fusion inputs and lacks zero-shot capabilities.
[0004] 2. Inaccurate text-action alignment: Mainstream methods often use coarse-grained natural language descriptions as input, while action modeling is based on the entire body. This granularity mismatch makes it difficult to achieve fine control of actions, and the generated actions may deviate from the textual meaning.
[0005] 3. Obvious semantic gap between modalities: Although some works have attempted to incorporate multiple modalities (text, speech, music, and action) into a unified modeling framework through a unified discrete token representation, this type of "token full unification" approach only achieves superficial unification and cannot truly eliminate the modal gap at the semantic level, limiting the model's ability to understand and reason about multimodal conditions.
[0006] 4. Lack of support for multi-modal joint drive: Although some methods support multi-modal input, they are still unable to achieve multi-modal fusion-driven action generation (such as text + voice joint control) in a single model, and lack controllability of action details (such as body part hierarchy).
[0007] These shortcomings limit the adaptability of existing technologies to complex scenes, multimodal fusion, and tasks requiring high detail accuracy in practical applications.
[0008] It should be noted that the information disclosed in the above background technology section is only used to understand the background of this application, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention
[0009] The main purpose of the present invention is to overcome the defects existing in the above-mentioned background technology and provide a multi-modal driven human motion generation method based on a large language model.
[0010] To achieve the above object, the present invention adopts the following technical solutions:
[0011] A multimodal driven human motion generation method based on a large language model comprises the following steps:
[0012] S1. Constructing a multimodally aligned fine-grained motion dataset: By extracting structured information from 3D human motion data and combining it with a large language model to generate atomic motion semantic descriptions at the body part level, we then integrate text, speech, music, and motion data to form a uniformly aligned MotionWords dataset.
[0013] S2, semantically aware decoupled action discretization: decouples full-body actions by body parts, performs residual quantization on each part's actions through independent vector quantization encoders, generates semantically aligned atomic action tokens, and introduces fine-grained text descriptions as conditions to guide token generation;
[0014] S3. Constructing mixed action sentences: Concatenate the fine-grained text description and the action tokens of each body part in a preset order to form a unified sequence input format with special start and end markers;
[0015] S4. Multimodal action generation and reasoning: A language-guided action generation model is built based on a large language model. It receives multimodal input and converts it into mixed action sentences. Through joint modeling, it generates fine-grained text descriptions and atomic action token sequences. After decoding, they are restored to semantically consistent 3D human actions, supporting zero-sample generation and part-level action control.
[0016] Another multimodal driven human motion generation method based on a large language model includes the following steps:
[0017] A1. Constructing a multimodal aligned motion dataset: By extracting structured information from 3D human motion data, we integrate text, speech, music, and motion data to form a uniformly aligned MotionWords dataset. The text descriptions are either manually written or coarse-grained descriptions from existing datasets.
[0018] A2. Semantic-aware decoupled action discretization: Decouple the whole-body action by body part, perform residual quantization on each part’s action using independent vector quantization encoders, and generate atomic action tokens.
[0019] A3. Constructing a mixed action sentence: Concatenate the coarse-grained text description and the action tokens of each body part in a preset order to form a unified sequence input format containing special start and end marks;
[0020] A4. Multimodal action generation and reasoning: Build an action generation model based on a large language model, receive multimodal input and convert it into mixed action sentences, generate atomic action token sequences through joint modeling, and restore them to 3D human actions after decoding, supporting generation control based on action tokens.
[0021] Another multimodal driven human motion generation method based on a large language model includes the following steps:
[0022] B1. Constructing a multimodal aligned motion dataset: By extracting structured information from 3D human motion data and combining it with coarse-grained text descriptions from traditional datasets, we generate semantic information about atomic motions at the body part level. Furthermore, we integrate text, speech, music, and motion data to form a uniformly aligned MotionWords dataset.
[0023] B2. Semantic-aware decoupled action discretization: Decouple the whole-body action by body part, perform residual quantization on each part’s action using independent vector quantization encoders, generate semantically aligned atomic action tokens, and use the coarse-grained text description as a condition to guide token generation.
[0024] B3. Constructing mixed action sentences: Concatenate the coarse-grained text description and the action tokens of each body part in a preset order to form a unified sequence input format with special start and end markers;
[0025] B4. Multimodal action generation and reasoning: Build a language-guided action generation model based on a large language model, receive multimodal input and convert it into mixed action sentences, generate coarse-grained text descriptions and atomic action token sequences through joint modeling, and decode them to restore them to semantically consistent 3D human actions, supporting part-decoupling-based action control.
[0026] The present invention has the following beneficial effects:
[0027] The present invention proposes a multimodal driven human motion generation method based on a large language model. The preferred embodiment significantly improves the semantic accuracy and controllability of motion generation by introducing a hybrid action sentence structure of fine-grained action text description and atomic-level body action tokens. Traditional methods rely on coarse-grained natural language descriptions, which makes it difficult to achieve fine alignment of actions and texts, resulting in the generated actions easily deviating from semantic intent. The present invention automatically generates atomic action semantic descriptions at the body part level through a large language model, and establishes a one-to-one mapping relationship with the decoupled action tokens, so that the language model can not only understand the overall action intent, but also independently control the local actions of specific parts. For example, the fine-grained description of atomic actions such as "right arm stretched upward" and the strong association with the corresponding token ensure that the generated action is highly consistent with the input semantics at the detail level, effectively solving the industry problem of imprecise text-action alignment.
[0028] In terms of action representation and modeling, the present invention proposes the semantic-aware decoupled motion discretization (SDMT) technology, which breaks through the limitations of traditional methods that treat whole-body actions as a whole. By decoupling the actions according to parts such as the torso and limbs, and using independent vector quantization encoders for residual quantization, the actions of each body part are encoded as atomic tokens with clear semantic orientation. This structured representation not only retains the local semantic features of the action, but also guides token generation by introducing text conditions, thereby achieving cross-modal semantic alignment. Compared with the existing "token full unification" method, this technology significantly enhances the expressive power of the action token space, enabling the model to generate clearer, stable and diverse action sequences, while having zero-sample generalization capabilities, and can adapt to the generation needs of new scenarios and new action categories.
[0029] In addition, the innovative hybrid action sentence format of the present invention provides a unified modeling framework for multimodal conditional fusion. By splicing fine-grained text descriptions and atomic action tokens into natural language sequences, the language model can process multimodal inputs such as text, speech, and music under a unified autoregressive prediction paradigm. This design not only simplifies the complexity of the multimodal fusion mechanism, but also realizes cross-modal semantic association modeling through the self-attention mechanism. For example, in the complex scenario of "talking and playing football", the model can autonomously infer the spatiotemporal relationship between speech content and body movements, and generate coordinated movements that meet multimodal conditions. This technology breaks through the limitations of single-modal drive of traditional methods, significantly improves the quality of action generation and multimodal adaptability in complex scenarios, and provides more powerful technical support for applications in virtual reality, intelligent interaction and other fields.
[0030] Other beneficial effects of the embodiments of the present invention will be further described below. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1This is a module overview diagram of a multi-modal driven human motion generation method according to an embodiment of the present invention.
[0032] Figure 2 This is an algorithm framework diagram of a multi-modal driven human motion generation method according to an embodiment of the present invention.
[0033] Figure 3 This is an overall flow chart of a multi-modal driven human motion generation method according to an embodiment of the present invention.
[0034] Figure 4 This is an overall flow chart of a multi-modal driven human motion generation method according to another embodiment of the present invention.
[0035] Figure 5 This is an overall flow chart of a multi-modal driven human motion generation method according to another embodiment of the present invention. DETAILED DESCRIPTION
[0036] The following is a detailed description of the embodiments of the present invention. It should be emphasized that the following description is only exemplary and is not intended to limit the scope of the present invention and its application.
[0037] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present invention, "plurality" means two or more, unless otherwise specifically defined.
[0038] In order to solve the problems existing in the prior art, the present invention proposes a multimodal-driven human motion generation method based on a large language model, which can process multiple modal inputs (text, voice, music) within a unified framework; support fine-grained body part-level motion modeling and control; establish a tighter and more accurate semantic alignment between modalities such as text and audio and motion; and at the same time have zero-sample generation capabilities, stronger generalization and reasoning capabilities. The preferred embodiment of the present invention aims to implement a unified motion generation method based on a large language model by introducing a hybrid motion sentence (Hybrid Motion Sentence) composed of fine-grained motion text descriptions and atomic-level body motion tokens, so as to solve the bottlenecks of the prior art in multimodal fusion, motion semantic alignment and fine control.
[0039] See Figures 1 to 3 The present invention provides a method for generating human motions driven by multimodality based on a large language model, comprising the following steps:
[0040] Step S1: Construct a multimodally aligned fine-grained motion dataset: By extracting the structured information of 3D human motion data and combining it with a large language model to generate atomic motion semantic descriptions at the body part level, and integrating text, speech, music, and motion data, a unified aligned MotionWords dataset is formed.
[0041] In some embodiments, step S1 specifically includes: extracting joint angles, relative positions, and motion trend features based on 3D human motion data; generating atomic motion semantic segments for each body part; using a large language model to fuse motion features with coarse-grained text to generate fine-grained natural language motion descriptions; and constructing a unified aligned dataset covering text, speech, music, and motion modalities.
[0042] Step S2: Semantic-aware decoupled action discretization: Decouple the whole-body action by body part, perform residual quantization on the action of each part through an independent vector quantization encoder, generate semantically aligned atomic action tokens, and introduce fine-grained text descriptions as conditions to guide token generation.
[0043] In some embodiments, step S2 specifically includes: decoupling and dividing the whole body movements into torso, limbs, head, and hands; configuring an independent vector quantization encoder for each body part, and using multi-layer residual quantization to generate discrete tokens; fusing fine-grained text descriptions of the corresponding parts during the token encoding process, and realizing semantic guidance through a diffusion transformation module; generating atomic action tokens that are mapped one-to-one with the text descriptions.
[0044] In some embodiments, step S2 further includes: in the decoding stage, fusing the text description with the action token through a diffusion transformation module to ensure that the generated token has clear semantic reference and location attribution.
[0045] Step S3: Construct a mixed action sentence: Concatenate the fine-grained text description and the action tokens of each body part in a preset order to form a unified sequence input format containing special start and end marks.
[0046] In some embodiments, the construction format of the mixed action sentence in step S3 is: a fine-grained text description sequence followed by action tokens for each body part, arranged in a preset body structure order, and separated by special start and end marks.
[0047] Step S4, multimodal action generation and reasoning: Build a language-guided action generation model based on the large language model, receive multimodal input and convert it into mixed action sentences, generate fine-grained text descriptions and atomic action token sequences through joint modeling, and decode and restore them to semantically consistent 3D human actions, supporting zero-sample generation and part-level action control.
[0048] In some embodiments, step S4 specifically includes: based on the large language model architecture, learning text generation action, action generation text and action continuation tasks through multi-task training; in the multimodal instruction tuning stage, integrating text, voice and music input combinations for joint modeling; during reasoning, generating fine-grained text descriptions and action token sequences through mixed action sentences, and decoding them into continuous 3D action data.
[0049] In a further preferred embodiment, the action generation and reasoning process of the large language model specifically includes: cross-modal semantic association modeling of input text features through a self-attention mechanism to generate a fine-grained text description sequence; in the action decoding stage, the diffusion transformation module is used to fuse the action token and text description, and the atomic action features of various parts of the body are parsed through a feedforward neural network; based on the hierarchical structure of a multi-layer residual vector quantization encoder, the continuous action features are gradually discretized into a semantically aligned atomic token sequence, and are jointly input into the language model with the text description for joint autoregressive prediction, thereby achieving end-to-end alignment of action generation and semantic control.
[0050] In some embodiments, the multimodal instruction tuning preferably includes: using low-rank adaptation technology (LoRA) to efficiently tune the parameters of the pre-trained language model to adapt to multimodal input and action generation tasks.
[0051] The multimodal-driven human action generation method described in the above embodiment utilizes multimodal joint input during the inference phase, generating fine-grained text descriptions to guide the generation of action token sequences, enabling independent control and editing of specific body parts. This method, through a unified sequence input format for mixed action sentences, enables a large language model to achieve autonomous multimodal reasoning and action generation within the natural language prediction paradigm, while maintaining zero-shot generalization capabilities.
[0052] See Figure 4 The present invention also provides another embodiment of a multi-modal driven human motion generation method based on a large language model, comprising the following steps:
[0053] Step A1: Construct a multimodal aligned motion dataset: By extracting the structured information of 3D human motion data, we integrate text, speech, music, and motion data to form a unified aligned MotionWords dataset, where the text descriptions are coarse-grained descriptions written manually or from existing datasets.
[0054] Step A2: Semantic-aware decoupling and action discretization: Decouple the whole-body action by body part, perform residual quantization on the action of each part through an independent vector quantization encoder, and generate atomic action tokens;
[0055] Step A3: Constructing a mixed action sentence: concatenate the coarse-grained text description and the action tokens of each body part in a preset order to form a unified sequence input format containing special start and end marks;
[0056] Step A4, multimodal action generation and reasoning: Build an action generation model based on the large language model, receive multimodal input and convert it into mixed action sentences, generate atomic action token sequences through joint modeling, and restore them to 3D human actions through decoding, supporting generation control based on action tokens.
[0057] The difference from the first embodiment is that this embodiment does not introduce fine-grained text descriptions as semantic guiding conditions in the action token generation process, and the text descriptions are manually written or coarse-grained descriptions from existing data sets. Although the semantic consistency and control accuracy of the generated actions are reduced, the basic advantages of action decomposition are still retained through body part decoupling and atomic tokenization structure, and it is suitable for action structure verification scenarios.
[0058] See Figure 5 The present invention also provides another embodiment of a multi-modal driven human motion generation method based on a large language model, comprising the following steps:
[0059] Step B1: Construct a multimodal aligned motion dataset: By extracting the structured information of 3D human motion data and combining it with the coarse-grained text descriptions of traditional datasets, we generate atomic motion semantic information at the body part level. We then integrate text, speech, music, and motion data to form a unified aligned MotionWords dataset.
[0060] Step B2, semantically aware decoupled action discretization: decouple the whole-body action by body part, perform residual quantization on the action of each part using an independent vector quantization encoder, generate semantically aligned atomic action tokens, and use the coarse-grained text description as a condition to guide token generation;
[0061] Step B3: Constructing mixed action sentences: Concatenate the coarse-grained text description and the action tokens of each body part in a preset order to form a unified sequence input format with special start and end markers;
[0062] Step B4, multimodal action generation and reasoning: Build a language-guided action generation model based on the large language model, receive multimodal input and convert it into mixed action sentences, generate coarse-grained text descriptions and atomic action token sequences through joint modeling, and decode them to restore them to semantically consistent 3D human actions, supporting part-decoupling-based action control.
[0063] The difference from the first embodiment is that this embodiment uses coarse-grained text descriptions of traditional data sets to replace the fine-grained text generation module, and performs weak semantic alignment based on coarse-grained text during the action token generation process. Although the ability to control local details of the generated action is limited, the basic semantic guidance function is still retained through the unified format of mixed action sentences, which is suitable for lightweight deployment scenarios with limited annotation resources.
[0064] Specific embodiments of the present invention are further described below.
[0065] In one embodiment, a unified human action generation method based on mixed action sentences effectively bridges the modality gap between text and action by introducing a combination of fine-grained text descriptions and atomic-level body action tokens, and supports multimodal human action generation. The core structure and workflow of this invention are described in detail below.
[0066] The overall system consists of three core modules, such as Figure 1 As shown, they are:
[0067] (1) MotionWords dataset construction module
[0068] This module is used to build large-scale multimodal alignment data resources. Specifically includes:
[0069] Automatically extract structural information such as joint angles and relative positions from 3D human motion data;
[0070] Design rules to generate atomic action semantic fragments;
[0071] Using large language models (such as GPT-4o), coarse-grained text and action analysis results are integrated to generate high-quality fine-grained action descriptions;
[0072] Build a unified aligned dataset covering four modalities: text, speech, music, and action.
[0073] The "fine-grained action description + action data" pair output by this module serves as the basis for subsequent training.
[0074] (2) Semantic-aware Decoupled Motion Tokenization
[0075] This module is responsible for encoding action sequences into semantically aligned discrete tokens and constructing a "mixed action sentence" in a unified format:
[0076] First, decouple the whole-body movement into its parts (trunk, limbs, head, hands);
[0077] Each part is quantized through 3 layers of residuals through an independent VQ-VAE encoder to generate atomic action tokens;
[0078] Introducing text descriptions as conditions in the tokenization process to achieve semantic guidance;
[0079] Finally, the text description and action token are spliced together to generate a structured "mixed action sentence".
[0080] This module realizes the atomic-level alignment of text and action and is the key bridge between model understanding and generation.
[0081] (3) Language Informed Motion Model
[0082] This module builds a unified multimodal action generation system MotionUPG based on a large language model (such as LLaMA). Its functions include:
[0083] receiving input text, voice, music, or a combination thereof;
[0084] Convert the input modality into a token sequence and jointly model it with the mixed action sentence;
[0085] Support multi-task training (such as text generation action, action generation text, action continuation, etc.);
[0086] The inference stage is able to generate high-quality, semantically consistent 3D action sequences under zero-shot conditions.
[0087] This module realizes the understanding and generation of human motion driven by multimodal conditions, and has strong generalization ability and controllability.
[0088] Figure 2 The workflow of each module is shown and is described in detail below.
[0089] Fine-grained action text description generation
[0090] To address the problem that action text in existing datasets is too coarse, this paper designs a process to automatically extract atomic-level semantic descriptions from the original 3D action. Specifically, it includes:
[0091] 1. Extract local motion features using human joint angles, relative positions, and motion trends;
[0092] 2. Generate atomic action statements for body parts (such as left arm, right leg, head, etc.), such as "raise left arm more than 90 degrees";
[0093] 3. Input the above structured information and action images into a large visual language model (such as GPT-4o) to generate fine-grained action description text in natural language form;
[0094] 4. Each description corresponds to the action of a body part, with clear action-language alignment.
[0095] Semantic-aware disentangled motion discretization (SDMT)
[0096] To achieve more refined action representation and semantic alignment, this paper proposes a decoupled body part action tokenization structure and introduces text semantic guidance:
[0097] 1. Divide the whole body movements into 8 categories: trunk, limbs, head, and hands;
[0098] 2. Each part uses an independent VQ-VAE encoder and introduces three layers of residual vector quantization (RVQ) to obtain discrete tokens;
[0099] 3. In the token decoding stage, the corresponding text description is fused through the Diffusion Transformer module (DiT), so that the generated token has a clear semantic reference;
[0100] 4. This method ensures that each token corresponds to a semantic action unit of a certain part of the body, namely the "atomic action token".
[0101] Mixed Action Sentence Construction
[0102] Based on the output of the previous two steps, a mixed action sentence with a unified input format is constructed for training the large language model. The format is as follows:
[0103] [Fine-grained text description] <som>[Torso token][Left leg token]...[Right hand token] <eom>
[0104] in:
[0105] <som>and <eom>Mark the start and end of the action token respectively;
[0106] All tokens are arranged in order of body structure;
[0107] This sentence structure aligns with the "next token prediction" mechanism of large language models, transforming action generation into a language modeling task.
[0108] MotionUPG model training and inference
[0109] Based on mixed action sentences, the present invention constructs a unified multimodal action generation model MotionUPG. The training and inference process is as follows:
[0110] 1. Training phase:
[0111] Pre-training tasks include: text-to-action (FT2M), action-to-text (M2FT), and action-to-continuation (M2M).
[0112] Use MotionWords data to optimize multimodal commands, including different input combinations such as text, voice, and music;
[0113] The model is based on the LLaMA architecture and combined with LoRA tuning to achieve efficient multi-task learning.
[0114] 2. Reasoning stage:
[0115] Can accept multiple modal inputs (single or combined);
[0116] First, generate a fine-grained text description, and then generate an action token based on it;
[0117] Finally, SDMT decoding restores high-quality 3D motion;
[0118] Supports zero-sample generation and controllable part-level editing.
[0119] In summary, the present invention proposes a multimodal human motion generation method that constructs fine-grained, aligned mixed motion sentences to guide the language model to achieve accurate motion generation and semantic understanding. Compared with existing technologies, the present invention has the following significant technical effects and advantages:
[0120] (1) Fine-grained semantic modeling significantly improves the accuracy and controllability of action generation
[0121] Existing methods generally rely on coarse-grained textual descriptions of actions, which fail to establish a precise mapping between actions and semantics. This results in ambiguous or imprecise control of generated actions. In contrast, this paper introduces automatically constructed fine-grained textual descriptions of actions. Based on atomic action units, this paper constructs natural language text with clear semantic boundaries based on the local movements of multiple body parts, and establishes a one-to-one mapping relationship with the corresponding action tokens.
[0122] In this way, the language model can not only "understand" the overall action intention, but also perform fine-grained control and editing of the movement status of specific body parts, greatly improving the consistency between the generated action and the input semantics.
[0123] (2) Decoupled semantic-aware action encoding enhances generalization and action quality
[0124] Traditional action tokenization methods often treat the entire action as a unified whole, ignoring the layered structure and local semantic differences of the body. This results in low-quality generated actions and difficulty generalizing. The semantically-aware decoupled motion discretization (SDMT) proposed in this paper encodes actions by body part and introduces textual semantics as a condition to guide token generation. This ensures that each token has clear semantic meaning and location attribution, enabling the expression of atomic actions such as "right arm extended upward."
[0125] This design effectively expands the expressive power of the action token space while avoiding the problem of excessive coupling between tokens, significantly improving the clarity, stability, and diversity of generated actions, and supporting zero-shot generalization capabilities in new scenarios and new action categories.
[0126] (3) Unified mixed action sentence format realizes the integration of multimodal conditions and improves reasoning ability
[0127] Existing methods often train on a single modality (such as text or music), making it difficult to integrate multimodal conditions to jointly drive action generation. This invention constructs a unified hybrid action sentence structure, incorporating language descriptions and atomic action tokens into the same input sequence, enabling language models to process multimodal input within the natural language prediction paradigm.
[0128] This sentence structure is perfectly aligned with the training methods of large language models, significantly simplifying the design cost of multimodal fusion mechanisms and enabling the model to autonomously infer and generate actions in complex scenarios. For example, given the input of "a person imitating kicking a ball while speaking," the model can generate continuous, appropriate actions that match the voice intonation and body language.
[0129] In general, the embodiments of the present invention solve the problems of existing technologies such as inaccurate alignment of actions and semantics, insufficient modal fusion capabilities, and weak generalization performance through three key innovations: fine-grained semantic enhancement, structured token representation, and unified sequence modeling. It has achieved significant technological progress in the accuracy, controllability, diversity, and multimodal adaptability of action generation.
[0130] The present invention proposes a multimodal human motion generation method, with preferred embodiments featuring fine-grained semantic alignment and strong generalization performance. Key technical points and innovative contributions include:
[0131] A method for automatically constructing fine-grained action text descriptions, including extracting atomic action semantics based on skeletal kinematic information, and using a large language model (such as GPT-4o) to generate natural language text descriptions with body part-level semantics.
[0132] A semantic-aware decoupled motion tokenization method (SDMT) decouples human motion by body part and introduces textual semantics as a guiding condition in the motion token encoding process, so that the generated discrete tokens have clear atomic motion semantics.
[0133] A method for constructing "hybrid action sentences" combines fine-grained text descriptions with body part action tokens into a unified input format. Special tokens are used to mark the start and end positions, making them compatible with the token sequence processing mechanism of large language models.
[0134] A language-informed motion model based on mixed action sentences. This model uses mixed sentences as training input, supports joint modeling of multimodal conditions (text, speech, and music), and has the capabilities of action generation, action continuation, and action-text conversion.
[0135] A multimodal action generation method that supports zero-shot reasoning and part-level action control, through structured input and semantic guidance mechanism, achieves generalization generation of unseen texts, scenes or modal combinations.
[0136] A method for constructing the MotionWords dataset automatically generates cross-modal aligned data, including action sequences, atomic-level text descriptions, and audio information, to support multimodal training and unified representation learning.
[0137] The following alternatives provide variations from different perspectives. Although these alternatives differ from the optimal solution in terms of effect, their core principles—decoupling human actions into atomic tokens and semantically aligning them through natural language—remain consistent.
[0138] Alternative Example 1: Action Tokenization without Semantic Guidance
[0139] In this implementation, actions are still decoupled by body part and tokenized using a separate VQ-VAE model. However, fine-grained text descriptions are not introduced as semantic guidance during the tokenization process. The text is either handwritten or directly used from existing datasets, without any structure. This solution already possesses the structural advantage of decoupled tokens.
[0140] Alternative Example 2: The text uses the original coarse-grained description, and the rest remains the same
[0141] In this embodiment, the fine-grained text description generation module is omitted, and the coarse-grained text description provided in the traditional dataset is directly used as the language input content, while the action still adopts the decoupled token method proposed in the present invention, and constructs a mixed action sentence format. This solution retains the semantic advantages of decoupled tokens to a certain extent, and the structure is still aligned with the training paradigm of the language model. It has the following characteristics: compared with the traditional "text → whole action" solution, it has better action generation stability; due to the low semantic granularity of the text, the semantic alignment accuracy decreases during the token generation process; it is difficult to obtain corresponding mappings at the part level during model training, and the control accuracy of the generated action is weaker than the optimal solution. This solution can be used as a feasible variation in lightweight deployment or scenarios without high-quality text annotation resources, and has certain practical value.
[0142] An embodiment of the present invention further provides a storage medium for storing a computer program, which at least performs the above method when executed.
[0143] An embodiment of the present invention further provides a control device, comprising a processor and a storage medium for storing a computer program; wherein the processor is configured to execute at least the method described above when executing the computer program.
[0144] An embodiment of the present invention further provides a processor, which executes a computer program and at least performs the method described above.
[0145] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface memory, an optical disc or a read-only optical disc (CD-ROM); the magnetic surface memory can be a magnetic disk memory or a magnetic tape memory. The storage medium described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.
[0146] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0147] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0148] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0149] Those skilled in the art will understand that all or part of the steps of the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc. Various media that can store program codes.
[0150] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.
[0151] The methods disclosed in the several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.
[0152] The features disclosed in several product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.
[0153] The features disclosed in several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0154] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. Those skilled in the art will recognize that, without departing from the scope of the present invention, several equivalent substitutions or obvious variations can be made, and the performance or use of the same should be considered to fall within the scope of protection of the present invention.< / eom> < / som> < / eom> < / som>
Claims
1. A multimodal driven human motion generation method based on a large language model, characterized in that: The following steps are involved: S1. Constructing a multimodally aligned fine-grained motion dataset: By extracting structured information from 3D human motion data and combining it with a large language model to generate atomic motion semantic descriptions at the body part level, we then integrate text, speech, music, and motion data to form a uniformly aligned MotionWords dataset. S2, semantically aware decoupled action discretization: decouples full-body actions by body parts, performs residual quantization on each part's actions through independent vector quantization encoders, generates semantically aligned atomic action tokens, and introduces fine-grained text descriptions as conditions to guide token generation; S3. Constructing mixed action sentences: Concatenate the fine-grained text description and the action tokens of each body part in a preset order to form a unified sequence input format with special start and end markers; S4. Multimodal action generation and reasoning: A language-guided action generation model is built based on a large language model. It receives multimodal input and converts it into mixed action sentences. Through joint modeling, it generates fine-grained text descriptions and atomic action token sequences. After decoding, they are restored to semantically consistent 3D human actions, supporting zero-sample generation and part-level action control.
2. The multimodal driven human motion generation method based on a large language model according to claim 1, characterized in that: Step S1 specifically includes: Extract joint angles, relative positions and motion trend features based on 3D human motion data; Generate atomic action semantic fragments for each body part; Use a large language model to fuse action features with coarse-grained text to generate fine-grained natural language action descriptions; Build a unified aligned dataset covering text, speech, music, and motion modalities.
3. The multimodal driven human motion generation method based on a large language model according to claim 1 or 2, characterized in that: Step S2 specifically includes: Decouple and divide the whole body movements into trunk, limbs, head and hands; Configure an independent vector quantization encoder for each body part and use multi-layer residual quantization to generate discrete tokens; In the token encoding process, fine-grained text descriptions of corresponding parts are integrated, and semantic guidance is achieved through the diffusion transformation module; Generate atomic action tokens that are mapped one-to-one with text descriptions.
4. The multimodal driven human motion generation method based on a large language model according to claim 3, characterized in that: Step S2 also includes: In the decoding stage, the text description and action token are fused through the diffusion transformation module to ensure that the generated token has clear semantic reference and location attribution.
5. The multimodal driven human motion generation method based on a large language model according to any one of claims 1 to 4, characterized in that: The construction format of the mixed action sentence in step S3 is: The fine-grained text description sequence is followed by the action tokens of each body part, arranged in the preset body structure order, and special start and end markers are used to separate the text and action token parts.
6. The multimodal driven human motion generation method based on a large language model according to any one of claims 1 to 5, characterized in that: Step S4 specifically includes: Based on the large language model architecture, we learn text generation actions, action generation text, and action continuation tasks through multi-task training; In the multimodal instruction tuning phase, text, voice, and music inputs are combined for joint modeling; During inference, fine-grained text descriptions and action token sequences are generated by mixing action sentences and decoded into continuous 3D action data.
7. The multimodal driven human motion generation method based on a large language model according to claim 6, characterized in that: The action generation and reasoning process of the large language model specifically includes: The self-attention mechanism is used to model cross-modal semantic associations of input text features and generate fine-grained text description sequences. In the action decoding stage, the diffusion transformation module is used to fuse the action token and text description, and the atomic action features of each part of the body are analyzed through a feedforward neural network; Based on the hierarchical structure of a multi-layer residual vector quantization encoder, continuous action features are gradually discretized into semantically aligned atomic token sequences, and input together with text descriptions into a language model for joint autoregressive prediction, achieving end-to-end alignment of action generation and semantic control.
8. The method for generating human motions based on a multimodal drive and a large language model according to any one of claims 1 to 7, wherein: The multimodal instruction tuning further includes: using low-rank adaptation technology (LoRA) to efficiently tune the parameters of the pre-trained language model to adapt to multimodal input and action generation tasks.
9. A multimodal driven human motion generation method based on a large language model, characterized in that: The following steps are involved: A1. Constructing a multimodal aligned motion dataset: By extracting structured information from 3D human motion data, we integrate text, speech, music, and motion data to form a uniformly aligned MotionWords dataset. The text descriptions are either manually written or coarse-grained descriptions from existing datasets. A2. Semantic-aware decoupled action discretization: Decouple the whole-body action by body part, perform residual quantization on each part’s action using independent vector quantization encoders, and generate atomic action tokens. A3. Constructing a mixed action sentence: Concatenate the coarse-grained text description and the action tokens of each body part in a preset order to form a unified sequence input format containing special start and end marks; A4. Multimodal action generation and reasoning: Build an action generation model based on a large language model, receive multimodal input and convert it into mixed action sentences, generate atomic action token sequences through joint modeling, and restore them to 3D human actions after decoding, supporting generation control based on action tokens.
10. A multimodal driven human motion generation method based on a large language model, characterized in that: The following steps are involved: B1. Constructing a multimodal aligned motion dataset: By extracting structured information from 3D human motion data and combining it with coarse-grained text descriptions from traditional datasets, we generate semantic information about atomic motions at the body part level. Furthermore, we integrate text, speech, music, and motion data to form a uniformly aligned MotionWords dataset. B2. Semantic-aware decoupled action discretization: Decouple the whole-body action by body part, perform residual quantization on each part’s action using independent vector quantization encoders, generate semantically aligned atomic action tokens, and use the coarse-grained text description as a condition to guide token generation. B3. Constructing mixed action sentences: Concatenate the coarse-grained text description and the action tokens of each body part in a preset order to form a unified sequence input format with special start and end markers; B4. Multimodal action generation and reasoning: Build a language-guided action generation model based on a large language model, receive multimodal input and convert it into mixed action sentences, generate coarse-grained text descriptions and atomic action token sequences through joint modeling, and decode them to restore them to semantically consistent 3D human actions, supporting part-decoupling-based action control.
Citation Information
Patent Citations
Training method, generation method and device for controllable music-driven three-dimensional dance movement generation model
CN119516051B
A text-driven method for generating digital human motion
CN119579743B
Cited By
Vision and action combined characterization method and system for human body action generation
CN121505186A
Text-to-action generation method and system based on fine-grained representation of body parts
CN121544767A