A diffusion model driven multi-language human motion generation method

By combining multilingual text processing and customized networks with efficient samplers for joint training, the problems of multilingual adaptation, inference efficiency, and motion realism in diffusion models were solved, achieving efficient, natural, and semantically consistent multilingual human motion generation, thus expanding the scope of applications.

CN122289309APending Publication Date: 2026-06-26NANJING UNIV OF SCI & TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING UNIV OF SCI & TECH
Filing Date
2026-03-19
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing diffusion models suffer from insufficient adaptability to multilingual inputs, low inference efficiency, and defects in motion realism, especially the prominent problem of foot sliding, which limits the application scenarios and promotion of human motion generation technology.

Method used

Employing a multilingual text processing module, a customized CondUNet1D motion denoising network, and a second-order DPMSolver++ sampler, combined with a joint training strategy of EMA and CFG, this method eliminates foot slippage through time-step noise addition and gradient descent optimization, thereby efficiently generating multilingual human motion sequences that conform to the laws of human movement.

Benefits of technology

It achieves efficient generation of multilingual text input, significantly improves reasoning efficiency, eliminates foot slippage problems, ensures the naturalness of generated motion and semantic alignment, and broadens application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122289309A_ABST
    Figure CN122289309A_ABST
Patent Text Reader

Abstract

This invention discloses a diffusion model-driven method for generating multilingual human motion, comprising: constructing a sample library; performing multilingual text translation and feature encoding using the Ali Tongyi Qianwen model; constructing a customized CondUNet1D motion denoising network with residual linear multi-head cross-attention, Dropout layer, and Rearrange layer optimization; and training using a joint strategy of exponential moving average and classifier-free guidance. During the inference phase, starting with pure noise, iterative denoising is achieved using a variant of the second-order DPMSolver++ sampler combined with training-free acceleration techniques. Simultaneously, foot slippage is identified and corrected using a vGRFs model transfer model. After secondary optimization of the posture using the diffusion model, a slippage-free motion sequence is obtained. Finally, a video is generated by a remote server and transmitted back to the local terminal. This invention achieves multilingual-driven, efficient, and highly realistic human motion generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and artificial intelligence, and in particular to a method for generating human motion based on a diffusion model. Background Technology

[0002] Human motion generation technology is a core supporting technology in fields such as digital entertainment, virtual robotics, and human-computer interaction. Its core objective is to generate semantically consistent human motion sequences and videos based on natural language descriptions. Most current mainstream human motion generation methods are based on diffusion models. These methods achieve motion generation by simulating a "forward noise addition-backward noise reduction" process, offering advantages such as high generation quality and strong semantic alignment.

[0003] However, existing technologies still suffer from three major problems that hinder their practical application: First, insufficient language adaptability. Existing diffusion models are mostly trained based on English text instructions, resulting in extremely poor semantic understanding of multilingual inputs such as Chinese and Japanese. They cannot directly parse non-English text instructions, making it difficult to meet the needs of multilingual users. Second, low inference efficiency. Traditional diffusion models use conventional sampling strategies, requiring many iterations for denoising and consuming a lot of computation time. This makes it impossible to achieve real-time generation and difficult to adapt to scenarios with stringent requirements for generation speed, such as games and real-time interactions. Third, significant defects in motion realism. The generated motion sequences generally suffer from foot slippage, meaning that the feet do not slide naturally on the ground, which seriously undermines the naturalness and realism of the motion and fails to meet the accuracy requirements for motion details in practical applications.

[0004] To address the aforementioned issues, existing technologies have attempted to improve and optimize individual components, such as optimizing the text encoder to enhance language understanding or improving the sampler to increase inference speed. However, none of these approaches have formed a systematic solution and cannot simultaneously address the three core problems of multilingual adaptation, low efficiency, and foot slippage, thus limiting the application scenarios and scope of human motion generation technology. Summary of the Invention

[0005] Purpose of the invention: To address the aforementioned existing technologies, this invention proposes a diffusion model-driven multilingual human motion generation method, which supports multilingual text input, has high inference efficiency, and generates human motion sequences and videos without foot slippage.

[0006] Technical solution: A diffusion model-driven multilingual human motion generation method, comprising:

[0007] S1: Obtain a dataset of real human motion sequences and add noise step by step to build a noisy motion sample library;

[0008] S2: Complete multilingual text processing and time-step encoding, construct a customized CondUNet1D motion denoising network, and obtain a denoising model by joint training with exponential moving average (EMA) and classifier-free guided CFG.

[0009] S3: Based on the trained model, starting with pure noise, the SDE variant of the second-order DPMSolver++ sampler is combined with training-free acceleration techniques to iteratively denoise the model and simultaneously eliminate foot slippage to obtain the target human motion sequence.

[0010] S4: Input the motion sequence into the remote server to generate video and send it back to the local terminal.

[0011] Furthermore, in S1, the noise addition formula is: ;in, The cumulative noise figure is calculated over time steps. Increase gradually decrease; To follow a standard normal distribution Gaussian noise; For the first Noisy motion samples at time steps The range of values ​​is , This represents the maximum time step.

[0012] Furthermore, in S2, multilingual text processing involves translating non-English instructions into English using the Ali Tongyi Qianwen model, then obtaining text feature vectors through random masking and text encoding, while simultaneously padding the motion sequence to unify the dimensions.

[0013] Furthermore, in S2, a customized CondUNet1D motion denoising network is constructed based on the standard 1D Unet structure. It includes 4 downsampling stages and 4 upsampling stages, with 2 residual Conv1D blocks in each stage. Residual linear multi-head cross-attention is added after the residual Conv1D blocks. The residual Conv1D blocks embed a Dropout layer, and Rearrange layers are added before and after GroupNorm.

[0014] Furthermore, in S2, the EMA formula is: , Indicates the preceding The average value of the network parameters in each iteration ( = 0), Indicates the weighted value. For the network The weight parameters for the next iteration; the CFG formula is... , This represents the model output used to predict the original motion sequence after being guided by CFC. This indicates that the diffusion model under given conditions empty set The model output that predicts the original motion sequence is given. This indicates that the diffusion model under given conditions The model output for predicting the original motion sequence when the set is not empty. Provides CFG guidance scale parameters.

[0015] Furthermore, in S3, training-free acceleration techniques include embedded text caching, parallel CFG computation, and FP32 to FP16 low-precision inference.

[0016] Furthermore, in S3, foot slippage elimination is achieved by: pre-training the 23-joint vGRF model. Transferred to a 22-joint model The algorithm identifies foot sliding joints and frame ranges in motion sequences, constructs a multi-loss term weighted function, corrects it with gradient descent, and optimizes the pose twice using a diffusion model.

[0017] Furthermore, the weighted loss function is ,in For attitude loss, For foot contact loss, For trajectory loss, For vGRFs loss, The weighting coefficients for each loss term are used, and the foot position in the middle frame of the sliding frame is used as a fixed point during correction.

[0018] Furthermore, in S3, the weighted loss function, , , ;in, It is a combination of foot gliding joints. For a set of frame ranges, The position of joint j, It is a model The target anchor point is calculated. The key point is the result after the sliding cleanup. This represents the root bone position from frame 1 to frame H in the original motion sequence. This represents the root bone position from frame 0 to frame (H-1) in the original motion sequence.

[0019] Furthermore, S4 connects to a remote server via SSH to complete text transmission, model inference, and video rendering, generating video that is then sent back to the local terminal.

[0020] Beneficial effects: 1. Achieve multilingual input adaptation: Through the Alibaba Qwen3 model, a unified conversion of multilingual text commands to English is achieved, breaking the limitations of existing technologies on English input, supporting motion generation driven by multilingual text commands such as Chinese and Japanese, and broadening the application scenarios of the technology.

[0021] 2. Significantly improves inference efficiency: It integrates a second-order DPM Solver++ sampler and embedded text cache, parallel CFG computation, and low-precision inference training-free acceleration techniques. No additional model training is required, which greatly reduces the number of iteration denoising steps and reduces computing power consumption, enabling efficient generation of human motion sequences. It can be adapted to scenarios with high generation speed requirements, such as real-time interaction and games.

[0022] 3. Ensure the authenticity of motion generation: By using a foot slip elimination module based on the vGRFs model, foot slip problems are identified and eliminated from the entire motion sequence generation process. Combined with gradient descent optimization and secondary correction by the diffusion model, the generated motion sequence is ensured to conform to the laws of human movement and have no natural defects, thereby improving the practical value of motion generation.

[0023] 4. Improve the semantic alignment between generated motion and text: The customized CondUNet1D network achieves deep fusion of text and motion features through residual linear multi-head cross attention. Combined with the EMA+CFG joint training strategy, it effectively improves the semantic consistency between generated motion and text instructions, and avoids the generated motion deviating from the text description. Attached Figure Description

[0024] Figure 1 This is a schematic diagram of the overall framework of the diffusion model used in this invention;

[0025] Figure 2 This is a diagram of the CondUNet1D motion denoising network framework in this invention;

[0026] Figure 3 The evaluation results of the model of this invention and several existing models on the HumanML3D dataset are shown.

[0027] Figure 4 The results show the evaluation of the model of this invention and several existing models on the KIT-ML dataset. Detailed Implementation

[0028] The invention will now be further explained with reference to the accompanying drawings.

[0029] Example 1

[0030] This embodiment provides a method for generating human motion based on Chinese text commands, such as... Figure 1 As shown, the specific steps are as follows:

[0031] Step 1: Forward Noise Addition Process

[0032] Obtain a dataset of real human motion sequences, and analyze each real motion sequence. Noise is added step-by-step to construct a noisy motion sample library.

[0033] This embodiment collects a dataset of real human motion sequences containing everyday actions such as waving, walking, and running, totaling 1000 motion sequences. Each motion sequence contains 200 temporal frames of skeletal keypoints. For each real motion sequence... According to the formula Noise is added step-by-step, where The cumulative noise figure is calculated over time steps. Increase gradually decrease; To follow a standard normal distribution Gaussian noise; For the first Noisy motion samples at time steps The range of values ​​is , For the maximum time step, when hour, These are pure noise motion samples that perfectly follow a standard normal distribution.

[0034] In this embodiment, the time step The value ranges from 0 to 1000, resulting in 1000 noisy motion samples and their corresponding time step information, thus constructing a noisy motion sample library. All noisy motion samples and time step information corresponding to real motion sequences are used as input data for the reverse denoising training process.

[0035] Step 2: Reverse Denoising Training Process

[0036] A multilingual adaptive diffusion model training system is constructed, the core of which includes a multilingual text processing module, a customized CondUNet1D motion denoising network, and a joint training strategy, specifically divided into:

[0037] 1. Multilingual Text Processing: Construct a Chinese text instruction dataset containing 1,000 Chinese instructions such as "the person is waving", "the person is walking slowly", "the person is running fast". Input each Chinese instruction into the Alibaba Tongyi Qianwen (Qwen3) model and translate it into the corresponding English instruction, such as "the person is waving", "the person is walking slowly", "the person is running fast". Perform random masking (Rand Mask) processing (randomly masking 10% of the text characters) and text encoding (Text Encoder) processing on the translated English instructions, and obtain a 768-dimensional text feature vector through the BERT text encoder. At the same time, perform motion padding processing on the motion sequences, and uniformly pad all motion sequences to 200 frames to obtain motion padding features.

[0038] 2. Time Step Encoding: Input the time step in Step 1 into the time step encoder (Time Step Encoder) and convert it into a 256-dimensional time step feature vector, which is used to represent the noise level of the current denoising stage.

[0039] 3. Customized CondUNet1D Network Training: Based on the standard 1D Unet, construct a customized CondUNet1D motion denoising network. As shown in Figure 2 the figure, in the figure, Mish is the activation function; GroupNorm is used for normalization; Conv1D is used for feature extraction and transformation; Scale and Shift are a conditional modulation mechanism, and the time step passes the parameters into Scale and Shift through a multi-layer perceptron (MLP) and applies them to the motion features; the constructed customized CondUNet1D motion denoising network contains 4 downsampling stages and 4 upsampling stages; each down / upsampling stage is configured with 2 residual Conv1D blocks, and residual linear multi-head cross-attention (Linear Multi-Head Cross-Attention) is added after the residual Conv1D blocks to fuse text features and motion features; a Dropout layer is embedded in the residual Conv1D block to randomly deactivate neurons to prevent model overfitting; at the same time, rearrangement (Rearrange) operations are added before and after GroupNorm respectively. By rearranging the data before and after GroupNorm, it is avoided that Conv1D confuses the padded data and the non-padded data during the time dimension operation, and the accuracy of feature extraction is improved.

[0040] 4. Joint Training: The noisy motion samples A customized CondUNet1D motion denoising network is constructed by inputting time step features, text features, and motion imputation features. It is trained using a joint training strategy of exponential moving average (EMA) and classifier-free guided generation (CFG), with the loss function being motion sequence reconstruction loss. The training objective is to minimize the network's predicted output. Compared with real motion sequences The error is calculated until the network converges, resulting in the trained CondUNet1D motion denoising model. The EMA formula is: , Indicates the preceding The average value of the network parameters in each iteration ( = 0), Indicates the weighted value. For the network The weight parameters for the next iteration; the CFG formula is... , This represents the model output used to predict the original motion sequence after being guided by CFG. This indicates that the diffusion model under given conditions empty set The model output that predicts the original motion sequence is given. This indicates that the diffusion model under given conditions The model output for predicting the original motion sequence when the set is not empty. The CFG guides the scaling parameters. In this embodiment, the EMA weighting is used. CFG guidance scale parameters After 1000 training iterations, the trained CondUNet1D model is obtained.

[0041] Step 3: Reasoning and Generation Process

[0042] Based on the trained CondUNet1D motion denoising model, efficient generation of target human motion sequences from pure noise is achieved, and the foot slippage problem is eliminated. Specifically, it includes:

[0043] 1. Initial noise initialization: from a standard normal distribution Medium-sampled pure noise samples , as the initial input for reasoning generation.

[0044] 2. Efficient iterative denoising: Initial pure noise samples... and current time step The input high-efficiency sampling module uses an SDE variant of the second-order DPMSolver++ sampler for iterative denoising. In this embodiment, the input... and initial time step Iterative denoising was performed using an SDE variant of the second-order DPMSolver++ sampler (SDE DPm-Solver++ 2M Karras), while incorporating three training-free acceleration techniques: (1) embedding text cache, directly reusing the generated text embedding features during the forward propagation of each grid to avoid redundant calculations; (2) parallel CFG calculation, executing conditional denoising calculation and unconditional denoising calculation in parallel to improve calculation speed; (3) low-precision inference, converting FP32 floating-point calculation to FP16 low-precision floating-point calculation to reduce computational power consumption; the next time step was calculated through the sampling module. and denoised motion samples Repeatedly perform iterative noise reduction operations until... The initial motion sequence generation is completed.

[0045] In this embodiment, by simultaneously integrating embedded text caching, parallel CFG computation, and FP16 low-precision inference techniques, the number of iteration steps is reduced from 1000 steps in the traditional diffusion model to 20 steps, generating an initial motion sequence.

[0046] 3. Foot Slide Elimination: The preliminary motion sequence generated by iterative denoising is input into the FootskateCleanup module. Based on the number of skeletal joints in the target human motion dataset, a pre-trained 23-joint vGRFs (vertical ground reaction force) prediction model is used. Transfer to a 22-joint vGRFs model adapted to the target dataset. The optimization objective for migration is: ,in The key point of the sliding step movement, for Results of redirection to the 23-joint skeleton.

[0047] Using the transferred vGRFs model Identify the set of foot gliding joints in a motion sequence. and frame range set This includes the right ankle, right toes, left ankle, and left toes; a model incorporating pose loss is constructed. Foot contact loss Trajectory loss vGRFs loss Weighted loss function Among them, China These are the weighting coefficients for each loss term. , , ;in, The position of joint j, It is a model The calculated target anchor point, is the key point of the result after slide cleaning, is the root bone position from the 1st frame to the Hth frame in the original motion sequence, is the root bone position from the 0th frame to the (H - 1)th frame in the original motion sequence.

[0048] Taking the foot position at the middle frame of the slide frame as the fixed point, the foot slide problem is corrected through the gradient descent algorithm, and then the trained diffusion model is used to perform secondary optimization on the corrected unreasonable posture to obtain a target human motion sequence without foot slide.

[0049] In this embodiment, based on the 22 - joint vGRFs model to identify slides, when constructing the weighted loss function, , , , , the slide is corrected through gradient descent, and then the diffusion model is used for secondary optimization to obtain a non - slide motion sequence.

[0050] 4. Video rendering and output: The motion sequence is input into the rendering module of the remote server to generate a human motion video with a resolution of 1080P, which is transmitted back to the local terminal through SSH connection to complete the generation.

[0051] Embodiment 2

[0052] This embodiment provides a human motion generation method based on Japanese text instructions. The specific steps are basically the same as those in Embodiment 1, except that:

[0053] In the multi - language text processing step, the text instruction dataset is Japanese instructions, such as "人が手を振っている"、"人がゆっくり歩いている". The Japanese instructions are translated into English instructions through the Alibaba Tongyi Qianwen (Qwen3) model, and the subsequent process is the same as that in Embodiment 1, and finally a human motion video that conforms to the Japanese instruction description is generated.

[0054] Embodiment 3

[0055] The model of this invention is compared with several state-of-the-art models, including T2M, MDM, MLD, MotionDiffuse, T2M-GPT, MotionGTP, ReMoDiffuse, M2DM, and fg-T2M. Evaluations are performed on the HumanML3D and KIT-ML datasets, respectively. The evaluation metrics are four key aspects: 1) Motion Realism: Frechet InceptionDistance (FID), which uses feature vectors extracted by a pre-trained motion encoder to evaluate the similarity between generated motion sequences and real motion sequences. 2) Text Matching: R Precision calculates the average top-k accuracy of matching generated motions with text descriptions using a pre-trained contrastive model. 3) Generation Diversity: Diversity measures the average joint difference between generated sequences from all test texts. Multi-Modality is quantified as the diversity of motions generated from the same text. 4) Time Cost: Average Inference Time per Sentence (AITS) measures the inference efficiency of the diffusion model in seconds, considering a generation batch size of 1, and ignoring model or data loading time. The comparison results are as follows: Figure 3 , Figure 4 As shown, Figure 3 For quantitative results in HumanML3D, Figure 4 For quantitative results on the KIt-ML dataset, the right arrow indicates that the closer to the true value, the better, and red and blue indicate the best and second-best results, respectively.

[0056] This method achieves state-of-the-art results in FID and R-precision (top k) on the HumanML3D dataset, and also yields good results on the KIT-ML dataset: best R-precision (top k) and second-best FID. This demonstrates the ability of the diffusion model of this invention to generate high-quality motion aligned with text cues. On the other hand, while some methods excel in diversity and multimodality, anchoring these aspects with accuracy (R-precision) and precision (FID) is crucial to strengthening their persuasiveness. Otherwise, diversity or multimodality becomes meaningless if the generated motion is poor. Therefore, the model of this invention achieves state-of-the-art experimental results on both datasets and demonstrates robustness in model performance.

[0057] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A diffusion model-driven method for generating multilingual human motion, characterized in that, include: S1: Obtain a dataset of real human motion sequences and add noise step by step to build a noisy motion sample library; S2: Complete multilingual text processing and time-step encoding, construct a customized CondUNet1D motion denoising network, and obtain a denoising model by joint training with exponential moving average (EMA) and classifier-free guided CFG. S3: Based on the trained model, starting with pure noise, the SDE variant of the second-order DPMSolver++ sampler is combined with training-free acceleration techniques to iteratively denoise the model and simultaneously eliminate foot slippage to obtain the target human motion sequence. S4: Input the motion sequence into the remote server to generate video and send it back to the local terminal.

2. The method according to claim 1, characterized in that, In S1, the noise addition formula is: ;in, The cumulative noise figure is calculated over time steps. Increase gradually decrease; To follow a standard normal distribution Gaussian noise; For the first Noisy motion samples at time steps The range of values ​​is , This represents the maximum time step.

3. The method according to claim 1, characterized in that, In S2, multilingual text processing involves translating non-English instructions into English using the Ali Tongyi Qianwen model, then obtaining text feature vectors through random masking and text encoding, while simultaneously padding the motion sequence to unify the dimensions.

4. The method according to claim 1, characterized in that, In S2, a customized CondUNet1D motion denoising network is constructed based on the standard 1D Unet structure. It includes 4 downsampling stages and 4 upsampling stages. Each stage is equipped with 2 residual Conv1D blocks. Residual linear multi-head cross-attention is added after the residual Conv1D blocks. The residual Conv1D blocks embed a Dropout layer, and Rearrange layers are added before and after GroupNorm.

5. The method according to claim 1, characterized in that, In S2, the EMA formula is: , Indicates the preceding The average value of the network parameters in each iteration. Indicates the weighted value. For the network The weight parameters for the next iteration; the CFG formula is... , This represents the model output used to predict the original motion sequence after being guided by CFG. This indicates that the diffusion model under given conditions empty set The model output that predicts the original motion sequence is given. This indicates that the diffusion model under given conditions The model output for predicting the original motion sequence when the set is not empty. Provides CFG guidance scale parameters.

6. The method according to claim 1, characterized in that, In S3, training-free acceleration techniques include embedded text caching, parallel CFG computation, and FP32 to FP16 low-precision inference.

7. The method according to claim 1, characterized in that, In S3, foot slippage elimination is achieved by: modifying the 23-joint vGRFs pre-trained model. Transferred to a 22-joint model The algorithm identifies foot sliding joints and frame ranges in motion sequences, constructs a multi-loss term weighted function, corrects it with gradient descent, and optimizes the pose twice using a diffusion model.

8. The method according to claim 7, characterized in that, The weighted loss function is: ,in For attitude loss, For foot contact loss, For trajectory loss, For vGRFs loss, The weighting coefficients for each loss term are used, and the foot position in the middle frame of the sliding frame is used as a fixed point during correction.

9. The method according to claim 8, characterized in that, In S3, the weighted loss function, , , ;in, It is a combination of foot gliding joints. For a set of frame ranges, The position of joint j, It is a model The target anchor point is calculated. The key point is the result after the sliding cleanup. This represents the root bone position from frame 1 to frame H in the original motion sequence. This represents the root bone position from frame 0 to frame (H-1) in the original motion sequence.

10. The method according to claim 1, characterized in that, S4 connects to a remote server via SSH to complete text transmission, model inference, and video rendering, generating video that is then sent back to the local terminal.