Text-Driven Motion Generation With Two-Stage Diffusion Refinement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing text-driven motion synthesis methods using auto-encoders and variational auto-encoders generate human motion sequences with insufficient details due to limited expression capability in low-dimensional feature spaces.

Innovation Solution

A motion generation model comprising a text encoder, a first diffusion model operating in a low-dimensional feature space, and a second diffusion model in a higher-dimensional space to refine the intermediate motion sequence, enhancing detail processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a low-dimensional feature space is used for motion sequence generation, then the generation process is computationally efficient, but the expression capability and detail richness are insufficient

Engineering Contradiction:
Improvegeneration efficiencyVSAvoiddetail richness
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent transitions from a low-dimensional feature space to a high-dimensional feature space to enhance the expression capability and detail richness of generated motion sequences. By increasing the dimensionality of the feature space, the model can capture more nuanced motion details while maintaining computational efficiency through the diffusion model framework.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent employs a two-stage generation process: first generating an intermediate motion sequence in a compressed representation, then refining it in a high-dimensional feature space. This segmentation allows the system to benefit from both computational efficiency in the first stage and detail richness in the second stage, resolving the contradiction between efficiency and quality.

Inventive Principle:
Principle #1Segmentation

2Device complexity

If auto-encoders or variational auto-encoders are used for motion synthesis, then the model structure is simple, but the generated motion sequences lack sufficient details

Engineering Contradiction:
Improvemodel structure complexityVSAvoidmotion detail quality
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent changes the fundamental parameters of the generative model by introducing a diffusion model framework instead of traditional auto-encoder architectures. This parameter change enables the model to generate motion sequences with sufficient details while maintaining a manageable structure through the progressive denoising process inherent to diffusion models.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If a high-dimensional feature space is used directly for motion generation, then the detail richness is improved, but the computational complexity increases significantly

Engineering Contradiction:
Improvedetail richnessVSAvoidcomputational complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary generation of an intermediate motion sequence in a lower-dimensional space before refining it in the high-dimensional feature space. This preliminary action reduces the computational burden by handling the bulk of the generation task in a compressed representation, allowing the high-dimensional refinement stage to focus only on adding detailed nuances without the full computational cost.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20260004501A1Motion generation model-based motion generation method and apparatus, and device
Publication Date: 2026.01.01 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US20260004501A1 patent drawing
  • US20260004501A1 patent drawing
  • US20260004501A1 patent drawing

AI summary

Motion generation model-based motion generation method, device, and storage medium relate to the field of artificial intelligence technologies. The method includes: obtaining a text containing motion information; generating a text feature of the text through a text encoder; generating an intermediate motion sequence in a feature space of a first dimension based on the text feature through a first diffusion model; and performing detail enhancement processing on the intermediate motion sequence in a feature space of a second dimension through a second diffusion model, to obtain an output motion sequence matching the text, the second dimension being greater than the first dimension. In this application, the intermediate motion sequence is preliminarily generated through the first diffusion model, and detail enhancement processing is performed on the intermediate motion sequence through the second diffusion model, thereby improving the richness of details in the output motion sequence.