Text-Driven Motion Generation With Two-Stage Diffusion Refinement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing text-driven motion synthesis methods using auto-encoders and variational auto-encoders generate human motion sequences with insufficient details due to limited expression capability in low-dimensional feature spaces.
Innovation Solution
A motion generation model comprising a text encoder, a first diffusion model operating in a low-dimensional feature space, and a second diffusion model in a higher-dimensional space to refine the intermediate motion sequence, enhancing detail processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a low-dimensional feature space is used for motion sequence generation, then the generation process is computationally efficient, but the expression capability and detail richness are insufficient
Solution Approach 1:
The patent transitions from a low-dimensional feature space to a high-dimensional feature space to enhance the expression capability and detail richness of generated motion sequences. By increasing the dimensionality of the feature space, the model can capture more nuanced motion details while maintaining computational efficiency through the diffusion model framework.
Solution Approach 2:
The patent employs a two-stage generation process: first generating an intermediate motion sequence in a compressed representation, then refining it in a high-dimensional feature space. This segmentation allows the system to benefit from both computational efficiency in the first stage and detail richness in the second stage, resolving the contradiction between efficiency and quality.
2Device complexity
If auto-encoders or variational auto-encoders are used for motion synthesis, then the model structure is simple, but the generated motion sequences lack sufficient details
Solution Approach 1:
The patent changes the fundamental parameters of the generative model by introducing a diffusion model framework instead of traditional auto-encoder architectures. This parameter change enables the model to generate motion sequences with sufficient details while maintaining a manageable structure through the progressive denoising process inherent to diffusion models.
3Manufacturing precision
If a high-dimensional feature space is used directly for motion generation, then the detail richness is improved, but the computational complexity increases significantly
Solution Approach 1:
The patent performs preliminary generation of an intermediate motion sequence in a lower-dimensional space before refining it in the high-dimensional feature space. This preliminary action reduces the computational burden by handling the bulk of the generation task in a compressed representation, allowing the high-dimensional refinement stage to focus only on adding detailed nuances without the full computational cost.
Data Source
AI summary
Motion generation model-based motion generation method, device, and storage medium relate to the field of artificial intelligence technologies. The method includes: obtaining a text containing motion information; generating a text feature of the text through a text encoder; generating an intermediate motion sequence in a feature space of a first dimension based on the text feature through a first diffusion model; and performing detail enhancement processing on the intermediate motion sequence in a feature space of a second dimension through a second diffusion model, to obtain an output motion sequence matching the text, the second dimension being greater than the first dimension. In this application, the intermediate motion sequence is preliminarily generated through the first diffusion model, and detail enhancement processing is performed on the intermediate motion sequence through the second diffusion model, thereby improving the richness of details in the output motion sequence.


