Text-Tagged Character Motion Generation for AI Training Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies require high-quality training data for artificial intelligence models to infer character motion from text, and there is a need for efficient methods to generate text-tagged motion data for training these models.
Innovation Solution
A motion generation device that converts animation data into intermediate data, generates character motion, and uses a language model to tag text to the motion, allowing for the creation of text-tagged motion data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If high-quality training data with labeled character motion is used to train AI models, then model inference performance is improved, but data collection and preparation complexity increases
Solution Approach 1:
The system automatically generates training data by processing existing animation data through a pipeline that includes motion extraction, text generation via language models, and automatic labeling. This self-service approach eliminates the need for manual annotation of training data, significantly reducing preparation complexity while maintaining data quality for model training
Solution Approach 2:
The system creates synthetic training data by copying and transforming existing animation data through automated processing pipelines. Existing animation data is converted into labeled motion-text pairs through multiple processing stages, generating training data that mirrors real-world requirements without needing actual filmed performances
2Manufacturing precision
If manual annotation of training data is performed, then data quality is improved, but time consumption increases
Solution Approach 1:
The system replaces manual mechanical annotation processes with automated computational pipelines. Language models generate text descriptions automatically, motion extraction algorithms process animation data automatically, and the entire labeling process is performed through software automation rather than human annotators, dramatically reducing time consumption while maintaining quality through multiple validation stages
3Quantity of substance
If existing animation data is directly used for training, then data availability is improved, but data relevance to text-conditioned motion decreases
Solution Approach 1:
The system introduces an intermediary processing pipeline between existing animation data and training data generation. This pipeline includes motion extraction modules, text generation components, and labeling mechanisms that transform raw animation data into text-conditioned motion training data, ensuring relevance while utilizing available data resources
Solution Approach 2:
The system changes the parameters and format of existing animation data through automated processing. Animation data is transformed into intermediate representations, then converted into text-label pairs with modified parameters that make the data suitable for text-conditioned motion inference training, bridging the gap between raw data availability and training requirements
Data Source
AI summary
Disclosed is a motion generation device tagged with text and an operation method thereof. The motion generation device may include a memory configured to store at least one instruction; and at least one processor configured to execute the at least one instruction stored in the memory, wherein the at least one processor is configured to: obtain animation data including a character, convert the obtained animation data into intermediate data for generating a motion of the character over a plurality of frames, generate the motion of the character based on the converted intermediate data, generate a caption for each of the plurality of frames included in the generated motion of the character, generate a text corresponding to the motion of the character by providing the plurality of generated captions to a language model, and generate the text-tagged motion by labeling the generated text to the motion of the character.


