Motion Estimation Learning Model Using Animation Transfer
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing motion estimation deep learning models are limited by the number of key points defined in standard datasets, such as COCO and MPII, which restricts their ability to precisely estimate human motion and pose.
Innovation Solution
A method and apparatus for providing a precise motion estimation learning model through transfer learning, using a standard dataset with a predetermined number of key points and an animation dataset that easily expands key points, followed by fine-tuning with a photorealistic dataset.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a standard dataset with a limited number of key points is used for training, then the training process is simple and fast, but the motion estimation precision is limited
Solution Approach 1:
The patent applies preliminary action by pre-processing animation data to generate expanded key point information before the main training process. The animation dataset is prepared in advance with additional key points (21 key points including hand, foot, and body center points) that are then used to initialize and guide the photorealistic dataset training, eliminating the need to collect and annotate large amounts of complex photorealistic data from scratch.
Solution Approach 2:
The patent uses animation data as an intermediary between the standard dataset and the final photorealistic motion estimation model. The animation dataset with expanded key points serves as a bridge that transfers learning from the limited standard dataset to achieve more precise motion estimation with fewer photorealistic annotations required.
2Measurement precision
If the number of key points is increased to improve motion estimation accuracy, then the precision is improved, but the time and effort for data collection and annotation increases significantly
Solution Approach 1:
The patent uses copying by training the model on animation data (which can be easily generated and annotated) before fine-tuning on photorealistic data. The animation dataset with 21 key points serves as a copy or proxy that allows the model to learn expanded key point detection without requiring equivalent amounts of complex photorealistic annotated data.
Solution Approach 2:
The patent applies parameter changes by modifying the key point configuration from the standard 17 key points to an expanded 21 key point system. This parameter change is achieved by leveraging animation data that naturally contains the additional key point information, thereby improving detection accuracy without proportionally increasing annotation effort.
3Reliability
If a large amount of photorealistic data with expanded key points is collected, then the model can achieve better performance, but the data preparation process becomes extremely time-consuming and resource-intensive
Solution Approach 1:
The patent performs preliminary action by pre-training on animation data that contains expanded key points. This preliminary training phase establishes a strong foundation that requires minimal photorealistic data for fine-tuning, dramatically improving data preparation efficiency while maintaining high model reliability.
Solution Approach 2:
Animation data serves as an intermediary that enables the model to achieve high reliability without requiring large amounts of photorealistic data. The intermediary animation dataset bridges the gap between limited standard data and the need for high-performance motion estimation, maintaining productivity by avoiding direct collection of extensive photorealistic annotated data.
Data Source
AI summary
The present disclosure is an apparatus and a method for providing a precise motion estimation learning model including, a database unit which stores a standard dataset labeled according to a first number of key points, an animation dataset labeled according to a second number of key points which is larger than the first number, and a photorealistic dataset having the second number of key points, a standard learning unit which learns the standard dataset for motion estimation to generate a standard learning model, an animation learning unit which retrains the animation dataset based on a weight of the standard learning model to generate an animation learning model, and a motion estimation learning unit which trains the photorealistic dataset based on the weight of the animation learning model to finely tune to generate a precise motion estimation learning model.


