Motion Retargeting via Implicit 3D Warping Fields
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current human motion retargeting systems require extensive training for each new subject, involving large numbers of training frames and significant time, especially when only a few reference images are available, and they struggle with non-rigid human motion and varying poses without the use of explicit 3D models or high-fidelity data.
Innovation Solution
A motion retargeting system utilizing a transformable bottleneck network (TBN) that learns implicit 3D representations without explicit 3D supervision, enabling flexible and expressive modeling by encoding and manipulating image content across different resolutions and poses, and separating foreground and background processing to handle non-rigid human motion and varying poses.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If traditional motion retargeting systems use extensive training with large numbers of training frames for each new subject, then the quality of motion transfer is improved, but the training time and computational resources required increase significantly
Solution Approach 1:
The patent uses a pre-trained source model that has learned general motion patterns from extensive training data. This source model serves as a template that can be rapidly adapted to new subjects through flow-guided transformation, avoiding the need to retrain from scratch for each new subject while maintaining high motion transfer quality
Solution Approach 2:
The system transforms the pre-trained source model into a target model by changing key parameters through flow-guided motion retargeting. Instead of retraining all parameters, the system selectively adjusts parameters related to motion flows and transformations, significantly reducing training time while preserving motion quality
2Measurement precision
If the system uses explicit 3D models or high-fidelity data for motion retargeting, then the accuracy of non-rigid human motion representation is improved, but the complexity of data collection and processing increases
Solution Approach 1:
The patent replaces complex explicit 3D modeling and high-fidelity data collection processes with a learned implicit representation system. Neural networks automatically learn to represent non-rigid motions from image sequences, substituting mechanical 3D scanning and modeling processes with automated learning-based approaches that reduce complexity while maintaining accuracy
Solution Approach 2:
The system transforms 2D image sequences into representations that capture 3D non-rigid motion information through learned features. By processing images through neural networks that learn temporal and spatial patterns, the system extracts 3D motion semantics from 2D observations without requiring explicit 3D data
3Manufacturing precision
If the system processes images at high resolution to retain fine-grained details, then the quality of generated motion videos is improved, but the computational resources and processing time increase
Solution Approach 1:
The patent divides the image processing into multiple stages with different resolution requirements. Early stages process lower-resolution images to capture global motion patterns, while later stages process higher-resolution images to refine local details. This segmented approach maintains fine-grained detail quality while reducing overall computational energy consumption
Data Source
AI summary
Systems and methods herein describe a motion retargeting system. The motion retargeting system accesses a plurality of two-dimensional images comprising a person performing a plurality of body poses, extracts a plurality of implicit volumetric representations from the plurality of body poses, generates a three-dimensional warping field, the three-dimensional warping field configured to warp the plurality of implicit volumetric representations from a canonical pose to a target pose, and based on the three-dimensional warping field, generates a two-dimensional image of an artificial person performing the target pose.


