Monocular 3D Pose Correction for Continuous Virtual Object Animation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing motion capture technologies, such as optical and inertial methods, are costly and suffer from poor motion continuity and generalization issues, while end-to-end video methods lack sufficient training data and precision.
Innovation Solution
A method and apparatus that utilize monocular video to determine 3D and 2D pose information, incorporate foot ground contact, and apply motion priors for correcting pose information, ensuring continuity and appropriateness through iterative optimization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If optical motion capture method is used, then motion capture precision is improved, but device complexity and cost increase
Solution Approach 1:
The patent extracts the motion capture function from complex optical systems and implements it through a simplified monocular video processing approach. By using a single camera and leveraging deep learning models to infer 3D poses from 2D video frames, the system eliminates the need for expensive optical motion capture equipment while maintaining acceptable precision through iterative optimization and motion prior constraints.
Solution Approach 2:
The patent replaces the mechanical/optical motion capture system with a computational vision-based system. Instead of using physical optical sensors and markers, the system uses a monocular camera combined with deep learning algorithms (specifically a 3D pose estimation model) to extract motion information from video data, substituting mechanical measurement with computational inference.
2Measurement precision
If inertial motion capture method is used, then motion capture capability is improved, but device complexity and cost increase
Solution Approach 1:
The patent removes the inertial measurement units and associated hardware from the motion capture system. By extracting the motion capture function and implementing it through monocular video analysis with deep learning models, the system eliminates complex inertial sensing equipment while maintaining the ability to capture human motion through computational methods.
Solution Approach 2:
The patent substitutes the inertial measurement system with a vision-based computational system. Instead of using accelerometers and gyroscopes, the system uses a monocular camera and iterative optimization algorithms to infer motion parameters, replacing mechanical sensing with optical observation and computational processing.
3Device complexity
If end-to-end video motion capture method is used, then device complexity is reduced, but measurement precision and generalization deteriorate due to insufficient training data
Solution Approach 1:
The patent applies preliminary action by pre-training a 3D pose estimation model using extensive training data before actual motion capture operations. This pre-trained model serves as a foundation that encodes learned patterns from diverse motion scenarios, enabling the system to handle various motion types and generalization cases without requiring retraining for each specific application, thus improving precision while maintaining simplicity.
Solution Approach 2:
The patent implements feedback through an iterative optimization process where the initial pose estimates are continuously refined using motion priors and consistency constraints. The system iteratively adjusts pose parameters by comparing predicted motions with observed video data, using feedback loops to converge on more accurate pose estimates, thereby improving measurement precision without increasing device complexity.
4Productivity
If motion retargeting is applied to virtual objects, then animation generation is improved, but motion continuity and execution accuracy deteriorate
Solution Approach 1:
The patent applies preliminary action by pre-computing motion priors and pose change information before the actual retargeting process. By preparing motion constraints and temporal consistency rules in advance, the system ensures that motion continuity is maintained during retargeting operations, preventing common issues like motion discontinuities and ensuring smooth transitions when adapting motions to different virtual objects.
Solution Approach 2:
The patent implements feedback in the retargeting process by iteratively refining pose parameters and motion trajectories. The system uses feedback from the corrected pose information and motion priors to adjust the retargeted animation, ensuring continuity and accuracy. This iterative refinement loop allows the system to maintain motion quality while generating animations for diverse virtual objects.
Data Source
AI summary
A virtual object animation generation method and apparatus including obtaining a to-be-processed video, determining initial 3D pose information, two-dimensional (2D) pose information, and foot ground contact information of a target object in the to-be-processed video, determining pose change information between two adjacent video frames in the to-be-processed video based on the initial 3D pose information of the target object, determining to-be-corrected pose information of the target object based on the initial 3D pose information of the target object and the pose change information, performing correction processing on the to-be-corrected pose information of the target object based on the 2D pose information and the foot ground contact information of the target object, to obtain corrected pose information of the target object, and retargeting the corrected pose information to a virtual object by means of motion retargeting, to generate a virtual object animation video corresponding to the to-be-processed video.


