3D Pose Estimation Data Augmentation via Ground Plane Rotation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional 3D pose estimation models face challenges with unseen views, occlusion, and shaking poses, particularly in multi-person scenarios, and are limited by high costs and spatial and temporal constraints in generating training data.
Innovation Solution
A training data augmentation method and device that collects foot coordinates to estimate a ground plane, generates 3D pose data by rotating or moving individuals or the ground plane, and maps this data to 2D pose data, using a processor to create diverse training pairs for AI models, thereby addressing occlusion and increasing data scope and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If motion capture system is used to generate training data, then data accuracy is improved, but cost and spatial-temporal constraints increase
Solution Approach 1:
The patent uses 2D pose data from existing datasets as copies to generate synthetic 3D training data, avoiding the need for expensive motion capture systems. By projecting 2D poses onto 3D virtual environments and rendering new views, the system creates training data that mimics real capture data without requiring physical motion capture equipment
Solution Approach 2:
The patent replaces the mechanical motion capture system with a computational approach using 2D pose estimation algorithms and 3D rendering software. Instead of using physical cameras and markers to capture motion, the system uses image processing to estimate 2D poses and then synthesizes 3D data through virtual rendering
2Productivity
If existing models are trained with limited data, then training cost is reduced, but performance on unseen views degrades
Solution Approach 1:
The patent transforms 2D pose data into 3D training data by projecting poses onto a virtual ground plane and rendering from multiple camera angles. This dimensional transformation allows the model to learn 3D spatial relationships from 2D input, improving generalization to unseen views while maintaining training efficiency
Solution Approach 2:
The patent pre-processes existing 2D pose data by estimating ground plane parameters and generating 3D pose hypotheses before training. This preliminary preparation creates a diverse set of training examples that cover various viewing angles and scenarios, enabling the model to adapt to unseen views without requiring extensive additional data collection
3Measurement precision
If conventional models estimate 3D pose from 2D data, then pose estimation is achieved, but occlusion causes ambiguity and inconsistent estimates
Solution Approach 1:
The patent inverts the traditional approach by first generating 3D pose data from 2D poses, then projecting it back to 2D for training. This inversion allows the model to learn the relationship between 3D pose and 2D projection, making it more robust to occlusion by understanding the underlying 3D structure even when parts are hidden in the 2D input
4Ease of manufacture
If single-person data is used for training, then data collection is simplified, but multi-person pose estimation performance is limited
Solution Approach 1:
The patent creates a universal training framework that can handle both single-person and multi-person scenarios using the same data generation process. By generating 3D poses on a virtual ground plane and rendering from various angles, the system creates training data that is applicable to multiple persons simultaneously, enabling the model to generalize from single-person to multi-person estimation
Data Source
AI summary
A training data augmentation method for three-dimensional (3D) pose estimation includes collecting foot coordinates of each of persons appearing in a two-dimensional image, estimating a ground plane in a three-dimensional space based on the collected foot coordinates, generating three-dimensional pose data by moving or rotating at least one person or moving or rotating the ground plane based on two basis vectors perpendicular to a normal vector of the ground plane, mapping the 3D pose data to two-dimensional pose data based on a focal length of a camera obtained by capturing the two-dimensional image and a principal point of coordinates of the two-dimensional image, and acquiring a pair of the 3D pose data and two-dimensional pose data as training data.


