3D Pose Estimation Data Augmentation via Ground Plane Rotation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional 3D pose estimation models face challenges with unseen views, occlusion, and shaking poses, particularly in multi-person scenarios, and are limited by high costs and spatial and temporal constraints in generating training data.

Innovation Solution

A training data augmentation method and device that collects foot coordinates to estimate a ground plane, generates 3D pose data by rotating or moving individuals or the ground plane, and maps this data to 2D pose data, using a processor to create diverse training pairs for AI models, thereby addressing occlusion and increasing data scope and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If motion capture system is used to generate training data, then data accuracy is improved, but cost and spatial-temporal constraints increase

Engineering Contradiction:
Improvedata accuracyVSAvoidspatial-temporal constraints
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses 2D pose data from existing datasets as copies to generate synthetic 3D training data, avoiding the need for expensive motion capture systems. By projecting 2D poses onto 3D virtual environments and rendering new views, the system creates training data that mimics real capture data without requiring physical motion capture equipment

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical motion capture system with a computational approach using 2D pose estimation algorithms and 3D rendering software. Instead of using physical cameras and markers to capture motion, the system uses image processing to estimate 2D poses and then synthesizes 3D data through virtual rendering

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If existing models are trained with limited data, then training cost is reduced, but performance on unseen views degrades

Engineering Contradiction:
Improvetraining efficiencyVSAvoidperformance on unseen views
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent transforms 2D pose data into 3D training data by projecting poses onto a virtual ground plane and rendering from multiple camera angles. This dimensional transformation allows the model to learn 3D spatial relationships from 2D input, improving generalization to unseen views while maintaining training efficiency

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent pre-processes existing 2D pose data by estimating ground plane parameters and generating 3D pose hypotheses before training. This preliminary preparation creates a diverse set of training examples that cover various viewing angles and scenarios, enabling the model to adapt to unseen views without requiring extensive additional data collection

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If conventional models estimate 3D pose from 2D data, then pose estimation is achieved, but occlusion causes ambiguity and inconsistent estimates

Engineering Contradiction:
Improvepose estimation accuracyVSAvoidocclusion ambiguity
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent inverts the traditional approach by first generating 3D pose data from 2D poses, then projecting it back to 2D for training. This inversion allows the model to learn the relationship between 3D pose and 2D projection, making it more robust to occlusion by understanding the underlying 3D structure even when parts are hidden in the 2D input

Inventive Principle:
Principle #13The other way round (Inversion)

4Ease of manufacture

If single-person data is used for training, then data collection is simplified, but multi-person pose estimation performance is limited

Engineering Contradiction:
Improvedata collection simplicityVSAvoidmulti-person estimation capability
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal training framework that can handle both single-person and multi-person scenarios using the same data generation process. By generating 3D poses on a virtual ground plane and rendering from various angles, the system creates training data that is applicable to multiple persons simultaneously, enabling the model to generalize from single-person to multi-person estimation

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240378749A1Training data augmentation device and method for 3D pose estimation
Publication Date: 2024.11.14 SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
  • US20240378749A1 patent drawing
  • US20240378749A1 patent drawing
  • US20240378749A1 patent drawing

AI summary

A training data augmentation method for three-dimensional (3D) pose estimation includes collecting foot coordinates of each of persons appearing in a two-dimensional image, estimating a ground plane in a three-dimensional space based on the collected foot coordinates, generating three-dimensional pose data by moving or rotating at least one person or moving or rotating the ground plane based on two basis vectors perpendicular to a normal vector of the ground plane, mapping the 3D pose data to two-dimensional pose data based on a focal length of a camera obtained by capturing the two-dimensional image and a principal point of coordinates of the two-dimensional image, and acquiring a pair of the 3D pose data and two-dimensional pose data as training data.