3D Human Pose Estimation via Epipolar Geometry
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current human pose estimation systems struggle to accurately estimate three-dimensional (3D) poses from a single image without requiring 3D supervision or camera extrinsics, and they often rely on limited and costly 3D datasets.
Innovation Solution
A system that uses a data processor to execute computer executable instructions for receiving images, extracting features, and generating 3D poses using a neural network, which also retrieves a weight matrix from an annotation server to produce a desired keypoint set, thereby creating its own 3D supervision using epipolar geometry and 2D ground-truth poses.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If 3D pose estimation is performed using traditional methods with limited 3D datasets, then the system can produce 3D pose estimates, but the accuracy and reliability are insufficient due to lack of supervision and small dataset size
Solution Approach 1:
The patent creates synthetic 3D pose data by projecting 2D pose annotations into 3D space using epipolar geometry and camera matrices. This copying approach generates virtual 3D supervision signals from existing 2D data, eliminating the need for costly manual 3D annotations while providing sufficient training data for accurate pose estimation
Solution Approach 2:
The system performs self-supervised learning by automatically generating its own 3D supervision signals from 2D pose data through geometric projections. The model creates its own training labels using epipolar constraints and camera extrinsics, enabling it to improve accuracy without external 3D annotations or human intervention
2Ease of manufacture
If weakly or self-supervised approaches are used to obtain accurate 3D pose estimator with minimal supervision, then the quantity of supervision data is reduced, but the measurement precision and reliability of 3D pose estimation deteriorate
Solution Approach 1:
The patent introduces epipolar geometry and camera projection matrices as intermediary elements that bridge 2D pose data and 3D pose estimation. These mathematical constraints act as mediators that transform minimal 2D supervision into reliable 3D pose predictions, maintaining accuracy while reducing supervision requirements
Solution Approach 2:
The system changes the parameter space by introducing camera extrinsics (intrinsic matrix, rotation, translation) as additional parameters that enable accurate 3D reconstruction from 2D data. By optimizing these parameters alongside pose estimates, the model achieves high precision with minimal supervision
3Measurement precision
If additional supervision or extrinsic camera parameters are used in multi-view settings, then the 3D pose estimation accuracy improves, but the device complexity and data processing requirements increase
Solution Approach 1:
The patent extracts and utilizes only the essential camera extrinsic parameters (intrinsic matrix, rotation, translation) needed for epipolar geometry calculations, separating these critical parameters from other complex multi-view processing requirements. This extraction approach maintains accuracy while reducing system complexity by focusing on minimal necessary parameters
Data Source
AI summary
A system for estimating a three dimensional pose of one or more persons in a scene is disclosed herein. The system includes one or more cameras and a data processor configured to execute computer executable instructions. The computer executable instructions include: (i) receiving the one or more images of the scene from the one or more cameras; (ii) extracting features from the one or more images of the scene for providing inputs to a three dimensional pose estimation neural network; (iii) generating, by using the three dimensional pose estimation neural network, vertices of a canonical human mesh model for the one or more images of the scene; (iv) retrieving a weight matrix from an annotation server that corresponds to a desired keypoint set; and (v) generating the desired keypoint set by multiplying the retrieved weight matrix with the vertices of the canonical human mesh model.


