Monocular 3D Pose Estimation for Interacting Humans
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer vision techniques struggle to accurately estimate 3D pose of multiple individuals interacting in close proximity, as they often ignore contextual information and result in degraded performance due to occlusions and ambiguities.
Innovation Solution
A layered model that combines bottom-up observations with top-down prior knowledge and context, using a multi-aspect flexible pictorial structure model to infer 2D and 3D pose estimates from monocular images, accounting for human-human interactions and correlated activities like dancing, and employing Gaussian Process Dynamical Models for robust inference.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If independent pose estimation is used for multiple individuals, then computational simplicity is maintained, but pose estimation accuracy degrades due to competing estimates and ignored contextual information
Solution Approach 1:
The patent merges the pose estimation processes of multiple individuals into a unified contextual framework. Instead of estimating poses independently, the system combines information from all individuals in the scene, using each person as contextual information for others. This merging resolves the contradiction by maintaining computational tractability while improving accuracy through joint estimation that leverages mutual contextual relationships.
Solution Approach 2:
The patent introduces contextual information as an intermediary that mediates between individual pose estimates. By using each individual as contextual information for others, the system creates a bridging mechanism that connects previously independent estimation problems. This intermediary approach allows the system to maintain relative computational simplicity while significantly improving pose estimation accuracy through the mediating contextual relationships.
2Ease of operation
If tracking-by-detection methods are used, then individual tracking is achieved, but contextual information from scene, objects, and other people is ignored
Solution Approach 1:
The patent merges tracking capability with contextual information utilization by combining detection results with contextual relationships. The system integrates information from scene, objects, and other people into the tracking process, resolving the contradiction by maintaining ease of operation through unified processing while preventing loss of contextual information through joint estimation that explicitly incorporates these elements.
3Measurement precision
If pose estimation is performed for well-separated subjects, then estimation accuracy is improved, but performance degrades when two people are in close proximity due to occlusions and part-person ambiguities
Solution Approach 1:
The patent applies dynamics by making the pose estimation model adaptive to different spatial configurations. The system dynamically adjusts its estimation approach based on the proximity and interaction context of individuals. This dynamic adaptation resolves the contradiction by maintaining high accuracy for well-separated subjects while automatically adjusting to handle close interactions with occlusions and ambiguities through contextual relationships.
Solution Approach 2:
The patent changes the parameters of the estimation model based on interaction context. By using contextual information from correlated activities, the system adjusts estimation parameters dynamically. This parameter change approach resolves the contradiction by maintaining accuracy across different scenarios - using appropriate parameters for well-separated subjects while adapting parameters to handle the complexities of close interactions.
Data Source
AI summary
Techniques are disclosed for the automatic recovery of two dimensional (2D) and three dimensional (3D) poses of multiple subjects interacting with one another, as depicted in a sequence of 2D images. As part of recovering 2D and 3D pose estimates, a pose recovery tool may account for constraints on positions of body parts of the first and second person resulting from the correlated activity. That is, individual subjects in the video are treated as mutual context for one another.


