Monocular 3D Pose Estimation for Interacting Humans

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer vision techniques struggle to accurately estimate 3D pose of multiple individuals interacting in close proximity, as they often ignore contextual information and result in degraded performance due to occlusions and ambiguities.

Innovation Solution

A layered model that combines bottom-up observations with top-down prior knowledge and context, using a multi-aspect flexible pictorial structure model to infer 2D and 3D pose estimates from monocular images, accounting for human-human interactions and correlated activities like dancing, and employing Gaussian Process Dynamical Models for robust inference.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If independent pose estimation is used for multiple individuals, then computational simplicity is maintained, but pose estimation accuracy degrades due to competing estimates and ignored contextual information

Engineering Contradiction:
Improvecomputational simplicityVSAvoidpose estimation accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent merges the pose estimation processes of multiple individuals into a unified contextual framework. Instead of estimating poses independently, the system combines information from all individuals in the scene, using each person as contextual information for others. This merging resolves the contradiction by maintaining computational tractability while improving accuracy through joint estimation that leverages mutual contextual relationships.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces contextual information as an intermediary that mediates between individual pose estimates. By using each individual as contextual information for others, the system creates a bridging mechanism that connects previously independent estimation problems. This intermediary approach allows the system to maintain relative computational simplicity while significantly improving pose estimation accuracy through the mediating contextual relationships.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If tracking-by-detection methods are used, then individual tracking is achieved, but contextual information from scene, objects, and other people is ignored

Engineering Contradiction:
Improvetracking capabilityVSAvoidcontextual information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent merges tracking capability with contextual information utilization by combining detection results with contextual relationships. The system integrates information from scene, objects, and other people into the tracking process, resolving the contradiction by maintaining ease of operation through unified processing while preventing loss of contextual information through joint estimation that explicitly incorporates these elements.

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If pose estimation is performed for well-separated subjects, then estimation accuracy is improved, but performance degrades when two people are in close proximity due to occlusions and part-person ambiguities

Engineering Contradiction:
Improvepose estimation accuracyVSAvoidhandling of close interactions
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies dynamics by making the pose estimation model adaptive to different spatial configurations. The system dynamically adjusts its estimation approach based on the proximity and interaction context of individuals. This dynamic adaptation resolves the contradiction by maintaining high accuracy for well-separated subjects while automatically adjusting to handle close interactions with occlusions and ambiguities through contextual relationships.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameters of the estimation model based on interaction context. By using contextual information from correlated activities, the system adjusts estimation parameters dynamically. This parameter change approach resolves the contradiction by maintaining accuracy across different scenarios - using appropriate parameters for well-separated subjects while adapting parameters to handle the complexities of close interactions.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9058663B2Modeling human-human interactions for monocular 3D pose estimation
Publication Date: 2015.06.16 DISNEY ENTERPRISES INC
  • US9058663B2 patent drawing
  • US9058663B2 patent drawing
  • US9058663B2 patent drawing

AI summary

Techniques are disclosed for the automatic recovery of two dimensional (2D) and three dimensional (3D) poses of multiple subjects interacting with one another, as depicted in a sequence of 2D images. As part of recovering 2D and 3D pose estimates, a pose recovery tool may account for constraints on positions of body parts of the first and second person resulting from the correlated activity. That is, individual subjects in the video are treated as mutual context for one another.