Human Body Orientation Estimation via Skeleton and Image Feature Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for estimating the orientation of a person's entire body from video images struggle to accurately distinguish between different orientations, such as sitting and squatting, due to reliance on joint information alone, which can lead to incorrect interpretations.

Innovation Solution

An information processing apparatus that acquires images, detects the entire human body, estimates its skeleton, extracts feature quantities from both the skeleton and clipped images, and uses a neural network to estimate orientation by connecting these features, incorporating background information to improve accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If orientation estimation is performed based only on joint information, then the estimation process is simple, but different orientations cannot be distinguished accurately

Engineering Contradiction:
Improveestimation process complexityVSAvoidorientation estimation accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent combines skeleton information (joint positions) with clipped image information to create a comprehensive feature quantity for orientation estimation. This merging of multiple information sources allows the system to distinguish between different orientations (such as sitting vs. squatting) that would be indistinguishable using joint information alone, thereby resolving the contradiction between simplicity and accuracy.

Inventive Principle:
Principle #5Merging (Combining)

2Productivity

If only skeleton information is used for orientation estimation, then processing is faster, but estimation accuracy deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidorientation estimation accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the feature extraction process into two parallel paths: one extracting features from skeleton information and another extracting features from clipped images. These segmented features are then combined to form the complete feature quantity used for orientation estimation, allowing the system to maintain processing efficiency while improving accuracy.

Inventive Principle:
Principle #1Segmentation

3Ease of manufacture

If joint-based orientation estimation is used, then the method is easy to implement, but misinterpretation of orientations occurs

Engineering Contradiction:
Improveimplementation easeVSAvoidorientation interpretation reliability
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent introduces clipped image information as an intermediary element that mediates between the simple joint-based estimation and the need for accurate orientation interpretation. This intermediary provides additional contextual information that helps correctly interpret orientations that would otherwise be ambiguous, maintaining implementation ease while improving reliability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240331192A1Information processing apparatus, orientation estimation method, and storage medium
Publication Date: 2024.10.03 CANON KK
  • US20240331192A1 patent drawing
  • US20240331192A1 patent drawing
  • US20240331192A1 patent drawing

AI summary

An information processing apparatus includes at least one processor, and at least one memory storing executable instructions which, when executed by the at least one processor, cause the at least one processor to perform operations including acquiring an image, detecting an entire human body from the acquired image, estimating a skeleton of the detected entire human body and generating skeleton information about the skeleton of the entire human body, extracting a first feature quantity based on the generated skeleton information, extracting a second feature quantity based on a clipped image including the detected entire human body, and estimating an orientation of the detected entire human body based on a third feature quantity in which the first and second feature quantities are connected.