Multi-View Pose Estimation via Selective Image Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing pose estimation techniques for multiple objects captured by a multi-view camera system face challenges such as increased processing time, estimation errors due to overlapping objects, and the need to determine the camera that shot each object.

Innovation Solution

A method and apparatus for generating a pose model of an object by selecting appropriate images from multiple cameras, considering factors like object positions, overlapping objects, and camera angles, to reduce processing time and estimation errors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If all multi-view camera images from multiple cameras are used to estimate pose models, then the completeness of object capture is improved, but the processing time increases significantly

Engineering Contradiction:
Improvepose estimation completenessVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the set of all camera images into multiple candidate groups, where each group contains images from a specific subset of cameras. This segmentation allows the system to process only relevant image subsets for each object rather than all images simultaneously, reducing processing time while maintaining estimation accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by selecting only the necessary subset of camera images required for accurate pose estimation of each object, rather than processing all available images. The system determines the minimum required image set that provides sufficient information for reliable pose modeling, avoiding unnecessary computational overhead.

Inventive Principle:
Principle #16Partial or excessive action

2Area of stationary object

If images from multiple cameras capturing overlapping objects are processed together, then the coverage of the scene is improved, but estimation errors increase due to object overlap

Engineering Contradiction:
Improvescene coverageVSAvoidpose estimation accuracy
Core Design Contradiction:
Area of stationary objectVSMeasurement precision

Solution Approach 1:

The patent segments the scene into distinct object regions and assigns specific camera image subsets to each object based on their spatial relationships. This segmentation prevents mixing of overlapping objects in the same processing group, eliminating estimation errors caused by object overlap while maintaining comprehensive scene coverage.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by tailoring the image selection and processing parameters specifically for each object based on its position and overlap characteristics. Different objects receive customized image subsets and processing treatments optimized for their specific spatial relationships, improving overall estimation accuracy.

Inventive Principle:
Principle #3Local quality

3Measurement precision

If the system determines which camera shot each object, then the accuracy of object identification is improved, but the device complexity increases

Engineering Contradiction:
Improveobject identification accuracyVSAvoidcamera selection complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements self-service by enabling the system to automatically determine camera assignments and select appropriate image subsets without manual intervention. The algorithm autonomously analyzes object positions, camera angles, and image content to identify which cameras captured each object, reducing operational complexity while maintaining high identification accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent employs feedback mechanisms where the system continuously refines camera selection and object identification based on processing results. The feedback loop allows the system to adjust camera assignments and image subsets dynamically, improving identification accuracy while managing computational complexity through iterative optimization.

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP3951715B1Generation apparatus, generation method, and program
Publication Date: 2025.02.19 CANON KK
  • EP3951715B1 patent drawingFigure 1~2
  • EP3951715B1 patent drawingFigure 3A~3B
  • EP3951715B1 patent drawingFigure 4

AI summary

A pose estimation apparatus (110) obtains a plurality of images captured by a plurality of image capturing apparatuses (100) from different directions, specifies an image that is to be used for generating a three-dimensional pose model indicating a plurality of joint positions of an object, from the plurality of images, and generates a three-dimensional pose model of the object based on the specified image.