Learning Model Selection for Geometric Estimation Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for estimating the position and orientation of an image capturing apparatus in mixed reality and other applications face accuracy issues when the scene of the captured image differs from the scene used in training the learning model.

Innovation Solution

An information processing apparatus that selects a learning model based on evaluation values indicating the suitability of the model to the input image scene, using methods such as pHash, object detection, and position information to accurately estimate geometric information and calculate the position and orientation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a single learning model trained on a specific scene is used, then the estimation accuracy is improved for that scene, but the adaptability to different scenes deteriorates

Engineering Contradiction:
Improvegeometric information estimation accuracyVSAvoidscene adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent divides the learning models into multiple specialized models, each trained on a specific scene type (indoor, outdoor, urban, rural). This segmentation allows each model to specialize in its training domain, achieving high accuracy for that specific scene while the system as a whole maintains adaptability through model selection based on scene evaluation.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If multiple learning models are maintained for different scenes, then the adaptability to different scenes is improved, but the device complexity increases

Engineering Contradiction:
Improvescene adaptabilityVSAvoidlearning model management complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces a scene evaluation unit as an intermediary that automatically assesses the input image characteristics and selects the most appropriate learning model. This intermediary component manages the complexity of having multiple models by providing an automated selection mechanism, reducing the burden on users to manually manage model complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Device complexity

If a learning model trained on dissimilar scenes is used, then the device complexity is reduced, but the measurement precision deteriorates

Engineering Contradiction:
Improvemodel selection simplicityVSAvoidposition and orientation calculation accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary scene evaluation and learning model selection before the actual geometric information estimation. The scene evaluation unit pre-assesses the input image characteristics and selects the appropriate model in advance, ensuring that the most suitable model is used without requiring complex real-time adjustments during the estimation process.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11144786B2Information processing apparatus, method for controlling information processing apparatus, and storage medium
Publication Date: 2021.10.12 CANON KK
  • US11144786B2 patent drawing
  • US11144786B2 patent drawing
  • US11144786B2 patent drawing

AI summary

An information processing apparatus comprising: a holding unit configured to hold a plurality of learning models for estimating geometric information based on an input image captured by an image capturing apparatus; a selection unit configured to calculate, for each of the learning models, an evaluation value that indicates suitability of the learning model to a scene of the input image, and select a learning model from the plurality of learning models based on the evaluation values; and an estimation unit configured to estimate first geometric information using the input image and the selected learning model.