Camera Pose Estimation Using Virtual Scene Appearance Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for estimating the position and orientation of an image capturing device, such as in robots or augmented reality, face accuracy issues when the appearances of scenes differ significantly.
Innovation Solution
A system that generates a virtual space image and corresponding geometric information, using a learning model to calculate the position and orientation of an image capturing device based on input images, by aligning the virtual and physical space appearances through viewpoint and environment parameter adjustment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a learning model is trained using image information from a physical space, then position and orientation can be calculated, but the calculation accuracy deteriorates when the appearance of the captured scene differs significantly from the training scene
Solution Approach 1:
The patent creates a virtual space that copies the physical space geometry but allows independent appearance control. The learning model is trained using rendered images from this virtual space, where appearances can be systematically varied to match different capture conditions, thereby improving both accuracy and adaptability without requiring additional physical scene data
Solution Approach 2:
The patent changes appearance parameters (illumination, material properties, camera parameters) of the virtual space independently from the geometric structure. By adjusting these parameters during rendering, the system can generate training data that matches the appearance characteristics of actual captured images, resolving the contradiction between maintaining geometric accuracy and adapting to varying appearances
Data Source
AI summary
As learning data, an image of a virtual space corresponding to a physical space and geometric information of the virtual space is generated. Learning processing of a learning model is performed using the learning data. A position and/or orientation of an image capturing device is calculated based on geometric information output from the learning model when a captured image of the physical space captured by the image capturing device is input to the learning model.


