Camera Pose Estimation Using Displayed Visual Content
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current camera pose estimation methods, particularly in vision-based systems, face challenges in determining accurate scale factors in environments with poor lighting or textureless conditions, and often rely on inaccurate location sensors, making them unreliable for robust camera pose determination relative to a reference object or environment.
Innovation Solution
A method and system that utilize a second camera to capture images of visual content displayed on a display device with a known spatial relationship to a first camera, allowing for robust estimation of the first camera's pose by determining the spatial relationship between the visual content and the camera, even in environments where traditional methods fail, and enabling correct scale factor calculation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If vision based methods are used to compute camera pose, then the method is robust and popular, but it requires detectable visual features and known geometry which are not always available
Solution Approach 1:
The patent introduces a display device as an intermediary object between the camera and the environment. The display device presents visual content with known geometric properties, serving as a mediator that provides the necessary reference information for pose estimation without requiring features in the natural environment. This resolves the contradiction by enabling vision-based methods to work in environments where natural visual features are absent.
Solution Approach 2:
The patent uses a virtual copy of the display device rendered in the augmented reality scene to establish correspondence with the physical display device. By comparing the rendered virtual display with the captured image of the physical display, the system can compute camera pose without requiring natural environment features. This copying approach enables the method to work universally across different environments.
2Measurement precision
If calibration objects with known geometry are introduced to determine correct scale factors, then the scale factor can be determined accurately, but the device complexity increases
Solution Approach 1:
The patent makes the display device serve multiple functions: it acts as both the user interface for augmented reality content display and as the calibration object for pose estimation. The visual content displayed on the screen provides the known geometric information needed for scale factor determination, eliminating the need for separate calibration objects. This multi-functionality reduces device complexity while maintaining measurement precision.
Solution Approach 2:
The display device performs self-calibration by using its own displayed visual content as the reference for pose estimation. The system captures images of the display device, renders a virtual copy, and computes the transformation between them to determine camera pose and scale factor. This self-service approach eliminates the need for external calibration objects or additional calibration equipment, reducing overall system complexity.
3Ease of operation
If location sensors such as GPS are used to initialize vision based pose estimation, then initialization can be provided, but the accuracy is poor especially in indoor environment
Solution Approach 1:
The patent replaces GPS location sensors with the display device as an intermediary reference object for initialization. Instead of relying on inaccurate location data from GPS, the system uses the known geometric properties of the display device and its rendered virtual copy to establish accurate camera pose. This intermediary approach provides both initialization capability and high precision, especially in indoor environments where GPS fails.
Data Source
AI summary
The invention is related to a method and system for determining a pose of a first camera, comprising providing or receiving a spatial relationship (Rvc 1) between a visual content displayed on a display device and the first camera, receiving image information associated with an image (B1) of at least part of the displayed visual content captured by a second camera, and determining a pose of the first camera according to the image information associated with the image (B1) and the spatial relationship (Rvc1).


