Camera Pose Estimation Using Cross-Domain Feature Embedding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing augmented reality systems face limitations in outdoor environments due to the inability of device sensors to provide adequate information for estimating camera pose at larger distances and non-planar geometries, restricting the operational range and accuracy of virtual content overlay.
Innovation Solution
A neural network-based system that learns data-driven cross-domain feature embedding to match images with a rendered terrain model, allowing for the estimation of camera pose using cross-domain feature descriptors and GPS information, enabling accurate feature matching and localization for large-scale augmented reality applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If device sensors (active depth sensors, stereo camera, multiview geometry) are used for camera pose tracking, then tracking accuracy is improved, but operational range is limited due to light falloff, stereo baselines, and camera parallax constraints
Solution Approach 1:
The patent introduces terrain models as an intermediary reference system. Instead of relying solely on direct sensor measurements between camera and scene points, the system mediates pose estimation by matching image features against pre-built terrain models with known poses, extending operational range beyond sensor limitations
Solution Approach 2:
The patent transitions from 2D image plane matching to 3D terrain model matching. By lifting feature matching from the image plane into the third dimension using terrain elevation data, the system overcomes the parallax and baseline limitations that constrain traditional 2D stereo vision approaches
2Device complexity
If traditional feature matching methods are used between images and terrain models, then system complexity is reduced, but matching accuracy deteriorates due to domain differences between real images and rendered terrain
Solution Approach 1:
The patent transforms the feature representation parameters by projecting both image features and terrain model features into a common embedding space using learned projection matrices. This parameter transformation enables accurate cross-domain matching despite differences between photorealistic images and synthetic terrain renders
Solution Approach 2:
The patent replaces traditional mechanical/optical feature matching approaches with data-driven neural network-based embedding. Instead of relying on hand-crafted descriptors and geometric constraints, the system uses learned representations that automatically adapt to the specific domain characteristics
3Measurement precision
If more keypoints are used for camera pose estimation, then pose accuracy is improved, but computational complexity increases
Solution Approach 1:
The patent employs a two-stage approach where an initial rough pose estimate is computed first, then used to selectively refine the pose. This partial action strategy avoids the computational cost of processing all possible keypoints while still achieving high accuracy through targeted refinement on the most relevant features
Data Source
AI summary
Methods and systems are provided for facilitating large-scale augmented reality in relation to outdoor scenes using estimated camera pose information. In particular, camera pose information for an image can be estimated by matching the image to a rendered ground-truth terrain model with known camera pose information. To match images with such renders, data driven cross-domain feature embedding can be learned using a neural network. Cross-domain feature descriptors can be used for efficient and accurate feature matching between the image and the terrain model renders. This feature matching allows images to be localized in relation to the terrain model, which has known camera pose information. This known camera pose information can then be used to estimate camera pose information in relation to the image.


