Camera Pose Estimation Using Cross-Domain Feature Embedding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing augmented reality systems face limitations in outdoor environments due to the inability of device sensors to provide adequate information for estimating camera pose at larger distances and non-planar geometries, restricting the operational range and accuracy of virtual content overlay.

Innovation Solution

A neural network-based system that learns data-driven cross-domain feature embedding to match images with a rendered terrain model, allowing for the estimation of camera pose using cross-domain feature descriptors and GPS information, enabling accurate feature matching and localization for large-scale augmented reality applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If device sensors (active depth sensors, stereo camera, multiview geometry) are used for camera pose tracking, then tracking accuracy is improved, but operational range is limited due to light falloff, stereo baselines, and camera parallax constraints

Engineering Contradiction:
Improvecamera pose tracking accuracyVSAvoidoperational range
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent introduces terrain models as an intermediary reference system. Instead of relying solely on direct sensor measurements between camera and scene points, the system mediates pose estimation by matching image features against pre-built terrain models with known poses, extending operational range beyond sensor limitations

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transitions from 2D image plane matching to 3D terrain model matching. By lifting feature matching from the image plane into the third dimension using terrain elevation data, the system overcomes the parallax and baseline limitations that constrain traditional 2D stereo vision approaches

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Device complexity

If traditional feature matching methods are used between images and terrain models, then system complexity is reduced, but matching accuracy deteriorates due to domain differences between real images and rendered terrain

Engineering Contradiction:
Improvesystem complexityVSAvoidfeature matching accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent transforms the feature representation parameters by projecting both image features and terrain model features into a common embedding space using learned projection matrices. This parameter transformation enables accurate cross-domain matching despite differences between photorealistic images and synthetic terrain renders

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces traditional mechanical/optical feature matching approaches with data-driven neural network-based embedding. Instead of relying on hand-crafted descriptors and geometric constraints, the system uses learned representations that automatically adapt to the specific domain characteristics

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If more keypoints are used for camera pose estimation, then pose accuracy is improved, but computational complexity increases

Engineering Contradiction:
Improvecamera pose estimation accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent employs a two-stage approach where an initial rough pose estimate is computed first, then used to selectively refine the pose. This partial action strategy avoids the computational cost of processing all possible keypoints while still achieving high accuracy through targeted refinement on the most relevant features

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11568642B2Large-scale outdoor augmented reality scenes using camera pose based on learned descriptors
Publication Date: 2023.01.31 ADOBE INC
  • US11568642B2 patent drawing
  • US11568642B2 patent drawing
  • US11568642B2 patent drawing

AI summary

Methods and systems are provided for facilitating large-scale augmented reality in relation to outdoor scenes using estimated camera pose information. In particular, camera pose information for an image can be estimated by matching the image to a rendered ground-truth terrain model with known camera pose information. To match images with such renders, data driven cross-domain feature embedding can be learned using a neural network. Cross-domain feature descriptors can be used for efficient and accurate feature matching between the image and the terrain model renders. This feature matching allows images to be localized in relation to the terrain model, which has known camera pose information. This known camera pose information can then be used to estimate camera pose information in relation to the image.