Visual Localization Map Generation Using 3D Model Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current visual localization methods require extensive map generation processes, which are time-consuming and costly, especially for outdoor environments, and struggle with precision due to the inclusion of hindering factors like trees and roads.

Innovation Solution

A method utilizing 3D model data based on aerial photos to generate a feature point map, allowing for visual localization by rendering images from a virtual camera pose and excluding unnecessary objects, enabling precise 3D position and pose estimation with reduced data noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional visual localization methods are used with extensive map generation processes, then comprehensive map coverage is achieved, but time consumption and cost increase significantly

Engineering Contradiction:
Improvelocalization precisionVSAvoidmap generation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-generating 3D model data from aerial photos and pre-extracting feature points before actual localization is needed. The 3D model data including depth information is prepared in advance, and feature points are extracted and stored in a feature point map beforehand, so that when localization is required, the system can quickly perform matching without time-consuming map generation processes.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If traditional map generation includes all objects in the scene, then complete environmental representation is achieved, but precision deteriorates due to hindering factors like trees and roads

Engineering Contradiction:
Improveposition estimation precisionVSAvoidmap data complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies the extraction principle by selectively removing hindering objects from the 3D model data. Objects that obstruct visual localization such as trees, roads, and other irrelevant elements are identified and excluded from the feature point map. Only relevant and useful feature points are extracted and stored, creating a simplified yet effective map for accurate localization without the noise from hindering factors.

Inventive Principle:
Principle #2Taking out (Extraction)

3Productivity

If 3D model data from aerial photos is used directly, then map generation efficiency is improved, but data quality deteriorates due to noise and irrelevant information

Engineering Contradiction:
Improvemap generation efficiencyVSAvoidfeature point quality
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent converts the potential harm of noise and irrelevant information in aerial photo-based 3D model data into a benefit by systematically filtering and processing the data. The method transforms the raw, noisy 3D model data into a refined feature point map by extracting only meaningful feature points and removing irrelevant information, thereby turning the initial data quality issue into an opportunity for creating a high-quality, optimized localization map.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Data Source

PatentUS11941755B2Method of generating map and visual localization system using the map
Publication Date: 2024.03.26 NAVER LABS CORP
  • US11941755B2 patent drawing
  • US11941755B2 patent drawing
  • US11941755B2 patent drawing

AI summary

A method of generating a map for visual localization includes specifying a virtual camera pose by using 3-dimensional (3D) model data which is based on an image of an outdoor space captured from the air; rendering the image of the outdoor space from a perspective of the virtual camera, by using the virtual camera pose and the 3D model data; and generating a feature point map by using the rendered image and the virtual camera pose.