Semantic-aware Visual Localization via Feature Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current visual localization methods, such as 3D structure-based and 2D image-based localization, are not robust to changes in viewing conditions like illumination, weather, and dynamic scene changes, and face scalability and computational complexity issues.

Innovation Solution

The implementation of a semantically-aware image-based localization system using neural networks with spatial attention to extract and fuse appearance and semantic features, projecting them into a shared embedding space for improved robustness and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If low-level appearance descriptors are used for visual localization, then pose estimation can be performed, but the method is not robust to changes in viewing conditions such as illumination, weather, and seasons

Engineering Contradiction:
Improverobustness to viewing condition changesVSAvoidlocalization accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent segments the image into multiple semantic regions (e.g., sky, building, road, vegetation) and extracts appearance descriptors for each region separately. This segmentation allows the system to focus on stable, discriminative regions while ignoring dynamic or less informative areas, thereby improving robustness to viewing condition changes without sacrificing localization accuracy

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different processing strategies to different semantic regions based on their stability and discriminative power. Stable regions (e.g., buildings, roads) are weighted more heavily in the fusion process, while dynamic regions (e.g., vegetation, sky) are either down-weighted or processed differently. This local quality approach ensures that the most reliable features drive the localization decision

Inventive Principle:
Principle #3Local quality

2Measurement precision

If 3D structure-based localization methods are used, then localization can be achieved, but the methods are less scalable and present greater computational complexity

Engineering Contradiction:
Improvelocalization accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the essential 2D appearance descriptors from images and uses these directly for localization, avoiding the need to reconstruct full 3D models or perform computationally intensive 3D-2D matching. By taking out only the necessary features (appearance descriptors from segmented regions) and using them in a 2D embedding space, the system achieves comparable accuracy to 3D methods with significantly reduced computational complexity

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates a 2D embedding space that captures the essential geometric and appearance information normally requiring 3D models. By copying the relevant information into this simplified 2D representation space, the system avoids the computational burden of 3D processing while maintaining localization accuracy

Inventive Principle:
Principle #26Copying

3Device complexity

If traditional 2D image-based localization methods are used, then the process is simpler, but they are not robust to dynamic scene changes and occlusions

Engineering Contradiction:
Improvemethod simplicityVSAvoidrobustness to dynamic changes
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent segments the image into semantic regions and identifies which regions are stable versus dynamic. By separating stable regions (e.g., buildings, infrastructure) from dynamic regions (e.g., pedestrians, vehicles, vegetation), the system can focus matching on stable regions, making the simple 2D method robust to dynamic scene changes and occlusions

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent dynamically adjusts the weighting of different semantic regions based on their stability characteristics. Stable regions receive higher weights in the fusion process, while dynamic regions receive lower weights or are excluded. This parameter change (weight adjustment) allows the simple 2D method to adapt to dynamic conditions and maintain reliability

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11361470B2Semantically-aware image-based visual localization
Publication Date: 2022.06.14 SRI INTERNATIONAL
  • US11361470B2 patent drawing
  • US11361470B2 patent drawing
  • US11361470B2 patent drawing

AI summary

A method, apparatus and system for visual localization includes extracting appearance features of an image, extracting semantic features of the image, fusing the extracted appearance features and semantic features, pooling and projecting the fused features into a semantic embedding space having been trained using fused appearance and semantic features of images having known locations, computing a similarity measure between the projected fused features and embedded, fused appearance and semantic features of images, and predicting a location of the image associated with the projected, fused features. An image can include at least one image from a plurality of modalities such as a Light Detection and Ranging image, a Radio Detection and Ranging image, or a 3D Computer Aided Design modeling image, and an image from a different sensor, such as an RGB image sensor, captured from a same geo-location, which is used to determine the semantic features of the multi-modal image.