Visual Localization Using Semantic Error Images for Scene Changes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current visual localization methods face challenges in maintaining accuracy due to significant changes in scene environments, such as light and season variations, and structural overlaps, leading to reduced localization effectiveness.

Innovation Solution

A visual localization method based on semantic error images, which involves feature extraction, semantic segmentation, and the construction of reprojection and semantic error images to determine a hypothesized pose with minimum errors, using a three-dimensional scene model and the PNP algorithm to improve localization accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional visual localization methods (3D structure, image-based, or learning model) are used, then localization can be performed in general conditions, but localization accuracy significantly degrades when scene environment changes significantly

Engineering Contradiction:
Improvelocalization reliabilityVSAvoidlocalization accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent segments the image matching process by introducing semantic segmentation to divide the image into different semantic regions (sky, building, road, etc.). This segmentation allows the system to focus on semantically consistent regions for matching, reducing the impact of environmental changes in non-critical areas and improving both reliability and accuracy under varying conditions.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter space by introducing semantic information as an additional dimension for comparison. Instead of only comparing geometric or photometric parameters, the system now compares semantic labels alongside traditional image features. This parameter expansion makes the localization more robust to environmental changes while maintaining accuracy.

Inventive Principle:
Principle #35Parameter changes

2Ease of manufacture

If image similarity retrieval is used for pose estimation, then the method is simple to implement, but the localization accuracy is low due to reduced structural overlaps between images

Engineering Contradiction:
Improvemethod simplicityVSAvoidlocalization accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent introduces semantic information as an intermediary element that mediates between image similarity retrieval and pose estimation. By incorporating semantic labels as an additional comparison dimension, the system enhances the reliability of matching without significantly complicating the overall workflow, thus improving accuracy while maintaining relative simplicity.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If a learning model is trained for each specific scene, then pose estimation can be performed using the model, but the method lacks generality when applied to large scenes

Engineering Contradiction:
Improvepose estimation accuracyVSAvoidscene generality
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent enhances the universality of the localization system by introducing semantic information that is scene-agnostic. Semantic labels such as sky, building, road, and vegetation can be applied across different scenes and environments, making the system more adaptable and generalizable without requiring scene-specific retraining, thus improving both versatility and maintaining accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20220138484A1Visual localization method and apparatus based on semantic error image
Publication Date: 2022.05.05 NAT UNIV OF DEFENSE TECH
  • US20220138484A1 patent drawing
  • US20220138484A1 patent drawing

AI summary

The present disclosure provides a visual localization method and apparatus based on a semantic error image. The method includes: performing feature extraction for a target image, and obtaining at least one matching pair by performing feature matching for each extracted feature point and each three-dimensional point of a constructed three-dimensional scene model; obtaining a two-dimensional semantic image of the target image by performing semantic segmentation for the target image; and determining semantic information of each matching pair according to semantic information of each pixel of the two-dimensional semantic image; constructing a hypothesized pose pool including at least one hypothesized pose according to at least one matching pair; for each hypothesized pose, constructing a reprojection error image and a semantic error image; determining a hypothesized pose with a minimum reprojection error and a minimum semantic error as a pose estimation according to the reprojection error image and the semantic error image of each hypothesized pose. Optimal pose screening is performed using the semantic error image constructed based on a semantic error, so as to achieve good localization effect even in a case of significant change of a scene.