Visual Localization Using Semantic Error Images for Scene Changes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current visual localization methods face challenges in maintaining accuracy due to significant changes in scene environments, such as light and season variations, and structural overlaps, leading to reduced localization effectiveness.
Innovation Solution
A visual localization method based on semantic error images, which involves feature extraction, semantic segmentation, and the construction of reprojection and semantic error images to determine a hypothesized pose with minimum errors, using a three-dimensional scene model and the PNP algorithm to improve localization accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional visual localization methods (3D structure, image-based, or learning model) are used, then localization can be performed in general conditions, but localization accuracy significantly degrades when scene environment changes significantly
Solution Approach 1:
The patent segments the image matching process by introducing semantic segmentation to divide the image into different semantic regions (sky, building, road, etc.). This segmentation allows the system to focus on semantically consistent regions for matching, reducing the impact of environmental changes in non-critical areas and improving both reliability and accuracy under varying conditions.
Solution Approach 2:
The patent changes the parameter space by introducing semantic information as an additional dimension for comparison. Instead of only comparing geometric or photometric parameters, the system now compares semantic labels alongside traditional image features. This parameter expansion makes the localization more robust to environmental changes while maintaining accuracy.
2Ease of manufacture
If image similarity retrieval is used for pose estimation, then the method is simple to implement, but the localization accuracy is low due to reduced structural overlaps between images
Solution Approach 1:
The patent introduces semantic information as an intermediary element that mediates between image similarity retrieval and pose estimation. By incorporating semantic labels as an additional comparison dimension, the system enhances the reliability of matching without significantly complicating the overall workflow, thus improving accuracy while maintaining relative simplicity.
3Measurement precision
If a learning model is trained for each specific scene, then pose estimation can be performed using the model, but the method lacks generality when applied to large scenes
Solution Approach 1:
The patent enhances the universality of the localization system by introducing semantic information that is scene-agnostic. Semantic labels such as sky, building, road, and vegetation can be applied across different scenes and environments, making the system more adaptable and generalizable without requiring scene-specific retraining, thus improving both versatility and maintaining accuracy.
Data Source
AI summary
The present disclosure provides a visual localization method and apparatus based on a semantic error image. The method includes: performing feature extraction for a target image, and obtaining at least one matching pair by performing feature matching for each extracted feature point and each three-dimensional point of a constructed three-dimensional scene model; obtaining a two-dimensional semantic image of the target image by performing semantic segmentation for the target image; and determining semantic information of each matching pair according to semantic information of each pixel of the two-dimensional semantic image; constructing a hypothesized pose pool including at least one hypothesized pose according to at least one matching pair; for each hypothesized pose, constructing a reprojection error image and a semantic error image; determining a hypothesized pose with a minimum reprojection error and a minimum semantic error as a pose estimation according to the reprojection error image and the semantic error image of each hypothesized pose. Optimal pose screening is performed using the semantic error image constructed based on a semantic error, so as to achieve good localization effect even in a case of significant change of a scene.

