Cross-View Image Localization Using Ground-Plane Feature Aggregation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image-based localization methods face challenges in accurately aligning ground-based and aerial images due to significant viewpoint differences, leading to domain gaps that compromise localization accuracy, particularly in handling off-ground features and occlusions, and relying on image-level feature matching restricts precision.
Innovation Solution
A method and system that aggregates elevated pixels onto a ground plane, generates query feature maps using attention mechanisms, and adjusts computational model parameters for domain and perspective variations to enhance feature alignment and keypoint detection, leveraging off-ground features for improved localization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If geometry-alignment-based methods are used to bridge the domain gap, then feature alignment between ground-based and aerial images is improved, but off-ground features are neglected and visual occlusions cannot be handled effectively
Solution Approach 1:
The patent introduces a vertical dimension by aggregating elevated pixels onto the ground plane, transforming 3D spatial information into 2D ground-plane representations. This dimensional transformation enables the system to capture off-ground features (streetlights, trees, buildings) that traditional 2D-to-2D matching methods overlook, while maintaining compatibility with geometry-alignment-based approaches for on-ground feature matching.
Solution Approach 2:
The patent merges two previously separate feature sets: on-ground features handled by geometry-alignment methods and off-ground features captured through elevated pixel aggregation. By combining these complementary feature types into a unified feature map, the system achieves both precise feature alignment for on-ground pixels and robust handling of occlusions through off-ground landmark detection.
2Adaptability or versatility
If image-level feature matching is used for localization, then the domain gap between ground-based and aerial images is addressed, but localization precision is restricted
Solution Approach 1:
The patent segments the feature matching process into two distinct levels: image-level feature matching for bridging the domain gap between ground-based and aerial views, and pixel-level feature aggregation for achieving precise localization. This hierarchical segmentation allows each level to specialize - image-level methods handle view transformation while pixel-level aggregation delivers sub-meter localization accuracy.
Solution Approach 2:
The patent transitions from image-level feature matching to pixel-level feature aggregation, adding a finer granularity dimension to the localization process. By aggregating elevated pixels onto ground-plane pixels, the system creates detailed pixel-correspondence maps that enable precise pose estimation while maintaining the cross-view matching capabilities provided by image-level features.
3Reliability
If elevated features are utilized for localization, then robustness to road mark degradation and occlusions is improved, but feature alignment complexity increases
Solution Approach 1:
The patent introduces the ground plane as an intermediary representation that mediates between elevated features (streetlights, trees) and on-ground features (road marks, lane lines). By projecting elevated pixels onto their corresponding ground-plane locations, the system creates a unified feature space where diverse feature types can be aligned and matched, managing complexity through this intermediate abstraction layer.
Data Source
AI summary
A method for localising a ground-based query image captured by an imaging device with respect to an aerial image of a physical environment within which the imaging device is disposed, enabling a pose of the imaging device to be determined within the physical environment. The method includes obtaining query features identified in the ground-based query image, identifying elevated features from the query features, generating aggregated pixels by aggregating elevated pixels of the elevated features onto a ground plane of the ground-based query image and generating a query feature map using the aggregated pixels, wherein the query feature map is usable in localising the ground-based query image by mapping the query feature map to an aerial feature map of aerial features within the aerial image.


