Learned Correspondence Masks for Cross-Perspective Image Registration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing location-based services face challenges in maximizing the use of image sources with varying accuracy levels, such as overhead and street-level imagery, leading to reduced positioning accuracy due to sparse and computationally difficult feature matching across different perspectives.
Innovation Solution
A machine learning-based approach using deep neural networks to generate correspondence masks and implicit coordinate transforms between images from different viewpoints, enabling accurate alignment and registration of imagery with varying perspectives.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional feature matching methods are used to register images from different perspectives, then the process is computationally difficult and time-consuming, but the accuracy is limited due to sparse feature correspondences
Solution Approach 1:
The patent replaces traditional mechanical feature matching algorithms with a machine learning-based deep neural network model. The neural network is trained to predict correspondence masks between images from different perspectives, substituting complex computational geometry operations with learned patterns from training data. This approach achieves high positioning accuracy while reducing computational complexity compared to traditional methods.
Solution Approach 2:
The patent changes the approach from explicit geometric parameter estimation to implicit coordinate transforms learned through neural networks. Instead of calculating transformations based on detected features and their geometric relationships, the system uses a neural network to directly map coordinates between different perspectives, achieving more accurate and efficient registration.
2Area of stationary object
If images from multiple sources with varying accuracy levels are used, then the coverage area increases, but the overall positioning accuracy decreases due to mixing high and low accuracy sources
Solution Approach 1:
The patent changes the registration approach by using a neural network that learns to handle varying accuracy levels implicitly. The model is trained on diverse image pairs and automatically adapts to the quality and characteristics of input images, allowing accurate registration across different data sources without manual intervention or quality filtering.
Solution Approach 2:
The neural network-based registration system provides universal functionality to handle multiple image sources with different accuracy levels simultaneously. The single model can process both high-accuracy overhead images and lower-accuracy street-level images, as well as images from various other perspectives, maintaining performance across diverse inputs without requiring separate processing pipelines.
Data Source
AI summary
An approach is provided for machine learning-based registration of imagery with different perspectives. The approach, for example, involves retrieving a first training image and a second training image. The first training image depicts a geographic area from a first perspective and the second training image depicts the geographic area from a second perspective. The approach also involves initiating a labeling of one or more ground truth correspondence masks between the first training image and the second training image. The one or more ground truth correspondence masks denote an image region of the first training image that matches a corresponding image region of the second training image or vice versa. The approach further involves using the one or more ground truth correspondence masks to train a machine learning model to determine one or more predicted correspondence masks between a first input image and a second input image.


