Image Correspondence Learning With Pyramidal Matching Priors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing learning-based methods for extracting local image descriptors require annotated training data sets with pixel-level correspondences, limiting the potential of available image pairs and restricting the quality of training results.
Innovation Solution
A method for unsupervised learning of local image descriptors using a pyramidal structure to enforce local consistency and uniqueness priors, allowing the neural network to learn optimal descriptors without supervision, enabling the identification of corresponding portions in images with differing parameters.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If supervised learning approaches are used to extract local image descriptors, then the quality of identified correspondences is improved, but the requirement for large amounts of annotated training data increases
Solution Approach 1:
The system performs self-supervised learning by automatically generating training labels from the image data itself through consistency regularization across augmented views, eliminating the need for external annotated training data while maintaining high correspondence quality
Solution Approach 2:
The method pre-computes image augmentations and their transformations to generate pseudo-labels before the actual learning process, allowing the model to be trained on unlabeled data with pre-generated supervision signals
2Measurement precision
If ground-truth correspondences are required for training, then the accuracy of local image descriptors is improved, but the limitation of available training data increases
Solution Approach 1:
The system generates its own training labels by computing consistency relationships between augmented views of images, allowing any image pair to be used for training without requiring pre-existing ground-truth correspondences
Solution Approach 2:
The training framework can process any image pair regardless of whether ground-truth correspondences are available, making the approach universally applicable to both labeled and unlabeled data for descriptor learning
Data Source
AI summary
A computer-implemented method includes: obtaining a pair of images depicting a same scene, the pair of images including a first image with a first pixel grid and a second image with a second pixel grid, the first pixel grid different than the second pixel grid; by a neural network module having a first set of parameters: generating a first feature map based on the first image; and generating a second feature map based on the second image; determining a first correlation volume based on the first and second feature maps; iteratively determining a second correlation volume based on the first correlation volume; determining a loss for the first and second feature maps based on the second correlation volume; generating a second set of the parameters based on minimizing a loss function using the loss; and updating the neural network module to include the second set of parameters.


