A feature matching method based on the idea of iterative matching

CN117036744BActive Publication Date: 2025-07-18BEIJING UNIV OF POSTS & TELECOMM +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310416628.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-09-16
Filing Date
2023-04-18
Publication Date
2025-07-18
Estimated Expiration
2043-04-18

AI Technical Summary

Technical Problem

[0004]本发明提出了一种基于迭代匹配思路的特征匹配方法,要解决的技术问题是在复杂多变的场景下面对来自于大尺度的视角变化、光照变化、模糊、重复的纹理信息,甚至是低纹理信息的挑战,通过迭代匹配精准控制基于transformer的注意力机制的感受野,从而保证特征匹配的鲁棒性以及良好的实时性

Benefits of technology

[0033] A feature matching method based on an iterative matching idea of the present invention ensures that the matching quality is not affected while greatly reducing the computational overhead of the attention mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117036744B_ABST
    Figure CN117036744B_ABST
Patent Text Reader

Abstract

The present invention proposes a feature matching method based on the idea of iterative matching, belonging to the field of visual image processing. Specifically, first, two original images to be matched are subjected to feature extraction to obtain feature points and descriptors. Then, the matching scores between the feature points are calculated using the descriptors to construct a similarity score matrix, and the Sinkhorn algorithm is used to optimize and obtain a matching relationship distribution matrix. Next, the matching relationship probability distribution is obtained from the matching relationship distribution matrix, and through the NMS non-maximum suppression method, groups of non-adjacent matching feature points are selected and their probabilities are marked as 0. Finally, it is judged whether the number of matching feature point pairs meets the requirements. If so, the matching result is output; otherwise, new descriptors are regenerated for each feature point, and the similarity score matrix is reconstructed. By reducing the non-maximum suppression radius, new matching feature points are obtained. The present invention greatly reduces the computational overhead of the attention mechanism while ensuring that the matching quality is not affected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of visual image processing, and specifically relates to a feature matching method based on the idea of iterative matching. Background Art

[0002] Feature point-based image matching is of great significance to computer vision. In the fields of image stitching, 3D reconstruction, and visual positioning, feature matching is required to determine the spatial position relationship between images so as to achieve the purposes of stitching, reconstruction, and positioning. Existing traditional matching methods cannot cope with challenges from large-scale perspective changes, illumination changes, blurring, repetitive texture information, or even low-texture information.

[0003] Existing deep learning-based feature matching methods are mainly divided into two categories: one type aims to solve the existing challenges in feature matching by extracting more robust descriptors, such as hardnet and superpoint; the other type is to design a deep neural network that directly takes the feature matching relationship as the network output, such as superglue. The former can only extract limited information around feature points and cannot obtain global information and information of corresponding matching images, so it is impossible to obtain a truly robust and accurate feature. Although the latter can obtain more feature information through the transformer attention mechanism, since this process is a broad attention process (feature points need to pay attention to all feature points), it inevitably requires deep iteration (which often means a large computational overhead) to gradually converge the attention (screen out useful information) in order to obtain a good matching result. Summary of the Invention

[0004] The present invention proposes a feature matching method based on the idea of iterative matching. The technical problem to be solved is to face challenges from large-scale perspective changes, illumination changes, blurring, repetitive texture information, or even low-texture information in complex and variable scenarios, and precisely control the receptive field of the transformer-based attention mechanism through iterative matching, so as to ensure the robustness and good real-time performance of feature matching.

[0005] The feature matching method based on the idea of iterative matching is specifically as follows:

[0006] Step 1: Undistort two original RGB images to be matched and convert them into grayscale images, and input them into the feature extraction network superpoint to obtain each image's feature points and corresponding descriptors;

[0007] Feature points are a type of pixel feature that is insensitive to illumination changes and perspective changes. Descriptors are vectors composed of 256-dimensional floating-point numbers, which are used to describe the feature points and their surrounding texture features.

[0008] Step 2: Calculate the matching score and non-matching score between the feature points of two images using descriptors, and construct a similarity score matrix;

[0009] The matching score P between feature points i,j and the non-matching score of feature points The calculation formulas are as follows:

[0010]

[0011]

[0012] where f i represents the descriptor corresponding to feature point i, and f j is the descriptor corresponding to feature point j; A and B are the sets of feature points corresponding to the original RGB images being matched respectively; <> represents the inner product operation of descriptors; θ is a learnable parameter generated by training and fitting; R is a real number; N A , N B are the numbers of feature points in the sets of feature points A and B respectively.

[0013] The calculation formula for the similarity score matrix is:

[0014] Step 3: Optimize the similarity score matrix using the Sinkhorn algorithm to obtain a matching relationship distribution matrix;

[0015] The matching relationship distribution matrix is a probability distribution matrix, reflecting the probability of matching and non-matching between feature point pairs; the matching probability that satisfies the following constraint conditions is regarded as a successful match:

[0016]

[0017] Form a matching relationship distribution matrix with all the matching probabilities of successful matches;

[0018] Step 4: From the matching relationship distribution matrix, obtain the probability distribution of the matching relationship between each group of feature points, and select the non-adjacent matching feature points in each group through the NMS non-maximum suppression method, and mark their probabilities as 0;

[0019] Feature points that satisfy the following conditions belong to non-adjacent matching feature points:

[0020] P i1,j1 = 0 (if P i1,j1 < P i2,j2 && distance(P i1,j1 P i2,j2 ) < r)

[0021] Pi1,j1 represents the matching probability between feature points i1 and j1; P i2,j2 represents the matching probability between feature points i2 and j2;

[0022] i1, i2 ∈ A, j1, j2 ∈ B; distance() calculates the minimum value of the pixel distances of two groups of feature points in two images, and r is the non-maximum suppression radius.

[0023] Step 5: Determine whether the number of matching feature point pairs in the matching relationship probability distribution meets the requirements or reaches the maximum number of iterations. If it meets the requirements, output the matching result; if it does not meet the requirements, go to Step 6;

[0024] Step 6: For each feature point, generate new descriptors based on self-attention and cross-attention of the transformer within their respective receptive fields;

[0025] For feature point i, generate a new descriptor Fi for this feature point i i,new , and the formula is as follows:

[0026] Fi i,new = mlp(Fi i,old | ∑ j∈B Attention i,j * V j )

[0027] Fi i,new represents the new descriptor of feature point i, and Fi i,old represents the old descriptor of feature point i, and (·|·) represents the vector concatenation operation;

[0028] V j is obtained by mapping the feature vector fi j through the multi-layer perceptron mlp and is used for feature vector fusion; Attention i,j represents the attention weight and is calculated as follows:

[0029]

[0030] q i represents the attention mechanism based on the transformer, which maps the old descriptor corresponding to feature point i through the multi-layer perceptron into a q vector; k j represents the attention mechanism based on the transformer, which maps the old descriptor corresponding to feature point j through the multi-layer perceptron into a k vector;

[0031] Step 7: Return to Step 2, construct a similarity score matrix from the new descriptors generated by all feature points, obtain a matching relationship distribution matrix, reduce the non-maximum suppression radius, and obtain new matching feature points.

[0032] The beneficial effects of the present invention are as follows:

[0033] A feature matching method based on an iterative matching idea of the present invention ensures that the matching quality is not affected while greatly reducing the computational overhead of the attention mechanism. Description of the Drawings

[0034] Figure 1 It is a flowchart of a feature matching method based on an iterative matching idea of the present invention;

[0035] Figure 2 It is a construction diagram of an embodiment of a feature matching method based on an iterative matching idea of the present invention. Detailed Embodiments

[0036] For the convenience of those of ordinary skill in the art to understand and implement the present invention, the following further describes the present invention in detail and in-depth with reference to the accompanying drawings.

[0037] Step 1: Undistort two original RGB images to be matched and convert them into grayscale images, and input them into the feature extraction network superpoint to obtain the image feature points and corresponding descriptors of each image;

[0038] Superpoint is a fully convolutional deep neural network that can calculate the descriptors and feature points of an image in parallel. Feature points are a type of pixel feature that is insensitive to illumination changes and perspective changes. A descriptor is a vector composed of 256-dimensional floating-point numbers, which is used to describe the feature points and their surrounding texture features.

[0039] Step 2: Use the descriptors to calculate the matching scores and non-matching scores between the feature points of the two images, and construct a similarity score matrix;

[0040] Since the matching relationship between feature points should be that there is only a unique corresponding feature point or there is no corresponding feature point (in this case, the point is assigned to the dustbin - trash can), this constitutes the uniqueness constraint of the matching.

[0041] ∑ j∈N2 P i,j =1,i∈N1&&P i,j =0(i≠j)&&P i,j =1(i==j)

[0042] i≠j indicates that there is no matching relationship between two feature points, i = j indicates that there is a matching relationship between two feature points or j is a dustbin; there is no corresponding matching feature point for i; N1 and N2 represent two sets of feature points from two images, and a dustbin reflecting the non - existent matching feature points.

[0043] The matching score P between feature points i,j And the non - matching score of feature points The calculation formula is as follows:

[0044]

[0045]

[0046] Among them, f i represents the descriptor corresponding to feature point i, and f j is the descriptor corresponding to feature point j; A and B are respectively the sets of feature points corresponding to the original RGB images of the match; <> is the inner - product operation of descriptors; θ is a learnable parameter generated by training and fitting; R is a real number; N A , N B are respectively the numbers of feature points in the feature - point sets A and B.

[0047] The calculation formula for the similarity - score matrix is:

[0048] Step 3: Use the Sinkhorn algorithm to optimize the similarity - score matrix and obtain the matching - relationship distribution matrix;

[0049] The Sinkhorn algorithm is a differentiable optimization algorithm for solving the Optimal Transport Problem. The Optimal Transport Problem is to find an optimal mapping between two probability distributions such that one distribution is mapped to the other with the minimum total cost.

[0050] The Sinkhorn algorithm is based on the method of iterative weight balancing. Its basic idea is to minimize the cost function while ensuring that the constraint conditions are met, so as to obtain the optimal transport plan;

[0051] Using the similarity - score matrix and the matching - uniqueness constraint, the problem is transformed into an optimal transport problem: The aim is to solve how to convert one probability distribution into another with the minimum cost. Regarding the distribution of the score matrix as the original distribution, and the matching - uniqueness constraint constitutes a probability distribution of the true matching relationship. The scores of the score matrix reflect the transport cost. Then, with the help of the Sinkhorn algorithm, a matching - relationship distribution matrix that satisfies both the matching - uniqueness constraint and is similar to the distribution of the score matrix is obtained.

[0052] The matching relationship distribution matrix is a probability distribution matrix, which reflects the probability of matching and non-matching between feature point pairs; a reasonable threshold is induced through machine learning to determine whether the matching probability can be adopted.

[0053] The matching probability that satisfies the following constraints is regarded as a successful match:

[0054]

[0055] All the matching probabilities of successful matches are formed into a matching relationship distribution matrix;

[0056] Step Four: From the matching relationship distribution matrix, obtain the matching relationship probability distribution between each group of feature points. Through the NMS (Non-Maximum Suppression) method, select the non-adjacent matching feature points in each group and mark their probabilities as 0;

[0057] Feature points that satisfy the following conditions belong to non-adjacent matching feature points:

[0058] P i1,j1 = 0 (if P i1,j1 < P i2,j2 && distance(P i1,j1 P i2,j2 ) < r)

[0059] P i1,j1 represents the matching probability between feature points i1, j1; P i2,j2 represents the matching probability between feature points i2, j2;

[0060] i1, i2 ∈ A, j1, j2 ∈ B; distance() calculates the minimum value of the pixel distances of two groups of feature points in two images, and r is the non-maximum suppression radius.

[0061] After that, the remaining feature points are divided into the receptive fields of different feature point pairs.

[0062] Step Five: Judge whether the number of matching feature point pairs in the matching relationship probability distribution meets the requirements or reaches the maximum number of iterations. If it meets the requirements, output the matching result; if it does not meet the requirements, go to Step Six;

[0063] Step Six: For each feature point, generate new descriptors within its respective receptive field based on self-attention and cross-attention of the transformer;

[0064] Self-attention is an attention mechanism among feature points within the same image; cross-attention is an attention mechanism between two sets of feature points that match image pairs; the attention mechanism based on the transformer first maps descriptors into q vectors, k vectors, and v vectors through a multi-layer perceptron. Then, the similarity between the q vector and all k vectors is calculated and normalized to obtain the attention weights for each descriptor.

[0065] Specifically:

[0066] For feature point i, a new descriptor F for this feature point i is generated i,new , and the formula is as follows:

[0067] F i,new = mlp(F i,old |∑ j∈B Attention i,j *V j )

[0068] F i,new represents the new descriptor of feature point i, F i,old represents the old descriptor of feature point i, (·|·) represents the vector concatenation operation;

[0069] V j is obtained by mapping the feature vector f j through the multi-layer perceptron mlp and is used for feature vector fusion; Attention i,j represents the attention weight and is calculated as follows:

[0070]

[0071] q i represents the attention mechanism based on the transformer, which maps the old descriptor corresponding to feature point i into a q vector through a multi-layer perceptron; k j represents the attention mechanism based on the transformer, which maps the old descriptor corresponding to feature point j into a k vector;

[0072] The attention mechanism updates the descriptor for subsequent new rounds of matching point screening or finally calculates the final matching relationship distribution matrix.

[0073] Step seven: Return to step two, construct a similarity score matrix from the new descriptors generated for all feature points, obtain the matching relationship distribution matrix, reduce the non-maximum suppression radius, and obtain new matching feature points.

[0074] For example Figure 2In the shown example, the specific experimental data obtained are as follows:

[0075] The test of the matching accuracy rate on the standard dataset Hpatches, and the results are shown in Table 1:

[0076]

[0077] The real-time test of the algorithm, and the results are shown in Table 2:

[0078] Method FPS superpoint 100 Our Method 60 Superpoint+Superglue 40

Claims

1. A feature matching method based on the idea of iterative matching, characterized in that The specific steps are as follows: Step 1: Undistort the two original RGB images to be matched and convert them into grayscale images, and input them into the feature extraction network superpoint to obtain the feature points and corresponding descriptors of each image; Step 2: Calculate the matching scores and non-matching scores between the feature points of the two images using the descriptors, and construct a similarity score matrix; The matching score P between feature points i,j and the non-matching score of feature points The calculation formula is as follows: Among them, f i represents the descriptor corresponding to feature point i, and f j is the descriptor corresponding to feature point j; A and B are respectively the sets of feature points corresponding to the original RGB images of the match; <> represents the inner product operation of the descriptors; θ is a learnable parameter generated by training and fitting; R is a real number; N A , N B are respectively the numbers of feature points in the sets of feature points A and B; The calculation formula for the similarity score matrix is as follows: Step 3: Optimize the similarity score matrix using the Sinkhorn algorithm to obtain a matching relationship distribution matrix; The matching relationship distribution matrix is a probability distribution matrix that reflects the probability of matching and non-matching between feature point pairs; regard the matching probability that satisfies the following constraints as a successful match: Form a matching relationship distribution matrix with all the successful matching probabilities; Step 4: From the matching relationship distribution matrix, obtain the matching relationship probability distribution between each group of feature points, and select the non-adjacent matching feature points in each group through the NMS non-maximum suppression method, and mark their probabilities as 0; Feature points that satisfy the following conditions belong to non-adjacent matching feature points: P i1,j1 = 0 (if P i1,j1 < P i2,j2 && distance(P i1,j1 P i2,j2 ) < r) P i1,j1 represents the matching probability between feature points i1 and j1; P i2,j2 represents the matching probability between feature points i2 and j2; i1,i2∈A,j1,j2∈B; distance() calculates the minimum value of the pixel distances of two groups of feature points in the two images, and r is the non-maximum suppression radius; Step 5: Judge whether the number of matching feature point pairs in the matching relationship probability distribution meets the requirements or reaches the maximum number of iterations. If it meets, output the matching result; if it does not meet, go to Step 6; Step 6: For each feature point, generate new descriptors based on self-attention and cross-attention of the transformer within its respective receptive field; Step 7: Return to Step 2, construct a similarity score matrix with the new descriptors generated by all feature points, obtain a matching relationship distribution matrix, reduce the non-maximum suppression radius, and obtain new matching feature points.

2. The feature matching method based on the iterative matching idea according to claim 1, characterized in that The feature points are a type of pixel feature that is insensitive to illumination changes and perspective changes; The descriptor is a vector composed of 256-dimensional floating-point numbers, which is used to describe the feature points and their surrounding texture features.

3. The feature matching method based on the iterative matching idea as described in claim 1, wherein In the sixth step, for feature point i, a new descriptor F of the feature point i is generated i,new , and the formula is as follows: F i,new = mlp(F i,old |∑ j∈B Attention i,j *V j ) F i,new represents the new descriptor of feature point i, F i,old represents the old descriptor of feature point i, and (·|·) represents the vector concatenation operation; V j is obtained by mapping the eigenvector f j through the multi-layer perceptron mlp and is used for eigenvector fusion; Attention i,j represents the attention weight and is calculated as follows: q i represents the attention mechanism based on the transformer, mapping the old descriptor corresponding to feature point i to the q vector through a multi-layer perceptron; k j represents the attention mechanism based on the transformer, mapping the old descriptor corresponding to feature point j to the k vector.

Citation Information

Patent Citations

  • Image matching method for similar image retrieval

    CN114491122A

  • Feature matching method based on attention mechanism and neighborhood consistency

    CN114758152A