Feature matching method and system based on image block comparison score
Through the feature matching method based on image block comparison score, image blocks are used as visual information, combined with Transformer and ResNet blocks for feature enhancement, the outlier problem of feature matching in complex scenarios is solved, and the accuracy and robustness of feature matching is improved. It is suitable for tasks such as pose estimation, point cloud registration and image stitching.
Patent Information
- Application Number
- CN202510711454.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-05
AI Technical Summary
The existing feature matching methods perform poorly in complex scenarios. The initial matching pairs contain a large number of outliers, which seriously affects the performance of downstream tasks. The learning-based method ignores the visual information provided by the original image.
The feature matching method based on image block comparison score is adopted, and the network is trained by calculating matching pair classification loss and intrinsic matrix regression loss, using image blocks as visual information, combining Transformer and ResNet blocks for feature enhancement, and capturing global context information through self-attention operations, and pruning high-performance matching pairs.
It effectively removes outliers, improves the accuracy and robustness of feature matching, and is suitable for pose estimation, point cloud registration, image stitching and motion segmentation.
Smart Images

Figure CN120599301A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a feature matching method and system based on image block comparison scoring. Background Art
[0002] Obtaining reliable matching pairs (i.e., feature matches) between two images is a key step in many computer vision tasks, such as image registration, image fusion, and stereo matching. Current mainstream feature matching methods follow a three-stage process, including detecting salient keypoints and extracting descriptors, generating an initial set of matching pairs, and removing incorrect matching pairs (outliers) to construct a reliable set of matching pairs. However, commonly used feature detection and description methods (such as SIFT and SuperPoint) often perform poorly in complex scenes involving repetitive textures, illumination variations, and large changes in viewpoint. In these cases, the initial set of matching pairs often contains a large number of outliers, which severely degrades the performance of downstream tasks. Therefore, removing outliers and obtaining correct matching pairs (inliers) has become a key research topic.
[0003] Outlier removal methods can be roughly divided into traditional methods and learning-based methods. Among traditional methods, RANSAC adopts an iterative sampling strategy for outlier removal. However, traditional methods show performance limitations in sets of matching pairs with a high proportion of outliers. To address these limitations, learning-based methods have emerged to achieve more efficient and reliable outlier removal. PointCN is regarded as a pioneering work, which formulates outlier removal as a joint task of essential matrix regression and binary classification. It is based on the multi-layer perceptron (MLP) framework to predict the inlier probability of matching pairs. CLNet combines graph convolutional networks (GCN) with MLP and introduces a progressive pruning framework. Based on CLNet, NCMNet enhances feature representation by aggregating neighbor information in coordinate space, feature space and global graph space.
[0004] Although learning-based methods have made significant contributions to the development of this field, they mainly operate on sets of matching pairs (i.e., spatial information of keypoints), while ignoring the visual information that the original image can provide. This omission may limit their performance in complex scenes. VSFormer combines visual cues and spatial information as scene perception before pruning, specifically by embedding visual cues into matching pairs for pruning. However, directly using the entire image may introduce irrelevant information (such as false features) because it performs feature-based matching rather than region-based matching. To address this issue, we propose to use image patches containing matching pairs instead of the entire image as visual information. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a feature matching method and system based on image block contrast scoring, which guides network training by calculating matching pair classification loss and essential matrix regression loss, thereby obtaining a high-performance feature matching model.
[0006] To achieve the above object, the present invention adopts the following technical solution: a feature matching method based on image block comparison scoring, comprising the following steps:
[0007] Step 1: Given a set of image pairs, use the SIFT algorithm to extract the key points of the two images and generate their corresponding descriptors;
[0008] Step 2: Cut out image blocks of different sizes of each matching pair from the original image using the key point coordinates of the matching pair;
[0009] Step 3: First, select the matching pairs with small differences, that is, the matching pairs with high PCScores. Then, a vanilla Transformer performs a cross-attention operation to allow the selected matching pairs to guide the overall matching pairs. Finally, a ResNet block is used to further integrate and enhance the feature representations of the guided matching pairs.
[0010] Step 4: First, construct a neighbor graph based on the feature representation F″ obtained after sampling and perform a cross-attention operation to capture local context information;
[0011] Step 5: Enhance the current feature representation through a three-branch feature enhancer. The results obtained from the three branches are integrated through a connection operation and the dimension reduction operation is performed through an encoder.
[0012] Step 6: Perform a self-attention operation to capture global context information, and use the global consistency perception matrix GCWM as the weight matrix of the self-attention operation.
[0013] In a preferred embodiment, in step 1, an initial matching pair set S = [s1, s2, ..., s N ]∈R N×4 , s i =(x i ,y i ,x′ i ,y′ i ), the set S contains some wrong matching pairs; a multi-layer perceptron MLP is used to elevate the matching pairs to a high-dimensional space to obtain the initial feature representation F.
[0014] In a preferred embodiment, in step 2, the negative absolute value of the difference between the image blocks is calculated and reduced to one dimension through a plurality of multi-layer perceptrons (MLPs), thereby obtaining PCScores to evaluate the difference between the image blocks of each matching pair;
[0015] PCScores = MLPs(-|p1-p2|), (Formula 1)
[0016] Here, p1 and p2 represent image patches captured from the two images, and MLPs represents multiple MLP operations. This process is performed three times, i.e., three different image patch sizes are selected, and the final PCScores are obtained by summing the PCScores calculated at different sizes. The image patches of the matching pairs with high PCScores have smaller differences, that is, these matching pairs are more likely to be inliers.
[0017] In a preferred embodiment, in step 3,
[0018] F′=sampling(F,topk(PCScores)), (Formula 2)
[0019] F″=ResNets(TF(F,F′)), (Formula 3)
[0020] Where F′ represents the feature representation obtained after sampling, F represents the initial feature, TF represents vanilla Transformer, and ResNets represents multiple ResNet block operations; topk is a filtering function that returns the features corresponding to the top k matching pairs in PCScores.
[0021] In a preferred embodiment, in step 4, a shallow KNN neighbor graph and a deep KNN neighbor graph are first constructed using the feature representation F″ obtained after sampling, and then the shallow KNN neighbor graph is used as Q, and the deep KNN neighbor graph is used as K and V, and a cross attention operation is performed; finally, the context information is aggregated through the circular convolution and ReLU activation function and the residual summation is performed on the input feature representation to capture the local context information;
[0022] G att =MLP(SF(Q⊙K T )⊙V), (Formula 4)
[0023] F agg =F″+Re(AnnC(G att )), (Formula 5)
[0024] Among them G att It represents the result obtained after cross attention of two graphs, SF represents the softmax function, KT represents the transpose of K, ⊙ represents a Hadamard product operation; Re represents the ReLU activation function, F agg It represents the feature representation obtained by aggregating local context information, and AnnC represents the ring convolution.
[0025] In a preferred embodiment, in step 5, the current feature representation is enhanced by a feature enhancer with a three-branch structure, which mainly includes an OABlock and two squeeze-excitation operations on different dimensions; the results obtained by the three branches are integrated through a concatenation operation and reduced to 128 dimensions through an encoder;
[0026] F sd =PC(F agg )*MLP(AP(PC(F agg ),sd)+MP(PC(F agg ),sd)), (Formula 6)
[0027] F cd =PC(F agg )*MLP(AP(PC(F agg ),cd)+MP(PC(F agg ),cd)), (Formula 7)
[0028] F′ agg =cat(F sd ,F cd ,OABlock(F agg )), (Formula 8)
[0029] Among them, F sd It represents the feature representation obtained after the squeeze excitation operation in the spatial dimension, sd represents the spatial dimension, F cd represents the feature representation obtained after the squeeze excitation operation in the channel dimension, cd represents the channel dimension, (,*) represents the operation along the dimension specified by *; PC represents PointCN, AP represents average pooling, and MP represents maximum pooling; F′ agg It represents the feature representation after feature enhancement, and cat represents the connection operation.
[0030] In a preferred embodiment, in step 6, for two matching pairs S i =(u i ,v i ) and S j =(u j ,v j ), their length consistency is calculated as follows:
[0031] LCi,j =1-Sigmoid(|||u i -u j ||-||v i -v j |||) (Formula 9)
[0032] The global consistency perception matrix is constructed as follows:
[0033]
[0034] The feature representation before and after the self-attention operation is concatenated and the dimension is reduced to 128 dimensions through a ResNet Block to obtain the output feature representation.
[0035] In a preferred embodiment, steps 3 to 6 are regarded as a matching pair pruning module, which is iteratively executed twice in each round of network training. After each round of training, the matching pair classification loss and the essential matrix regression loss are calculated to guide the network training, thereby obtaining a high-performance feature matching model.
[0036] The present invention also provides a feature matching system based on image block contrast scoring, which runs the feature matching method based on image block contrast scoring.
[0037] Compared with the existing technology, the present invention has the following advantages: it can effectively remove outliers. The method proposed in the present invention can be applied to multiple fields such as posture estimation, point cloud registration, image stitching and motion segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 A flowchart of a preferred embodiment of the present invention;
[0039] Figure 2 This is a comparison chart of the effects of the preferred embodiment of the present invention on the YFCC100M dataset and the SUN3D dataset with other methods, where GTPCS is the method proposed in this patent. DETAILED DESCRIPTION
[0040] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0041] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present application belongs.
[0042] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form, and it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or their combinations.
[0043] The feature matching method based on image block comparison score of the present invention refers to Figure 1-2 , including the following steps:
[0044] A. Given a set of image pairs, use the SIFT algorithm to extract the key points of the two images and generate their corresponding descriptors. According to the similarity between the descriptors, an initial matching pair set S = [s1, s2, ..., s N ]∈R N×4 , s i =(x i ,y i ,x′ i ,y′ i ), which contains some incorrect matching pairs. A multi-layer perceptron (MLP) is used to elevate the matching pairs to a high-dimensional space (128 dimensions) to obtain the initial feature representation F.
[0045] B. Using the keypoint coordinates of the matching pairs, we intercept image patches of different sizes from the original image for each matching pair, calculate the negative absolute value of the difference between the image patches, and reduce it to one dimension through several MLPs to obtain PCScores to evaluate the difference between the image patches of each matching pair.
[0046] PCScores = MLPs(-|p1-p2|), (Formula 1)
[0047] Where p1 and p2 represent image patches taken from the two images, and MLPs represents multiple MLP operations. This process is performed three times, selecting three different image patch sizes. The final PCScores are calculated by summing the PCScores calculated at different sizes. Pairs with higher PCScores have smaller differences between the image patches, indicating that these pairs are more likely to be inliers.
[0048] C. First, select the matching pairs with the smallest differences, that is, those with higher PCScores. Then, a vanilla Transformer performs a cross-attention operation, allowing the selected matching pairs to guide the overall matching pairs. Finally, a ResNet block further integrates and enhances the feature representations of these guided matching pairs.
[0049] F′=sampling(F,topk(PCScores)), (Formula 2)
[0050] F″=ResNets(TF(F,F′)), (Formula 3)
[0051] Where F′ represents the feature representation obtained after sampling, TF represents the vanilla Transformer, and ResNets represents multiple ResNetblock operations.
[0052] D. First, construct a shallow KNN neighbor graph and a deep KNN neighbor graph through F″. Then, use the shallow KNN neighbor graph as Q and the deep KNN neighbor graph as K and V to perform cross attention operation. Finally, aggregate the context information through the ring convolution and ReLU activation function and perform residual summation on the input feature representation to capture local context information.
[0053] G att =MLP(SF(Q⊙K T )⊙V), (Formula 4)
[0054] F agg =F″+Re(AnnC(G att )), (Formula 5)
[0055] Where SF represents the softmax function, ⊙ represents a Hadamard product operation, Re represents the ReLU activation function, and AnnC represents the circular convolution.
[0056] E. Enhance the current feature representation through a three-branch feature enhancer, which mainly includes an OABlock and two squeeze-excitation operations on different dimensions. The results obtained from the three branches are integrated through a concatenation operation and the dimension is reduced to 128 dimensions through an encoder.
[0057] F sd =PC(F agg )*MLP(AP(PC(F agg ),sd)+MP(PC(F agg ),sd)), (Formula 6)
[0058] F cd =PC(F agg )*MLP(AP(PC(F agg ),cd)+MP(PC(F agg ),cd)), (Formula 7)
[0059] F′ agg=cat(F sd ,F cd ,OABlock(F agg )), (Formula 8)
[0060] Where sd represents the spatial dimension, cd represents the channel dimension, and (,*) represents the operation along the dimension specified by *. PC represents PointCN, AP represents average pooling, and MP represents maximum pooling.
[0061] F. Perform self-attention operation to capture global context information and use the global consistency perception matrix (GCWM) as the weight matrix of the self-attention operation. i =(u i ,v i ) and S j =(u j ,v j ), their length consistency is calculated as follows:
[0062] LC i,j =1-Sigmoid(|||u i -u j ||-||v i -v j |||). (Formula 9)
[0063] The global consistency perception matrix is constructed as follows:
[0064]
[0065] The feature representation before and after the self-attention operation is concatenated and the dimensionality is reduced to 128 dimensions through a ResNet Block to obtain the output feature representation.
[0066] Steps C through F are considered a matching pair pruning module, which is iteratively executed twice in each round of network training. After each round of training, the matching pair classification loss and essential matrix regression loss are calculated to guide network training, resulting in a high-performance feature matching model.
Claims
1. Feature matching method based on image block contrast scoring, characterized in that: The following steps are involved: Step 1: Given a set of image pairs, use the SIFT algorithm to extract the key points of the two images and generate their corresponding descriptors; Step 2: Cut out image blocks of different sizes of each matching pair from the original image using the key point coordinates of the matching pair; Step 3: First, select the matching pairs with small differences, that is, the matching pairs with high PCScores; then perform a cross-attention operation through a vanilla Transformer to let the selected matching pairs guide the total matching pairs; finally, use a ResNet block to further integrate and enhance the feature representations of the guided matching pairs; Step 4: First, construct a neighbor graph based on the feature representation F' obtained after sampling and perform a cross-attention operation to capture local context information; Step 5: Enhance the current feature representation through a three-branch feature enhancer. The results obtained from the three branches are integrated through a connection operation and the dimension reduction operation is performed through an encoder. Step 6: Perform a self-attention operation to capture global context information, and use the global consistency perception matrix GCWM as the weight matrix of the self-attention operation.
2. The feature matching method based on image block contrast scoring according to claim 1, characterized in that: In step 1, an initial matching pair set S = [s1, s2, ..., s N ]∈R N×4 , s i =(x i ,y i ,x′ i ,y′ i ), the set S contains some wrong matching pairs; a multi-layer perceptron MLP is used to elevate the matching pairs to a high-dimensional space to obtain the initial feature representation F.
3. The feature matching method based on image block contrast scoring according to claim 1, characterized in that: In step 2, the negative absolute value of the difference between the image blocks is calculated and reduced to one dimension through several multi-layer perceptrons (MLPs), thereby obtaining PCScores to evaluate the difference between the image blocks of each matching pair; PCScores = MLPs(-|p1-p2|), (Formula 1) Here, p1 and p2 represent image patches captured from the two images, and MLPs represents multiple MLP operations. This process is performed three times, i.e., three different image patch sizes are selected, and the final PCScores are obtained by summing the PCScores calculated at different sizes. The image patches of the matching pairs with high PCScores have smaller differences, that is, these matching pairs are more likely to be inliers.
4. The feature matching method based on image block contrast scoring according to claim 1, characterized in that: In the step 3, F′=sampling(F,topk(PCScores)), (Formula 2) F' = ResNets(TF(F,F')), (Formula 3) Where F′ represents the feature representation obtained after sampling, F represents the initial features, TF represents the vanilla Transformer, and ResNets represents multiple ResNet block operations; topk is a filtering function that returns the features corresponding to the top k matching pairs in PCScores.
5. The feature matching method based on image block contrast scoring according to claim 1, characterized in that: In step 4, a shallow KNN neighbor graph and a deep KNN neighbor graph are first constructed using the feature representation F' obtained after sampling. Then, the shallow KNN neighbor graph is used as Q, and the deep KNN neighbor graph is used as K and V for cross attention operation. Finally, the context information is aggregated by the circular convolution and ReLU activation function, and the residual summation of the input feature representation is performed to capture the local context information. G att =MLP(SF(Q⊙K T )⊙V), (Formula 4) F agg =F"+Re(AnnC(G att )), (Formula 5) Among them G att It represents the result obtained after cross attention of two graphs, SF represents the softmax function, K T represents the transpose of K, ⊙ represents a Hadamard product operation; Re represents the ReLU activation function, F agg It represents the feature representation obtained by aggregating local context information, and AnnC represents the ring convolution.
6. The feature matching method based on image block contrast scoring according to claim 1, characterized in that: In step 5, the current feature representation is enhanced by a feature enhancer with a three-branch structure. The feature enhancer mainly includes an OABlock and a squeeze-excitation operation on two different dimensions. The results obtained by the three branches are integrated through a concatenation operation and a dimensionality reduction operation is performed through an encoder to reduce the dimension to 128 dimensions. F sd =PC(F agg )*MLP(AP(PC(F agg ),sd)+MP(PC(F agg ),sd)), (Formula 6) F cd =PC(F agg )*MLP(AP(PC(F agg ),cd)+MP(PC(F agg ),cd)), (Formula 7) F′ agg =cat(F sd ,F cd ,OABlock(F agg )), (Formula 8) Among them, F sd It represents the feature representation obtained after the squeeze excitation operation in the spatial dimension, sd represents the spatial dimension, F cd represents the feature representation obtained after the squeeze excitation operation in the channel dimension, cd represents the channel dimension, (,*) represents the operation along the dimension specified by *; PC represents PointCN, AP represents average pooling, and MP represents maximum pooling; F′ agg It represents the feature representation after feature enhancement, and cat represents the connection operation.
7. The feature matching method based on image block contrast scoring according to claim 1, characterized in that: In step 6, for two matching pairs S i =(u i ,v i ) and S j =(u j ,v j ), their length consistency is calculated as follows: LC i,j =1-Sigmoid(|||u i -u j ||-||v i -v j |||) (Formula 9) The global consistency perception matrix is constructed as follows: The feature representation before and after the self-attention operation is concatenated and the dimension is reduced to 128 dimensions through a ResNet Block to obtain the output feature representation.
8. The feature matching method based on image block contrast scoring according to any one of claims 1 to 7, characterized in that: Steps 3 to 6 are considered a matching pair pruning module. In each round of network training, the matching pair pruning module is iteratively executed twice. After each round of training, the matching pair classification loss and essential matrix regression loss are calculated to guide network training, thereby obtaining a high-performance feature matching model.
9. Feature matching system based on image block contrast scoring, characterized in that: The feature matching method based on image block contrast scoring described in any one of claims 1 to 7 is run.