Double-view feature matching method based on channel and space joint interaction and bidirectional consensus interaction

By enhancing feature representation through a dual-path attention mechanism and a local-global consensus interaction module, the problem of outlier removal in feature matching in complex scenarios is solved, and matching accuracy is improved. It is suitable for pose estimation, point cloud registration, and image stitching.

CN120635510APending Publication Date: 2025-09-12MINJIANG UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510792947.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

When dealing with feature matching in complex scenarios, existing methods have poor outlier removal effects and insufficient interaction between local and global consensus, which affects matching accuracy.

Method used

A channel-space interaction module with a dual-path attention mechanism is adopted, combined with a local consensus mining module and a global consensus-aware attention module. Feature representation is enhanced through channel adaptive attention and spatial attention, achieving two-way interaction between local and global consensus.

Benefits of technology

It improves the accuracy and precision of feature matching, effectively captures the global information of ordered motion vectors, and is suitable for tasks such as pose estimation, point cloud registration, and image stitching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635510A_ABST
    Figure CN120635510A_ABST
Patent Text Reader

Abstract

The invention relates to a double-view feature matching method based on channel and space joint interaction and bidirectional consensus interaction, and belongs to the field of computer vision. The method comprises the following steps: providing a network backbone-channel-space interaction module, adopting a dual-path attention mechanism framework, and enhancing feature representation capability by integrating channel adaptive attention and space attention, thereby avoiding information loss of a Point CN module caused by independent processing of a spatial position; a local consensus mining module and a global consensus perception attention module are provided, the local consensus mining module jointly models geometric structures and spatial continuity of matching pairs, and local consensus is extracted; and the global consensus perception attention module constructs global consensus by aggregating the high-confidence correct matching pairs, and realizes interaction of local and global consensus by using a cross attention mechanism. The method can effectively capture the global information of the ordered motion vector, and can be applied to multiple fields of attitude estimation, point cloud registration, image splicing, motion segmentation and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and in particular relates to a dual-view feature matching method based on channel and space joint interaction and two-way consensus interaction. Background Art

[0002] In computer vision tasks, establishing accurate point-to-point correspondences between related images (i.e., feature matching) is the basis for tasks such as 3D reconstruction, visual simultaneous localization, and image stitching. Existing methods (such as SIFT

[19] and SuperPoint[6]) can be used to establish an initial correspondence set. However, due to complex factors such as drastic changes in viewpoints, different lighting conditions, motion blur, or repetitive structures in real scenes, the initial correspondence set usually contains a large number of incorrect matches (i.e., outliers). Therefore, to ensure the accuracy of the matching results, it is crucial to design effective methods to identify and remove these outliers while retaining the correct matching pairs (i.e., inliers) as much as possible.

[0003] Outlier removal (i.e., two-view matching pair screening) methods can be divided into two categories: traditional methods and learning-based methods. Among the traditional methods, RANSAC[1] and its improved algorithms are the most widely used, mainly relying on sampling and verification strategies. Although they perform well in simple scenes, their performance decreases significantly as the proportion of outliers increases. With the development of deep learning, learning-based outlier removal methods have emerged. The pioneering work PointCN[2] redefined the outlier removal problem as a combination of inlier-outlier classification and essential matrix regression, using a multi-layer perceptron as the backbone network, capturing the implicit associations between matching pairs through context normalization, and using a weighted eight-point algorithm for camera pose estimation.

[0004] Based on PointCN, CLNet[3] achieved significant performance improvement by mining the consistency of matching pairs. Specifically, CLNet introduced a progressive learning strategy to gradually learn local-global consistency: first, it searches for the nearest neighbor for each matching pair in the feature space and captures local consistency (i.e., local consensus) by aggregating neighborhood information; then, it constructs a global graph connecting all matching pairs and uses graph convolutional networks (GCN)[4] to model global consistency (i.e., global consensus). Based on this progressive strategy, subsequent methods have focused on designing diverse modules to better capture consistency information. For example, NCMNet[5] enhances local consensus representation by searching for the nearest neighbor in the geometry, feature, and graph spaces; GCTNet[6] extracts local consensus using two aggregation methods, maximum pooling and ring convolution, and explores the interaction of different local consensuses; VSFormer[7] simultaneously captures local and global consensus by fusing image and spatial information.

[0005] Although existing methods have made some progress, they still have the following limitations: 1) Current methods mainly use the PointCN module as the feature extraction backbone. Its spatial position independent processing method limits the interaction between the channel dimension and the spatial dimension, resulting in insufficient matching feature representation ability, which in turn affects the screening accuracy; 2) Existing progressive learning strategies achieve one-way transmission from local to global through cascade operations, ignoring the two-way interaction between local and global consensus, and failing to integrate the two into a more cohesive feature representation. Summary of the Invention

[0006] The purpose of the present invention is to provide a dual-view feature matching method based on channel and space joint interaction and two-way consensus interaction, which can effectively capture the global information of ordered motion vectors. The method of the present invention can be applied to multiple fields such as posture estimation, point cloud registration, image stitching and motion segmentation.

[0007] To achieve the above objectives, the technical solution of the present invention is: a dual-view feature matching method based on channel and space joint interaction and two-way consensus interaction, comprising:

[0008] A network backbone, the channel-space interaction module, is proposed. It uses a dual-path attention mechanism framework to enhance feature representation by integrating channel-adaptive attention and spatial attention, thus avoiding information loss caused by independent spatial processing in the PointCN module.

[0009] A local consensus mining module and a global consensus-aware attention module are proposed. The local consensus mining module jointly models the geometric structure and spatial continuity of matching pairs to extract local consensus; the global consensus-aware attention module builds global consensus by aggregating high-confidence correct matching pairs, and uses the cross-attention mechanism to realize the interaction between local and global consensus.

[0010] Furthermore, the channel-space interaction module specifically implements the following functions:

[0011] 1) In the channel dimension, adaptive attention dynamically learns the strength of inter-channel correlation, strengthens key channels and suppresses redundant noise;

[0012] 2) In the spatial dimension, spatial attention uses global average pooling and maximum pooling to generate spatial statistics and optimize the feature space distribution;

[0013] 3) The original input, channel attention and spatial attention output are fused through multiple residual connections to extract information-rich feature representations.

[0014] Furthermore, the local consensus mining module constructs the geometric structure of matching pairs through the K-nearest neighbor graph to capture explicit dynamic features, and adopts OABlock to extract sequence-aware features to capture spatial continuity.

[0015] Furthermore, the global consensus-aware attention module has the following specific functions:

[0016] 1) Screen high-confidence matching pairs through linear projection to reduce the interference of erroneous local consensus;

[0017] 2) Capturing global consensus using multi-vector tagging and enhancing semantic understanding through self-attention;

[0018] 3) Reverse cross attention propagates the global consensus to all matching pairs, achieving two-way interaction between local and global consensus.

[0019] Furthermore, the method comprises the following steps:

[0020] A. Given a pair of images, we first extract key points and their descriptors through a feature extraction network, and then construct an initial matching set S based on the similarity of the descriptors.

[0021] B. Map the initial matching set to C-dimensional space through a multi-layer perceptron to obtain a C-dimensional vector set F;

[0022] C. Use channel-space joint interaction technology to design channel-space interaction blocks to enhance feature expression and obtain F out :

[0023] F out =F+F ch +F sp

[0024] Among them F ch represents the output of channel adaptive attention, F sp represents the output of spatial attention, F represents the initial input; the specific formula of channel adaptive attention is as follows:

[0025]

[0026] Where Conv1 and Conv2 represent convolution operations with output channels of 1 and C / 2 respectively, BN represents batch normalization, and ⊙ represents element-by-element multiplication. represents matrix multiplication, SoftMax and ReLU are activation functions, Squeeze represents a compression operation with a normalization layer, a Sigmoid layer, and a 1×1 convolutional layer; the specific definition of spatial attention is as follows:

[0027] F sp =F ch ⊙MLP(APool(F ch )+MPool(F ch ))+F ch

[0028] APool and MPool represent average pooling and maximum pooling, and MLP represents multi-layer perceptron.

[0029] D. Design a local consensus mining block using feature fusion technology to extract reliable local consensus F. local :

[0030]

[0031] in Indicates that based on F out An explicit dynamic graph is constructed by selecting K nearest neighbors. Conv3 and Conv4 represent convolutions with different kernels, respectively. [·||·] represents feature concatenation along the channel dimension, and OA represents OABlock.

[0032] E. Use channel-space interaction blocks to enhance feature expression, and use cross attention to design global consensus-aware attention blocks to achieve local-global consensus bidirectional interaction:

[0033] F top =TopK(Project(F g ), s p )

[0034] F b =CA(F g ,SA(CA(F global , F top )))

[0035] Among them F g By the local consensus feature F local Obtained through the channel-space interaction block, Project means F g Projected into one-dimensional space, the probability distribution p of each matching point pair is generated. TopK means filtering F according to the probability p. g Top s p The proportion of high confidence feature subsets, F top represents the selected high-quality feature subset, F global Represents a learnable global tag, which is a learnable variable introduced during network initialization. b After the feature expression is enhanced by the local consensus-aware attention block, SA and CA represent self-attention and cross-attention respectively;

[0036] F. Decode the optimized features through MLP, predict the probability of each match being a correct match, and generate a binary classification result;

[0037] G. Based on the classification results, the essential matrix is ​​estimated using the weighted eight-point algorithm.

[0038] Furthermore, in step A, the feature extraction network is SIFT or SuperPoint.

[0039] Furthermore, steps C to G will be iteratively executed twice.

[0040] Furthermore, each iteration uses the predicted classification results and the actual category results to calculate the cross entropy loss, and uses the predicted essential matrix and the actual essential matrix to calculate the regression loss to guide network training, and finally obtain a high-performance feature matching model.

[0041] The present invention also provides a dual-view feature matching system based on channel and space joint interaction and two-way consensus interaction, including a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, it can implement any of the method steps described above.

[0042] The present invention also provides a computer-readable storage medium on which computer program instructions that can be executed by a processor are stored. When the processor executes the computer program instructions, any of the method steps described above can be implemented.

[0043] Compared with the existing technology, the present invention has the following beneficial effects: the method of the present invention can effectively capture the global information of ordered motion vectors, and the method of the present invention can be applied to multiple fields such as posture estimation, point cloud alignment, image stitching and motion segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 Flowchart of an embodiment of the present invention.

[0045] Figure 2 This is a comparison chart of the results of the embodiment of the present invention on the YFCC and SUN3D datasets with other methods (CSBCNet is the method proposed in the present invention). DETAILED DESCRIPTION

[0046] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.

[0047] The present invention provides a dual-view feature matching method based on channel and space joint interaction and two-way consensus interaction, comprising:

[0048] First, we propose a novel and efficient network backbone: the Channel-Spatial Interaction Module. This module employs a dual-path attention framework, integrating channel-adaptive attention and spatial attention to enhance feature representation, avoiding the information loss caused by the independent processing of spatial positions in the PointCN module. Specifically: 1) In the channel dimension, adaptive attention dynamically learns the strength of inter-channel correlations, strengthening key channels and suppressing redundant noise; 2) In the spatial dimension, spatial attention utilizes global average pooling and maximum pooling to generate spatial statistics and optimize the spatial distribution of features; 3) Multiple residual connections are used to fuse the original input, the output of channel attention, and the output of spatial attention to extract information-rich feature representations.

[0049] In addition, to explore the two-way interaction between local consensus and global consensus and generate more cohesive feature representations, we propose a local consensus mining module and a global consensus-aware attention module. On the one hand, the local consensus mining module constructs the geometric structure of matching pairs through the K-nearest neighbor graph to capture explicit dynamic features; on the other hand, it uses OABlock[8] to extract sequence-aware features to capture spatial continuity. By fusing the two types of features, the LCM module can jointly model the geometric structure and spatial continuity of matching pairs, reliably extract local consensus, and lay the foundation for global consensus extraction. The GCAA module constructs global consensus by aggregating high-confidence correct matching pairs and uses the cross-attention mechanism to achieve the interaction between local and global consensus: 1) High-confidence matching pairs are screened through linear projection to reduce the interference of erroneous local consensus; 2) Multi-vector labeling is used to capture global consensus and semantic understanding is enhanced through self-attention; 3) Inverse cross-attention is used to propagate global consensus to all matching pairs, achieving a two-way interaction between local and global consensus.

[0050] Based on the above modules, the present invention designs a new type of efficient network - channel-space interaction and two-way consensus interaction network.

[0051] like Figure 1 As shown, an embodiment of the present invention provides a dual-view feature matching method based on channel and space joint interaction and two-way consensus interaction, including the following steps:

[0052] A. Given an image pair consisting of two images, first extract key points and their descriptors through a feature extraction network (such as SIFT or deep learning descriptors), and construct an initial matching set S based on the descriptor similarity;

[0053] B. Map the initial matching set to C-dimensional space through a multi-layer perceptron to obtain a C-dimensional vector set F;

[0054] C. Use channel-space joint interaction technology to design channel-space interaction blocks to enhance feature expression and obtain F out :

[0055] F out =F+F ch +F sp

[0056] Among them F ch represents the output of channel adaptive attention, F sp represents the output of spatial attention, F represents the initial input; the specific formula of channel adaptive attention is as follows:

[0057]

[0058] Where Conv1 and Conv2 represent convolution operations with output channels of 1 and C / 2 respectively, BN represents batch normalization, and ⊙ represents element-by-element multiplication. represents matrix multiplication, SoftMax and ReLU are activation functions, Squeeze represents a compression operation with a normalization layer, a Sigmoid layer, and a 1×1 convolutional layer; the specific definition of spatial attention is as follows:

[0059] F sp =F ch ⊙MLP(APool(F ch )+MPool(F ch ))+F ch

[0060] APool and MPool represent average pooling and maximum pooling, and MLP represents multi-layer perceptron.

[0061] D. Design a local consensus mining block using feature fusion technology to extract reliable local consensus F. local :

[0062]

[0063] in Indicates that based on F out An explicit dynamic graph is constructed by selecting K nearest neighbors. Conv3 and Conv4 represent convolutions with different kernels, respectively. [·||·] represents feature concatenation along the channel dimension, and OA represents OABlock.

[0064] E. Use channel-space interaction blocks to enhance feature expression, and use cross attention to design global consensus-aware attention blocks to achieve local-global consensus bidirectional interaction:

[0065] F top =TopK(Project(F g ), s p )

[0066] F b =CA(F g,SA(CA(F global , F top )))

[0067] Among them F g By the local consensus feature F local Obtained through the channel-space interaction block, Project means F g Projected into one-dimensional space, the probability distribution p of each matching point pair is generated. TopK means filtering F according to the probability p. g Top s p The proportion of high confidence feature subsets, F top represents the selected high-quality feature subset, F global Represents a learnable global tag, which is a learnable variable introduced during network initialization. b After the feature expression is enhanced by the local consensus-aware attention block, SA and CA represent self-attention and cross-attention respectively;

[0068] F. Decode the optimized features through MLP, predict the probability of each match being a correct match, and generate a binary classification result;

[0069] G. Based on the classification results, the essential matrix is ​​estimated using the weighted eight-point algorithm.

[0070] Steps C to G are iterated twice. In addition, each iteration uses the predicted classification results and the actual classification results to calculate the cross entropy loss, and uses the predicted essence matrix and the actual essence matrix to calculate the regression loss to guide network training, ultimately obtaining a high-performance feature matching model.

[0071] Figure 2 This is a comparison chart of the results of the embodiment of the present invention on the YFCC and SUN3D datasets with other methods (CSBCNet is the method proposed in the present invention).

[0072] References:

[0073] [1]Lowe D G.Distinctive image features from scale-invariant keypoints[J].Internationaljournal ofcomputer vision,2004,60:91-110.

[0074] [2]Kwang Moo Yi,Eduard Trulls,Yuki Ono,Vincent Lepetit,MathieuSalzmann,and Pascal Fua.2018.Learning to find good correspondences.InProceedings of the IEEE Conference on Computer Vision and PatternRecognition.2666–2674.

[0075] [3]Chen Zhao,Yixiao Ge,Feng Zhu,Rui Zhao,Hongsheng Li,and MathieuSalzmann.2021.Progressive Correspondence Pruning by Consensus Learning.InProceedings ofthe IEEE Conference on International Conference on ComputerVision.6464–6473.

[0076] [4]Thomas N.Kipf and Max Welling.2017.Semi-Supervised Classificationwith Graph Convolutional Networks.In International Conference on LearningRepresentations.1–14.

[0077] [5]Xin Liu and Jufeng Yang.2023.Progressive Neighbor ConsistencyMining for Correspondence Pruning.In Proceedings of the IEEE Conference onComputer Vision and Pattern Recognition.9527–9537.Zhang S,Ma J.ConvMatch:Rethinking Network Design for Two-View Correspondence Learning[C].Proceedingsof the AAAI Conference on Artificial Intelligence,2023:3472-3479.

[0078] [6]Junwen Guo,Guobao Xiao,Shiping Wang,and Jun Yu.2024.Graph contexttransformation learning for progressive correspondence pruning.In Proceedingsofthe AAAI Conference onArtificial Intelligence,Vol.38.1968–1975.

[0079] [7]Tangfei Liao,Xiaoqin Zhang,Li Zhao,Tao Wang,and GuobaoXiao.2024.VSFormer:Visual-Spatial Fusion Transformer for CorrespondencePruning.In Proceedings of the AAAI Conference onArtificial Intelligence.3369–3377.

[0080] [8]Jiahui Zhang,Dawei Sun,Zixin Luo,Anbang Yao,Lei Zhou,Tianwei Shen,Yurong Chen,Hongen Liao,and Long Quan.2019.Learning Two-View Correspondencesand Geometry Using Order-Aware Network.In Proceedings of the IEEE Conferenceon International Conference on Computer Vision.5844–5853.

[0081] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions and effects do not exceed the scope of the technical solution of the present invention, shall fall within the scope of protection of the present invention.

Claims

1. A dual-view feature matching method based on channel and space joint interaction and two-way consensus interaction, characterized in that: include: A network backbone, the channel-space interaction module, is proposed. It uses a dual-path attention mechanism framework to enhance feature representation by integrating channel-adaptive attention and spatial attention, thus avoiding information loss caused by independent spatial processing in the PointCN module. A local consensus mining module and a global consensus-aware attention module are proposed. The local consensus mining module jointly models the geometric structure and spatial continuity of matching pairs to extract local consensus; the global consensus-aware attention module builds global consensus by aggregating high-confidence correct matching pairs, and uses the cross-attention mechanism to realize the interaction between local and global consensus.

2. A dual-view feature matching method based on channel and space joint interaction and two-way consensus interaction according to claim 1, characterized in that: The channel-space interaction module has the following specific functions: 1) In the channel dimension, adaptive attention dynamically learns the strength of inter-channel correlation, strengthens key channels and suppresses redundant noise; 2) In the spatial dimension, spatial attention uses global average pooling and maximum pooling to generate spatial statistics and optimize the feature space distribution; 3) The original input, channel attention and spatial attention output are fused through multiple residual connections to extract information-rich feature representations.

3. The dual-view feature matching method based on channel and space joint interaction and two-way consensus interaction according to claim 1 is characterized in that: The local consensus mining module constructs the geometric structure of matching pairs through the K-nearest neighbor graph to capture explicit dynamic features, and adopts OABlock to extract sequence-aware features to capture spatial continuity.

4. The dual-view feature matching method based on channel and space joint interaction and two-way consensus interaction according to claim 1 is characterized in that: The global consensus-aware attention module has the following specific functions: 1) Screen high-confidence matching pairs through linear projection to reduce the interference of erroneous local consensus; 2) Capturing global consensus using multi-vector tagging and enhancing semantic understanding through self-attention; 3) Reverse cross attention propagates the global consensus to all matching pairs, achieving two-way interaction between local and global consensus.

5. A dual-view feature matching method based on channel and space joint interaction and two-way consensus interaction according to any one of claims 1 to 4, characterized in that: The method comprises the following steps: A. Given a pair of images, we first extract key points and their descriptors through a feature extraction network, and then construct an initial matching set S based on the similarity of the descriptors. B. Map the initial matching set to C-dimensional space through a multi-layer perceptron to obtain a C-dimensional vector set F; C. Use channel-space joint interaction technology to design channel-space interaction blocks to enhance feature expression and obtain F out : F out =F+F ch +F sp Among them F ch represents the output of channel adaptive attention, F sp represents the output of spatial attention, F represents the initial input; the specific formula of channel adaptive attention is as follows: Where Conv1 and Conv2 represent convolution operations with output channels of 1 and C / 2 respectively, BN represents batch normalization, and ⊙ represents element-by-element multiplication. represents matrix multiplication, SoftMax and ReLU are activation functions, Squeeze represents a compression operation with a normalization layer, a Sigmoid layer, and a 1×1 convolutional layer; the specific definition of spatial attention is as follows: F sp =F ch ⊙MLP(APool(F ch )+MPool(F ch ))+F ch APool and MPool represent average pooling and maximum pooling, and MLP represents multi-layer perceptron. D. Design a local consensus mining block using feature fusion technology to extract reliable local consensus F. local : in Indicates that based on F out An explicit dynamic graph is constructed by selecting K nearest neighbors. Conv3 and Conv4 represent convolutions with different kernels, respectively. [·‖·] represents feature concatenation along the channel dimension, and OA represents OABlock. E. Use channel-space interaction blocks to enhance feature expression, and use cross attention to design global consensus-aware attention blocks to achieve local-global consensus bidirectional interaction: F top =TopK(Project(F g ),s p ) F b =CA(F g ,SA(CA(F global ,F top ))) Among them F g By the local consensus feature F local Obtained through the channel-space interaction block, Project means F g Projected into one-dimensional space, the probability distribution p of each matching point pair is generated. TopK means filtering F according to the probability p. g Top s p The proportion of high confidence feature subsets, F top represents the selected high-quality feature subset, F global Represents a learnable global tag, which is a learnable variable introduced during network initialization. b After the feature expression is enhanced by the local consensus-aware attention block, SA and CA represent self-attention and cross-attention respectively; F. Decode the optimized features through MLP, predict the probability of each match being a correct match, and generate a binary classification result; G. Based on the classification results, the essential matrix is ​​estimated using the weighted eight-point algorithm.

6. The dual-view feature matching method based on channel and space joint interaction and two-way consensus interaction according to claim 5 is characterized in that: In step A, the feature extraction network is SIFT or SuperPoint.

7. The dual-view feature matching method based on channel and space joint interaction and two-way consensus interaction according to claim 5 is characterized in that: Steps C to G will be iterated twice.

8. The dual-view feature matching method based on channel and space joint interaction and two-way consensus interaction according to claim 7 is characterized in that: Each iteration uses the predicted classification results and the actual category results to calculate the cross entropy loss, and uses the predicted essential matrix and the actual essential matrix to calculate the regression loss to guide network training, and finally obtain a high-performance feature matching model.

9. A dual-view feature matching system based on channel and space joint interaction and two-way consensus interaction, characterized in that: The method comprises a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the method steps according to any one of claims 1 to 8 can be implemented.

10. A computer-readable storage medium storing computer program instructions that can be executed by a processor, wherein when the processor executes the computer program instructions, the method steps according to any one of claims 1 to 8 can be implemented.

Citation Information

Cited By

  • Cross-modal small-strand pedestrian re-identification method based on optimal subset learning

    CN122135306A

  • A cross-modal small-group pedestrian re-identification method based on optimal subset learning

    CN122135306B