Feature matching method based on geometric perception
By extracting geometric features in the image and combining SuperPoint and SuperGlue methods, the problem of lack of geometric structure information in the existing feature matching methods is solved, and more robust and accurate feature matching is achieved, improving the accuracy of camera pose estimation.
Patent Information
- Application Number
- CN202510446271.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-29
AI Technical Summary
The existing feature matching methods lack geometric structure information when facing challenges such as poor texture, occlusion and repetitive patterns, resulting in unstable matching, especially in scenarios where viewing angle changes or texture repetitions, affecting the accuracy of camera pose estimation.
By extracting the curvature of key points, key line segments and plane points from the image as geometric features, combining SuperPoint to generate semantic feature descriptors, constructing graphs and integrating local and global feature information, using the optimal transmission filtering matching pairs of SuperGlue and LoFTR to enhance geometric consistency.
It improves the robustness and accuracy of feature matching, enhances the local feature matching effect in challenging scenarios, and improves the performance of camera pose estimation.
Smart Images

Figure CN120388193A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a feature matching method based on geometric perception. Background Art
[0002] Deep learning has brought feature matching into a new stage of development. Feature matching can be divided into two key parts according to the presence of detectors. Among them, one type of method based on detectors has outstanding performance, and its representatives are SuperGlue, Gluestick, and Lightglue. SuperGlue is an attention-based GNN network. This method constructs a complete graph using key points within and between images, and updates node representations through attention-based information transmission. The GlueStick method jointly matches key points and line segments, and processes them through a unified graph neural network and wireframe structure. The CVG team proposed the LightGlue algorithm based on SuperGlue, which realizes performance improvement under various operating conditions by dynamically adjusting the scale of the network instead of simply reducing the size of the attention mechanism. Another type is the category that does not use detectors, including CNN-based, Transformer-based, and Patch-based methods. The detector-free method eliminates the feature detector and directly extracts visual descriptors in the dense grid on the image to generate dense matches. Therefore, it can capture key points that are repeatable between image pairs.
[0003] The existing technical methods have the following disadvantages: Self-attention and cross-attention mechanisms usually focus on the significant features of the matching region, but ignore the real geometric structure information within these regions. Therefore, when facing challenges such as poor texture, occlusion, and repetitive patterns, traditional local feature matching methods perform poorly, and the lack of geometric structure information (such as curvature) will affect the accuracy of camera pose estimation, resulting in the matched features not being robust enough, especially performing poorly in scenes with perspective changes or texture repetitions. Summary of the Invention
[0004] The purpose of the present invention is to provide a feature matching method based on geometric perception to solve the above existing problems.
[0005] The technical solution adopted by a feature matching method based on geometric perception disclosed by the present invention is:
[0006] A feature matching method based on geometric perception, characterized by comprising:
[0007] Constructor: Extract key points, key line segments, and planar point curvatures from the image as true geometric features. Use SuperPoint to generate semantic feature descriptors. After applying filters, we combine these features to construct a graph for each image;
[0008] Information Transmitter: Used to integrate information between all local and global features while ignoring unnecessary representations in irrelevant foreground or background regions. Its output consists of updated feature vectors;
[0009] Matcher: To generate the final matching matrix, use optimal transport in SuperGlue to screen matching pairs, and use two independent softmax outputs of LoFTR to generate the final matching score matrix.
[0010] As an optimal solution, the graph constructor:
[0011] Since the subsequent information transmitter needs to take a pair of graphs and their nodes and edges (ε a , ε b ) in the form of structured data input, therefore, in this module, it is necessary to construct node features and edge features Consequently, it is necessary to determine the positions of key points and filter out the correct point-to-point connectivity from the output of the edge extractor to form an entire graph. For the feature description of key points, it is necessary to combine position encoding and the planar curvature of the curve where the key points are located;
[0012] Use SuperPoint to extract the positions of image key points, their corresponding feature descriptors, and related confidences. Subsequently, use the ELSED method to segment the true line segments in the image, and filter the key points by reprojection according to the relative pose between the two images;
[0013] Since the line extraction method does not generate a feature vector corresponding to the line, use the EdgeConv module to take the feature vectors of the two endpoints of the line as input, and then map the feature vectors of the two endpoints through a multi-layer perceptron to generate a new vector, which represents the edge feature between the two endpoints The parameters of the multi-layer perceptron model represent the possibility of connection between the two endpoints;
[0014] Point-Line Filter: During the filtering process, key points in the distant background area are not helpful for the relative pose. Therefore, use depth information and camera parameters to extract the overlapping areas of the two images. For the filtering of line features and their endpoints, use the DBSCAN method to cluster and filter the endpoints that are close enough to form a connected subgraph Among which K a represents the figure the number of sub - figures;
[0015] Since the extracted line segments may contain instances that are far apart but still have points connected, it is necessary to then filter the connectivity relationship between the end - points within each sub - figure;
[0016] Directly calculate the distance between each pair of points in the sub - figure, and select the 1 to 2 closest points as valid connection points, and other connectivity relationships are cut off to ensure that the connectivity of each point is consistent with the real structure in the image;
[0017] Finally, in order to obtain an undirected graph, the generated adjacency matrix must be symmetric to ensure that the graph is undirected.
[0018] As a preferred solution, for the information transmitter: For a pair of constructed graphs, it is essential to exchange geometric and semantic information between adjacent nodes. It is also necessary to verify the consistency of features between the two graphs. Combining with the SuperGlue framework, an intra - graph and inter - graph message - passing layer is given.
[0019] The intra - graph message - passing layer: The purpose is to facilitate the feature exchange between adjacent points so that each key point can contain semantic and geometric information from its surroundings;
[0020] The inter - graph message - passing layer: The goal is to establish an accurate correspondence between the nodes of two graphs. The matching between graphs needs to identify the corresponding relationships in their grid - like topological structures. From the perspective of fused feature matching, this involves enriching the semantic features of each node to enhance the matching accuracy. Considering that our method reduces the input key points, in this layer, multi - head attention operations are used to update the node features, using the inherent relationship between key points to optimize their representations, and promoting more accurate matching.
[0021] As a preferred solution, for the intra - graph message - passing layer: Define the neighbor node representation as where represents the set of neighbor nodes of node i. For the update strategy of node information, the whole process is divided into two steps. First, in order to construct the feature vector of the edge, the EdgeConv module is proposed. The role of this module is to use the information difference and spatial distance between adjacent nodes to generate learnable parameters to represent the soft connection between these two nodes. For EdgeConv, there is the following definition:
[0022]
[0023] where θ represents the parameters of the multi - layer perceptron neural network,
[0024] To ensure that the feature vectors updated through the connectivity relationship are not affected by significant feature differences between adjacent key points, after the EdgeConv operation, a graph attention feature aggregation operation is added. After obtaining the updated node features z' from the EdgeConv operation i , graph attention is used to update it, aiming to calculate the score of each relationship weight for each edge (i, j), which implies the importance of neighbors to the node features z' i . The specific formula is as follows:
[0025]
[0026] where and represent learnable weights, d represents the feature embedding dimension, the symbol [·] represents the concatenation operation on the channel dimension, Act represents the activation function. According to the generation of the score s(z' i , z' j ), the attention weight α ij is calculated as follows:
[0027]
[0028] After the graph attention operation, the updated feature z' i ' of node i is obtained, which contains both local feature differences and spatial relationships with neighboring nodes, ensuring robustness to changes in feature dissimilarity between neighboring key points:
[0029]
[0030] where, W z represents the output projection of the feature vector, including the aggregation operation between nodes i and j.
[0031] As a preferred solution, in the stage of the inter-graph message passing layer, for node i in graph , there is an output feature from the inter-graph message passing layer and a feature corresponding to a neighboring node in graph . First, a query vector Q i ' is generated from the feature z' i , and then a key vector K j ' and a value vector V i ' are respectively generated from the feature z' i . The specific formulas are as follows:
[0032] Q i =W Q z' i ', K i =WK z' j ', V i = W V z' j '
[0033] W Q , W K and W V represent learnable projection weights, treat the output of attention as cross messages, and then pass them to Z j , this layer can be expressed as:
[0034]
[0035] As a preferred solution, the matcher includes: first calculate the inner product between all pairs of nodes in two graphs and convert it into a confidence matrix. Considering the important role of Dual-softmax in optimizing the matching matrix, this operation is introduced to calculate the joint probability of each pair of key point matches. According to the method of SuperGlue, in order to store the unmatched key point pairs, a storage point is added to both the rows and columns of the matrix.
[0036] As a preferred solution, in image analysis, given a set of images, denoted as (I a , I b ), where each image is accompanied by its associated set of key points (P a , P b ), their respective score vectors (S a , S b ) and feature descriptors (F a , F b ). In addition, the supplement of these descriptors is a set of explicit geometric features, denoted as (G a , G b ); by fusing the image descriptors with the geometric features, a feature vector for each key point p is constructed. It should be noted that each pixel coordinate is associated with a corresponding position confidence c. Among them, the image visual descriptors are extracted through a convolutional neural network architecture, while the geometric information is obtained through traditional digital image processing techniques. Therefore, images I a and I b generate two different sets of feature descriptors, denoted as and The matching of image features is defined as identifying the optimal one-to-one correspondence matrix as shown in the formula:
[0037]
[0038] The aim is to minimize the relative pose error of image pairs by utilizing all matching pairs. Each feature descriptor d is allowed to be matched at most once for pose estimation. Therefore, the key points or descriptors that still do not match are considered outliers. Thus, an alignment matrix can be derived. It represents the one-to-one correspondence between the features of two images.
[0039] The beneficial effects of a geometric-awareness-based feature matching method disclosed by the present invention are as follows: introducing real geometric information such as curvature to enhance the CNN-based image feature descriptor, improving the geometric expression ability of features to provide robust matching results, constructing the edges of a graph using the real texture features in the image and generating representations for these edges, enabling the network to explicitly learn geometric structure information, strengthening the geometric consistency of features by combining geometric information and learning schemes, improving the matching accuracy, enhancing the effect of local feature matching in challenging scenarios, and ultimately improving the performance of camera pose estimation. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is the network structure diagram of a geometric-awareness-based feature matching method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0041] The present invention will be further described and explained below in conjunction with specific embodiments and the accompanying drawings of the specification:
[0042] In image analysis, given a set of images, denoted as (I a , I b ), where each image is accompanied by its associated set of key points (P a , P b ), their respective score vectors (S a , S b ) and feature descriptors (F a , F b ). In addition, the supplement of these descriptors is a set of explicit geometric features, denoted as (G a , G b ); by fusing the image descriptors with the geometric features to construct the feature vector of each key point p. It should be noted that each pixel coordinate is associated with a corresponding position confidence c. Among them, the image visual descriptor is extracted through a Convolutional Neural Network (CNN) architecture, while the geometric information is obtained through traditional digital image processing techniques. Therefore, images I a and I b generate two different sets of feature descriptors, denoted as and The matching of image features is defined as identifying the optimal one-to-one correspondence matrix As shown in the formula:
[0043]
[0044] The purpose is to minimize the relative pose error of the image pair by utilizing all matching pairs. Each feature descriptor d is allowed to be matched at most once for pose estimation. Therefore, the keypoints or descriptors that still do not match are considered outliers. Thus, an alignment matrix can be derived It represents the one-to-one correspondence between the features of two images.
[0045] Please refer to Figure 1 , a geometric perception-based feature matching method, including:
[0046] Constructor: Extract keypoints, key line segments, and planar point curvatures from the image as real geometric features. Use SuperPoint to generate semantic feature descriptors. After applying filters, we combine these features to construct a graph for each image;
[0047] Since the subsequent information transmitter needs to take a pair of graphs and their nodes and edges (ε a , ε b ) in the form of structured data input. Therefore, node features and edge features need to be constructed simultaneously in this module. Thus, it is necessary to determine the positions of the keypoints and filter out the correct point-to-point connectivity from the output of the edge extractor to form an entire graph. For the feature description of the keypoints, position encoding and the planar curvature of the curve where the keypoints are located need to be combined;
[0048] Use SuperPoint to extract the positions of the image keypoints, their corresponding feature descriptors, and related confidences. Subsequently, use the ELSED method to segment the real line segments in the image and filter the keypoints by reprojection according to the relative pose between the keypoints and the two images;
[0049] Since the line extraction method does not generate feature vectors corresponding to the lines, the EdgeConv module is used to take the feature vectors of the two endpoints of the line as input, and then map the feature vectors of the two endpoints through a multi-layer perceptron to generate a new vector, which represents the edge feature between the two endpoints The parameters of the multi-layer perceptron model represent the possibility of connection between the two endpoints;
[0050] Point-line filter: During the filtering process, the key points in the distant background area are not helpful for the relative pose. Therefore, the overlapping area of two images is extracted using depth information and camera parameters. For the filtering of line features and their endpoints, the DBSCAN method is used to cluster and filter the endpoints that are close enough to form a connected subgraph. Where K a represents the number of subgraphs in the graph; Different from GlueStick which enhances the connectivity of line segments through interpolation, the key points with significant distance steps in the graph do not need to be aggregated with each other.
[0051] Since the extracted line segments may contain instances that are far apart but still have points connected, it is necessary to further filter the connectivity relationship between the endpoints within each subgraph.
[0052] Directly calculate the distance between each pair of points in the subgraph, and select the 1 to 2 points with the closest distance as valid connection points. Other connectivity relationships are cut off to ensure that the connectivity of each point is consistent with the true structure in the image.
[0053] Finally, in order to obtain an undirected graph, the generated adjacency matrix must be symmetric to ensure that the graph is undirected.
[0054] Information transmitter: Used to integrate the information between all local and global features, while ignoring the unnecessary representations in the irrelevant foreground or background areas. Its output consists of updated feature vectors.
[0055] For a pair of constructed graphs, it is essential to exchange geometric and semantic information between adjacent nodes, and it is also necessary to verify the consistency of features between the two graphs. Combining the SupeGlue framework, the message passing layers within and between graphs are given.
[0056] Message passing layer within the graph: The purpose is to facilitate the exchange of features between adjacent points, so that each key point can contain the semantic and geometric information from its surroundings. Define the neighbor nodes as Where represents the set of neighbor nodes of node i. For the update strategy of node information, the whole process is divided into two steps. First, in order to construct the feature vector of the edge, the EdgeConv module is proposed. The role of this module is to use the information difference and spatial distance between adjacent nodes to generate learnable parameters to represent the soft connection between these two nodes. For EdgeConv, there is the following definition:
[0057]
[0058] Where θ represents the parameters of the multi-layer perceptron neural network.
[0059] To ensure that the feature vectors updated through the connectivity relationship are not affected by significant feature differences between adjacent key points, after the EdgeConv operation, a graph attention feature aggregation operation is added. After obtaining the updated node features z' from the EdgeConv operation i it is updated using graph attention with the aim of calculating the score for each relationship weight of each edge (i, j), which implies the importance of the neighbor to the node feature z' i The specific formula is as follows:
[0060]
[0061] where and represent learnable weights, d represents the feature embedding dimension, the symbol [·] represents the concatenation operation on the channel dimension, Act represents the activation function. According to the generation of the score s(z' i ,z' j ), the attention weight α ij is calculated as follows:
[0062]
[0063] After the graph attention operation, the updated feature z' i ' of node i is obtained, which contains both local feature differences and spatial relationships with neighboring nodes, ensuring robustness to changes in feature dissimilarity between neighboring key points:
[0064]
[0065] where, W z represents the output projection of the feature vector, including the aggregation operation between nodes i and j.
[0066] Inter - graph message passing layer: The goal is to establish an accurate correspondence between the nodes of two graphs. The matching between graphs requires identifying the corresponding relationships in their grid - like topological structures. From the perspective of fused feature matching, this involves enriching the semantic features of each node to enhance the matching accuracy. Considering that our method reduces the input key points, in this layer, multi - head attention operations are used to update the node features, leveraging the inherent relationships between key points to optimize their representations and promote more accurate matching.
[0067] For node i in graph there is the output feature from the inter - graph message passing layer and the feature corresponding to a neighboring node in graph First, from the feature z' i '
[0068] Generate the query vector Q i , and then generate the key vector K j ' and the value vector V i ' respectively from the feature z' i . The specific formulas are as follows:
[0069] Q i = W Q z' i ', K i = W K z' j ', V i = W V z' j '
[0070] W Q , W K and W V represent learnable projection weights. Treat the output of the attention as cross messages and then pass them to Z j . This layer can be expressed as:
[0071]
[0072] Matcher: To generate the final matching matrix, the optimal transport in SuperGlue is used to filter matching pairs, and the two independent softmax outputs of LoFTR are used to generate the final matching score matrix
[0073] The matcher includes: First, calculate the inner product between all node pairs in the two graphs and convert it into a confidence matrix. Considering the important role of Dual-softmax in optimizing the matching matrix, this operation is introduced to calculate the joint probability of each pair of key point matches. According to the method of SuperGlue, in order to store unmatched key point pairs, a storage point is added to both the rows and columns of the matrix, similar to the role of a trash can.
[0074] Introduce real geometric information such as curvature to enhance the CNN-based image feature descriptor, improve the geometric expression ability of the features, provide robust matching results, use the real texture features in the image to construct the edges of the graph, and generate representations for these edges, enabling the network to explicitly learn geometric structure information. By combining geometric information and learning schemes, strengthen the geometric consistency of the features, improve the matching accuracy, enhance the effect of local feature matching in challenging scenarios, and ultimately improve the performance of camera pose estimation.
[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting the protection scope of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. A geometric perception-based feature matching method, characterized in that, Including: Constructor: Extract key points, key line segments, and planar point curvatures from the image as true geometric features. Use SuperPoint to generate semantic feature descriptors. After applying filters, we combine these features to construct a graph for each image; Information Transmitter: Used to integrate information between all local and global features, while ignoring unnecessary representations in irrelevant foreground or background regions. Its output consists of updated feature vectors; Matcher: To generate the final matching matrix, use optimal transport in SuperGlue to screen matching pairs, and use two independent softmax outputs of LoFTR to generate the final matching score matrix.
2. The feature matching method based on geometric perception according to claim 1, wherein The graph constructor: Since the subsequent information transmitter requires structured data input in the form of a pair of graphs and their nodes and edges Therefore, in this module, it is necessary to construct node features and edge features Consequently, it is necessary to determine the positions of the key points and filter out the correct connectivity between points from the output of the edge extractor to form an entire graph. For the feature description of the key points, it is necessary to combine the positional encoding and the planar curvature of the curve where the key points are located; Use SuperPoint to extract the positions of image key points, their corresponding feature descriptors, and related confidences. Subsequently, use the ELSED method to segment the true line segments in the image, and filter the key points by reprojection according to the relative pose between the key points and the two images; Since the line extraction method does not generate feature vectors corresponding to the lines, the feature vectors of the two endpoints of the line are used as inputs to the EdgeConv module, and then the feature vectors of the two endpoints are mapped by a multi-layer perceptron to generate a new vector, which represents the edge feature between the two endpoints. The parameters of the multi-layer perceptron model represent the possibility of connection between the two endpoints; Point-line filter: During the filtering process, key points in the distant background area are not helpful for relative pose. Therefore, the overlapping area of two images is extracted using depth information and camera parameters. For the filtering of line features and their endpoints, the DBSCAN method is used to cluster and filter endpoints that are close enough to form a connected subgraph. where K a represents the number of subgraphs in the graph; Since the extracted line segments may contain instances that are far apart but still have points connected, it is necessary to then filter the connectivity relationship between the endpoints within each subgraph; Directly calculate the distance between each pair of points in the subgraph, and select the 1 to 2 closest points as valid connection points. Other connectivity relationships are cut off to ensure that the connectivity of each point is consistent with the true structure in the image; Finally, to obtain an undirected graph, the generated adjacency matrix must be symmetric to ensure that the graph is undirected.
3. The geometric perception-based feature matching method according to claim 1, wherein The information transmitter: For a pair of constructed graphs, it is essential to exchange geometric and semantic information between adjacent nodes. It is also necessary to verify the consistency of features between the two graphs. Combining with the SupeGlue framework, the intra-graph and inter-graph message passing layers are given; The intra-graph message passing layer: Aims to facilitate the exchange of features between adjacent points, enabling each key point to contain semantic and geometric information from its surroundings; The inter-graph message passing layer: The goal is to establish an accurate correspondence between the nodes of the two graphs. The matching between graphs requires identifying the corresponding relationships in their grid-like topological structures. From the perspective of fusing feature matching, this involves enriching the semantic features of each node to enhance the accuracy of matching. Considering that our method reduces the input key points, in this layer, use multi-head attention operations to update the node features, utilize the inherent relationship between key points to optimize their representations, and promote more accurate matching.
4. The feature matching method based on geometric perception according to claim 3, wherein The in-graph message passing layer: Defines neighbor nodes as where represents the set of neighbor nodes of node i. For the update strategy of node information, the whole process is divided into two steps. First, in order to construct the feature vector of the edge, the EdgeConv module is proposed. The role of this module is to utilize the information difference and spatial distance between adjacent nodes to generate learnable parameters to characterize the soft connection between these two nodes. For EdgeConv, the following definition is given: Where θ represents the parameters of the multi-layer perceptron neural network; To ensure that the feature vectors updated through the connectivity relationship are not affected by significant feature differences between adjacent key points, after the EdgeConv operation, a graph attention feature aggregation operation is added. After obtaining the updated node features z' from the EdgeConv operation i it is updated using graph attention, aiming to calculate the score of each relationship weight for each edge (i, j), which implies the importance of neighbors to the node features z' i The specific formula is as follows: Among them and represent learnable weights, d represents the feature embedding dimension, the symbol [·] represents the concatenation operation on the channel dimension, Act represents the activation function. According to the generation of the score s(z' i , z' j ), it is necessary to calculate the attention weight α ij as follows: After the graph attention operation, the updated feature z' of node i is obtained i ' , which contains both local feature differences and spatial relationships with neighboring nodes, ensuring robustness to changes in feature dissimilarity between neighboring key points: Among them, W z represents the output projection of the feature vector and includes the aggregation operation between nodes i and j.
5. The feature matching method based on geometric perception according to claim 3, wherein In the stage of the inter-graph message passing layer, for node i in graph , there is an output feature from the inter-graph message passing layer and a feature corresponding to an adjacent node in graph . First, a query vector Q i ' is generated from the feature z' i . Then, a key vector K j ' and a value vector V i i are respectively generated from the feature z', and the specific formulas are as follows: Q i = W Q z' i ' , K i = W K z' j ' , V i = W V z' j ' W Q ,W K and W V represent learnable projection weights, treat the output of attention as cross messages, and then pass them to Z j , and this layer can be expressed as:
6. The feature matching method based on geometric perception according to claim 1, wherein The matcher includes: First, calculate the inner product between all pairs of nodes in the two graphs and convert it into a confidence matrix. Considering the important role of Dual-softmax in optimizing the matching matrix, this operation is introduced to calculate the joint probability of matching for each pair of key points. According to the method of SuperGlue, to store the unmatched key point pairs, a storage point is added to both the rows and columns of the matrix.
7. The geometric perception-based feature matching method according to claim 6, wherein In image analysis, given a set of images, denoted as (I a , I b ), where each image is accompanied by its associated set of key points (P a , P b ), their respective score vectors (S a , S b ), and feature descriptors (F a , F b ). In addition, the complement of these descriptors is a set of explicit geometric features, denoted as (G a , G b ); by fusing the image descriptors with the geometric features to construct a feature vector for each key point p, it is noted that each pixel coordinate is associated with a corresponding position confidence c, where the image visual descriptors are extracted through a convolutional neural network architecture, while the geometric information is obtained through traditional digital image processing techniques. Therefore, images I a and I b produce two different sets of feature descriptors, denoted as and The matching of image features is defined as identifying the optimal one-to-one correspondence matrix as shown in the equation: The goal is to minimize the relative pose error of image pairs by utilizing all matching pairs. Each feature descriptor d is allowed to be matched at most once for pose estimation. Therefore, the key points or descriptors that still do not match are considered outliers. Thus, an alignment matrix can be derived. It represents the one-to-one correspondence between the features of two images.