Image feature matching model considering geometric prior information attention calculation

By introducing an attention computer system and a multi-stage feature matching framework that takes into account geometric prior information into image matching, the lack of matching accuracy and robustness in complex scenarios is solved, and high-precision and efficient image feature matching is achieved.

CN120147666AActive Publication Date: 2025-06-13HUBEI UNIV OF TECH

Patent Information

Application Number
CN202510194510.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-13
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

Existing image matching methods are difficult to achieve high robustness and high precision matching in complex scenarios, especially in sparse texture areas, repeated texture patterns and large-view angle changes.

Method used

An image feature matching model that takes into account the attention calculation of geometric prior information is adopted. This model improves the accuracy and robustness of the matching through a multi-stage feature matching framework and window division with affine transformation compensation.

Benefits of technology

It significantly improves the accuracy and efficiency of image matching, and can achieve high-precision feature matching in complex geometric deformation and large-view angle changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147666A_ABST
    Figure CN120147666A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of computer vision and remote sensing image processing, and relates to a method for constructing an image feature matching model considering geometric prior information attention calculation and the model, and the method comprises the steps: 1) constructing a data set; 2) constructing an image feature matching model by adopting a geometric prior information attention calculation mode; and 3) constructing a loss function of the image feature matching model, and training the image feature matching model by using the constructed data set number to obtain the weight of the image feature matching model. The invention provides the construction method of the image feature matching model considering the geometric prior information attention calculation and the model, which can gradually improve the matching precision and robustness in complex geometric deformation and large view angle change scenes and can cope with diversified matching scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and image processing, and relates to a method for constructing an image feature matching model and the model, in particular to a method for constructing an image feature matching model and the model that takes into account geometric prior information attention calculation. Background Art

[0002] Image matching is one of the core tasks in the field of computer vision, and it is widely applied to various geometric vision tasks such as 3D reconstruction, visual positioning, and simultaneous localization and mapping. These tasks need to complete the restoration of the spatial structure, pose estimation, etc. by establishing reliable feature correspondence relationships between image pairs. However, due to various challenges existing in the real scene, such as large perspective changes, different lighting conditions, texture-sparse regions, repetitive texture patterns, and motion blur, etc., the accuracy of image matching is often significantly limited. Therefore, how to achieve high-robustness and high-precision image matching in complex scenes has become a key issue in current research.

[0003] Traditional image matching methods usually adopt detector-based methods, including three stages: feature point detection, feature description, and feature matching. In this mode, the algorithm first extracts sparse key points in the image through a detector, then characterizes the key points through a descriptor, and finally searches for the optimal match in the feature space. However, the performance of such methods highly depends on the quality of key point detection and description. When facing scenes with scarce textures or image pairs with large geometric deformations and lighting interferences, the key point detector often fails to extract sufficient repeatable interest points, thereby resulting in matching failures.

[0004] To address the limitations of traditional methods, a series of detector-free matching methods that abandon detectors have emerged in recent years. These methods generate global and local feature correspondences directly from the original images, avoiding the dependence on key-point detection. This not only expands the applicable scenarios of matching but also significantly improves the adaptability to extreme scenarios. For example, the matching method based on Transformer can achieve more robust feature alignment through long-range dependence modeling and local detail attention, showing stronger robustness especially in regions with sparse textures or repetitive patterns. As a representative study among them, LoFTR updates cross-view features through self-attention blocks and cross-attention blocks, and uses a linear transformer to replace global attention to achieve more controllable computational costs and at the same time achieve very good matching results. Although LoFTR performs well in practical applications, its main problem lies in the lack of detailed local interaction between pixel labels, which may limit its ability to extract highly accurate and well-localized correspondences. New research results reveal that the cross-attention maps of the linear transformers generated by LoFTR tend to spread between larger regions and fail to effectively focus on the actual corresponding regions. And in the actual image matching process, not all regions of the images to be matched are overlapping and matchable. In fact, only a small part of the regions have effective matching relationships, while the vast majority of regions are irrelevant or uncorrelated. This means that when performing feature matching, more attention needs to be paid to those regions with potential matching relationships rather than blindly comparing the entire image. By identifying and filtering out irrelevant regions, the accuracy and efficiency of matching can be significantly improved, thus achieving more accurate feature matching and better application effects. Summary of the Invention

[0005] To solve the above technical problems existing in the background art, the present invention provides a method for constructing an image feature matching model and the model, which can gradually improve the accuracy and robustness of matching and can handle diverse matching scenarios that take into account geometric prior information attention in scenarios of complex geometric deformation and large viewing angle changes.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A method for constructing an image feature matching model that takes into account geometric prior information attention, characterized in that: the method for constructing the image feature matching model that takes into account geometric prior information attention includes the following steps:

[0008] 1) Construct a data set;

[0009] 2) Construct an image feature matching model using geometric prior information attention;

[0010] 3) Construct the loss function of the image feature matching model, and use the dataset constructed in step 1) to train the image feature matching model to obtain the weights of the image feature matching model.

[0011] Preferably, step 1) is specifically: select images of multiple scenes in the Megadepth dataset as the training dataset for training real landmark information; each scene includes at least 100 pairs of images.

[0012] Preferably, the specific implementation manner of step 2) is:

[0013] 2.1) Extract local image features from the images in the dataset obtained in step 1);

[0014] 2.2) Perform rough feature matching on the results obtained in step 2.1) to obtain rough feature matching results;

[0015] 2.3) Perform medium-grained feature matching on the rough feature matching results obtained in step 2.2) to obtain medium-grained feature matching results;

[0016] 2.4) Perform fine feature matching on the medium-grained feature matching results obtained in step 2.3) to obtain fine feature matching results;

[0017] 2.5) Construct the architecture of the image feature matching model according to rough feature matching, medium-grained feature matching, and fine feature matching.

[0018] Preferably, in step 2.1), based on the ResNet backbone network and combined with the Feature Pyramid Network, multi-scale local image features are extracted from the dataset obtained in step 1); preferably, the specific implementation manner of step 2.1) is:

[0019] 2.1.1) Pass the images in the dataset obtained in step 1) through the four layers of the ResNet backbone network to gradually extract features, and respectively obtain intermediate feature maps of 1 / 2 scale, 1 / 4 scale, 1 / 8 scale, and 1 / 16 scale;

[0020] 2.1.2) Use the Feature Pyramid Network to perform fusion processing on the intermediate feature maps. Starting from the deepest intermediate feature map of 1 / 16 scale, adjust the number of channels through 1×1 convolution, and then use the bilinear interpolation method to upsample the intermediate feature map of 1 / 16 scale to obtain the same spatial resolution as the intermediate feature map of 1 / 8 scale. Add the feature map obtained by upsampling the intermediate feature map of 1 / 16 scale to the intermediate feature map of 1 / 8 scale to form a medium-grained feature map of 1 / 8 scale;

[0021] 2.1.3) Further extract features from the medium-grained feature map at 1 / 8 scale obtained in step 2.1.2) through a 3×3 convolution. After convolution, use Batch Normalization and the LeakyReLU activation function to enhance the non-linear expression ability, adjust the number of output channels through a 1×1 convolution, and use bilinear interpolation to upsample to 1 / 4 scale, then add it to the intermediate feature map at 1 / 4 scale to form a feature map at 1 / 4 scale;

[0022] 2.1.4) Based on the feature map at 1 / 4 scale obtained in step 2.1.3), repeat step 2.1.3) to obtain a fine-grained feature map at 1 / 2 scale;

[0023] 2.1.5) Directly use the intermediate feature map at 1 / 16 scale obtained in step 2.1.1) as the rough feature map at 1 / 16 scale; use the rough feature map at 1 / 16 scale, the medium-grained feature map at 1 / 8 scale obtained in step 2.1.2), and the fine-grained feature map at 1 / 2 scale obtained in step 2.1.4) as the results of local image feature extraction, and denote the rough feature map at 1 / 16 scale, the medium-grained feature map at 1 / 8 scale, and the fine-grained feature map at 1 / 2 scale as and,

[0024] Preferably, the specific implementation manner of step 2.2) is:

[0025] 2.2.1) Transform the rough feature map at 1 / 16 scale obtained in step 2.1.5) to obtain the transformed rough feature; preferably, the specific implementation manner of step 2.2.1) is:

[0026] First, perform position encoding on the rough feature map at 1 / 16 scale obtained in step 2.1.5):

[0027]

[0028] Perform a vector addition operation on the position encoding and the rough feature map at 1 / 16 scale to obtain the rough feature after position encoding; then perform linear self-attention and cross-attention calculations on the rough feature after position encoding, and the calculation methods of the linear self-attention and cross-attention are:

[0029]

[0030] Attention(Q, K, V) = D -1 ·(φ(Q)·(φ(K) T V))

[0031] Where:

[0032] R is the numerical type of the tensor;

[0033] K is the key in the attention mechanism;

[0034] V is the value matrix;

[0035] n and dk respectively represent the dimensions of Q and K after being processed by the function φ(x);

[0036] T is the transpose of the matrix;

[0037] D is the normalization factor;

[0038] φ(x) is the kernel function that converts Q and K into non - negative mappings, and is specifically defined as:

[0039] φ(x) = elu(x) + 1

[0040] After calculating the attention, the calculated attention is multiplied element - by - element with the coarse feature map at a scale of 1 / 16 to obtain the transformed coarse feature and the and are respectively and transformed features;

[0041] 2.2.2) Use the double - softmax method to perform coarse feature matching on the transformed coarse features and to obtain the coarse feature matching result; Preferably, the specific implementation manner of the step 2.2.2) is:

[0042] First, flatten the transformed coarse features and feature maps to obtain the feature vectors feat 0 and feat 1 , and further perform L2 - norm normalization processing. The expressions of the feature vectors feat 0 and feat 1 are:

[0043] feat 0 = Normalize(feat 0 );

[0044] feat 1 = Normalize(feat 1 )

[0045] where:

[0046] feat 0 and feat 1They are the feature vectors obtained after flattening the rough conversion feature maps at the 1 / 16 scale respectively;

[0047] feat 0 correspond to the flattened feature vectors;

[0048] feat 1 correspond to the flattened feature vectors;

[0049] Normalize refers to the L2 norm normalization operation;

[0050] Secondly, the flattened feature matrix is obtained according to the normalized feature vectors:

[0051]

[0052] Where:

[0053] R is the numerical type of the tensor;

[0054] N 0 and N 1 respectively represent that the length of the flattened sequence is H 0 ×W 0 and H 1 and W 1 ;

[0055] H 0 and W 0 respectively represent the height and width of the rough feature ;

[0056] H 1 and W 1 respectively represent the height and width of the rough feature ;

[0057] C represents the number of channels, and the number of channels of the rough feature and the rough feature are equal;

[0058] The similarity matrix 0 between the feature matrices F 1 is: For:

[0059] S i,j = F 0 [i, :]·F 1 [j, :] T

[0060] Where:

[0061] S i,j represents F 0The similarity between the feature corresponding to the i-th row and F 1 and the feature corresponding to the j-th row;

[0062] T is the transpose of the matrix;

[0063] Apply softmax normalization to the similarity matrix S in both the row and column directions, and its formula is as follows:

[0064]

[0065] where:

[0066] S i,j represents the similarity between the feature corresponding to the i-th row and F 0 and the feature corresponding to the j-th row; 1

[0067] t is the temperature parameter used to control the smoothness of softmax;

[0068] The softmax in the row direction and the column direction are respectively:

[0069]

[0070] The final bidirectional normalized similarity is: sim i,j =row i,j ·col i,j i,j ,from which the normalized similarity matrix is obtained:

[0071]

[0072] Apply the threshold τ to the normalized similarity matrix, and only retain the point pairs higher than the threshold, and eliminate the low-confidence matching points:

[0073]

[0074] Finally, apply the mutual nearest neighbor constraint to the retained point pairs: If the nearest neighbor of feature point i is j and the nearest neighbor of j is also i, then (i, j) is considered a valid match, otherwise it is regarded as a false match and deleted to ensure the uniqueness of the matching result; Based on the obtained matching result, use the RANSAC algorithm to robustly estimate the affine transformation matrix between the images to be matched.

[0075] Preferably, the specific implementation method of step 2.3) is:

[0076] 2.3.1) Use the converted rough features obtained in step 2.2.1) to perform medium-grained feature conversion on the medium-grained features; Preferably, the specific implementation method of step 2.3.1) is: Upsample the converted rough features and and perform upsampling on them, and then combine them with the medium-grained features and fuse, and then perform medium-grained feature transformation;

[0077] Preferably, the self-attention calculation in the medium-grained scale feature transformation is completed using the linear attention method; the specific implementation method of the self-attention calculation during the medium-grained feature transformation is as follows:

[0078] a.1) Window partitioning and affine transformation guidance:

[0079] First, the query feature map needs to be partitioned, and the query feature map is decomposed into query windows of a fixed size; each query window of a fixed size is represented by its four vertex coordinates, namely the upper left corner, the upper right corner, the lower right corner, and the lower left corner;

[0080] To obtain the vertex coordinates of the key-value window guided by the affine transformation matrix, a key-value window is opened at the center position of the query window, and the four vertex coordinates of the key-value window are input into the affine transformation matrix; the formula for affine transformation is as follows:

[0081] x′ = A·x + b

[0082] where:

[0083] x represents the vertex coordinates of the query window;

[0084] A and b are the linear part and the translation part of the affine transformation matrix respectively;

[0085] x′ is the coordinate of each vertex, used to lock the key-value window;

[0086] Through affine transformation, the vertex coordinates of the key-value window are mapped into the space of the key-value feature map, generating the four vertices of the key-value window;

[0087] a.2) Key-value window validity check:

[0088] After the window transformation is completed, it is necessary to verify whether the key-value window is within the valid area of the key-value feature map; for each key-value window, check whether its vertices fall within the boundaries of the feature map, and count the number of valid vertices; if the number of valid vertices meets a certain proportion, it is regarded as a valid window; otherwise, it is excluded; at the same time, it will also be jointly screened according to the valid mask of the query window;

[0089] a.3) Feature sampling and attention calculation:

[0090] For the query window, feature extraction is directly performed from the query feature matrix; for each valid key-value window, feature sampling is performed from the key-value feature map using bilinear interpolation; for each valid key-value window, the extracted query features, key features, and value features are reorganized into the form of multi-head attention, and the attention is calculated in the manner of standard attention:

[0091]

[0092] Where:

[0093] are the query, key, and value feature matrices respectively;

[0094] d is a scaling factor to alleviate the problem of vanishing gradients caused by large numerical values; after calculating the self-attention and cross-attention, the transformed medium-grained features are further processed and

[0095] 2.3.2) The double softmax method is used to perform medium-grained feature matching on the results obtained in step 2.3.1) to obtain the medium-grained feature matching result; preferably, the specific implementation manner of step 2.3.2) is exactly the same as the specific implementation manner of step 2.2.2).

[0096] Preferably, the specific implementation manner of step 2.4) is:

[0097] First, use the medium-grained feature matching result to locate and open a window of a certain size in the fine features and and calculate the heat map of the corresponding area using the features in the opened window, and process the heat map through the spatial expectation method to finally obtain the accurate coordinates of the matching points and refine the matching result.

[0098] Preferably, the specific implementation manner of step 3) is:

[0099] 3.1) Construct the total loss function of the model, and the expression of the total loss function of the model is: L = L c + L m + L f ;

[0100] Where:

[0101] L c 、L m and L f represent the loss function of the rough features, the loss function of the medium-grained features, and the loss function of the fine features respectively;

[0102] Where: L c and Lm The negative log-likelihood loss is used to measure the difference between the predicted feature matching similarity and the true matching label, and its calculation formula is:

[0103]

[0104] Where:

[0105] y i,j represents the true label of feature i and feature j;

[0106] ρ i,j is the matching similarity of feature i and feature j, that is, sim i,j ;

[0107] w ij is the matching degree weight, calculated according to the ratio of positive and negative samples;

[0108] For L f , the L2 loss function is used to optimize the exact coordinates of the matching points, and the specific formula is:

[0109]

[0110] Where:

[0111] represents the exact coordinates of the feature points predicted by the matching model;

[0112] o i represents the true coordinates of the feature points;

[0113] w i is the matching accuracy weight, measuring the uncertainty of each feature point matching, represented by the variance of the feature point matching heat map;

[0114] 3.2) Use the total loss function of the model obtained in step 3.1) to train the architecture of the image feature matching model obtained in step 2), iteratively update the weights of the model, and after the image feature matching model finally converges, obtain the weights of the image feature matching model.

[0115] An image feature matching model considering geometric prior information attention calculation constructed based on the construction method of the image feature matching model considering geometric prior information attention calculation as described above.

[0116] Application of the image feature matching model considering geometric prior information attention calculation in remote sensing image processing, autonomous driving, and / or 3D reconstruction.

[0117] The advantages of the present invention are:

[0118] The present invention provides an image feature matching model that takes into account geometric prior information attention calculation. The construction method of this model includes: 1) constructing a data set; 2) constructing an image feature matching model by means of geometric prior information attention calculation; 3) constructing a loss function for the image feature matching model, and training the image feature matching model using the data set to obtain the weights of the image feature matching model. The present invention introduces geometric prior-guided feature matching: by using the affine transformation matrix obtained from the rough matching result in the initial matching stage, important geometric prior information between images is obtained to guide subsequent attention calculation to focus on potential matching regions, thereby improving the matching success rate; at the same time, a multi-stage feature matching framework is adopted: a feature matching framework consisting of three stages of rough feature matching, medium-grained feature matching, and fine feature matching is constructed to realize the matching inference from coarse-grained to high-precision, and finally a fine and efficient matching result is obtained; in addition, through a cross-attention calculation mechanism of window division and affine transformation compensation of the matching window, robust feature association under different scene conditions is realized, especially showing stronger robustness in the presence of large viewing angle and scale changes. Compared with the existing technologies, the construction method of the image feature matching model provided by the present invention and the model obtained based on this construction method improve the attention calculation mechanism by obtaining important geometric prior information between the images to be matched, and improve the matching success rate, that is, by using the affine transformation matrix obtained in the rough feature matching stage, important geometric prior information is obtained, and through the cross-attention calculation strategy of window division and affine transformation compensation of the matching window, the accuracy of attention calculation is improved, and then the accuracy and efficiency of feature matching are significantly improved, which is applicable to high-precision image feature matching in complex scenes such as geometric deformation and large viewing angle changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0119] Figure 1 It is the overall technical framework diagram of the model provided by the present invention.

[0120] Figure 2 It is the comparison diagram of the actual effects of the model (upper) and the LoFTR model (lower) provided by the present invention under 0.95 high confidence. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0121] Example 1:

[0122] See Figure 1 , the present invention provides a construction method for an image feature matching model that takes into account geometric prior information attention calculation, and the implementation steps are elaborated in detail as follows:

[0123] The first step, construction of the data set

[0124] In this step, the main work is to establish a dataset. 368 scenes from the Megadepth dataset are selected, and 100 pairs of images in each scene are used as training data. The information composed of their depth maps and camera poses is used as the training true landmark information.

[0125] Step 2: Construction of an image feature matching model considering geometric prior information attention calculation

[0126] The overall architecture diagram of the image feature matching model considering geometric prior information attention calculation is as Figure 1 shown. The model mainly consists of four parts: local image feature extraction, rough feature matching, medium-grained feature matching, and fine feature matching. By processing the features extracted at low, medium, and high resolutions in a staged manner from coarse to fine, the final matching result is obtained. The specific implementation steps are as follows:

[0127] 2.1 Local image feature extraction (Feature Encoder)

[0128] The images in the dataset obtained in the first step are processed using a standard convolutional neural network architecture and combined with a Feature Pyramid Network (FPN) to extract multi-scale local image features. Specifically, three different-scale features are extracted from the input image, including a coarse-grained feature map at a scale of 1 / 16, a medium-grained feature map at a scale of 1 / 8, and a fine-grained feature map at a scale of 1 / 2, denoted as and This module consists of multiple convolutional layers and basic blocks, gradually extracting image features and performing downsampling. During the processing, the features are integrated in channels through 1×1 convolutions and further feature extraction and non-linear mapping are performed through 3×3 convolutions. To achieve feature fusion, bilinear interpolation is used to integrate the upsampled features, and finally, a multi-scale feature output with rich semantic information is generated.

[0129] Exemplarily, the process of local image feature extraction may specifically include:

[0130] The images in the dataset obtained in the first step are gradually passed through four layers of the ResNet backbone network to obtain intermediate feature maps at scales of 1 / 2, 1 / 4, 1 / 8, and 1 / 16 respectively; the Feature Pyramid Network is used to fuse the intermediate feature maps. Starting from the deepest intermediate feature map at the 1 / 16 scale, the number of channels is adjusted through 1×1 convolution, and then the bilinear interpolation method is used to upsample the intermediate feature map at the 1 / 16 scale to obtain the same spatial resolution as the intermediate feature map at the 1 / 8 scale. The feature map obtained by upsampling the intermediate feature map at the 1 / 16 scale is added to the intermediate feature map at the 1 / 8 scale to form a medium-grained feature map at the 1 / 8 scale; the medium-grained feature map at the 1 / 8 scale is further processed through a 3×3 convolution to extract features. After convolution, the Batch Normalization and LeakyReLU activation functions are used to enhance the non-linear expression ability, the number of output channels is adjusted through 1×1 convolution, and it is upsampled to the 1 / 4 scale using bilinear interpolation and added to the intermediate feature map at the 1 / 4 scale to form a feature map at the 1 / 4 scale; based on the feature map at the 1 / 4 scale, the previous steps are repeated, and finally a fine-grained feature map at the 1 / 2 scale is obtained; the intermediate feature map at the 1 / 16 scale is directly used as a rough feature map at the 1 / 16 scale; the rough feature map at the 1 / 16 scale, the medium-grained feature map at the 1 / 8 scale, and the fine-grained feature map at the 1 / 2 scale are used as the results of local image feature extraction, and the rough feature map at the 1 / 16 scale, the medium-grained feature map at the 1 / 8 scale, and the fine-grained feature map at the 1 / 2 scale are respectively denoted as and,

[0131] 2.2 Coarse Feature Matching

[0132] The main purpose of coarse feature matching is to obtain a relatively reliable affine transformation matrix to acquire important geometric prior information between the image pairs to be matched, providing guidance for subsequent medium feature matching. A method similar to LoFTR for feature enhancement and matching is adopted, mainly divided into two steps: coarse feature transformation and coarse feature matching.

[0133] 2.2.1 Coarse Feature Transformation

[0134] First, the following form of position encoding is performed on the rough feature map at the 1 / 16 scale:

[0135]

[0136] This is used to assign unique position information to each pixel element, which is crucial for generating matching capabilities in visually unobvious areas. Then, linear self-attention and cross-attention calculations are performed on the feature after position encoding, and its calculation formula is:

[0137]

[0138] Attention(Q, K, V) = D -1 ·(φ(Q)·(φ(K) T V))

[0139] where D is a normalization factor. φ(x) is a kernel function, defined as:

[0140] φ(x) = elu(x) + 1

[0141] Use φ(x) to transform Q and K into non - negative mappings:

[0142]

[0143] After calculating the attention, perform an element - wise multiplication of the calculated attention with the coarse feature map at the 1 / 16 scale to obtain the transformed coarse feature and and are respectively and 's transformed features.

[0144] 2.2.2 Coarse Feature Matching

[0145] After obtaining the transformed coarse features with enhanced attention, the present invention uses the double - softmax method for feature matching.

[0146] For convenient matching, first flatten the transformed coarse features and feature maps to obtain feature vectors feat 0 and feat 1 , and normalize them to the L2 norm:

[0147] feat 0 = Normalize(feat 0 ), feat 1 = Normalize(feat 1 )

[0148] Obtain the flattened feature matrix:

[0149]

[0150] where:

[0151] N 0 = H 0 ×W 0 , N 1 = H 1 ×W1 Then, the similarity matrix of the two images is calculated by dot product of the feature points.

[0152] S i,j = F 0 [i, :]·F 1 [j, :] T

[0153] Here, S i,j represents the similarity between the i-th feature and the j-th feature.

[0154] To ensure the bidirectionality of the matching, that is, each feature has a normalized probability in both directions, the softmax normalization is applied to the similarity matrix S in both the row and column directions, which is called double softmax normalization. The formula is as follows:

[0155]

[0156] where S i,j is the similarity value, and t is the temperature parameter used to control the smoothness of softmax. The softmax in the row direction and the column direction are respectively:

[0157]

[0158] The final bidirectional normalized similarity is: sim i,j = row i,j ·col i,j , and thus the normalized similarity matrix is obtained

[0159]

[0160] Furthermore, a threshold τ is applied to the normalized similarity matrix, and only the point pairs higher than the threshold (taking 0.2 in this embodiment) are retained, and the low-confidence matching points are eliminated:

[0161]

[0162] Finally, the mutual nearest neighbor constraint is applied to the retained point pairs. If the nearest neighbor of feature point i is j and the nearest neighbor of j is also i, then (i, j) is considered a valid match, otherwise it is regarded as a false match and deleted to ensure the uniqueness of the matching result.

[0163] Based on the obtained matching results, the affine transformation matrix between the images can be robustly estimated using the RANSAC algorithm.

[0164] 2.3 Medium Feature Matching

[0165] Medium feature matching mainly includes two steps: medium feature transformation and medium feature matching.

[0166] 2.3.1 Medium Feature Transformation

[0167] Transform the rough features and perform upsampling and fuse with the medium features and and then perform medium feature transformation.

[0168] The self-attention calculation in the medium-scale feature transformation is completed using the linear attention method. For the cross-attention calculation in the medium-scale feature transformation, the present invention proposes a cross-attention method based on affine transformation window compensation to improve the accuracy of cross-attention under large rotation angles and scale scaling deformations. Its main steps include: window partitioning and affine transformation guidance, window validity check, feature sampling, and attention calculation.

[0169] 1) Window partitioning and affine transformation guidance:

[0170] In the calculation of cross-attention, first, the query feature map needs to be partitioned into query windows of a fixed size (4×4 in this example). Each query window is represented by the coordinates of its four vertices, namely the upper left corner, upper right corner, lower right corner, and lower left corner. These vertex coordinates define the spatial range of the query window in the feature map and are used for subsequent attention calculations.

[0171] To obtain the vertex coordinates of the key-value window guided by affine transformation, an 8×8 key-value window is opened at the center position of the query window, and the four vertex coordinates of the key-value window are input into the affine transformation module. The formula for affine transformation is as follows:

[0172] x′ = A·x + b

[0173] where x represents the vertex coordinates of the query window, and A and b are the linear part and translation part of the affine transformation matrix, respectively. Through affine transformation, the vertex coordinates are mapped into the space of the key-value feature map, generating the four vertices of the key-value window.

[0174] 2) Key-value window validity check:

[0175] After the window transformation is completed, it is necessary to verify whether the key-value window is within the valid region of the key-value feature map. For each key-value window, check whether its vertices fall within the boundaries of the feature map and count the number of valid vertices. If the number of valid vertices meets a certain proportion (such as more than 50%), the window is considered valid; otherwise, it is excluded. At the same time, it will also be jointly screened according to the valid mask of the query window. The purpose of this step is to ensure that the subsequent attention calculation is performed within the valid region.

[0176] 3) Feature sampling and attention calculation:

[0177] The query features are directly extracted from the query feature matrix, and for each valid key-value window, feature sampling is performed from the key-value feature map using bilinear interpolation.

[0178] Then, for each valid window, the extracted query features, key features, and value features are reorganized into the form of multi-head attention. Then, the attention is calculated in the standard attention manner:

[0179]

[0180] where are the query, key, and value matrices respectively. d is a scaling factor to alleviate the problem of vanishing gradients caused by large numerical values. After calculating the self-attention and cross-attention, the transformed features can be obtained through further processing and

[0181] 2.3.2 Medium Feature Matching

[0182] The medium feature matching method uses the same method as "2.2.2 Coarse Feature Matching" to perform matching processing on and to obtain the medium feature matching result.

[0183] 2.4 Fine Feature Matching:

[0184] For fine feature matching, the same method as the fine matching of the LoFTR algorithm is adopted. Specifically, it first locates a window of a certain size (5×5 in this embodiment) in the feature map using the medium feature matching result, and combines the feature map to calculate the heat map of the corresponding region to characterize the potential distribution of the matching points. The heat map is deduced by the spatial expectation method, and the accurate coordinates of the matching points are calculated to refine the matching result.

[0185] The third step is to train the image feature matching model

[0186] Construct the total loss function of the model L = L c + L m + L f . L c , L m and L f represent the loss functions at the coarse scale, medium scale, and fine scale respectively. L c , L m use the negative log-likelihood loss to measure the difference between the predicted matching confidence and the true matching label, and its calculation formula is:

[0187]

[0188] Among them, y i,j represents the true label of feature i and feature j, and ρ i,j is the matching similarity sim between feature i and feature j i,j , and w ij is the matching degree weight, which is calculated according to the ratio of positive and negative samples. For L f , the L2 loss function is used to optimize the accurate coordinate offset of the matching points. The specific formula is:

[0189]

[0190] Among them, represents the accurate coordinate of the feature point predicted by the matching model, and o i represents the true coordinate of the feature point, and w i is the matching accuracy weight, which measures the uncertainty of each feature point matching and is represented by the variance of the feature point matching heat map. In this embodiment, the Adam optimizer is used for training, and the initial learning rate is set to 1×10 3 . The batch size is 1, and the model converges after 10 days of training on a 4090 graphics card. Attached Figure 2 shows the comparison result graph between the method of the present invention and the LoFTR algorithm. In the case of high confidence (0.95), the number of successfully matched points of LoFTR is very small, but the method of the present invention can still obtain a relatively large number of matched points and can still obtain relatively satisfactory results in the case of large geometric deformation of the image.

[0191] Meanwhile, the present invention also provides an image feature matching model constructed based on the method described above, and this image feature matching model can have good applications in remote sensing image processing, autonomous driving, and / or 3D reconstruction.

Claims

1. A method for constructing an image feature matching model taking into account geometric prior information attention calculation, characterized in that: The method for constructing an image feature matching model taking into account geometric prior information attention calculation comprises the following steps: 1) Build a dataset; 2) Use geometric prior information attention calculation to build an image feature matching model; 3) Constructing a loss function of the image feature matching model, using the data set constructed in step 1) to train the image feature matching model, and obtaining the image feature matching model weights.

2. The method for constructing an image feature matching model taking into account geometric prior information attention calculation according to claim 1, characterized in that: The step 1) specifically includes: selecting images of multiple scenes under the Megadepth dataset as a training dataset for training real landmark information; each scene includes at least 100 pairs of images.

3. The method for constructing an image feature matching model taking into account geometric prior information attention calculation according to claim 2, characterized in that: The specific implementation method of step 2) is: 2.1) Extracting local image features from the images in the data set obtained in step 1); 2.2) performing rough feature matching on the result obtained in step 2.1) to obtain a rough feature matching result; 2.3) performing medium-grained feature matching on the coarse feature matching result obtained in step 2.2) to obtain a medium-grained feature matching result; 2.4) performing fine feature matching on the medium-granularity feature matching result obtained in step 2.3) to obtain a fine feature matching result; 2.5) Construct an image feature matching model based on coarse feature matching, medium-grained feature matching, and fine feature matching.

4. The method for constructing an image feature matching model taking into account geometric prior information attention calculation according to claim 3, characterized in that: The step 2.1) is to extract multi-scale local image features from the data set obtained in step 1) based on the ResNet backbone network and in combination with the feature pyramid network; preferably, the specific implementation of the step 2.1) is: 2.1.1) The images in the dataset obtained in step 1) are gradually extracted with features through the four layers of the ResNet backbone network to obtain intermediate feature maps of 1 / 2 scale, 1 / 4 scale, 1 / 8 scale and 1 / 16 scale respectively; 2.1.2) Use the feature pyramid network to fuse the intermediate feature maps. Starting from the deepest 1 / 16 scale intermediate feature map, adjust the number of channels through 1×1 convolution, and then use the bilinear interpolation method to upsample the 1 / 16 scale intermediate feature map to obtain the same spatial resolution as the 1 / 8 scale intermediate feature map. Add the feature map obtained by upsampling the 1 / 16 scale intermediate feature map to the 1 / 8 scale intermediate feature map to form a 1 / 8 scale medium-granularity feature map. 2.1.3) The 1 / 8 scale medium-sized feature map obtained in step 2.1.2) is further extracted through a 3×3 convolution. After the convolution, Batch Normalization and LeakyReLU activation functions are used to enhance the nonlinear expression ability. The number of output channels is adjusted through 1×1 convolution, and bilinear interpolation is used to upsample to 1 / 4 scale, and then added to the 1 / 4 scale intermediate feature map to form a 1 / 4 scale feature map; 2.1.4) Based on the 1 / 4 scale feature map obtained in step 2.1.3), repeat step 2.1.3) and obtain a 1 / 2 scale fine feature map; 2.1.5) The 1 / 16 scale intermediate feature map obtained in step 2.1.1) is directly used as the 1 / 16 scale coarse feature map; the 1 / 16 scale coarse feature map, the 1 / 8 scale medium granularity feature map obtained in step 2.1.2) and the 1 / 2 scale fine feature map obtained in step 2.1.4) are used as the results of local image feature extraction, and the 1 / 16 scale coarse feature map, the 1 / 8 scale medium granularity feature map and the 1 / 2 scale fine feature map are respectively recorded as and, 5. The method for constructing an image feature matching model taking into account geometric prior information attention calculation according to claim 4, characterized in that: The specific implementation method of step 2.2) is: 2.2.1) Convert the 1 / 16 scale coarse feature map obtained in step 2.1.5) to obtain the converted coarse features; preferably, the specific implementation of step 2.2.1) is: First, the 1 / 16 scale coarse feature map obtained in step 2.1.5) is positionally encoded: Perform vector addition operation on the position code and the 1 / 16 scale coarse feature map to obtain the coarse feature after position coding; Then, linear self-attention and cross-attention are calculated for the rough features after position encoding. The calculation method of linear self-attention and cross-attention is: Attention(Q, K, V)=D -1 ·(φ(Q)·(φ(K) T V)) in: R is the numeric type of the tensor; K is the key feature matrix in the attention mechanism; V is the value feature matrix; n and dk represent the dimensions of Q and K after being processed by the function φ(x); T is the transpose of the matrix; D is the normalization factor; φ(x) is the kernel function that transforms Q and K into non-negative mappings, which is specifically defined as: φ(x)=elu(x)+1 After calculating the attention, the calculated attention is multiplied element-wise with the coarse feature map of 1 / 16 scale to obtain the converted coarse feature and Said and They are and The conversion characteristics of 2.2.2) Use the double softmax method to transform the rough features obtained in step 2.2.1) and Perform rough feature matching to obtain a rough feature matching result; preferably, the specific implementation of step 2.2.2) is: First, convert the coarse features and The feature map is flattened to obtain feature vectors feat0 and feat1, and further L2 norm normalization is performed. The expressions of the feature vectors feat0 and feat1 are: feat0 = Normalize(feat0); feat1=Normalize(feat1) in: feat0 and feat1 are the feature vectors obtained after flattening the coarsely transformed feature map at 1 / 16 scale; feat0 corresponds Flattened eigenvector; feat1 corresponds Flattened eigenvector; Normalize refers to the L2 norm normalization operation; Secondly, the flattened feature matrix is ​​obtained based on the normalized feature vector: in: R is the numeric type of the tensor; N0 and N1 indicate that the length of the flattened sequence is H0×W0 and H1 and W1 respectively; H0 and W0 represent the rough features respectively The height and width of H1 and W1 represent the rough features respectively The height and width of C represents the number of channels, the rough features and coarse features The number of channels is equal; Similarity matrix between feature matrices F0 and F1 for: S i,j =F0[i,:]·F1[j,:] T in: S i,j Indicates the similarity between the feature corresponding to the i-th row of F0 and the feature corresponding to the j-th row of F1; T is the transpose of the matrix; Apply softmax normalization to the similarity matrix S in the row and column directions respectively, and the formula is as follows: in: S i,j Indicates the similarity between the feature corresponding to the i-th row of F0 and the feature corresponding to the j-th row of F1; t is the temperature parameter, which is used to control the smoothness of softmax; The softmax in the row direction and column direction are: The final bidirectional normalized similarity is: sim i,j =row i,j ·col i,j , thus obtaining the normalized similarity matrix: Apply a threshold τ to the normalized similarity matrix, retaining only point pairs above the threshold and removing low-confidence matching points: Finally, the mutual nearest neighbor constraint is applied to the retained point pairs: if the nearest neighbor of feature point i is j, and the nearest neighbor of j is also i, then (i, j) is considered a valid match, otherwise it is considered a false match and deleted to ensure the uniqueness of the matching result; based on the obtained matching results, the RANSAC algorithm is used to robustly estimate the affine transformation matrix between the images to be matched.

6. The method for constructing an image feature matching model taking into account geometric prior information attention calculation according to claim 5 is characterized in that: the specific implementation method of step 2.3) is: 2.3.1) Using the converted coarse features obtained in step 2.2.1), the medium-grained features are converted into medium-grained features; preferably, the specific implementation method of step 2.3.1) is: and Upsample and combine with medium-grained features and Fusion followed by medium-grained feature transformation; Preferably, the self-attention calculation in the medium-granularity feature conversion is completed using a linear attention method; the specific implementation method of the self-attention calculation during the medium-granularity feature conversion is: a.1) Window division and affine transformation guidance: First, the query feature graph needs to be divided into fixed-size query windows. Each fixed-size query window is represented by its four vertex coordinates, namely, the upper left corner, the upper right corner, the lower right corner, and the lower left corner. To obtain the vertex coordinates of the key window guided by the affine transformation matrix, a key window is opened at the center of the query window, and the four vertex coordinates of the key window are input into the affine transformation matrix; the formula of the affine transformation is as follows: x′=A·x+b in: x represents the vertex coordinates of the query window; A and b are the linear part and translation part of the affine transformation matrix respectively; x′ is the coordinate of each vertex, which is used to lock the key value window; Through affine transformation, the vertex coordinates of the key-value window are mapped to the space of the key-value feature map to generate four vertices of the key-value window; a.2) Key value window validity check: After the window transformation is completed, it is necessary to verify whether the key-value window is within the valid area of ​​the key-value feature map; for each key-value window, check whether its vertices fall within the boundary of the feature map and count the number of valid vertices; if the number of valid vertices meets a certain ratio, the window is considered valid; otherwise, it is removed; at the same time, it will also be screened together according to the valid mask of the query window; a.3) Feature sampling and attention calculation: For the query window, feature extraction is performed directly from the query feature matrix; for each valid key-value window, feature sampling is performed from the key-value feature map using bilinear interpolation; for each valid key-value window, the extracted query features, key features, and value features are reorganized into the form of multi-head attention, and attention is calculated in the standard attention manner: in: They are query, key, and value matrices respectively; d is a scaling factor to alleviate the gradient vanishing problem caused by large values; after calculating the self-attention and cross-attention, further processing is performed to obtain the converted medium-granularity features and 2.3.2) Use the double softmax method to perform medium-granularity feature matching on the result obtained in step 2.3.1) to obtain a medium-granularity feature matching result; preferably, the specific implementation method of step 2.3.2) is exactly the same as the specific implementation method of step 2.2.2).

7. The method for constructing an image feature matching model taking into account geometric prior information attention calculation according to claim 6 is characterized in that: the specific implementation method of step 2.4) is: First, the medium-grained feature matching results are used to and A window of a certain size is located and opened in the image, and the features in the opened window are used to calculate the heat map of the corresponding area. The heat map is processed by the spatial expectation method to finally obtain the precise coordinates of the matching points and refine the matching results.

8. The method for constructing an image feature matching model taking into account geometric prior information attention calculation according to claim 7, characterized in that: The specific implementation method of step 3) is: 3.1) Construct the total loss function of the model, the expression of the total loss function of the model is: L = L c +L m +L f ; in: L c , L m and L f They represent the loss function of coarse features, the loss function of medium-grained features, and the loss function of fine features respectively; Where: L c and L m The negative log-likelihood loss is used to measure the difference between the predicted feature matching similarity and the true matching label, which is calculated as: in: y i,j Represents the true label of feature i and feature j; ρ i,j is the matching similarity between feature i and feature j, i.e., sim i,j ; w ij is the matching weight, which is calculated according to the ratio of positive and negative samples; For L f , the L2 loss function is used to optimize the precise coordinates of the matching points. The specific formula is: in: Indicates the precise coordinates of the feature points predicted by the matching model; o i Represents the real feature point coordinates; w i It is the matching accuracy weight, which measures the uncertainty of each feature point matching and is represented by the variance of the feature point matching heat map; 3.2) Based on the total loss function of the model obtained in step 3.1), the image feature matching model obtained in step 2) is trained, and the weights of the model are iteratively updated. After the architecture of the image feature matching model finally converges, the image feature matching model weights are obtained.

9. An image feature matching model taking into account geometric prior information and attention calculation is constructed based on the method for constructing an image feature matching model taking into account geometric prior information and attention calculation as described in any one of claims 1 to 8.

10. Application of the image feature matching model based on geometric prior information attention calculation as described in claim 9 in remote sensing image processing, autonomous driving and / or three-dimensional reconstruction.

Citation Information

Patent Citations

  • General target detection method of adaptive attention guidance mechanism

    CN111259930A

  • Construction method and system of new coronal pneumonia severe detection model based on deep learning

    CN114821254A

  • Monocular three-dimensional reconstruction method fusing attention mechanism

    CN115375844A

  • Template matching method based on heterogeneous image feature fusion, medium and equipment

    CN117115483A

  • Path planning method based on local self-attention moving window algorithm

    CN117889867A

Cited By

  • Remote sensing image feature matching and splicing method based on improved LoFTR algorithm

    CN120634880A

  • A method, apparatus, and storage medium for generating a medical image report

    CN122619241A

  • A method, apparatus, and storage medium for generating a medical image report

    CN122619241B