A multi-scale image matching method and system

By constructing a multi-scale image matching method, using the visual spatial clue extraction and fusion module, combined with Transformer and graph construction module, the problems of matching errors and high computational complexity of traditional methods in complex scenarios are solved, and efficient and accurate image matching is achieved.

CN119888285BActive Publication Date: 2025-07-22XIAMEN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510362652.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-22
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

The existing image matching methods have problems of matching errors and high computational complexity when dealing with complex scenes, especially when noise and image deformation are obvious. The deep learning-based methods are costly to achieve real-time applications during training.

Method used

A multi-scale image matching method is constructed, including the visual-spatial clue extraction module, the visual-spatial fusion module and the visual-spatial latent graph transformation module. Feature extraction and fusion are performed through convolutional neural network and Transformer layer, combining differential pooling and differential solution pooling operations, and using the multi-head self-attention mechanism and graph construction module to capture multi-scale context information to reduce the computational complexity.

Benefits of technology

It improves the accuracy and robustness of image matching, reduces the computational complexity, and makes the model more efficient when processing large-scale image data, and is suitable for image matching in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119888285B_ABST
    Figure CN119888285B_ABST
Patent Text Reader

Abstract

The present invention relates to a multi-scale image matching method and system. The method comprises the following steps: S1, obtaining a data set for image matching training; S2, constructing an image matching model, which includes a visual space cue extraction module, a visual space fusion module and a visual space latent map transformation module. The visual space cue extraction module extracts visual cues and spatial cues from the input image pair. The visual space fusion module fuses the extracted visual cues and spatial cues into the same space. The visual space latent map transformation module extracts local features from the visual space fusion features by constructing a feature map and fuses the local features with global features, so as to effectively capture multi-scale context information corresponding to the visual space; training the image matching model through the data set; S3, performing image matching on the image pair to be matched through the trained image matching model. The method and system can improve the speed and accuracy of image matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image matching, and particularly relates to a multi-scale image matching method and system. Background Art

[0002] Image matching is a technique for aligning different images, aiming to align and fuse multiple images from different times, sensors, perspectives, or modalities. The goal of matching is to align the corresponding positions of the same object in the images so that more effective analysis and processing can be carried out. It occupies an important position in the fields of computer vision and pattern recognition. Traditional methods often extract image feature points through SIFT and SURF and use algorithms such as RANSAC for image feature matching. However, these methods often perform poorly when noise and image deformation are obvious, and the computational complexity is relatively high. With the development of deep learning, deep learning-based methods such as SuperGlue have made significant progress in image feature matching. These methods improve the accuracy and robustness of matching through graph neural networks and attention mechanisms.

[0003] Image feature matching has a wide range of applications in practical applications. In Simultaneous Localization and Mapping (SLAM), through image feature matching between consecutive frames, not only can a three-dimensional map of the environment be constructed and updated in real time, but also real-time perception and response to a dynamic environment can be achieved. In terms of robot navigation, image feature matching can help the robot accurately understand the structure and features of the surrounding environment, so as to formulate precise path planning and obstacle avoidance strategies. For autonomous driving technology, image feature matching technology is the key to achieving high-precision positioning and real-time environment perception. By analyzing and processing binocular images, the autonomous driving system can quickly and accurately identify roads, traffic signs, and obstacles, thus ensuring driving safety and efficiency.

[0004] Some models aim to achieve robust visual localization by learning geometric consistency. This method is based on a graph neural network and significantly improves the reliability and accuracy of image feature matching by utilizing the geometric information in the image. For example, the key of OANet lies in its ability to accurately identify and match feature points in complex scenes, thereby achieving high-precision pose estimation.

[0005] On the other hand, ACNet and MS 2 DG-Net is a method based on the attention mechanism and multi-scale feature extraction. These two methods introduce the attention mechanism and improve the accuracy of visual position recognition by focusing on key regions in the image. ACNet uses the attention mechanism to enhance the information aggregation ability in the feature extraction process, enabling the network to more effectively identify and match important features in the image. In this way, ACNet not only improves the positioning accuracy but also enhances the efficiency of the algorithm when processing large-scale data. MS2 DG-Net was proposed by Dai et al. It is a multi-scale and multi-depth geometric network, aiming to achieve robust visual localization through multi-level feature extraction and geometric information fusion. MS 2 By extracting features at different scales and depths, DG-Net effectively captures rich geometric information in images, thereby enhancing the accuracy and robustness of image feature matching. This method is particularly suitable for complex and dynamic scenes and can achieve high-precision pose estimation in various environments. ACNet has poor performance in some low-level network layers including its asymmetric convolution, which may lead to a slight performance degradation. In addition, although there is no additional computational burden during inference, the training process may affect performance due to different gradient flows and random initializations. MS 2 The main drawback of DG-Net is the high computational cost brought by its complex graph construction and dynamic update mechanism, which makes it less efficient in real-time applications. At the same time, on some datasets (especially those with fewer outliers), its performance may not be significantly better than simple models. Summary of the Invention

[0006] The object of the present invention is to provide a multi-scale image matching method and system, which can improve the speed and accuracy of image matching.

[0007] To achieve the above object, the technical solution adopted by the present invention is: a multi-scale image matching method, including the following steps:

[0008] S1. Obtain a dataset for image matching training;

[0009] S2. Construct an image matching model, the image matching model includes a visual space cue extraction module, a visual space fusion module, and a visual space latent map transformation module. The visual space cue extraction module extracts visual cues and spatial cues from the input image pair to reduce the influence caused by occlusion, large viewing angle changes, and illumination changes between the two images. The visual space fusion module fuses the extracted visual cues and spatial cues into the same space and projects them jointly into the original space. The visual space latent map transformation module extracts local features from the visual space fusion features by constructing a feature map and fuses the local features with the global features, thereby effectively capturing the multi-scale context information corresponding to the visual space; train the image matching model through the dataset;

[0010] S3. Perform image matching on the image pair to be matched through the trained image matching model.

[0011] Further, the dataset obtained for image matching training includes the YFCC100M dataset and the SUN3D dataset.

[0012] Furthermore, the visual spatial cue extraction module includes a convolutional neural network, a feature correspondence module, a cross-attention layer, and a multi-layer perceptron MLP. The visual spatial cue extraction module extracts high-dimensional local features {F A , F B} from the input image pair {I A , I B} through the convolutional neural network, and then flattens the extracted high-dimensional local features into one-dimensional vectors through the feature correspondence module and transfers them to the cross-attention layer to generate the initial visual cues of the scene. Then, the initial visual cues are embedded into the MLP to obtain the visual cues F v for fusion. In addition, the input image pair is preprocessed by a feature detector, and then an initial correspondence set I C is established through the nearest neighbor matching strategy. The obtained initial correspondence set I C is embedded into another MLP to extract deep features and used as the spatial cue F s .

[0013] Furthermore, the visual spatial fusion module performs interactive modeling on the visual cues and spatial cues through the Transformer layer, and combines the differential pooling and differential unpooling operations to generate the fused visual spatial features. The specific implementation method is as follows:

[0014] Input the visual cue F v and the spatial cue F s ; reduce the dimension of the spatial cue and extract the compressed spatial features through differential pooling:

[0015] S pooled = DiffPool(F s )

[0016] where S pooled is the spatial feature after dimension reduction, and DiffPool( ) represents the differential pooling operation;

[0017] Concatenate the visual cue and the spatial feature after dimension reduction to form the joint feature VS:

[0018] VS = Concat(S pooled , F v )

[0019] where Concat( ) represents concatenation in the feature dimension, and VS is the concatenated joint feature;

[0020] Then, use the Transformer layer to perform global modeling on the joint feature of the visual and spatial cues to generate the interacted feature:

[0021] VS' = Transformer(VS)

[0022] Among them, Transformer( ) represents the Transformer function, which models the relationship between feature points using the multi-head self-attention mechanism, and VS' is the feature after interaction;

[0023] Then, the feature VS' after interaction is decomposed into updated visual cues and spatial cues:

[0024] VS' = {S' updated , V' updated}

[0025] Among them, S' updated represents the updated spatial cue, and V' updated represents the updated visual cue;

[0026] Fuse the initial cues with the updated cues:

[0027] S updated = F s + S' updated

[0028] V updated = F v + V' updated

[0029] Extract features from the updated and fused spatial cues and visual cues through the ResNet encoder:

[0030] S encoded = ResNetBlock(S updated )

[0031] V encoded = ResNetBlock(V updated )

[0032] Among them, S encoded is the encoded spatial feature, and V encoded is the encoded visual feature;

[0033] Add the encoded spatial feature and visual feature to generate a joint feature:

[0034] VS encoded = S encoded + V encoded

[0035] Map the joint feature back to the original space through de-pooling:

[0036] F VS = DiffUnpool(Scues , VS encoded )

[0037] Among them, DiffPool( ) represents the differential de-pooling operation, which is used to map the joint feature back to the original spatial resolution; F VS is the visual spatial fusion feature finally output by the visual spatial fusion module.

[0038] Furthermore, the visual spatial latent graph transformation module includes a graph construction module, a graph boundary node extraction module BDB, an aggregation module, an encoding module, a multi-head self-attention mechanism module MHSA, and an FFN network; among them, the graph construction module constructs a feature graph by calculating the relationship between feature points to capture the local geometric relationship between nodes; the BDB module extracts multi-scale context information through the fusion of local and global features; the aggregation module fuses the local information and global information in the feature to enhance the overall feature expression; the encoding module performs deep encoding on the input feature to extract high-level semantic features while retaining important input information; the MHSA module is used to capture the global context information and the complex relationship between feature points; the FFN network is used to further extract and non-linearly transform the channel dimension information of each feature point.

[0039] Furthermore, the implementation method of the graph construction module is as follows:

[0040] Construct a KNN-based graph according to the Euclidean distance between feature points in the joint visual space;

[0041] First, determine the K neighbors of the feature points: for each feature point in the joint visual space f i , select the K feature points with the closest Euclidean distance to it as its nearest neighbor feature points, and its nearest neighbor feature set is represented as follows:

[0042] V i = { f i1 ,f i2 ,…,f ik}

[0043] Among them, V i represents the nearest neighbor feature set of the feature point f i , f ik represents the k-th nearest neighbor feature point of the feature point f i ;

[0044] Then calculate the edges between each feature point and its neighbors:

[0045] Connect the feature points e ij through an edge f i to its nearest neighbor feature point f ij , and define the content of the edge by calculating the difference between the two feature points:

[0046] e ij = f i , f ij − f i

[0047] where f ij is the j-th nearest neighbor feature point of the feature point f i , and f ij − f i is the difference vector between the two feature points;

[0048] Construct a global graph G in from the obtained edges and feature points, which is represented as:

[0049] G in = {G1, ..., G i , ..., G N}

[0050] where G1, ..., G i , ..., G N are the local feature graphs corresponding to N feature points respectively, which contain the connections of each feature point and its nearest neighbor feature points.

[0051] Furthermore, the implementation method of the BDB module is as follows:

[0052] The BDB module includes a boundary enhancement module, a KNN projection branch, a local attention branch, a global attention branch, and a cross-layer fusion module;

[0053] The boundary enhancement module is used for global feature extraction, its input is G in , and its output is x1; the processing process of the boundary enhancement module is represented as:

[0054]

[0055] where ​It is denoted that the number of channels is mapped from C to C using a 1×1 convolution, BN( ) represents batch normalization processing, ReLU( ) represents the ReLU activation function, and x1 represents the output global feature;

[0056] The KNN projection branch is used for local feature extraction, with its input being x1 and output being x local ; The processing process of the KNN projection branch is expressed as:

[0057]

[0058] where, It is denoted that the number of channels is mapped from C to C / 2 using a 1×1 convolution, BN( ) represents batch normalization processing, ReLU( ) represents the ReLU activation function, and x local represents the output local feature;

[0059] The local attention branch is used for local graph convolution feature extraction, with its input being x1 and output being x localatt ; The processing process of the local attention branch is expressed as:

[0060] x localatt = DGCNN Layer (G(x1))

[0061] where, G( ) represents constructing local graph features and extracting local structure information; DGCNN Layer ( ) represents performing graph convolution operations using the DGCNN layer;

[0062] The global attention branch is used for global context feature extraction, with its input being x1 and output being x globalatt ; The processing process of the global attention branch is expressed as:

[0063]

[0064] First, perform adaptive average pooling to compress the spatial size to 1×1, then use a 1×1 convolution to map the number of channels from C to C / r; then perform batch normalization processing, followed by passing through the ReLU activation function, and then use a 1×1 convolution to map the number of channels from C / r back to C / 2, and after performing batch normalization processing, output x globalatt ;

[0065] The cross-layer fusion module is used for feature fusion, which performs an element-wise addition operation on the three features x globalatt 、x localatt and x local to output the fused feature map x lg , expressed as:

[0066] x lg= x globalatt + x localatt + x local

[0067] Then, for the fused feature map x lg apply Sigmoid, which is expressed as:

[0068] G out = σ·x lg

[0069] where σ represents the Sigmoid function, and G out is the output of the BDB module.

[0070] Furthermore, the implementation methods of the aggregation module, encoding module, MHSA module, and FFN network are as follows:

[0071] The aggregation module performs a neighborhood aggregation operation on the feature G output by the BDB module out to capture more geometric and semantic information:

[0072] F Agg = Agg(G out )

[0073] where F Agg represents the feature output by the aggregation module, and Agg( ) represents the neighborhood aggregation operation; the feature after aggregation by the aggregation module contains the relationship between the current feature and its neighbors;

[0074] The encoding module performs deep encoding on the output feature F of the aggregation module Agg to extract high-level semantic features and maintain the integrity of the input information:

[0075] F Encoded = Encoder(F Agg )

[0076] where F Encoded represents the feature output by the encoding module, and Encoder( ) represents the encoding operation; the encoder consists of multiple ResNet blocks, and each ResNet block performs a convolution operation to extract more feature information;

[0077] The MHSA module calculates the geometric relationship between feature points by introducing a length similarity matrix to ensure the accuracy of feature matching:

[0078] F MHSA = MHSA(F Encoded , M ls )

[0079] where F MHSADenote the features output by the MHSA module, MHSA( ) represents the multi-head self-attention mechanism operation, M ls represents the length similarity matrix; the length similarity matrix M ls is calculated as follows:

[0080] Preprocess the input image pair {I A , I B} using a feature detector, and then establish an initial correspondence set I C ={c1, c2, ..., c i , ..., c j ,..., c N} through the nearest neighbor matching strategy, where c i and c j represent the i-th and j-th feature point pairs in the initial correspondence set I C respectively, N represents the number of feature point pairs in the initial correspondence set I C ; c i =(p i A , p i B ), c j =(p j A , p j B ), where p i A and p i B represent the feature points in the input image I A and the input image I B respectively. Then, the element m ls in the length similarity matrix M i,j is calculated according to the following formula:

[0081]

[0082] where, || || represents the norm calculation, and | | represents the absolute value;

[0083] The features output by the MHSA module are split into two paths. One path is input to the FFN network, and the other path is concatenated with the features output by the FFN network to obtain the final output features;

[0084] The FFN network receives the features output by the MHSA module and performs a series of transformations to further extract features:

[0085]

[0086] Among them, Linear1 is the first linear transformation layer, which maps the features output by the MHSA module to a higher-dimensional space to enhance the representation ability; ReLU( ) is the ReLU activation function, which processes the features output by Linear1 to enhance the non-linear characteristics and obtains the activated features; Linear2 is the second linear transformation layer, which is used to map the activated features back to the final output space.

[0087] Furthermore, a hybrid loss function is used to supervise the training process of the image matching model:

[0088]

[0089] Among them, represents the classification loss; α is a hyperparameter that balances the classification loss and the essential matrix loss; represents the essential matrix loss, which is used to measure the difference between the predicted essential matrix and the true essential matrix. After feature extraction by the image matching model, the predicted essential matrix is calculated based on the matched feature point pairs, and then the essential matrix loss is calculated;

[0090] The classification loss is expressed as:

[0091]

[0092] Among them, H( ) represents the binary cross-entropy loss function; o t represents the relevant weight at the t-th iteration; y t represents the weakly supervised label; ω t is the adaptive temperature vector, ⊙ represents the Hadamard product, and λ is the number of iterations.

[0093] The present invention also provides a multi-scale image matching system, including a memory, a processor, and computer program instructions stored on the memory and capable of being run by the processor. When the processor runs the computer program instructions, the above-mentioned method can be implemented.

[0094] Compared with the prior art, the present invention has the following beneficial effects: The present invention provides a multi-scale image matching method and system. By constructing an image matching model that combines visual space cue extraction, visual space fusion module, and visual space potential map transformation module, the method and system can extract and fuse visual cues and spatial cues for the input image pair. On this basis, local features are further extracted from the visual space fusion features and fused with the global features, thereby effectively capturing the multi-scale context information corresponding to the visual space, making full use of the local geometric information and global context information in the image, solving the matching error problem of traditional methods when dealing with complex scenes, and improving the robustness and accuracy of the method in complex environments; in addition, through local and global cross-layer multi-branch fusion, the computational complexity is effectively reduced, making the model more efficient when processing large-scale image data. BRIEF DESCRIPTION OF THE DRAWINGS

[0095] Figure 1 is the architecture diagram of the image matching model in the embodiment of the present invention;

[0096] Figure 2 is the specific implementation structure diagram of the BDB module in the embodiment of the present invention;

[0097] Figure 3 is the visualization comparison diagram in the YFCC100M outdoor scene in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0098] The present invention will be further described below with reference to the drawings and embodiments.

[0099] It should be noted that the following detailed description is exemplary and is intended to provide further description of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0100] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0101] This embodiment provides a multi-scale image matching method, including the following steps:

[0102] S1. Obtain a data set for image matching training;

[0103] S2. Construct as Figure 1The image matching model shown, the image matching model includes a visual-spatial clue extraction module, a visual-spatial fusion module, and a visual-spatial latent graph transformation module. The visual-spatial clue extraction module extracts visual clues and spatial clues from the input image pair to reduce the impact caused by occlusion, large view angle changes, and illumination changes between the two images. The visual-spatial fusion module fuses the extracted visual clues and spatial clues into the same space and projects them together into the original space. The visual-spatial latent graph transformation module (VSLG Former, Visual-Spatial Latent Graph Transformer) extracts local features from the visual-spatial fusion features by constructing a feature map and fuses the local features with the global features, thereby effectively capturing the multi-scale context information corresponding to the visual space; the image matching model is trained through a dataset;

[0104] S3. Perform image matching on the image pair to be matched through the trained image matching model.

[0105] In this embodiment, the dataset obtained for image matching training includes the YFCC100M dataset and the SUN3D dataset. The scenes included in the dataset are divided into two categories: known scene groups and unknown scene groups. In this embodiment, evaluations are performed on both unknown scenes and known scenes.

[0106] The visual-spatial clue extraction module includes a convolutional neural network, a correspondence feature module (Correspondence Feature Module), a cross-attention layer, and a multi-layer perceptron MLP. The visual-spatial clue extraction module extracts high-dimensional local features {F A , F B} from the input image pair {I A , I B} through the convolutional neural network, and then flattens the extracted high-dimensional local features into one-dimensional vectors through the correspondence feature module and passes them to the cross-attention layer to generate the initial visual clues of the scene; then, the initial visual clues are embedded into the MLP to obtain the visual clues F v for fusion; in addition, the input image pair is preprocessed using feature detectors (such as SIFT and SuperPoint), and then an initial correspondence set I C is established through the nearest neighbor matching strategy; the obtained initial correspondence set I C is embedded into another MLP to extract deep features and used as spatial clues F s . In this embodiment, the convolutional neural network adopts the ResNet34 convolutional neural network architecture.

[0107] The visual-spatial fusion module models the interaction between visual cues and spatial cues through Transformer layers, and combines differential pooling and differential unpooling operations to generate fused visual-spatial features. The specific implementation method of the visual-spatial fusion module is as follows.

[0108] Input visual cues and spatial cues , where B is the batch size, D v is the dimension of the visual feature, D s is the dimension of the spatial feature.

[0109] Reduce the dimension of the spatial cues, and extract the compressed spatial features through differential pooling:

[0110] S pooled = DiffPool(F s )

[0111] where S pooled is the spatial feature after dimension reduction, and DiffPool( ) represents the differential pooling operation.

[0112] Concatenate the visual cues and the spatial features after dimension reduction to form the joint feature VS:

[0113] VS = Concat(S pooled , F v )

[0114] where Concat( ) represents concatenation in the feature dimension, is the joint feature after concatenation.

[0115] Then, use the Transformer layer to globally model the joint feature of visual and spatial cues to generate the feature after interaction:

[0116] VS' = Transformer(VS)

[0117] where Transformer( ) represents the Transformer function, which uses the multi-head self-attention mechanism to model the relationship between feature points, is the feature after interaction.

[0118] Subsequently, decompose the feature VS' after interaction into updated visual cues and spatial cues:

[0119] VS' = {S' updated , V' updated}

[0120] where S' updated represents the updated spatial cues, V'updated Represents updated visual cues.

[0121] Fuse the initial cues with the updated cues:

[0122] S updated = F s + S' updated

[0123] V updated = F v + V' updated

[0124] Extract features from the updated and fused spatial and visual cues through a ResNet encoder:

[0125] S encoded = ResNetBlock(S updated )

[0126] V encoded = ResNetBlock(V updated )

[0127] Where S encoded is the encoded spatial feature and V encoded is the encoded visual feature.

[0128] Add the encoded spatial and visual features to generate a combined feature:

[0129] VS encoded = S encoded + V encoded

[0130] Map the combined feature back to the original space through unpooling:

[0131] F VS = DiffUnpool(S cues , VS encoded )

[0132] Where DiffPool( ) represents the differential unpooling operation for mapping the combined feature back to the original spatial resolution; is the visually spatially fused feature finally output by the visual spatial fusion module, and D is the fused dimension.

[0133] The visual space potential map transformation module includes a graph construction module, a graph boundary node extraction module BDB (Boundary Detection and Extraction Block), an aggregation module, an encoding module, a multi-head self-attention mechanism (MHSA) module, and an FFN network. Among them, the graph construction module constructs a feature map by calculating the relationships between feature points, capturing the local geometric relationships between nodes. The BDB module extracts multi-scale context information through the fusion of local and global features. The aggregation module fuses the local information and global information in the features to enhance the overall feature representation. The encoding module performs deep encoding on the input features to extract high-level semantic features while retaining important input information. The MHSA module is used to capture global context information and complex relationships between feature points. The FFN network is used to further extract and non-linearly transform the channel dimension information of each feature point.

[0134] The specific implementation method of the graph construction module is as follows.

[0135] Construct a KNN-based graph according to the Euclidean distance between feature points in the joint visual space.

[0136] a) First, determine the K neighbors of the feature points: For each feature point in the joint visual space f i , select the K feature points with the closest Euclidean distance to it as its nearest neighbor feature points, and its nearest neighbor feature set is represented as follows:

[0137] V i = { f i1 ,f i2 ,…,f ik}

[0138] Among them, V i represents the nearest neighbor feature set of the feature point f i , f ik represents the k-th nearest neighbor feature point of the feature point f i .

[0139] b) Then calculate the edges between each feature point and its neighbors:

[0140] Connect the feature point e ij to its nearest neighbor feature point f i through the edge f ij, and define the content of the edge by calculating the difference between two feature points:

[0141] e ij = f i , f ij - f i

[0142] wherein, f ij is the j-th nearest neighbor feature point of the feature point f i , and f ij - f i is the difference vector between two feature points.

[0143] Construct a global graph G from the obtained edges and feature points in , expressed as:

[0144] G in = {G1, ..., G i , ..., G N}

[0145] wherein, G1, ..., G i , ..., G N are the local feature graphs corresponding to N feature points respectively, which contain the connections of each feature point and its nearest neighbor feature points.

[0146] As Figure 2 shown, the BDB module includes a boundary enhancement module (BE, Boundary Extraction Grid), a KNN projection branch, a local attention branch, a global attention branch, and a cross-layer fusion module. The specific implementation method of the BDB module is as follows.

[0147] The boundary enhancement module is used for global feature extraction, and its input is G in ∈ℝ B×C×H×W , and the output is x1∈ℝ B×C×H×W , where B is the batch size, C is the number of channels, H and W are the height and width of the image; the processing process of the boundary enhancement module is expressed as:

[0148]

[0149] wherein, ​It is denoted that the number of channels is mapped from C to C using a 1×1 convolution, BN( ) represents batch normalization processing, ReLU( ) represents the ReLU activation function, and x1 represents the output global feature.

[0150] The KNN projection branch is used for local feature extraction, with its input being x1 and output being x local ; The processing process of the KNN projection branch is expressed as:

[0151]

[0152] Among them, It is denoted that the number of channels is mapped from C to C / 2 using a 1×1 convolution, BN( ) represents batch normalization processing, ReLU( ) represents the ReLU activation function, and x local represents the output local feature.

[0153] The local attention branch is used for local graph convolution feature extraction, with its input being x1 and output being x localatt ∈ℝ B×C / 2×H×W ; The processing process of the local attention branch is expressed as:

[0154] x localatt = DGCNN Layer (G(x1))

[0155] Among them, G( ) represents constructing local graph features through the get_graph_feature function to extract local structure information; DGCNN Layer ( ) represents performing graph convolution operations using the DGCNN layer.

[0156] The global attention branch is used for global context feature extraction, with its input being x1 and output being x globalatt ∈ℝ B×C / 2×1×1 ; The processing process of the global attention branch is expressed as:

[0157]

[0158] First, perform adaptive average pooling to compress the spatial size to 1×1, then use a 1×1 convolution to map the number of channels from C to C / r; then perform batch normalization processing, and then pass through the ReLU activation function, and then use a 1×1 convolution to map the number of channels from C / r back to C / 2, and after performing batch normalization processing, output x globalatt .

[0159] The cross-layer fusion module is used for feature fusion, which performs element-wise addition operations on the three features x globalatt , x localatt and x local to output the fused feature map xlg ∈ ℝ B×C / 2×1×1 is expressed as:

[0160] x lg = x globalatt + x localatt + x local

[0161] Then, for the fused feature map x lg apply Sigmoid, which is expressed as:

[0162] G out = σ·x lg

[0163] where σ represents the Sigmoid function, and G out is the output of the BDB module.

[0164] The specific implementation methods of the aggregation module, encoding module, MHSA module, and FFN network are described as follows respectively.

[0165] The aggregation module performs a neighborhood aggregation operation on the feature G output by the BDB module out to capture more geometric and semantic information:

[0166] F Agg = Agg(G out )

[0167] where F Agg represents the feature output by the aggregation module, and Agg( ) represents the neighborhood aggregation operation; the feature after aggregation by the aggregation module contains the relationship between the current feature and its neighbors.

[0168] The encoding module performs deep encoding on the output feature F of the aggregation module Agg to extract high-level semantic features and maintain the integrity of the input information:

[0169] F Encoded = Encoder(F Agg )

[0170] where F Encoded represents the feature output by the encoding module, and Encoder( ) represents the encoding operation; the encoder consists of multiple ResNet blocks, and each ResNet block performs a convolution operation to extract more feature information.

[0171] In this embodiment, a multi-head self-attention mechanism (MHSA) module is designed to localize and fuse the global to each correspondence. The MHSA module calculates the geometric relationship between feature points by introducing a length similarity matrix to ensure the accuracy of feature matching:

[0172] F MHSA = MHSA(F Encoded , M ls )

[0173] Among them, F MHSA represents the feature output by the MHSA module, MHSA( ) represents the multi-head self-attention mechanism operation, and M ls represents the length similarity matrix.

[0174] For the length similarity matrix M ls , the calculation method is as follows:

[0175] Preprocess the input image pair {I A , I B} using a feature detector (such as SIFT and SuperPoint), and then establish an initial correspondence set I C = {c1, c2,..., c i ,..., c j ,..., c N} through the nearest neighbor matching strategy, where c i , c j respectively represent the i-th and j-th feature point pairs in the initial correspondence set I C , and N represents the number of feature point pairs in the initial correspondence set I C ; c i = (p i A , p i B ), c j = (p j A , p j B ), where p i A , p i B respectively represent the feature points in the input image I A , the input image I B . Then, the element m ls in the length similarity matrix M i,j is calculated according to the following formula:

[0176]

[0177] Among them, || || represents the norm calculation, that is, calculating the length of the vector or the Euclidean distance between two points, and | | represents taking the absolute value to ensure that the calculated value is non-negative.

[0178] The features output by the MHSA module are split into two paths. One path is input into the FFN network, and the other path is concatenated with the features output by the FFN network to obtain the final output features.

[0179] The FFN network receives the features output by the MHSA module and performs a series of transformations to further extract features:

[0180]

[0181] Among them, Linear1 is the first linear transformation layer, which maps the features output by the MHSA module to a higher-dimensional space to enhance the representation ability; ReLU( ) is the ReLU activation function, which processes the features output by Linear1 to enhance the non-linear characteristics and obtains the activated features; Linear2 is the second linear transformation layer, which is used to map the activated features back to the final output space.

[0182] In this embodiment, a hybrid loss function is adopted to supervise the training process of the image matching model:

[0183]

[0184] Among them, represents the classification loss; α is a hyperparameter that balances the classification loss and the essential matrix loss; represents the essential matrix loss, which is used to measure the difference between the predicted essential matrix and the true essential matrix. After feature extraction by the image matching model, the predicted essential matrix is calculated based on the matched feature point pairs, and then the essential matrix loss is calculated.

[0185] The classification loss is expressed as:

[0186]

[0187] Among them, H( ) represents the binary cross-entropy loss function; o t represents the relevant weight at the t-th iteration; y t represents the weakly supervised label, which is selected as positive under the epipolar distance threshold 10 −4 ; ω t is the adaptive temperature vector, ⊙ represents the Hadamard product, and λ is the number of iterations.

[0188] In this embodiment, this method is compared with existing deep learning methods. This embodiment selects multiple benchmark models, including the traditional RANSAC algorithm, graph-based methods such as MS 2DG-Net, as well as a series of deep learning-based models such as PointNet++, PointCN, and MSA-Net, etc. Our goal is to evaluate the performance of this method and existing methods under different scenarios, especially their performance in known and unknown scenarios. Through testing on the YFCC100M and SUN3D datasets, it can be seen that this method is superior to existing methods in terms of the number of parameters and accuracy, and the comparison results are shown in Table 1. In the YFCC100M dataset, whether in known or unknown scenarios, this method is significantly higher than other methods. In contrast, the traditional RANSAC algorithm is inferior to this model in both the number of parameters and accuracy. For example, the accuracy of RANSAC in the YFCC100M dataset is only 9.07%. And for graph-based methods such as MS 2 DG-Net, although the number of parameters is reduced, it still fails to exceed in terms of accuracy. This method shows strong performance in both known and unknown scenarios. Specifically, this method reaches accuracies of 47.95% and 61.85% in known and unknown scenarios of the YFCC100M dataset respectively. This method also shows sub-optimal performance in the SUN3D dataset. The high efficiency of this method benefits from the aggregation strategies in the global and local fusion modules. These strategies not only improve the efficiency of the sampling process, but also improve the computational efficiency by identifying multi-branch connections in the graph, and effectively reduce the impact of pseudo-outliers.

[0189] Table 1 Performance comparison table between this method and existing methods

[0190]

[0191] This embodiment also gives the area under the cumulative error curve (AUC) under the maximum rotation and displacement errors, and sets thresholds of 5°, 10°, and 20°. We choose the nearest neighbor (NN) method as the baseline method, which retains all assumed correspondences. Then, this method is compared with several classic outlier removal methods. In addition, it includes learning-based methods that directly classify correspondences, and other methods that consider consistency and smoothness. During the evaluation process, we tested different robust model estimators, including the weighted eight-point algorithm and the RANSAC algorithm, for calculating the pose relationship of predicted inliers. It should be noted that the weighted eight-point algorithm needs to calculate the inlier probability for each correspondence, so it is not applicable to classic methods. All results are shown in Table 2. Considering that the accuracy of camera pose estimation is affected by the quality of the initial correspondences, when performing RANSAC post-processing under the same conditions, this method is more accurate than all other methods in the same AUC parameter comparison.

[0192] Table 2 Results of relative pose estimation using the weighted eight-point algorithm / RANSAC (AUC scores at 5°, 10°, and 20° thresholds in the unknown scenes of YFCC100M)

[0193]

[0194] Figure 3 This is the visualization comparison chart in the outdoor scene of YFCC100M in this embodiment. Figure 3 In it, (a), (b), and (c) are the results of OANet, VSformer, and this method in sequence. If the correspondence represents a true positive, it is drawn in green, and if it represents a false positive, it is drawn in red.

[0195] As Figure 3 shown, the visualization of this method achieves the best performance in this comparison. This is because of the advantages of the global and local strategies of this method, making this method more effective in dealing with details.

[0196] This embodiment also provides a multi-scale image matching system, including a memory, a processor, and computer program instructions stored on the memory and capable of being run by the processor. When the processor runs the computer program instructions, the above-mentioned method can be implemented.

[0197] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0198] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0199] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes Figure 1 one or more processes and / or blocks Figure 1 specified in the block or blocks.

[0200] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes Figure 1 one or more processes and / or blocks Figure 1 specified in the block or blocks.

[0201] As described above, the above are only the preferred embodiments of the present invention, and are not intended to limit the present invention in any other form. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A multi-scale image matching method, characterized in that, It includes the following steps: S1. Obtain a dataset for image matching training; S2. Construct an image matching model, which includes a visual-spatial cue extraction module, a visual-spatial fusion module, and a visual-spatial latent map transformation module. The visual-spatial cue extraction module extracts visual cues and spatial cues from the input image pair. The visual-spatial fusion module fuses the extracted visual cues and spatial cues into the same space and projects them jointly into the original space. The visual-spatial latent map transformation module extracts local features from the visual-spatial fusion features by constructing a feature map and fuses the local features with the global features; train the image matching model with the dataset; S3. Perform image matching on the image pair to be matched by the trained image matching model; The visual-spatial latent map transformation module includes a graph construction module, a graph boundary node extraction module BDB, an aggregation module, an encoding module, a multi-head self-attention mechanism module MHSA, and an FFN network. Among them, the graph construction module constructs a feature map by calculating the relationships between feature points and captures the local geometric relationships between nodes. The BDB module extracts multi-scale context information through the fusion of local and global features. The aggregation module fuses the local information and global information in the features to enhance the overall feature expression. The encoding module performs deep encoding on the input features to extract high-level semantic features while retaining important input information. The MHSA module is used to capture global context information and the complex relationships between feature points. The FFN network is used to further extract and perform non-linear transformation on the channel dimension information of each feature point; The implementation method of the BDB module is: The BDB module includes a boundary enhancement module, a KNN projection branch, a local attention branch, a global attention branch, and a cross-layer fusion module; The boundary enhancement module is used for global feature extraction, and its input is the global graph G in , and the output is x1; the processing process of the boundary enhancement module is expressed as: Among them, represents mapping the number of channels from C to C using a 1×1 convolution, BN( ) represents batch normalization processing, ReLU( ) represents the ReLU activation function, and x1 represents the output global feature; The KNN projection branch is used for local feature extraction, with its input being x1 and output being x local ; The processing process of the KNN projection branch is expressed as: Among them, means mapping the number of channels from C to C / 2 using a 1×1 convolution, BN( ) represents batch normalization processing, ReLU( ) represents the ReLU activation function, and x local represents the output local feature; The local attention branch is used for local graph convolution feature extraction, with its input being x1 and output being x localatt ; The processing process of the local attention branch is expressed as: x localatt = DGCNN Layer (G(x1)) Among them, G( ) represents constructing local graph features and extracting local structure information; DGCNN Layer ( ) represents performing graph convolution operations using DGCNN layers; The global attention branch is used for global context feature extraction. Its input is x1 and its output is x globalatt ; The processing process of the global attention branch is as follows: First, perform adaptive average pooling to compress the spatial size to 1×1, and then use a 1×1 convolution to map the number of channels from C to C / r; then perform batch normalization, and then pass through the ReLU activation function. Then use a 1×1 convolution to map the number of channels from C / r back to C / 2. After batch normalization, the output is x globalatt ; The cross-layer fusion module is used for feature fusion, which performs an element-wise addition operation on three features x globalatt , x localatt and x local to output the fused feature map x lg , expressed as: x lg = x globalatt + x localatt + x local Then apply the Sigmoid function to the fused feature map x lg which is expressed as: G out = σ·x lg Among them, σ represents the Sigmoid function, and G out is the output of the BDB module.

2. The multi-scale image matching method according to claim 1, characterized in that The dataset obtained for image matching training includes the YFCC100M dataset and the SUN3D dataset.

3. A multi-scale image matching method according to claim 1, characterized in that, The visual spatial cue extraction module includes a convolutional neural network, a feature correspondence module, a cross-attention layer, and a multi-layer perceptron MLP. The visual spatial cue extraction module extracts high-dimensional local features {F A , F B} from the input image pair {I A , I B} through the convolutional neural network, and then flattens the extracted high-dimensional local features into one-dimensional vectors and passes them to the cross-attention layer to generate the initial visual cues of the scene. Then, the initial visual cues are embedded into the MLP to obtain the visual cues F v for fusion. In addition, the input image pair is preprocessed by a feature detector, and then an initial correspondence set I C is established through the nearest neighbor matching strategy. The obtained initial correspondence set I C is embedded into another MLP to extract deep features and used as the spatial cues F s .

4. A multi-scale image matching method according to claim 1, characterized in that The visual-spatial fusion module performs interactive modeling on visual cues and spatial cues through a Transformer layer and generates fused visual-spatial features by combining difference pooling and difference unpooling operations. Its specific implementation method is: Input visual cue F v and spatial cue F s ; perform dimensionality reduction on the spatial cue and extract compressed spatial features through differential pooling: S pooled = DiffPool(F s ) Among them, S pooled is the spatial feature after dimensionality reduction, and DiffPool( ) represents the differential pooling operation; Concatenate the visual cues and the dimension-reduced spatial features to form a joint feature VS: VS = Concat(S pooled , F v ) Among them, Concat( ) represents concatenation in the feature dimension, and VS is the concatenated joint feature; Then, use the Transformer layer to perform global modeling on the joint feature of visual and spatial cues to generate an interactive feature: VS' = Transformer(VS) Among them, Transformer( ) represents the Transformer function, which uses a multi-head self-attention mechanism to model the relationships between feature points, and VS' is the interactive feature; Then, decompose the interactive feature VS' into updated visual cues and spatial cues: VS' = {S' updated , V' updated} Among them, S' updated represents the updated spatial cue, and V' updated represents the updated visual cue; Fuse the initial cues with the updated cues: S updated = F s + S' updated V updated = F v + V' updated Extract features from the updated and fused spatial cues and visual cues through a ResNet encoder: S encoded = ResNetBlock(S updated ) V encoded = ResNetBlock(V updated ) Among them, S encoded is the encoded spatial feature, and V encoded is the encoded visual feature; Add the encoded spatial features and visual features to generate a joint feature: VS encoded = S encoded + V encoded Map the joint features back to the original space through unpooling: F VS = DiffUnpool(S cues , VS encoded ) Among them, DiffPool( ) represents the differential pooling operation, which is used to map the joint feature back to the original spatial resolution; F VS is the visually spatially fused feature finally output by the visually spatially fused module.

5. A multi-scale image matching method according to claim 1, characterized in that The implementation method of the graph construction module is as follows: Construct a KNN-based graph according to the Euclidean distance between feature points in the joint visual space; First, determine the K nearest neighbors of the feature points: For each feature point in the joint visual space f i , select the K feature points with the closest Euclidean distance to it as its nearest neighbor feature points. The set of its nearest neighbor features is represented as follows: V i = { f i1 ,f i2 ,…,f ik} Among them, V i represents the nearest neighbor feature set of the feature point f i ; f ik represents the k-th nearest neighbor feature point of the feature point f i ; Then calculate the edges of each feature point and its neighbors: Through the edge e ij Connect feature points f i with its nearest neighbor feature points f ij , and define the content of the edge by calculating the difference between the two feature points: e ij = [ f i , f ij - f i ] Among them, f ij is the j-th nearest neighbor feature point f i of the feature point, f ij - f i is the difference vector between two feature points; A global graph G is constructed from the obtained edges and feature points in , denoted as: G in = {G1, ..., G i , ..., G N} Among them, G1, ..., G i , ..., G N are the local feature maps corresponding to N feature points, respectively, which contain the connections of each feature point and its nearest neighbor feature points.

6. A multi-scale image matching method according to claim 1, characterized in that, The implementation methods of the aggregation module, encoding module, MHSA module, and FFN network are as follows: The aggregation module performs neighborhood aggregation operations on the feature G output by the BDB module out to capture more geometric and semantic information: F Agg = Agg(G out ) Among them, F Agg represents the feature output by the aggregation module, and Agg( ) represents the neighborhood aggregation operation; the feature after aggregation by the aggregation module contains the relationship between the current feature and its neighbors; The encoding module performs deep encoding on the output feature F of the aggregation module Agg to extract high-level semantic features and maintain the integrity of the input information: F Encoded = Encoder(F Agg ) Among them, F Encoded represents the feature output by the encoding module, and Encoder( ) represents the encoding operation; the encoder consists of multiple ResNet blocks, and each ResNet block performs a convolution operation to extract more feature information; The MHSA module calculates the geometric relationship between feature points by introducing a length similarity matrix to ensure the accuracy of feature matching: F MHSA = MHSA(F Encoded , M ls ) Among them, F MHSA represents the feature output by the MHSA module, MHSA( ) represents the multi-head self-attention mechanism operation, M ls represents the length similarity matrix; the calculation method of the length similarity matrix M ls is as follows: Preprocess the input image pair {I A , I B} using a feature detector, and then establish an initial correspondence set I C = {c1, c2, ..., c i , ..., c j , ..., c N} through a nearest neighbor matching strategy, where c i and c j represent the i-th and j-th feature point pairs in the initial correspondence set I C respectively, N represents the number of feature point pairs in the initial correspondence set I C ; c i = (p i A , p i B ), c j = (p j A , p j B ), where p i A and p i B represent the feature points in the input image I A and the input image I B respectively. Then, the elements m ls in the length similarity matrix M i,j are calculated according to the following formula: Among them, || || represents norm calculation, and | | represents taking the absolute value; The features output by the MHSA module are divided into two paths. One path is input to the FFN network, and the other path is concatenated with the features output by the FFN network to obtain the final output features; The FFN network receives the features output by the MHSA module and performs a series of transformations to further extract features: Among them, Linear1 is the first linear transformation layer, which maps the features output by the MHSA module to a higher-dimensional space to enhance the representation ability; ReLU( ) is the ReLU activation function, which processes the features output by Linear1 to enhance the non-linear characteristics and obtain the activated features; Linear2 is the second linear transformation layer, which is used to map the activated features back to the final output space.

7. A multi-scale image matching method according to claim 1, characterized in that, Adopt a hybrid loss function to supervise the training process of the image matching model: Among them, represents the classification loss; α is a hyperparameter that balances the classification loss and the essential matrix loss; represents the essential matrix loss, which is used to measure the difference between the predicted essential matrix and the true essential matrix. After feature extraction by the image matching model, the predicted essential matrix is calculated based on the matched feature point pairs, and then the essential matrix loss is calculated; Classification loss Expressed as: Among them, H( ) represents the binary cross-entropy loss function; o t represents the relevant weight of the t-th iteration; y t represents the weakly supervised label; ω t is the adaptive temperature vector, ⊙ represents the Hadamard product, and λ is the number of iterations.

8. A multi-scale image matching system, characterized in that, It includes a memory, a processor, and computer program instructions stored on the memory and executable by the processor. When the processor runs the computer program instructions, the method described in any one of claims 1-7 can be implemented.

Citation Information

Patent Citations

  • Two-way image-text matching method based on cross-modal global and local attention mechanism

    CN116610778A

  • End-to-end image matching method based on mixed attention mechanism

    CN118351334A