A Cross-View and Cross-Modal Image Geolocation Method Combining CNN and Cross-Layer Interaction Transformer

Through the pyramid splitting attention CNN and cross-layer interactive Transformer method, the problem of lack of multi-scale feature information and feature interaction fusion in cross-view image geolocation is solved, and a higher precision cross-view geolocation is achieved.

CN116310866BActive Publication Date: 2025-07-11NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310211410.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-07
Publication Date
2025-07-11
Estimated Expiration
2043-03-07

AI Technical Summary

Technical Problem

The existing cross-view image geolocation methods lack the use of multi-scale spatial feature information, and the lack of interactive fusion in the multi-source image feature extraction process, which makes the network unable to pay attention to the close connection between multi-source images at the same time, limiting the improvement of geolocation task performance.

Method used

The method of pyramid split attention CNN and cross-layer interactive Transformer is adopted to extract multi-scale features through the pyramid split attention module, and feature fusion is combined with the cross-layer interactive Transformer module to establish a long-term dependence relationship between attention in multi-scale channels, and promote the joint positioning of global spatial layout and local detail features.

Benefits of technology

The accuracy of geolocation of cross-view images is improved, the network enhances the common feature enhancement of the same position image pair and the feature difference recognition ability of different position image pairs, and significantly improves the geolocation performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116310866B_ABST
    Figure CN116310866B_ABST
Patent Text Reader

Abstract

The present invention relates to a cross - perspective and cross - modal image geolocation method that synergizes CNN and cross - layer interaction Transformer. A high - precision cross - perspective geolocation network model is designed, which optimizes local detail features using a pyramid split attention module, captures global dependency relationships across layers using Transformer, and applies a multi - source fusion mechanism method, thereby improving the performance of the cross - perspective geolocation network. Each branch not only pays attention to the changes in its own features through the cross - layer interaction mechanism, but also can pay attention to the important features of the source image in the other branch through the multi - source fusion mechanism, promoting the flow of important information useful for positioning between the two source images, and then extracting more discriminative features to obtain better positioning accuracy. The positioning accuracy rate obtained by the present invention when only retrieving one image is 3 - 7 times that of the existing geolocation methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a geolocation method, in particular to a cross-view and cross-modal image geolocation method that collaborates CNN and cross-layer interaction Transformer. Background Art

[0002] Geolocation is a very important research field in computer vision. Image-based geolocation is an important auxiliary positioning method under the conditions of weak or interfered GPS signals and large positioning errors of base stations. By using image retrieval technology to match a query image at an unknown location with a reference image database with geographical tags, the location of the unknown location can be achieved. Cross-view geolocation based on satellite overhead view-ground street view is an important research direction of image-based geolocation. Since satellite overhead views are easy to collect and cover a wide area, and ground street views can be obtained in real time, this technology plays an important role in fields such as autonomous driving, robot navigation and tracking, and 3D reconstruction.

[0003] The problem of cross-view geolocation of multi-source images based on satellite overhead view and ground street view usually refers to the matching query of street view images obtained by ground cameras and overhead view images obtained by satellite equipment, which is generally defined as an image retrieval task. Traditional methods mainly align multiple images of the same scene through manually designed feature descriptors, first generate feature descriptors for each image, and then calculate the similarity matching between the descriptor of the query image and the descriptor of each image in the reference image database to achieve retrieval and positioning. For example, Jegou et al. used VLAD descriptors in Aggregating local descriptors into a compact image representation [C] / / 2010IEEE computer society conference on computer vision and pattern recognition. IEEE, 2010: 3304-3311. to aggregate the residuals of local features into cluster centers as image representation, but traditional methods have poor performance in cross-view image feature representation. In recent years, with the development of deep learning technology, deep neural networks have shown powerful image representation capabilities. Image geolocation methods using CNN models have continuously improved the accuracy of cross-view geolocation tasks. At the same time, the emergence of large-scale cross-view geolocation datasets has also provided conditions for model training. The CNN-based image geolocation model adopts a deep metric learning method and uses a deep neural network to map image representations to the same metric space. In the metric space, positive pairs of images with the same geographic tags are brought closer and negative pairs of images with different geographic tags are alienated, so that the deep network can find corresponding image pairs through similarity matching to achieve geolocation. For example, Hu et al. fused the NetVLAD layer in the VGG network in Cvm-net: Cross-view matching network for image-based ground-to-aerialgeo-localization[C] / / Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.2018:7258-7267. to extract the view-invariant feature descriptors of image pairs, realize cross-view image matching, and improve the accuracy of geolocation. With the introduction of the Transformer network, it has also achieved remarkable results in the field of image processing. For example, the introduction of networks such as DETR, ViT and SETR has further improved the task performance in related fields.The self-attention mechanism of the Transformer network has a powerful ability to extract global context information of images. Combining CNN and Transformer and applying them to geolocation can further improve the accuracy of location. Yang et al. proposed a new cross-attention mechanism in Cross-view geo-localization with layer-to-layer transformer [J]. Advances in Neural Information Processing Systems, 2021, 34: 29009-29020., which enables the flow of effective information between Transformer blocks, enhances the generalization ability of the network, allows the network to extract more discriminative features, and thus improves the geolocation accuracy.

[0004] However, the existing cross-view geolocation technologies still have certain limitations: ① Traditional methods focus on extracting the overall features of images, easily neglect local detail information, and it is difficult to achieve higher-precision location. There is a lack of research on location that combines global spatial features and local detail information of images. ② Existing methods mainly use a multi-branch network structure. The feature extraction processes between multi-source images lack interaction and fusion, and only single-view images are concerned, making the network unable to simultaneously focus on the close connections between multi-source images, which limits the improvement of the performance of the geolocation task. Summary of the Invention

[0005] Technical Problems to be Solved

[0006] In the existing cross-view image geolocation methods, there are problems such as the lack of use of multi-scale spatial feature information and the lack of interaction and fusion in the feature extraction process of multi-source images. To avoid the deficiencies of the existing technologies, the present invention provides a cross-view cross-modal image geolocation method that coordinates CNN and cross-layer interaction Transformer.

[0007] Technical Solution

[0008] A cross-view cross-modal image geolocation method that coordinates CNN and cross-layer interaction Transformer, characterized by the following steps:

[0009] Step 1: A multi-source image feature extraction module based on pyramid split attention CNN;

[0010] The satellite image I s and the ground panoramic image I gThey are respectively input into two branches of the feature extraction network. Each branch consists of ResNet-50 and the Pyramid Split Attention module PSA, that is, EPSANet-50. EPSANet-50 replaces the 3×3 convolutional kernel in ResNet-50 with the Pyramid Split Attention PSA module and outputs a feature map with inter-channel relationships. and Their sizes are all [H, W, C], where H is the height, W is the width, and C is the number of channels.

[0011] Step 2: Fuse the features based on the cross-layer interaction and multi-source fusion Transformer module.

[0012] Step 2-1: Flatten the feature maps and into a column of feature blocks respectively, for and then The formula for mapping the feature block X p1 into a column of sequences is The formula for mapping the feature block X p2 into a column of sequences is where, X class-p1 is the classification embedding marker added to the feature block X p1 , X class-p2 is the classification embedding marker added to the feature block X p2 . denotes the 1st, 2nd, …, N p1 th feature blocks in X p1 , denotes the 1st, 2nd, …, N p2 th feature blocks in X p2 , E p1 and E p2 respectively represent the projection parameters used for the feature blocks X p1 and X p2 , with a size of [1, 1, C]; E pos1 represents the positional encoding PE feature embedding of X p1 , E pos2 represents the positional encoding PE feature embedding of X p2 ; The positional encoding PE is implemented using sine and cosine functions with different frequencies, and the formula for the positional encoding is as follows:

[0013]

[0014] Among them, pos represents the position of each feature block, and the range of pos is [1, N], i represents the i-th feature number, and the range of i is [1, C]; that is, each dimension of the position encoding corresponds to a sine curve; the wavelength forms a geometric series from 2π to 10000·2π; this step is based on Z p1 and Z p2 is the output;

[0015] Step 2-2: Z p1 and Z p2 In the input cross-layer interaction module, the interaction of adjacent layer feature blocks of Transformer is used to learn the global context information of the image; Transformer has 12 layers in total, and the cross-layer interaction module is applied in the first 8 layers; the attention map of the lth layer is not only based on the learning of the l-1th layer feature blocks, but also based on the learning of the l-2th layer feature blocks; after matrix mapping and attention calculation, the attention of the lth layer can be obtained as Att l , and then we get Z p1-cl and Z p2-cl As output; where l ranges from [1, 8]; in particular, when l = 1, the first layer attention map is based on Z p1 and Z p2 Learning; when l = 2, the second layer attention map is based on Z p1 , Z p2 and the first layer feature block learning;

[0016] Step 2-3: Z p1-cl and Z p2-cl The input is fed into the multi-source fusion module, and the multi-source fusion module is applied to the last 4 layers of Transformer; that is, the attention map of the first layer is not only based on the learning of a branch feature block of the first-1 layer, but also based on the learning of another branch feature block of the first-1 layer, and the value range of l is [9, 12]; in this way, each branch not only pays attention to the changes of its own features through the cross-layer interaction mechanism, but also pays attention to the important features of the source image in another branch through the multi-source fusion mechanism, which promotes the deep interaction of information useful for positioning between ground panoramic images and satellite images in the deep layer of the network, and obtains the final global descriptor. and

[0017] Step 2-4: Calculate feature similarity

[0018] Use Euclidean distance to calculate the similarity between features, using and Construct a weighted soft margin triplet objective function to shorten the distance between matching image pairs and make the distance between unmatched image pairs as far as possible;

[0019]

[0020] Among them, d p and d n respectively represent the Euclidean distances between the anchor point and the positive and negative samples. α is a hyperparameter that accelerates the network convergence during the training phase;

[0021] Step 3: Train the constructed network

[0022] Put the data in the training set into the network batch by batch to generate the top K satellite images that are most similar to each ground panoramic image. Calculate the loss using the predicted labels and the truly matched labels. Specifically, use a weighted soft margin triplet loss function and optimize it using the Adam optimizer until the value of the objective function no longer decreases to end the training;

[0023] Step 4: Test the image set

[0024] Input the test images into the image matching network trained in Step 3. Calculate the similarity score between the ground panoramic image and the satellite image by using the Euclidean distance to obtain the query results of the top K satellite images that are the most similar, and evaluate them using the recall metric Recall@K;

[0025] Step 5: Locate the ground panoramic image

[0026] Complete the positioning task of the ground panoramic image by querying the GPS longitude and latitude position information corresponding to the top K satellite images that are the most similar.

[0027] A further technical solution of the present invention: The pyramid split attention PSA module in Step 1 includes four steps: First, use convolutional kernels with receptive fields of 3×3, 5×5, 7×7, and 9×9 to divide the input feature map into 4 groups from the channels to obtain feature maps with different scales; Second, use channel attention SE to extract the weighted values of each group of channels and extract the attention of the feature maps with different scales; Then, after re-calibration by Softmax, obtain the re-calibration weights of the multi-scale channels; Finally, perform element-wise multiplication on the obtained channel attention and the corresponding feature maps to obtain the final feature map including the relationship between channels.

[0028] A computer system, characterized in that it includes: one or more processors, a computer-readable storage medium for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the above method.

[0029] A computer-readable storage medium, characterized in that it stores computer-executable instructions, and the instructions are used to implement the above method when executed.

[0030] Beneficial effects

[0031] A cross - perspective and cross - modal image geolocation method that combines CNN and cross - layer interaction Transformer is provided by the present invention. Compared with the prior art, the beneficial effects of the present invention are as follows:

[0032] 1) The present invention proposes a cross - perspective and cross - modal geolocation network based on pyramid - split attention CNN and cross - layer interaction collaborative Transformer for cross - perspective geolocation of satellite - ground images;

[0033] 2) The present invention can effectively establish long - term dependencies between multi - scale channel attentions while processing multi - scale feature space information, and perform cross - perspective geolocation by combining global spatial layout and local detail features;

[0034] 3) The present invention applies a cross - layer interaction multi - source fusion strategy, which enables the network to strengthen the common important features learned by each branch for cross - perspective image pairs at the same location, and for cross - perspective image pairs at different locations, it also prompts each branch to increase this feature difference, thereby improving the network performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The drawings are only for the purpose of showing specific embodiments and are not considered as limitations of the present invention. Throughout the drawings, the same reference signs represent the same components.

[0036] Figure 1 is the network structure diagram of the embodiment of the present invention.

[0037] Figure 2 is the EPSANet - 50 structure diagram and the residual unit structure diagram of the present invention.

[0038] Figure 3 is the structure diagram of the pyramid - split attention module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0040] The present invention provides a cross - perspective and cross - modal image geolocation method that combines CNN and cross - layer interaction Transformer, and designs a cross - perspective and cross - modal image geolocation model that combines CNN and cross - layer interaction Transformer, as Figure 1As shown, the model mainly includes a feature extraction module based on a pyramid split attention CNN and a cross-layer interaction and multi-source fusion Transformer module. The pyramid split attention module can effectively establish long-term dependencies between multi-scale channel attentions while processing multi-scale feature space information; the cross-layer interaction strategy models the global dependencies between adjacent layers, combines the relationship between the feature blocks of the front and back layers while the features in the previous layer evolve, enabling the network to extract more discriminative features; applying the multi-source fusion strategy to multi-source data enables the single-source image to not only focus on its own features but also on the features useful for localization in another source image, thereby improving the cross-view image geolocation performance.

[0041] The specific method includes the following steps:

[0042] Step 1: A multi-source image feature extraction module based on a pyramid split attention CNN.

[0043] Input the ground panoramic image Ig and the satellite image Is into different branch networks respectively. The images input into each branch network pass through the EPSANet-50 network. The EPSANet-50 network replaces the 3×3 convolution in the ResNet50 network with a pyramid split attention (PSA) module. Obtain the feature maps and as outputs.

[0044] The EPSANet-50 network (as shown in Figure 2 the left) consists of 50 convolutional layers and fully connected layers. It sequentially includes: a convolutional layer with a 7×7 convolutional kernel and a stride of 2, a max pooling layer with a 3×3 pooling kernel and a stride of 2, Residual Block 1 (consisting of 3 residual units, and each residual unit (as shown in Figure 2 the right) sequentially includes a 1×1 convolutional layer, a PSA module, a 1×1 convolutional layer, and an identity mapping skip connection), Residual Block 2 (consisting of 4 residual units), Residual Block 3 (consisting of 6 residual units), Residual Block 4 (consisting of 3 residual units), a global average pooling layer (7×7 pooling kernel), a fully connected layer, and a Softmax layer.

[0045] The pyramid split attention module (as shown in Figure 3As shown in the figure, it mainly includes four steps: First, use the convolution kernels with receptive fields of 3×3, 5×5, 7×7, and 9×9 to divide the input feature map into 4 groups from the channels to obtain feature maps with different scales; Second, use channel attention (SE) to extract the weighted values of each group of channels and extract the attention of feature maps with different scales; Then, through Softmax recalibration, obtain the recalibrated weights of multi-scale channels; Finally, perform element-wise multiplication on the obtained channel attention and the corresponding feature map to obtain the final feature map containing the relationships between channels.

[0046] Step 2: Cross-layer interaction and multi-source fusion Transformer module;

[0047] Step 2-1: Flatten the feature maps output by the feature extraction module and into a column of feature blocks, which are and (each feature block has a size of [1,1,C], N p1 represents the number of blocks of the feature block obtained by flattening the feature map , N p2 represents the number of blocks of the feature block obtained by flattening the feature map , and C represents the number of channels of each feature block), then The formula for mapping the feature block X p1 to a column of sequences is The formula for mapping the feature block X p2 to a column of sequences is where X class-p1 is the classification embedding mark added to the feature block X p1 , X class-p2 is the classification embedding mark added to the feature block X p2 , represents the 1st, 2nd, …, N p1 th feature blocks in X p1 , represents the 1st, 2nd, …, N p2 th feature blocks in X p2 , E p1 and E p2 represent the projection parameters used for the feature blocks X p1 and X p2 respectively, with a size of [1,1,C]. E pos1 represents the positional encoding (PE) feature embedding of X p1 , and E pos2 represents the positional encoding (PE) feature embedding of X p2 . The positional encoding (PE) is implemented using sine and cosine functions with different frequencies, and the formula for the positional encoding is as follows:

[0048]

[0049] Among them, pos represents the position of each feature block, and the range of pos is [1, N]. i represents the i-th feature number, and the range of i is [1, C]. That is, each dimension of the position encoding corresponds to a sine curve. The wavelengths form a geometric progression from 2π to 10,000·2π. This step takes Z p1 and Z p2 as the output.

[0050] Step 2-2: Input Z p1 and Z p2 into the Transformer encoder. The encoder module is stacked by 12 encoder layers, and the cross-layer interaction module is applied to the first 8 layers. Specifically, the attention map of the l-th layer is learned not only based on the feature blocks of the (l - 1)-th layer but also based on the feature blocks of the (l - 2)-th layer. After matrix mapping and attention calculation, the attention of the l-th layer can be obtained as Att l , and then Z p1-cl and Z p2-cl are obtained as the output. Among them, the value range of l is [1, 8]. In particular, when l = 1, the attention map of the first layer is learned based on Z p1 and Z p2 . When l = 2, the attention map of the second layer is learned based on Z p1 , Z p2 and the feature blocks of the first layer. In the l-th layer, the cross-layer interaction module can be expressed as:

[0051] Q l = LN(Z l-1 )W l q , K l = LN(Z l-2 )W l k , V l = LN(Z l-1 )W l v

[0052]

[0053] Among them, LN represents layer normalization, Z l-1 represents the sequence mapped from the feature blocks of the (l - 1)-th layer, Z l-2 represents the sequence mapped from the feature blocks of the l-th layer, W l q , W l k , W l v represent random weight matrices, and D represents the dimension.

[0054] Step 2-3: Z p1-cl and Z p2-cl Input into the multi-source fusion module. The multi-source fusion module is applied to the last 4 layers of the Transformer encoder. The formula of the multi-source fusion module is as follows:

[0055]

[0056]

[0057] Among them, the value range of l is [9, 12], LN represents layer normalization, and represents the sequence of the l-1th layer, W l q , W l k , W l v represents the random weight matrix, and D represents the dimension. Then we get the final global descriptor and

[0058] Step 2-4: Calculate feature similarity.

[0059] Use Euclidean distance to calculate the similarity between features, using and A weighted soft margin triplet objective function is constructed to shorten the distance between matching image pairs and make the distance between unmatched image pairs as far as possible.

[0060]

[0061] Here, d p and d n They represent the Euclidean distance from the anchor point to the positive sample and the negative sample respectively. α is a hyperparameter, which is generally set to 10 to accelerate network convergence during the training phase.

[0062] Step 3: Train the constructed network. Put the ground panoramic images and satellite images in the training set into the network in batches, generate the top K satellite images most similar to each ground panoramic image, and use the predicted labels and the real matching labels to calculate the loss. Specifically, a weighted soft margin triplet loss function is used, and the Adam optimizer is used for optimization until the value of the objective function does not decrease, and the training ends.

[0063] Step 4: Test images. The ground panoramic images and satellite images of the test set are input into the image matching network trained in step 3. The similarity scores between the ground panoramic images and satellite images are calculated using the Euclidean distance to obtain the top K satellite image query results that are most similar, and the recall rate indicator Recall@K is used for evaluation.

[0064] Step 5, locate the ground panoramic image. By querying the GPS longitude and latitude position information corresponding to the top K satellite images that are closest to the ground panoramic image, the correct positioning of the ground panoramic image is completed. On the CVUSA and CVACT test sets, the Recall@1 accuracy rates are 94.35% and 85.40% respectively. Compared with the results of 31.99% and 11.50% in the reference document "CVMNet", the accuracy is improved by 3 to 7 times.

[0065] Table 1 Comparison chart of the test results of the method of the present invention and other existing methods in the CVUSA dataset in the embodiments of the present invention.

[0066]

[0067] As described above, the above are only the specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention.

Claims

1. A cross - perspective and cross - modal image geolocation method that combines CNN and cross - layer interaction Transformer, characterized in that The steps are as follows: Step 1: A multi-source image feature extraction module based on a pyramid split attention CNN; Input the satellite image I s and the ground panoramic image I g into two branches of the feature extraction network respectively. Each branch consists of ResNet-50 and the Pyramid Split Attention module PSA, that is, EPSANet-50; EPSANet-50 replaces the 3×3 convolution kernel in ResNet-50 with the Pyramid Split Attention PSA module and outputs feature maps with inter-channel relationships and their sizes are all [H, W, C], where H is the height, W is the width, and C is the number of channels; Step 2: Fuse features based on a cross-layer interaction and multi-source fusion Transformer module; Step 2-1: Flatten the feature maps and into a column of feature blocks respectively, for and then The formula for mapping the feature block X p1 into a column of sequences is The formula for mapping the feature block X p2 to a sequence of columns is where X class-p1 is the classification embedding marker added to the feature block X p1 is the classification embedding marker added to the feature block X class-p2 is the classification embedding marker added to the feature block X p2 is the classification embedding marker added to the feature block X represents the 1st, 2nd, …, Nth p1 feature blocks in X p1 feature blocks; represents the 1st, 2nd, …, Nth p2 feature blocks in X p2 feature blocks, E p1 and E p2 respectively represent the projection parameters used for the feature blocks X p1 and X p2 with a size of [1, 1, C]; E pos1 represents the position encoding PE feature embedding of X p1 and E pos2 represents the position encoding PE feature embedding of X p2 The position encoding PE is implemented using sine and cosine functions with different frequencies. The formula for the position encoding is as follows: where pos represents the position of each feature block, the range of pos is [1, N], i represents the i-th feature number, and the range of i is [1, C]; that is, each dimension of the positional encoding corresponds to a sine curve; the wavelengths form a geometric progression from 2π to 10000·2π; this step takes Z p1 and Z p2 as the output; Step 2-2: Input Z p1 and Z p2 into the cross-layer interaction module, and utilize the interaction between adjacent layer feature blocks of the Transformer to learn the global context information of the image; The Transformer has 12 layers, and the cross-layer interaction module is applied to the first 8 layers. The attention map of the l-th layer is learned not only based on the feature blocks of the (l-1)-th layer but also based on the feature blocks of the (l-2)-th layer; the attention of the l-th layer, Att, can be obtained through matrix mapping and attention calculation l , and then Z p1-cl and Z p2-cl are obtained as outputs; where the value range of l is [1, 8]; in particular, when l = 1, the attention map of the first layer is learned based on Z p1 and Z p2 ; when l = 2, the attention map of the second layer is learned based on Z p1 , Z p2 and the feature blocks of the first layer; Step 2-3: Input Z p1-cl and Z p2-cl into the multi-source fusion module, and apply the multi-source fusion module to the last 4 layers of the Transformer; that is, the attention map of the l-th layer is learned not only based on one branch feature block of the (l-1)-th layer, but also based on the other branch feature block of the (l-1)-th layer, where the value range of l is [9, 12]; in this way, each branch can not only focus on the changes of its own features through the cross-layer interaction mechanism, but also can focus on the important features of the source image in the other branch through the multi-source fusion mechanism, promoting the deep interaction of the information useful for localization between the ground panoramic image and the satellite image in the deep layer of the network to obtain the final global descriptor and Step 2-4: Calculate feature similarity Use Euclidean distance to calculate the similarity between features, using and Construct a weighted soft margin triplet objective function to shorten the distance between matching image pairs and make the distance between unmatched image pairs as far as possible; where d p and d n represent the Euclidean distances between the anchor point and the positive and negative samples respectively, and α is a hyperparameter that accelerates network convergence during the training phase; Step 3: Train the constructed network Put the data in the training set into the network in batches to generate the top K satellite images that are most similar to each ground panoramic image. Calculate the loss using the predicted labels and the truly matched labels. Specifically, use a weighted soft margin triplet loss function and optimize it using the Adam optimizer until the value of the objective function no longer decreases, then end the training; Step 4: Test image set Input the test images into the image matching network trained in Step 3. Calculate the similarity score between the ground panoramic image and the satellite image by using the Euclidean distance to obtain the query result of the top K satellite images that are the most similar, and evaluate it using the recall metric Recall@K; Step 5: Locate the ground panoramic image Complete the positioning task of the ground panoramic image by querying the GPS longitude and latitude position information corresponding to the top K satellite images that are the most similar.

2. The cross - perspective and cross - modal image geolocation method combining CNN and cross - layer interaction Transformer according to claim 1, characterized in that: In Step 1, the pyramid split attention PSA module includes four steps: First, use convolutional kernels with receptive fields of 3×3, 5×5, 7×7, and 9×9 to divide the input feature map into 4 groups from the channels to obtain feature maps with different scales; Secondly, use channel attention SE to extract the weighted values of each group of channels to extract the attention of feature maps with different scales; then, re-calibrate through Softmax to obtain the re-calibrated weights of multi-scale channels; finally, perform element-wise multiplication on the obtained channel attention and the corresponding feature map to obtain the final feature map containing the relationship between channels.

3. A computer system, characterized in that Including: One or more processors, a computer-readable storage medium for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in claim 1.

4. A computer-readable storage medium, characterized in that Store computer-executable instructions that, when executed, are used to implement the method described in claim 1.

Citation Information

Patent Citations

  • Aerial image geographic positioning method based on spatial scale attention mechanism and vector map

    CN113239952A

  • Fine-grained costume retrieval method based on CNN-Transform double-flow network

    CN115410067A