Cross-visual-angle geographic positioning method based on attention mechanism, medium and equipment

By introducing attention mechanism and multi-dimensional feature fusion into the neural network model, the problem of insufficient attention to the target building in cross-view matching is solved, and higher matching accuracy and positioning accuracy are achieved.

CN120047795AActive Publication Date: 2025-05-27XIAMEN UNIV

Patent Information

Application Number
CN202510209652.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-27
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

Existing cross-view angle matching methods cannot accurately focus on target buildings in images of different perspective angles, resulting in a decrease in matching accuracy and affecting the accuracy of building positioning.

Method used

Using a neural network model based on attention mechanism, through multi-dimensional feature fusion, deep convolution optimization and cross-view feature alignment, the robustness of attention to target buildings or important areas during transformations such as image offset and rotation in geolocation tasks is significantly improved.

Benefits of technology

It improves the accuracy and robustness of cross-view image matching, enhances the adaptability to image translation, rotation and other transformations, and significantly improves the accuracy of building positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047795A_ABST
    Figure CN120047795A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-view geographic positioning method based on an attention mechanism, a medium and equipment. The method comprises the following steps: receiving image information shot by an unmanned aerial vehicle or a satellite; and inputting the image information into the trained neural network model, and outputting geographical location information corresponding to the image information. When the neural network model is trained, through modes of multi-dimensional feature fusion, deep convolution optimization, cross-view feature alignment and the like, robustness of attention to a target building or an important area during transformation of image offset, rotation and the like in a geographic positioning task is remarkably improved; therefore, the difficulty in positioning in complex scenes such as view angle change and geometric shape difference is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly to a cross-view geolocation method, medium and device based on an attention mechanism. Background Art

[0002] In the fields of computer vision and pattern recognition, cross-view image matching is a task with important research significance, which plays a key role in many practical applications such as drone image analysis, satellite remote sensing image processing, and scene reconstruction. However, due to the significant geometric deformations and illumination differences between cross-view images, how to extract robust and discriminative features from them has become a challenge.

[0003] Currently, the attention mechanism has become an important tool for improving model performance. By simulating the selective attention mechanism of the human visual system, it can enhance the expressive ability of significant regions in the feature map. However, existing attention mechanisms such as SENet, CBAM, and SKNet have certain limitations when dealing with cross-view tasks. For example, SENet only focuses on channel dimension information and ignores the importance of spatial attention; although CBAM combines channel and spatial attention, it processes them separately and fails to achieve the fusion of cross-dimensional features; although SKNet captures some directional features using multi-scale convolutions, it fails to fully integrate channel and spatial information. Summary of the Invention

[0004] In view of the above problems, the present invention provides a technical solution for cross-view geolocation based on an attention mechanism to solve the problem that existing cross-view matching methods cannot accurately focus on target buildings in images from different perspectives, resulting in a decrease in the matching accuracy of target buildings and further affecting the accuracy of building location.

[0005] To achieve the above object, in a first aspect, the present application provides a cross-view geolocation method based on an attention mechanism, including the following steps:

[0006] Receiving image information captured by a drone or a satellite;

[0007] Inputting the image information into a trained neural network model and outputting the geographical location information corresponding to the image information;

[0008] When the neural network model is being trained, it includes the following steps:

[0009] S1: Receiving a first sample image collected from a satellite perspective or a second sample image collected from a drone perspective, and extracting a feature map from the first sample image or the second sample image using a pre-trained ConvNeXt network;

[0010] S2: Divide the extracted feature map into a first-branch tensor, a second-branch tensor, and a third-branch tensor. After performing ECA operations on the first-branch tensor, the second-branch tensor, and the third-branch tensor respectively, perform global average pooling to obtain a first-channel description vector, a second-channel description vector, and a third-channel description vector;

[0011] S3: Use the first-channel description vector to perform weighted calculation on the first-branch tensor, use the second-channel description vector to perform weighted calculation on the second-branch tensor, and use the third-channel description vector to perform weighted calculation on the third-branch tensor to obtain first attention feature maps corresponding to the three branches. Perform average calculation on the three first attention feature maps and output a second attention feature map;

[0012] S4: Perform depth convolution operation on the second attention feature map to obtain an enhanced third attention feature map;

[0013] S5: Determine the geographical location information based on the third attention feature map.

[0014] Optionally, the feature map extracted from the first sample image or the second sample image by using a pre-trained ConvNeXt network is represented by the following formula:

[0015]

[0016] where \(i\in\{1,2\}\), \(x\) 1 represents the first sample image, \(x\) 2 represents the second sample image, \(L\) i \(\in\)

[0017] \(R\) C×H×W represents the extracted feature map, \(C\) is the number of channels, and \(H\) and \(W\) respectively represent the height and width of the image.

[0018] Furthermore, the shape of the feature map is \(C\times H\times W\), \(C\) is the number of channels, and \(H\) and \(W\) respectively represent the height and width of the image;

[0019] The first-branch tensor is represented as \(W\times H\times C\), the second-branch tensor is represented as \(H\times C\times W\), and the third-branch tensor retains the original shape of the feature map and is represented as \(C\times H\times W\);

[0020] Steps S2 and S3 are represented by the following formula:

[0021]

[0022] where \(L\) D represents the input first-branch tensor, second-branch tensor, and third-branch tensor, \(X\) DThe first attention feature map is denoted as output, X represents the second attention feature map, and the symbol σ represents the Sigmoid activation function.

[0023] Further, step S4 includes:

[0024] Performing depth convolution operations using convolutional kernels of different lengths, and realizing information interaction between channels through a 1×1 convolutional layer to obtain an enhanced third attention feature map;

[0025] Step S4 is represented by the following formula:

[0026]

[0027] Among them, DwConv represents 5×5 depth convolution, Pathi is the result of two depth strip convolutions and the body after 5×5 depth convolution, and Y is the enhanced third attention feature map.

[0028] Further, when the neural network model is being trained, the following steps are also included:

[0029] When calculating the KL divergence, ignore the prediction probability of the positive class, and reapply the softmax function to calculate the prediction probability of the negative class, specifically including:

[0030] Pass the third attention feature map through an average pooling operation and then transfer it to the classifier. Obtain the logit scores for each class through the classifier, and transfer the logit scores for each class to the Log-Softmax layer to obtain the logarithm of the prediction probability for each class, specifically represented by the following formula:

[0031] L = Log_Softmax(Classifier(AvgPool(Y)));

[0032] Among them, L = {l 1 ,l 2 ,...,l p ,...,l N} ∈ R 1×N represents the logarithm of the prediction probability for each class, N represents the total number of classes, and the subscript p represents the positive class.

[0033] To calculate the negative class KL divergence loss, ignore lp, so L is rewritten as:

[0034]

[0035] Then, L^ obtains the predicted probability distribution through the Softmax layer, expressed as follows:

[0036]

[0037] Among them, P represents the predicted probability distribution of all negative classes.

[0038] Furthermore, the method further includes:

[0039] Mutually learn the predicted probability distributions of the first sample image collected from the satellite perspective and the second sample image collected from the drone perspective, so as to align the depth feature distributions of the two.

[0040] The calculation process of the negative class KL divergence is expressed by the following formula:

[0041]

[0042] Among them, P1 and P2 respectively represent the negative class predicted probability distributions of the first sample image and the second sample image.

[0043] Furthermore, the method includes:

[0044] Comprehensively calculate according to the KL divergence loss, contrast loss and cross-entropy loss to obtain the final loss function of the neural network model.

[0045] Furthermore, the contrast loss is expressed as follows:

[0046]

[0047] Among them, q is the feature vector of the query image, R is the set of reference images, r + is the positive sample corresponding to the query image, and τ is the temperature parameter used to control the scaling of the similarity score;

[0048] The cross-entropy loss is expressed as follows:

[0049]

[0050] Among them, p^(y∣x) is the predicted probability of the neural network model, and lp(y) is the logarithm of the predicted probability of the positive class;

[0051] The calculation formula of the final loss function of the neural network model is as follows:

[0052] Loss = L infoNCE + λ 1 L CE + λ 2 L NCKL ;

[0053] Among them, λ1 and λ2 are adjustable hyperparameters.

[0054] In a second aspect, the present invention further provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described in the first aspect is implemented.

[0055] In a third aspect, the present invention further provides an electronic device, including a memory and a processor, where the memory is used to store one or more computer program instructions, and wherein the one or more computer program instructions are executed by the processor to implement the method described in the first aspect.

[0056] Different from the prior art, the above solution provides a cross-perspective geolocation method, medium and device based on an attention mechanism. The method includes: receiving image information captured by a drone or a satellite; inputting the image information into a trained neural network model, and outputting the geographical location information corresponding to the image information. When training the neural network model, the present application significantly improves the robustness of paying attention to target buildings or important areas during image offset and rotation and other transformations in the geolocation task through multi-dimensional feature fusion, deep convolution optimization, cross-perspective feature alignment, etc., thereby overcoming the positioning difficulties in complex scenarios such as perspective changes and geometric shape differences.

[0057] The above relevant descriptions of the invention content are only an overview of the technical solution of the present invention. In order to enable those of ordinary skill in the art to more clearly understand the technical solution of the present invention, and then can be implemented according to the content recorded in the description and the drawings, and in order to make the above objects, other objects, features and advantages of the present invention more easily understood, the following is described in conjunction with the specific embodiments and drawings of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] The drawings are only used to illustrate the principles, implementation methods, applications, features and effects of the specific embodiments of the present invention and other related contents, and should not be considered as a limitation to the present invention.

[0059] In the accompanying drawings of the specification:

[0060] Figure 1 is a flowchart of the cross-perspective geolocation method based on an attention mechanism according to the first exemplary embodiment of the present invention;

[0061] Figure 2 is a flowchart of the training process of the neural network model according to the second exemplary embodiment of the present invention;

[0062] Figure 3 is a framework diagram of the cross-perspective geolocation method based on an attention mechanism according to a specific embodiment of the present invention;

[0063] Figure 4Schematic diagram of the principle of the attention mechanism involved in a specific embodiment of the present invention;

[0064] Figure 5 Schematic diagram of the visualization model heat map generated by the cross-view matching method related to the prior art;

[0065] Figure 6 Schematic diagram of the visualization model heat map generated by the cross-view geolocation method based on the attention mechanism involved in a specific embodiment of the present invention;

[0066] Figure 7 Comparison chart of the output results of the method involved in a specific embodiment of the present invention and the method related to the prior art on the dataset University-1652;

[0067] Figure 8 Comparison chart of the output results of the method involved in a specific embodiment of the present invention and the method related to the prior art on the dataset SUES-200;

[0068] Figure 9 Comparison chart of the processing results of the attention mechanism involved in a specific embodiment of the present invention and the attention mechanism related to the prior art;

[0069] Figure 10 Comparison chart of the processing results of the loss function involved in a specific embodiment of the present invention and the loss function related to the prior art;

[0070] Figure 11 Schematic diagram of the electronic device involved in a specific embodiment of the present invention.

[0071] The description of the reference numerals involved in the above-mentioned drawings is as follows:

[0072] 10. Electronic device;

[0073] 101. Processor;

[0074] 102. Storage medium. Specific Embodiment

[0075] To describe in detail the possible application scenarios, technical principles, implementable specific solutions, achievable purposes and effects of the present invention, etc., the following will be described in detail with reference to the specific examples listed and in conjunction with the drawings. The embodiments described herein are only used to more clearly illustrate the technical solutions of the present invention, so they are only examples and cannot be used to limit the protection scope of the present invention.

[0076] References to "embodiments" in this specification mean that the particular features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present invention. The term "embodiment" as used in various places in the specification does not necessarily refer to the same embodiment, nor does it particularly limit its independence or relevance to other embodiments. In principle, in the present invention, as long as there is no technical contradiction or conflict, the technical features mentioned in each embodiment can be combined in any way to form corresponding implementable technical solutions.

[0077] Unless otherwise defined, the meanings of the technical terms used herein are the same as those commonly understood by those skilled in the technical field to which the present invention belongs; the use of the relevant terms herein is only for describing specific embodiments and is not intended to limit the present invention.

[0078] In the description of the present invention, the phrase "and / or" is an expression used to describe the logical relationship between objects, indicating that there can be three relationships, for example, A and / or B, which means: the existence of A, the existence of B, and the simultaneous existence of A and B. In addition, the character " / " in this article generally represents an "or" logical relationship between the associated objects before and after.

[0079] In the present invention, terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual quantitative, primary-secondary, or sequential relationships between these entities or operations.

[0080] Without further limitation, in the present invention, the open-ended expressions such as "including", "comprising", "having" or other similar expressions used in the statements are intended to cover non-exclusive inclusion. These expressions do not exclude the possibility that there may be additional elements in the process, method, or product including the said elements, so that the process, method, or product including a series of elements may include not only those defined elements, but also other elements not explicitly listed, or elements inherent to such process, method, or product.

[0081] In the present invention, expressions such as "greater than", "less than", "exceeding" are understood not to include the number itself; expressions such as "above", "below", "within" are understood to include the number itself. In addition, in the description of the embodiments of the present invention, the meaning of "multiple" is two or more (including two), and similar expressions related to "many" are also understood in this way, such as "multiple groups", "multiple times", etc., unless otherwise specifically defined.

[0082] In a first aspect, as Figure 1 shown, the present application provides a cross-view geolocation method based on an attention mechanism, including the following steps:

[0083] First, enter step S101: Receive the image information captured by the drone or satellite.

[0084] Then, enter step S102: Input the image information into the trained neural network model and output the geographical location information corresponding to the image information.

[0085] In this embodiment, the image information is the image information containing the target building. By accurately matching the image information, the specific location of the target building can be accurately marked.

[0086] Such as Figure 2 As shown, when the neural network model is being trained, it includes the following steps:

[0087] S1: Receive the first sample image collected from the satellite perspective or the second sample image collected from the drone perspective, and use the pre-trained ConvNeXt network to extract the feature map from the first sample image or the second sample image.

[0088] S2: Divide the extracted feature map into the first branch tensor, the second branch tensor, and the third branch tensor, perform ECA operations on the first branch tensor, the second branch tensor, and the third branch tensor respectively, and then perform global average pooling to obtain the first channel description vector, the second channel description vector, and the third channel description vector.

[0089] S3: Use the first channel description vector to perform weighted calculation on the first branch tensor, use the second channel description vector to perform weighted calculation on the second branch tensor, use the third channel description vector to perform weighted calculation on the third branch tensor, obtain the first attention feature maps corresponding to the three branches, perform average calculation on the three first attention feature maps, and output the second attention feature map.

[0090] S4: Perform depth convolution operation on the second attention feature map to obtain the enhanced third attention feature map.

[0091] S5: Determine the geographical location information based on the third attention feature map.

[0092] In some embodiments, the present invention uses the ConvNeXt network pre-trained on ImageNet as the feature extraction module to extract features from the images taken from the satellite perspective and the drone perspective respectively. Considering that the images from the satellite perspective and the drone perspective have a similar pattern structure, the feature extraction part is implemented by sharing weights. Specifically, the neural network model framework of the present application is as Figure 3 As shown, the extraction of the feature map from the first sample image or the second sample image by using the pre-trained ConvNeXt network is represented by the following formula:

[0093]

[0094] where \(i\in\{1,2\}\), \(x\) 1 represents the first sample image, \(x\) 2 represents the second sample image, \(L\) i \(\in\)

[0095] \(R\) C×H×W represents the extracted feature map, \(C\) is the number of channels, and \(H\) and \(W\) represent the height and width of the image respectively.

[0096] To overcome the challenges brought by the geometric shape changes between images, this application proposes an innovative Shape Self - Adaptation and Significance - Aware Attention module (SSA), which includes two key sub - modules: Cross - Dimensional Feature Fusion (CFF) module and Direction - Guided Depth Convolution (DDC) module. They enhance the fusion and alignment ability of cross - perspective features through operations in different dimensions. The explanatory diagram of the Shape Self - Adaptation and Significance - Aware Attention module is as Figure 4 shown.

[0097] The main role of the CFF module is to perform cross - dimensional operations on the input feature map to achieve effective feature fusion. In some embodiments, the shape of the feature map is \(C\times H\times W\), where \(C\) is the number of channels, and \(H\) and \(W\) represent the height and width of the image respectively;

[0098] The first - branch tensor is represented as \(W\times H\times C\), the second - branch tensor is represented as \(H\times C\times W\), and the third - branch tensor retains the original shape of the feature map, represented as \(C\times H\times W\);

[0099] ECA (Efficient Channel Attention) operations are respectively performed on each part of the tensors, and then global average pooling is carried out to obtain three channel description vectors, including the first channel description vector, the second channel description vector, and the third channel description vector. Then, the first channel description vector is used to perform weighted calculation on the first - branch tensor, the second channel description vector is used to perform weighted calculation on the second - branch tensor, and the third channel description vector is used to perform weighted calculation on the third - branch tensor to obtain the channel attention map and the spatial attention map (i.e., the first attention feature map). Finally, the three feature maps are simply averaged to obtain the output \(X\) of the CFF module. The processing process of the CFF module is the steps described in the previous steps S2 and S3, and steps S2 and S3 are represented by the following formulas:

[0100]

[0101] where \(L\) D represents the input first - branch tensor, second - branch tensor, and third - branch tensor, \(X\)D The first attention feature map of the output is denoted as, X represents the second attention feature map, and the symbol σ represents the Sigmoid activation function.

[0102] The DDC module uses a 5×5 depth convolution operation to capture spatial feature relationships. This operation can reduce the computational complexity while maintaining the relationships between channels. To enhance the ability of the convolution operation to extract spatial features, the DDC module also uses convolution kernels of different lengths and realizes information interaction between channels through a 1×1 convolution layer. Therefore, in some embodiments, step S4 includes:

[0103] Performing depth convolution operations using convolution kernels of different lengths and realizing information interaction between channels through a 1×1 convolution layer to obtain an enhanced third attention feature map;

[0104] Step S4 is represented by the following formula:

[0105]

[0106] where DwConv represents a 5×5 depth convolution, Pathi is the result of two depth strip convolutions and the body after the 5×5 depth convolution, and Y is the enhanced third attention feature map.

[0107] In this way, through the collaborative effect of the SSA module, the model can effectively extract and fuse deep features across dimensions from images from different perspectives, thereby achieving more accurate matching.

[0108] The KL divergence measures the difference between two probability distributions. In the traditional KL divergence loss, the model tends to focus more on the positive class, thus ignoring the importance of the negative class. To enhance the model's attention to the negative class, the present invention proposes a negative class KL divergence loss (NCKL). In some embodiments, when the neural network model is being trained, the following steps are further included: when calculating the KL divergence, the predicted probability of the positive class is ignored, and the softmax function is reapplied to calculate the predicted probability of the negative class.

[0109] Specifically, it includes:

[0110] The third attention feature map (i.e., the output feature map of the SSA module) is passed to the classifier after an average pooling operation, and the logit score of each class is obtained through the classifier. The logit scores of each class are passed to the Log-Softmax layer to obtain the logarithm of the predicted probability of each class, which is specifically represented by the following formula:

[0111] L = Log_Softmax(Classifier(AvgPool(Y)));

[0112] where L = {l1 ,l 2 ,...,l p ,...,l N}, ∈ R 1×N represents the logarithm of the predicted probability for each class, N represents the total number of classes, and the subscript p represents the positive class (i.e., the true class).

[0113] To calculate the negative class KL divergence loss, lp is ignored, so L is rewritten as:

[0114]

[0115] Next, L^ obtains the predicted probability distribution through the Softmax layer, which is expressed as follows:

[0116]

[0117] where P represents the predicted probability distribution for all negative classes.

[0118] In some embodiments, the method further includes:

[0119] Mutually learning the predicted probability distributions of the output results of the first sample image collected from the satellite perspective and the second sample image collected from the drone perspective, so as to align the depth feature distributions of the two;

[0120] The calculation process of the negative class KL divergence is expressed by the following formula:

[0121]

[0122] where P1 and P2 respectively represent the negative class predicted probability distributions of the first sample image and the second sample image.

[0123] In some embodiments, the method includes:

[0124] Comprehensively calculating according to the KL divergence loss, contrast loss, and cross-entropy loss to obtain the final loss function of the neural network model.

[0125] In short, to optimize the feature alignment effect of cross-perspective images, the present invention introduces a contrast loss (InfoNCE loss) and a cross-entropy loss (Cross-Entropy loss). Through these loss functions, the model can effectively learn the difference between the positive class and the negative class, and further improve the accuracy of cross-perspective geolocation.

[0126] The contrast loss is expressed as follows:

[0127]

[0128] Among them, q is the feature vector of the query image, R is the set of reference images, and r + is the positive sample corresponding to the query image, and τ is the temperature parameter used to control the scaling of the similarity score;

[0129] The cross-entropy loss is expressed as follows:

[0130]

[0131] Among them, p^(y∣x) is the predicted probability of the neural network model, and lp(y) is the logarithm of the predicted probability of the positive class;

[0132] The optimization objective of the present invention is to enable the network to learn an effective representation of cross-view deep features through the combination of contrastive learning, cross-entropy loss, and negative-class KL divergence loss, thereby improving the accuracy of cross-view image matching. The calculation formula of the final loss function of the neural network model is as follows:

[0133] Loss=L infoNCE +λL CE +λ 2 L NCKL ;

[0134] Among them, λ1 and λ2 are adjustable hyperparameters.

[0135] The above solution can effectively improve the matching accuracy of satellite-view and drone-view images by comprehensively using the proposed shape-adaptive and importance-aware attention module (SSA) and negative-class KL divergence loss (NCKL). Through multi-dimensional feature fusion, deep convolution optimization, and cross-view feature alignment, the present invention significantly improves the robustness of attention to target buildings or important regions during image offset and rotation in the geolocation task.

[0136] The present invention proposes a deep learning method based on satellite-view and drone-view image matching. By introducing a novel SSA module and NCKL loss function, the performance of the cross-view image matching task is significantly improved. The specific technical effects are as follows:

[0137] 1. Enhanced robustness to image transformations such as translation and rotation: By using the SSA module, the present invention effectively realizes the cross-dimensional interaction of channel and spatial features, enabling the model to better focus on the features of the target building. Especially when facing image displacement and rotation, the model can still continuously focus on the target building area of the image. At the same time, the model can accurately perceive the building shapes in satellite-view and drone-view images, improving the robustness of image matching. The comparison result of the model heat map generated by the attention mechanism method of this application and the MCCG involved in the prior art during image processing is as Figures 5 - 6 shown.

[0138] Figure 5 and Figure 6 The top two rows and the bottom two rows of images are respectively selected from two scene categories in the University-1652 dataset. The images in the first column are the input images, and the images in the second, third, and fourth columns are the generated heatmaps. The heatmap in the third column is generated after the input image is translated 80 pixels to the right, while the heatmap in the fourth column is generated after the input image is rotated 90 degrees clockwise. By comparing the images in the second column with those in the third column, it can be observed that MCCG focuses on the inconsistent areas between the original input image and the translated image, and its attention still remains concentrated in the middle of the transformed image. In addition, by comparing the images in the third column with those in the fourth column, MCCG fails to focus on the target building areas in the original image and the rotated image, showing its poor perception ability of the target building shape.

[0139] 2. Improvement in cross-view image matching performance: Compared with the existing technologies, the present invention has achieved significant performance improvements on multiple benchmark datasets (such as University-1652 and SUES-200). On the University-1652 dataset, the method proposed by the present invention is superior to the methods involved in the existing technologies in all evaluation metrics. Especially in the case of image rotation and displacement, it can effectively maintain the perception of the building shape. On the SUES-200 dataset, the present invention shows superior performance compared with traditional methods (such as CCR and Sample4Geo) under all settings, especially the improvement in accuracy in the UAV target positioning task. The results are as Figure 7 and Figure 8 shown.

[0140] 3. Comparison with some attention mechanisms: As Figure 9 shown, ablation experiments on different attention mechanisms are carried out in this application. Compared with the multi-scale convolutional kernels and channel attention used in SKNet, the SSA proposed in this application integrates the channel attention map and the spatial attention map, and uses multi-scale bar-shaped convolutions to capture the directional features of the feature map, thus achieving better performance. The performance of SENet is not as good as that of SSA because it ignores the importance of spatial attention. Although CBAM combines channel attention and spatial attention, it separates the two and fails to achieve cross-dimensional feature fusion. In contrast, the proposed SSA integrates channel and spatial information simultaneously in two stages, thus making the expression of information more abundant.

[0141] 4. Comparison of Different Loss Functions: The traditional KL divergence loss considers the alignment of the positive and negative class distributions. However, in cross-view tasks, the probability distribution of the positive class often dominates too much, masking the negative class information. The proposed NCKL loss only focuses on the negative class distribution and reduces the interference of the positive class on feature learning by excluding the positive class. As Figure 10 shown, compared with the traditional KL divergence loss, using the NCKL loss significantly improves the performance on the SUES-200 dataset. At the same time, compared with the Triplet loss, the proposed NCKL loss combines more effectively with the cross-entropy loss and can learn discriminative features of cross-view images. Compared with these two loss functions, in the Drone→Satellite task on the SUES-200 dataset, the Recall@1 using the NCKL loss is increased by about +4.5%, and the AP is increased by +4%.

[0142] In a second aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the cross-view geolocation method based on the attention mechanism as described in the first aspect of the present invention.

[0143] Among them, the computer-readable storage medium may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories.

[0144] The non-volatile memory may be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD ROM, Compact Disc Read Only Memory); the magnetic surface memory may be a disk memory or a tape memory.

[0145] The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), sync link dynamic random access memory (SLDRAM), direct rambus random access memory (DRRAM). The computer-readable storage medium described in the embodiments of the present invention is intended to include these and any other suitable types of memory.

[0146] As Figure 11 shown, in a third aspect, the present invention provides an electronic device 10, including a processor 101 and a storage medium 102, where a computer program is stored on the storage medium, and when the computer program is executed by the processor, the cross-perspective geolocation method based on the attention mechanism as described in the first aspect of the present invention is implemented.

[0147] In some embodiments, the processor may be implemented by software, hardware, firmware, or a combination thereof, and at least one of a circuit, a single or multiple application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), central processing units (CPUs), controllers, microcontrollers, and microprocessors may be used, so that the processor can execute some steps, all steps, or any combination of the steps in the cross-perspective geolocation method based on the attention mechanism in each embodiment of the present application.

[0148] Finally, it should be noted that although the above embodiments have been described in the text and drawings of the specification of the present invention, the patent protection scope of the present invention cannot be limited thereby. Any equivalent structure or equivalent process substitution or modification made by using the content recorded in the text and drawings of the specification of the present invention based on the substantial concept of the present invention, as well as the technical solutions directly or indirectly implemented in other related technical fields of the above embodiments, are all included in the patent protection scope of the present invention.

Claims

1. A cross-view geolocation method based on attention mechanism, characterized in that: The following steps are involved: Receive image information taken by drones or satellites; Inputting the image information into a trained neural network model, and outputting geographic location information corresponding to the image information; When the neural network model is trained, the following steps are included: S1: receiving a first sample image captured from a satellite perspective or a second sample image captured from a drone perspective, and extracting a feature map from the first sample image or the second sample image using a pre-trained ConvNeXt network; S2: Divide the extracted feature map into a first branch tensor, a second branch tensor and a third branch tensor, perform ECA operation on the first branch tensor, the second branch tensor and the third branch tensor respectively, and then perform global average pooling to obtain a first channel description vector, a second channel description vector and a third channel description vector; S3: Use the first channel description vector to perform weighted calculation on the first branch tensor, use the second channel description vector to perform weighted calculation on the second branch tensor, and use the third channel description vector to perform weighted calculation on the third branch tensor to obtain first attention feature maps corresponding to the three branches, average the three first attention feature maps, and output a second attention feature map; S4: performing a deep convolution operation on the second attention feature map to obtain an enhanced third attention feature map; S5: Determine geographic location information based on the third attention feature map.

2. The cross-view geolocation method based on the attention mechanism as claimed in claim 1, characterized in that: The feature map extracted from the first sample image or the second sample image using the pre-trained ConvNeXt network is expressed by the following formula: Where i∈{1,2}, x1 represents the first sample image, x2 represents the second sample image, L i ∈R C×H×W represents the extracted feature map, C is the number of channels, and H and W represent the height and width of the image respectively.

3. The cross-view geolocation method based on the attention mechanism as claimed in claim 1, characterized in that: The shape of the feature map is C×H×W, where C is the number of channels, and H and W represent the height and width of the image respectively; The first branch tensor is expressed as W×H×C, the second branch tensor is expressed as H×C×W, and the third branch tensor retains the original shape of the feature map and is expressed as C×H×W; Steps S2 and S3 are expressed by the following formula: Among them, L D Represents the first branch tensor, the second branch tensor, and the third branch tensor of the input, X D represents the first attention feature map of the output, X represents the second attention feature map, and the symbol σ represents the Sigmoid activation function.

4. The cross-view geolocation method based on the attention mechanism as claimed in claim 1, characterized in that: Step S4 includes: Use convolution kernels of different lengths to perform deep convolution operations, and use a 1×1 convolution layer to achieve information interaction between channels to obtain an enhanced third attention feature map; Step S4 is represented by the following formula: Among them, DwConv represents a 5×5 deep convolution, Pathi is the result of two deep strip convolutions and the body after 5×5 deep convolution, and Y is the enhanced third attention feature map.

5. The cross-view geolocation method based on the attention mechanism as claimed in claim 1, characterized in that: When the neural network model is being trained, the following steps are also included: When calculating KL divergence, ignore the predicted probability of the positive class and reapply the softmax function to calculate the predicted probability of the negative class, including: The third attention feature map is passed to the classifier after the average pooling operation, and the logit score of each category is obtained by the classifier. The logit score of each category is passed to the Log-Softmax layer to obtain the logarithm of the predicted probability of each category, which is specifically expressed by the following formula: L=Log_Softmax(Classifier(AvgPool(Y))); Where L = {l1,l2,...,l p ,...,l N }∈R 1×N represents the logarithm of the predicted probability of each category, N represents the total number of categories, and the subscript p represents the positive class; To compute the negative class KL divergence loss, lp is ignored, so L is rewritten as: Next, L^ passes through the Softmax layer to obtain the predicted probability distribution, which is expressed as follows: Where P represents the predicted probability distribution of all negative classes.

6. The cross-view geolocation method based on the attention mechanism as claimed in claim 5, characterized in that: The method further comprises: The output results of the first sample image collected from the satellite perspective and the second sample image collected from the drone perspective learn each other's prediction probability distribution, thereby aligning the deep feature distributions of the two. The negative KL divergence calculation process is expressed by the following formula: Wherein, P1 and P2 represent the negative class prediction probability distribution of the first sample image and the second sample image respectively.

7. The cross-view geolocation method based on the attention mechanism as claimed in claim 6, characterized in that: The method comprises: Based on the comprehensive calculation of KL divergence loss, contrast loss and cross entropy loss, the final loss function of the neural network model is obtained.

8. The cross-view geolocation method based on the attention mechanism as claimed in claim 7, characterized in that: The contrast loss is expressed as follows: Among them, q is the feature vector of the query image, R is the set of reference images, and r + is the positive sample corresponding to the query image, τ is the temperature parameter used to control the scaling of the similarity score; The cross entropy loss is expressed as follows: Among them, p^(y|x) is the predicted probability of the neural network model, and lp(y) is the logarithm of the predicted probability of the positive class; The calculation formula of the final loss function of the neural network model is as follows: Loss=L infoNCE +λ1L CE +λ2L NCKL ; Among them, λ1 and λ2 are adjustable hyperparameters.

9. A computer-readable storage medium storing computer program instructions, characterized in that: The computer program instructions implement the method of any one of claims 1 to 8 when executed by a processor.

10. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Cross-view image geographic positioning method and device based on attention efficiency network

    CN116310701A

  • Cross-view cross-modal image geographic positioning method based on cooperation of CNN and cross-layer interaction Transform

    CN116310866A

  • Video anomaly detection method based on full attention memory module, medium and equipment

    CN117710861A

  • Cross-view geographic positioning method based on unit dot product attention mechanism

    CN118261970A

  • Geographic positioning method and system for cross-view image

    CN118968006A

Cited By

  • Cross-view target geographic positioning method and device based on cross-task knowledge migration

    CN121861497A