Cross-view geographic positioning method for multi-scale frequency perception attention fusion

By fusing low-frequency global structure and high-frequency detail information through a multi-scale frequency-aware attention fusion network model, the problem of fragmented feature representation in traditional cross-view geolocation methods is solved, the feature representation and recognition capabilities of cross-view images are improved, and the robustness and adaptability of the model are enhanced.

CN120932240APending Publication Date: 2025-11-11HARBIN INST OF TECH

Patent Information

Application Number
CN202511219773.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Traditional cross-view geolocation methods divide image content or feature maps by fixed sizes or preset rules, lacking awareness of the natural semantic boundaries of images. This leads to fragmented feature representation and affects the model's accurate modeling and matching of cross-view semantic correspondences.

Method used

A multi-scale frequency-aware attention fusion network model is adopted, which combines the visual basic model EVA02, a multi-scale attention fusion module, an average pooling layer, and a multi-classifier module. By fusing low-frequency global structure and high-frequency detail information through a multi-frequency branch module and a frequency-aware spatial attention module, the modeling capability of fine-grained texture features of images is optimized.

Benefits of technology

It significantly enhances the feature representation and recognition capabilities of cross-view images, improves the model's robustness to changing viewpoints and adaptability to large-scale datasets, and increases matching accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932240A_ABST
    Figure CN120932240A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-view-angle geographic positioning method for multi-scale frequency perception attention fusion, and relates to a cross-view-angle geographic positioning method for an unmanned aerial vehicle scene. The objective of the invention is to improve the cross-view-angle geographic positioning accuracy of an existing unmanned aerial vehicle scene. The method comprises the following steps of: 1, acquiring cross-view-angle image pairs in different scenes and corresponding geographic position label data sets; 2, constructing a cross-view multi-scale frequency perception attention fusion network model; 3, obtaining a trained cross-view multi-scale frequency perception attention fusion network model; 4, inputting a to-be-detected unmanned aerial vehicle visual angle image without a label and a satellite visual angle image with a label into the trained model, and outputting a feature vector of the unmanned aerial vehicle visual angle image and a feature vector of the satellite visual angle image by the trained model; and 5, selecting the position corresponding to the satellite image with the highest similarity with the unmanned aerial vehicle image as the geographic position of the unmanned aerial vehicle. The method is applied to the field of cross-view geographic positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a cross-view geolocation method for drone scenarios. Background Technology

[0002] With the rapid development of drone technology and the continuous reduction in production costs, it has been widely applied in many fields. Traditional drone geolocation technology usually relies on the Global Positioning System (GPS), but GPS signals are easily interfered with by factors such as weather and buildings, and may be completely lost in some complex environments, failing to provide accurate location information. Cross-view geolocation in drone scenarios, as an emerging positioning method, can accurately determine the drone's geographical location by matching drone images with satellite images carrying GPS signals, without relying on the drone's GPS signal. The cross-view geolocation task based on visual algorithms aims to extract a unified feature representation of image pairs in a high-dimensional feature space, and measure the similarity of image pairs by calculating the similarity of high-dimensional features. However, due to the difference in shooting angles between drone and satellite images, images of the same location often have significant differences in appearance. Therefore, how to effectively reduce the feature differences between images from different perspectives has become one of the main challenges facing cross-view geolocation.

[0003] Currently, cross-view geolocation methods typically learn region-level feature representations by spatially partitioning the complete image content or its feature maps. Such partitioning is usually based on fixed dimensions or pre-defined rules, lacking awareness and adaptability to natural semantic boundaries within the image. Because this "mechanical" segmentation ignores the semantic integrity of objects in the image, when the partition boundary crosses key targets, such as a building, it disrupts the structural continuity of the target, leading to fragmented feature representations and affecting the model's accurate modeling and matching of cross-view semantic correspondences. Summary of the Invention

[0004] The purpose of this invention is to improve the accuracy of cross-view geolocation in existing drone scenarios, and a multi-scale frequency perception attention fusion cross-view geolocation method is proposed.

[0005] The specific process of a multi-scale frequency-aware attention fusion cross-view geolocation method is as follows:

[0006] Step 1: Obtain cross-view image pairs and corresponding geographic location label datasets from different scenarios;

[0007] Cross-view image pairs include drone-view images and satellite-view images;

[0008] Step 2: Construct a cross-view, multi-scale frequency-aware attention fusion network model;

[0009] The cross-view multi-scale frequency perception attention fusion network model consists of the visual base model EVA02, a multi-scale attention fusion module, an average pooling layer, and a multi-classifier module.

[0010] Step 3: Input the cross-view image pairs and corresponding geographic location label datasets obtained in Step 1 into the cross-view multi-scale frequency-aware attention fusion network model. The cross-view multi-scale frequency-aware attention fusion network model outputs the classification results until the loss function converges, and the trained cross-view multi-scale frequency-aware attention fusion network model is obtained.

[0011] Step 4: Input the unlabeled UAV view image and the labeled satellite view image into the trained cross-view multi-scale frequency perception attention fusion network model. The multi-classifier module in the trained cross-view multi-scale frequency perception attention fusion network model outputs the feature vectors of the UAV view image and the satellite view image.

[0012] Step 5: Calculate the cosine similarity between the feature vectors of the UAV view image and the feature vectors of the satellite view image, until all the feature vectors of the labeled satellite view images have been traversed. Select the location corresponding to the satellite image with the highest similarity as the UAV's geographical location.

[0013] The beneficial effects of this invention are as follows:

[0014] This invention addresses the aforementioned problems by constructing a multi-scale attention fusion module. (Currently, cross-view geolocation methods typically learn region-level feature representations by spatially partitioning the complete image content or its feature maps. Such partitioning is usually based on fixed dimensions or preset rules, lacking awareness and adaptability to natural semantic boundaries in the image. Because this "mechanical" segmentation ignores the semantic integrity of objects in the image, when the partition boundary crosses a key target, such as a building, it disrupts the structural continuity of the target, leading to fragmented feature representation and affecting the model's accurate modeling and matching of cross-view semantic correspondences.) By combining a multi-frequency branch module and a frequency-aware spatial attention module to fully integrate low-frequency global structure and high-frequency detail information, the ability to model fine-grained texture features of the image is optimized, significantly enhancing the feature representation and recognition capabilities of cross-view images and improving robustness to changing viewpoints.

[0015] The method of this invention introduces the visual basic model EVA02 into the cross-view geolocation task and combines it with a multi-scale attention fusion module, thereby improving the model's dual modeling ability for global semantic information and local detailed features, and enhancing the model's adaptability and performance on large-scale datasets.

[0016] To verify the performance of this invention, it was validated on a dataset of real UAV-satellite image pairs. Experimental results show that the proposed method achieves higher matching accuracy compared to current representative methods. The experimental results validate the effectiveness of the multi-scale frequency-aware attention fusion cross-view geolocation method proposed in this invention. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating the implementation of the present invention;

[0018] Figure 2a This is a flowchart of the multi-scale attention fusion module in this invention;

[0019] Figure 2b This is a detailed flowchart of the multi-scale low-frequency branch in the multi-scale attention fusion module of this invention;

[0020] Figure 2c This is a flowchart detailing the mixed high-frequency branches in the multi-scale attention fusion module of this invention;

[0021] Figure 2d This is a detailed flowchart of the frequency-aware spatial attention module in the multi-scale attention fusion module of this invention;

[0022] Figure 3a This is a comparison chart of the performance indicators of different methods in UAV positioning tasks;

[0023] Figure 3b This is a comparison chart of the performance indicators of different methods in UAV navigation tasks;

[0024] Figure 4 This is a qualitative analysis and comparison chart of positioning results from different methods in UAV navigation missions;

[0025] Figure 5 This is a qualitative analysis and comparison chart of the positioning results of different methods in UAV positioning tasks. Detailed Implementation

[0026] Specific implementation method one: Combining Figure 1 This embodiment describes a multi-scale frequency-sensing attention fusion cross-view geolocation method, the specific process of which is as follows:

[0027] Step 1: Obtain cross-view image pairs and corresponding geographic location label datasets from different scenarios;

[0028] Cross-view image pairs include drone-view images and satellite-view images;

[0029] Step 2: Construct a cross-view, multi-scale frequency-aware attention fusion network model;

[0030] The cross-view multi-scale frequency perception attention fusion network model consists of the visual base model EVA02, a multi-scale attention fusion module, an average pooling layer, and a multi-classifier module.

[0031] Step 3: Input the cross-view image pairs and corresponding geographic location label datasets obtained in Step 1 into the cross-view multi-scale frequency-aware attention fusion network model. The cross-view multi-scale frequency-aware attention fusion network model outputs the classification results until the loss function converges, and the trained cross-view multi-scale frequency-aware attention fusion network model is obtained.

[0032] Step 4: Input the unlabeled UAV view image and the labeled satellite view image into the trained cross-view multi-scale frequency perception attention fusion network model. The multi-classifier module in the trained cross-view multi-scale frequency perception attention fusion network model outputs the feature vectors of the UAV view image and the satellite view image.

[0033] Step 5: Calculate the cosine similarity between the feature vectors of the UAV view image and the feature vectors of the satellite view image, until all the feature vectors of the labeled satellite view image have been traversed. Select the location (GPS information) corresponding to the satellite image with the highest similarity as the UAV's geographical location, thereby achieving cross-view geolocation and evaluating the positioning accuracy based on the real labels.

[0034] Specific Implementation Method Two: This implementation method differs from Specific Implementation Method One in that the multi-scale attention fusion module in step two includes a multi-frequency branch module and a frequency-aware spatial attention module;

[0035] The multi-frequency branching module includes multi-scale low-frequency branches, hybrid high-frequency branches, and original branches;

[0036] The frequency-aware spatial attention module includes mean pooling layer, max pooling layer, convolutional layer, batch normalization (BN) layer, ReLU activation function layer, convolutional layer, and sigmoid activation function layer.

[0037] The other steps and parameters are the same as in Specific Implementation Method 1.

[0038] Specific Implementation Method 3: This implementation method differs from Specific Implementation Method 1 or 2 in that the multi-classifier module includes a first feature extraction module, a first classification module, a second feature extraction module, a second classification module, a third feature extraction module, and a third classification module;

[0039] The first feature extraction module includes, in sequence, a linear layer, a batch normalization (BN) layer, and a dropout layer;

[0040] The second feature extraction module includes, in sequence, a linear layer, a batch normalization (BN) layer, and a dropout layer;

[0041] The third feature extraction module includes, in sequence, a linear layer, a batch normalization (BN) layer, and a dropout layer;

[0042] The first classification module includes, in sequence, a linear layer and a Softmax activation function layer;

[0043] The second classification module includes a linear layer and a Softmax activation function layer in sequence;

[0044] The third classification module includes a linear layer and a Softmax activation function layer.

[0045] Other steps and parameters are the same as in specific implementation method one or two.

[0046] Specific implementation method four: This implementation method differs from one of the specific implementation methods one to three in that, in step three, the cross-view image pairs under different scenarios and the corresponding geographic location label datasets obtained in step one are input into the cross-view multi-scale frequency perception attention fusion network model, the cross-view multi-scale frequency perception attention fusion network model outputs the classification result until the loss function converges, and the trained cross-view multi-scale frequency perception attention fusion network model is obtained.

[0047] The specific process is as follows:

[0048] Step 3: Input the UAV view image and corresponding geographic location label data obtained in Step 1 into the visual base model EVA02 in the cross-view multi-scale frequency perception attention fusion network model. The visual base model EVA02 outputs the UAV view image features. , ; Indicates the height of the image. Indicates the width of the image. Indicates the number of channels in an image;

[0049] Step 3.2: Input the satellite view image and corresponding geographic location label data obtained in Step 1 into the visual base model EVA02 in the cross-view multi-scale frequency perception attention fusion network model. The visual base model EVA02 outputs the satellite view image features. , ; Indicates the height of the image. The width of the table image, Indicates the number of channels in the image;

[0050] Step 33: Output the UAV viewpoint image features from the visual base model EVA02. Input to a multi-scale attention fusion module, output features from the multi-scale attention fusion module. , , ;

[0051] Steps 3 and 4: Output satellite view image features from the visual basic model EVA02 Input to a multi-scale attention fusion module, output features from the multi-scale attention fusion module. , , The method is the same as step three.

[0052] Step 35

[0053] Features Input to average pooling layer, output features from average pooling layer ;

[0054] Features Input to average pooling layer, output features from average pooling layer ;

[0055] Features Input to average pooling layer, output features from average pooling layer ;

[0056] Step 36

[0057] Features Input to average pooling layer, output features from average pooling layer ;

[0058] Features Input to average pooling layer, output features from average pooling layer ;

[0059] Features Input to average pooling layer, output features from average pooling layer ;

[0060] Step 37: Features ,feature ,feature Input to the multi-classifier module, and the multi-classifier module outputs the classification results. , , ;

[0061] Step 38: Features ,feature ,feature Input to the multi-classifier module, and the multi-classifier module outputs the classification results. , , The method is the same as step three seven;

[0062] Step 39: Repeat steps 31 to 39 until the loss function converges, and obtain the trained cross-view multi-scale frequency-aware attention fusion network model.

[0063] The other steps and parameters are the same as those in one of the specific implementation methods one to three.

[0064] Specific Implementation Method Five: This implementation method differs from Specific Implementation Methods One to Four in that, in step three, the UAV view image and corresponding geographic location tag data obtained in step one are input into the visual base model EVA02 in the cross-view multi-scale frequency perception attention fusion network model. The visual base model EVA02 outputs the UAV view image features. , ; Indicates the height of the image from the drone's perspective. Indicates the width of the image from the drone's perspective. This indicates the number of channels in the image viewed from the drone's perspective;

[0065] The specific process is as follows:

[0066] Step 3.11: The UAV view image and corresponding geolocation label data obtained in Step 1 are represented as follows: ;

[0067] in, Indicates the first Images from the perspective of drones in various categories; Representation and Image Corresponding scene category tags;

[0068] Indicates the first One category; Indicates the total number of categories;

[0069] Step 3.12: Input the UAV view image and corresponding geographic location label data obtained in Step 1 into the visual base model EVA02 in the cross-view multi-scale frequency perception attention fusion network model. The visual base model EVA02 outputs the UAV view image features. , ; indicates as:

[0070]

[0071] Indicates the image Input the basic visual model EVA02.

[0072] The other steps and parameters are the same as those in one of the specific implementation methods one to four.

[0073] Specific Implementation Method Six: This implementation method differs from Specific Implementation Methods One to Five in that, in step three-two, the satellite view image obtained in step one and the corresponding geographic location label data are input into the visual base model EVA02 in the cross-view multi-scale frequency perception attention fusion network model, and the visual base model EVA02 outputs satellite view image features. , ; Indicates the altitude of the satellite image. Indicates the width of the satellite-view image. This indicates the number of channels in a satellite-view image.

[0074] The specific process is as follows:

[0075] Step 3, Step 2, and Step 1: The satellite view images and corresponding geographic location label data obtained are represented as follows: ;

[0076] in, Indicates the first Satellite view images of various categories; Representation and Image Corresponding scene category tags;

[0077] Indicates the first One category; Indicates the total number of categories;

[0078] Step 3.22: Input the satellite view image and corresponding geographic location label data obtained in Step 1 into the visual base model EVA02 in the cross-view multi-scale frequency perception attention fusion network model. The visual base model EVA02 outputs the satellite view image features. , ; indicates as:

[0079]

[0080] Indicates the image Input the basic visual model EVA02.

[0081] For example, when inputting image pairs with dimensions of 448×448×3 (H×W×C), This will generate two feature maps with a size of 8×1025×768. EVA02 is a Transformer-based visual foundational model that optimizes the traditional Transformer architecture by replacing the activations in the traditional visual Transformer with a SwiGLU (Sigmoid-Gated Linear Unit) gating mechanism. To ensure consistency in parameters and floating-point operations, EVA02 sets the hidden dimension of the position feedforward network to 2 / 3 of that of a traditional multilayer perceptron, thus achieving efficiency optimization through computational balance. Furthermore, EVA02 employs a Xavier Normal initialization strategy to address the performance degradation that may occur during random initialization of the SwiGLU gating mechanism, thereby improving the stability of the training process. In terms of network structure, EVA02 uses a post-normalization structure instead of the traditional input-first normalization, thus mitigating the gradient vanishing problem in deep networks and enhancing the stability of gradient flow. Unlike traditional visual Transformers that use absolute or standard relative embedding, EVA02 employs a two-dimensional RoPE (Relative Positional Encoding) scheme. This maintains spatial relationship modeling capabilities while avoiding instability during pre-training. EVA02 achieves image-text alignment visual feature learning by reconstructing masked regions based on visible image patches. During the pre-training phase, the model is trained using a large-scale dataset composed of multiple large datasets to fully explore and unleash the potential of large models in visual representation learning.

[0082] The other steps and parameters are the same as those in one of the specific implementation methods one to five.

[0083] Specific Implementation Method Seven: This implementation method differs from Specific Implementation Methods One through Six in that, in step three, the visual basic model EVA02 outputs the UAV viewpoint image features. Input to the multi-scale attention fusion module, output of the multi-scale attention fusion module , , ;

[0084] The specific process is as follows:

[0085] Input feature map After multi-scale low-frequency branching and hybrid high-frequency branching, features of interest in low-frequency information are obtained respectively. and high-frequency information ;

[0086] Step 331: Output the UAV viewpoint image features from the basic visual model EVA02. Input multi-scale low-frequency branch, output multi-scale low-frequency branch features The specific process is as follows:

[0087] 1) Hierarchical information extraction is achieved through average pooling operations at three scales, expressed as:

[0088]

[0089]

[0090]

[0091] in,

[0092] Indicates the kernel size as Average pooling operation, These are learnable channel weight parameters. This indicates a channel-by-channel multiplication operation. Indicates the kernel size; Indicates characteristics;

[0093] Indicates the kernel size as Average pooling operation, These are learnable channel weight parameters. Indicates the kernel size; Indicates characteristics;

[0094] Indicates the kernel size as Average pooling operation, These are learnable channel weight parameters. Indicates the kernel size; Indicates characteristics;

[0095] 2) Features ,feature ,feature By fusing, features are obtained. ,feature As a multi-scale low-frequency branch output feature The formula is:

[0096]

[0097] This design creates a hierarchical structure representation from local to global: 3×3 pooling can capture small-scale texture patterns, 5×5 pooling integrates medium-scale structural information, and 7×7 pooling helps to capture global semantic layout, thus achieving multi-scale structure representation.

[0098] Step 332: Output the UAV viewpoint image features from the visual base model EVA02. Input hybrid high-frequency branch, hybrid high-frequency branch output features The specific process is as follows:

[0099] 1) Four-directional Sobel operators are used to extract high-frequency edge features. The frequency response characteristics of the four-directional Sobel operators are as follows:

[0100]

[0101] in,

[0102] express The frequency response of the Sobel operator in the direction;

[0103] express The frequency response of the Sobel operator in the direction;

[0104] express The frequency response of the Sobel operator in the direction is used to detect edges or textures in an image along the diagonal (trend from top left to bottom right);

[0105] express The frequency response of the Sobel operator in the direction is used to detect edges or textures in an image along the diagonal (trend from the upper right to the lower left).

[0106] Direction: A second-order mixed direction, first horizontal and then vertical;

[0107] Direction: A second-order mixed direction, first vertical and then horizontal;

[0108] 2) Output the UAV perspective image features from the visual base model EVA02 Convolving the edge response features with the frequency response characteristics of the Sobel operators in the four directions respectively yields the edge response features in the four directions; represented as:

[0109]

[0110]

[0111]

[0112]

[0113] in,

[0114] express Edge response characteristics of the direction; express Edge response characteristics of the direction;

[0115] express Edge response characteristics of the direction; express Edge response characteristics of the direction;

[0116] Represents convolution; Indicates taking the absolute value;

[0117] 3) To enhance the discriminative power of edge features, Edge response characteristics of direction , Edge response characteristics of direction , Edge response characteristics of direction , Edge response characteristics of direction By stitching along the channel dimension, the features are obtained. , ;

[0118] 4) Features The process proceeds sequentially through a convolutional layer, a ReLU activation function layer, another convolutional layer, and a sigmoid function, with the sigmoid function generating channel weights. ;

[0119] 5) Edge response characteristics of direction , Edge response characteristics of direction , Edge response characteristics of direction , Edge response characteristics of direction Pixel summation yields the feature. , ;

[0120] 6) Adjust channel weights With features Pixel-by-pixel multiplication yields the hybrid high-frequency branch output features. , :

[0121]

[0122] in,

[0123] This indicates pixel-by-pixel multiplication;

[0124] Step 333: Output the UAV viewpoint image features from the visual base model EVA02. As original features , ;

[0125] Step 334: Output features of multi-scale low-frequency branches Input frequency-aware spatial attention module, output frequency-aware spatial attention module features The specific process is as follows:

[0126] 1) Calculate the multi-scale low-frequency branch output characteristics The mean pooling value and max pooling value are used to generate features for two channels. :

[0127]

[0128] in,

[0129] This indicates that the output features of the multi-scale low-frequency branch will be... The average pooled value obtained after the average pooling layer;

[0130] This indicates that the output features of the multi-scale low-frequency branch will be... The maximum pooling value obtained after the maximum pooling layer;

[0131] 2) Features containing two channels Input convolutional layer, batch normalized (BN) layer, ... Activation function layer, convolutional layer, sigmoid activation function layer; sigmoid activation function layer generates attention weights. (Using convolution to extract features from two channels) Dimensionality is reduced, then increased, and finally attention weights are generated through activation functions.

[0132] 3) Based on attention weights and characteristics Obtain features ; indicates as:

[0133]

[0134] Step 335: Output features of multi-scale high-frequency branches Input frequency-aware spatial attention module, output frequency-aware spatial attention module features The specific process is as follows:

[0135] 1) Calculate the output characteristics of multi-scale high-frequency branches The mean pooling value and max pooling value are used to generate features for two channels. :

[0136]

[0137] in,

[0138] This indicates that the multi-scale high-frequency branch output features will be used. The average pooled value obtained after the average pooling layer;

[0139] This indicates that the multi-scale high-frequency branch output features will be used. The maximum pooling value obtained after the maximum pooling layer;

[0140] 2) Features containing two channels Input convolutional layer, batch normalized (BN) layer, ... Activation function layer, convolutional layer, sigmoid activation function layer; sigmoid activation function layer generates attention weights. (Using convolution to extract features from two channels) Dimensionality is reduced, then increased, and finally attention weights are generated through activation functions.

[0141] 3) Based on attention weights and characteristics Obtain features ; indicates as:

[0142]

[0143] Step 336: Extract multi-scale original branch features As a feature .

[0144] The other steps and parameters are the same as those in one of the specific implementation methods one to six.

[0145] Specific Implementation Method Eight: This implementation method differs from Specific Implementation Methods One to Seven in that, in step three-seven, the feature... ,feature ,feature Input to the multi-classifier module, and the multi-classifier module outputs the classification results. , , ;

[0146] The specific process is as follows:

[0147] The multi-classifier module includes a feature extraction block and a classification block;

[0148] 1) Features Input feature extraction block, output feature ;feature Input a classification block, and the classification block outputs the classification result. The specific process is as follows:

[0149] 11) Features Input feature extraction block, output feature The specific process is as follows:

[0150] feature The input layers are sequentially a Linear layer, a Batch Normalization (BN) layer, and a Dropout layer. The Dropout layer outputs the features. ;

[0151] 12) Features Input a classification block, and the classification block outputs the classification result. The specific process is as follows:

[0152] feature The input is sequentially processed by a linear layer and a softmax activation function layer. The softmax activation function layer outputs the classification result. ;

[0153] 2) Features Input feature extraction block, output feature ;feature Input a classification block, and the classification block outputs the classification result. The specific process is as follows:

[0154] 21) Features Input feature extraction block, output feature The specific process is as follows:

[0155] feature The input layers are sequentially a Linear layer, a Batch Normalization (BN) layer, and a Dropout layer. The Dropout layer outputs the features. ;

[0156] 22) Features Input a classification block, and the classification block outputs the classification result. The specific process is as follows:

[0157] feature The input is sequentially processed by a linear layer and a softmax activation function layer. The softmax activation function layer outputs the classification result. ;

[0158] 3) Features Input feature extraction block, output feature ;feature Input a classification block, and the classification block outputs the classification result. The specific process is as follows:

[0159] 31) Features Input feature extraction block, output feature The specific process is as follows:

[0160] feature The input layers are sequentially a Linear layer, a Batch Normalization (BN) layer, and a Dropout layer. The Dropout layer outputs the features. ;

[0161] 32) Features Input a classification block, and the classification block outputs the classification result. The specific process is as follows:

[0162] feature The input is sequentially processed by a linear layer and a softmax activation function layer. The softmax activation function layer outputs the classification result. .

[0163] The input to each branch in the multi-classifier module is the feature vector obtained through the multi-scale attention fusion module. The feature vector is obtained after average pooling. Subsequently, these feature vectors are passed through a multi-classifier module consisting of three classifier layers to generate discriminative feature representations. During the training phase, each feature vector is sequentially processed by a feature extraction module and a classifier module, transforming it into a class vector. , , The model is optimized based on cross-entropy loss. During testing, the model only uses the feature extraction module to process the input image, generating a 3×512 feature representation. , , These three features are then concatenated into a 1536-dimensional feature vector, which is used for UAV localization and UAV navigation tasks.

[0164] The other steps and parameters are the same as those in any of the specific implementation methods one to seven.

[0165] Specific Implementation Method Nine: This implementation method differs from Specific Implementation Methods One through Eight in that the loss function acquisition process in step three nine is as follows:

[0166] loss function ;

[0167] in,

[0168] The loss function represents the cross-view, multi-scale frequency-aware attention fusion network model.

[0169] Represents the cross-entropy loss function;

[0170] Represents the cross-domain triplet loss function;

[0171]

[0172] in,

[0173] Features representing images from the drone's perspective Features corresponding to satellite view images (Including features corresponding to satellite view images of the same scene and different scenes) The cross-domain triplet loss function;

[0174] Features representing images from the drone's perspective Features corresponding to satellite view images (Including features corresponding to satellite view images of the same scene and different scenes) The cross-domain triplet loss function;

[0175] Features representing images from the drone's perspective Features corresponding to satellite view images (Including features corresponding to satellite view images of the same scene and different scenes) The cross-domain triplet loss function;

[0176]

[0177] in,

[0178] This represents a sample in the image from the drone's perspective, serving as an anchor point;

[0179] Indicates selecting from the satellite view image domain and Samples of the same category within the same scene are considered positive samples.

[0180] Indicates selecting from the satellite view image domain and Samples belonging to different categories in different scenarios are used as negative samples;

[0181] express The process sequentially passes through the visual base model EVA02, the multi-scale attention fusion module, and the average pooling layer, with the average pooling layer outputting features.

[0182] express The process sequentially passes through the visual base model EVA02, the multi-scale attention fusion module, and the average pooling layer, with the average pooling layer outputting features.

[0183] express The process sequentially passes through the visual base model EVA02, the multi-scale attention fusion module, and the average pooling layer, with the average pooling layer outputting features.

[0184] The function represents the distance between samples in the feature space. This invention uses Euclidean distance, i.e., Euclidean distance.

[0185] Indicates interval, ;

[0186]

[0187] in,

[0188] This represents the cross-entropy loss function corresponding to the image from the drone's perspective.

[0189] This represents the cross-entropy loss function corresponding to satellite view images;

[0190]

[0191] in, This represents the classification result output by the multi-classifier module. The cross-entropy loss function is calculated from the truth value; Indicates the true label;

[0192]

[0193] in, This represents the classification result output by the multi-classifier module. The cross-entropy loss function is calculated from the truth value;

[0194]

[0195] in, This represents the classification result output by the multi-classifier module. The cross-entropy loss function is calculated from the truth value;

[0196]

[0197] in, This represents the classification result output by the multi-classifier module. The cross-entropy loss function is calculated from the truth value;

[0198]

[0199] in, This represents the classification result output by the multi-classifier module. The cross-entropy loss function is calculated from the truth value;

[0200]

[0201] in, This represents the classification result output by the multi-classifier module. The cross-entropy loss function is calculated from the truth value;

[0202] The mathematical expression for the cross-entropy loss function is: ;in, The sample belongs to the category The true probability, Indicates the model predicts samples Category The probability. Since the true label is a "hot encoded" vector, assume the true label of the sample is... The model predicts the class probability distribution as follows: Then the cross-entropy loss can be simplified to:

[0203]

[0204] Cross-domain triplet loss functions aim to reduce distributional differences between different domains, addressing the significant differences arising from different data domains in cross-view geolocation and guiding the model to learn feature representations with cross-domain consistency. In traditional triplet loss, each sample triplet consists of an anchor sample, a positive sample of the same class, and a negative sample of a different class. Its core objective is to minimize the distance between the anchor sample and the positive sample in the feature space, while ensuring that the distance between the anchor sample and the negative sample is sufficiently large. Typically, this objective is constrained by setting a hyperparameter "margin," ensuring that the distance between the negative sample and the anchor is at least greater than the distance between the positive sample and the anchor. The mathematical expression for this objective is:

[0205]

[0206] in, Indicates anchor point sample, Indicates a positive sample. Indicates a negative sample; This invention represents a function for calculating the distance between samples in the feature space, and samples Euclidean distance. The hyperparameter "interval" mentioned above is used in experiments. By minimizing this loss function, the model can learn feature representations that bring similar samples closer together and push different samples further apart, thereby improving the performance of tasks such as classification or retrieval. In the context of cross-domain learning, the cross-domain triplet loss further expands this concept. Cross-view geolocation tasks involve two different view domains—the UAV view image domain. and satellite view image domain Their data distributions are as follows: and The goal of the model is to learn a feature map. This ensures that similar samples in the UAV-view image domain and the satellite-view image domain remain close in the mapped feature space, while dissimilar samples are far apart, thus overcoming the challenge posed by cross-domain data distribution differences. For example, in this invention, a sample in the UAV-view image domain is selected as an anchor point. And select from the satellite view image domain with Samples of the same category are considered positive samples. At the same time, choose with Samples belonging to different categories As a negative sample, the cross-domain triplet loss function can then be expressed as:

[0207] .

[0208] The other steps and parameters are the same as those in one of the specific implementation methods one to eight.

[0209] Specific Implementation Method Ten: This implementation method differs from Specific Implementation Methods One to Nine in that, in step four, the unlabeled UAV view image and the labeled satellite view image are input into the trained cross-view multi-scale frequency perception attention fusion network model. The multi-classifier module in the trained cross-view multi-scale frequency perception attention fusion network model outputs the feature vectors of the UAV view image and the satellite view image; the specific process is as follows:

[0210] Step 4: Input the unlabeled drone view image to be tested into the visual base model EVA02 in the trained cross-view multi-scale frequency perception attention fusion network model. The visual base model EVA02 outputs features and inputs them into the multi-scale attention fusion module. The multi-scale attention fusion module outputs 3 features and inputs them into the average pooling layer. The average pooling layer outputs 3 features and inputs them into the feature extraction block in the multi-classifier module. The feature extraction block outputs 3 feature vectors. The 3 feature vectors output by the feature extraction block are concatenated to obtain 1 feature vector corresponding to the unlabeled drone view image to be tested.

[0211] Step 4.2: Input the labeled satellite view image into the visual base model EVA02 in the trained cross-view multi-scale frequency perception attention fusion network model. The visual base model EVA02 outputs features and inputs them into the multi-scale attention fusion module. The multi-scale attention fusion module outputs three features and inputs them into the average pooling layer. The average pooling layer outputs three features and inputs them into the feature extraction block in the multi-classifier module. The feature extraction block outputs three feature vectors. The three feature vectors output by the feature extraction block are concatenated to obtain one feature vector corresponding to the labeled satellite view image.

[0212] The other steps and parameters are the same as those in any of the specific implementation methods one to nine.

[0213] The beneficial effects of the present invention are verified using the following embodiments:

[0214] Example 1:

[0215] A multi-scale frequency-aware attention fusion cross-view geolocation method is prepared according to the following steps:

[0216] The University-1652 dataset used in the experiment is a multi-view, multi-source benchmark dataset for UAV geolocation, containing data from satellite and UAV platforms. It collects data from 1652 buildings across 72 universities globally, with an average of 54 UAV-view images and one satellite image per building. The dataset is evenly divided into training and testing sets, containing 701 buildings from 33 and 39 universities respectively, with the remaining 250 buildings added as distractors to the image library. This dataset presents two key tasks: UAV localization (UAV view to satellite view) and UAV navigation (satellite view to UAV view). In the UAV localization task during the testing phase, each UAV-view query image corresponds to a uniquely matching satellite view image. The query set contains 37,855 UAV images, while the image library contains 701 true matching satellite images and 250 satellite distractors. In the UAV navigation task, the query set consists of 701 satellite images, the image library contains 37,855 corresponding UAV images, and 13,500 UAV distractors. Before inputting the model, the image size was uniformly adjusted to 448×448, and a series of data augmentation techniques were applied, including random fill, random rotation, random cropping, random flipping, and random erasure.

[0217] like Figure 3a and 3b As shown, the experiment compares the retrieval performance of existing methods (MCCG, FSRA) with those of UAV localization and UAV navigation, presenting comparative results for recall@1 (R@1) and average precision (AP) in both tasks. It can be seen that our proposed method (MFAF) achieves the highest accuracy in both UAV localization and UAV navigation tasks, and our method significantly outperforms the comparative methods in all metrics across both tasks.

[0218] Figure 4 , Figure 5 The image shows representative search results for two tasks, when the query image contains a clear and unique target (such as...). Figure 4 and Figure 5 (Second line in the image), all three methods can successfully retrieve the target building; however, when the target building is blurred or confused by a visually similar background (e.g., ... Figure 4 The first and fourth lines of the text Figure 5 In the fourth line of the image, MCCG and FSRA showed retrieval errors, while MFAF consistently identified correct matches, demonstrating its superior ability to distinguish visually similar content. This advantage is attributed to the proposed multi-frequency feature fusion strategy, which preserves global representations while capturing fine-grained details, enhancing the discriminability of cross-view features. Although in some challenging cases (such as...), MFAF... Figure 4 The third line and Figure 5(In the first row of the diagram), the performance of MFAF slightly decreases, but existing methods completely fail under such extreme viewpoint differences, further highlighting the robustness of MFAF. Experimental results verify the effectiveness of the multi-scale frequency-aware attention fusion method proposed in this invention in cross-viewpoint geolocation tasks.

[0219] This invention may have other embodiments. Without departing from the spirit and essence of this invention, those skilled in the art can make various corresponding changes and modifications according to this invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims.

Claims

1. A cross-view geolocation method based on multi-scale frequency-aware attention fusion, characterized in that: The specific process of the method is as follows: Step 1: Obtain cross-view image pairs and corresponding geographic location label datasets from different scenarios; Cross-view image pairs include drone-view images and satellite-view images; Step 2: Construct a cross-view, multi-scale frequency-aware attention fusion network model; The cross-view multi-scale frequency perception attention fusion network model consists of the visual base model EVA02, a multi-scale attention fusion module, an average pooling layer, and a multi-classifier module. Step 3: Input the cross-view image pairs and corresponding geographic location label datasets obtained in Step 1 into the cross-view multi-scale frequency-aware attention fusion network model. The cross-view multi-scale frequency-aware attention fusion network model outputs the classification results until the loss function converges, and the trained cross-view multi-scale frequency-aware attention fusion network model is obtained. Step 4: Input the unlabeled UAV view image and the labeled satellite view image into the trained cross-view multi-scale frequency perception attention fusion network model. The multi-classifier module in the trained cross-view multi-scale frequency perception attention fusion network model outputs the feature vectors of the UAV view image and the satellite view image. Step 5: Calculate the cosine similarity between the feature vectors of the UAV view image and the feature vectors of the satellite view image, until all the feature vectors of the labeled satellite view images have been traversed. Select the location corresponding to the satellite image with the highest similarity as the UAV's geographical location.

2. The cross-view geolocation method based on multi-scale frequency-sensing attention fusion according to claim 1, characterized in that: The multi-scale attention fusion module in step two includes a multi-frequency branch module and a frequency-aware spatial attention module. The multi-frequency branching module includes multi-scale low-frequency branches, hybrid high-frequency branches, and original branches; The frequency-aware spatial attention module includes mean pooling layer, max pooling layer, convolutional layer, batch normalization (BN) layer, ReLU activation function layer, convolutional layer, and sigmoid activation function layer.

3. The cross-view geolocation method based on multi-scale frequency-aware attention fusion according to claim 2, characterized in that: The multi-classifier module includes a first feature extraction module, a first classification module, a second feature extraction module, a second classification module, a third feature extraction module, and a third classification module; The first feature extraction module includes, in sequence, a linear layer, a batch normalization (BN) layer, and a dropout layer; The second feature extraction module includes, in sequence, a linear layer, a batch normalization (BN) layer, and a dropout layer; The third feature extraction module includes, in sequence, a linear layer, a batch normalization (BN) layer, and a dropout layer; The first classification module includes, in sequence, a linear layer and a Softmax activation function layer; The second classification module includes a linear layer and a Softmax activation function layer in sequence; The third classification module includes a linear layer and a Softmax activation function layer.

4. The cross-view geolocation method based on multi-scale frequency-aware attention fusion according to claim 3, characterized in that: In step three, the cross-view image pairs and corresponding geographic location label datasets obtained in step one under different scenarios are input into the cross-view multi-scale frequency-aware attention fusion network model. The cross-view multi-scale frequency-aware attention fusion network model outputs the classification results until the loss function converges, and the trained cross-view multi-scale frequency-aware attention fusion network model is obtained. The specific process is as follows: Step 3: Input the UAV view image and corresponding geographic location label data obtained in Step 1 into the visual base model EVA02 in the cross-view multi-scale frequency perception attention fusion network model. The visual base model EVA02 outputs the UAV view image features. , ; Indicates the height of the image. Indicates the width of the image. Indicates the number of channels in an image; Step 3.2: Input the satellite view image and corresponding geographic location label data obtained in Step 1 into the visual base model EVA02 in the cross-view multi-scale frequency perception attention fusion network model. The visual base model EVA02 outputs the satellite view image features. , ; Indicates the height of the image. The width of the table image, Indicates the number of channels in an image; Step 33: Output the UAV viewpoint image features from the visual base model EVA02. Input to a multi-scale attention fusion module, output features from the multi-scale attention fusion module. , , ; Steps 3 and 4: Output satellite view image features from the visual basic model EVA02 Input to a multi-scale attention fusion module, output features from the multi-scale attention fusion module. , , ; Step 35 Features Input to average pooling layer, output features from average pooling layer ; Features Input to average pooling layer, output features from average pooling layer ; Features Input to average pooling layer, output features from average pooling layer ; Step 36 Features Input to average pooling layer, output features from average pooling layer ; Features Input to average pooling layer, output features from average pooling layer ; Features Input to average pooling layer, output features from average pooling layer ; Step 37: Features ,feature ,feature Input to the multi-classifier module, and the multi-classifier module outputs the classification results. , , ; Step 38: Features ,feature ,feature Input to the multi-classifier module, and the multi-classifier module outputs the classification results. , , ; Step 39: Repeat steps 31 to 39 until the loss function converges, and obtain the trained cross-view multi-scale frequency-aware attention fusion network model.

5. The cross-view geolocation method based on multi-scale frequency-aware attention fusion according to claim 4, characterized in that: In step three, the UAV view image and corresponding geographic location tag data obtained in step one are input into the visual base model EVA02 in the cross-view multi-scale frequency perception attention fusion network model. The visual base model EVA02 outputs the UAV view image features. , ; Indicates the height of the image from the drone's perspective. Indicates the width of the image from the drone's perspective. This indicates the number of channels in the image viewed from the drone's perspective; The specific process is as follows: Step 3.11: The UAV view image and corresponding geolocation label data obtained in Step 1 are represented as follows: ; in, Indicates the first Images from the perspective of drones in various categories; Representation and Image Corresponding scene category tags; Indicates the first One category; Indicates the total number of categories; Step 3.12: Input the UAV view image and corresponding geographic location label data obtained in Step 1 into the visual base model EVA02 in the cross-view multi-scale frequency perception attention fusion network model. The visual base model EVA02 outputs the UAV view image features. , ; indicates as: Indicates the image Input the basic visual model EVA02.

6. The cross-view geolocation method based on multi-scale frequency-aware attention fusion according to claim 5, characterized in that: In step three, the satellite view image and corresponding geographic location tag data obtained in step one are input into the visual base model EVA02 in the cross-view multi-scale frequency perception attention fusion network model. The visual base model EVA02 outputs the satellite view image features. , ; Indicates the altitude of the satellite image. Indicates the width of the satellite-view image. Indicates the number of channels in a satellite-view image; The specific process is as follows: Step 3, Step 2, and Step 1: The satellite view images and corresponding geographic location label data obtained are represented as follows: ; in, Indicates the first Satellite view images of various categories; Representation and Image Corresponding scene category tags; Indicates the first One category; Indicates the total number of categories; Step 3.22: Input the satellite view image and corresponding geographic location label data obtained in Step 1 into the visual base model EVA02 in the cross-view multi-scale frequency perception attention fusion network model. The visual base model EVA02 outputs the satellite view image features. , ; indicates as: Indicates the image Input the basic visual model EVA02.

7. The cross-view geolocation method based on multi-scale frequency-aware attention fusion according to claim 6, characterized in that: In step three, the visual basic model EVA02 outputs the drone's perspective image features. Input to the multi-scale attention fusion module, output of the multi-scale attention fusion module , , ; The specific process is as follows: Step 331: Output the UAV viewpoint image features from the basic visual model EVA02. Input multi-scale low-frequency branch, output multi-scale low-frequency branch features The specific process is as follows: 1) Hierarchical information extraction is achieved through average pooling operations at three scales, expressed as: in, Indicates the kernel size as Average pooling operation, These are learnable channel weight parameters. This indicates a channel-by-channel multiplication operation. Indicates the kernel size; Indicates characteristics; Indicates the kernel size as Average pooling operation, These are learnable channel weight parameters. Indicates the kernel size; Indicates characteristics; Indicates the kernel size as Average pooling operation, These are learnable channel weight parameters. Indicates the kernel size; Indicates characteristics; 2) Features ,feature ,feature By fusing, features are obtained. ,feature As a multi-scale low-frequency branch output feature The formula is: Step 332: Output the UAV viewpoint image features from the visual base model EVA02. Input hybrid high-frequency branch, hybrid high-frequency branch output features The specific process is as follows: 1) The frequency response characteristics of the Sobel operator in four directions are as follows: in, express The frequency response of the Sobel operator in the direction; express The frequency response of the Sobel operator in the direction; express The frequency response of the Sobel operator in the direction; express The frequency response of the Sobel operator in the direction; Direction: A second-order mixed direction, first horizontal and then vertical; Direction: A second-order mixed direction, first vertical and then horizontal; 2) Output the UAV perspective image features from the visual base model EVA02 Convolving the edge response features with the frequency response characteristics of the Sobel operators in the four directions respectively yields the edge response features in the four directions; represented as: in, express Edge response characteristics of the direction; express Edge response characteristics of the direction; express Edge response characteristics of the direction; express Edge response characteristics of the direction; Represents convolution; Indicates taking the absolute value; 3) Edge response characteristics of direction , Edge response characteristics of direction , Edge response characteristics of direction , Edge response characteristics of direction By stitching along the channel dimension, the features are obtained. , ; 4) Features The process proceeds sequentially through a convolutional layer, a ReLU activation function layer, another convolutional layer, and a sigmoid function, with the sigmoid function generating channel weights. ; 5) Edge response characteristics of direction , Edge response characteristics of direction , Edge response characteristics of direction , Edge response characteristics of direction Pixel summation yields the feature. , ; 6) Adjust channel weights With features Pixel-by-pixel multiplication yields the hybrid high-frequency branch output features. , : in, This indicates pixel-by-pixel multiplication; Step 333: Output the UAV viewpoint image features from the visual base model EVA02. As original features , ; Step 334: Output features of multi-scale low-frequency branches Input frequency-aware spatial attention module, output frequency-aware spatial attention module features The specific process is as follows: 1) Calculate the multi-scale low-frequency branch output characteristics The mean pooling value and max pooling value are used to generate features for two channels. : in, This indicates that the output features of the multi-scale low-frequency branch will be... The average pooled value obtained after the average pooling layer; This indicates that the output features of the multi-scale low-frequency branch will be... The maximum pooling value obtained after the maximum pooling layer; 2) Features containing two channels Input convolutional layer, batch normalized (BN) layer, ... Activation function layer, convolutional layer, sigmoid activation function layer; sigmoid activation function layer generates attention weights. ; 3) Based on attention weights and characteristics Obtain features ; indicates as: Step 335: Output features of multi-scale high-frequency branches Input frequency-aware spatial attention module, output frequency-aware spatial attention module features The specific process is as follows: 1) Calculate the multi-scale low-frequency branch output characteristics The mean pooling value and max pooling value are used to generate features for two channels. : in, This indicates that the multi-scale high-frequency branch output features will be used. The average pooled value obtained after the average pooling layer; This indicates that the multi-scale high-frequency branch output features will be used. The maximum pooling value obtained after the maximum pooling layer; 2) Features containing two channels Input convolutional layer, batch normalized (BN) layer, ... Activation function layer, convolutional layer, sigmoid activation function layer; sigmoid activation function layer generates attention weights. ; 3) Based on attention weights and characteristics Obtain features ; indicates as: Step 336: Extract multi-scale original branch features As a feature .

8. The cross-view geolocation method based on multi-scale frequency-aware attention fusion according to claim 7, characterized in that: In step three-seven, the features ,feature ,feature Input to the multi-classifier module, and the multi-classifier module outputs the classification results. , , ; The specific process is as follows: The multi-classifier module includes a feature extraction block and a classification block; 1) Features Input feature extraction block, output feature ;feature Input a classification block, and the classification block outputs the classification result. The specific process is as follows: 11) Features Input feature extraction block, output feature The specific process is as follows: feature The input layers are sequentially a Linear layer, a Batch Normalization (BN) layer, and a Dropout layer. The Dropout layer outputs the features. ; 12) Features Input a classification block, and the classification block outputs the classification result. ; The specific process is as follows: feature The input is sequentially processed by a linear layer and a softmax activation function layer. The softmax activation function layer outputs the classification result. ; 2) Features Input feature extraction block, output feature ;feature Input a classification block, and the classification block outputs the classification result. The specific process is as follows: 21) Features Input feature extraction block, output feature ; The specific process is as follows: feature The input layers are sequentially a Linear layer, a Batch Normalization (BN) layer, and a Dropout layer. The Dropout layer outputs the features. ; 22) Features Input a classification block, and the classification block outputs the classification result. The specific process is as follows: feature The input is sequentially processed by a linear layer and a softmax activation function layer. The softmax activation function layer outputs the classification result. ; 3) Features Input feature extraction block, output feature ;feature Input a classification block, and the classification block outputs the classification result. ; The specific process is as follows: 31) Features Input feature extraction block, output feature ; The specific process is as follows: feature The input layers are sequentially a Linear layer, a Batch Normalization (BN) layer, and a Dropout layer. The Dropout layer outputs the features. ; 32) Features Input a classification block, and the classification block outputs the classification result. The specific process is as follows: feature The input is sequentially processed by a linear layer and a softmax activation function layer. The softmax activation function layer outputs the classification result. .

9. The cross-view geolocation method based on multi-scale frequency-sensing attention fusion according to claim 8, characterized in that: The loss function acquisition process in step 39 is as follows: loss function in, The loss function represents the cross-view, multi-scale frequency-aware attention fusion network model. Represents the cross-entropy loss function; Represents the cross-domain triplet loss function; in, Features representing images from the drone's perspective Features corresponding to satellite view images The cross-domain triplet loss function; Features representing images from the drone's perspective Features corresponding to satellite view images The cross-domain triplet loss function; Features representing images from the drone's perspective Features corresponding to satellite view images The cross-domain triplet loss function; in, This represents a sample in the image from the drone's perspective, serving as an anchor point; Indicates selecting from the satellite view image domain and Samples of the same category within the same scene are considered positive samples. Indicates selecting from the satellite view image domain and Samples belonging to different categories in different scenarios are used as negative samples; express The process sequentially passes through the visual base model EVA02, the multi-scale attention fusion module, and the average pooling layer, with the average pooling layer outputting features. express The process sequentially passes through the visual base model EVA02, the multi-scale attention fusion module, and the average pooling layer, with the average pooling layer outputting features. express The process sequentially passes through the visual base model EVA02, the multi-scale attention fusion module, and the average pooling layer, with the average pooling layer outputting features. This represents a function that calculates the distance between samples in the feature space; Indicates interval, ; in, This represents the cross-entropy loss function corresponding to the image from the drone's perspective. This represents the cross-entropy loss function corresponding to satellite view images; in, This represents the classification result output by the multi-classifier module. The cross-entropy loss function is calculated from the truth value; Indicates the true label; in, This represents the classification result output by the multi-classifier module. The cross-entropy loss function is calculated from the truth value; in, This represents the classification result output by the multi-classifier module. The cross-entropy loss function is calculated from the truth value; in, This represents the classification result output by the multi-classifier module. The cross-entropy loss function is calculated from the truth value; in, This represents the classification result output by the multi-classifier module. The cross-entropy loss function is calculated from the truth value; in, This represents the classification result output by the multi-classifier module. The cross-entropy loss function calculated from the truth value.

10. The cross-view geolocation method based on multi-scale frequency-aware attention fusion according to claim 9, characterized in that: In step four, the unlabeled UAV view image and the labeled satellite view image are input into the trained cross-view multi-scale frequency perception attention fusion network model. The multi-classifier module in the trained cross-view multi-scale frequency perception attention fusion network model outputs the feature vectors of the UAV view image and the satellite view image. The specific process is as follows: Step 4: Input the unlabeled drone view image to be tested into the visual base model EVA02 in the trained cross-view multi-scale frequency perception attention fusion network model. The visual base model EVA02 outputs features and inputs them into the multi-scale attention fusion module. The multi-scale attention fusion module outputs 3 features and inputs them into the average pooling layer. The average pooling layer outputs 3 features and inputs them into the feature extraction block in the multi-classifier module. The feature extraction block outputs 3 feature vectors. The 3 feature vectors output by the feature extraction block are concatenated to obtain 1 feature vector corresponding to the unlabeled drone view image to be tested. Step 4.2: Input the labeled satellite view image into the visual base model EVA02 in the trained cross-view multi-scale frequency perception attention fusion network model. The visual base model EVA02 outputs features and inputs them into the multi-scale attention fusion module. The multi-scale attention fusion module outputs three features and inputs them into the average pooling layer. The average pooling layer outputs three features and inputs them into the feature extraction block in the multi-classifier module. The feature extraction block outputs three feature vectors. The three feature vectors output by the feature extraction block are concatenated to obtain one feature vector corresponding to the labeled satellite view image.

Citation Information

Patent Citations

  • Cross-view geographic positioning method based on unit dot product attention mechanism

    CN118261970A

  • Cross-view geographic positioning method based on spatial frequency domain attention model

    CN119169466A

Cited By

  • Unmanned aerial vehicle position heading regression method and device for cross-view scene

    CN121297867A

  • Unmanned aerial vehicle assisted ground and satellite cross-view image geographic positioning method and system

    CN122115579A

  • Multi-modal geographic positioning method and system based on three-dimensional condition prompt learning

    CN122244168A