A cross-view geolocalization method based on spatial-frequency attention model
By adopting the dual-branch network structure of the spatial frequency domain attention model in cross-view geolocation, extracting and combining multi-scale spatial structure features and spatial frequency domain cross-dimensional interaction features, the problem of insufficient interaction attention in the multi-scale spatial structure information and frequency domain texture information in the existing technology is solved, and higher matching accuracy and computing efficiency are achieved.
Patent Information
- Application Number
- CN202411314418.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-20
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-09-20
AI Technical Summary
Existing cross-view geolocation technology is difficult to effectively pay attention to the cross-dimensional interaction of multi-scale spatial structure information and frequency domain texture information, resulting in insufficient matching accuracy.
A dual-branch network structure based on the spatial frequency domain attention model is adopted, multi-scale spatial structure features are extracted through the spatial attention module, and cross-dimensional interaction features of the spatial frequency domain are captured through the frequency domain attention module, and classification and training are combined with the multi-classifier module.
Cross-dimensional interaction between spatial dimensions and frequency domain dimensions is realized, the accuracy and computing efficiency of image matching are improved, and the matching accuracy of cross-view geolocation is significantly improved.
Smart Images

Figure CN119169466B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of photogrammetry and remote sensing technology, and in particular relates to a cross-viewing angle geographic positioning method based on a spatial frequency domain attention model. Background Art
[0002] The cross-view geolocation task was originally used for ground-to-air image matching. By inputting a ground image to retrieve a satellite image, the geographic location of the ground image is obtained. However, since the perspectives of the ground image and the satellite image are almost orthogonal, matching is difficult. With the advancement of drone technology, cross-view geolocation research in recent years believes that increasing viewpoints can improve the accuracy of cross-view geolocation, which is used to carry out research on drone positioning and drone navigation. Drone positioning is to find the most similar satellite view image based on a given drone view image to locate the target. Drone navigation takes the satellite view image as input and returns the most relevant drone view image to guide the drone to navigate to the target location. However, due to the differences in drone and satellite imaging angles and altitudes, cross-view geolocation based on the two images still faces huge challenges.
[0003] The key to successful cross-view geo-localization is to obtain discriminative feature representations of images. In 2020, Zheng published a paper titled "University-1652: A multi-view multi-source benchmark for drone-based geo-localization" in the conference proceedings "Proceedings of the 28th ACM International Conference on Multimedia". In the study, the University-1652 dataset was constructed, including satellite view images, ground view images, and drone view images. They regarded all view images of the same location as one category, completed the geo-localization task in a classified manner, and optimized the model using instance loss. However, this method only focuses on global information and does not consider the impact of detailed information on cross-view geo-localization. In 2022, Wang et al. published the article "Each part matters: Local patterns facilitate cross-view geo-localization" in the journal "IEEE Transactions on Circuits and Systems for Video Technology", Volume 32, proposing a ring segmentation strategy to segment the feature image, so that the network focuses on the surrounding environment of the target building, thereby obtaining more detailed information, and achieved significant performance improvement on the University-1652 dataset. Recent studies have improved the accuracy of image matching by introducing attention mechanisms to establish global associations between local features. In 2023, Shen et al. published the article "MCCG: AConvNeXt-based multiple-classifier method for cross-view geo-localization" in the journal "IEEE Transactions on Circuits and Systems for Video Technology", proposing a new method for joint feature representation in the spatial domain and channel domain, which can effectively associate contextual information with global and local information in cross-view images.
[0004] However, in current cross-view geolocation research, less attention has been paid to the cross-dimensional interaction of multi-scale spatial structure information and frequency domain texture information. The appearance between satellite images and drone images is very different, but the spatial structures such as the shape of buildings and the layout of roads are usually consistent. These consistent spatial structure information helps to achieve accurate matching under different perspectives. At the same time, existing feature extraction methods usually use operations such as pooling and dilated convolution to reduce computational costs, but these operations easily lead to the loss of texture details in the image. Therefore, how to pay attention to multi-scale spatial structure information and frequency domain texture information at the same time is an important consideration to improve the accuracy of cross-view geolocation matching. Summary of the invention
[0005] In order to solve the above technical problems, the present invention proposes a cross-view geolocation method based on a spatial frequency domain attention model to solve the problems existing in the above-mentioned prior art.
[0006] To achieve the above object, the present invention provides a cross-viewpoint geolocation method based on a spatial frequency domain attention model, comprising:
[0007] Constructing a spatial frequency domain attention model, wherein the spatial frequency domain attention model adopts a dual-branch network structure, wherein the two branches of the dual-branch network structure respectively process the drone image and the satellite image, wherein one branch network structure in the dual-branch network structure comprises a backbone network, a spatial frequency domain attention module, and a classifier connected in sequence; wherein the backbone network is used to extract multi-scale semantic backbone features from the drone image or the satellite image, the spatial frequency domain attention module is used to process the multi-scale semantic backbone features, and extract spatial structural features and spatial frequency domain cross-dimensional interaction features, and the classifier is used to classify the spatial features and the spatial frequency domain features, and obtain the classification result, i.e., the label corresponding to the drone image or the satellite image;
[0008] Obtain training samples including drone images and satellite images;
[0009] According to the training samples, the spatial frequency domain attention model is trained by mixing the triplet loss function and the cross entropy loss function to obtain a trained model;
[0010] Obtain the image to be tested, identify the image to be tested through the trained model, obtain labels, match the labels, and obtain another perspective image of the same geographical location to achieve cross-perspective geographic positioning.
[0011] Optionally, the backbone network adopts a ConvNext backbone network, wherein the multi-scale semantic backbone features extracted by the backbone network include overall structural features and local area features.
[0012] Optionally, the spatial frequency domain attention module includes a spatial attention module and a frequency domain attention module connected in sequence;
[0013] The multi-scale semantic backbone features are evenly divided by the spatial attention module to obtain evenly distributed features, each evenly distributed feature is subjected to maximum pooling and average pooling to obtain a first feature, the first feature is processed by 1*1 convolution and Sigmoid function to obtain a second feature, the second feature is multiplied by the corresponding evenly distributed feature to obtain a third feature, the third feature is subjected to average pooling and softmax function processing to obtain a first coding feature; each evenly distributed feature is subjected to 3*3 convolution processing to obtain a fourth feature, the fourth feature is subjected to average pooling and softmax function processing to obtain a second coding feature; the third feature is multiplied by the second coding feature to obtain a first combination feature, the fourth feature is multiplied by the first coding feature to obtain a second combination feature, the first combination feature and the second combination feature are summed to obtain a combination feature, the combination feature is multiplied by the evenly distributed feature to obtain a spatial structure feature;
[0014] The spatial structural features are Fourier transformed through the frequency domain attention module to obtain frequency domain features, and the frequency domain features are weighted through the complex weight matrix. The weighted frequency domain features are inverse Fourier transformed to obtain spatial-frequency domain cross-dimensional interaction features.
[0015] Optionally, the process of training the spatial-frequency domain attention model includes:
[0016] Connecting a classifier after the backbone network, inputting the training samples into the spatial frequency domain attention model in which the classifier is added after the backbone network, and performing a pooling operation before inputting the classifier, obtaining the predicted probabilities of different categories of results output by different classifiers through the classifier connected after the salient feature extraction module and the classifier connected after the backbone network, and calculating the cross entropy loss value according to the predicted probabilities and the actual category probabilities through the cross entropy loss function;
[0017] The training samples are input into a spatial frequency domain attention model in which a classifier is added after the backbone network, and feature groups corresponding to the drone images and satellite images in the training samples are extracted, wherein the feature groups include: multi-scale semantic backbone features, spatial structure features, and spatial frequency domain cross-dimensional interaction features; a pooling operation is performed on the feature groups corresponding to the drone images and satellite images in the training samples to obtain a pooled feature group, and a triplet loss value is calculated according to the pooled feature group through a triplet loss function;
[0018] According to the cross entropy loss value and the triplet loss value, the final loss is calculated, and the spatial frequency domain attention model is optimized according to the final loss to realize the training of the spatial frequency domain attention model.
[0019] Optionally, the final loss is the sum of the cross entropy loss value and the triplet loss value.
[0020] Optionally, the cross entropy loss function is:
[0021]
[0022] in, represents the cross entropy loss value of the i-th class, p(x ir ) represents the probability that the rth training sample in the i-th category belongs to this category, q(x ir ) represents the probability that the model predicts that the rth training sample in the i-th category belongs to this category, r represents the training sample number, and R represents the total number of training samples.
[0023] Optionally, the triple loss function is:
[0024]
[0025] Among them, f(x) represents the embedding function of the model, which maps the input image x to the embedding space. is the anchor point sample in the i-th triplet, is the positive sample in the i-th triplet, is the negative sample in the i-th triplet, and α is the margin, which is used to control the minimum distance difference between positive and negative samples.
[0026] Compared with the prior art, the present invention has the following advantages and technical effects:
[0027] (1) The present invention realizes cross-dimensional interaction between spatial dimension and frequency domain dimension. The spatial attention module is good at extracting global spatial structural features, while the frequency domain attention module can capture detailed texture features. These two features are complementary in image representation. The former provides macroscopic structural information, while the latter supplements microscopic detail information.
[0028] (2) The feature extraction module in the present invention can significantly enhance important features, suppress noise, highlight key information, dynamically adjust weights, enhance cross-view matching capabilities, and improve computational efficiency, making the model more effective and robust when processing complex geolocation tasks.
[0029] (3) The present invention provides a new multi-classifier structure for multiple feature representations of cross-view geolocation tasks, realizing a comprehensive representation of features, making the model more suitable for cross-view geolocation tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0031] Figure 1 It is an overall framework diagram of a cross-viewpoint geolocation method based on a spatial frequency domain attention model according to an embodiment of the present invention;
[0032] Figure 2 A network structure diagram of a spatial frequency domain attention module according to an embodiment of the present invention;
[0033] Figure 3 The present invention is a flowchart of a cross-view geolocation method based on a spatial frequency domain attention model according to an embodiment of the present invention. DETAILED DESCRIPTION
[0034] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0035] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0036] The present invention discloses a cross-view geolocation method based on a spatial frequency domain attention model. A cross-view geolocation network architecture based on a spatial frequency domain attention model is constructed, and multi-scale semantic backbone features are extracted through the ConvNext backbone network; multi-scale spatial structure features and spatial frequency domain cross-dimensional interaction features are generated through the spatial frequency domain attention module. These feature groups are classified using a multi-classifier module, and the model is trained by a function of a mixed triplet loss and a cross entropy loss. UAV images or satellite images are input into the trained network, and corresponding labels are output. Through the labels, another view image of the same geographical location can be matched and obtained, thereby realizing cross-view geolocation. The present invention obtains rich discriminant information through cross-dimensional interactions of spatial dimensions and frequency domain dimensions, obtains multiple feature representations, realizes comprehensive representation of features, and improves the accuracy of cross-view matching.
[0037] The technical problem to be solved by the present invention is that the channel or spatial attention mechanism is significantly effective in generating more recognizable feature representations during feature extraction. However, modeling cross-channel relationships through channel dimensionality reduction may have side effects on extracting deep visual representations; frequency domain texture information can identify and extract detail information in images, but is ignored by cross-view geolocation tasks, and thus existing technologies cannot complete the subsequent cross-view geolocation tasks.
[0038] In order to solve the above technical problems, such as Figure 1 , 3 As shown, the present invention provides a cross-viewpoint geolocation method based on a spatial frequency domain attention model, comprising the following steps:
[0039] Step 1: Construct a cross-view geolocation network architecture based on the spatial frequency domain attention model;
[0040] Step 2: Extract multi-scale semantic backbone features through the ConvNext backbone network. The multi-scale semantic backbone features include overall structural features and local area information.
[0041] Step 3: Generate multi-scale spatial structure features and spatial frequency domain cross-dimensional interaction features through the spatial frequency domain attention module;
[0042] Step 4: Use the multi-classifier module to classify these feature groups and train the model by mixing the triplet loss and cross entropy loss function;
[0043] Step 5: Input the drone image or satellite image into the trained network. You can match the labels and obtain another view image of the same geographic location, thereby achieving cross-view geolocation.
[0044] In order to further understand the content of the present invention, the present invention is described in detail in conjunction with the accompanying drawings.
[0045] Step 1: Build a cross-view geolocation network architecture based on the spatial frequency domain attention model
[0046] The images from the UAV perspective and the satellite perspective are obtained respectively, and a dual-branch network architecture for cross-perspective geolocation is established. The dual branches have the same structure and shared weights. Each branch includes a backbone network, a spatial frequency domain attention module, and a classifier.
[0047] Step 2: Extract multi-scale semantic backbone features through the ConvNext backbone network. These features include overall structural features and local area information.
[0048] The Convnext backbone network pre-trained on ImageNet21k is used to extract features. The Convnext backbone network processes the operation F ConvNextLayer Output Where B represents the batch size, N represents the number of feature map elements, C represents the number of channels, and the superscript v indicates that it is applicable to drone (d) and satellite (a) views. The details are as follows:
[0049] F v =F ConvNxetLayer (X v )
[0050] like Figure 2 As shown in the figure, step 3: Generate multi-scale spatial structure features and spatial frequency domain cross-dimensional interaction features through the spatial frequency domain attention module
[0051] Satellite images or drone images are input into the backbone network Convnext to obtain feature F, and then the spatial structure feature F is obtained through the spatial attention module. s , and then F s Input into the frequency domain attention module to obtain the spatial frequency domain cross-dimensional interaction feature F sf .
[0052] In order to preserve the information of each channel, the spatial attention module reconstructs the channel into the batch dimension and groups the channel dimension into multiple sub-feature groups so that the spatial semantic features are evenly distributed within each group. Specifically, in the 1x1 branch, the channels are encoded along two spatial directions through global average pooling and global maximum pooling; in the 3x3 branch, only one 3x3 convolution is stacked to capture multi-scale feature representation. The output features of the two parallel branches are further aggregated through cross-dimensional interactions to capture the detailed relationship between pixels.
[0053] Specifically, the multi-scale semantic backbone features are evenly divided by the spatial attention module, the multi-scale semantic backbone dimension is H*W*C, divided into G, and uniformly distributed features are obtained, each uniformly distributed feature size is H*W*C / G, and each uniformly distributed feature is subjected to maximum pooling and average pooling to obtain a first feature, the first feature includes features in different directions, and the dimensions are H*2*C / G and 2*W*C / G, the first feature is processed by 1*1 convolution and Sigmoid function to obtain a second feature, the second feature also includes features in different directions, which are H*1*C / G and 1*W*C / G, the second feature is multiplied by the corresponding uniformly distributed feature to obtain a third feature, the dimension is H*W*C / G, and the third feature is processed. Perform average pooling and softmax function processing to obtain a first coding feature with a dimension of 1*1*C / G; perform 3*3 convolution processing on each uniformly distributed feature to obtain a fourth feature with a dimension of H*W*C / G, perform average pooling and softmax function processing on the fourth feature to obtain a second coding feature with a dimension of 1*1*C / G; perform product processing on the third feature and the second coding feature to obtain a first combined feature, perform product processing on the fourth feature and the first coding feature to obtain a first combined feature, sum the first combined feature and the second combined feature to obtain a combined feature with a dimension of H*W*1, and multiply the combined feature with the uniformly distributed feature to obtain a spatial structure feature with a dimension of H*W*C.
[0054] After the frequency domain attention module converts the image to the Fourier domain, it uses learnable weight parameters to multiply the elements in the frequency domain to adjust and control the different frequency components of the signal, thereby achieving a gating effect, and finally restores the signal through the inverse Fourier transform. The specific method is as follows:
[0055] a) Perform Fourier transform on the input spatial structure feature Fs to convert the spatial domain features into frequency domain representation. The Fourier transform formula is:
[0056]
[0057] Among them, F freq (k) represents the kth frequency component in the frequency domain, F s (n) represents the spatial domain representation of the input feature, n represents the sample index, ranging from 0 to N-1, and N represents the total number of samples of the signal. is a complex number whose real and imaginary parts correspond to different amplitudes and phases at frequency k and time n.
[0058] b) Use a complex weight matrix to weight the frequency domain features. The weighting formula is:
[0059] G freq (k) = Ffreq (k)·W(k) ←
[0060] Among them, G freq (k) represents the weighted frequency domain feature, and W(k) represents the kth frequency component in the complex weight matrix.
[0061] c) The weighted frequency domain features are converted back to the spatial domain through inverse Fourier transform. The inverse Fourier transform formula is:
[0062]
[0063] Among them, F sf (n) represents the spatial domain feature after inverse Fourier transform. Through these steps, the frequency domain module can effectively extract and utilize frequency domain information, enhance the expressiveness and discrimination of features, and thus improve the performance of cross-view geolocation tasks.
[0064] Step 4: Use the multi-classifier module to classify these feature groups and train the model by mixing the triplet loss and cross entropy loss function;
[0065] Multiple classifier modules are applied to the pooled backbone features, spatial features, and spatial frequency domain features. Although the three classifiers have the same structure, their parameters are not shared in order to capture discriminative image features. Each classifier module consists of a fully connected layer (FC), a batch normalization layer (BN), a dropout layer, and a classification layer (Cls), where the classification layer is a fully connected layer. Through these classifier modules, the model can predict the geo-tag of the image based on the backbone features, spatial features, and spatial frequency domain features of the image.
[0066] The three sets of features sent to the classifier from drone images and satellite images are F, F after pooling. s and F sf , and calculate the cross entropy loss based on these features. The feature group for calculating the triplet loss of the drone image and the satellite image includes the following five groups of features after pooling: F, F s 、F sf The details are as follows:
[0067] a) Calculate the cross entropy loss of a single classifier:
[0068]
[0069] in, represents the cross entropy loss value of the i-th class, p(x i r) is the true distribution, which indicates the probability that the rth training sample in the i-th category belongs to this category, q(x ir) is the predicted probability of the model, which indicates the probability that the model predicts that the rth training sample in the i-th category belongs to this category. The sum of the cross entropy losses of the three classifiers is taken as the final cross entropy loss.
[0070] b) Calculate triplet loss
[0071]
[0072] Among them, f(x) represents the embedding function of the model, which maps the input image x to the embedding space, that is, it represents the relevant processing operations of the input image x to generate the above features. is the anchor point sample in the i-th triplet, is the positive sample in the i-th triplet, that is, the image from another perspective corresponding to the input image, is the negative sample in the i-th triplet, that is, the image from another perspective that does not correspond to the input image. α is the margin, which is used to control the minimum distance difference between positive and negative samples, and its value is 0.3.
[0073] c) Final loss
[0074] L=L triplet +L cl
[0075] The cross entropy loss and triplet loss are superimposed to optimize the model. Through these loss functions, the model can maximize the match between the predicted location and the true location, thereby minimizing the classification error and improving the performance of the cross-view geolocation task.
[0076] The model training details are as follows: During training, we resized the input image to 256×256 and performed image augmentation such as random padding, random cropping, and random flipping. For the optimizer, we adopted stochastic gradient descent (SGD) with momentum of 0.9, weight decay of 0.0005, and a mini-batch of 8. The learning rate of the backbone parameters was 0.003, and the learning rate of the remaining layers was 0.01. We trained for 200 epochs, and the learning rate was reduced by 0.1 after 80 and 120 epochs.
[0077] Step 5: Inputting drone images or satellite images into the trained network can match and obtain another perspective image of the same geographical location, thereby achieving cross-perspective geolocation.
[0078] The image of the drone to be tested is input into the network, the label of the image is output, and according to the label, the satellite image with the same label is retrieved in the database, and the satellite image of the corresponding geographical location is returned to achieve accurate positioning of the drone. The satellite image to be tested is input into the network, the label of the image is output, and according to the label, the drone image with the same label is retrieved in the database, and the corresponding drone image is returned to achieve drone navigation.
[0079] In order to verify the effectiveness of the algorithm in cross-view geographic tasks, the present invention is tested on the University-1652 dataset. According to the experimental results, the algorithm proposed in the present invention has Recall@1 and AP results of 87.65% and 89.54% respectively when retrieving satellite images from drone images, and 93.07% and 87.57% respectively when retrieving drone images from satellite images.
[0080] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A cross-view geolocation method based on a spatial frequency domain attention model, characterized in that: include: Constructing a spatial frequency domain attention model, wherein the spatial frequency domain attention model adopts a dual-branch network structure, wherein the two branches of the dual-branch network structure respectively process the drone image and the satellite image, wherein one branch network structure in the dual-branch network structure comprises a backbone network, a spatial frequency domain attention module, and a classifier connected in sequence; wherein the backbone network is used to extract multi-scale semantic backbone features from the drone image or the satellite image, the spatial frequency domain attention module is used to process the multi-scale semantic backbone features, and extract spatial structural features and spatial frequency domain cross-dimensional interaction features, and the classifier is used to classify the spatial features and the spatial frequency domain features, and obtain the classification result, i.e., the label corresponding to the drone image or the satellite image; Obtain training samples including drone images and satellite images; According to the training samples, the spatial frequency domain attention model is trained by mixing the triplet loss function and the cross entropy loss function to obtain a trained model; Obtain the image to be tested, identify the image to be tested through the trained model, obtain labels, match the labels, and obtain another perspective image of the same geographical location to achieve cross-perspective geographic positioning; The spatial frequency domain attention module includes a spatial attention module and a frequency domain attention module connected in sequence; The multi-scale semantic backbone features are evenly divided by the spatial attention module to obtain evenly distributed features, and each evenly distributed feature is subjected to maximum pooling and average pooling to obtain a first feature, the first feature includes features in different directions, the first feature is processed by 1*1 convolution and Sigmoid function to obtain a second feature, the second feature also includes features in different directions, the second feature is multiplied by the corresponding evenly distributed feature to obtain a third feature, the third feature is subjected to average pooling and softmax function processing to obtain a first coding feature; each evenly distributed feature is subjected to 3*3 convolution processing to obtain a fourth feature, the fourth feature is subjected to average pooling and softmax function processing to obtain a second coding feature; the third feature is multiplied by the second coding feature to obtain a first combination feature, the fourth feature is multiplied by the first coding feature to obtain a second combination feature, the first combination feature and the second combination feature are summed to obtain a combination feature, the combination feature is multiplied by the evenly distributed feature to obtain a spatial structure feature; The spatial structural features are Fourier transformed through the frequency domain attention module to obtain frequency domain features, and the frequency domain features are weighted through the complex weight matrix. The weighted frequency domain features are inverse Fourier transformed to obtain spatial-frequency domain cross-dimensional interaction features.
2. The method according to claim 1, characterized in that The backbone network adopts the ConvNext backbone network, wherein the multi-scale semantic backbone features extracted by the backbone network include overall structural features and local area features.
3. The method according to claim 1, characterized in that The process of training the spatial-frequency domain attention model includes: A classifier is connected after the backbone network, the training sample is input into the spatial frequency domain attention model in which the classifier is added after the backbone network, and a pooling operation is performed before input into the classifier, and the prediction probability of different category results output by different classifiers is obtained through the classifier connected after the spatial frequency domain attention module and the classifier connected after the backbone network, and a cross entropy loss value is calculated according to the prediction probability and the actual category probability through a cross entropy loss function; The training samples are input into a spatial frequency domain attention model in which a classifier is added after the backbone network, and feature groups corresponding to the drone images and satellite images in the training samples are extracted, wherein the feature groups include: multi-scale semantic backbone features, spatial structure features, and spatial frequency domain cross-dimensional interaction features; a pooling operation is performed on the feature groups corresponding to the drone images and satellite images in the training samples to obtain a pooled feature group, and a triplet loss value is calculated according to the pooled feature group through a triplet loss function; According to the cross entropy loss value and the triplet loss value, the final loss is calculated, and the spatial frequency domain attention model is optimized according to the final loss to realize the training of the spatial frequency domain attention model.
4. The method according to claim 3, characterized in that The final loss is the sum of the cross entropy loss value and the triple loss value.
5. The method according to claim 1, characterized in that The cross entropy loss function is: in, Indicates The cross entropy loss value of the class, Indicates No. The probability that a training sample belongs to this category is It means that the model predicts No. The probability that a training sample belongs to this category is represents the training sample label, Represents the total number of training samples.
6. The method according to claim 1, characterized in that The triplet loss function is: in, Represents the embedding function of the model, which takes the input image Mapped to the embedding space, It is Anchor samples in triplets, It is The positive samples in the triples are It is Negative samples in triplets, is the margin, which is used to control the minimum distance difference between positive and negative samples.
Citation Information
Patent Citations
Unmanned aerial vehicle-remote sensing image cross-view geographic positioning method with high positioning precision
CN118097406A
Three-path Fourier-time domain modulation network framework for high-precision intra-operative navigation
CN118570423A