Unmanned aerial vehicle visual geolocation method based on transformer and view distribution alignment
By building a location classifier and view discriminator based on Transformer and view distribution alignment, the problems of view angle and distribution differences in UAV visual geo-positioning are solved, achieving higher positioning accuracy and feature aggregation capabilities.
Patent Information
- Application Number
- CN202511093424.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-08-06
AI Technical Summary
Existing UAV visual geolocation methods have insufficient image matching accuracy when facing changes in viewing angle, altitude and season, and ignore the distribution differences between UAV and satellite viewing angles.
A method based on Transformer and view distribution alignment is adopted to construct a position classifier and view discriminator through region segmentation and feature aggregation optimization. The multi-layer Transformer Encoder is used to enhance the feature aggregation capability, and the position classification loss and adversarial loss are used to train the model.
It significantly improves the accuracy of UAV visual geolocation, narrows the domain difference between UAV and satellite views, enhances feature representation capabilities, and achieves more accurate image matching.
Smart Images

Figure CN120599507B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a UAV visual geolocation method based on Transformer and view distribution alignment. Background Art
[0002] With the rapid development of microelectronics, intelligence, digital communications, sensing, and virtual reality technologies, drone technology has rapidly become popular and advanced. Visual positioning technology, as a key component of drone technology, has also made significant progress. Vision-based positioning methods rely on inexpensive, lightweight, and compact onboard cameras to achieve high-precision positioning, and have become a current research focus both domestically and internationally. Drones are widely used in agriculture, logistics, surveying and mapping, environmental monitoring, public safety, infrastructure maintenance, film and television, tourism, and the military, where precise positioning is crucial. Traditional GPS positioning methods may be subject to signal interference in special environments (such as inside buildings or near large structures), while visual positioning can provide more accurate and reliable information. Therefore, vision-based drone geolocation technology has broad development prospects and important application value.
[0003] Drone visual geolocation essentially determines geographic location through cross-view image retrieval between drone-viewed images and geotagged satellite-viewed images. It has applications in a variety of fields, including precision agriculture, rescue, and environmental monitoring. Using a drone-viewed image as a query, the retrieval system can find the most relevant satellite-viewed candidate images, thereby determining the geographic location of the target from the drone's view. Early drone visual geolocation methods employed two-branch CNN models for cross-view matching between drone-viewed and satellite-viewed images. These models were optimized within a location classification framework to learn view-invariant but location-dependent features. However, these learned features focus on the entire image, relying on polar coordinate transformations and limited receptive fields of convolutional layers, neglecting fine-grained details crucial for distinguishing images from different locations. Furthermore, due to variations in viewpoint, altitude, and season, drone and satellite image pairs exhibit significant differences in appearance, posing significant challenges for accurate image matching. Previous studies mapped drone and satellite images into a shared feature space and employed a classification framework to learn location-dependent features, neglecting the overall distributional shifts between drone and satellite views. Summary of the Invention
[0004] In view of the shortcomings of the existing technology, the purpose of the present invention is to propose a UAV visual geolocation method based on Transformer and view distribution alignment. The technical problem to be solved is how to improve the accuracy of UAV visual geolocation.
[0005] In order to achieve the above technical effects, the technical solution adopted by the present invention is:
[0006] A UAV visual geolocalization method based on Transformer and view distribution alignment includes the following steps:
[0007] S1, obtains drone images and satellite images, the satellite images include real location tags, and performs feature extraction on the drone images and satellite images to obtain drone image features and satellite image features;
[0008] S2, performing regional segmentation on the UAV image features and the satellite image features to obtain multiple sub-regions of UAV image features and multiple sub-regions of satellite image features. The real location label is correspondingly segmented into the real location labels of multiple sub-regions;
[0009] S3, respectively, aggregate and optimize the drone image features of all sub-regions and the satellite image features of all sub-regions to obtain drone region map features and satellite region map features; the drone region map features include multiple drone sub-region feature vectors, and the satellite region map features include multiple satellite sub-region feature vectors;
[0010] S4: Build a UAV visual geolocation model. The UAV visual geolocation model includes a location classifier and a view discriminator. The location classifier is used to perform branch processing on the features of the UAV area map and the satellite area map, with each sub-area corresponding to a branch, to obtain the location probability distribution of multiple sub-areas. The view discriminator is used to predict the view category.
[0011] S5, input the pre-processed drone area map features and satellite area map features obtained in S3 and the real location labels of all sub-areas obtained in S2 into the drone visual geolocation model, and train the drone visual geolocation model with the location classification loss function (1) and the adversarial loss function to obtain a trained location classifier;
[0012] (1)
[0013] in, M Indicates the number of branches processed by the branch, Indicates the k The prediction probability of each branch for the UAV segmentation map feature / satellite segmentation map feature, Indicates the k The true location label of the sub-region corresponding to each branch;
[0014] S6, performing S1-S3 on the to-be-positioned unmanned aerial vehicle image to obtain to-be-positioned unmanned aerial vehicle region map features, the to-be-positioned unmanned aerial vehicle region map features including a plurality of to-be-positioned unmanned aerial vehicle sub-region feature vectors, and inputting the to-be-positioned unmanned aerial vehicle region map features into the trained position classifier obtained in S5 to obtain a position probability distribution of each sub-region of the to-be-positioned unmanned aerial vehicle image;
[0015] S7, for the satellite image obtained in S1, extracting a candidate positioning region satellite image with a size consistent with the to-be-positioned unmanned aerial vehicle image from the satellite image according to the maximum value of the position probability distribution of the to-be-positioned sub-region obtained in S6, N wherein the number of candidate positioning region satellite images is an integer greater than or equal to 1; N wherein the number of candidate positioning region satellite images is an integer greater than or equal to 1;
[0016] S8, performing cross matching between the to-be-positioned unmanned aerial vehicle image and all candidate positioning region satellite images to obtain a similarity score of the to-be-positioned unmanned aerial vehicle image and each candidate positioning region satellite image, and selecting a candidate positioning region satellite image with the largest similarity score to obtain a positioning point of the to-be-positioned unmanned aerial vehicle image.
[0017] Preferably, the position classifier includes a plurality of parallel branch processing paths, each branch processing path including a cross-correlation layer, a Dropout layer and a fully connected layer connected in series, and a Softmax function is used to predict the position probability distribution of each sub-region after the fully connected layer.
[0018] Preferably, the view discriminator includes a convolution layer, an average pooling layer and a fully connected layer connected in series; the convolution layer is used to optimize the features of the unmanned aerial vehicle region map features and the satellite region map features, and the optimized unmanned aerial vehicle region map features and the optimized satellite region map features are input into the global average pooling and the fully connected layer for view category prediction.
[0019] Preferably, M =4.
[0020] A computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the unmanned aerial vehicle visual geographic positioning method based on the Transformer and view distribution alignment disclosed in the application.
[0021] A computer program product includes a computer program / instruction, and the computer program / instruction is executed by a processor to implement the unmanned aerial vehicle visual geographic positioning method based on the Transformer and view distribution alignment disclosed in the application.
[0022] Due to the adoption of the above technical solutions, the following beneficial effects are achieved:
[0023] (1) The present invention proposes a UAV visual geolocation method based on Transformer and view distribution alignment, which aligns the distribution of UAV view and satellite view in a common feature space, effectively narrowing the gap between the two and successfully solving the domain migration problem, especially the conversion problem between UAV view images and satellite view images.
[0024] (2) The present invention proposes a UAV visual geolocation method based on Transformer and view distribution alignment, which adopts a progressive adversarial learning strategy to train a location classifier and a view discriminator. It not only achieves view distribution alignment but also performs location classification simultaneously, thereby significantly reducing the domain difference between UAV views and satellite views.
[0025] (3) The present invention proposes a UAV visual geo-localization method based on Transformer and view distribution alignment. Based on the powerful feature aggregation and refinement capabilities of the multi-layer Transformer Encoder, it fully utilizes its advantages in capturing long-range dependencies and feature relationships, and further enhances the feature representation capability. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 This is a partial example of the University-1652 dataset in the embodiment (where (a) represents a set of drone views, and (b) represents a set of satellite views);
[0027] Figure 2 Visualization of the UAV visual geolocation results in the embodiment;
[0028] The present invention will be described in detail below with reference to the accompanying drawings and specific implementation methods. DETAILED DESCRIPTION
[0029] The present invention will be described in detail below with reference to the accompanying drawings and embodiments to facilitate a better understanding of the present invention by those skilled in the art. It should be noted that in the following description, detailed descriptions of known functions and designs will be omitted when they might obscure the main aspects of the present invention.
[0030] Other structures and functions of the method of the present invention are known to those skilled in the art and will not be described in detail to reduce redundancy.
[0031] Example
[0032] This embodiment discloses a method for visual geolocation of a drone based on Transformer and view distribution alignment, including the following steps:
[0033] S1, obtains drone images and satellite images, the satellite images include real location tags, and performs feature extraction on the drone images and satellite images to obtain drone image features and satellite image features;
[0034] This embodiment uses ResNet-50 for feature extraction. ResNet50 is a shared backbone network.
[0035] S2, perform regional segmentation on the UAV image features and satellite image features respectively, and obtain the UAV image features of 4 sub-regions and the satellite image features of 4 sub-regions. The real location label is correspondingly divided into the real location labels of the 4 sub-regions;
[0036] S3, respectively, aggregate and optimize the drone image features of all sub-regions and the satellite image features of all sub-regions to obtain drone region map features and satellite region map features; the drone region map features include multiple drone sub-region feature vectors, and the satellite region map features include multiple satellite sub-region feature vectors;
[0037] This embodiment performs aggregation optimization through a multi-layer Transformer Encoder.
[0038] S4, building a UAV visual geo-localization model, which includes a location classifier and a view discriminator;
[0039] The location classifier is used to perform branch processing on the features of the drone area map and the satellite area map, with each sub-area corresponding to a branch, to obtain the location probability distribution of multiple sub-areas; the view discriminator is used to predict the view category;
[0040] S5, input the pre-processed drone area map features and satellite area map features obtained in S3 and the real location labels of all sub-areas obtained in S2 into the drone visual geolocation model, and train the drone visual geolocation model with the location classification loss function (1) and the adversarial loss function to obtain a trained location classifier;
[0041] (1)
[0042] in, M Indicates the number of branches processed by the branch, M =4; Indicates the k The prediction probability of each branch for the UAV segmentation map feature / satellite segmentation map feature, Indicates the k The true location label of the sub-region corresponding to each branch;
[0043] S6, executing S1-S3 on the drone image to be located to obtain a map feature of the drone region to be located, wherein the map feature of the drone region to be located includes feature vectors of multiple sub-regions of the drone to be located, and inputting the map feature of the drone region to be located into the trained position classifier obtained in S5 to obtain a position probability distribution of multiple sub-regions of the drone image to be located;
[0044] S7, for the satellite image obtained in S1, according to the maximum value of the position probability distribution of the sub-area to be located obtained in S6, extract the area in the satellite image that is the same size as the drone image to be located N Satellite images of candidate positioning areas, N is an integer greater than or equal to 1;
[0045] S8, cross-matching the drone image to be located with all candidate positioning area satellite images to obtain similarity scores between the drone image to be located and the satellite images of each candidate positioning area, selecting the candidate positioning area satellite image with the largest similarity score, and obtaining the positioning point of the drone image to be located.
[0046] The position classifier of this embodiment includes multiple parallel branch processing paths, each branch processing path includes a cross-correlation layer, a Dropout layer, and a fully connected layer connected in series.
[0047] The cross-correlation layer is used to fuse the correlation information of the drone area map features and the satellite area map features; the Dropout layer is used to reduce overfitting and improve the generalization ability of the model; the fully connected layer is used to map the correlation information after feature mapping;
[0048] After the fully connected layer, the Softmax function is used to predict the position probability distribution of each sub-region.
[0049] The view discriminator of this embodiment includes a convolutional layer, an average pooling layer, and a fully connected layer connected in series. The convolutional layer is used to optimize the features of the drone area map and the satellite area map. The optimized drone area map features and the optimized satellite area map features are then used through global average pooling and a fully connected layer to predict the view category, that is, to determine whether the view is a drone image or a satellite image.
[0050] This embodiment provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned drone visual geolocation method based on Transformer and view distribution alignment.
[0051] This embodiment provides a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, it implements the above-mentioned UAV visual geolocation method based on Transformer and view distribution alignment.
[0052] The following simulation further illustrates:
[0053] 1. Simulation conditions
[0054] To test the effectiveness of our present invention and verify the effectiveness of our Transformer-based and view distribution alignment-based UAV visual geolocalization method, we conducted the following experiments. We used the large-scale University-1652 dataset, specifically designed for UAV view target localization, to ensure the accuracy and reliability of our results. The experimental environment was configured using the efficient PyTorch framework, and the network was trained on a high-performance NVIDIA GeForce 2080Ti graphics card, ensuring a fully trained and optimized model, thereby verifying the effectiveness of our present invention.
[0055] 2. Simulation experiment
[0056] Figure 1 Some examples of visual matching and localization in the University-1652 dataset are shown. Table 1 compares the accuracy of different methods for cross-view image matching. The proposed method performs well in drone-satellite image matching. Figure 2 The visualization shows the results of drone visual geolocation, showing the first five candidate positioning area satellite images retrieved in the drone view target positioning task. The first and second rows show the matching results of a single query, and the actual matching satellite images are marked with bold borders; the third row uses multiple drone view images as queries. Figure 2 As shown in the figure, even when the drone and satellite image pairs have significant appearance differences, the present invention can still accurately retrieve the satellite view image (as shown in the first query). However, the second row also shows a failure case where irrelevant surroundings interfere with the image matching. However, when using multiple drone images as queries, a more accurate target representation is extracted and satellite images that match the ground truth are retrieved sequentially, verifying the effectiveness of the present invention.
[0057] Table 1 Some examples of visual matching positioning in the University-1652 dataset
[0058]
[0059] The above descriptions are merely examples of various embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any modifications or substitutions that can be readily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included within the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A UAV visual geolocation method based on Transformer and view distribution alignment, characterized by: The following steps are involved: S1, obtaining a drone image and a satellite image, wherein the satellite image includes a real location tag, and performing feature extraction on the drone image and the satellite image to obtain drone image features and satellite image features; S2, performing region segmentation on the drone image features and the satellite image features to obtain drone image features of multiple sub-regions and satellite image features of multiple sub-regions, and the real location label is correspondingly segmented into the real location labels of the multiple sub-regions; S3, aggregate and optimize the drone image features of all sub-regions and the satellite image features of all sub-regions respectively to obtain the drone area map features and the satellite area map features; The drone area map features include multiple drone sub-area feature vectors, and the satellite area map features include multiple satellite sub-area feature vectors; S4, constructing a UAV visual geolocation model, wherein the UAV visual geolocation model includes a location classifier and a view discriminator; The location classifier is used to perform branch processing on the features of the drone area map and the satellite area map, with each sub-area corresponding to a branch, to obtain the location probability distribution of multiple sub-areas; the view discriminator is used to predict the view category; S5, input the pre-processed drone area map features and satellite area map features obtained in S3 and the real location labels of all sub-areas obtained in S2 into the drone visual geolocation model, and train the drone visual geolocation model with the location classification loss function (1) and the adversarial loss function to obtain a trained location classifier; (1) in, M Indicates the number of branches processed by the branch, Indicates the k The prediction probability of each branch for the UAV segmentation map feature / satellite segmentation map feature, Indicates the k The true location label of the sub-region corresponding to each branch; S6, executing S1-S3 on the drone image to be located to obtain a map feature of the drone region to be located, wherein the map feature of the drone region to be located includes multiple feature vectors of the drone sub-regions to be located, and inputting the map feature of the drone region to be located into the trained position classifier obtained in S5 to obtain a position probability distribution of the multiple sub-regions of the drone image to be located; S7, for the satellite image obtained in S1, according to the maximum value of the position probability distribution of the sub-area to be located obtained in S6, extract the area in the satellite image that is the same size as the drone image to be located. N Satellite images of candidate positioning areas, N is an integer greater than or equal to 1; S8, cross-matching the drone image to be located with all candidate positioning area satellite images to obtain similarity scores between the drone image to be located and the satellite images of each candidate positioning area, selecting the candidate positioning area satellite image with the largest similarity score, and obtaining the positioning point of the drone image to be located.
2. The UAV visual geolocation method based on Transformer and view distribution alignment according to claim 1, characterized in that: The location classifier includes multiple parallel branch processing paths, each of which includes a cross-correlation layer, a Dropout layer, and a fully connected layer connected in series. After the fully connected layer, the location probability distribution of each sub-region is predicted by the Softmax function.
3. The UAV visual geolocation method based on Transformer and view distribution alignment according to claim 1, characterized in that: The view discriminator includes a convolutional layer, an average pooling layer, and a fully connected layer connected in series; The convolutional layer is used to optimize the features of the drone area map and the satellite area map. The optimized drone area map features and the optimized satellite area map features are used to predict the view category through global average pooling and a fully connected layer.
4. The UAV visual geolocation method based on Transformer and view distribution alignment according to any one of claims 1 to 3, characterized in that: M =4。 5. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the drone visual geolocation method based on Transformer and view distribution alignment as described in any one of claims 1-4.
6. A computer program product, characterized in that The present invention comprises a computer program / instruction, which, when executed by a processor, implements the UAV visual geolocation method based on Transformer and view distribution alignment as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Cross-view-angle geographic positioning method based on unmanned aerial vehicle-satellite
CN113361508A
Cross-view geographic positioning method based on optimal transmission theory
CN114926827A