Network for cross-view geographic positioning of unmanned aerial vehicle in severe weather
By adding a weather simulation module to the drone's viewing feature extraction, and using technologies such as ConvNeXt twin network and directional attention, the viewing angle difference and weather impact problems of drone's cross-view geolocation in bad weather are solved, achieving higher accuracy and robust positioning effects.
Patent Information
- Application Number
- CN202510148519.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-24
AI Technical Summary
In severe weather, drones cross-view geolocation face challenges such as perspective differences, scale changes and weather impacts, and existing methods are difficult to achieve accurate matching in complex environments.
By adding a weather simulation module in the drone perspective feature extraction process, the drone perspective in severe weather is simulated, and features are extracted using a twin network based on ConvNeXt, combining directional attention and local mode network to integrate environmental information and context information to generate a more comprehensive feature expression.
It significantly improves the positioning accuracy of the drone in bad weather, enhances the robustness of the network, and can achieve more accurate image matching and positioning in complex environments.
Smart Images

Figure CN120196781A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of geographic information, and specifically relates to a method for cross-view geolocation of unmanned aerial vehicles (UAVs) in adverse weather conditions. Background Art
[0002] Currently, UAVs mainly rely on GNSS (Global Navigation Satellite System) signals to calculate the current coordinates of the device and achieve positioning and navigation tasks. However, the positioning accuracy of GNSS is affected by factors such as multipath effects, atmospheric delays, and satellite geometric distributions, and it is unable to perform positioning tasks normally.
[0003] Cross-view geolocation of UAVs uses top-down images such as satellite images and aerial images with geographic tags as a reference image library, and matches the UAV images at unknown locations with them to obtain the accurate geographical location of the UAV images. This method can assist UAVs in positioning under GNSS denial conditions. However, there are still some challenges in cross-view geolocation of UAVs. The first is the perspective difference, which is huge between the UAV perspective and the satellite perspective. The differences in height, angle, and distance between these perspectives will cause significant changes in the image content and structure, making the matching process more complex. The second is the scale change, where objects in the image will appear at different scales under different perspectives. How to effectively handle these scale changes is an important challenge. The third is the weather impact, as Figure 2 shown. Weather conditions such as fog, rain, and snow will significantly reduce the visibility of the image, resulting in blurred images or loss of important details. Moreover, the lighting changes under different weather conditions, such as between sunny and cloudy days, and changes in sunlight intensity, will affect the color and contrast of the image, thereby affecting the quality and consistency of the image.
[0004] Existing cross-view geolocation methods for UAVs mainly include traditional methods and deep learning methods. Traditional methods mainly rely on manually designed features, such as SIFT (Scale-Invariant Feature Transform), SURF (Speeded-Up Robust Features), etc. These features usually perform poorly in cross-view and cross-scale situations and are difficult to effectively match aerial and ground images. Deep learning methods use neural networks, especially networks based on contrast learning and cross-view matching, which can extract more distinguishable features from aerial and ground images, thereby achieving more accurate matching in the case of inconsistent perspectives. However, most existing methods directly extract fine-grained features of UAV images or satellite images or only consider features at a single scale, which may lead to information loss or overemphasis on certain features, and do not consider the impact of weather on the UAV perspective in the actual environment, thus causing major obstacles in complex matching tasks.
[0005] To overcome the differences in perspectives between drones and satellites and the impacts brought by bad weather, the present invention fuses environmental information with context information in different directions to obtain better feature representations. During the process of extracting features from the drone perspective, a weather simulation module is added to simulate the drone perspective under real weather conditions, enhancing the robustness of the network so that the network can still complete the positioning task of the drone perspective under bad weather. Summary of the Invention
[0006] To overcome or reduce the deficiencies of the methods described in the above background art, the present invention designs and proposes a method for cross-perspective geolocation of drones under bad weather.
[0007] The present invention provides the following technical solutions to achieve: A method for cross-perspective geolocation of drones under bad weather. This method enables the network to adapt to images under bad weather by simulating the drone perspective under bad weather. The drone and satellite images are used to extract feature maps through a Siamese network based on ConvNeXt. The extracted feature maps pass through two branches, namely the local pattern network and the direction attention, to obtain feature vectors containing environmental information and context information in different directions, and these feature vectors are fused to obtain better feature representations. The multiple features generated after fusion are subjected to feature extraction by multiple classifiers and spliced into a feature vector. A database is constructed with the feature vectors extracted from satellite images, and the feature vectors of drone images are quickly matched with the feature vectors of satellite images in the database. The images stored in the database usually carry known geographical location information, and the location information of the matched image is the positioning result of the drone.
[0008] The method specifically includes the following steps: Step 1: Use the weather simulation module to enhance the drone images in the cross-perspective geolocation dataset of drones. This module defines various weather conditions using the imgaug library to simulate different environments: dark conditions are simulated by reducing brightness, while fog is simulated using clouds with specific intensity and transparency. Rain and snow are reproduced by different raindrop and snowflake sizes and speeds. Combinations such as fog and rain, fog and snow, and rain and snow blend these elements together. Lighting conditions are simulated by increasing brightness, and wind is created by applying motion blur.
[0009] Step 2: Use the Siamese network to share weights and extract features from the drone and satellite perspectives respectively. The backbone network uses the pre-trained model of ConvNeXt, which has been trained and fine-tuned on a large-scale dataset and has the ability to extract the most valuable features.
[0010] Step 3: To make the feature information more abundant, the present invention uses a method of fusing directional attention and local patterns. The directional attention module extracts spatial features in different directions and enhances the model's ability to capture spatial context information by fusing these features with weights. The main function of the local segmentation module is to capture detailed spatial structure information and multi-scale context information. Finally, the features extracted by the two modules are fused. The directional attention process is as follows:
[0011] Wherein, X As the feature map used as the input, angle Is the affine transformation function, Are multiple angles, Are the multiple feature maps output after affine transformation.
[0012] Perform global average pooling on Along the height and width directions respectively, and the following can be obtained:
[0013] Wherein, And Are The feature vectors generated by performing global average pooling along the height and width directions respectively.
[0014]
[0015] Wherein, Is Connected to The output feature vector, Conv Is the convolution, BN Is the normalization, h_swish Is the activation function.
[0016] Perform splitting on And perform convolution and activation to obtain:
[0017] Wherein, And Are The feature vectors output by splitting, And Are the attention weights in the height and width directions respectively, Is the sigmoid activation function.
[0018] Perform processing on With the attention weights and restore with the affine function to obtain:
[0019] Among them, is the feature map obtained after attention weight processing, is the feature map restored by the affine function, is element-wise multiplication.
[0020] The feature map output after weighted fusion of multiple feature maps is:
[0021] Among them, F is the feature map output after weighted fusion of multiple feature maps, is the learnable weight parameter.
[0022] Step 4: Use a multi-classifier for classification and feature extraction.
[0023] Step 5: Construct a satellite perspective feature vector database, and quickly match the feature vector of the UAV image with the database. The matching result is the positioning result of the UAV.
[0024] Compared with the prior art, the present invention has the following advantages: 1) The present invention proposes a network for cross-view geolocation of UAVs in bad weather. Based on ConVnext, the network can comprehensively capture multi-directional and multi-scale information in the image by fusing context information and environmental information in different directions, making the feature representation more comprehensive and rich, thereby improving the accuracy of image matching from complex cross-view images.
[0025] 2) The present invention adds a weather simulation module to the UAV perspective to simulate the impact of bad weather that may be encountered during the actual flight of the UAV on the UAV perspective. Bad weather may affect the image quality captured by the UAV camera, thereby affecting the accuracy of geolocation. By introducing weather simulation, the network can better handle the image quality degradation caused by weather changes, thereby significantly improving the positioning accuracy.
[0026] 3) The present invention proposes direction attention, a novel attention module that applies an attention mechanism in different directions to make the model pay more attention to different local regions in the feature map, and can more comprehensively understand the details and complex structures in the input data. By the process of weight addition, the features in different directions are fused together to form a comprehensive feature representation. In the face of different angles, illumination changes or partial occlusions, multi-directional attention can help the model more accurately extract key features, thereby enhancing the overall feature expression ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is the overall flowchart of cross-view geolocation of UAVs in bad weather Figure 2 Satellite and UAV perspective views Figure 3 Schematic diagram of the proposed network Figure 4 Schematic diagram of the proposed direction attention Figure 5 Schematic diagram of the multi-classifier Figure 6 UAV perspective positioning result diagram Specific implementation manners
[0028] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0029] The embodiments of the present invention conduct experimental evaluations on UAV cross-perspective geolocation. The datasets used in the experiments are University-1652 and University 160Kwx. University-1652 is a dataset containing UAV images and satellite images of 1652 university campuses, covering 72 countries around the world. University160k expands the current University-1652 dataset, adding 167,486 satellite view gallery interference items and introducing weather effects including fog, rain, snow, and datasets made of various weather combinations, aiming to simulate the perspective of UAVs in real weather. The overall flowchart of the present invention is as Figure 1 shown. First, through training, the network obtains the ability to extract feature maps of UAV and satellite perspectives. Then, the UAV perspective is used as the query image to match the images in the satellite perspective database. Finally, the satellite perspective image corresponding to the UAV perspective is output. The satellite perspective image contains geographical location information, thereby realizing the positioning of the UAV. The schematic diagram of the network structure of the present invention is as Figure 3 shown. The specific steps are as follows: Step 1: The input image first adds weather simulation to the UAV perspective, and then through a siamese network with shared weights, ConvNext is used as the backbone network to extract the features of the target image.
[0030] Step 2: After the backbone network finishes extracting features, the extracted feature maps enter two branches: direction attention and local pattern. Four feature vectors are created through local pattern circular segmentation and fused with the feature vectors obtained through direction attention. The mechanism schematic diagram of direction attention is as Figure 4 shown.
[0031] Step 3: Send the fused feature vector to the multi-classifier. The multi-classifier consists of four classifiers, and the weights of each classifier are not shared. Each classifier layer is responsible for generating a different feature vector. These feature vectors are processed by the feature extraction module and the classification module to map their dimensions to the number of categories. During the training process, the multi-classifier calculates the ternary loss for the feature vectors generated by the drone perspective and the satellite perspective, and calculates the cross entropy loss for the mapped category vector. In the process of querying the drone perspective, only the feature extraction module is used to process the feature vector, and a 4×512 feature representation is obtained. The structure of the multi-classifier is as follows Figure 5 shown.
[0032] In order to further analyze the effectiveness of the network proposed in the present invention, quantitative tests were carried out on the University-1652 and University160Kwx datasets, and the performance comparisons are shown in Tables 1 and 2.
[0033] As shown in Table 1, the University-1652 benchmark is used to evaluate the capabilities of the proposed UAV cross-view geo-network based on cutting-edge methods in the fields of UAV image localization tasks and UAV navigation tasks. In the UAV image localization task, the network achieves a Recall@1 accuracy of 92.95% and an average precision (AP) of 94.09%. All results are based on an image input size of 384×384 pixels. These results exceed the performance of the current state-of-the-art MCCG network, with the Recall@1 accuracy increasing from 89.64% to 92.93%, an increase of approximately 3%, and the AP increasing from 89.39% to 94.09%, an increase of approximately 4.5%, in the UAV image localization task. Overall, these results demonstrate the effectiveness of the network in UAV cross-view geo-location.
[0034]
[0035] As shown in Table 2, the University160k-WX test set is used to simulate the cross-view geolocation of drones in severe weather. The accuracy of drone positioning is greatly reduced when the LPN method is used on the University160k test set, and the Recall@1 accuracy is reduced from 75.93% to 7.98%. In order to verify the accuracy of the cross-view geolocation of drones in severe weather, the University-1652 training set is used for training in the training phase, and then University160k-WX is used as the test set. Compared with the LPN method, the proposed network improves Reall@1 from 7.94% to 82.25%, which proves the effectiveness and robustness of the present invention in the cross-view geolocation of drones in severe weather environments.
[0036] Partial result diagrams of the cross-view geolocation of the UAV are shown in the appendix Figure 6 As shown, the first column of images on the left is the view of the UAV in normal weather and adverse weather. The second to fifth columns on the right are the retrieved satellite view images, where the green boxes indicate correct retrieval and the red boxes indicate incorrect retrieval. It can be seen that the first retrieved satellite images are all correct, indicating the correct positioning of the UAV. Therefore, by combining the comparison table and the result diagrams, the network proposed in the present invention for cross-view geolocation of UAVs in adverse weather can not only improve the distinguishability of features but also provide comprehensive feature representations. Moreover, the positioning of the UAV can be achieved in both normal weather and adverse weather.
[0037] As described above, the above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed by the present invention, and these modifications or substitutions should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for cross-viewing geolocation of unmanned aerial vehicles in bad weather, characterized in that: The method specifically includes the following five steps.
2. A method for cross-viewing geolocation of a drone in bad weather according to claim 1, characterized in that: In step 1), the UAV imagery from the UAV cross-view geolocation dataset is enhanced using the weather simulation module. This module uses the imgaug library to define various weather conditions to simulate different environments: dark conditions are simulated by reducing the brightness, while fog uses clouds with specific intensity and transparency. Rain and snow are reproduced by different raindrop and snowflake sizes and speeds. Combinations such as fog and rain, fog and snow, and rain and snow blend these elements together. Lighting conditions are simulated by increasing the brightness, while wind is created by applying motion blur.
3. A method for cross-viewing geolocation of a drone in bad weather according to claim 1, characterized in that: In step 2), the twin network shared weights are used to extract the features of the drone and satellite perspectives respectively. The backbone network uses the ConvNeXt pre-trained model, which has been trained and fine-tuned on large-scale datasets and is capable of extracting the most valuable features.
4. A method for cross-viewing geolocation of a drone in bad weather according to claim 1, characterized in that: In step 3), in order to enrich the feature information, the present invention uses the method of directional attention and local pattern fusion. The directional attention module extracts spatial features in different directions and fuses these features by weighted fusion, thereby improving the model's ability to capture spatial context information. The main function of the local segmentation module is to capture detailed spatial structure information and multi-scale context information. Finally, the features extracted by the two modules are fused. Its directional attention process is expressed as follows:
5. Among them, X As the feature map taken as input, angle is the affine transformation function, For multiple angles, are multiple feature maps output after affine transformation.
6. Yes Performing global average pooling along the height and width directions respectively, we can get:
7. Among them, and yes Feature vectors generated by global average pooling along the height and width directions respectively.
8. Yes and After connecting, output and activate, we can get:
9. Among them, yes and The feature vector output after concatenation is, Conv is convolution, BN It is normalization. h_ swish is the activation function.
10. Yes Split, convolve and activate to get:
11. Among them, and yes The feature vector of the split output, and are the attention weights in height and width directions respectively, is the sigmoid activation function.
12. Yes After attention weight processing and affine function restoration, we can get:
13. Among them, yes The feature map obtained after attention weight processing, yes The feature map after restoration by affine function, is element-wise multiplication.
14. The output feature map after weighted fusion of multiple feature maps is:
15. Among them, F It is the feature map output after weighted fusion of multiple feature maps. is a learnable weight parameter.
16. A method for cross-viewing geolocation of a drone in bad weather according to claim 1, characterized in that: In step 4), multiple classifiers are used for classification and feature extraction.
17. A method for cross-viewing geolocation of a drone in bad weather according to claim 1, characterized in that: In step 5), a satellite view feature vector database is constructed, and the feature vector of the drone image is quickly matched with the database. The matching result is the positioning result of the drone.
Citation Information
Cited By
Geographic surveying and mapping data processing method and system based on big data
CN121211362A
Geographic mapping data processing method and system based on big data
CN121211362B