Cross-view-angle image geographic positioning method and system for unmanned aerial vehicle and satellite image
By using the Dinov2 model fine-tuned by Conv-LoRA in cross-view image geolocation and the spatial relationship-aware feature aggregator MSRA based on the Mamba module, the problems of scarcity of high-quality annotation samples and insufficient spatial layout feature capture are solved, and higher accuracy and faster cross-view image matching positioning is achieved.
Patent Information
- Application Number
- CN202510280349.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-13
Smart Images

Figure CN120147424A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cross-view image geolocation, and particularly to a cross-view image geolocation method and system for unmanned aerial vehicles and satellite images. Background Art
[0002] The cross-view image geolocation technology based on computer vision, as a second source of accurate position information in addition to the Global Navigation Satellite System (GNSS), can work independently in environments where GNSS signals are blocked or even absent, providing important technical support for fields such as remote sensing information utilization, unmanned driving, and augmented reality.
[0003] However, the huge visual appearance differences caused by the drastic perspective changes between unmanned aerial vehicle and satellite images pose significant challenges to this task. Therefore, how to extract highly discriminative feature representations and establish reliable feature associations between the two to eliminate such image differences is a key goal of the cross-view image geolocation task. In recent years, with the continuous development of space remote sensing technology and deep learning methods, relevant scholars have been able to use deep neural networks to extract robust features with viewpoint invariance from high-quality remote sensing images, thus achieving higher accuracy and better generalization performance. However, this advantage is based on the training of the model on a large number of unmanned aerial vehicle and satellite image pairs. Currently, although the acquisition of satellite images has become increasingly convenient and fast, and covers the global scope, the acquisition of unmanned aerial vehicle images is relatively much more difficult. Limited by the difficulty of obtaining flight permits, the complexity of operation technology, and the annotation challenges brought by flexible perspectives, high-quality labeled data is still relatively lacking. Therefore, the development and application of this technology have inevitably been restricted to a certain extent. But with the rise of Vision Foundation Models (VFMs) in the field of computer vision, it provides new possibilities for solving this problem. VFMs (such as SAM, FastSAM, Dino, Dinov2) can improve model performance with less or no labeled data by leveraging the knowledge obtained from large-scale datasets. Their strong generalization ability for different imaging conditions and visual objects promotes their application in real-world scenarios. However, due to the inductive biases learned in natural images, VFMs show limitations when applied to images in certain specific fields, such as medical images and remote sensing images. Therefore, such models need to be fine-tuned to adapt to new tasks or new datasets. Traditional fine-tuning methods usually require updating all or most of the model's parameters, which is not only computationally expensive but also prone to overfitting problems, especially when the dataset is small. In this context, it is particularly important to try a new fine-tuning strategy.
[0004] In addition, drone images usually capture the stereoscopic information of scenes at an oblique angle. Compared with satellite images that are completely perpendicular to the ground plane, the key visual features between the two do not fully match. Therefore, global feature descriptors are usually used as the final feature representation. The feature aggregator plays a crucial role. For example, NetVLAD adds a trainable network module on the basis of VLAD to learn the clustering centers and aggregation weights. Radenovic et al. proposed Generalized Mean Pooling (GeM) to aggregate features. This aggregation technique is an extension of the classical Average Pooling, and its concise and efficient characteristics have made it highly favored. Recently, MixVPR has presented the best results in this field by combining deep features with a multi-layer perceptron (MLP) layer. However, the above aggregation methods do not consider embedding spatial configuration features into the feature descriptor. Spatial configuration features are the most stable features when the image perspective changes. Especially when dealing with obliquely captured drone images, due to a large number of geometric and radiometric distortions, such images often lose most of their visual features and detailed information. In this case, spatial configuration features are particularly important for cross-perspective matching and positioning. They can provide relatively reliable and stable information against the background of image distortion and detail loss, and are the key elements for achieving high-precision matching and positioning. Summary of the Invention
[0005] The present invention aims to solve the problems that the scarcity of high-quality labeled samples limits the generalization ability of supervised learning models, and they often perform poorly when facing challenging practical application scenarios; secondly, the existing methods are insufficient in capturing spatial layout features, making it difficult to effectively compensate for the significant domain differences between cross-perspective image pairs. A cross-perspective image geolocation method and system for drone and satellite images are proposed, which can effectively address the large domain differences between drone images and satellite images, thereby improving the accuracy of cross-perspective image matching and positioning.
[0006] To achieve the above object, the technical solutions adopted are as follows:
[0007] The present invention provides a cross-perspective image geolocation method for drone and satellite images, including:
[0008] Using the Dinov2 large model fine-tuned by Conv-LoRA as a feature encoder to extract drone image feature vectors and satellite image feature vectors respectively;
[0009] Designing a Spatial Relationship Aware Feature Aggregator (MSRA) based on the Mamba module to aggregate image features and embed spatial configuration features into the global descriptor;
[0010] Adopting the InfoNCE loss function to train the model;
[0011] After using the trained model to extract the feature vectors of the image to be queried and each image in the reference image library, the cosine similarity is used to calculate the similarity scores between the image to be queried and all reference images, and the UAV-satellite image pairs with high matching degrees are selected to achieve the geolocation of the image to be queried.
[0012] According to the cross-view image geolocation method for UAVs and satellite images of the present invention, further, the Dinov2 large model includes a Transformer encoder, and the Transformer encoder is composed of multiple Vision Transformer encoding blocks, and each encoding block includes layer normalization, multi-head self-attention, and a multi-layer perceptron; the working process of the Dinov2 large model is as follows:
[0013] Given an input image F ∈ R c×h×w , where c represents the number of feature channels, and h and w represent the height and width of the image respectively; first, a convolutional layer is used to divide the input image into multiple image patches; then each image patch is converted into a one-dimensional vector through a linear mapping; then it is input into the Transformer encoder, and after being processed by the Transformer encoder, a feature matrix of size c × D is output, where c represents the number of feature channels and D represents the dimension of the feature vector; subsequently, this feature matrix is processed by layer normalization to be converted into a 1 × n feature vector, where n is the number of columns.
[0014] According to the cross-view image geolocation method for UAVs and satellite images of the present invention, further, the Conv-LoRA uses an encoder-decoder structure to impose a low-rank constraint on the weight update of the Dinov2 large model, specifically:
[0015] Given a pre-trained weight W 0 ∈ R d×k , first, a low-rank decomposition is constructed according to the size of this matrix to represent the parameter update ΔW ∈ R d×k , and ΔW is obtained by multiplying two low-rank matrix encoder A and decoder B; in addition, a mixture-of-experts model is also applied to process the encoder A; during the training process, the original parameter W 0 is frozen, and only the internal parameters of the Conv-LoRA are trained; finally, after the training is completed, the pre-trained weight W 0 and ΔW are added together as the fine-tuned model parameter W′, and the expression is: W′ = W 0 + ΔW; correspondingly, the forward propagation process changes from X′ = W 0 x to:
[0016]
[0017] where x represents the input matrix, denotes the mixture of experts model, and X′ denotes the output matrix after being weighted by the Vision Transformer encoding block.
[0018] According to the cross-view image geolocation method for drones and satellite images of the present invention, further, the working process of the Spatial Relationship Aware Feature Aggregator MSRA based on the Mamba module is as follows:
[0019] First, perform a max pooling operation on the input feature image F l ∈R c×h×w along the channel axis. The max pooling operation uses the max(·) function to aggregate the feature map, obtaining the aggregated feature matrix M ∈ R h×w and its corresponding position index matrix M idx ∈R h ×w ;
[0020] Next, use the first embedding layer to map the aggregated feature map M to K projection vectors E = [e 1 , e 2 , ···, e k . The dimension of each projection vector e i is h×w; based on the standard learnable position embedding E pe , an additional index-aware position embedding is introduced. The vector after the index-aware position embedding is expressed by the mathematical formula:
[0021]
[0022] where W LN represents a learnable linear transformation that maps M idx to K different subspaces, and HardTanh represents the activation function;
[0023] Then, input into the Mamba encoder to extract the high-level feature representation;
[0024] Finally, use the second embedding layer to project the output of the Mamba encoder into a set of K geometric layout descriptors E 2 = [e 1 , e 2 , ···, e k ; perform an element-wise multiplication operation on these geometric layout descriptors and E, and perform flattening and normalization processing to finally obtain the global descriptor f that contains both the spatial structure and retains the key feature information.
[0025] According to the cross-view image geolocation method for UAV and satellite images of the present invention, further, the mathematical expression of the global descriptor f is:
[0026] f = BN(flatten(E 2 ×E))
[0027] Wherein, BN represents batch normalization, and flatten represents flattening processing.
[0028] According to the cross-view image geolocation method for UAV and satellite images of the present invention, further, the mathematical expression of the InfoNCE loss function is:
[0029]
[0030] Wherein, q represents the feature encoding of the image to be queried, B represents the set of feature encodings of all reference images within a batch, and there is a positive sample r in this set that matches q + and B - 1 negative samples r that do not match q i , (·) represents using the dot product to calculate the similarity between the query image and the reference image, and the temperature coefficient τ is a hyperparameter.
[0031] Further, the present invention also provides a cross-view image geolocation system for UAV and satellite images, which is used to implement the cross-view image geolocation method for UAV and satellite images as described above. The system includes:
[0032] A feature extraction module, which is used to use the Dinov2 large model fine-tuned by Conv-LoRA as a feature encoder to extract the UAV image feature vector and the satellite image feature vector respectively;
[0033] A feature aggregation module, which is used to design a spatial relationship-aware feature aggregator MSRA based on the Mamba module to aggregate image features and embed the spatial configuration features into the global descriptor;
[0034] A training module, which is used to train the model using the InfoNCE loss function;
[0035] A similarity calculation module, which is used to use the trained model to extract the feature vectors of the image to be queried and each image in the reference image library, and then use the cosine similarity to calculate the similarity scores between the image to be queried and all reference images, and select the UAV-satellite image pairs with high matching degrees to achieve the geolocation of the image to be queried.
[0036] Adopting the above technical solutions, the beneficial effects obtained are:
[0037] The present invention proposes a cross-view image geolocation architecture - DINO-MSRA for unmanned aerial vehicle (UAV) and satellite imagery. First, the Dinov2 large model fine-tuned by Conv-LoRA is used as a feature encoder, aiming to enhance the model's feature extraction ability with fewer parameters. Second, a spatial relationship-aware feature aggregator (MSRA) based on the Mamba module is designed to aggregate image features, and by embedding spatial configuration features into the global descriptor, the large domain difference problem between UAV-satellite image pairs is thus compensated. Finally, the InfoNCE loss function is used to train the model. A large number of experiments conducted on the SUES-200 and University-1652 datasets show that DINO-MSRA is superior to the state-of-the-art methods in cross-view image matching, achieving higher accuracy and faster inference speed, demonstrating its robustness and practical application potential in challenging scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments of the present invention will be briefly introduced below. Among them, the accompanying drawings are only used to show some embodiments of the present invention, rather than limiting all embodiments of the present invention thereto.
[0039] Figure 1 is a framework diagram of the cross-view image geolocation method (DINO-MSRA) for unmanned aerial vehicle and satellite imagery according to an embodiment of the present invention;
[0040] Figure 2 is an architecture diagram of Dinov2-CL according to an embodiment of the present invention;
[0041] Figure 3 is a network architecture diagram of the Dinov2 large model according to an embodiment of the present invention;
[0042] Figure 4 is an architecture diagram of the spatial relationship-aware feature aggregator based on the Mamba module according to an embodiment of the present invention;
[0043] Figure 5 are examples of two datasets, University-1652 and SUES-200, according to an embodiment of the present invention. The satellite imagery is on the left side of the dashed line, and the UAV imagery corresponding to the satellite imagery is on the right side of the dashed line;
[0044] Figure 6 is a comparative experiment of each algorithm on the SUES-200 dataset according to an embodiment of the present invention;
[0045] Figure 7It is the visualization of the retrieval results of the embodiments of the present invention. Among them, the left side of the dotted line is the query image, and the right side of the dotted line is the top Top-5 images retrieved using the algorithm of the present invention. The yellow box indicates the correctly retrieved image, and the blue box indicates the incorrectly retrieved image. No-150, no-200, no-250, and no-300 respectively represent the UAV images taken at flight altitudes of 150m, 200m, 250m, and 300m. Detailed implementation manners
[0046] In the following, the exemplary solutions of the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the specific embodiments of the present invention. Unless otherwise defined, the technical terms or scientific terms used in the present invention should be of the ordinary meaning understood by those of ordinary skill in the art.
[0047] Cross-view image geolocation refers to a technology that infers the geographical location of a query image by matching it with reference images from different perspectives with accurate location information. This technology has been widely applied in practical tasks such as UAV navigation and target positioning. In recent years, cross-view image geolocation algorithms based on deep learning have made breakthrough progress, but generally still face two major challenges: the scarcity of high-quality labeled samples limits the generalization ability of supervised learning models, and they often perform poorly when facing challenging practical application scenarios; secondly, the existing methods are insufficient in capturing spatial layout features, making it difficult to effectively compensate for the significant domain differences between cross-view image pairs. To address the above problems, this embodiment discloses a cross-view image geolocation method for UAV and satellite images, as Figure 1 shown, the method specifically includes:
[0048] Step S101: Use Dinov2, which has been frozen and has its norms and head layers removed, as a feature encoder to extract visual features that are more robust and generalizable than traditional models. At the same time, in order to better adapt to the cross-view geolocation task, the Conv-LoRA parameter fine-tuning strategy is introduced to fine-tune the Dinov2 large model.
[0049] So far, in the cross-view image geolocation task, the most outstanding methods all rely on deep learning techniques, which can extract highly discriminative image features, thus effectively overcoming the challenges of huge domain differences. However, as a data-driven technology, deep learning requires a large amount of sample data to support model training and optimization. In practical application scenarios, the high cost of UAV image acquisition, combined with the annotation difficulties caused by the diversity of shooting angles and heights, has jointly led to the scarcity of high-quality sample pairs. This situation limits the training effect and performance improvement of supervised learning models in the cross-view image geolocation task. At the same time, the computer vision field is experiencing a new trend led by the "Segment Anything" model (SAM) - exploring Visual Foundation Models (VFMs). SAM has demonstrated excellent zero-shot generalization ability through training on millions of annotated images, which provides new possibilities for solving the problem of sample scarcity. Although the SAM model has strong generalization ability, its provided general segmentation ability may not meet the needs of fine-tuning and optimization for specific domains or scenarios, and it is difficult to directly adapt to various downstream tasks. Compared with SAM, the Dinov2 model can extract powerful vision features independent of the task and has a wider range of applications. Therefore, this solution uses Dinov2 with Vision Transformer as the backbone architecture to extract vision features that are more generalizable and versatile than traditional models.
[0050] However, since Dinov2 is trained on natural images, there are often certain limitations when applied to remote sensing images. To make the model adapt to the cross-view remote sensing image matching and geolocation task, it is necessary to fine-tune the model. Dinov2 contains 11 Vision Transformer (ViT) encoding blocks. If all or most of the parameters in these modules are retrained, it will consume a huge amount of computing power, which may not be feasible under the condition of limited hardware devices. To address this problem, this solution introduces an efficient parameter fine-tuning strategy - Conv-LoRA. The specific fine-tuning method is as Figure 2 shown in DINO-CL in. By applying Conv-LoRA to the 11 ViT encoding blocks in the Dinov2 feature encoder, the fine-tuning of the large model is achieved, and the original settings of other modules in Dinov2 remain unchanged during this process.
[0051] The network architecture of the Dinov2 large model is as Figure 3 shown. Given an input image F ∈ R c×h×w, where \(c\) represents the number of feature channels, and \(h\) and \(w\) represent the height and width of the image respectively. First, a convolutional layer with a kernel size of \(14\times14\) and a stride of \(14\) is used to divide the input image into multiple image patches, each with a size of \(14\times14\). Then, each image patch is converted into a one-dimensional vector (i.e., an embedded patch) through a linear mapping. Next, it is input into a Transformer encoder, which consists of multiple Vision Transformer (ViT) encoding blocks, and the number of encoding blocks depends on the specific scale of the model. Each encoding block includes three parts: layer normalization, multi-head self-attention, and a multi-layer perceptron (MLP). After being processed by the Transformer encoder, a feature matrix of size \(c\times D\) (number of channels \(\times\) feature vector dimension) is output. Subsequently, this feature matrix is processed by layer normalization and converted into a \(1\times n\) feature vector, which is a vector with one row and \(n\) columns.
[0052] Conv-LoRA combines a convolutional neural network (CNN) and low-rank adaptation techniques. LoRA is a parameter-efficient fine-tuning strategy used in large language models, which minimizes latency and memory usage by introducing trainable low-rank matrices. Conv-LoRA extends this strategy to convolutional neural networks and achieves efficient parameter fine-tuning by integrating ultra-lightweight convolutional parameters into LoRA. The architecture of Conv-LoRA is as Figure 2 shown in the right box. It uses an encoder-decoder structure to impose low-rank constraints on weight updates. Specifically, given a pre-trained weight \(W\) 0 \(\in\mathbb{R}\) d×k , first, a low-rank decomposition is constructed according to the size of this matrix to represent the parameter update \(\Delta W\in\mathbb{R}\) d×k , and \(\Delta W\) is obtained by multiplying two low-rank matrices, an encoder \(A\) and a decoder \(B\), where \(A\in\mathbb{R}\) d×r , \(B\in\mathbb{R}\) r×k , and \(r\ll\min(d,k)\). In addition, a mixture of experts (MoE) model is applied to process the encoder \(A\) so that the model can dynamically select different experts for processing, thereby improving the flexibility and performance of the model. During training, the original parameters \(W\) 0 are frozen, and only the internal parameters of Conv-LoRA are trained. In this way, the number of parameters to be trained can be greatly reduced. Finally, after training is completed, the pre-trained weight \(W\) 0 and \(\Delta W\) are added together as the fine-tuned model parameters \(W'\), and this process can be expressed by the mathematical formula:
[0053] \(W' = W\) 0 +\(\Delta W\)
[0054] Correspondingly, the forward propagation process changes from \(X' = W\) 0 \(x\) to:
[0055]
[0056] Among them, x represents the input matrix, represents the mixture-of-experts model, and X′ represents the output matrix after being weighted by the VisionTransformer encoding block.
[0057] However, since the dimension of the feature matrix output by the last ViT block in Dinov2 is c×D, while the feature aggregator MSRA designed in this scheme requires the dimension of the input feature matrix to be h×w×S. Therefore, a transform module is inserted between them to convert the feature dimension from the dimension output by Dinov2 to the input dimension required by MSRA. The specific conversion process is shown in the following formula:
[0058]
[0059] Among them, D and C represent the length of the feature vector and the number of feature channels of the output feature matrix, and h, w, and S represent the height, width, and number of feature channels of the input feature matrix respectively.
[0060] Step S102: Design a spatial relationship-aware feature aggregator (MSRA) based on the Mamba module. This aggregator can, while extracting the visual content features of the image, deeply mine the geometric spatial configuration information between local features, thereby making up for the huge domain difference problem between pairs of UAV-satellite images.
[0061] In order to generate high-quality image latent representations, previous cross-view image geolocation studies have emphasized the spatial configuration of visual features and low-level features. Because the spatial structure not only reflects the positions between visual features in the image, but also reflects the global context information between visual features, and these geometric information are considered to be stable information during the perspective transformation process. In this study, even though UAV images and satellite images have many visually similar features, the change in image content features caused by the perspective difference is still an interference factor that cannot be ignored. Therefore, in addition to using the basic feature extraction network to extract visual content features, the mining and application of the spatial configuration relationship between these features also become very important. To achieve this goal, this scheme designs a new type of spatial relationship-aware feature aggregator to bridge the perspective difference existing between cross-domain images through the mining of the spatial configuration relationship, and at the same time embed the target features into a discriminative global image descriptor for image matching.
[0062] To achieve this goal, the spatial relationship-aware feature aggregator is built on the Mamba module. Compared with the advantages of the traditional Transformer architecture in modeling long-sequence data, the Mamba module not only inherits this ability well but also greatly improves the training and inference capabilities on large-scale data, which obviously conforms to the original intention of designing this model. Combining this feature aggregator with the previous Dinov2 large model technology can achieve the purpose of improving training efficiency and making the model operation lightweight.
[0063] As Figure 4 shown, first, apply max pooling operation to the input feature image F l ∈R c×h×w along the channel axis. Specifically, use the max(·) function to aggregate the feature maps to obtain the aggregated feature matrix M∈R h×w and its corresponding position index matrix M idx ∈R h×w . This step simplifies the feature representation by aggregating information from different channels while retaining the key position information.
[0064] M, M idx = F l .max(1)
[0065] Next, use the first embedding layer to map the aggregated feature map M to K projection vectors E = [e 1 , e 2 , ···, e k . The dimension of each projection vector e i is h×w. To further enhance the spatial configuration relationship of the feature representation, this scheme introduces an additional index-aware position embedding on the basis of the standard learnable position embedding E pe . The vector after the index-aware position embedding can be expressed by the mathematical formula:
[0066]
[0067] where W LN represents a learnable linear transformation that can map M idx to K different subspaces, and HardTanh represents the activation function.
[0068] Then, input into the Mamba encoder, and use its lightweight features to efficiently capture the correlation between features, thereby further extracting high-level feature representations.
[0069] Finally, use the second embedding layer to project the output of the Mamba encoder into a set of K geometric layout descriptors E2 = [e 1 , e 2 , ···, e k . Element - wise multiply these geometric layout descriptors with E, and then perform flattening and normalization operations to finally obtain the global descriptor f that contains both spatial structure and preserves key feature information.
[0070] f = BN(flatten(E 2 × E))
[0071] Where, BN represents batch normalization, and flatten represents flattening processing.
[0072] Step S103: Use the InfoNCE loss function to train the model. The InfoNCE loss function further improves the generalization and overall performance of the model by effectively utilizing all negative samples within the training batch.
[0073] For cross - view image geolocation tasks based on metric learning, the triplet loss and its various variants (such as the soft - margin triplet loss, etc.) are usually used as the objective function for training. This type of loss function trains the model by constructing a triplet consisting of an anchor, a positive sample, and a negative sample, with the aim of reducing the distance between the anchor and the positive sample while increasing the distance between the anchor and the negative sample. Although this method has been proven effective, there are still some limitations. In particular, the selection of negative samples is random, which may lead to instability in the training process. In addition, since each training batch usually contains only one or a small number of negative samples, the generalization ability of the model may also be affected to a certain extent.
[0074] To address these problems, this solution considers using the InfoNCE loss function for model training. This loss function makes full use of all negative sample pairs within the batch, not only weakening the potential impact of the randomness of negative sample selection on the training process but also increasing the sensitivity of the model to the differences between different negative samples, thereby improving the scalability and generalization performance of the model. Its mathematical expression is as follows:
[0075]
[0076] Where, q represents the feature encoding of the query image, B represents the set of feature encodings of all reference images within a batch, and there is a positive sample r in this set that matches q + and B - 1 negative samples r that do not match q i , (·) represents calculating the similarity between the query image and the reference image using the dot product, and the temperature coefficient τ is a hyperparameter that can be set to be learnable or a static value.
[0077] In step S104, after using the trained model to extract the feature vectors of the image to be queried and each image in the reference image library, the cosine similarity is used to calculate the similarity scores between the image to be queried and all reference images, and the UAV-satellite image pairs with high matching degrees are selected to achieve the geolocation of the image to be queried.
[0078] Correspondingly, this embodiment also discloses a cross-view image geolocation system for UAVs and satellite images. The system includes:
[0079] A feature extraction module, which uses the Dinov2 large model fine-tuned by Conv-LoRA as a feature encoder to extract UAV image feature vectors and satellite image feature vectors respectively.
[0080] A feature aggregation module, which is used to design a spatial relationship-aware feature aggregator MSRA based on the Mamba module to aggregate image features and embed the spatial configuration features into the global descriptor.
[0081] A training module, which is used to train the model using the InfoNCE loss function.
[0082] A similarity calculation module, which is used to use the trained model to extract the feature vectors of the image to be queried and each image in the reference image library, and then use the cosine similarity to calculate the similarity scores between the image to be queried and all reference images, and select the UAV-satellite image pairs with high matching degrees to achieve the geolocation of the image to be queried.
[0083] To verify the effectiveness of this solution, further explanation will be given below in combination with experimental data.
[0084] (1) Datasets
[0085] In this experiment, relevant experiments were carried out on two datasets, University-1652 and SUES-200. Figure 5 Some examples of these two datasets are shown, and Table 1 shows the sizes and partitions of these two datasets.
[0086] The University-1652 dataset was collected and organized by Zheng et al., integrating data from three platforms: drones, satellites, and the ground. It is the first geolocation dataset containing drone-view images. This dataset can be widely applied to tasks such as drone-view target localization and drone navigation, and can also serve as an intermediate medium in the process of ground-satellite cross-view image matching. During the data collection process, the research team selected a total of 72 universities and sampled 1652 building scenes. For each building scene, the geographical location in Google Maps was first used for projection to obtain satellite view images. Then, a 3D model provided by Google Earth was used to simulate a real drone camera, and a spiral flight path was adopted to gradually approach the building for shooting, thus capturing multi-scale and multi-view drone-view images.
[0087] The SUES-200 dataset is a multi-height and multi-scene cross-view image benchmark dataset created by Shanghai University of Engineering Science. This dataset contains images taken by drones at four different heights (150m, 200m, 250m, 300m), as well as satellite view images corresponding to the same target scenes. These images were all collected from Shanghai University of Engineering Science and its surrounding areas, covering various scenes such as parks, schools, lakes, and public buildings. With its multi-height and multi-scene characteristics, SUES-200 brings higher practical value and challenges to cross-view image matching tasks.
[0088] Table 1 Details of the experimental datasets
[0089]
[0090] Note: There is no overlap between the buildings in the training set and the test set.
[0091] (2) Evaluation Metrics
[0092] The top-K recall accuracy (R@K) is selected as the model evaluation metric. Specifically, given a query image, if its true matching image is among the top K retrieved images, it is considered a successful match. Finally, the percentage of all successfully matched query images is R@K.
[0093] In addition, the average precision (AP) is also used as an evaluation metric. AP represents the area under the precision-recall (PR) curve, where the recall rate is the abscissa and the accuracy rate is the ordinate.
[0094] (3) Experimental Details
[0095] This algorithm is based on the PyTorch architecture and is trained and tested on an NVIDIA GeForce RTX 3090 graphics card. The training rounds are 20 rounds and the training optimizer is the Adam optimizer. The staged learning rate adjustment strategy is used, that is, the overall training process of the model is divided into two stages. In the first 10 rounds of training, the learning rate is set to e -2 ; The learning rate for subsequent training rounds is set to e -3 For the loss function, the infoNCE loss function was selected to train the model. In addition, for ease of processing, the size of the drone and satellite images input to the network was uniformly adjusted to 224 pixels × 224 pixels.
[0096] (4) Comparative experiment
[0097] In order to prove the advanced nature of the algorithm proposed in this invention, comparative experiments were conducted on the University-1652 and SUES-200 datasets with other similar algorithms. Among them, the University-1652 baseline model, the SUES-200 baseline model, and the LPN algorithm are relatively basic cross-view geolocation algorithms, DWDR is an algorithm based on the CNN architecture, FSRA is an algorithm based on the Transformer architecture, and Sample4Geo is an algorithm based on the ConvNeXt architecture. Comparison with algorithms of different architectures is to highlight the superiority of the network proposed in this invention in terms of computing speed and accuracy.
[0098] (4.1) Experimental results on the University-1652 dataset
[0099] The experimental results on the University-1652 dataset are shown in Table 2. Since most algorithms have reached saturation in terms of accuracy in the R@5, R@10 and R@1% indicators, this experiment selects the more discriminative R@1 and AP as evaluation indicators to evaluate the model of the present invention. It can be seen from the results that the algorithm of the present invention is fully ahead in these two indicators and has achieved the current best accuracy. When it comes to UAV positioning (UAV→satellite), the accuracy in R@1 and AP is improved by 2.49% and 2.11% respectively compared with the current best algorithm Sample4geo. When it comes to UAV navigation tasks (satellite→UAV), the accuracy in R@1 and AP is improved by 2.15% and 2.42% respectively compared with the current best algorithm Sample4Geo.
[0100] During the model training process, the total number of parameters represents both learnable parameters and some non-learnable parameters. Among them, the non-learnable parameters do not participate in the update during the training process. Therefore, they have a relatively small impact on the training speed. The number of learnable parameters refers to the parameters that are updated through the optimization algorithm during the training process. These parameters directly participate in the calculations of the forward and backward propagations of the model and directly determine the speed and time of model training. Therefore, in this experiment, an additional evaluation metric of the number of learnable parameters was set to evaluate the lightweight degree of the model. Among them, the LPN and SAIG-D models do not share weights during the training process, while Sample4Geo and the algorithm of the present invention share weights during the training process. The experimental results show that the number of learnable parameters involved in the training of the model of the present invention is much lower than that of other algorithms, and the number of parameters is reduced by 260.39M, 14.19M, and 71.59M compared with the LPN, SAIG-D, and Sample4Geo models respectively.
[0101] The above experimental results strongly prove that the algorithm proposed by the present invention not only demonstrates excellent performance but also has significant lightweight advantages, achieving the purpose of the designed algorithm of the present invention. Through analysis, the main reason for the model's advantage is that, compared with traditional supervised models, the present invention adopts the self-supervised model Dinov2 trained on a large-scale dataset. This model can learn the general visual features between satellite images and drone images without the need for learnable parameters, thereby achieving effective cross-view transfer and reducing the negative impact brought by view differences. In addition, the Conv-LoRA fine-tuning strategy can efficiently adjust the cross-view matching and positioning tasks by compressing the weight update into a low-rank matrix method, using only a small amount of memory and computing resources, further improving the model performance.
[0102] Table 2 Precision comparison of each algorithm on the University-1652 dataset
[0103]
[0104] (4.2) Experimental results on the SUES-200 dataset
[0105] The experimental results on the SUES-200 dataset are shown in Tables 3 and 4. The more discriminative R@1 and AP are also used as evaluation metrics to evaluate the model of the present invention. The results show that in the UAV positioning task (UAV → satellite), the algorithm of the present invention achieves Recall@1 accuracies of 97.2%, 98.75%, 99.38%, and 99.63% at four heights respectively; in the UAV navigation task (satellite → UAV), the algorithm of the present invention achieves Recall@1 accuracies of 98.75%, 99.08%, 99.38%, and 99.42% at four heights respectively, and its performance is better than the current best algorithm Sample4Geo. In addition, the results also show that as the height of the UAV increases, the Recall@1 and AP of the model are also correspondingly improved (in the UAV positioning task, when the shooting height of the UAV rises from 150m to 300m, the Recall@1 accuracy increases from 97.2% to 99.63%; in the UAV navigation task, as the shooting height of the UAV rises from 150m to 300m, the Recall@1 accuracy increases from 98.75% to 99.42%). The reason for this phenomenon is that in the range of 150 meters and 200 meters at low altitudes, the images captured by the UAV are more affected by the surrounding environment and the camera attitude, resulting in a large difference between the UAV images and the satellite images. Therefore, the matching accuracy is relatively low. Then, as the height increases, the influence of the surrounding environment and the camera field of view on the UAV decreases, and the images captured by the camera are more similar to the satellite images. This reduces the domain difference between the UAV images and the satellite images, thereby improving the Recall@1 and AP metrics of the model. The complete algorithm performance comparison chart is as Figure 6 shown.
[0106] Table 3 Comparison of accuracy results of each algorithm on the SUES-200 dataset (UAV → satellite)
[0107]
[0108] Note: The bold font is the best value in each column. UAV → satellite indicates the UAV positioning task, where the UAV image is the query image and the satellite image is the reference image.
[0109] Table 4 Comparison of accuracy results of each algorithm on the SUES-200 dataset (satellite → UAV)
[0110]
[0111] Note: The bold font is the best value in each column. Satellite → UAV indicates the UAV navigation task, where the satellite image is the query image and the UAV image is the reference image.
[0112] (5), Ablation experiment
[0113] A series of ablation experiments are carried out below. First, in Section (5.1), the effects of different hyperparameter configurations in the Dinov2 backbone network and the parameter fine-tuning strategy Conv-LoRA on the model performance are explored and quantified. Then, in Section (5.2), relevant experiments are designed to verify the effectiveness of the Mamba-based spatial relationship-aware feature aggregator proposed in the present invention.
[0114] (5.1) Hyperparameter Ablation Experiments
[0115] During the fine-tuning process using Conv-LoRA, the rank parameter is a crucial hyperparameter, which determines the dimension of the low-rank matrix introduced during the fine-tuning process. Generally speaking, the larger the rank, the more trainable parameters are introduced, which can enhance the model's adaptability to new data, but correspondingly also increases the consumption of computing and memory resources. On the contrary, the smaller the rank, the fewer trainable parameters are introduced. Although it reduces the computing and memory requirements, it may not be able to fully adapt to new data due to insufficient parameters, thus affecting the model performance. In order to find the optimal solution for the performance and efficiency of the Conv-LoRA fine-tuning strategy and the internal parameter rank setting in the cross-view matching and localization tasks, ViTb14 is used as the backbone network, and multiple control groups are set up for testing on the University-1652 and SUES-200 datasets. The specific experimental configurations are shown in Tables 5, 6, and 7 below.
[0116] The experimental results show that, when the original parameters in Dinov2 are frozen and no fine-tuning strategy is adopted {number of frozen layers = [0 - 11], LoRA = ×}, although Dino-MSRA consumes the least number of parameters on the two datasets, its test accuracy is also the lowest. However, after fine-tuning using the Conv-LoRA strategy {number of frozen layers = [0 - 11], Conv-LoRA = r}, the test accuracy of the model is significantly improved. It is worth noting that, compared with the traditional methods of updating all the parameters of the model {number of frozen layers = ×, Conv-LoRA = ×} or updating some parameters {number of frozen layers = [0 - 7], Conv-LoRA = ×} (it has been proven in SALAD that freezing the parameters of layers [0 - 07] in the Dinov2 model can achieve the best fine-tuning effect, so this configuration is also selected as one of the control groups in this experiment), the Conv-LoRA strategy requires fewer parameter numbers while achieving higher accuracy. This result highlights the key role played by the Conv-LoRA strategy in improving model accuracy and efficiency. As the value of r gradually increases (from r = 2 to r = 8, 16), the model accuracy continuously improves and finally reaches the peak. However, when r continues to increase, the test accuracy of the model shows a slow downward fluctuation on the dataset. Based on the above analysis, the present invention selects r = 16 as the optimal parameter configuration of the Conv-LoRA strategy.
[0117] Table 5 Ablation experiments of different hyperparameter configurations on the University-1652 dataset
[0118]
[0119] Note: The bold font is the optimal value for each column. The parameter Conv-LoRA indicates whether the Conv-LoRA fine-tuning strategy is used, × indicates not used, otherwise it indicates used. r represents the size of the set rank. The number of learnable parameters represents the number of parameters actually participating in learning during the model training process.
[0120] Table 6 Ablation experiments of different hyperparameter configurations on the SUES-200 dataset (drone → satellite)
[0121]
[0122] Note: The bold font is the optimal value for each column; the parameter Conv-LoRA indicates whether the Conv-LoRA fine-tuning strategy is used, × indicates not used, otherwise it indicates used, and r represents the size of the set rank; the number of learning parameters represents the number of parameters actually participating in training during the model training process.
[0123] Table 7 Ablation experiments of different hyperparameter configurations on the SUES-200 dataset (satellite → drone)
[0124]
[0125] Note: The bold font represents the optimal value for each column; the parameter Conv-LoRA indicates whether to use the Conv-LoRA fine-tuning strategy, where × means not used, otherwise it means used, and r represents the size of the set rank; the learning parameter quantity represents the actual parameter quantity involved in the model training process.
[0126] (5.2), Ablation Experiment on the Effectiveness of the MSRA Aggregation Module
[0127] Relevant experiments were designed to verify the effectiveness of the MSRA aggregation module proposed in the present invention. In the experiment, ViTb14 was selected as the backbone network, and the query image and the panoramic image were used as the network inputs. The optimal parameters determined in the experiment of section (5.1) were fixed. The three classic aggregation modules, namely GEM, NetVLAD, and MixVPR, and the MSRA aggregation module proposed in the present invention were respectively injected into the last layer of the backbone network to aggregate the global feature vectors, and training and testing were carried out on the University-1652 dataset. The comparison results are shown in Tables 8, 9, and 10.
[0128] The experimental results show that, compared with the classic aggregation algorithms, the aggregation strategy MSRA proposed in the present invention exhibits the most excellent performance when combined with Dinov2. This indicates that in the cross-view image matching and positioning task, the accuracy of the model can be effectively improved by embedding the spatial configuration relationship between features during the process of aggregating global feature vectors. The main reason is that due to the tilted shooting angle of the UAV images, geometric distortion, radiometric distortion, and an increase in side texture information are often inevitably generated. If solely relying on visual features for cross-view matching, this information unrelated to the matching will cause interference to the matching result to a great extent. And the geometric layout information is the most stable information during the perspective transformation and can provide effective gain for cross-view matching.
[0129] Table 8 Verification of the Effectiveness of the MSRA Aggregation Strategy on the Univeristy-1652 Dataset
[0130]
[0131] Note: The bold font represents the optimal value for each column
[0132] Table 9 Verification of the Effectiveness of the MSRA Aggregation Strategy on the SUES-200 Dataset (UAV → Satellite)
[0133]
[0134] Note: The bold font represents the optimal value for each column
[0135] Effectiveness Verification of the 10MSRA Aggregation Strategy on the SUES-200 Dataset (Satellite → UAV)
[0136]
[0137] Note: The bold font indicates the optimal value in each column
[0138] (6) Visualization of Test Retrieval Results
[0139] To visually demonstrate the retrieval and positioning efficiency of the algorithm of the present invention, representative image retrieval cases were selected from the two datasets University-1652 and SUES-200 for visualization. The retrieval cases include not only simple images of scenes with sparse buildings and simple layouts, but also complex scene images with dense buildings and severe vegetation coverage. The retrieval results are as Figure 7 shown. The query image is on the left side of the dotted line, and the Top-1 to Top-5 image sequences retrieved are on the right side of the dotted line. The images correctly retrieved are marked with yellow borders, and the images wrongly retrieved are marked with blue borders. The results in the figure show that for the two types of tasks (UAV positioning task, UAV navigation task), most query images can retrieve the correct image at Top1, and all query images can retrieve the correct image at Top5. This indicates that the algorithm of the present invention can adapt to the retrieval and positioning requirements in different scenarios and can provide accurate and robust retrieval and positioning capabilities. In addition, the retrieval results of the query images at four different heights in the SUES-200 dataset show that the algorithm of the present invention can still provide robust retrieval and positioning capabilities when the UAV flight height and viewing angle change continuously.
[0140] In summary, the present invention designs a novel DINO-MSRA network architecture for the problems of lack of high-quality samples and insufficient mining of spatial layout features in the UAV-satellite cross-view image matching task. The architecture mainly includes two key components: one is to fine-tune the Dinov2 backbone network using Conv-LoRA to enhance the feature extraction ability while only using a small number of learnable parameters; the other is to design a spatial relationship-aware feature aggregator based on Mamba to achieve efficient integration and utilization of image features. Through extensive experimental verification, the algorithm proposed by the present invention has achieved excellent performance exceeding the current optimal algorithm Sample4Geo in terms of R@1 and AP accuracy on the two public datasets University-1652 and SUES-200. In addition, while ensuring the accuracy, the number of parameters required by DINO-MSRA is significantly reduced compared to Sample4Geo.
[0141] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any technician familiar with the technical field of the present invention can still modify the technical solutions recorded in the foregoing embodiments within the technical scope disclosed by the present invention, or perform equivalent substitution on some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A cross-view image geolocation method for unmanned aerial vehicles and satellite images, characterized in that: Include: The Dinov2 large model fine-tuned by Conv-LoRA is used as the feature encoder to extract the feature vectors of drone images and satellite images respectively; A spatial relationship-aware feature aggregator MSRA based on the Mamba module is designed to aggregate image features and embed spatial configuration features into the global descriptor. The InfoNCE loss function is used to train the model; After using the trained model to extract the feature vectors of the query image and each image in the reference image library, the cosine similarity is used to calculate the similarity score between the query image and all reference images, and the UAV-satellite image pairs with high matching degree are selected to realize the geographic positioning of the query image.
2. The cross-view image geolocation method for drone and satellite images according to claim 1, characterized in that: The Dinov2 large model includes a Transformer encoder, which consists of multiple VisionTransformer encoding blocks, each of which includes layer normalization, multi-head self-attention and a multi-layer perceptron; The working process of the Dinov2 large model is: Given an input image F∈R c×h×w , c represents the number of feature channels, h and w represent the image height and width respectively; first, a convolutional layer is used to divide the input image into multiple image blocks; then each image block is converted into a one-dimensional vector through linear mapping; then it is input into the Transformer encoder, and after being processed by the Transformer encoder, a feature matrix of size c×D is output, where c represents the number of feature channels and D represents the feature vector dimension; Subsequently, the feature matrix is converted into a 1×n feature vector through layer normalization, where n is the number of columns.
3. The cross-view image geolocation method for drone and satellite images according to claim 2, characterized in that: The Conv-LoRA uses an encoder-decoder structure to impose low-rank constraints on the weight update of the Dinov2 large model, specifically: Given a pre-trained weight W0∈R d×k , first construct a low-rank decomposition based on the size of the matrix to represent the parameter update ΔW∈R d×k , ΔW is obtained by multiplying two low-rank matrices, encoder A and decoder B. In addition, the hybrid expert model is applied to the encoder A. The original parameter W0 is frozen during the training process, and only the internal parameters of Conv-LoRA are trained. Finally, after the training is completed, the pre-trained weight W0 and ΔW are added as the fine-tuned model parameter W′, expressed as: W′=W0+ΔW; accordingly, the forward propagation process changes from X′=W0x to: Where x represents the input matrix, represents the mixed expert model, and X′ represents the output matrix after weighting by the VisionTransformer encoding block.
4. The cross-view image geolocation method for unmanned aerial vehicles and satellite images according to claim 1, characterized in that: The working process of the spatial relationship-aware feature aggregator MSRA based on the Mamba module is as follows: First, the input feature image F l ∈R c×h×w The maximum pooling operation is used along the channel axis. The maximum pooling operation uses the max(·) function to aggregate the feature map to obtain the aggregated feature matrix M∈R h×w and its corresponding position index matrix M idx ∈R h×w ; Next, the first embedding layer is used to map the aggregated feature map M to K projection vectors E = [e1, e2, ···, e k ], each projection vector e i The dimension is h×w; embedding E in the standard learnable position pe Based on this, we introduce additional index-aware position embedding. After the index-aware position embedding, the vector The mathematical formula is: Among them, W LN Represents a learnable linear transformation, M idx Mapped to K different subspaces, HardTanh represents the activation function; Then Input into the Mamba encoder to extract high-level feature representation; Finally, the second embedding layer is used to project the output of the Mamba encoder into a set of K geometric layout descriptors E2 = [e1, e2, ···, e k ]; These geometric layout descriptors are element-wise multiplied with E, flattened and normalized, and finally a global descriptor f is obtained that contains both the spatial structure and the key feature information.
5. The cross-view image geolocation method for unmanned aerial vehicles and satellite images according to claim 4, characterized in that: The mathematical expression of the global descriptor f is: f = BN (flatten (E2 × E)) Among them, BN means batch normalization and flatten means flattening.
6. The cross-view image geolocation method for unmanned aerial vehicles and satellite images according to claim 1, characterized in that: The mathematical expression of the InfoNCE loss function is: Where q represents the feature code of the query image, B represents the feature code set of all reference images in a batch, and there is a positive sample r matching q in the set. + And B-1 negative samples r that do not match q i , (·) indicates that the dot product is used to calculate the similarity between the query image and the reference image, and the temperature coefficient τ is a hyperparameter.
7. A cross-view image geolocation system for drones and satellite images, characterized in that: The system is used to implement the cross-view image geolocation method for unmanned aerial vehicles and satellite images as described in any one of claims 1 to 6, and comprises: The feature extraction module is used to use the Dinov2 large model fine-tuned by Conv-LoRA as a feature encoder to extract the feature vectors of drone images and satellite images respectively; Feature aggregation module, which is used to design the spatial relationship-aware feature aggregator MSRA based on the Mamba module to aggregate image features and embed spatial configuration features into the global descriptor; The training module is used to train the model using the InfoNCE loss function; The similarity calculation module is used to use the trained model to extract the feature vectors of the image to be queried and each image in the reference image library, and then use cosine similarity to calculate the similarity score between the image to be queried and all reference images, and select the drone-satellite image pair with a high matching degree to realize the geographic positioning of the image to be queried.
8. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Any-inclination-angle unmanned aerial vehicle-satellite geographic positioning method and system
CN121582347A
An arbitrary inclination unmanned aerial vehicle-satellite geolocation method and system
CN121582347B
Low-altitude ground feature recognition sample automatic labeling method and system applying SAM fine tuning
CN121837832A
Unmanned aerial vehicle assisted ground and satellite cross-view image geographic positioning method and system
CN122115579A
UAV-assisted geolocation method and system using ground and satellite images from different perspectives
CN122115579B