Cross-view-angle geographic positioning method based on satellite-ground coupling
Through the combination of hemispheric projection and feature extraction network, the problem of high complexity of geolocation methods in the existing technology is solved, and efficient and accurate ground target positioning is achieved in actual satellite images.
Patent Information
- Application Number
- CN202311838073.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-08
AI Technical Summary
Existing cross-view geolocation methods usually require strict data conditions and additional semantic information, resulting in high application complexity and difficulty in effectively applying in actual scenarios.
The hemispherical projection principle is used to convert the ground query image into a satellite perspective, and the feature extraction network is used to match features, combining the feature extraction network and a multi-scale fusion algorithm to achieve matching and positioning of the ground image and the satellite perspective.
It reduces substantial viewing angle differences in cross-view positioning, improves the convenience and accuracy of positioning, and can directly accurately locate ground targets in actual satellite images.
Smart Images

Figure CN120276005A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a cross - perspective geolocation method based on space - ground coupling, belonging to the technical field of cyberspace mapping. Background Art
[0002] Geolocation greatly expands the scope of image geolocation. It has important value for the recognition of the location of network images. It can not only help people understand the image content more accurately and systematically, but also provide important support for the acquisition of government target information and the recognition of important targets in the military field. Since satellite images can cover the whole world and carry corresponding geographical tags, using them as reference images for images. However, due to the drastic changes in the perspectives of ground and satellite images, the overlap of images from different perspectives is small, the changes are large, and there are usually also illumination and seasonal changes, resulting in huge visual content differences between ground and satellite images. This also makes cross - perspective geolocation a very challenging task.
[0003] In order to improve the effect of cross - perspective geolocation, many scholars focus on improving the accuracy of cross - perspective geolocation by reducing the huge perspective differences. The methods adopted by many scholars not only focus on feature representation, but also explicitly solve the viewpoint differences between satellite and street - view images by transforming the input images.
[0004] (1) Using geometric relationships for perspective transformation: Initially, generating street views from satellite images by using depth maps and semantic labels to perform perspective transformation. For example, street - level information was synthesized from top - view satellite inputs and used to create street - view information. Although they all transformed satellite - perspective images into ground - perspective images, they did not explicitly apply it to the geolocation problem and all required some additional semantic information, increasing the usage cost. Later, the method of polar - coordinate transformation emerged, which reduced the difference by transforming satellite - perspective images into a structure similar to ground - panoramic images and was widely applied. However, this method requires strict data conditions, requiring that the satellite - perspective image and the ground query image have the same range, the same center, and the same direction. In addition, the direction of the landscape image was used as supplementary information for alignment to reduce the influence of parallax. Besides, there was also the use of a geometric projection module to project the features of the satellite map onto the ground view based on the relative camera pose, and explicitly establish the geometric correspondence between the two - view images to extract geometric details from the initial features and obtain the spatial link between visual elements to assist the transformation process. However, these methods assume that the true - north direction of the satellite image is aligned with the starting direction of the panoramic image, which is usually not feasible for image geolocation because the camera azimuth is usually unknown. Therefore, the universality of these methods in real - world scenarios is limited, restricting their applicability.
[0005] (2) Use the GAN network for perspective conversion: Use the aerial images synthesized by GAN to extract complementary features for cross-view image matching, and generate natural scene images from arbitrary viewpoints according to the scene images and novel semantic maps. Most existing cross-view synthesis methods require semantic mapping to fully preserve the content. For example, by learning the latent conversion clues between the image code and the semantic map code, or by combining global image-level and local classes for calculation. In addition to using semantic segmentation, edge maps are also combined as input images. And use conditional GAN to create ground images from satellite views to make the generated street view images more realistic. Currently, these cross-perspective geographical location identifications based on the GAN network are all conversions from satellite perspective to ground perspective and require semantic image information of the corresponding ground images as inputs, which are difficult to apply in actual geographical positioning.
[0006] In summary, the existing perspective conversion methods mainly focus on converting satellite perspectives to ground perspectives. These methods have strict requirements for data, requiring the reference data to be consistent with the ground images in terms of range and center, which is often unrealistic. In terms of the processing flow, when positioning the ground target image, all the reference satellite images need to be converted to the ground perspective, increasing the complexity of the application. Summary of the Invention
[0007] The purpose of the present invention is to provide a cross-perspective geographical positioning method based on space-ground coupling to solve the problem of high complexity caused by currently using the conversion from satellite perspective to ground perspective for positioning.
[0008] The present invention provides a cross-perspective geographical positioning method based on space-ground coupling to solve the above technical problems. The positioning method includes the following steps:
[0009] 1) Use the principle of hemisphere projection to convert the ground query image into a satellite perspective view;
[0010] 2) Use the feature extraction network to extract features from the converted satellite perspective view and the original perspective view;
[0011] 3) Based on the extracted features, achieve the matching between the ground image and the original satellite perspective view, and realize the positioning of the ground query image according to the positioning data of the matched original satellite perspective view.
[0012] Further, in the step 1), it is converted using the projection relationship between the ground panoramic image coordinates and the satellite image coordinates, and the projection conversion relationship is:
[0013]
[0014]
[0015]
[0016]
[0017]
[0018] (x1, y1, z1) are the spatial coordinates of the satellite image, and (x2, y2, z2) are the spatial coordinates of the ground panoramic image, where H is the height of the satellite from the ground. The satellite camera projects a point (x1, y1, z1) onto its image coordinates through parallel projection. where s is the resolution of the satellite image, and (u0, v0) is the center of the satellite image. are the image coordinates of the ground panoramic image, and H g is the height of the panoramic image, and W g is the width of the panoramic image, and θ is the elevation angle relative to the z2 axis in the ground camera coordinate system. is the azimuth angle.
[0019] Furthermore, the feature extraction network in step 2) includes a backbone module and an aggregation module. The backbone module is used to extract local features of the input image; the aggregation module is used to aggregate the extracted local features.
[0020] Furthermore, the backbone module uses a Resnet34 network with the original classifier removed, and the output of the last convolutional layer of the Resnet34 network is used as the output of the backbone module.
[0021] Furthermore, the aggregation module uses a SAFA aggregation module.
[0022] Furthermore, the feature extraction network is trained using a metric loss function with hard sample mining.
[0023] Furthermore, the feature extraction network is trained using a weighted soft margin triplet loss function.
[0024] Furthermore, step 2) also includes performing multi-scale fusion processing on the features extracted by the feature extraction network. The process is as follows: The feature map extracted by the feature extraction network is decomposed into two one-dimensional features, and pooling is performed along the x and y directions respectively; Feature extraction at different scales is performed on the results of the pooling, and the features at each scale are concatenated; The concatenated features are multiplied element-wise with the feature map extracted by the feature extraction network.
[0025] The beneficial effects of the present invention are as follows: The present invention couples satellite images and ground images using the hemispherical projection relationship, converts the query image from the ground perspective to the satellite perspective, avoids the limitation of idealized reference data in the current method of converting satellite images to the ground perspective, and improves the convenience of application. Moreover, through the satellite-ground image coupling method of the present invention, the substantial perspective differences in cross-view positioning are reduced, making the converted features closer to the corresponding features in other domains, which helps to more accurately compare similarities, and thus achieve more reliable and accurate cross-view geolocation. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is a flowchart of the cross-perspective geolocation method based on satellite-ground coupling of the present invention;
[0027] Figure 2 is a schematic diagram of the satellite-ground coupling relationship established by the present invention using hemispherical projection;
[0028] Figure 3 is a schematic diagram of the image conversion effect obtained by satellite-ground coupling in an embodiment of the present invention;
[0029] Figure 4 is a schematic diagram of the principle of the multi-scale fusion mechanism of spatial features adopted by the present invention;
[0030] Figure 5 is a comparison schematic diagram between the present invention and the existing method on the CVACT dataset during the experiment;
[0031] Figure 6 is a comparison schematic diagram between the present invention and the existing method on the OP dataset during the experiment;
[0032] Figure 7 is a qualitative analysis of the method used in the present invention;
[0033] Figure 8 is a comparison schematic diagram of whether satellite-ground coupling is performed on the CVUSA, CVACT, and OP datasets during the experiment;
[0034] Figure 9a is a comparison schematic diagram of whether satellite-ground coupling is performed on the CVUSA dataset during the experiment;
[0035] Figure 9b is a comparison schematic diagram of whether satellite-ground coupling is performed on the CVACT dataset during the experiment;
[0036] Figure 9c is a comparison schematic diagram of whether satellite-ground coupling is performed on the OP dataset during the experiment;
[0037] Figure 10It is a schematic diagram of the effectiveness of whether to use the location-based multi-scale fusion module on the CVUSA, CVACT, and OP datasets during the experiment;
[0038] Figure 11 It is a schematic diagram for verifying the effect of the location-based multi-scale fusion module of the present invention on the heat map. Detailed implementation manners
[0039] The following further describes the detailed implementation manners of the present invention with reference to the accompanying drawings.
[0040] By establishing a hemispherical projection relationship, satellite images and ground images are effectively coupled, thereby realizing the conversion of ground images to the satellite perspective. This method significantly reduces the substantial perspective differences in cross-view localization.
[0041] The present invention first converts the ground query image into a satellite perspective view using the hemispherical projection principle; then uses a feature extraction network to extract features from the converted satellite perspective view and the original perspective view; finally, based on the extracted features, the ground image is matched with the original satellite perspective view, and the localization of the ground query image is realized according to the localization data of the matched original satellite perspective view. The implementation principle of this method is as Figure 1 shown, and the following is a detailed description.
[0042] 1. Convert the ground query image into a satellite perspective view.
[0043] The present invention effectively couples satellite images and ground images by establishing a hemispherical projection relationship, thereby realizing the conversion of ground images to the satellite perspective. The following is a brief introduction to the hemispherical projection principle.
[0044] As Figure 2 shown, the ground panoramic image is regarded as a cylindrical projection and is unfolded along the side of the cylinder, while the satellite perspective image is regarded as the top surface, and these two perspectives are connected by a hemispherical projection. By using satellite-to-satellite coupling, the query image is converted and mapped into the hemispherical space aligned with the satellite image, which can achieve seamless integration of the ground and satellite perspectives, thereby facilitating accurate cross-view geolocation.
[0045] As Figure 2 shown, the image space coordinates of the satellite image are (x1, y1, z1), and the corresponding ground space coordinates are (x2, y2, z2). The homogeneous coordinate expressions of the same position between the two coordinate systems are as follows:
[0046]
[0047] where H is the height of the satellite from the ground.
[0048] Satellite camera coordinate system, the satellite camera projects a point (x1, y1, z1) to its image coordinates through parallel projection The projection relationship is as follows:
[0049]
[0050] Where S is the resolution of the satellite image (usually 2), and (u0, v0) is the center of the satellite image.
[0051] Ground equidistant cylindrical projection: In the ground camera coordinate system, θ is defined as the elevation angle relative to the z2 axis, and Defined as the azimuth, to be consistent with the polar coordinate transformation, is the north direction, corresponding to the negative direction of the x2 axis. Clockwise is the positive direction. The projection formula between and (x2, y2, z2) is:
[0052]
[0053] where z2 represents the scene height at pixel.
[0054] Coordinates of ground panorama images The projection of is expressed as:
[0055]
[0056] From formula 1, formula 2, and formula 3, we can get:
[0057]
[0058] s is the resolution of the satellite image, (u0, v0) is the center of the satellite image, is the image coordinates of the ground panoramic image, H g is the height of the panoramic image, W g is the width of the panoramic image. Projection between the ground panoramic image and the satellite view image: Solve the previous formulas together and finally substitute formula (5) into formula (4) to establish the coordinates of the ground panoramic image With satellite image coordinates The projection relationship between them.
[0059] In practical applications, the scene height is much smaller than the span between the ground camera and the satellite camera. The height of the ground camera relative to the ground is z2. To achieve this goal, a geometric correspondence between the ground panoramic image and the satellite image is constructed. Space-ground coupling plays a crucial role in effectively aligning the two feature domains, resulting in closer proximity between the corresponding feature maps in the feature space. The main goal of space-ground coupling is to transform the ground perspective features into the satellite domain. By doing so, the transformed features are closer to the corresponding features in other domains, as Figure 3 shown. This proximity helps to more accurately compare similarities, thereby enhancing the discriminability of features. The space-ground coupling process effectively aligns the features in the ground domain with the satellite domain, thus achieving more reliable and accurate cross-view geolocation.
[0060] 2. Feature extraction is performed using a feature extraction network.
[0061] After the space-ground coupling process, feature extraction and loss function calculation need to be performed on the new satellite perspective image obtained by coupling the ground image and the original satellite perspective image to identify the most similar satellite image and achieve precise positioning. The overall feature extraction network is as Figure 1 shown in step2.
[0062] The feature extraction network in this embodiment includes a Backbone module and an aggregation module. The backbone module is used to extract local features of the input image, and the aggregation module is used to aggregate the extracted local features. Among them, the backbone module adopts the Resnet34 network with the original classifier removed, that is, the output of the last convolutional layer of the Resnet34 network is used as the output of the backbone module. In cross-view location recognition, the aggregation modules that appear are Netvlad and SAFA, and SAFA has a good effect. Therefore, this embodiment uses multiple aggregation modules as the final aggregation module. For example, 8 can be used.
[0063] The triplet CNN trained with the soft margin triplet loss is suitable for extracting deep features from cross-view image pairs. Subsequently, the triplet loss is widely used as the objective function for training deep neural networks for image localization and matching tasks. Therefore, this embodiment can use a weighted soft margin triplet loss function to train the feature extraction network combined with multi-scale fusion. Among them, the triplet consists of an anchor point, a positive (matching) example, and a negative (non-matching) example. The purpose of the triplet loss is to learn a distance metric that makes the positive sample closer to the anchor point while pushing the negative sample farther away. Following most geolocation methods, we use a weighted soft margin triplet loss to train our network. If the exponential power in the soft margin triplet loss is multiplied by a scalar, its expression is:
[0064]
[0065] where a is a parameter that controls the convergence rate of the training process, and d pos and d neg are the distances from the positive and negative examples (original satellite view) to the selected anchor (image after ground view transformation), respectively.
[0066] The above metric learning loss function can make similar samples close to each other and different samples far from each other. These traditional metric loss functions randomly sample image pairs from the training data for learning. Although this approach is relatively simple, most of the sampled pairs are easy-to-distinguish sample pairs. If a large number of training sample pairs are simple sample pairs, it is not conducive to the network learning better representations. Training the network with more difficult samples can improve the generalization ability of the network. Therefore, this embodiment adopts a metric loss function based on hard sample mining. Hard sample mining is a training strategy. Still based on the above loss function, during the sample training process, the training results will be used to calculate the IOU with the GroundTruth. Usually, a threshold (0.5) is set. If the result exceeds the threshold, it is considered a positive sample, and if it is below a certain threshold, it is considered a negative sample, and then they are thrown into the network for training.
[0067] 3. Multi-scale fusion of spatial features
[0068] After coupling the satellite and ground images, the new image pairs obtained are more homogeneous, especially in terms of spatial features. Traditional CNN methods usually ignore location information when extracting feature representations, which is crucial for forming spatially selective visual effects. Therefore, the present invention adopts a multi-scale fusion algorithm based on spatial location features, as Figure 4 shown. The present invention embeds the multi-scale spatial feature fusion module directly into the BasicBlock according to the attention insertion method of the traditional Resnet34.
[0069] First, the extracted multi-dimensional feature map is decomposed into two one-dimensional feature encodings, and average pooling is performed along the x and y directions respectively; then, the values in the x and y directions generated by the average pooling are multiplied to amplify the location features of the object of interest; on the basis of amplifying the location features, in order to make the object of interest more prominent, the present invention uses multiple receptive fields of different scales to obtain feature information in space, obtaining multi-scale spatial features. In this embodiment, three convolutional kernels of 3×3, 5×5, and 7×7 are used to obtain features of different scales; finally, the three feature maps of different scales are concatenated along the channel dimension, and a 1×1 convolution is used to generate an output of the same size as the input image. By performing the sigmoid activation function, spatial location features of different scales can be obtained.
[0070] Assuming that the finally obtained multi-scale spatial location feature map is F′ and the input feature map is F, then the following relationship exists:
[0071]
[0072] P(F) = σ(f 1×1 (f 3×3 C(F), f 5×5 C(F), f 7×7 C(F)))
[0073]
[0074] During the spatial location feature fusion process, and respectively represent the avg - pooling operations along the x - axis and y - axis, f n×n is an n×n convolution, σ is the sigmoid activation function, represents element - wise multiplication. The attention mechanism is placed in the residual module of each layer, re - weighted according to the positions of important features, and learns more discriminative feature representations.
[0075] 4. Use the obtained spatial location features of different scales to identify the most similar satellite images for positioning.
[0076] Calculate based on the spatial location features of the new satellite perspective image coupled with the extracted ground image and the spatial location features of the original satellite perspective image. According to the loss function, the image with the minimum loss is considered the consistent image, and the original satellite perspective image most similar to the new satellite perspective image coupled with the ground image is found, and the geographical location of this original satellite perspective image is used as the geographical location of the ground image to achieve the positioning of the ground image.
[0077] Experimental verification
[0078] To further illustrate the effect of the present invention in using the satellite perspective image as a reference image to complete the ground image positioning, a simulation experiment of the invention is carried out below.
[0079] In this experiment, satellite perspective images and corresponding ground images were used as the test and validation datasets to establish evaluation metrics for assessing the proposed algorithm. The training weights of Resnet34 on ImageNet were used as the initial weights, and the parameters in SAFA and the attention module were randomly initialized. The ground panoramic image and the satellite perspective image (or the satellite perspective image obtained by geometric transformation of the ground panoramic image) were respectively adjusted to 256×512 and 320×320 pixels, and then input into the network. The triplet loss was set to 10, and the network was trained using the stochastic gradient descent optimizer with a learning rate of 0.001 and a batch size of 32. The experiment was conducted using PyTorch, and a model was trained for 150 epochs using 4 3090 graphics processing units. To be consistent with existing methods, according to previous studies, K = 1, 5, 10, and 1% were selected, and the recall rate of the highest K (r@K) was used for evaluation.
[0080] (1) Comparison with other methods
[0081] To evaluate the feasibility and effectiveness of the present invention, this experiment tested it in the cross-view geolocation task. Using the public datasets CVACT and OP datasets as benchmarks, a comprehensive quantitative evaluation and application analysis of the image representation effects of different network models were carried out according to the metrics in the evaluation system.
[0082] 1) Quantitative evaluation
[0083] Quantitative evaluations were carried out on suburban data (CVACT) and urban data (OP). The method proposed in the present invention was evaluated by comparing it with some state-of-the-art cross-view location recognition methods on the test dataset CVACT. As Figure 5 shown, the method proposed in the present invention achieved a high recall rate. In the public CVACT_val dataset, the algorithm achieved the following performance metrics: the values of r@1, r@5, r@10, and r@1% reached 75.08, 89.34, 92.40, and 97.33 respectively. Similarly, in the CVACT_test dataset, the values of r@1, r@5, r@10, and r@1% achieved by the algorithm were 39.92, 61.69, 67.38, and 92.36 respectively. These results indicate that the method proposed in the present invention is superior to most existing algorithms in terms of accuracy and top retrieval performance.
[0084] This experiment also compared the present invention with the publicly available application results of the public OP dataset, as Figure 6 shown. The method of the present invention achieved significant positioning accuracy in urban areas. For the public dataset OP, its r@1 was 9.00, which was better than most other algorithms.
[0085] According to the above quantitative evaluation, it can be seen that the method of the present invention has achieved good accuracy results for both suburban and urban datasets, verifying its feasibility.
[0086] 2) Qualitative evaluation
[0087] Cross-view image localization involves using satellite images with global positioning system coordinates to locate ground images. As Figure 7 shown, the process has the following steps: (1) Determine the satellite image that is most similar to the ground image to complete the localization. However, due to the large domain difference between the two, there are obvious visual differences between them. Thus, cross-view geolocation becomes challenging. (2) To reduce the domain difference, many researchers perform polar coordinate transformation on satellite images to make their structures similar to those of ground images to improve the accuracy of localization. This significantly reduces the domain difference between the two. However, in practical applications, it is required that the satellite images used must have the same coverage and consistent orientation as the ground images to ensure that the transformed satellite images are consistent with the corresponding ground images in structure. However, in practical applications, the ground query images are unknown, and it is difficult to ensure that the satellite images and ground images have exactly the same coverage, etc., which limits the use of this method. In addition, during the localization process, all satellite images must be transformed, which requires a large amount of work, especially when dealing with large-scale datasets. (3) The research uses the hemispherical coupling relationship to transform the unknown ground images so that they are converted into perspectives similar to satellite images, thereby reducing the domain difference and improving the accuracy of localization. In practical applications, only a simple transformation of the target image is required to obtain an image similar to a satellite image, breaking through the previous restrictions on data coverage, etc., and also greatly increasing the convenience of application. This method also makes it possible to directly use actual satellite images for localization in the future.
[0088] In summary, the method proposed by the present invention reduces the domain difference and improves the accuracy of localization by using the hemispherical coupling relationship to transform the ground images into perspectives similar to satellite images, which can promote direct localization based on actual satellite images.
[0089] Ablation experiment
[0090] An ablation analysis was conducted to evaluate the efficiency of the proposed design.
[0091] (1) Space-ground coupling effectiveness
[0092] To study the success of the proposed satellite-ground coupling method, Resnet34 was selected as the underlying network for feature extraction. To keep up with the subsequent aggregation module SAFA, the number of classes was adjusted to 2048. The models were named: Resnet34 (using only Resnet34 for feature extraction) and SG_Resnet34 (transforming ground images based on satellite-ground coupling and then using Resnet34 for feature extraction). As Figure 8 shown, using ground images and satellite images transformed by hemispherical projection as inputs, at r@1, the performance of CVUSA_P2S increased by 13.5%, the performance of CVACT increased by 15.76%, and the performance on the OP dataset increased by 1.39%. As Figure 9a , Figure 9b and Figure 9c shown, it can be seen that the satellite-ground coupling technology can improve the performance of the two datasets.
[0093] (2) Effectiveness of the multi-scale spatial feature fusion mechanism
[0094] Comparing the multi-scale fusion of spatial features in the spatial feature module with the baseline network demonstrated its effectiveness. As Figure 10 shown, with the addition of the multi-scale fusion module of spatial features, the performance metrics of the network improved. In particular, our method achieved a recall rate of 26.58% for CVUSA_P2S at r@1, a relative increase of 3.57%; the recall rate of CVACT_val was 75.08%, a relative increase of 1.82%; and the recall rate of OP was 9.00%, a relative increase of 0.83%.
[0095] To further illustrate the effectiveness of the multi-scale spatial feature fusion mechanism proposed in the present invention, this experiment used Grad-CAM to visually interpret the heatmaps to show which regions contributed the most to the cosine similarity of the view embedding features. As Figure 11 shown, activation maps of satellite perspective images and transformed images were generated for comparison. In the heatmap, the bright colors of the features indicate a higher degree of activation, and the colors close to red indicate. In this study, the multi-scale fusion spatial feature module was inserted into the fifth layer of the backbone network Resnet34 for display. As shown in the figure, the network focuses on the road and suppresses irrelevant objects. Through our design based on the multi-scale spatial feature fusion mechanism, the proposed method successfully focuses on meaningful features.
Claims
1. A cross - perspective geolocation method based on satellite - ground coupling, characterized in that The positioning method includes the following steps: 1) Convert the ground query image into a satellite perspective view using the hemisphere projection principle; 2) Use a feature extraction network to extract features from the converted satellite perspective view and the original perspective view; 3) Based on the extracted features, implement the matching of the ground image and the original satellite perspective view, and realize the positioning of the ground query image according to the positioning data of the matched original satellite perspective view.
2. The cross-view geolocation method based on satellite-ground coupling according to claim 1, wherein In step 1), it is converted using the projection relationship between the ground panoramic image coordinates and the satellite image coordinates, and the projection conversion relationship is: (x1, y1, z1) are the spatial coordinates of the satellite image, and (x2, y2, z2) are the spatial coordinates of the ground panoramic image. Here, H is the height of the satellite from the ground. The satellite camera projects a point (x1, y1, z1) onto its image coordinates through parallel projection. where s is the resolution of the satellite image, and (u0, v0) is the center of the satellite image. are the image coordinates of the ground panoramic image, and H g is the height of the panoramic image, and W g is the width of the panoramic image. θ is the elevation angle relative to the z2 axis in the ground camera coordinate system. is the azimuth angle.
3. The cross-view geolocation method based on space-ground coupling according to claim 1 or 2, characterized in that The feature extraction network in step 2) includes a backbone module and an aggregation module. The backbone module is used to extract local features of the input image; the aggregation module is used to aggregate the extracted local features.
4. The cross-view geolocation method based on satellite-ground coupling according to claim 3, characterized in that The backbone module uses a Resnet34 network with the original classifier removed, and the output of the last convolutional layer of the Resnet34 network is used as the output of the backbone module.
5. The cross-view geolocation method based on satellite-ground coupling according to claim 3, characterized in that The aggregation module uses a SAFA aggregation module.
6. The cross-view geolocation method based on space-ground coupling according to claim 3, characterized in that, The feature extraction network is trained using a metric loss function with hard sample mining.
7. The cross-view geolocation method based on satellite-ground coupling according to claim 3, characterized in that The feature extraction network is trained using a weighted soft margin triplet loss function.
8. The cross-view geolocation method based on satellite-ground coupling according to claim 1 or 2, characterized in that, Step 2) also includes performing multi-scale fusion processing on the features extracted by the feature extraction network. The process is as follows: The feature map extracted by the feature extraction network is decomposed into two one-dimensional features, and pooling is performed along the x and y directions respectively; Feature extraction at different scales is performed on the results of pooling, and the features at each scale are concatenated; The concatenated features are multiplied element-wise with the feature map extracted by the feature extraction network.