Cross-view-angle target geographic positioning method based on view generation

Through polar coordinate conversion and Pix2Pix GAN network, the domain gap between cross-view images is narrowed, and the ConvNeXtV2 network is used for feature matching, which solves the shortcomings of the existing technology of cross-view geolocation methods in feature matching and real-time, and achieves higher positioning accuracy and real-time.

CN119992357AActive Publication Date: 2025-05-13NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411952180.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-13
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

The existing cross-view geolocation method is difficult to effectively realize cross-view image feature matching in feature extraction networks, and network computing redundancy results in insufficient real-time.

Method used

Polar coordinate transformation is used to generate panoramic viewing images, and the conditional generation adversarial network Pix2Pix GAN narrows the domain gap between cross-viewing images. Then, a geolocation matching model is constructed based on the pre-trained ConvNeXtV2 network, and the model is trained through InfoNCE comparison loss, the distance between positive sample feature vectors is reduced and the distance between negative sample feature vectors is expanded.

Benefits of technology

Effectively narrow the domain gap between cross-view imagery, improve the accuracy and real-timeness of cross-view object positioning, and is better than past work and methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992357A_ABST
    Figure CN119992357A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-view target geographic positioning method based on view generation, and the method comprises the steps: employing a polar coordinate conversion method for cross-view pictures obtained through satellite remote sensing and unmanned aerial vehicle aerial photographing, and preliminarily eliminating the gap between two view domains; a conditional generative adversarial network Pix2Pix GAN is utilized to train a satellite-unmanned aerial vehicle image pair, the generative network is used for generating an unmanned aerial vehicle view angle picture from a satellite remote sensing image, and a satellite view angle image library is constructed; the generated pictures are in one-to-one correspondence with the initial unmanned aerial vehicle view angle pictures, and a training set and a verification set are divided for training and evaluation of subsequent geographic positioning matching; a deep learning model ConvNeXtV2 subjected to ImageNet data set pre-training is used for discriminating feature extraction and feature similarity matching, so that cross-view geographic positioning is realized. The method still shows relatively high robustness even under the conditions of large view angle span (an unmanned aerial vehicle 45-degree oblique shooting view angle and a satellite aerial view view angle) and variable shooting angles of the unmanned aerial vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of geographic positioning, and in particular relates to a cross-viewing angle target geographic positioning method based on view generation. Background Art

[0002] Cross-view target geolocation can improve the accuracy and robustness of target positioning by combining information from different perspectives. Data from different perspectives provides multi-dimensional information about the target, reduces the error that may be caused by a single perspective, and improves the accuracy of positioning. It can enhance target detection and tracking capabilities through the collaboration of different perspectives in a dynamically changing complex environment.

[0003] Cross-view geolocation aims to determine the target location through cross-platform image matching, including matching ground front view images, oblique view images of drones at nearly 45 degrees, and satellite overhead view images with geotags. Its core lies in transcending the perspective difference and establishing the association between these images and the actual geographic space. Cross-view geolocation usually involves a variety of technologies, including feature extraction, deep learning modeling, large-scale image retrieval, and spatial geometric reasoning. It not only focuses on low-level features such as texture, shape, and color of the image content itself, but also needs to combine contextual information and semantic understanding to capture potential correlation features in cross-view images. Existing technologies have been mainly applied to scenarios such as drone positioning, disaster assessment, smart city construction, and military reconnaissance, that is, by providing drone image queries to match candidate satellite images and then locate the current geographic location. At the same time, this technology also provides important support for improving cross-domain cognitive capabilities in the visual field.

[0004] The research on cross-view geolocation can be divided into two categories: one is cross-view matching based on traditional manual features, and the other is based on deep learning. In the cross-view matching method based on traditional manual features, manual features such as oriented gradient histogram and color histogram were initially used for matching. Later, some studies distorted street view panoramic images into overhead images and used feature detection algorithms such as SIFT, SURF and FREAK for matching to locate the geographic location of the image to be queried. However, the image features extracted by the manual image feature extraction method are too different, and the retrieval accuracy is not high.

[0005] In the cross-view matching methods based on deep learning, some methods try to use convolutional neural networks (CNNs) to build network models, pre-train cross-view image pairs, use the trained networks to extract features of ground query images and overhead images in the database, and perform feature retrieval to obtain the geographic location of the query image. For example, LPN (local pattern network) achieves end-to-end learning of context information through feature-level division strategies, and SAIG-D (simple attention image geolocation network) adopts a fully connected form of MLP-mixer. This method can effectively represent the long-range interaction and cross-view correspondence between image blocks through multi-head self-attention layers. Some recent studies have tried to use vision transformers (ViT) for cross-view image retrieval. However, these methods all try to directly extract image features between different domains for comparison through more powerful neural network models. They only learn feature representation based on image content, but ignore the spatial correspondence between drone and satellite images, so the positioning ability is difficult to be effectively improved.

[0006] In order to reduce the difference between different perspectives, many studies have been devoted to converting cross-perspective images to the same perspective before image retrieval. The main perspective conversion methods currently include perspective projection transformation, polar coordinate transformation, generative adversarial network (GAN), etc. The PCL method uses conditional GAN ​​to generate corresponding satellite perspective images for drone perspective images, and then uses the ResNet network for image matching and geolocation. However, during the use phase of this method, it is necessary to first input the drone perspective image into the GAN to generate the corresponding satellite view, and then input the result into the feature extraction network. The extracted features are used for retrieval and comparison with the features in the satellite image library. This process is serial and there is computational redundancy, which may be problematic for some scenarios that require real-time performance.

[0007] Therefore, in view of the above shortcomings, there is an urgent need to provide a new cross-perspective geolocation method. Summary of the invention

[0008] The purpose of the present invention is to provide a cross-view target geolocation method based on view generation, aiming to overcome the shortcomings of existing geolocation methods that it is difficult to effectively achieve cross-view image feature matching only by relying on feature extraction networks, and network calculation redundancy makes the real-time performance insufficient.

[0009] The technical solution to achieve the purpose of the present invention is: a cross-view target geographic positioning method based on view generation, comprising the following steps:

[0010] Step 1: Obtain satellite and drone cross-view geolocation images of the area to be queried, and construct image pairs that match each other from different viewpoints. The image pairs include drone view images and satellite view images of the same geographic location, and divide the drone view images into training sets and validation sets.

[0011] Step 2: For each image in the image pair in step 1, use polar coordinate transformation, take the center point of the image as the origin of the polar coordinate system, remap each pixel of the image to the polar coordinate system, and generate the corresponding panoramic view image, where the panoramic view image generated by the drone view image is uniformly denoted as D, and the panoramic view image generated by the satellite view image is uniformly denoted as S, and D is constructed as a drone image query library, denoted as Q;

[0012] Step 3: Construct a conditional generative adversarial network Pix2Pix GAN, take S and D in step 2 as input at the same time, output the pseudo drone view image generated by S, train the conditional generative adversarial network Pix2Pix GAN, and construct the pseudo drone view image generated by S as a satellite image retrieval library, denoted as R;

[0013] Step 4: Match Q in step 2 with R in step 3 one by one, and use the successfully matched image pairs as positive sample pairs, and the rest as negative sample pairs;

[0014] Step 5: Build a geolocation matching model based on the pre-trained ConvNeXtV2 network. Take S and D in step 2 as input at the same time. Each image is input to the ConvNeXtV2 network in turn. Output the feature vector corresponding to each image. Use InfoNCE contrast loss for back propagation to reduce the feature similarity distance between the feature vectors corresponding to the positive sample pairs and expand the distance between the feature vectors corresponding to the negative sample pairs. Train the geolocation matching model and use the satellite image features of the positive sample pairs to build a satellite image feature retrieval library.

[0015] Step 6: Use the features of the drone images to be tested in the validation set to perform feature matching and sorting with the satellite image feature retrieval library, match the K satellite image features with the highest feature similarity, and obtain the location of the target in the drone's view image through the target detection model. For the identified target area, cross-check the satellite remote sensing image corresponding to the matched satellite image features, compare and analyze the spatial geometric relationship and pixel distribution characteristics of the target, recalculate the pixel geometric center of the target in the satellite view image, and realize cross-view target positioning.

[0016] Further, step 2: for each image in the image pair in step 1, polar coordinate conversion is used, with the center point of the image as the origin of the polar coordinate system, and each pixel of the image is remapped to the polar coordinate system to generate a corresponding panoramic view image, wherein the panoramic view image generated by the drone view image is uniformly recorded as D, and the panoramic view image generated by the satellite view image is uniformly recorded as S, and D is constructed as a drone image query library, recorded as Q, wherein the conversion formula is as follows:

[0017]

[0018] in and Represent the pixel coordinates of the original view image and the panoramic view image respectively, and their image sizes correspond to (W s ,H s ) and (W ps ,H ps ).

[0019] Further, step 3: construct the conditional generative adversarial network Pix2Pix GAN, take S and D in step 2 as input at the same time, output the pseudo drone perspective image generated by S, train the conditional generative adversarial network Pix2Pix GAN, and construct the pseudo drone perspective image generated by S as a satellite image retrieval library, denoted as R, where:

[0020] The conditional generative adversarial network Pix2Pix GAN includes a generator and a discriminator. The generator network model adopts the U-Net architecture, which contains an encoder and a decoder, which are used for downsampling feature extraction and upsampling recovery, respectively. Specifically, given a satellite view image In the encoder, the input is transformed into Then the feature map is obtained by downsampling through the first encoding block. Similarly, in each round of downsampling, the size of the feature map is halved and the number of channels is doubled. After the second, third, and fourth encoding blocks, the downsampling obtains feature maps X2, X3, and X4 respectively. In the decoder, feature map X4 is concatenated with feature map X3 through the nearest neighbor upsampling layer and then input into the convolution layer to obtain feature map Y4. Y4 is concatenated with X2 through the upsampling layer and then input into the convolution layer to obtain feature map Y3. Similarly, Y2 and Y1 are obtained. Finally, feature map Y1 is convolved with 1×1 to obtain the final output. Y0 is the final pseudo drone perspective image;

[0021] The discriminator network model uses the PatchGAN structure, which consists of multiple layers of convolution operations. It concatenates the input real drone view image with the pseudo drone view image generated by the generator. Then it goes through three convolutional layers for downsampling, each time the feature map channel is doubled and the size is halved, and then goes through padding, 1×1 convolution, BN normalization, LeakyReLU activation function, padding, 1×1 convolution, and finally the output matrix is ​​obtained. Used to judge the authenticity of each area of ​​the pseudo-UAV perspective image.

[0022] Further, step 3: construct the conditional generative adversarial network Pix2Pix GAN, take S and D in step 2 as input at the same time, output the pseudo drone perspective image generated by S, train the conditional generative adversarial network Pix2Pix GAN, and construct the pseudo drone perspective image generated by S as a satellite image retrieval library, denoted as R, where the total loss function formula of the training conditional generative adversarial network Pix2PixGAN is:

[0023]

[0024] L L1 =E[||yG(x)||1]#

[0025] L total =λL L1 +L GAN

[0026] Among them, y is the real target image, G(x) is the image generated by the generator, D(x,y) is the discriminant value of the real image, D(x,G(x)) is the discriminant value of the generated image, L L1 is the L1 loss between the generated image and the target image, L GAN is the adversarial loss between the generator and the discriminator, L total is the final total loss function, and the parameter λ is used to balance the importance of L1 loss and adversarial loss.

[0027] Further, step 3: construct the conditional generative adversarial network Pix2Pix GAN, take S and D in step 2 as input at the same time, output the pseudo drone perspective image generated by S, train the conditional generative adversarial network Pix2Pix GAN, and construct the pseudo drone perspective image generated by S as a satellite image retrieval library, denoted as R. In the process of training the conditional generative adversarial network Pix2PixGAN:

[0028] The optimizer uses the Adam method, the learning rates of the generator and discriminator are both 0.0001, the momentum parameters β1 and β2 are set to 0.5 and 0.999 respectively, and the batch size is set to 8.

[0029] Further, step 5: construct a geolocation matching model based on the pre-trained ConvNeXtV2 network, take S and D in step 2 as input at the same time, pass each image through the ConvNeXtV2 network in turn as input, output the feature vector corresponding to each image, use InfoNCE contrast loss for back propagation, reduce the feature similarity distance between the feature vectors corresponding to the positive sample pair, expand the distance between the feature vectors corresponding to the negative sample pair, train the geolocation matching model, and use the satellite image features of the positive sample pair to construct a satellite image feature retrieval library, where:

[0030] The ConvNeXtV2 network consists of four stages, each of which stacks several ConvNeXtV2Blocks and one convolutional downsampling block. The number of ConvNeXtV2 Blocks in the four stages are 3, 3, 27, and 3, respectively. Each ConvNeXtV2Block is connected in sequence to the depthwise separable convolution DW-Conv, the pointwise convolution Conv, and the GELU activation function, which are used to extract spatial features, reorganize feature channels, and improve nonlinear expression capabilities, respectively. The downsampling block is implemented using convolution. Except for the first stage downsampling block which uses a 4-fold downsampling rate, the remaining stages all use a 2-fold downsampling rate.

[0031] The initial input image is After each stage, the number of feature channels gradually increases to 128, 256, 512, and 1024, respectively, and the image size is halved at each stage. After 4 stages, we get Finally, a global average pooling GAP and a fully connected layer Linear are used to obtain a feature vector output with a length of 1000.

[0032] Further, step 5: construct a geolocation matching model based on the pre-trained ConvNeXtV2 network, take S and D in step 2 as input at the same time, pass each image through the ConvNeXtV2 network in turn as input, output the feature vector corresponding to each image, use InfoNCE contrast loss for back propagation, reduce the feature similarity distance between the feature vectors corresponding to the positive sample pair, expand the distance between the feature vectors corresponding to the negative sample pair, train the geolocation matching model, and use the satellite image features of the positive sample pair to build a satellite image feature retrieval library, where the calculation formula of the InfoNCE loss function for training the geolocation matching model is:

[0033]

[0034] Where N represents the size of each batch, A + represents the unique query positive sample in each batch, B -Represents the (N-1) negative samples in each batch. During the training process, the query vector continuously narrows the distance between the positive samples and widens the distance between the query vector and the negative samples, thereby reducing the loss value.

[0035] Further, step 5: construct a geolocation matching model based on the pre-trained ConvNeXtV2 network, take S and D in step 2 as input at the same time, pass each image through the ConvNeXtV2 network in turn as input, output the feature vector corresponding to each image, use InfoNCE contrast loss for back propagation, reduce the feature similarity distance between the feature vectors corresponding to the positive sample pair, expand the distance between the feature vectors corresponding to the negative sample pair, train the geolocation matching model, and use the satellite image features of the positive sample pair to build a satellite image feature retrieval library. In the process of training the geolocation matching model:

[0036] The AdamW optimizer was used for training, the initial learning rate was set to 0.001, the learning rate decay adopted the cosine annealing algorithm, the minimum learning rate was set to 0.0001, the batch size was set to 64, and the maximum number of training epochs was 10.

[0037] Furthermore, the target detection model adopts the YOLOv5 network.

[0038] A cross-view target geographic positioning system realizes cross-view target geographic positioning based on the cross-view target geographic positioning method, and is divided into 6 modules, which respectively execute steps 1 to 6.

[0039] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, cross-view target geolocation is achieved based on the cross-view target geolocation method.

[0040] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the cross-viewing angle target geolocation is realized based on the cross-viewing angle target geolocation method.

[0041] Compared with the prior art, the present invention has the following significant advantages: 1) Fully utilizing the spatial correspondence between drone aerial images and satellite remote sensing images. First, polar coordinate transformation is used to restore the scene distribution and object arrangement between cross-perspective images, and initially narrow the domain gap between satellite perspective and drone perspective images. Secondly, the Pix2Pix GAN network is used to train the image pairs generated by polar coordinate transformation, so as to achieve the purpose of inputting satellite perspective images and generating drone perspective images by learning the detailed feature mapping relationship. After these two steps, the domain gap between cross-perspective image pairs is further narrowed. 2) The ConvNeXtV2 model is used, which has stronger multi-scale feature extraction capabilities and is more capable of capturing long-distance dependencies and fine-grained features. The features of the drone perspective query set are matched with the features in the satellite perspective retrieval library, thereby achieving the purpose of precise geographic positioning. 3) The drone view image is fed into GAN to generate the corresponding satellite view, and then the result is fed into the feature extraction network. This process is serial and has computational redundancy. Our method is to feed the satellite remote sensing image into GAN to generate pseudo drone view images, and then build a satellite image retrieval database based on the generated images. This step can be performed in parallel with the input drone aerial image retrieval and positioning, and it no longer requires the input drone aerial image to be generated again, which greatly reduces computational redundancy and improves the real-time performance of cross-view geolocation. 4) In general, the present invention is superior to previous work and methods in terms of the accuracy of cross-view geolocation matching, and opens up a new direction for the construction of future intelligent combat systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 The figure is an overall flow chart of cross-view target geographic positioning in an embodiment of the present invention.

[0043] Figure 2 This is a structural diagram of generating a network model in an embodiment of the present invention.

[0044] Figure 3 4 is a structural diagram of the discriminant network model in an embodiment of the present invention.

[0045] Figure 4 Schematic diagram of the feature extraction network model in an embodiment of the present invention. DETAILED DESCRIPTION

[0046] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0047] like Figure 1As shown in the figure, a cross-view target geolocation method based on view generation is presented. The whole process mainly includes the following steps: 1. Constructing drone-satellite image pairs; 2. Polar coordinate conversion; 3. Constructing the conditional generative adversarial network Pix2PixGAN; 4. Constructing the matching data required for the geolocation model; 5. Constructing the geolocation matching model ConvNeXtV2; 6. Loading the trained ConvNeXtV2 model for inference and prediction.

[0048] Step 1, build drone-satellite image pairs:

[0049] We collected images from multiple public datasets (including CVUSA, CUACT, and University-1652), totaling 3,000 pairs of images. Each pair of images contained one satellite-view image and one drone-view image of the same geographic location. We divided the 3,000 drone-view images into training and validation sets according to the ratio of 8:2, that is, the training set and validation set contained 2,400 and 600 drone-view images, respectively.

[0050] Step 2, polar coordinate conversion:

[0051] For each image in the image pair in step 1, polar coordinate transformation is used, with the center point of the image as the origin of the polar coordinate system, and each pixel of the image is remapped to the polar coordinate system to generate the corresponding panoramic view image. The panoramic view image generated by the drone view image is uniformly recorded as D, and the panoramic view image generated by the satellite view image is uniformly recorded as S. D is constructed as a drone image query library, recorded as Q. The polar coordinate transformation formula is as follows:

[0052]

[0053] in and Represent the pixel coordinates of the original view image and the panoramic view image respectively, and their image sizes correspond to (W s ,H s ) and (W ps ,H ps ), the image size in the experiment was set to 512×512. Through polar coordinate transformation, the scene distribution and object arrangement of the original image are basically restored. For satellite images, polar coordinate transformation can map the parallel characteristics in the top view into a circular distribution, retaining the global information of the image; for drone images, polar coordinate transformation adjusts the trapezoidal distortion to a radial distribution, which is close to the characteristics of satellite images. However, this alone cannot completely eliminate the domain gap between cross-viewpoint views, and more detailed features cannot be restored in this way.

[0054] Step 3: Build the conditional generative adversarial network Pix2Pix GAN:

[0055] Construct a conditional generative adversarial network Pix2Pix GAN, take S and D in step 2 as input at the same time, output the pseudo drone perspective image generated by S, obtain the trained conditional generative adversarial network Pix2Pix GAN, and construct the pseudo drone perspective image generated by S as a satellite image retrieval library, denoted as R.

[0056] The conditional generative adversarial network Pix2Pix GAN includes a generator and a discriminator. The generator network model adopts the U-Net architecture, which contains an encoder and a decoder, which are used for downsampling feature extraction and upsampling recovery, respectively. Specifically, given a satellite view image The input is transformed into Then the feature map is obtained by downsampling through the first encoding block. Similarly, in each round of downsampling, the size of the feature map is halved and the number of channels is doubled. After the second, third, and fourth encoding blocks, the downsampling obtains feature maps X2, X3, and X4 respectively. In the decoder, feature map X4 is concatenated with feature map X3 through the nearest neighbor upsampling layer and then input into the convolution layer to obtain feature map Y4. Y4 is concatenated with X2 through the upsampling layer and then input into the convolution layer to obtain feature map Y3. Similarly, Y2 and Y1 are obtained. Finally, feature map Y1 is convolved with 1×1 to obtain the final output. Y0 is the final pseudo drone perspective image.

[0057] The discriminator network model uses the PatchGAN structure, which consists of multiple layers of convolution operations. It concatenates the input real drone view image with the pseudo drone view image generated by the generator. Then it goes through three convolutional layers for downsampling, each time the feature map channels are doubled and the size is halved, and then goes through padding, 1×1 convolution, BN normalization, LeakyReLU activation function, padding, 1×1 convolution, and finally gets The output matrix is ​​used to judge the authenticity of each area of ​​the pseudo drone perspective image.

[0058] The total loss function formula for training conditional generative adversarial network Pix2Pix GAN is:

[0059]

[0060] L L1 =E[||yG(x)||1]#

[0061] L total =λL L1 +L GAN

[0062] Among them, y is the real target image, G(x) is the image generated by the generator, D(x,y) is the discriminant value of the real image, D(x,G(x)) is the discriminant value of the generated image, L L1 is the L1 loss between the generated image and the target image, L GAN is the adversarial loss between the generator and the discriminator, L total is the final total loss function, and the parameter λ is used to balance the importance of L1 loss and adversarial loss.

[0063] In the process of training the conditional generative adversarial network Pix2Pix GAN, the optimizer uses the Adam method, the learning rates of the generator and discriminator are both 0.0001, the momentum parameters β1 and β2 are set to 0.5 and 0.999 respectively, and the batch size is set to 8.

[0064] Step 4: Build the matching data required for the geolocation model:

[0065] Match Q in step 2 with R in step 3 one by one for training (denoted as T, numbered as T Q1 ,T Q2 ,…T Q2400 and T R1 ,T R2 ,…T R3000 ), the image pairs that constitute the positive sample pairs in training (e.g. T Q1 and T R1 ), and the rest are used as negative sample pairs (e.g. T Q1 and T R2 );

[0066] Step 5: Build the geolocation matching model ConvNeXtV2:

[0067] Using T in step 4, each image is input into the ConvNeXtV2 network in turn, and the feature vector corresponding to each image is output. The InfoNCE contrast loss is used for back propagation to reduce the feature similarity distance between the feature vectors corresponding to the positive sample pairs and expand the distance between the feature vectors corresponding to the negative sample pairs, thus obtaining a trained geolocation matching model, where:

[0068] The ConvNeXtV2 network consists of four stages, each of which stacks several ConvNeXtV2Blocks (the number of blocks in the four stages are 3, 3, 27, and 3 respectively) and one convolutional downsampling block. Each ConvNeXtV2 Block is sequentially connected to the depthwise separable convolution DW-Conv, the pointwise convolution Conv, and the GELU activation function, which are used to extract spatial features, reorganize feature channels, and improve nonlinear expression capabilities, respectively. The downsampling block is implemented using convolution. Except for the first stage downsampling block which uses a 4-fold downsampling rate, the remaining stages all use a 2-fold downsampling rate.

[0069] The initial input image is After each stage, the number of feature channels gradually increases to 128, 256, 512, and 1024, respectively, and the image size is halved at each stage. After 4 stages, we get Finally, a global average pooling GAP and a fully connected layer Linear are used to obtain a feature vector output with a length of 1000.

[0070] The calculation formula of the InfoNCE loss function for training the geolocation matching model is:

[0071]

[0072] Where N represents the size of each batch, A + represents the unique query positive sample in each batch, B - Represents the (N-1) negative samples in each batch. During the training process, the query vector continuously narrows the distance between the positive samples and widens the distance between the query vector and the negative samples, thereby reducing the loss value.

[0073] In the process of training the geolocation matching model, the AdamW optimizer was used for training, the initial learning rate was set to 0.001, the learning rate decay adopted the cosine annealing algorithm, the minimum learning rate was set to 0.0001, the batch size was set to 64, and the maximum number of training epochs was 10.

[0074] Step 6: Load the trained ConvNeXtV2 model for inference prediction:

[0075] Enter the T in T from step 4 R1 ,T R2 ,…T R3000, get the corresponding satellite image features, build a satellite image feature retrieval library, input the drone view image to be tested in the verification set, first undergo polar coordinate conversion, and then pass through the trained ConvNeXtV2 model to get the drone image features, and use the drone image features to perform feature matching and sorting from the satellite image feature retrieval library, match the K satellite image features with the highest feature similarity, and obtain the target location in the drone view image through the target detection model. For the identified target area, cross-check with the satellite remote sensing image corresponding to the matched satellite image features, compare and analyze the spatial geometric relationship and pixel distribution characteristics of the target, recalculate the pixel geometric center of the target in the satellite view image, and realize cross-view target positioning.

[0076] Top-K (denoted as R@K) is used as an evaluation indicator to measure the cross-view geographic target positioning model. R@K is calculated by calculating the proportion of the correct location in the top K positions with the highest probability among the predicted positions. The calculation formula is as follows:

[0077]

[0078] in, Indicates that when the true category y i Exists in the top K categories of the prediction The value is 1 when the value is in the middle, otherwise it is 0. K can be 1, 5, 10 or the top 1%.

[0079] As can be seen from Table 1, compared with other methods (LPN is a local pattern network that uses a feature level division strategy to learn contextual information, SAIG-D is a simple image geolocation network based on an attention mechanism, and PCL is a method that combines conditional GAN ​​with a ResNet network), the method of the present invention has improved the matching accuracy and average precision of the first 1, 5, 10, and 1%, among which R@1 reached 86.31%, R@5 and R@10 were 95.30% and 96.72% respectively, and further improved to 96.88% in the retrieval of R@1%. These results prove the feasibility and efficiency of the strategy of first narrowing the cross-domain perspective gap and then matching with a powerful feature extraction network in improving the adaptability of the model to the cross-perspective target geolocation task.

[0080] Table 1. UAV image query results

[0081]

[0082] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0083] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A cross-view target geographic positioning method based on view generation, characterized in that: The following steps are involved: Step 1: Obtain satellite and drone cross-view geolocation images of the area to be queried, and construct image pairs that match each other from different viewpoints. The image pairs include drone view images and satellite view images of the same geographic location, and divide the drone view images into training sets and validation sets. Step 2: For each image in the image pair in step 1, use polar coordinate transformation, take the center point of the image as the origin of the polar coordinate system, remap each pixel of the image to the polar coordinate system, and generate the corresponding panoramic view image, where the panoramic view image generated by the drone view image is uniformly denoted as D, and the panoramic view image generated by the satellite view image is uniformly denoted as S, and D is constructed as a drone image query library, denoted as Q; Step 3: Construct a conditional generative adversarial network Pix2Pix GAN, take S and D in step 2 as input at the same time, output the pseudo drone view image generated by S, train the conditional generative adversarial network Pix2Pix GAN, and construct the pseudo drone view image generated by S as a satellite image retrieval library, denoted as R; Step 4: Match Q in step 2 with R in step 3 one by one, and use the successfully matched image pairs as positive sample pairs, and the rest as negative sample pairs; Step 5: Build a geolocation matching model based on the pre-trained ConvNeXtV2 network. Take S and D in step 2 as input at the same time. Each image is input to the ConvNeXtV2 network in turn. Output the feature vector corresponding to each image. Use InfoNCE contrast loss for back propagation to reduce the feature similarity distance between the feature vectors corresponding to the positive sample pairs and expand the distance between the feature vectors corresponding to the negative sample pairs. Train the geolocation matching model and use the satellite image features of the positive sample pairs to build a satellite image feature retrieval library. Step 6: Use the features of the drone images to be tested in the validation set to perform feature matching and sorting with the satellite image feature retrieval library, match the K satellite image features with the highest feature similarity, and obtain the location of the target in the drone's view image through the target detection model. For the identified target area, cross-check the satellite remote sensing image corresponding to the matched satellite image features, compare and analyze the spatial geometric relationship and pixel distribution characteristics of the target, recalculate the pixel geometric center of the target in the satellite view image, and realize cross-view target positioning.

2. The cross-view target geographic positioning method according to claim 1, characterized in that: Step 2: For each image in the image pair in step 1, polar coordinate transformation is used, with the center point of the image as the origin of the polar coordinate system, and each pixel of the image is remapped to the polar coordinate system to generate the corresponding panoramic view image. The panoramic view image generated by the drone view image is uniformly recorded as D, and the panoramic view image generated by the satellite view image is uniformly recorded as S. D is constructed as a drone image query library, recorded as Q, where the conversion formula is as follows: in and Represent the pixel coordinates of the original view image and the panoramic view image respectively, and their image sizes correspond to (W s , H s ) and (W ps , H ps ).

3. The cross-view target geographic positioning method according to claim 1, characterized in that: Step 3: Construct a conditional generative adversarial network Pix2Pix GAN, take S and D in step 2 as input at the same time, output the pseudo drone view image generated by S, train the conditional generative adversarial network Pix2Pix GAN, and construct the pseudo drone view image generated by S as a satellite image retrieval library, denoted as R, where: The conditional generative adversarial network Pix2Pix GAN includes a generator and a discriminator. The generator network model adopts the U-Net architecture, which contains an encoder and a decoder, which are used for downsampling feature extraction and upsampling recovery, respectively. Specifically, given a satellite view image In the encoder, the input is transformed into Then the feature map is obtained by downsampling through the first encoding block. Similarly, in each round of downsampling, the size of the feature map is halved and the number of channels is doubled. After the second, third, and fourth encoding blocks, the downsampling obtains feature maps X2, X3, and X4 respectively. In the decoder, feature map X4 is concatenated with feature map X3 through the nearest neighbor upsampling layer and then input into the convolution layer to obtain feature map Y4. Y4 is concatenated with X2 through the upsampling layer and then input into the convolution layer to obtain feature map Y3. Similarly, Y2 and Y1 are obtained. Finally, feature map Y1 is convolved with 1×1 to obtain the final output. Y0 is the final pseudo drone perspective image; The discriminator network model uses the PatchGAN structure, which consists of multiple layers of convolution operations. It concatenates the input real drone view image with the pseudo drone view image generated by the generator. Then it goes through three convolutional layers for downsampling, each time the feature map channel is doubled and the size is halved, and then goes through padding, 1×1 convolution, BN normalization, LeakyReLU activation function, padding, 1×1 convolution, and finally the output matrix is ​​obtained. Used to judge the authenticity of each area of ​​the pseudo-UAV perspective image.

4. The cross-view target geographic positioning method according to claim 3, characterized in that: Step 3: Construct a conditional generative adversarial network Pix2Pix GAN, take S and D in step 2 as input at the same time, output the pseudo drone perspective image generated by S, train the conditional generative adversarial network Pix2Pix GAN, and construct the pseudo drone perspective image generated by S as a satellite image retrieval library, denoted as R. The total loss function formula of the training conditional generative adversarial network Pix2Pix GAN is: L L1 =E[||y-G(x)||1]# THE total =λL L1 +L GAN Among them, y is the real target image, G(x) is the image generated by the generator, D(x, y) is the discriminant value of the real image, D(x, G(x)) is the discriminant value of the generated image, and L L1 is the L1 loss between the generated image and the target image, L GAN is the adversarial loss between the generator and the discriminator, L total is the final total loss function, and the parameter λ is used to balance the importance of L1 loss and adversarial loss.

5. The cross-view target geographic positioning method according to claim 4, characterized in that: Step 3: Construct a conditional generative adversarial network Pix2Pix GAN, take S and D in step 2 as input at the same time, output the pseudo drone perspective image generated by S, train the conditional generative adversarial network Pix2Pix GAN, and construct the pseudo drone perspective image generated by S as a satellite image retrieval library, denoted as R. In the process of training the conditional generative adversarial network Pix2Pix GAN: The optimizer uses the Adam method, the learning rates of the generator and discriminator are both 0.0001, the momentum parameters β1 and β2 are set to 0.5 and 0.999 respectively, and the batch size is set to 8.

6. The cross-view target geographic positioning method according to claim 1, characterized in that: Step 5: Build a geolocation matching model based on the pre-trained ConvNeXtV2 network. Take S and D in step 2 as input at the same time. Each image is input to the ConvNeXtV2 network in turn. Output the feature vector corresponding to each image. Use InfoNCE contrast loss for back propagation to reduce the feature similarity distance between the feature vectors corresponding to the positive sample pairs and expand the distance between the feature vectors corresponding to the negative sample pairs. Train the geolocation matching model and use the satellite image features of the positive sample pairs to build a satellite image feature retrieval library, where: The ConvNeXtV2 network consists of 4 stages. In each stage, several ConvNeXtV2Blocks and 1 convolutional downsampling block are stacked. The number of ConvNeXtV2 Blocks in the 4 stages is 3, 3, 27, and 3 respectively. Each ConvNeXtV2 Block is connected to the depth-separable convolution DW-Con in turn. v , point-by-point convolution Conv and GELU activation functions are used to extract spatial features, reorganize feature channels and improve nonlinear expression capabilities respectively; the downsampling block is implemented using convolution. Except for the first stage downsampling block using a 4-fold downsampling rate, the remaining stages all use a 2-fold downsampling rate. The initial input image is After each stage, the number of feature channels gradually increases to 128, 256, 512, and 1024, respectively, and the image size is halved at each stage. After 4 stages, we get Finally, a global average pooling GAP and a fully connected layer Linear are used to obtain a feature vector output with a length of 1000.

7. The cross-view target geographic positioning method according to claim 6, characterized in that: Step 5: Construct a geolocation matching model based on the pre-trained ConvNeXtV2 network. Take S and D in step 2 as input at the same time. Each image is input through the ConvNeXtV2 network in turn. Output the feature vector corresponding to each image. Use InfoNCE contrast loss for back propagation to reduce the feature similarity distance between the feature vectors corresponding to the positive sample pairs and expand the distance between the feature vectors corresponding to the negative sample pairs. Train the geolocation matching model and construct a satellite image feature retrieval library with the satellite image features of the positive sample pairs. The calculation formula of the InfoNCE loss function for training the geolocation matching model is: Where N represents the size of each batch, A + represents the unique query positive sample in each batch, B - Represents the (N-1) negative samples in each batch. During the training process, the query vector continuously narrows the distance between the positive samples and widens the distance between the query vector and the negative samples, thereby reducing the loss value.

8. The cross-view target geographic positioning method according to claim 6, characterized in that: Step 5: Build a geolocation matching model based on the pre-trained ConvNeXtV2 network. Take S and D in step 2 as input at the same time. Each image is input through the ConvNeXtV2 network in turn. Output the feature vector corresponding to each image. Use InfoNCE contrast loss for back propagation to reduce the feature similarity distance between the feature vectors corresponding to the positive sample pairs and expand the distance between the feature vectors corresponding to the negative sample pairs. Train the geolocation matching model and use the satellite image features of the positive sample pairs to build a satellite image feature retrieval library. In the process of training the geolocation matching model: The AdamW optimizer was used for training, the initial learning rate was set to 0.001, the learning rate decay adopted the cosine annealing algorithm, the minimum learning rate was set to 0.0001, the batch size was set to 64, and the maximum number of training epochs was 10.

9. The cross-view target geographic positioning method according to claim 1, characterized in that: The target detection model uses the YOLOv5 network.

10. A cross-view target geographic positioning system, characterized in that: Based on the cross-perspective target geographic positioning method described in any one of claims 1 to 9, cross-perspective target geographic positioning is achieved, which is divided into 6 modules and executes steps 1 to 6 respectively.

Citation Information

Patent Citations

  • Cross-view-angle geographic positioning method based on unmanned aerial vehicle-satellite

    CN113361508A

  • Cross-view-angle image real-time matching geographic positioning method and system based on deep learning

    CN114241464A