Unified visual geographic positioning method and device based on spatial query learning

Through a method based on spatial query learning, a training model is constructed to directly extract features from the image and predict GPS coordinates, which solves the problem of high demand for storage space and computing resources in the prior art, and achieves efficient visual geolocation.

CN119807457BActive Publication Date: 2025-06-06ZHEJIANG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510283325.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-06
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

Existing visual geolocation methods rely on image vector libraries and GPS reference sets, resulting in high demands for storage space and computing resources and are difficult to work effectively in storage-sensitive and low-cost scenarios.

Method used

Using a unified visual geolocation method based on spatial query learning, a training model including image encoder, self-attention layer, cross-attention layer, frozen coordinate encoder and regressor is constructed, and features are directly extracted from the image and predicted GPS coordinates, avoiding dependence on image vector library and GPS reference set.

Benefits of technology

It realizes that without relying on the image vector library and GPS reference set, it achieves better GPS coordinate prediction accuracy while reducing storage requirements and improving inference speed, and is suitable for storage-sensitive and low-cost scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119807457B_ABST
    Figure CN119807457B_ABST
Patent Text Reader

Abstract

The present invention discloses a unified visual geographic positioning method and device based on spatial query learning. The method cross-attentions feature-enhanced spatial query and image feature sequence so as to extract fused features strongly correlated with space from the image feature sequence, and then obtains predicted GPS coordinates through a regressor. The loss function training model constructed by predicting GPS coordinates and converted GPS coordinates is suitable for storage-sensitive and low-cost scenarios, and does not rely on picture vector libraries and GPS reference sets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image retrieval, and in particular relates to a unified visual geographic positioning method and device based on spatial query learning. Background Art

[0002] Visual geolocation is a very challenging and valuable research task. Its core goal is to accurately infer the specific location where a given image was taken, that is, the corresponding GPS coordinates.

[0003] In this field, current research work can be summarized into three different methods. The first is the classification-based method, which divides the earth's surface into multiple discrete area categories, classifies the input image through training models, and then determines its approximate shooting location. The second is the retrieval-based method, which relies on a large image database and calculates the similarity between the query image and the image in the database to find the matching shooting location information. The retrieval-augmented generation (RAG) method is a relatively new idea. It combines the advantages of retrieval and generation. On the basis of retrieving relevant information, it uses the generation model to further optimize and generate more accurate location predictions, providing new research directions and solutions for visual geolocation.

[0004] Specifically, the entire geographic space is divided into fixed grids based on the classification method, and each image is classified into a specific grid. When predicting the image coordinates, the algorithm needs to provide a set of GPS reference points, and the classification determines which candidate point the image belongs to with a higher probability. However, the obvious flaw of this method is that it is highly dependent on the construction of the GPS reference set.

[0005] Retrieval-based methods transform the image localization problem into an image-to-image retrieval task, relying on a database containing GPS information and image features. First, the feature network extracts features from the query image; then, the top K images that are most similar to it are retrieved from the database; the coordinates of the image closest to the current image are used as the GPS prediction value of the current image, or an additional geometric verification stage is required to use the GPS information of these reference images and their overlapping areas with the query image to further infer the GPS of the query image.

[0006] The invention patent application with publication number CN110347854A discloses an image retrieval method based on target positioning. First, a training library similar to the image retrieval database is selected for manual annotation, and the location and size information of the target area required by the database is recorded. The SSD target detection model is trained with the annotated training library to obtain an SSD model that can detect the target area. Then, the feature vectors of the query image and the test image are extracted according to the obtained SSD target detection model. Finally, the cosine distance between the feature vector of the test image and the feature vector of the query image is calculated to measure the similarity between the query image and the test image. The minimum similarity score is taken as the final score of the test image, and the scores of all images in the test library are ranked to obtain the retrieval results. Although the method provided by the above patent application performs well in accuracy, the storage of large-scale image vectors and querying the test library place high demands on storage space and computing resources.

[0007] In recent years, methods based on generative large models have achieved state-of-the-art performance in this task. Such methods use the retrieval-augmented generation (RAG) process to fully leverage the powerful reasoning and generalization capabilities of large-scale multimodal models (LMMs). Specifically, they embed the retrieved GPS coordinates as reference information into the input prompt of the LMM to generate more accurate prediction results. However, since such methods are essentially based on retrieval technology, they still rely on a global database of image vectors. Summary of the invention

[0008] The present invention provides a unified visual geographic positioning method based on spatial query learning, which does not rely on a picture vector library and a GPS reference set, reduces storage and improves reasoning speed while achieving better GPS coordinate prediction accuracy.

[0009] A specific embodiment of the present invention provides a unified visual geographic positioning method based on spatial query learning, including:

[0010] Obtain multiple pictures and corresponding real GPS coordinates, and use the multiple pictures as a training sample set;

[0011] Constructing a training model, the training model includes an image encoder, a self-attention layer, a cross-attention layer, a frozen coordinate encoder and a regressor, slicing each sample into a sequence, extracting features from the sequence through the image encoder to obtain an image feature sequence, enhancing features of the spatial query through the self-attention layer, interactively fusing the feature-enhanced spatial query and the image feature sequence through the cross-attention layer to obtain a fused feature, converting the real GPS coordinates through the coordinate encoder and encoding them to obtain a coordinate code, and inputting the fused feature into the regressor to obtain a predicted GPS coordinate;

[0012] Constructing a loss function, wherein the loss function is a mean square error loss function constructed by predicting the GPS coordinates and the converted GPS coordinates;

[0013] Based on the training sample set, the training model is trained by the loss function to obtain a unified visual geographic positioning model. When applied, the query image is input into the unified visual geographic positioning model to obtain the first predicted GPS coordinates and fusion features.

[0014] Preferably, the loss function also includes a normalized temperature-scale cross entropy function constructed by dimensionally aligned fusion features and coordinate encoding, and the coordinate encoder is turned on at the same time.

[0015] Preferably, average pooling or maximum pooling is performed on the fused features so as to adjust the dimension of the fused features so that the fused features are aligned with the dimension of the coordinate encoding.

[0016] Preferably, obtaining the second predicted GPS coordinates by using an image retrieval method based on the fusion features includes:

[0017] Based on the retrieval method, the first K pictures most similar to the fusion feature are retrieved from a database containing GPS coordinates and image features, and the GPS coordinates of the most similar pictures are used as the second predicted GPS coordinates of the query picture, or the GPS coordinates of the first K pictures and the overlapping area of ​​the first K pictures and the query picture are used to infer the second predicted GPS coordinates of the query picture.

[0018] Preferably, obtaining the third predicted GPS coordinates by using a coordinate classification method includes:

[0019] The query image and GPS reference point set are input into the unified visual geographic positioning model to obtain dimensionally aligned fusion features and coordinate coding sets respectively. The coordinate coding with the highest similarity to the fusion feature is screened out from the coordinate coding set, and the GPS reference point corresponding to the screened coordinate coding is used as the third predicted GPS coordinate.

[0020] Preferably, the feature-enhanced spatial query and the image feature sequence are interactively fused through a cross-attention layer to obtain a fused feature, including:

[0021] The spatial query after feature enhancement is linearly transformed to obtain the query vector of the cross attention layer, and the image feature sequence is linearly transformed to obtain the key vector and value vector of the cross attention layer respectively;

[0022] The self-attention score matrix of the criss-cross attention layer is obtained based on the query vector and key vector of the criss-cross attention layer through the scaled dot product attention mechanism;

[0023] Multiply the self-attention score matrix of the cross-attention layer by the value vector to get the fused features.

[0024] Preferably, feature enhancement is performed on the spatial query through a self-attention layer, including:

[0025] The initial value of the spatial query is obtained by random initialization, and the spatial query is linearly transformed to obtain the query vector, key vector and value vector of the self-attention layer respectively;

[0026] The self-attention score matrix is ​​obtained based on the query vector and key vector of the self-attention layer through the scaled dot product attention mechanism;

[0027] Multiplying the self-attention score matrix with the value vector yields the feature-enhanced spatial query.

[0028] Preferably, each sample is sliced ​​and converted into a sequence, including:

[0029] Each sample is cut into non-overlapping image blocks of fixed size, each image block is converted into a one-dimensional vector, each one-dimensional vector is mapped through a linear layer to obtain a multi-dimensional vector, and the multi-dimensional vectors corresponding to the image blocks contained in each sample are combined into a sequence.

[0030] Preferably, the coordinate code is obtained by converting the real GPS coordinates through a coordinate encoder and then encoding them, including:

[0031] The coordinate encoder includes equal earth projection, random Fourier features and feedforward multi-layer perceptron;

[0032] The real GPS coordinates are converted by equal earth projection, and the GPS coding features of different granularities are obtained based on the converted GPS coordinates by adjusting the frequency of random Fourier features. The GPS coding features of different granularities are extracted by feedforward multi-layer perceptron and then added to obtain coordinate coding.

[0033] The present invention also provides a unified visual geographic positioning device based on spatial query learning, comprising: a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used for the unified visual geographic positioning method based on spatial query learning.

[0034] Compared with the prior art, the present invention has the following beneficial effects:

[0035] The present invention cross-attentions the feature-enhanced spatial query and the image feature sequence to extract fused features strongly correlated with space from the image feature sequence, and then obtains the predicted GPS coordinates through a regressor. The loss function training model is constructed by predicting the GPS coordinates and the converted GPS coordinates, so that the trained model has good GPS coordinate prediction accuracy, and does not rely on the picture vector library and the GPS reference set during the prediction process, and is suitable for storage-sensitive and low-cost scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 A flow chart of a unified visual geographic positioning method based on spatial query learning provided by a specific embodiment of the present invention;

[0037] Figure 2 A block diagram of a unified visual geographic positioning method based on spatial query learning provided in a specific embodiment of the present invention. DETAILED DESCRIPTION

[0038] In order to solve the technical problem that the existing image retrieval depends on the image vector library and the GPS reference set and has high requirements for storage space and computing resources, a specific embodiment of the present invention provides a unified visual geographic positioning method based on spatial query learning. The method can train a unified visual geographic positioning model based on the mean square error loss function. The model can directly obtain predicted GPS coordinates without relying on the image vector library and the GPS reference set. The specific embodiment of the present invention also provides a normalized temperature-scale cross entropy function training model, so that the trained model is applicable to multiple training scenarios, and has strong tolerance for the amount of training samples during the training process. When the amount of training samples is small and the storage space is limited, a mean square error loss function training model constructed based on the predicted GPS value obtained by the regressor and the real GPS coordinates can be used. If the amount of training samples is sufficient and the storage space is relatively abundant, a normalized temperature-scale cross entropy (NT-Xent) loss function training model is constructed through dimensionally aligned fusion features and coordinate encoding, or a NT-Xent loss function and a mean square error loss function collaborative training model is constructed, so that the trained model has higher accuracy. The trained model is applicable to a variety of scenarios such as classification, retrieval and regression, and has rich applicable scenarios.

[0039] The specific embodiment of the present invention provides a unified visual geographic positioning method based on spatial query learning, such as Figure 1 and Figure 2 As shown, including:

[0040] S1. Constructing a training sample set based on the obtained pictures and the corresponding real GPS coordinates: obtaining multiple pictures and the corresponding real GPS coordinates, and using the multiple pictures as the training sample set.

[0041] S2. Construct a training model. The training model provided in the specific embodiment of the present invention includes an image encoder, a self-attention layer, a cross-attention layer, a coordinate encoder and a regressor.

[0042] S21. In a specific embodiment of the present invention, each sample is sliced ​​into pieces and then expanded into a one-dimensional vector, the one-dimensional vector is mapped to obtain a multi-dimensional vector, and the multi-dimensional vectors are combined to obtain a sequence corresponding to each sample.

[0043] Specifically, the present invention divides each sample into non-overlapping image patches of fixed size. Assuming that the input image size is H×W×C (height, width and number of channels), and the size of each patch is P×P, then after division, Each image patch has a shape of P×P×C, which is flattened into a one-dimensional vector , the one-dimensional vector corresponding to each slice is mapped to a high-dimensional feature space through a shared linear layer. The result of the projection is a d-dimensional vector, which represents the features of each image block. Each sample can be represented as a sequence S consisting of n d-dimensional vectors (similar to the word embedding sequence in natural language).

[0044] S22, in a specific embodiment of the present invention, the image encoder extracts features from the sequence to obtain an image feature sequence , where n is the length of the sequence, i.e., the number of blocks of the image, and h is the dimension, i.e., the length of the image feature. The image encoder can adopt any model for extracting image features, such as Transformer, ResNet, VGG, etc. In a specific embodiment, the image encoder adopted in this embodiment is a Transformer model, and a Patchify module is placed in front of the Transformer model.

[0045] S23. In a specific embodiment of the present invention, a randomly initialized spatial query is subjected to feature enhancement through a self-attention layer to facilitate the aggregation of key features. The feature-enhanced spatial query and the image feature sequence are interactively fused through a cross-attention layer to obtain fused features, so as to more specifically search for information related to the feature-enhanced spatial query in the image feature sequence, that is, to accurately extract image features related to geographic positioning, remove redundant information, and improve the performance of the model in visual geographic positioning tasks.

[0046] Specifically, the method for enhancing the features of spatial queries by using a self-attention layer provided in a specific embodiment of the present invention includes: ,in n qis the number of spatial queries, d q The dimension is obtained by random initialization, and the query vector of the self-attention layer is obtained by linearly changing the spatial query as follows. Q 1. Key vector K 1 and value vector V 1: ,in , respectively, are the query, key, and value projection matrices of the self-attention mechanism, , h is the feature dimension of the hidden layer.

[0047] The specific embodiment of the present invention obtains the self-attention score matrix A1 based on the query vector and the key vector of the self-attention layer by scaling the dot product attention mechanism: ,in, d k1 is the key vector K Dimension of 1.

[0048] The specific embodiment of the present invention multiplies the self-attention score matrix by the value vector to obtain the feature-enhanced spatial query Z1: .

[0049] Specifically, the method provided in the specific embodiment of the present invention interactively fuses the feature-enhanced spatial query and the image feature sequence through a cross attention layer to obtain a fused feature, including:

[0050] The spatial query after feature enhancement is linearly transformed to obtain the query vector of the cross attention layer Q 2. Perform linear transformation on the image feature sequence to obtain the key vectors of the cross attention layer K 2 and value vector V 2: ,in, Query, key, and value projection matrices for the criss-cross attention mechanism.

[0051] A specific embodiment of the present invention obtains the self-attention score matrix of the cross attention layer based on the query vector and the key vector of the cross attention layer by scaling the dot product attention mechanism: ,in, d k2 is the key vector K Dimension 2.

[0052] The specific embodiment of the present invention multiplies the self-attention score matrix of the cross attention layer by the value vector to obtain the fused feature Z2: .

[0053] S24. In a specific embodiment of the present invention, the real GPS coordinates are encoded by a coordinate encoder to obtain a coordinate code, wherein the coordinate encoder includes an equal-earth projection, a random Fourier feature and a feedforward multi-layer perceptron; the real GPS coordinates are converted by the equal-earth projection, and GPS coding information of different granularities is obtained based on the converted GPS coordinates by adjusting the frequency of the random Fourier feature, and the GPS coding information of different granularities is respectively subjected to feature extraction by a feedforward multi-layer perceptron and then added to obtain the coordinate code.

[0054] The specific embodiment of the present invention refers to the approach of the paper "GeoCLIP: Clip-Inspired Alignment between Locations and Images for Effective Worldwide Geo-localization", and adopts different components in the encoder, namely, using equal earth projection (EEP) to represent GPS coordinates, encoding positions through random Fourier features, and enforcing hierarchical representation at different scales. It is noted that the specific construction of the GPS feature extraction module does not affect the innovativeness of the present invention, and its specific construction method can be replaced by other means.

[0055] In order to accurately represent two-dimensional GPS coordinates and minimize the inherent distortion existing in the standard coordinate system (for example, countries close to the poles are over-represented in the traditional longitude and latitude system), the specific embodiment of the present invention refers to the paper GeoCLIP, which adopts the Equal Earth Projection (EEP) to balance the distortion and ensure a more accurate representation of the earth's surface.

[0056] The specific embodiment of the present invention gives G i ∈R2, ∀i∈[1…N], N is the number of training samples, G i is the real GPS coordinate corresponding to the i-th training sample. In the specific embodiment of the present invention, G i Convert to : ; in, is the dimension value of the converted GPS coordinates corresponding to the i-th training sample, is the longitude value of the converted GPS coordinate corresponding to the i-th training sample, is the dimension value of the real GPS coordinate corresponding to the i-th training sample, is the longitude value of the real GPS coordinate corresponding to the i-th training sample. After applying EEP, the longitude value is scaled to the range of −1 to 1, and the latitude value is scaled proportionally.

[0057] In order to capture rich high-frequency details in the converted low-dimensional GPS input, the present invention performs The sine-based position coding technology is applied. Specifically, the specific embodiment of the present invention uses random Fourier features (RFF) to perform position coding to obtain GPS coding features. The formula is as follows: , whose matrix , whose mth row and nth column come from the sampling of normal distribution, that is, .matrix R It is set at the beginning of training and remains constant throughout training.

[0058] In the process of position encoding, the specific embodiment of the present invention obtains multiple high-dimensional representations from coarse-grained to fine-grained by adjusting the frequency σ, so as to capture information of different scales. Since the frequency range in RFF depends on the σ value, the specific embodiment of the present invention proposes an index allocation strategy for selecting the σ value in the position encoder. Specifically, for a given σ value range (i.e., [σmin, σmax]), the σ value of the M-level hierarchy is obtained using the following formula: , through the exponential allocation of σ values, the position encoder (GPSEnc) can effectively process the input two-dimensional coordinates at different resolutions. Therefore, this hierarchical setting enables GPSEnc to focus on capturing the features of a specific location at different scales. GPSEnc forces the model to effectively learn rich spatial information, making it suitable for a wide range of geospatial deep learning tasks. Then, the specific embodiment of the present invention converts the encoded hierarchical features into i Feed-Forward Multi-Layer Perceptron corresponding to the real GPS coordinates f i Each MLP processes the output of RFF independently and adds these outputs element by element to obtain a joint hierarchical representation, i.e. i The coordinate encoding of the real GPS coordinates of the training samples: .

[0059] In a specific embodiment, the specific embodiment of the present invention adjusts the dimension of the fused feature through query merging so that the fused feature is aligned with the dimension of the hierarchical representation.

[0060] Specifically, in order to align the fusion feature Z2 with the coordinate encoding dimension of the output of GPSEnc and to filter useless information, the specific embodiment of the present invention merges the Query in Z2 to obtain the fusion feature after adjusting the dimension. Specifically, the calculation method is as follows: The pooling operation can be averaging (i.e., average pooling), maximizing (i.e., max pooling), or any other operation that can achieve the same effect.

[0061] S25, inputting the fused features into a regressor to obtain predicted GPS coordinates, wherein the regressor is any network module having a dimensionality conversion function. In a specific embodiment, the regressor is a two-layer fully connected layer (MLP).

[0062] S3. Construct a loss function, train the training model through the loss function based on the training sample set to obtain a unified visual geographic positioning model, and when applied, input the query image into the unified visual geographic positioning model to obtain the predicted GPS coordinates, wherein the loss function includes a normalized temperature-scale cross entropy function and / or a mean square error loss function.

[0063] The normalized temperature-scale cross entropy function provided in the specific embodiment of the present invention is constructed by dimensionally aligned fusion features and coordinate encoding, wherein the first i The normalized temperature-scale cross entropy function value corresponding to the training samples for: ,in, For the i The fusion features corresponding to the training samples are For the i The coordinate encoding corresponding to the training samples is L j For the j The coordinate encoding corresponding to the training samples is is the size of the data batch, is a hyperparameter.

[0064] The mean square error loss function provided in the specific embodiment of the present invention is constructed by predicting the GPS coordinates and the converted real GPS coordinates, wherein the first i The mean square error loss function value corresponding to the training samples for: ,in, For the i The predicted GPS coordinates of training samples, No. i The true value of the GPS coordinates of the training samples in the EEP coordinate system, that is, the converted GPS coordinates.

[0065] The unified visual geographic positioning method based on spatial query learning provided by the specific embodiment of the present invention supports multiple training modes. The unified visual geographic positioning model obtained after each training mode can support different reasoning scenarios. The specific embodiment of the present invention uses a pure contrast training mode. In this training mode, the regressor is frozen, that is, the unified visual geographic positioning model obtained by training the training model based on the normalized temperature-scale cross entropy function. Since the regression training is not passed, the model only supports use in classification application scenarios and image retrieval application scenarios. The specific embodiment of the present invention uses a pure regression training mode. In this training mode, the parameters of the coordinate encoder are frozen so that the parameters do not participate in the training. During the process, the unified visual geographic positioning model obtained by training the training model based on the mean square error loss function cannot obtain coordinate encoding and cannot be applied in classification scenarios because no comparative training is performed. Therefore, it can only support use in regression application scenarios and image retrieval application scenarios. The specific embodiment of the present invention can also use regression and comparative training modes. In this training mode, the frozen coordinate encoder is turned on, that is, the parameters of the coordinate encoder are retrained. The unified visual geographic positioning model obtained by training the training model based on the normalized temperature-scale cross entropy function and the mean square error loss function can support use in classification application scenarios, image retrieval application scenarios and regression application scenarios.

[0066] Specifically, a method for image retrieval based on fusion features is used to obtain predicted GPS coordinates in a retrieval scenario, including: a specific embodiment of the present invention inputs a query image into a unified visual geographic positioning model to obtain fusion features, retrieves the top K images most similar to the fusion features from a database containing GPS information and image features based on a retrieval method, uses the GPS coordinates of the most similar image as the GPS coordinates of the query image, or uses the GPS coordinates of the top K images and the overlapping area of ​​the top K images and the query image to infer the GPS coordinates of the query image.

[0067] Specifically, a coordinate classification method is used to obtain predicted GPS coordinates in a classified scenario, including: inputting a query image and a GPS reference point set into a unified visual geographic positioning model, using a trained coordinate encoder and an image encoder to respectively obtain dimensionally aligned fusion features and a coordinate coding set, screening out the coordinate coding with the highest similarity to the fusion feature from the coordinate coding set, and using the GPS reference point corresponding to the screened coordinate coding as the third predicted GPS coordinate.

[0068] A specific embodiment of the present invention also provides a unified visual geographic positioning device based on spatial query learning, including: a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the above-mentioned unified visual geographic positioning method based on spatial query learning.

Claims

1. A unified visual geolocation method based on spatial query learning, characterized in that: include: Obtain multiple pictures and corresponding real GPS coordinates, and use the multiple pictures as a training sample set; Constructing a training model, the training model includes an image encoder, a self-attention layer, a cross-attention layer, a frozen coordinate encoder and a regressor, slicing each sample into a sequence, extracting features from the sequence through the image encoder to obtain an image feature sequence, enhancing features of the spatial query through the self-attention layer, interactively fusing the feature-enhanced spatial query and the image feature sequence through the cross-attention layer to obtain a fused feature, converting the real GPS coordinates through the coordinate encoder and encoding them to obtain a coordinate code, and inputting the fused feature into the regressor to obtain a predicted GPS coordinate; Constructing a loss function, wherein the loss function is a mean square error loss function constructed by predicting the GPS coordinates and the converted GPS coordinates; Based on the training sample set, the training model is trained by the loss function to obtain a unified visual geographic positioning model. When applied, the query image is input into the unified visual geographic positioning model to obtain the first predicted GPS coordinates and fusion features; The second predicted GPS coordinates are obtained by using an image retrieval method based on the fusion features, including: Retrieving the first K pictures most similar to the fusion feature from a database containing GPS coordinates and image features based on a retrieval method, using the GPS coordinates of the most similar pictures as the second predicted GPS coordinates of the query picture, or using the GPS coordinates of the first K pictures and the overlapping area of ​​the first K pictures and the query picture to infer the second predicted GPS coordinates of the query picture; The third predicted GPS coordinates are obtained by using a coordinate classification method, including: The query image and GPS reference point set are input into the unified visual geographic positioning model to obtain dimensionally aligned fusion features and coordinate coding sets respectively. The coordinate coding with the highest similarity to the fusion feature is screened out from the coordinate coding set, and the GPS reference point corresponding to the screened coordinate coding is used as the third predicted GPS coordinate.

2. The unified visual geographic positioning method based on spatial query learning according to claim 1, characterized in that: The loss function also includes a normalized temperature-scale cross entropy function constructed by dimensionally aligned fusion features and coordinate encoding, while turning on the coordinate encoder.

3. The unified visual geographic positioning method based on spatial query learning according to claim 2 is characterized in that: Average pooling or maximum pooling is performed on the fused features to adjust the dimension of the fused features so that the fused features are aligned with the dimension of the coordinate encoding.

4. The unified visual geographic positioning method based on spatial query learning according to claim 1, characterized in that: The feature-enhanced spatial query and image feature sequence are interactively fused through the cross-attention layer to obtain fused features, including: The spatial query after feature enhancement is linearly transformed to obtain the query vector of the cross attention layer, and the image feature sequence is linearly transformed to obtain the key vector and value vector of the cross attention layer respectively; The self-attention score matrix of the criss-cross attention layer is obtained based on the query vector and key vector of the criss-cross attention layer through the scaled dot product attention mechanism; Multiply the self-attention score matrix of the cross-attention layer by the value vector to get the fused features.

5. The unified visual geographic positioning method based on spatial query learning according to claim 1, characterized in that: Feature enhancement of spatial queries through self-attention layers, including: The initial value of the spatial query is obtained by random initialization, and the spatial query is linearly transformed to obtain the query vector, key vector and value vector of the self-attention layer respectively; The self-attention score matrix is ​​obtained based on the query vector and key vector of the self-attention layer through the scaled dot product attention mechanism; Multiplying the self-attention score matrix with the value vector yields the feature-enhanced spatial query.

6. The unified visual geographic positioning method based on spatial query learning according to claim 1, characterized in that: Each sample is sliced ​​and converted into a sequence, including: Each sample is cut into non-overlapping image blocks of fixed size, each image block is converted into a one-dimensional vector, each one-dimensional vector is mapped through a linear layer to obtain a multi-dimensional vector, and the multi-dimensional vectors corresponding to the image blocks contained in each sample are combined into a sequence.

7. The unified visual geographic positioning method based on spatial query learning according to claim 1, characterized in that: The coordinate encoder converts the real GPS coordinates and encodes them to obtain the coordinate code, including: The coordinate encoder includes equal earth projection, random Fourier features and feedforward multi-layer perceptron; The real GPS coordinates are converted by equal earth projection, and the GPS coding features of different granularities are obtained based on the converted GPS coordinates by adjusting the frequency of random Fourier features. The GPS coding features of different granularities are extracted by feedforward multi-layer perceptron and then added to obtain coordinate coding.

8. A unified visual geographic positioning device based on spatial query learning, characterized in that: include: It comprises a memory and one or more processors, wherein the memory stores executable codes, and when the one or more processors execute the executable codes, they are used to implement the unified visual geographic positioning method based on spatial query learning as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image retrieval method based on target positioning

    CN110347854A

  • Remote sensing image visual positioning method based on text guidance

    CN116958829A

  • Radar vision fusion data target detection method and system based on cross attention

    CN117523514A