Multi-source multi-scale remote sensing image positioning method based on vector map database
Through the multi-source, multi-scale remote sensing image positioning method based on vector map database, the feature extraction and matching is used to use the two-stage attention mechanism deep learning model to solve the problems of view angle differences, scale inconsistencies and linear feature extraction in remote sensing image positioning, and efficient and accurate positioning effect is achieved.
Patent Information
- Application Number
- CN202510251061.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-20
AI Technical Summary
The existing remote sensing image positioning methods face problems such as different perspective angles, inconsistent scales, linear feature extraction problems, and high computing resource consumption.
A multi-source, multi-scale remote sensing image positioning method based on vector map database is adopted to construct a multi-scale map image pyramid through regular automation of vector map generation and slicing algorithms, and a deep learning model is used to extract and match features.
It effectively avoids the impact of perspective differences and inconsistencies in scale on positioning accuracy, saves computing resources, overcomes the problem of linear feature extraction, and significantly improves the efficiency and accuracy of feature processing.
Smart Images

Figure CN120182373A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image positioning, and particularly to a multi-source and multi-scale remote sensing image positioning method based on a vector map database. Background Art
[0002] In recent years, the application of remote sensing image geolocation technology has become increasingly widespread, covering multiple fields such as agricultural monitoring, aerial navigation, disaster emergency response, and military reconnaissance. Remote sensing images contain a large amount of geospatial information. However, in practical applications, many remote sensing images lack accurate geolocation parameters or even any location information. This information loss is usually due to interference or loss of positioning signals in a non-cooperative environment. A common remote sensing geolocation method is through image retrieval, retrieving similar images containing geographical locations (e.g., images containing geographical landmarks or coordinates) from a database, thereby providing geographical information corresponding to the query image.
[0003] The positioning process of remote sensing images generally includes the following steps: First, the remote sensing image to be positioned is matched with an image database containing landmark information. During the matching process, similarity measurement is performed through features extracted by traditional manual features or deep learning models. Common similarity measurements include the Euclidean distance and cosine similarity between feature vectors. Finally, based on the matching result, the positioning of the remote sensing image is achieved based on the place name or coordinates of the first returned image. The existing positioning process of remote sensing images is as Figure 1 shown; the key to the whole process lies in the extraction and expression of visual features.
[0004] Existing remote sensing image geolocation methods can be classified into two categories according to their reference types: image-based and map-based. Image-based geolocation usually relies on the matching of a query image with a pre-existing remote sensing image with correct georeference information. Traditionally, matching methods based on manually extracted features (such as SIFT and SURF) are more common. With the breakthrough of deep learning technology in computer vision, more and more research has turned to using deep neural networks, such as convolutional neural networks (CNNs), to learn the matching relationship between query images and satellite images. However, such methods often face multiple challenges such as perspective differences and scale inconsistencies. In addition, satellite images are often affected by factors such as weather changes, large amounts of data information, insufficient accuracy, and poor real-time performance, thus restricting the matching accuracy and positioning effect.
[0005] Compared with image-based localization methods, map-based localization methods have significant advantages in terms of stability and reliability. Since the mapping process may already include feature representation, semantic information, and feature generalization, it can effectively eliminate low-level information, thus providing a more stable and simpler representation than remote sensing image references (fast geolocation based on road networks). Some of these maps are directly generated through semantic segmentation or object detection in deep learning techniques. Specifically, some methods establish feature correspondences between road intersections and remote sensing semantic maps through intelligent interpretation; other methods perform cross-modal matching between images and web map tiles, such as using building outlines and road information from Google Maps or OpenStreetMap (OSM). Although these methods are less affected by environmental and weather conditions, the maps they generate usually contain comprehensive symbols and mapping rules, making it very time-consuming to create map tiles and making it difficult for matching algorithms to extract linear features such as long straight lines from cluttered backgrounds, thus complicating the matching process. Summary of the Invention
[0006] The object of the present invention is to provide a multi-source multi-scale remote sensing image localization method based on a vector map database to solve the challenges existing in current remote sensing image localization, including the problems of satellite image perspective differences, time-consuming map tile creation, and difficulty in extracting linear features. As a semantic map based on a vector database, the basic constituent elements (such as points, lines, and surfaces) of the vector map mainly represent the geometric forms of the geographical space and have sparse characteristics. The present invention effectively avoids the influence of satellite image perspective differences and scale inconsistencies on the positioning accuracy and saves computing resources; at the same time, it overcomes the problem of linear feature extraction and overcomes the limitations in feature extraction of traditional methods.
[0007] To achieve the above object, the present invention provides the following technical solution: A multi-source multi-scale remote sensing image localization method based on a vector map database, comprising the following steps:
[0008] Step S1, realizing the construction of a multi-scale map image pyramid based on regularized automated mapping: According to vector data from different sources, using an algorithm for generating and slicing regularized automated vector maps, combining the attribute information and geometric information in the vector data, and combining different geographical features and preset scale parameters, generating a high-precision vector map, rasterizing the vector map into map images, and further constructing a multi-scale map image pyramid to form a vector map database;
[0009] Step S2, Multi-feature Extraction and Matching of the Dual-stage Attention Mechanism Deep Learning Model: Use the dual-stage attention mechanism deep learning model to perform deep feature extraction and cross-modal alignment on the vector map database and the query remote sensing image, and perform deep similarity quantification between the two modal features of the vector map and the query remote sensing image; the dual-stage attention mechanism deep learning model includes a dual-stage attention module and a deep association matching module; the dual-stage attention module includes an adaptive feature extraction stage and a cross-modal alignment stage;
[0010] Step S3, Retrieval and Location of Remote Sensing Images: Specifically includes the following steps:
[0011] Step S31, Match the query remote sensing image to be located with each map picture in the vector map database. Using the feature vectors extracted by the dual-stage attention mechanism deep learning model, calculate the cosine similarity to obtain a preliminary retrieval result list;
[0012] Step S32, Sort the results according to the relevance of the coordinate information carried by the map pictures in the retrieval results, and place the ones with high relevance in the front; after sorting, the map picture ranked first is the best matching result; combined with the geographical information carried by the map picture, achieve geographical location under the condition of lacking coordinate reference.
[0013] Preferably, in step S1, the rule-based automated vector map generation and slicing algorithm sets different colors according to different elements of geographical features such as roads, water bodies, and buildings, and adjusts the road line width according to geographical types such as main roads and highways. At the same time, the rule-based automated vector map generation and slicing algorithm can decompose the vector map data into multiple levels of pictures to support accurate display at different resolutions.
[0014] Preferably, in step S2, the adaptive feature extraction stage: Use a multi-level feature extraction module and a self-attention module to extract features from the query remote sensing image and the map pictures in the vector map database, and divide the query remote sensing image and the map pictures into several small pieces. Each small piece is mapped into a D-dimensional feature space to generate an embedding matrix. Combine the features obtained by the multi-level feature extraction module to construct the overall feature representations of the query remote sensing image and the vector map database. The dual-stage attention mechanism deep learning model introduces learnable classification tokens and position encodings in the embedding matrix. The encoder captures global correlations layer by layer through multi-head self-attention and a feed-forward neural network, and retains the original features through residual connections.
[0015] Preferably, in the step S2, in the cross-modal alignment stage: through self-attention alignment, semantic interaction between the query remote sensing image and the vector map in two modalities is achieved, the query vectors of the query remote sensing image and the vector map in two modalities are exchanged by using a cross-attention mechanism, and interactive update of the features of the two modalities is realized in each layer of the encoder; the exchanged query, key, and value vectors are recombined and the embedding is updated through the cross-attention mechanism.
[0016] Preferably, in the step S2, in the deep association matching module: the features of the aligned remote sensing image and the vector map are divided into a global feature vector and a local feature token matrix; the local feature token matrix is dimension-reduced through a two-dimensional convolutional layer to obtain a low-dimensional feature representation; a fine-grained similarity matrix is constructed by calculating the cosine similarity, the similarity matrix is input into a multi-layer fully-connected neural network, and the fine-grained information is compressed and fused layer by layer, and finally a global matching score is output; this score is mapped to the range of (0, 1) through a sigmoid activation function.
[0017] Preferably, in the multi-level feature extraction module, the feature extraction process of each layer can be expressed as:
[0018] X i+1 = f(Conv i (X i ))
[0019] where X i+1 is the feature map of the multi-level feature extraction module, X i is the input feature map of the i-th layer, Conv i represents the convolutional operation of the i-th layer, and f is the activation function.
[0020] Preferably, the encoder is composed of multiple stacked encoding blocks, and each encoding block contains a multi-head self-attention layer and a feed-forward neural network;
[0021] The output of the encoding block is expressed as:
[0022] Z (l+1) = MLP(MSA(LayerNorm(Z( l )))) + Z (l)
[0023] where Z (l+1) is the feature map extracted by the self-attention module, MLP is the feed-forward neural network, and MSA represents the multi-head self-attention mechanism, which is specifically as follows:
[0024]
[0025] where D is the self-attention module mapping to a D-dimensional feature space to generate an embedding matrix; Q, K, and V are the query, key, and value matrices of the Transformer;
[0026] After the feature extraction is completed, it enters the cross-modal alignment stage. The specific formula is as follows:
[0027]
[0028] Where Q I , K V , V V and Q V , K I , V I are the query, key, and value vectors for querying remote sensing images and map pictures respectively.
[0029] Preferably, the local feature token matrix is converted into a three-dimensional tensor and reduced in dimension through a two-dimensional convolutional layer to obtain a low-dimensional feature representation suitable for subsequent similarity calculation;
[0030] A fine-grained similarity matrix is constructed by calculating the cosine similarity between the remote sensing image and the map picture. The specific formula is as follows:
[0031]
[0032] Where im_fea(i) and vt_fea(j) represent the i-th and j-th local feature points respectively;
[0033] During the cross-modal training process of retrieving vector maps from remote sensing images, a triplet loss function and a deep association loss function are adopted. The final loss is the sum of the two, that is:
[0034]
[0035] Where L tri is the triplet loss, L dm is the deep association loss function. Therefore, the total loss is the sum of the two; im_p, vt_p, and vt_n represent the remote sensing image sample, the map picture positive sample, and the negative sample respectively; d represents the Euclidean distance between the image and the map picture features, T is the total number of triplets, m is the margin parameter of the margin area; N represents the number of samples, predict i and target i are the i-th predicted value and the true value respectively.
[0036] Preferably, in step S32, the relevance is calculated by calculating the distance between the center points of each map picture in the retrieval results and comparing the distance with a preset threshold to judge the level of relevance.
[0037] Compared with the prior art, the beneficial effects of the present invention are:
[0038] 1. Adaptability and flexibility: The vector map generation and slicing algorithm proposed by the present invention can be compatible with a variety of vector data formats, such as Shapefile, GeoJSON, and KML (Keyhole Markup Language), etc., and has high adaptability and flexibility.
[0039] 2. High data processing efficiency and accuracy: The vector map images generated by the present invention through the rule-based automated vector map generation and slicing algorithm are only composed of sparse lines. Compared with traditional images and map data, it greatly reduces the consumption of computing resources, speeds up data storage and transmission, meets the needs of large-scale data processing, and effectively realizes the rasterization of vector maps, solving the problems of perspective differences and scale inconsistencies based on satellite images in traditional remote sensing image positioning.
[0040] 3. Multi-feature extraction and matching of the dual-stage attention mechanism deep learning model: The dual-stage attention mechanism deep learning model of the present invention realizes in-depth feature extraction, semantic interaction and feature alignment between different modalities through the dual-stage attention module, effectively overcoming the limitations of traditional methods in feature extraction, and at the same time overcoming the problem of difficult linear feature extraction; significantly improving the efficiency and accuracy of feature processing; and the deep association matching module of this model further quantifies the alignment degree and deep similarity of different modality features in the feature space, showing stronger robustness and adaptability in sparse data and complex scenarios.
[0041] 4. Generalization of multiple data sources: The method adopted by the present invention can not only realize the positioning of traditional optical remote sensing images, but also has strong positioning ability for synthetic aperture radar (SAR) data, demonstrating excellent generalization of multiple data sources. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is the flow chart of the existing remote sensing image positioning technology of the present invention;
[0043] Figure 2 is the overall work flow chart of the embodiment of the present invention;
[0044] Figure 3 is the schematic diagram of the positioning method of the embodiment of the present invention;
[0045] Figure 4 is the framework diagram of the dual-stage attention mechanism deep learning model of the embodiment of the present invention;
[0046] Figure 5 is the preliminary retrieval result diagram of the embodiment of the present invention;
[0047] Figure 6 is the positioning result diagram of the embodiment of the present invention;
[0048] Figure 7 Other image data positioning example diagrams of the embodiments of the present invention. Detailed implementation manners
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0050] Please refer to Figure 2-7 , and the following further elaborates the present invention in combination with specific embodiments. A multi-source and multi-scale remote sensing image positioning method based on a vector map database. Please refer to Figure 2 for its overall work flow chart. The method steps are as follows:
[0051] (1) Implement the construction of a multi-scale map picture pyramid based on regularized automated cartography.
[0052] Please refer to Figure 3 As shown, according to vector data from different sources, using a regularized automated vector map generation and slicing algorithm, the attribute information and geometric information in the vector data are combined with different geographical features and preset scale parameters to generate a high-precision vector map, and the vector map is rasterized into a map picture, and further a multi-scale map picture pyramid is constructed to form a vector map database; the regularized automated vector map generation and slicing algorithm can set different colors according to different elements of geographical features such as roads, water bodies, and buildings, and adjust the road line width according to geographical types such as main roads and highways. At the same time, the regularized automated vector map generation and slicing algorithm can decompose the vector map data into multiple levels of pictures to support accurate display of different resolutions.
[0053] (2) Multi-feature extraction and matching of a two-stage attention mechanism deep learning model.
[0054] Please refer to Figure 4As shown, in order to extract the features of query remote sensing images and adapt to the sparsity of vector map features, the present invention establishes a deep learning model with a two-stage attention mechanism; uses the deep learning model with a two-stage attention mechanism to perform deep feature extraction and cross-modal alignment on the vector map database and query remote sensing images, and performs deep similarity quantification between the two modal features of the vector map and query remote sensing images; the deep learning model with a two-stage attention mechanism includes a two-stage attention module and a deep association matching module; the two-stage attention module includes an adaptive feature extraction stage and a cross-modal alignment stage; the adaptive feature extraction stage extracts highly discriminative high-dimensional features from the query remote sensing image and the map image, and the cross-modal alignment stage performs precise spatial alignment on the features of the two different modalities of the query remote sensing image and the vector map; the deep association matching module realizes the deep similarity quantification of the alignment degree between the query remote sensing image and the map image in the feature space; the present invention uses the two-stage attention module to obtain the overall feature representation of the query remote sensing image and the vector map database, and performs semantic interaction and feature alignment between the two modalities of the query remote sensing image and the vector map.
[0055] When extracting features, that is, in the adaptive feature extraction stage, the multi-level feature extraction module and the self-attention module are used to extract features from the query remote sensing image and the map image in the vector map database to obtain the overall feature representation of the remote sensing image and the vector map. In the multi-level feature extraction module, the feature extraction process of each layer can be expressed as:
[0056] X i+1 = f(Conv i (X i ))
[0057] where X i+1 is the feature map of the multi-level feature extraction module, X i is the input feature map of the i-th layer, Conv i represents the convolutional operation of the i-th layer, and f is the activation function.
[0058] To further improve the effect of feature extraction, the query remote sensing image and the map image are segmented into several small blocks, each small block represents a part of the query remote sensing image and the map image, and is mapped into a D-dimensional feature space, thereby generating an embedding matrix. In this process, some features of the query image and the map image will be randomly discarded; subsequently, combined with the features obtained by the multi-level feature extraction module, the overall feature representation of the query remote sensing image and the vector map database is constructed; in order to obtain global image features, a learnable classification token is added to the front end of the embedding matrix by the two-stage attention mechanism deep learning model, which is used to represent the overall features of the query remote sensing image and the vector map database. The learnable classification token is processed through the self-attention mechanism during the encoding process and finally serves as the comprehensive expression of the overall features of the query remote sensing image and the vector map database; in addition, in order to provide the unique position information of each small block of the image patch (including the learnable classification token), the two-stage attention mechanism deep learning model introduces position encoding and superimposes the position encoding on the embedding matrix to generate an input representation with position information. Through the multi-head self-attention and feed-forward neural network of the encoder, the global correlation is captured layer by layer, and the original features are retained through residual connections to ensure the efficient extraction of deep features, thereby realizing fine image understanding.
[0059] The encoder consists of multiple stacked encoding blocks, each encoding block contains a multi-head self-attention layer and a feed-forward neural network. This structure captures the global correlation and hierarchical information of the image layer by layer, thereby realizing more fine-grained image understanding; the input of the encoding block is first normalized by layer normalization to stabilize the data distribution and accelerate convergence; then, the input passes through the multi-head self-attention layer, which is used to calculate the feature correlation between the image patches of each small block and capture the mutual relationship between different regions in the image; in this process, in order to prevent overfitting, the output of some neurons will be randomly discarded; the output of the multi-head self-attention layer and the input retain the original feature information through residual connections, and then are passed to the feed-forward neural network for further processing; the feed-forward neural network consists of two linear layers and an activation function, which maps the feature space to a higher dimension to extract deep features; the output of the encoding block is expressed as:
[0060] Z (l+1) =MLP(MSA(LayerNorm(Z( l ))))+Z (l)
[0061] where Z (l+1) is the feature map extracted by the self-attention module, MLP is the feed-forward neural network, and MSA represents the multi-head self-attention mechanism, specifically as follows:
[0062]
[0063] Among them, D is mapped to a D-dimensional feature space by the self-attention module to generate an embedding matrix; Q, K, and V are the query, key, and value matrices of the Transformer.
[0064] After feature extraction, in order to construct the deep visual embedding correlation between remote sensing images and map pictures, it enters the cross-modal alignment stage; during this process, through self-attention alignment, semantic interaction between the two modalities of query remote sensing images and vector maps is achieved, and the cross-attention mechanism is used to exchange the query vectors of the two modalities of query remote sensing images and vector maps, and the interactive update of the features of the two modalities is realized in each layer of the encoder; the exchanged query, key, and value vectors are recombined, and the embedding is updated through the cross-attention mechanism, so as to obtain a closer feature alignment in the cross-modal space; the specific formula is as follows:
[0065]
[0066] Among them, Q I , K V , V V and Q V , K I , V I are the query, key, and value vectors of the query remote sensing image and the map picture respectively.
[0067] (3) Use the deep association matching module to quantify the deep similarity of the alignment degree between the remote sensing image and the vector map in the feature space.
[0068] The features of the remote sensing image and the vector map after cross-modal alignment processing are respectively divided into a global feature vector and a local feature token matrix. The global feature vector represents the overall image features including the query remote sensing image and the vector map database, and is used to describe the full map information, while the local feature token matrix contains rich local feature details; in order to achieve effective dimensionality reduction of the features while maintaining the spatial structure, the local feature token matrix is converted into a three-dimensional tensor and reduced in dimension through a two-dimensional convolutional layer to obtain a low-dimensional feature representation suitable for subsequent similarity calculation.
[0069] A fine-grained similarity matrix is constructed by calculating the cosine similarity between the query remote sensing image and the map picture to capture the matching degree of the two modalities at the local feature points. The specific formula is as follows:
[0070]
[0071] Among them, im_fea(i) and vt_fea(j) represent the i-th and j-th local feature points respectively; each element of the matrix reflects the similarity of the query remote sensing image and the map picture in the local features.
[0072] To obtain the global similarity metric, the fine-grained similarity matrix is input into a multi-layer fully-connected neural network for layer-by-layer compression, and then the fine-grained similarity information is fused into a global matching score, which is mapped to the range of (0, 1) through the sigmoid activation function and used to measure the global matching degree between the query remote sensing image and the map picture.
[0073] During the cross-modal training process of remote sensing image retrieval of vector maps, the triplet loss function and the deep association loss function are adopted, which are used to enhance the ability to distinguish positive and negative samples and improve the accuracy of predicting the association score respectively. The final loss is the sum of the two, that is:
[0074]
[0075] where L tri is the triplet loss, and L dm is the deep association loss, so the total loss is the sum of the two; im_p, vt_p, and vt_n represent the remote sensing image sample, the positive sample and the negative sample of the map picture respectively; d represents the Euclidean distance between the features of the remote sensing image and the map picture, T is the total number of triplets, and m is the margin parameter of the margin area; N represents the number of samples, and predict i and target i are the i-th predicted value and the true value respectively.
[0076] (4) Match according to the extracted features, and locate the geographical information carried by the top result obtained according to the re-ranking method.
[0077] First, match the remote sensing image to be located with each map picture in the vector map database one by one; extract the feature vectors through the two-stage attention mechanism deep learning model, and use the cosine similarity calculation to obtain the preliminary retrieval result list. For the preliminary retrieval result list, please refer to Figure 5 ; when constructing the vector map database, the vector map at each scale is divided into several tiles according to a fixed step size, so as to ensure that each geographical range has a map picture that is fully covered, that is, a map tile.
[0078] Sort the results according to the relevance of the coordinate information carried by the map pictures in the retrieval results. For the re-ranking results, please refer to Figure 6 ; put the ones with high relevance in the front. The specific calculation method of the relevance is to calculate the distances between the center points of each map picture in the retrieval results and check whether these distances are less than a preset threshold; for each group of map pictures, if the distances between the center points of all map pictures in the group are less than the threshold, then this group is considered a valid match;
[0079] Sort the retrieval results according to the number of pictures in each group, with the group having the largest number ranked first; finally, select the group of map pictures ranked first as the best matching result, and complete the positioning in combination with the geographical information of this group; if the distances between the centers of all map pictures in the retrieval results are greater than the preset threshold, maintain the original sorting, select the map picture ranked first as the best matching result, and complete the positioning in combination with its geographical information.
[0080] Please refer to Figure 7 , which is an example diagram of the application of the method of the present invention to the positioning of other image data in addition to the positioning of traditional optical remote sensing images, including Landsat 8, Sentinel 2, high-resolution series and synthetic aperture radar, demonstrating excellent multi-data-source generalization.
[0081] The vector map pictures generated by the present invention through the rule-based automated vector map generation and slicing algorithm are only composed of sparse lines. Compared with traditional image and map data, it greatly reduces the consumption of computing resources, speeds up data storage and transmission, meets the needs of large-scale data processing, and solves the problems of perspective difference and scale inconsistency based on satellite images in traditional remote sensing image positioning; through the two-stage attention module in the two-stage attention mechanism deep learning model of the present invention, deep extraction of features, semantic interaction and feature alignment between different modalities are realized, effectively overcoming the limitations of traditional methods in feature extraction, and at the same time overcoming the problem of difficult linear feature extraction, significantly improving the efficiency and accuracy of feature processing; and the deep association matching module of this model further quantifies the alignment degree and deep similarity of different modality features in the feature space, showing stronger robustness and adaptability in sparse data and complex scenarios.
[0082] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0083] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A multi-source and multi-scale remote sensing image positioning method based on a vector map database, characterized in that: The following steps are involved: Step S1, realizing multi-scale map image pyramid construction based on regularized automated mapping: according to vector data from different sources, using a regularized automated vector map generation and slicing algorithm, combining attribute information and geometric information in the vector data with different geographic features and preset scale parameters, generating a high-precision vector map, rasterizing the vector map into a map image, further constructing a multi-scale map image pyramid, and forming a vector map database; Step S2, dual-stage attention mechanism deep learning model multi-feature extraction and matching: using a dual-stage attention mechanism deep learning model to perform deep feature extraction and cross-modal alignment on the vector map database and the query remote sensing image, and perform deep similarity quantification between the two modal features of the vector map and the query remote sensing image; The dual-stage attention mechanism deep learning model includes a dual-stage attention module and a deep association matching module; the dual-stage attention module includes an adaptive feature extraction stage and a cross-modal alignment stage; Step S3, remote sensing image retrieval and positioning: specifically includes the following steps: Step S31, matching the query remote sensing image to be located with each map image in the vector map database, and obtaining a preliminary search result list by using cosine similarity calculation using feature vectors extracted by the dual-stage attention mechanism deep learning model; Step S32, sorting the results according to the relevance of the coordinate information carried by the map images in the search results, with the most relevant ones placed in front; after the sorting is completed, the first ranked map image is the best matching result; combined with the geographic information carried by the map image, geographic positioning is achieved in the absence of coordinate reference conditions.
2. The multi-source and multi-scale remote sensing image positioning method based on a vector map database according to claim 1, characterized in that: In step S1, the regularized automatic vector map generation and slicing algorithm sets different colors according to different elements of geographical features such as roads, water bodies, and buildings, and adjusts the road line width according to geographical types such as main roads and highways. At the same time, the regularized automatic vector map generation and slicing algorithm can decompose vector map data into multiple levels of pictures to support accurate display of different resolutions.
3. The multi-source and multi-scale remote sensing image positioning method based on a vector map database according to claim 1, characterized in that: In step S2, the adaptive feature extraction stage: extract features from the query remote sensing image and the map picture in the vector map database through a multi-level feature extraction module and a self-attention module, and divide the query remote sensing image and the map picture into several small blocks, each of which is mapped to a D-dimensional feature space, thereby generating an embedding matrix, combining the features obtained by the multi-level feature extraction module to construct an overall feature representation of the query remote sensing image and the vector map database, the dual-stage attention mechanism deep learning model introduces learnable classification tags and position encodings in the embedding matrix, and the encoder captures global correlation layer by layer through multi-head self-attention and feedforward neural networks, and retains the original features through residual connections.
4. The multi-source and multi-scale remote sensing image positioning method based on a vector map database according to claim 1, characterized in that: In step S2, the cross-modal alignment stage: through self-attention alignment, semantic interaction between the two modalities of querying remote sensing images and vector maps is realized, and the query vectors of the two modalities of querying remote sensing images and vector maps are exchanged using a cross-attention mechanism, and interactive updates of the two modal features are realized at each layer of the encoder; the exchanged query, key and value vectors are recombined and embedded updated through the cross-attention mechanism.
5. The multi-source and multi-scale remote sensing image positioning method based on a vector map database according to claim 1, characterized in that: In the step S2, the deep correlation matching module: divides the features of the aligned remote sensing image and the vector map into a global feature vector and a local feature token matrix; The local feature token matrix is reduced in dimension through a two-dimensional convolutional layer to obtain a low-dimensional feature representation; a fine-grained similarity matrix is constructed by calculating cosine similarity, and the similarity matrix is input into a multi-layer fully connected neural network. The fine-grained information is compressed and fused layer by layer, and finally a global matching score is output; the score is mapped to the (0,1) range through the sigmoid activation function.
6. The multi-source and multi-scale remote sensing image positioning method based on a vector map database according to claim 3, characterized in that: In the multi-level feature extraction module, the feature extraction process of each layer can be expressed as: X i+1 =f(Conv i (X i )) Where X i+1 is the feature map of the multi-level feature extraction module, X i is the input feature map of the i-th layer, Conv i represents the convolution operation of the i-th layer, and f is the activation function.
7. The multi-source and multi-scale remote sensing image positioning method based on a vector map database according to claim 3, characterized in that: The encoder is composed of multiple stacked encoding blocks, each encoding block includes a multi-head self-attention layer and a feedforward neural network; The output of the encoding block is represented as: WITH (l+1) =MLP(MSA(LayerNorm(Z( l ))))+Z (l) Where Z (l+1) It is the feature map extracted by the self-attention module, MLP is a feedforward neural network, and MSA represents the multi-head self-attention mechanism, as follows: Where D is the self-attention module mapped to the D-dimensional feature space to generate the embedding matrix; Q, K, V are the query, key, and value matrices of the Transformer; After feature extraction, the cross-modal alignment phase begins. The specific formula is as follows: Where Q I ,K V ,V V and Q V ,K I ,V I They are the query, key, and value vectors for querying remote sensing images and map images, respectively.
8. The multi-source and multi-scale remote sensing image positioning method based on a vector map database according to any one of claim 5, characterized in that: The local feature token matrix is converted into a three-dimensional tensor and reduced in dimension through a two-dimensional convolutional layer to obtain a low-dimensional feature representation suitable for subsequent similarity calculation; A fine-grained similarity matrix is constructed by calculating the cosine similarity between the remote sensing image and the map image. The specific formula is as follows: Among them, im_fea(i) and vt_fea(j) represent the i-th and j-th local feature points respectively; In the cross-modal training process of remote sensing image retrieval vector map, the triple loss function and the deep association loss function are used, and the final loss is the sum of the two, namely: Where L tri is the triplet loss, L dm is the deep correlation loss function, so the total loss is the sum of the two; im_p, vt_p, vt_n represent remote sensing image samples, map image positive samples and negative samples respectively; d represents the Euclidean distance between remote sensing image and map image features, T is the total number of triplets, m is the interval parameter of the edge area; N represents the number of samples, predict i and target i are the i-th predicted value and the true value respectively.
9. The multi-source and multi-scale remote sensing image positioning method based on a vector map database according to claim 1, characterized in that: In step S32, the correlation is calculated by calculating the distance between the center points of each map image in the search result, and comparing the distance with a preset threshold value to determine the correlation.