Remote sensing image searching and matching method and device based on feature fusion

By preprocessing and establishing feature libraries for test images of drones or heterologous satellites, and estimating the affine transformation matrix using a random sampling consensus algorithm, the problem of how to efficiently and accurately match images in massive remote sensing image data is solved, and an automated and efficient matching process is realized.

CN120451595APending Publication Date: 2025-08-08XI AN JIAOTONG UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510536978.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Faced with massive and multi-source remote sensing image data, how to efficiently and accurately find images in specific areas has become an urgent problem.

Method used

By obtaining test images of drones or heterologous satellites, preprocessing and establishing a feature library, estimating the affine transformation matrix using a random sampling consensus algorithm, combining preliminary and fine matching steps, iteratively performs the matching process, and returns the latitude and longitude coordinates of the matching result.

Benefits of technology

It realizes efficient and automated matching of remote sensing images, improves the accuracy and robustness of matching, supports matching at different rotation angles, and adapts to new data and scenario needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451595A_ABST
    Figure CN120451595A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image searching and matching method and device based on feature fusion, and relates to the technical field of remote sensing, a test image of an unmanned aerial vehicle or a heterogenous satellite is acquired, a feature library is established after preprocessing, and feature extraction and index establishment are performed on each base map in the base map library. And selecting a retrieval strategy according to the size of the test image, retrieving the name of the base image and cutting a related grid image. The method comprises the following steps: converting a test and grid image into gray scale and zooming, extracting feature points to form a set, estimating an affine transformation matrix by using a random sampling consensus algorithm, and preliminarily evaluating the number of matched inner points; if the first preset threshold value is exceeded, executing a fine matching step; and if the number of points in the fine matching step exceeds a second preset threshold value, returning to match the name of the base map and the longitude and latitude of the angular points. And if not, iterating precise matching until an optimal matching result and the longitude and latitude of the angular point are returned. The problem of how to efficiently and accurately find an image of a specific area for massive and multi-source remote sensing image data is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of remote sensing technology, and in particular to a remote sensing image search and matching method and device based on feature fusion. Background Art

[0002] Remote sensing technology, a key means of Earth observation, uses sensors aboard platforms such as satellites and aircraft to obtain information about the Earth's surface from a distance and in a non-contact manner. Remote sensing image processing and analysis aims to extract useful information from these images to meet the needs of various fields, including environmental monitoring, urban planning, disaster assessment, and resource exploration.

[0003] In recent years, with the rapid development of high-resolution imaging technology, the spatial and temporal resolution of remote sensing images has continued to improve, providing a richer information foundation for image processing and analysis. Existing methods for remote sensing image retrieval primarily include content-based retrieval (CBIR) and feature-based retrieval. CBIR extracts low-level image features such as color, texture, and shape for similarity matching; feature-based retrieval, on the other hand, focuses on high-level semantic features extracted using deep learning models. These features are more capable of capturing complex information such as objects and scenes within an image.

[0004] Although remote sensing image retrieval and positioning technology has made significant progress, facing massive, multi-source remote sensing image data, how to efficiently and accurately find images of specific areas has become an urgent problem to be solved. Summary of the Invention

[0005] In the embodiments of the present application, the problem of how to efficiently and accurately find images of a specific area in the face of massive, multi-source remote sensing image data is solved.

[0006] In a first aspect, an embodiment of the present application provides a remote sensing image search and matching method based on feature fusion, comprising: obtaining a first preset number of test images, the test images being sourced from a drone or a foreign satellite; preprocessing the test images, and establishing a feature library based on the preprocessed test images; performing feature extraction and indexing on each base map in the base map library; selecting a retrieval strategy based on the size information of the test image, and retrieving the base map name in the feature library; for each base map name retrieved, cropping a grid image associated with the base map name from the base map library; converting the test image and the grid image into a first grayscale image, and performing scaling processing; extracting feature points from the first grayscale image to form a first feature point set; and selecting inner points from the first feature point set using a random sampling consensus algorithm. , estimate the affine transformation matrix between the test image and the grid image; return the number of matched inliers as a preliminary evaluation, and if the number of matched inliers is greater than a first preset threshold, execute the fine matching step; if the number of inliers obtained in the fine matching step is greater than a second preset threshold, return the name of the base map that successfully matches the detection image, and the latitude and longitude coordinates of the four corner points of the detection image on the base map, and use them as the final evaluation; if the number of inliers obtained in the fine matching step is less than or equal to the second preset threshold, iteratively execute the fine matching step until the first second preset number of retrieved base map names at different rotation angles of the detection image are matched, return the matching result with the largest number of inliers as the best match, and return the latitude and longitude coordinates of the four corner points of the detection image on the base map in the matching result.

[0007] In a possible implementation, the method further includes: if the number of matched inliers is less than or equal to a first preset threshold, directly skipping the fine matching step and continuing to match the next base map name.

[0008] In one possible implementation, the test image is preprocessed and a feature library is established based on the preprocessed test image, including: finding and removing black pixels at the edge of the test image to eliminate the black edge area; rotating the test image after the black edge area is removed to generate image variants with different rotation angles; scaling the test image to the same resolution as the base map in the base map library; counting the length and width of the scaled test image, and drawing a distribution histogram and a quartile map, and identifying areas with dense size distribution by observing the distribution histogram and the quartile map; for areas with dense size distribution, establishing a feature library for every preset pixel.

[0009] In a possible implementation, the feature extraction and indexing of each base map in the base map library include: reading a configuration file to obtain a cropping size list and a step size list; cropping the base map according to the cropping size and step size; when the base map features are large and obvious, the step size is equal to the cropping window size, and each movement covers a new and non-overlapping area; when the base map features are small and dense, the step size is less than the cropping window size, and each movement covers a partially overlapping area; the cropped image data is converted into an array format and sent to a large model for feature extraction; the extracted features are associated with a unique key and stored in the form of character-typical data; wherein the unique key includes the base map name, cropping size, the column coordinate of the upper left corner of the cropping area, and the row coordinate of the upper left corner of the cropping area; when all cropping size processing of a base map is completed, the extracted features are classified according to the cropping size, and the features of different cropping sizes are written into feature files classified by the cropping size; based on the stored feature files, an index is established using the Faiss library for retrieval in the feature library.

[0010] In a possible implementation, the method selects a retrieval strategy based on the size information of the test image and retrieves the base map name in the feature library, including: traversing each size range in the feature library, checking whether the size difference between the test image and the size range is less than a preset size threshold; if the difference between the size of the test image and a certain size range in the feature library meets a preset condition, the retrieval strategy is to select the feature library of the size range for retrieval; if the difference between the size of the test image and multiple size ranges in the feature library all meet the preset condition, the retrieval strategy is to select retrieval in all feature libraries; the preset condition is that the size of the test image is less than a preset size threshold compared with the image of a certain size range in the feature library; the difference is defined as the difference between the size of the test image and the boundary of the size range; according to the selected retrieval strategy, retrieve a base map similar to the test image in the corresponding feature library, and return the base map name.

[0011] In one possible implementation, the precise matching step includes: calculating the homography matrix of the grid image and the base map, and multiplying the homography matrix with the affine transformation matrix to obtain a transformation matrix; using the transformation matrix to determine the position of the cropping area on the base map, and adding a preset filling area around the position to form a cropping frame; performing a cropping operation on the base map according to the cropping frame to obtain a cropped image; converting both the cropped image and the test image into a second grayscale image, and performing scaling processing; extracting feature points in the second grayscale image to form a second feature point set; using a random sampling consensus algorithm to select inliers from the second feature point set to estimate the homography matrix between the cropped image and the test image.

[0012] In a second aspect, an embodiment of the present application provides a remote sensing image search and matching device based on feature fusion, comprising: an acquisition module for acquiring a first preset number of test images, the test images being sourced from a drone or a foreign satellite; a preprocessing module for preprocessing the test images and establishing a feature library based on the preprocessed test images; a feature extraction module for extracting features and establishing an index for each base map in the base map library; a retrieval module for selecting a retrieval strategy based on the size information of the test image and retrieving the base map name in the feature library; a cropping module for cropping a grid image associated with each base map name retrieved from the base map library; a conversion module for converting the test image and the grid image into a first grayscale image and performing scaling processing; a feature point extraction module for extracting feature points from the first grayscale image to form a first feature point set; and an estimation module for using A random sampling consensus algorithm is used to select inliers from the first feature point set to estimate the affine transformation matrix between the test image and the grid image; a preliminary evaluation module is used to return the number of matched inliers as a preliminary evaluation, and if the number of matched inliers is greater than a first preset threshold, a fine matching step is executed; a fine matching step module is used to return the name of the base map that successfully matches the detection image, and the latitude and longitude coordinates of the four corner points of the detection image on the base map if the number of inliers obtained in the fine matching step is greater than a second preset threshold, and use this as a final evaluation; if the number of inliers obtained in the fine matching step is less than or equal to the second preset threshold, the fine matching step is iteratively executed until the first second preset number of retrieved base map names at different rotation angles of the detection image are matched, and the matching result with the largest number of inliers is returned as the best match, and the latitude and longitude coordinates of the four corner points of the detection image on the base map in the matching result are returned.

[0013] In the third aspect, an embodiment of the present application provides a remote sensing image search and matching server based on feature fusion, comprising a memory and a processor; the memory is used to store computer-executable instructions; the processor is used to execute the computer-executable instructions to implement the method described in the first aspect or any possible implementation method of the first aspect.

[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores executable instructions. When a computer executes the executable instructions, it can implement the method described in the first aspect or any possible implementation method of the first aspect.

[0015] One or more technical solutions provided in the embodiments of this application have at least the following technical effects:

[0016] The present embodiment provides a remote sensing image search and matching method based on feature fusion. By acquiring test images from drones or heterogeneous satellites, it can process remote sensing image data from multiple sources. Retrieval efficiency is further optimized by selecting a retrieval strategy based on the size information of the test image. By preprocessing the test image and establishing a feature library, as well as extracting features and indexing each base map in the base map library, a base map similar to the test image can be quickly located. The affine transformation matrix is estimated using a random sampling consensus algorithm, effectively reducing mismatches and improving matching accuracy. Furthermore, matching accuracy is further ensured through iterative execution of the preliminary evaluation and fine matching steps. The entire matching process is automated, requiring no human intervention, greatly improving work efficiency. Furthermore, by setting preset thresholds, the quality of the matching results can be intelligently judged. Based on the matching results, the latitude and longitude coordinates of the four corner points of the detected image on the base map are returned, providing strong support for subsequent geographic positioning and analysis. The present application supports matching of detection images at different rotation angles, improving matching robustness. Furthermore, the precision matching step is iterated until the first preset number of retrieved basemap names are matched, and the match with the largest number of inliers is returned as the best match, further enhancing the reliability and accuracy of the match. This scalability and iterative nature of the solution allows it to continuously adapt to new data and scenario requirements. It solves the problem of efficiently and accurately finding images of specific areas in the face of massive, multi-source remote sensing imagery data. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments of the present application or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 A flowchart of a remote sensing image search and matching method based on feature fusion provided in an embodiment of the present application;

[0019] Figure 2 A schematic diagram of a remote sensing image search and matching method based on feature fusion provided in an embodiment of the present application;

[0020] Figure 3 A schematic diagram of a remote sensing image search and matching device based on feature fusion provided in an embodiment of the present application;

[0021] Figure 4 A schematic diagram of a remote sensing image search and matching server based on feature fusion provided in an embodiment of the present application. DETAILED DESCRIPTION

[0022] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0023] The following description of some of the technologies involved in the embodiments of this application is provided to facilitate understanding and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for the sake of clarity and conciseness, some descriptions of well-known functions and structures are omitted from the following description.

[0024] The present invention provides a remote sensing image search and matching method based on feature fusion. Figure 1 As shown, the method includes steps S101 to S111. Figure 1 This is only an execution order shown in the embodiment of the present application, and does not represent the only execution order of a remote sensing image search and matching method based on feature fusion. Figure 1 The steps shown may be performed in parallel or reversed.

[0025] S101: Acquire a first preset number of test images, where the test images are from a drone or a satellite of a different source.

[0026] It should be noted that the first preset number can be 600. The test images come from drones or heterogeneous satellites and have different resolutions, including 0.05, 0.67, 0.89, and 2.00. These test images include multi-texture RGB images (red, green, and blue images) taken by drones, which usually have rich texture and color information. They also include grayscale and color images with less texture taken by heterogeneous satellites. The sizes of the test images range from 500 pixels to 10,000 pixels, and they are all heterogeneous images, that is, they come from different imaging devices and platforms.

[0027] S102: Preprocess the test image and establish a feature library based on the preprocessed test image.

[0028] Preprocess the test image and build a feature library based on the preprocessed test image. This includes: finding and removing black pixels from the edges of the test image to eliminate black borders. Rotate the test image after removing the black borders to generate image variants at different rotation angles. Scale the test image to the same resolution as the base image in the base image library. Calculate the length and width of the scaled test image and plot a distribution histogram and quartiles. By observing the distribution histogram and quartiles, identify areas with dense size distribution. For areas with dense size distribution, build a feature library every preset pixel.

[0029] Specifically, there are about 1,000 high-resolution color images in the base map library, all of which have a fine resolution of 0.5 meters and are saved in tif format. Each image has a size of 28,000*19,000 pixels, which not only ensures the clarity and detail of the image, but also makes the total coverage area of the entire base map library reach a vast area of about 80,000 square kilometers. The total amount of these image data is as high as 1.75TB, which fully demonstrates its richness and value as a high-resolution satellite image. Specifically, by detecting the four edges of the image, finding edge points and determining the actual boundaries of the detection image. Check the left, top, right, and bottom edges of the detection image and look for edge points with non-black pixel values to determine the actual size of the detection image. If the four detected edge points are sufficient and the position is correct, further calculate the homography matrix of the detection image to correct the geometric deformation to find and remove the black pixels at the edge of the test image. To reduce the impact of rotation on retrieval accuracy, the test image after removing the black borders is rotated by 0°, 90°, 180°, and 270° to generate image variants with different rotation angles. This process utilizes parallel computing to improve efficiency. Next, a hierarchical management strategy is implemented to address the impact of scale on retrieval and matching. To ensure consistency with the resolution of the basemaps in the basemap library, the test image is scaled to the same resolution as the basemaps in the basemap library. This is then used to calculate the length and width of the test image, and a distribution histogram and quartile plot are plotted. For densely distributed size regions, such as 2000-20,000 pixels, a more dense feature library construction strategy is developed. Specifically, starting at 2000 pixels, a feature library is constructed every 1000 pixels. In practical applications, the density of the feature library can be flexibly adjusted based on retrieval efficiency and accuracy requirements.

[0030] S103: Extract features and create indexes for each base map in the base map library.

[0031] Feature extraction and indexing are performed for each basemap in the basemap library. This involves reading a configuration file to obtain a list of cropping dimensions and a list of step sizes. The basemap is then cropped based on the cropping dimensions and step sizes. When the basemap features are large and distinct, a step size equal to the cropping window size is used, with each move covering a new, non-overlapping area. When the basemap features are small and dense, a step size smaller than the cropping window size is used, with each move covering a partially overlapping area. The cropped image data is converted to an array format and fed into the large model for feature extraction. The extracted features are associated with a unique key and stored as binary data. The unique key consists of the basemap name, cropping dimensions, the column coordinate of the top left corner of the cropping area, and the row coordinate of the top left corner of the cropping area. After all cropping dimensions for a basemap have been processed, the extracted features are categorized by cropping dimension, and features for different cropping dimensions are written to feature files categorized by cropping dimension. Based on the stored feature files, an index is created using the Faiss library for retrieval within the feature library.

[0032] Specifically, the size of a single base map is approximately 28,000*19,000 pixels. In this application, when the base map features are small and dense, the step size can be set to 1 / 2 the cropping window size. When the cropping window is at the image boundary, an indentation method is adopted. For example, for the right boundary, the cropped right boundary is the image width, and the left boundary is indented to the difference between the image width and the cropping window size. The cropped image data can be converted into numpy array format and sent to the large model for feature extraction. The feature file can be stored in json format for easy parsing and processing. The Faiss library is used to establish an index for efficient retrieval in the feature library. The Faiss library supports multiple index types, such as flat index (Flat Index), inverted file index (IVF) and product quantization index (PQ). In this application, considering efficiency and stability, a flat index method is used. The index establishment process is as follows. A corresponding index file is generated for each individual json feature file. When performing a retrieval task, for a test image, a feature library of appropriate size is selected for indexing. All json feature files are merged to generate a total index. For a test image, it can also be searched in the total feature library.

[0033] The large model used in this application is the Vision Transformer (ViT) architecture model (a deep learning model), and is optimized on its basis. Using fixed 2D sine and cosine position embedding, this position encoding method can effectively capture the spatial information of the image. The dimension of the position embedding is set to a multiple of 4, and the spatial position is encoded by sine and cosine functions, which helps the model better understand the features of different positions. In the nn.Linear layer (ie, the fully connected layer), special attention is paid to the initialization of the query (Q), key (K), and value (V) matrices. These matrices play a core role in the self-attention mechanism of the Transformer architecture.

[0034] To improve training stability and performance, the weights of these matrices are initialized using a uniform distribution. This means that each element of the Q, K, and V matrices is randomly drawn from a pre-set uniform distribution, providing a relatively balanced and random starting point for the model. The model's weight initialization strategy is carefully designed to improve training stability and performance. The Q, K, and V matrices in the nn.Linear layer (fully connected layers) are initialized using a uniform distribution. The weights of other layers in the model (such as other fully connected layers and convolutional layers) are initialized using a Xavier uniform distribution (a neural network weight initialization method). In the Vision Transformer (ViT) architecture, Patch Embedding is a key component responsible for partitioning the input image into a series of patches and mapping these patches into an embedding space. To maintain the consistency and stability of the Patch Embedding, its weights are initialized using a uniform distribution. Furthermore, to ensure that the output does not deviate from the center, the bias term is initialized to zero. In the ViT architecture, the Cls-token is a special token that represents the category information of the entire input sequence. To ensure the stability of the model's processing of Cls-tokens in the early stages of training, its weights are initialized using a normal distribution. This initialization method helps the model gradually learn more accurate category information during training.

[0035] This application uses Conv2d convolution as the model convolution front network, which cleverly integrates multi-layer convolution structures. Each layer has carefully designed convolution operations, batch normalization (Batch Normalization) and ReLU activation functions to ensure effective feature extraction and robust training of the model. In the first layer of the Conv2d network, it specializes in processing RGB images and captures the initial features of the image through carefully designed convolution kernels. As the network goes deeper, the subsequent convolution layers will gradually increase the number of channels of the feature map. This design aims to extract deeper and more abstract feature information from the image. Ultimately, these features will be mapped to a preset embedding dimension to provide high-quality input data for subsequent Transformer modules. The Conv2d convolution front network has excellent flexibility and can handle input images of various sizes. It can convert images into the specific format required by large models to ensure seamless transmission and efficient processing of information. In order to further improve the stability and training effect of the model, a regularization layer is added after the embedding layer. By default, the regularization layer uses the nn.Identity() function, which means no additional regularization is performed. However, this design leaves open the possibility of adding more complex regularization strategies in the future to accommodate different training requirements and scenarios.

[0036] S104: Select a retrieval strategy based on the size information of the test image and retrieve the base image name in the feature library.

[0037] According to the size information of the test image, a retrieval strategy is selected to retrieve the base map name in the feature library, including: traversing each size range in the feature library, and checking whether the size difference between the test image and the size range is less than the preset size threshold. If the difference between the size of the test image and a certain size range in the feature library meets the preset conditions, the retrieval strategy is to select the feature library of the size range for retrieval. If the difference between the size of the test image and multiple size ranges in the feature library all meet the preset conditions, the retrieval strategy is to select to search in all feature libraries. The preset condition is that the difference between the size of the test image and the image of a certain size range in the feature library is less than the preset size threshold. The difference is defined as the difference between the size of the test image and the boundary of the size range. According to the selected retrieval strategy, a base map similar to the test image is retrieved in the corresponding feature library, and the base map name is returned.

[0038] Specifically, this application uses the advanced Faiss library to implement an efficient vector retrieval mechanism. This mechanism is based on constructing and storing feature vectors for images in the base map library at different cropping sizes, aiming to quickly and accurately retrieve base maps similar to the test image. A large model is used to convert unstructured image data into structured feature vectors. These feature vectors are then imported into the Faiss library, indexed and stored (inserted) into the knowledge base (Knowledge Base), providing a basis for subsequent retrieval operations. When a search request (query) for a test image is received, the test image is first preprocessed, including removing black border areas and scaling to an appropriate resolution. Subsequently, to increase the robustness of the retrieval, the test image is rotated 0°, 90°, 180°, and 270° respectively, and these rotated images are converted into feature vectors using the same large model. Based on the size information of the test image, a hierarchical retrieval strategy or a merged retrieval strategy can be flexibly selected. Hierarchical retrieval is: if the difference between the size of the test image and a size range in the feature library meets the preset conditions, the retrieval strategy is to select the feature library of that size range for retrieval. For example, if the test image's width and height are 10,000 pixels, a basemap feature library cropped to 10,000 is assigned for retrieval. This strategy reduces unnecessary search space and improves retrieval efficiency. In a combined retrieval strategy, if the difference between the test image's dimensions and multiple size ranges in the feature library meets preset conditions, the retrieval strategy selects to search across all feature libraries. This strategy is suitable for situations where the test image's dimensions are similar to multiple size ranges in the feature library, ensuring comprehensive retrieval results. In actual retrieval, the top 100 images are typically retrieved as candidate objects for matching. Because each test image requires three rotation angles for retrieval, each image may generate up to 400 matching objects. Using the retrieved feature vector ID, the corresponding basemap name is found in the object database (Image Name) storing unstructured data images for the next stage of matching. This mechanism not only improves retrieval accuracy and efficiency but also provides strong support for subsequent image matching tasks.

[0039] S105: For each retrieved base map name, a grid image associated with the base map name is cropped from the base map library.

[0040] Specifically, an image area of a specific range is cropped from the base image, and this area is called a grid image.

[0041] S106: Convert the test image and the grid image into a first grayscale image and perform scaling processing.

[0042] S107: Extract feature points from the first grayscale image to form a first feature point set.

[0043] S108: Selecting inliers from the first feature point set using a random sampling consensus algorithm to estimate an affine transformation matrix between the test image and the grid image.

[0044] Specifically, since the matching model accepts grayscale images, both the test image and the grid image need to be converted to grayscale. This step ensures that the input image format is consistent with the model's requirements. To improve the efficiency of the matching process, the converted first grayscale image is scaled to a uniform size, such as (512, 512). This step reduces computational complexity while preserving the image's key features. Feature points are extracted from the scaled grayscale image using an appropriate feature extraction algorithm (such as SIFT or SURF), forming a first set of feature points. These feature points represent key information in the image, such as edges and corners. Using the Random Sample Consensus Algorithm (RANSAC), inliers—pairs of feature points with a high degree of matching—are selected from this first set of feature points. These inliers are used to estimate the affine transformation matrix between the test image and the grid image. An affine transformation is a linear transformation that maintains the linearity and parallelism of an image while allowing for scaling, rotation, and translation. As a preliminary assessment, the number of matching inliers is returned. A greater number of inliers indicates a higher degree of matching between the two images and a more accurate estimate of the affine transformation matrix.

[0045] Figure 2 A process diagram of a remote sensing image search and matching method based on feature fusion provided in an embodiment of the present application. Figure 2 The grid clipping in the example is a clipping operation, and re-clipping is a re-clipping operation. This application uses multi-task deep learning technology to fully explore the semantic features of the surface, extract global features for retrieval, and local features for matching.

[0046] It should be noted that the matching model used in this application is Efficient LoFTR, which aims to achieve accurate matching of image feature points through a series of efficient operation steps. This model is mainly compared with LoFTR, and a series of changes are made to make the matching method more efficient. The specific process is as follows. Given a pair of images I A with I B , respectively represent the test image and the grid image cropped from the base map library. First, the CNN network is used to extract its rough feature map and The coarse features are converted into more discriminative feature maps by interleaving the aggregated self-attention and cross-attention N times. In this step, feature aggregation is performed adaptively to reduce the token size before each attention operation, thereby improving computational efficiency. Unlike LoFTR, Efficient LoFTR first aggregates features on salient markers, significantly improving the efficiency of attention. The transformed coarse features are associated with the score matrix S, and then a mutual nearest neighbor (MNN) search is performed to establish a coarse matching set {M c This step preliminarily determines the corresponding relationship between the feature points of the image pairs. In order to refine the rough matching results, the transformed rough feature map and Fusion with backbone features to obtain discriminative fine features at full resolution and Then, the feature blocks are cut into each coarse matching M c The center of the sub-pixel level is obtained through two-stage refinement. f . This step further improves the matching accuracy. LoFTR uses all the markers of the feature map to calculate the attention and reduces the computational cost through linear attention. However, this method may still face high computational complexity. The attention module proposed by EfficientLoFTR first aggregates the features of the salient markers, significantly improving the efficiency of attention. Then, Vanilla attention (a self-attention mechanism) is used to transform the aggregated features and insert relative position encoding to capture spatial information. This method not only reduces the computational complexity but also maintains a high matching accuracy. In the fine matching stage, EfficientLoFTR fuses the transformed coarse features with the backbone features and obtains discriminative fine features at full resolution. This step makes full use of feature information at different levels and improves the robustness of matching. Through two-stage refinement operations, Efficient LoFTR can obtain sub-pixel level correspondences, further improving the matching accuracy.

[0047] S109: Return the number of matched inliers as a preliminary evaluation. If the number of matched inliers is greater than a first preset threshold, perform a fine matching step.

[0048] In an actual matching process, for a data set with sparse textures, the first preset threshold may be set to 50 for testing, and for a data set with dense textures, the first preset threshold may be set to 100 for testing.

[0049] The present application also includes: if the number of matched inliers is less than or equal to a first preset threshold, directly skipping the fine matching step and continuing to match the next base map name.

[0050] Specifically, if the number of matched inliers is less than or equal to the first preset threshold, it indicates that the current matching result is not accurate enough, and performing a fine match may further amplify the error, resulting in an even more inaccurate result. Skipping the fine match can avoid this situation.

[0051] The precise matching step includes: calculating the homography matrix between the grid image and the base map, and multiplying the homography matrix with the affine transformation matrix to obtain a transformation matrix. The transformation matrix is used to determine the position of the cropping area on the base map, and a preset padding area is added around the position to form a cropping frame. A cropping operation is performed on the base map according to the cropping frame to obtain a cropped image. The cropped image and the test image are both converted into a second grayscale image and scaled. Feature points are extracted from the second grayscale image to form a second feature point set. An inlier is selected from the second feature point set using a random sampling consensus algorithm to estimate the homography matrix between the cropped image and the test image.

[0052] S110: If the number of inliers obtained in the fine matching step is greater than a second preset threshold, the name of the base map that successfully matches the detection image and the latitude and longitude coordinates of the four corner points of the detection image on the base map are returned as the final evaluation.

[0053] S111: If the number of inliers obtained in the fine matching step is less than or equal to the second preset threshold, the fine matching step is iteratively executed until the first second preset number of retrieved base map names at different rotation angles of the detection image are matched, and the matching result with the largest number of inliers is returned as the best match, and the latitude and longitude coordinates of the four corner points of the detection image on the base map in the matching result are returned.

[0054] In an actual matching process, for a data set with sparse textures, the second preset threshold may be first set to 100 for testing; for a data set with dense textures, the second preset threshold may be first set to 200 for testing.

[0055] Specifically, in order to obtain the precise coordinates of the test image on the base map, it is first necessary to calculate the transformation relationship between the grid image and the base map. This step is achieved by calculating the homography matrix H between the grid image and the base map. The homography matrix is a geometric transformation that can describe the mapping relationship between two planes. After obtaining the homography matrix, it is necessary to multiply it with the previously calculated affine transformation matrix to obtain the transformation matrix. According to the transformation matrix, a cropping operation is performed on the base map. In order to ensure that the cropped area contains enough information to accommodate further precise matching, a preset padding area is added around the cropped area determined by the transformation matrix, such as a pixel width of 5, to form a cropping frame. Both the cropped image and the test image are converted into a second grayscale image and scaled. The second grayscale image can be scaled to a multiple of 32 to adjust the image size more flexibly.

[0056] In image matching tasks, the selection and merging strategies of feature libraries have a significant impact on matching accuracy and operational efficiency. To further explore this relationship, a series of experiments were designed and the experimental results were analyzed in detail.

[0057] In Example 1, as shown in Table 1, we experimented with feature library combinations of varying sizes, fixing the top N candidate set to 200. Experimental results show that with the appropriate selection and combination of feature library sizes, matching accuracy significantly improves while also effectively controlling runtime. In particular, when the feature library size ranges from 300 to 1000, the accuracy reaches 99.02%. While the runtime is slightly longer (4531 seconds), this significantly outperforms other combinations.

[0058] Table 1

[0059]

[0060] In Example 2, since the image resolution is known, we also need to consider the impact of image normalization and feature library merging on retrieval efficiency. Experiments have found that while merging feature libraries can improve operational efficiency, in some cases, the merging strategy may lead to mutual contamination between different features or waste of space, especially when using a planar index as the indexing method.

[0061] In order to meet these challenges, two sets of solutions are designed. Hierarchical solution: If the difference between the size of the test image and a certain size range in the feature library meets the preset conditions, the retrieval strategy is to select the feature library in this size range for retrieval. This strategy aims to reduce the mismatch and contamination problems between features. Merged solution: If the difference between the size of the test image and multiple size ranges in the feature library meets the preset conditions, the retrieval strategy is to select all feature libraries for retrieval to simplify the retrieval process and improve efficiency. As shown in Table 2, preliminary experimental results show that the two solutions have similar performance in accuracy (both 61.1%), but there are slight differences in running time, and the merged strategy takes slightly less time. Therefore, at the current stage, it is more inclined to choose the merged strategy as the preferred solution, while retaining the hierarchical strategy as an alternative.

[0062] Table 2

[0063]

[0064] Furthermore, as shown in Table 3, this application tests the success rate of different feature extraction models with Topk=100 and the time taken for single process and multi-process. It should be noted that Topk=100 refers to the top 100 results with the highest prediction probability.

[0065] Table 3

[0066]

[0067] As shown in Table 3, considering both efficiency and accuracy, the Mocov3 model demonstrated the best performance in the test, followed by the Dinov2-small model. Experimental results demonstrate that redundant feature preparation not only improves accuracy but also reduces processing time to a certain extent, thereby optimizing overall performance. Selecting a larger Topk value helps improve the success rate. Compared to single-process processing, multi-process processing offers significant time advantages and can significantly improve task processing efficiency. For consistent accuracy, smaller feature dimensions are preferred, as they result in faster retrieval.

[0068] As shown in Table 4, this application also explores the possibility of a dual-model strategy. The combination of dual models can further improve the accuracy. For example, the MAE_MOCO combination achieved an accuracy of 0.9574 in Top150, which is significantly better than the performance of a single model. There are differences in the effects between different model combinations, but in general, the combination of MOCO and other models performs better in terms of accuracy and stability. Except for MaMba, the entire or partial architecture of other models is ViT. DINO, MOCO, and MAE are pure vision models, and the overall architecture is ViT, while siglip and vil2-s are multimodal models, in which the architecture of the visual components is ViT.

[0069] Table 4

[0070]

[0071] As shown in Table 5, this application also conducted a test experiment. The baseline group used dino-s as the feature extraction model, set Topk to 100, and adopted a hierarchical strategy as the retrieval strategy. The control group: No special adjustments were made and the baseline group settings were used for testing. As a comparison benchmark, the accuracy rate under this strategy was 40%.

[0072] In the fine matching step, instead of resizing the cropped image to the same size as the test image, we resized it to a multiple of 32 of its resolution. This strategy improved the accuracy to 50% compared to the control group. This suggests that more detailed resizing of the cropped image during the second matching step can help improve matching accuracy.

[0073] During the retrieval process, the image is not resized and is retrieved directly at its original size. This strategy achieves an accuracy rate of 80%, significantly higher than the control group. However, it should be noted that the feature library size used for retrieval should be the size closest to the image after alignment with the base map.

[0074] The test image is aligned according to the resolution of the base map, and then the feature library closest to this size is selected as the retrieval feature library, that is, the feature library larger than this size and with the smallest size is selected as the retrieval feature library. The accuracy rate under this strategy is 40%.

[0075] Table 5

[0076]

[0077] The above tests yielded the following conclusions. During the fine matching step, more detailed adjustments to the cropped image size can help improve matching accuracy. Using the original image size for retrieval can significantly improve accuracy, but it's important to ensure that the feature library size matches the image size. When using a layered strategy, the feature library size selection strategy doesn't significantly affect accuracy, but you can choose the most appropriate strategy based on your specific situation to balance performance and storage requirements.

[0078] The embodiment of the present application also provides a remote sensing image search and matching device 300 based on feature fusion, such as Figure 3 As shown, the device includes: an acquisition module 301, a preprocessing module 302, a feature extraction module 303, a retrieval module 304, a cropping module 305, a conversion module 306, a feature point extraction module 307, an estimation module 308, a preliminary evaluation module 309 and a precise matching step module 310.

[0079] The acquisition module 301 is used to acquire a first preset number of test images, where the test images are from a drone or a satellite of a different source.

[0080] The preprocessing module 302 is used to preprocess the test image and establish a feature library based on the preprocessed test image.

[0081] The feature extraction module 303 is used to extract features and create indexes for each base map in the base map library.

[0082] The retrieval module 304 is used to select a retrieval strategy based on the size information of the test image and retrieve the base image name in the feature library.

[0083] The cropping module 305 is used to crop the grid image associated with each base map name retrieved from the base map library.

[0084] The conversion module 306 is used to convert the test image and the grid image into a first grayscale image and perform scaling processing.

[0085] The feature point extraction module 307 is used to extract feature points from the first grayscale image to form a first feature point set.

[0086] The estimation module 308 is configured to select inliers from the first feature point set using a random sampling consensus algorithm to estimate an affine transformation matrix between the test image and the grid image.

[0087] The preliminary evaluation module 309 is configured to return the number of matched inliers as a preliminary evaluation. If the number of matched inliers is greater than a first preset threshold, a fine matching step is performed.

[0088] If the number of inliers obtained in the fine matching step is greater than a second preset threshold, the fine matching step module 310 is configured to return the base map name that successfully matches the test image, as well as the latitude and longitude coordinates of the four corner points of the test image on the base map, as the final evaluation. If the number of inliers obtained in the fine matching step is less than or equal to the second preset threshold, the fine matching step is iteratively executed until the first second preset number of retrieved base map names at different rotation angles of the test image are matched. The matching result with the largest number of inliers is returned as the best match, along with the latitude and longitude coordinates of the four corner points of the test image on the base map in that matching result.

[0089] Some modules in the apparatus described herein may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0090] The devices or modules described in the above application embodiments can be implemented by computer chips or physical devices, or by products with certain functions. For ease of description, the above devices are described separately by function in various modules. When implementing the embodiments of this application, the functions of each module can be implemented in the same or multiple software and / or hardware. Of course, a module that implements a certain function can also be implemented by combining multiple sub-modules or sub-units.

[0091] The methods, devices, or modules described in this application can be implemented in the form of computer-readable program code. The controller can be implemented in any appropriate manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (such as software or firmware) that can be executed by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to implement the same function of the controller in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the means for realizing various functions may be considered to be both a software module for realizing the method and a structure within a hardware component.

[0092] like Figure 4 As shown, an embodiment of the present application also provides a remote sensing image search and matching server based on feature fusion, including a memory 401 and a processor 402; the memory 401 is used to store computer-executable instructions; the processor 402 is used to execute computer-executable instructions to implement a remote sensing image search and matching method based on feature fusion described above in an embodiment of the present application.

[0093] An embodiment of the present application also provides a computer-readable storage medium, which stores executable instructions. When a computer executes the executable instructions, it can implement the remote sensing image search and matching method based on feature fusion described above in the embodiment of the present application.

[0094] Through the description of the above implementation methods, it can be known that those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary hardware. Based on this understanding, the technical solution of the present application can be essentially or the part that contributes to the prior art can be embodied in the form of a software product, or it can be embodied in the implementation process of data migration. The computer software product can be stored in a storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a mobile terminal, a server, or a network device, etc.) to execute the method described in the embodiment of the present application.

[0095] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to in detail. Each embodiment focuses on the differences from other embodiments. All or part of this application can be used in many general or special computer system environments or configurations.

[0096] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit the present application. Although the present application has been described in detail with reference to the aforementioned embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the aforementioned embodiments can still be modified, or some or all of the technical features therein can be replaced by equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the present application.

Claims

1. A remote sensing image search and matching method based on feature fusion, characterized in that: include: Acquire a first preset number of test images, where the test images are from a drone or a satellite of a different source; Preprocessing the test image and establishing a feature library based on the preprocessed test image; Perform feature extraction and index creation on each base map in the base map library; According to the size information of the test image, select the retrieval strategy and search the base map name in the feature library; For each basemap name retrieved, a grid image associated with the basemap name is cropped from the basemap library; Convert the test image and the grid image into a first grayscale image and perform scaling processing; Extracting feature points from the first grayscale image to form a first feature point set; Selecting inliers from the first feature point set using a random sampling consensus algorithm to estimate the affine transformation matrix between the test image and the grid image; Return the number of matched inliers as a preliminary evaluation. If the number of matched inliers is greater than a first preset threshold, perform a fine matching step. If the number of inliers obtained in the fine matching step is greater than the second preset threshold, the name of the base map that successfully matches the detection image and the latitude and longitude coordinates of the four corner points of the detection image on the base map are returned as the final evaluation; If the number of inliers obtained in the fine matching step is less than or equal to the second preset threshold, the fine matching step is iteratively executed until the first second preset number of retrieved base map names at different rotation angles of the detection image are matched, and the matching result with the largest number of inliers is returned as the best match, and the latitude and longitude coordinates of the four corner points of the detection image on the base map in the matching result are returned.

2. The remote sensing image search and matching method based on feature fusion according to claim 1, characterized in that: Also includes: If the number of matched inliers is less than or equal to the first preset threshold, the fine matching step is directly skipped and the matching continues with the next base map name.

3. The remote sensing image search and matching method based on feature fusion according to claim 1, characterized in that: The preprocessing of the test image and establishing a feature library based on the preprocessed test image includes: Find and remove the black pixels at the edge of the test image to eliminate the black edge area; The test image after removing the black border area is rotated to generate image variants with different rotation angles; Scale the test image to the same resolution as the basemaps in the basemap gallery; Count the length and width of the scaled test image, draw the distribution histogram and quartile plot, and identify the area with dense size distribution by observing the distribution histogram and quartile plot; For areas with dense size distribution, a feature library is established every preset pixel.

4. The remote sensing image search and matching method based on feature fusion according to claim 3, characterized in that: The feature extraction and indexing of each base map in the base map library includes: Read the configuration file to obtain the cropping size list and step size list; Crop the base map according to the cropping size and step length; When the base map features are large and obvious, the step size is equal to the cropping window size, and each move covers a new and non-overlapping area; When the base map features are small and dense, the step size is smaller than the cropping window size, and each move covers partially overlapping areas; Convert the cropped image data into array format and feed it into the large model for feature extraction; The extracted features are associated with a unique key and stored in the form of word-typical data; wherein the unique key includes the base map name, cropping size, column coordinates of the upper left corner of the cropping area, and row coordinates of the upper left corner of the cropping area; After all cropping sizes of a base map are processed, the extracted features are classified according to the cropping size, and the features of different cropping sizes are written into feature files classified by cropping size. Based on the stored feature files, the Faiss library is used to create an index for retrieval in the feature library.

5. The remote sensing image search and matching method based on feature fusion according to claim 4, characterized in that: The method of selecting a retrieval strategy based on the size information of the test image and retrieving the base image name in the feature library includes: Traverse each size range in the feature library and check whether the size difference between the test image and the size range is less than the preset size threshold; If the difference between the size of the test image and a certain size range in the feature library meets the preset conditions, the retrieval strategy is to select the feature library of this size range for retrieval; If the difference between the size of the test image and multiple size ranges in the feature library all meet the preset conditions, the search strategy is to select the search in all feature libraries; The preset condition is that the size of the test image is smaller than the preset size threshold compared with the images in a certain size range in the feature library; The disparity is defined as the difference between the size of the test image and the size range boundary; According to the selected retrieval strategy, the base map similar to the test image is retrieved in the corresponding feature library and the base map name is returned.

6. The remote sensing image search and matching method based on feature fusion according to claim 1, characterized in that: The precise matching step comprises: Calculate the homography matrix between the grid image and the base map, and multiply the homography matrix with the affine transformation matrix to obtain the transformation matrix; The transformation matrix is used to determine the position of the cropping area on the base map, and a preset filling area is added around the position to form a cropping frame; Performing a cropping operation on the base map according to the cropping frame to obtain a cropped image; Convert the cropped image and the test image into a second grayscale image and perform scaling processing; Extracting feature points from the second grayscale image to form a second feature point set; The inliers are selected from the second set of feature points using a random sampling consensus algorithm to estimate the homography matrix between the cropped image and the test image.

7. A remote sensing image search and matching device based on feature fusion, characterized in that: include: An acquisition module is used to acquire a first preset number of test images, where the test images are from a drone or a satellite of a different source; A preprocessing module is used to preprocess the test image and establish a feature library based on the preprocessed test image; Feature extraction module, used to extract features and create indexes for each base map in the base map library; The retrieval module is used to select a retrieval strategy based on the size information of the test image and retrieve the base image name in the feature library; A cropping module is used for cropping a grid image associated with each base map name retrieved from the base map library; A conversion module, used for converting the test image and the grid image into a first grayscale image and performing scaling processing; A feature point extraction module, configured to extract feature points from the first grayscale image to form a first feature point set; an estimation module for selecting inliers from the first feature point set using a random sampling consensus algorithm and estimating an affine transformation matrix between the test image and the grid image; A preliminary evaluation module, configured to return the number of matched inliers as a preliminary evaluation, and if the number of matched inliers is greater than a first preset threshold, perform a fine matching step; A fine matching step module is configured to return the name of the base map that successfully matches the detection image, and the latitude and longitude coordinates of the four corner points of the detection image on the base map, as a final evaluation, if the number of inliers obtained in the fine matching step is greater than a second preset threshold; If the number of inliers obtained in the fine matching step is less than or equal to the second preset threshold, the fine matching step is iteratively executed until the first second preset number of retrieved base map names at different rotation angles of the detection image are matched, and the matching result with the largest number of inliers is returned as the best match, and the latitude and longitude coordinates of the four corner points of the detection image on the base map in the matching result are returned.

8. A remote sensing image search and matching server based on feature fusion, characterized in that: including memory and processor; The memory is used to store computer-executable instructions; The processor is configured to execute the computer-executable instructions to implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores executable instructions, and when a computer executes the executable instructions, the method according to any one of claims 1 to 6 can be implemented.

Citation Information

Cited By

  • Urban canyon positioning method and system integrating GNSS, vision and low-orbit satellites

    CN121254316A

  • Unmanned aerial vehicle image and multi-modal map area cutting method and system based on attitude information

    CN121527101A

  • Remote sensing image interpretation method for multi-modal information synchronization and model parameter optimization

    CN121600409A

  • A remote sensing image interpretation method of multi-modal information synchronization and model parameter optimization

    CN121600409B