Cross-View Image Retrieval Method under Limited Field of View Based on VMamba
The VMamba-based method addresses the challenges of limited field of view and orientation variability in cross-view angle image retrieval by dividing images into blocks, predicting orientation, and using dynamic embedding to enhance matching accuracy and reduce computational load.
Patent Information
- Application Number
- CN202510537730.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The prior art calculates overhead and search time in cross-view image retrieval under confined fields of view, and it is difficult to accurately match the features of ground images and remote sensing images in different directions.
Using a cross-view image retrieval method based on VMamba, by dividing ground images and remote sensing images into image blocks, local features are extracted using the backbone network, and image block similarity matrix is constructed. Combined with the direction prediction module and the position-embedded feature aggregation module, coarse-grained and fine matching is performed, reducing calculation overhead and improving retrieval accuracy.
It significantly reduces the computational overhead and duration of large-scale cross-view image retrieval, improves retrieval accuracy and recall, ensures that features can be accurately matched in different orientations, and improves the performance of cross-view geolocation.
Smart Images

Figure CN120067382B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of image retrieval, and particularly to a cross-view image retrieval method based on VMamba under a restricted field of view. Background Art
[0002] Cross-view geolocation technology aims to determine the geographical coordinates of the shooting location by matching ground photos and remote sensing images. This technology has important applications in fields such as autonomous driving and drone navigation. Traditional methods usually assume that ground photos are 360-degree panoramas and the shooting directions are known. However, in reality, photos taken by mobile phones or in-vehicle cameras often have a narrow field of view of only 70 degrees or 90 degrees, and the shooting directions are unknown, which poses the following challenges to positioning: 1) The field of view of ground images is limited and contains less information; 2) The orientations of ground images vary widely, and the corresponding regions between ground images and remote sensing images change.
[0003] Existing solutions are mainly divided into two categories: The first category of methods directly let deep learning models learn the feature matching between ground photos and satellite images, but ignore the direction information. When the ground photo turns to different directions, the content will change greatly (for example, it is a building complex when shooting eastward and becomes a forest when shooting westward), resulting in the model being difficult to establish a stable match. In the second category of methods, the model extracts features from both of them respectively, and predicts the orientation based on the extracted features to obtain a set of possible orientations. Then, the model crops all the images in the remote sensing image according to the predicted orientation, and performs a second feature extraction on the cropped remote sensing image before retrieving.
[0004] However, in the direction matching of this type of method, it is necessary to compare each angle one by one, which has extremely large computational overhead and retrieval time in large-scale cross-view image retrieval tasks and is unacceptable. At the same time, after predicting the orientation, it is necessary to crop all remote sensing images, and a second feature extraction must be performed on each image. On the one hand, performing feature extraction on all remote sensing images requires a large amount of computing resources and time; on the other hand, the remote sensing images have already undergone a feature extraction, and the second feature extraction seems repetitive and redundant. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide a cross-view image retrieval method based on VMamba under a restricted field of view that can improve the geolocation accuracy in complex scenarios.
[0006] A cross-view image retrieval method based on VMamba under a restricted field of view, the method includes:
[0007] Obtain a ground image and a remote sensing image; preprocess the remote sensing image to obtain a preprocessed remote sensing image;
[0008] Construct a cross - perspective image retrieval model; the cross - perspective image retrieval model includes a backbone network, an orientation prediction module, and a feature aggregation module based on position embedding;
[0009] Divide the ground image and the pre - processed remote sensing image into image patches of a predetermined proportional size and input them into the cross - perspective image retrieval model to perform local feature extraction and coarse - grained retrieval through the backbone network, obtaining multiple candidate remote sensing images; calculate the similarity between the ground image patches and the candidate remote sensing images, and construct an image patch similarity matrix; Figure 1 Divide the ground image and the pre - processed remote sensing image into image patches of a predetermined proportional size and input them into the cross - perspective image retrieval model to perform local feature extraction and coarse - grained retrieval through the backbone network, obtaining multiple candidate remote sensing images; calculate the similarity between the ground image patches and the candidate remote sensing images, and construct an image patch similarity matrix;
[0010] Input the image patch similarity matrix into the orientation prediction module to regress the shooting direction angle of the ground image using a multi - layer perception mechanism;
[0011] Input the shooting direction angle into the feature aggregation module based on position embedding to generate a dynamic embedding vector, add it to the candidate remote sensing images and perform feature aggregation, calculate the similarity according to the aggregated features and the features corresponding to the ground image patches, and output the remote sensing image corresponding to the highest similarity as the retrieval result;
[0012] Train the cross - perspective image retrieval model according to a preset loss function and test it using a preset test set, and implement cross - perspective image retrieval according to the well - tested cross - perspective image retrieval model.
[0013] The above - mentioned cross - perspective image retrieval method based on VMamba under a limited field of view first divides the ground image and the remote sensing image into image patches and inputs them into the backbone network of the cross - perspective image retrieval model to extract local features. In this way, the hidden information in the image is fully mined, and then candidate remote sensing images are screened out by local feature matching, overcoming the problem that it is difficult to match due to insufficient ground image information. Then, calculate the similarity between the ground image patches and the candidate remote sensing images, construct an image patch similarity matrix, and use the image patch similarity matrix to regress the shooting direction angle with the help of a multi - layer perception mechanism. At the same time, combined with the direction - adaptive cross - modal feature alignment mechanism in the cross - perspective image retrieval model, the image features are adaptively aligned according to the predicted orientation information, solving the problem of corresponding area changes and ensuring accurate feature matching under different orientations. Integrate the orientation information into the model features through variable position embedding to improve the retrieval accuracy. In addition, the coarse - to - fine two - stage retrieval process of the cross - perspective image retrieval model first performs coarse - grained retrieval to narrow the range and then performs fine - grained matching, greatly reducing the computational cost and solving the problem of excessive computational overhead and duration in large - scale retrieval. Finally, train the model through a preset loss function and test it using a test set to continuously optimize the model performance, enabling the model to stably and accurately output retrieval results in complex cross - perspective image retrieval tasks, comprehensively improving the performance of cross - perspective geolocation technology. Brief Description of the Drawings
[0014] Figure 1 1 is a flow chart of a cross-viewing image retrieval method under a restricted field of view based on VMamba in one embodiment;
[0015] Figure 2 A schematic diagram of training and testing a cross-view image retrieval model in one embodiment. DETAILED DESCRIPTION
[0016] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0017] In one embodiment, Figure 1 As shown, a cross-view image retrieval method under a restricted field of view based on VMamba is provided, comprising the following steps:
[0018] Step 102, acquiring a ground image and a remote sensing image; preprocessing the remote sensing image to obtain a preprocessed remote sensing image; constructing a cross-view image retrieval model; the cross-view image retrieval model includes a backbone network, an orientation prediction module, and a feature aggregation module based on position embedding.
[0019] The remote sensing image is converted from the Cartesian coordinate system to the polar coordinate system to eliminate the geometric perspective difference between the ground image and the satellite image.
[0020] Step 104, dividing the ground image and the pre-processed remote sensing image into original Figure 1 Image blocks of a certain scale are input into the cross-view image retrieval model, and local feature extraction and coarse-grained retrieval are performed through the backbone network to obtain multiple candidate remote sensing images; the similarity between the ground image blocks and the candidate remote sensing images is calculated, and the image block similarity matrix is constructed.
[0021] Divide the ground image and the preprocessed remote sensing image into image blocks of a certain size, and then input them into the backbone network of the cross-view image retrieval model for local feature extraction. The ground image with limited field of view and less information can be split into multiple small image blocks, each of which may contain relatively unique local features. By extracting these local features, the hidden information in the image can be fully mined and the model's ability to capture image information can be enhanced. After local feature extraction, coarse-grained retrieval is performed to obtain multiple candidate remote sensing images. This process uses the matching of local features to screen out the parts that may match the ground image from a large number of remote sensing images, avoiding the problem of difficulty in directly matching the overall remote sensing image due to insufficient ground image information.
[0022] Calculate the similarity between the ground image patches and the candidate remote sensing images, and construct an image patch similarity matrix. For the first time, a single inference calculation for direction prediction is realized through the ground-remote sensing image patch similarity matrix (PSM), replacing the traditional 360° sliding window comparison mode one by one, and significantly reducing the time consumption of direction prediction.
[0023] Step 106: Input the image patch similarity matrix into the orientation prediction module to regress the shooting direction angle of the ground image using a multi-layer perception mechanism; input the shooting direction angle into the feature aggregation module based on position embedding to generate a dynamic embedding vector, add it to the candidate remote sensing image, and perform feature aggregation. Calculate the similarity based on the aggregated features and the features corresponding to the ground image patches, and output the remote sensing image with the highest similarity as the retrieval result.
[0024] Input the image similarity matrix into the orientation prediction module to regress the shooting direction angle of the ground image using a multi-layer perception mechanism. By analyzing the similarity between image patches, this method can effectively capture the orientation information of the ground image. Even if the orientation of the ground image varies widely, it can accurately predict the shooting direction.
[0025] Use VMamba to construct the internal correlation of the image and predict the orientation of the ground image with the help of the image patch similarity matrix. VMamba can effectively capture the complex relationships within the image, providing richer information for orientation prediction and avoiding the problem of ignoring direction information in existing methods. Input the predicted shooting direction angle into the feature aggregation module based on position embedding to generate a dynamic embedding vector and add it to the candidate remote sensing image for feature aggregation. This variable position embedding method naturally integrates the orientation information into the model features, enabling the model to fully consider the direction factor during feature matching, solving the perspective deviation problem of fixed position encoding, and verifying that the recall rates are increased by 14.66% and 21.21% respectively on the CVUSA and CVACT datasets, improving the retrieval accuracy. The cross-view image retrieval model adopts a coarse-to-fine two-stage retrieval process. First, perform a coarse-grained retrieval to obtain multiple candidate remote sensing images, reducing the search range of subsequent fine-grained matching, thereby reducing the computational cost of retrieval. It avoids the huge computational overhead and retrieval time problem caused by comparing each angle one by one in large-scale cross-view image retrieval tasks in existing methods.
[0026] Step 108: Train the cross-view image retrieval model according to the pre-set loss function and test it using the pre-set test set, and realize cross-view image retrieval according to the tested cross-view image retrieval model.
[0027] The cross - view image retrieval model is trained using a pre - set loss function and tested using a pre - set test set. The loss function can guide the model to learn more accurate feature representations and orientation predictions, continuously optimizing the performance of the model. When the model faces complex cross - view image retrieval tasks, it can output retrieval results more stably and accurately, further enhancing the ability to solve the above - mentioned technical problems.
[0028] The above - mentioned cross - view image retrieval method based on VMamba under a restricted field of view first divides ground images and remote - sensing images into image patches, and inputs them into the backbone network of the cross - view image retrieval model to extract local features. In this way, the hidden information in the images is fully mined, and then candidate remote - sensing images are screened out through local feature matching, overcoming the problem of difficult matching due to insufficient ground - image information. Then, the similarity between the ground - image patches and the candidate remote - sensing images is calculated to construct an image - patch similarity matrix. Using the image - patch similarity matrix, the shooting direction angle is regressed with the help of a multi - layer perception mechanism. At the same time, combined with the direction - adaptive cross - modal feature alignment mechanism in the cross - view image retrieval model, the image features are adaptively aligned according to the predicted orientation information, solving the problem of corresponding - area changes and ensuring accurate feature matching under different orientations. The orientation information is integrated into the model features through variable - position embedding to improve the retrieval accuracy. In addition, the coarse - to - fine two - stage retrieval process of the cross - view image retrieval model first performs coarse - grained retrieval to narrow the range and then conducts fine - grained matching, significantly reducing the computational cost and solving the problem of excessive computational overhead and duration during large - scale retrieval. Finally, the model is trained using a pre - set loss function and tested using a test set, continuously optimizing the model performance, so that the model can stably and accurately output retrieval results in complex cross - view image retrieval tasks, comprehensively enhancing the performance of cross - view geolocation technology.
[0029] In one embodiment, the remote - sensing image is pre - processed to obtain a pre - processed remote - sensing image, including:
[0030] The remote - sensing image is converted from the Cartesian coordinate system to the polar coordinate system to obtain the pre - processed remote - sensing image.
[0031] In a specific embodiment, the remote - sensing image is converted from the Cartesian coordinate system to the polar coordinate system to eliminate the geometric view - angle difference between the ground image and the remote - sensing image (experiments show that the matching absolute error is reduced by more than 10%).
[0032] In one embodiment, local feature extraction and coarse - grained retrieval are performed through the backbone network to obtain multiple candidate remote - sensing images, including:
[0033] Extract the local features of the ground image patches and the preprocessed remote sensing image patches through the backbone network, map the local features into one-dimensional vectors using a linear mapping layer, and calculate the cosine similarity of the corresponding one-dimensional vectors of the ground image patches and the preprocessed remote sensing image patches to achieve coarse-grained retrieval. Select the top N candidate remote sensing images according to the similarity from high to low.
[0034] In one embodiment, input the shooting direction angle into the feature aggregation module based on position embedding to generate a dynamic embedding vector, including:
[0035] Dynamically adjust the generation strategy of position embedding based on the predicted direction, establish a direction-adaptive spatial encoding mechanism, and inject the shooting direction angle into the feature channels through a learnable weight matrix to generate a dynamic embedding vector with direction perception ability.
[0036] In one embodiment, the preset loss function includes the first retrieval loss, the second retrieval loss, and the orientation loss.
[0037] In one embodiment, the orientation loss is
[0038]
[0039] where N is the batch size, and represent the predicted orientation and the true orientation respectively.
[0040] In a specific implementation, the error of the model in orientation prediction is quantitatively measured by calculating the square of the difference between the predicted orientation and the true orientation and averaging all samples within the batch. The square operation makes the contribution of samples with larger errors to the loss more significant, enabling the model to pay more attention to those cases with large prediction deviations and prompting the model to optimize in the direction of reducing this error.
[0041] In one embodiment, the expression of the first retrieval loss is:
[0042]
[0043] where is the feature vector of the i th ground image patch, is the feature vector of the i th candidate remote sensing image patch, D represents the dimension of the feature vector, represents the m th dimension value of represents the m th dimension value of
[0044] In a specific embodiment, the present application sums the differences of each dimension of the feature vector, refines the feature differences of each dimension, ensures that the model does not ignore the feature changes of some important dimensions, can better adapt to the feature characteristics of different image patches, enables the model to more accurately screen out candidate remote sensing images related to the ground image in the coarse-grained retrieval stage, greatly promotes the effective matching of image features, and significantly improves the recall rate of the retrieval.
[0045] In one of the embodiments, the expression of the second retrieval loss is:
[0046]
[0047] Wherein, and are the gradients of the aggregated feature vectors, is the square of the Frobenius norm, which is used to measure the gradient difference, is the number of image pairs participating in the similarity calculation at this stage, is the k th feature vector of the ground image patch in the later aggregation stage, is the feature vector of the corresponding remote sensing image in the later aggregation stage.
[0048] In a specific embodiment, an additional gradient consistency term is designed to ensure that the change trends of the feature vectors are consistent during the feature aggregation process, optimize the feature aggregation process, further improve the accuracy of the final retrieval result, and enhance the performance of the model in the cross-view image retrieval task.
[0049] The total loss function of the cross-view image retrieval model is:
[0050]
[0051] Wherein, , and are the weights of the loss function respectively.
[0052] It should be understood that although each step in the Figure 1 flowchart is shown in sequence according to the indication of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 1At least a part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed and completed at the same moment, but can be executed at different moments. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turns with at least a part of other steps or sub-steps or stages of other steps.
[0053] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0054] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0055] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A cross-view image retrieval method under a restricted field of view based on VMamba, characterized in that The method includes: Obtaining a ground image and a remote sensing image; preprocessing the remote sensing image to obtain a preprocessed remote sensing image; Constructing a cross-view image retrieval model; the cross-view image retrieval model includes a backbone network, an orientation prediction module, and a feature aggregation module based on position embedding; Dividing the ground image and the preprocessed remote sensing image into image patches of a certain proportion of the original image size and inputting them into the cross-view image retrieval model. Local feature extraction and coarse-grained retrieval are performed through the backbone network to obtain multiple candidate remote sensing images; calculating the similarity between the ground image patches and the candidate remote sensing images, and constructing an image patch similarity matrix; Inputting the image patch similarity matrix into the orientation prediction module, using VMamba to construct the internal correlation of the image, predicting the ground image orientation by means of the image patch similarity matrix, and regressing the shooting direction angle of the ground image using a multi-layer perception mechanism; Inputting the shooting direction angle into the feature aggregation module based on position embedding to generate a dynamic embedding vector, adding it to the candidate remote sensing image and performing feature aggregation, calculating the similarity according to the aggregated feature and the feature corresponding to the ground image patch, and outputting the remote sensing image with the highest similarity as the retrieval result; Training the cross-view image retrieval model according to a preset loss function and testing it using a preset test set, and implementing cross-view image retrieval according to the well-tested cross-view image retrieval model.
2. The method according to claim 1, characterized in that Preprocessing the remote sensing image to obtain a preprocessed remote sensing image, including: Converting the remote sensing image from a Cartesian coordinate system to a polar coordinate system to obtain a preprocessed remote sensing image.
3. The method according to claim 1, wherein Performing local feature extraction and coarse-grained retrieval through the backbone network to obtain multiple candidate remote sensing images, including: Extract the local features of the ground image patches and the preprocessed remote sensing image patches through the backbone network, map the local features into one-dimensional vectors using a linear mapping layer, calculate the cosine similarity of the corresponding one-dimensional vectors of the ground image patches and the preprocessed remote sensing image patches to achieve coarse-grained retrieval, and screen the top N candidate remote sensing images according to the similarity level.
4. The method according to claim 1, wherein Inputting the shooting direction angle into the feature aggregation module based on position embedding to generate a dynamic embedding vector, including: Based on the predicted direction, dynamically adjusting the generation strategy of position embedding, establishing a direction-adaptive spatial encoding mechanism, and injecting the shooting direction angle into the feature channel through a learnable weight matrix to generate a dynamic embedding vector with direction perception ability.
5. The method according to claim 1, characterized in that The preset loss function includes a first retrieval loss, a second retrieval loss, and an orientation loss.
6. The method according to claim 5, wherein The orientation loss is ; Among them, N is the batch size, and respectively represent the predicted orientation and the true orientation.
7. The method according to claim 5, wherein The expression of the first retrieval loss is: ; Among them, is the feature vector of the i th ground image patch, is the feature vector of the i th candidate remote sensing image patch, D represents the dimension of the feature vector, represents the m th dimension value of represents the m th dimension value of 8. The method according to claim 5, characterized in that The expression of the second retrieval loss is: ; Among them, and are the gradients of the aggregated feature vectors, is the squared Frobenius norm, which is used to measure the gradient difference, is the logarithm of the number of images participating in the similarity calculation at this stage, is the k th eigenvector of the ground image patch in the later aggregation stage, is the eigenvector of the corresponding remote sensing image in the later aggregation stage.
Citation Information
Patent Citations
Cross-view image matching method and system fused with eye movement attention mechanism
CN119380052A
Learning data generation device, learning data generation method, and learning data generation program
JP2023163692A