VMama-based cross-view-angle image retrieval method in limited view field
By dividing the ground image and remote sensing image into image blocks, and using VMamba's cross-view image retrieval method, combining the orientation prediction module and the position-embedded feature aggregation module, the problem of large calculation overhead and inaccurate orientation matching in the traditional method is solved, and efficient cross-view geolocation is achieved.
Patent Information
- Application Number
- CN202510537730.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-27
AI Technical Summary
When traditional cross-view image retrieval methods process ground images under confined fields of view, the calculation overhead and search time are too large, and it is difficult to accurately match features under different orientations.
Using a cross-view image retrieval method under a confined field of view based on VMamba, the ground image and remote sensing image are divided into image blocks, and local feature extraction and coarse-grained search are used to combine the direction prediction module and the position-embedded feature aggregation module to realize direction adaptive cross-modal feature alignment.
This method significantly reduces the calculation overhead and retrieval time, improves geolocation accuracy, can accurately match features in different orientations, and improves the performance of cross-view geolocation technology.
Smart Images

Figure CN120067382A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image retrieval technology, and particularly to a cross-view image retrieval method based on VMamba under a restricted field of view. Background Art
[0002] Cross-view geolocation technology aims to determine the geographical coordinates of the shooting location by matching ground photos and remote sensing images. This technology has important applications in fields such as autonomous driving and drone navigation. Traditional methods usually assume that ground photos are 360-degree panoramas and the shooting directions are known. However, in reality, photos taken by mobile phones or vehicle-mounted cameras often have a narrow field of view of only 70 degrees or 90 degrees, and the shooting directions are unknown, which brings the following challenges to positioning: 1) The field of view of ground images is limited and contains less information; 2) The orientations of ground images vary greatly, and the corresponding areas between ground images and remote sensing images change.
[0003] Existing solutions are mainly divided into two categories: The first type of method directly allows a deep learning model to learn the feature matching between ground photos and satellite images, but ignores the direction information. When the ground photo turns to different directions, the content will change greatly (for example, when shooting eastward, it is a building complex, and when shooting westward, it becomes a forest), resulting in the model being difficult to establish a stable match. In the second type of method, the model extracts features from both of them respectively, and predicts the orientation based on the extracted features to obtain a set of possible orientations. Then, the model crops all the images in the remote sensing image according to the predicted orientation, and performs a second feature extraction on the cropped remote sensing image and then conducts the retrieval.
[0004] However, this type of method needs to compare each angle one by one in the direction matching, which has extremely large computational overhead and retrieval duration in large-scale cross-view image retrieval tasks and is unacceptable. At the same time, after predicting the orientation, all remote sensing images need to be cropped, and a second feature extraction must be performed on each image. On the one hand, performing feature extraction on all remote sensing images requires a large amount of computing resources and time; on the other hand, the remote sensing images have already undergone a first feature extraction, and the second feature extraction seems repetitive and redundant. Summary of the Invention
[0005] Based on this, it is necessary to provide a cross-view image retrieval method based on VMamba under a restricted field of view that can improve the geographical positioning accuracy in complex scenarios for the above technical problems.
[0006] A cross-view image retrieval method based on VMamba under a restricted field of view, the method includes: Obtain a ground image and a remote sensing image; perform preprocessing on the remote sensing image to obtain a preprocessed remote sensing image; Construct a cross - perspective image retrieval model; the cross - perspective image retrieval model includes a backbone network, an orientation prediction module, and a feature aggregation module based on position embedding; Divide the ground image and the pre - processed remote sensing image into image patches of a predetermined proportional size and input them into the cross - perspective image retrieval model for local feature extraction and coarse - grained retrieval through the backbone network to obtain multiple candidate remote sensing images; calculate the similarity between the ground image patches and the candidate remote sensing images to construct an image patch similarity matrix; Figure 1 Input the image patch similarity matrix into the orientation prediction module to regress the shooting direction angle of the ground image using a multi - layer perception mechanism; Input the shooting direction angle into the feature aggregation module based on position embedding to generate a dynamic embedding vector, add it to the candidate remote sensing image and perform feature aggregation, calculate the similarity based on the aggregated features and the features corresponding to the ground image patches, and output the remote sensing image with the highest similarity as the retrieval result; Train the cross - perspective image retrieval model according to a pre - set loss function and test it using a pre - set test set, and implement cross - perspective image retrieval according to the well - tested cross - perspective image retrieval model. For the above - mentioned cross - perspective image retrieval method based on VMamba under a limited field of view, first divide the ground image and the remote sensing image into image patches and input them into the backbone network of the cross - perspective image retrieval model to extract local features. In this way, the hidden information in the image is fully exploited, and then candidate remote sensing images are screened out by local feature matching, overcoming the problem that it is difficult to match due to insufficient ground image information. Then calculate the similarity between the ground image patches and the candidate remote sensing images to construct an image patch similarity matrix, and use the image patch similarity matrix to regress the shooting direction angle by means of a multi - layer perception mechanism. At the same time, combined with the direction - adaptive cross - modal feature alignment mechanism in the cross - perspective image retrieval model, the image features are adaptively aligned according to the predicted orientation information, solving the problem of corresponding area changes and ensuring accurate feature matching under different orientations. Integrate the orientation information into the model features through variable position embedding to improve the retrieval accuracy. In addition, the coarse - to - fine two - stage retrieval process of the cross - perspective image retrieval model first reduces the scope through coarse - grained retrieval and then performs fine - grained matching, greatly reducing the computational loss and solving the problem of excessive computational overhead and time duration in large - scale retrieval. Finally, train the model through a pre - set loss function, test it using the test set, and continuously optimize the model performance, so that the model can stably and accurately output retrieval results in complex cross - perspective image retrieval tasks, comprehensively improving the performance of cross - perspective geolocation technology.
[0007] Brief Description of the Drawings Brief Description of the Drawings
[0008] Figure 11 is a flow chart of a cross-viewing image retrieval method under a restricted field of view based on VMamba in one embodiment; Figure 2 A schematic diagram of training and testing a cross-view image retrieval model in one embodiment. DETAILED DESCRIPTION
[0009] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0010] In one embodiment, Figure 1 As shown, a cross-view image retrieval method under a restricted field of view based on VMamba is provided, comprising the following steps: Step 102, acquiring a ground image and a remote sensing image; preprocessing the remote sensing image to obtain a preprocessed remote sensing image; constructing a cross-view image retrieval model; the cross-view image retrieval model includes a backbone network, an orientation prediction module, and a feature aggregation module based on position embedding.
[0011] The remote sensing image is converted from the Cartesian coordinate system to the polar coordinate system to eliminate the geometric perspective difference between the ground image and the satellite image.
[0012] Step 104, dividing the ground image and the pre-processed remote sensing image into original Figure 1 Image blocks of a certain scale are input into the cross-view image retrieval model, and local feature extraction and coarse-grained retrieval are performed through the backbone network to obtain multiple candidate remote sensing images; the similarity between the ground image block and the candidate remote sensing image is calculated, and the image block similarity matrix is constructed.
[0013] The ground image and the preprocessed remote sensing image are divided into image blocks of a certain size, and then input into the backbone network of the cross-view image retrieval model for local feature extraction. The ground image with limited field of view and less information can be split into multiple small image blocks, each of which may contain relatively unique local features. By extracting these local features, the hidden information in the image can be fully mined and the model's ability to capture image information can be enhanced. After local feature extraction, coarse-grained retrieval is performed to obtain multiple candidate remote sensing images. This process uses the matching of local features to screen out the parts that may match the ground image from a large number of remote sensing images, avoiding the problem of difficulty in directly matching the overall remote sensing image due to insufficient ground image information.
[0014] Calculate the similarity between the ground image patches and the candidate remote sensing images, and construct an image patch similarity matrix. For the first time, single-inference calculation of direction prediction is realized through the ground-remote sensing image patch similarity matrix (PSM), replacing the traditional 360° sliding window comparison mode one by one, and greatly reducing the time consumption of direction prediction.
[0015] Step 106: Input the image patch similarity matrix into the orientation prediction module to regress the shooting direction angle of the ground image using a multi-layer perception mechanism; input the shooting direction angle into the feature aggregation module based on position embedding to generate a dynamic embedding vector, add it to the candidate remote sensing image, and perform feature aggregation. Calculate the similarity based on the aggregated features and the features corresponding to the ground image patches, and output the remote sensing image with the highest similarity as the retrieval result.
[0016] Input the image similarity matrix into the orientation prediction module, and regress the shooting direction angle of the ground image using a multi-layer perception mechanism. This method can effectively capture the orientation information of the ground image by analyzing the similarity between image patches. Even if the orientation of the ground image varies greatly, it can accurately predict the shooting direction.
[0017] Use VMamba to construct the internal correlation of the image, and predict the orientation of the ground image with the help of the image patch similarity matrix. VMamba can effectively capture the complex relationships within the image, providing richer information for orientation prediction and avoiding the problem of ignoring direction information in existing methods. Input the predicted shooting direction angle into the feature aggregation module based on position embedding to generate a dynamic embedding vector and add it to the candidate remote sensing image for feature aggregation. This variable position embedding method naturally integrates the orientation information into the model features, enabling the model to fully consider the direction factor during feature matching, solving the perspective deviation problem of fixed position encoding, and verifying that the recall rates are increased by 14.66% and 21.21% respectively on the CVUSA and CVACT datasets, improving the retrieval accuracy. The cross-view image retrieval model adopts a coarse-to-fine two-stage retrieval process, first performing coarse-grained retrieval to obtain multiple candidate remote sensing images, reducing the search range of subsequent fine-grained matching, thereby reducing the computational cost of retrieval. It avoids the huge computational overhead and retrieval time problem caused by comparing each angle one by one in large-scale cross-view image retrieval tasks in existing methods.
[0018] Step 108: Train the cross-view image retrieval model according to the preset loss function and test it using the preset test set, and realize cross-view image retrieval according to the well-tested cross-view image retrieval model.
[0019] Train the cross - view image retrieval model through a pre - set loss function and test it using a pre - set test set. The loss function can guide the model to learn more accurate feature representations and orientation predictions, continuously optimize the performance of the model, so that when the model faces complex cross - view image retrieval tasks, it can output retrieval results more stably and accurately, further improving the ability to solve the above - mentioned technical problems.
[0020] The above - mentioned cross - view image retrieval method based on VMamba under a restricted field of view first divides the ground image and the remote - sensing image into image patches, and inputs them into the backbone network of the cross - view image retrieval model to extract local features. In this way, the hidden information in the image is fully mined, and then candidate remote - sensing images are screened out by local feature matching, overcoming the problem that it is difficult to match due to insufficient ground - image information. Then, calculate the similarity between the ground - image patches and the candidate remote - sensing images, construct an image - patch similarity matrix, and use the image - patch similarity matrix to regress the shooting direction angle with the help of a multi - layer perception mechanism. At the same time, combined with the direction - adaptive cross - modal feature alignment mechanism in the cross - view image retrieval model, the image features are adaptively aligned according to the predicted orientation information, solving the problem of corresponding - region changes and ensuring accurate feature matching under different orientations. Integrate the orientation information into the model features through variable - position embedding to improve the retrieval accuracy. In addition, the coarse - to - fine two - stage retrieval process of the cross - view image retrieval model first performs coarse - grained retrieval to narrow the range, and then performs fine - grained matching, greatly reducing the computational cost and solving the problem of excessive computational overhead and duration in large - scale retrieval. Finally, train the model through a preset loss function and test it using a test set, continuously optimizing the model performance, so that the model can stably and accurately output retrieval results in complex cross - view image retrieval tasks, comprehensively improving the performance of cross - view geolocation technology.
[0021] In one embodiment, pre - process the remote - sensing image to obtain a pre - processed remote - sensing image, including: Convert the remote - sensing image from the Cartesian coordinate system to the polar coordinate system to obtain the pre - processed remote - sensing image.
[0022] In a specific embodiment, convert the remote - sensing image from the Cartesian coordinate system to the polar coordinate system to eliminate the geometric perspective difference between the ground image and the remote - sensing image (experiments show that the matching absolute error is reduced by more than 10%).
[0023] In one embodiment, perform local - feature extraction and coarse - grained retrieval through the backbone network to obtain multiple candidate remote - sensing images, including: Extract the local features of the ground image patches and the preprocessed remote sensing image patches through the backbone network, map the local features into one-dimensional vectors using a linear mapping layer, calculate the cosine similarity of the corresponding one-dimensional vectors of the ground image patches and the preprocessed remote sensing image patches to achieve coarse-grained retrieval, and screen the top N candidate remote sensing images according to the similarity.
[0024] In one embodiment, input the shooting direction angle into the feature aggregation module based on position embedding to generate a dynamic embedding vector, including: Dynamically adjust the generation strategy of position embedding based on the predicted direction, establish a direction-adaptive spatial coding mechanism, and inject the shooting direction angle into the feature channels through a learnable weight matrix to generate a dynamic embedding vector with direction perception ability.
[0025] In one embodiment, the preset loss function includes the first retrieval loss, the second retrieval loss, and the orientation loss.
[0026] In one embodiment, the orientation loss is
[0027] where N is the batch size, and represent the predicted orientation and the true orientation respectively.
[0028] In a specific implementation, the error of the model in orientation prediction is quantitatively measured by calculating the square of the difference between the predicted orientation and the true orientation and averaging all samples within the batch. The square operation makes the contribution of samples with larger errors to the loss more significant, enabling the model to pay more attention to those cases with large prediction deviations and prompting the model to optimize in the direction of reducing this error.
[0029] In one embodiment, the expression of the first retrieval loss is:
[0030] where is the feature vector of the i th ground image patch, is the feature vector of the i th candidate remote sensing image patch, D represents the dimension of the feature vector, represents the m th dimension value of represents the m th dimension value of
[0031] In a specific embodiment, the present application sums the differences in each dimension of the feature vector, refining the feature differences to each dimension, ensuring that the model does not ignore the feature changes in some important dimensions, being able to better adapt to the feature characteristics of different image patches, enabling the model to more accurately screen out candidate remote sensing images related to the ground image during the coarse-grained retrieval stage, greatly promoting the effective matching of image features, and significantly improving the recall rate of the retrieval.
[0032] In one of the embodiments, the expression of the second retrieval loss is:
[0033] Wherein, and are the gradients of the aggregated feature vectors, is the square of the Frobenius norm, used to measure the gradient difference, is the number of image pairs participating in the similarity calculation at this stage, is the k th feature vector of the ground image patch in the later aggregation stage, is the feature vector of the corresponding remote sensing image in the later aggregation stage.
[0034] In a specific embodiment, an additional gradient consistency term is designed to ensure that the change trends of the feature vectors are consistent during the feature aggregation process, optimize the feature aggregation process, further improve the accuracy of the final retrieval result, and enhance the performance of the model in the cross-view image retrieval task.
[0035] The total loss function of the cross-view image retrieval model is:
[0036] Wherein, , and are the weights of the loss function respectively.
[0037] It should be understood that although each step in the Figure 1 flowchart is shown in sequence according to the indication of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 1 at least a part of the steps in
[0038] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0039] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0040] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A cross-view image retrieval method under limited field of view based on VMamba, characterized in that: The method comprises: Acquire a ground image and a remote sensing image; preprocess the remote sensing image to obtain a preprocessed remote sensing image; Constructing a cross-view image retrieval model; the cross-view image retrieval model includes a backbone network, an orientation prediction module, and a feature aggregation module based on position embedding; The ground image and the preprocessed remote sensing image are divided into image blocks of a certain proportion of the original image, and input into the cross-view image retrieval model, and local feature extraction and coarse-grained retrieval are performed through the backbone network to obtain multiple candidate remote sensing images; the similarity between the ground image blocks and the candidate remote sensing images is calculated, and an image block similarity matrix is constructed; Inputting the image block similarity matrix into the orientation prediction module to regress the shooting direction angle of the ground image using a multi-layer perception mechanism; Input the shooting direction angle into the feature aggregation module based on position embedding to generate a dynamic embedding vector and add it to the candidate remote sensing image to perform feature aggregation, calculate the similarity between the aggregated features and the features corresponding to the ground image block, and output the remote sensing image corresponding to the highest similarity as the retrieval result; The cross-view image retrieval model is trained according to a preset loss function and tested using a preset test set, and cross-view image retrieval is implemented based on the tested cross-view image retrieval model.
2. The method according to claim 1, characterized in that: Preprocessing the remote sensing image to obtain a preprocessed remote sensing image includes: The remote sensing image is converted from a Cartesian coordinate system to a polar coordinate system to obtain a preprocessed remote sensing image.
3. The method according to claim 1, characterized in that Through the backbone network, local feature extraction and coarse-grained retrieval are performed to obtain multiple candidate remote sensing images, including: The local features of the ground image block and the preprocessed remote sensing image block are extracted through the backbone network, and the local features are mapped into a one-dimensional vector using a linear mapping layer. The cosine similarity calculation of the one-dimensional vector corresponding to the ground image block and the preprocessed remote sensing image block is calculated to realize coarse-grained retrieval, and the front-end is screened according to the similarity. N Candidate remote sensing images.
4. The method according to claim 1, characterized in that: Inputting the shooting direction angle into the position embedding-based feature aggregation module to generate a dynamic embedding vector includes: Based on the generation strategy of dynamically adjusting the position embedding of the predicted direction, a direction-adaptive spatial encoding mechanism is established, and the shooting direction angle is injected into the feature channel through a learnable weight matrix to generate a dynamic embedding vector with direction perception ability.
5. The method according to claim 1, characterized in that: The preset loss function includes a first retrieval loss, a second retrieval loss and an orientation loss.
6. The method according to claim 5, characterized in that The orientation loss is: in, N is the batch size, and represent the predicted orientation and the true orientation respectively.
7. The method according to claim 5, characterized in that The expressions of the first retrieval loss are: in, It is i The feature vector of the ground image patch, It is i The feature vector of candidate remote sensing image patches, D represents the dimension of the feature vector, express No. m The value of the dimension, express No. m The value of a dimension.
8. The method according to claim 5, characterized in that The expression of the second retrieval loss is: in, and is the gradient of the aggregated feature vector, is the squared Frobenius norm, which measures the gradient difference, is the number of image pairs involved in similarity calculation at this stage, It is k The feature vector of the ground image patch in the late aggregation stage, It is the feature vector of the corresponding remote sensing image in the later aggregation stage.
Citation Information
Patent Citations
Cross-view image matching method for unmanned aerial vehicle image and satellite image
CN112580546A
Cross-view-angle image matching and positioning method and system based on feature relation guidance
CN117495956A
Ocean remote sensing ship image retrieval method based on self-attention hashing
CN118312636A
Cross-view-angle image geographic positioning method and device under direction misalignment condition
CN118674785A
Feature extraction method and system based on difficult sample mining and multi-granularity division
CN118823371A