A method for rapid positioning of remote sensing images at any position
Through the combination of two-stage positioning strategies and VPR models, the rapid positioning of remote sensing images is used to use FAISS and LOFTR technology to solve the problems of low positioning accuracy and efficiency in the existing technology, and efficient and accurate remote sensing image positioning is achieved.
Patent Information
- Application Number
- CN202510013302.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-01-06
AI Technical Summary
When processing optical remote sensing images, the accuracy of TOP1 and TOP5 has significantly decreased, and the performance is poor for low-texture or few-feature scenarios, lack of verification mechanisms, large calculation volume, and high storage space occupancy.
Using a two-stage positioning strategy, combining advanced VPR model and RERANK technology, efficient global feature retrieval is carried out through the FAISS vector feature library, and fine local feature matching and reordering is used to achieve fast and accurate positioning of remote sensing images.
It significantly improves the fast and accurate positioning ability of remote sensing images, improves positioning accuracy and efficiency, reduces calculation and storage requirements, and enhances adaptability to low-texture scenes.
Smart Images

Figure CN119418223B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image positioning, and in particular to a method for quickly positioning remote sensing images at any position. Background Art
[0002] The research on Visual Place Recognition (VPR) technology has traditionally focused mainly on the field of street view positioning, and its application in multi-source optical remote sensing images is relatively less. Although the visual model pre-trained based on large-scale street view data can achieve a TOP1 accuracy rate of over 0.95 after fine-tuning, due to the unique overlooking shooting perspective of optical remote sensing images, there is a significant semantic gap between them and natural pictures. This difference results in a significant decline in the TOP1 and TOP5 accuracy rates of the VPR model directly fine-tuned with street view data when processing optical remote sensing images.
[0003] The Chinese invention patent application with the application number CN115630236A discloses a method, storage medium, and device for global rapid retrieval and positioning of passive remote sensing images. The positioning method includes: establishing a global reference image tile database, extracting local feature points and feature descriptors of the tiles, performing feature aggregation on the local feature point features to generate a global feature vector of the tile, which is the first global feature vector, establishing a global reference image tile data feature library, for the passive remote sensing image to be matched, generating a second global feature vector to be matched using the same method, comparing the second global feature vector with the first global feature vector and performing similarity checking, and calculating the image with the highest similarity as the matching image, and taking the geographical coordinates of this matching image as the geographical coordinates of the passive image to be matched. The disadvantages of this method are: 1) The generated global feature vector has a very high dimension, resulting in a very high storage space required for establishing the global reference image tile database and a very high computational amount for subsequent calculation of the similarity of the feature vectors; 2) It depends on the extraction and aggregation of local feature points, which leads to a strong dependence on low-level features, is easily affected by illumination and weather changes, has a weak adaptability to long-term changes, and performs poorly in low-texture or few-feature scenarios; 3) It lacks a verification mechanism, and global remote sensing image positioning only relies on the distance between two 256-dimensional vectors to determine whether the matching is successful, and the success rate is low. Summary of the Invention
[0004] In order to solve the above technical problems, the present invention proposes a method for quickly positioning remote sensing images at any position, adopting an innovative two-stage positioning strategy, integrating an advanced VPR model and RERANK technology, and effectively improving the rapid and accurate positioning of remote sensing images without any fine-tuning.
[0005] The object of the present invention is to provide a method for quickly positioning remote sensing images at any position, including constructing a database, and further comprising the following steps:
[0006] Step 1: Obtain the remote sensing image to be located and name it the test image;
[0007] Step 2: Preprocess the test image;
[0008] Step 3: Divide the preprocessed remote sensing image into blocks and perform feature extraction;
[0009] Step 4: Perform feature aggregation and sorting;
[0010] Step 5: Perform re - sorting and initial positioning;
[0011] Step 6: Perform secondary sorting and precise positioning;
[0012] Step 7: Evaluate the precise positioning result;
[0013] Step 8: When the positioning accuracy rate E is greater than the accuracy rate threshold, output the final positioning position of the test image.
[0014] Preferably, the database includes an offline reference image dataset, a test dataset, and a FAISS vector feature library.
[0015] In any of the above - mentioned solutions, preferably, the i th reference image T i in the offline reference image dataset is randomly cropped into m sub - region image blocks according to BLOCK_SIZE * BLOCK_SIZE pixels to generate a sub - region image block set T i,j ={ T i,1 , T i,2 ,…, T i,m}, where, j represents the j th sub - region, 1 ≤ j ≤ m .
[0016] In any of the above - mentioned solutions, preferably, each image in the sub - region image set T i,j is rotated, scaled, and noise - added, and stored in the test dataset.
[0017] In any of the above - mentioned solutions, preferably, there is an overlap region of overlap pixels between each sub - region image block. Global features are extracted from each sub - region image block and these features are stored in the FAISS vector feature library.
[0018] Preferably, in any of the above solutions, the preprocessing includes calculating the minimum circumscribed quadrilateral of the effective area of the test image, and accordingly, affine-transforming the test image into a "vertical" state.
[0019] Preferably, in any of the above solutions, the block division includes cutting the test image into N test image blocks according to BLOCK_SIZE*BLOCK_SIZE pixels.
[0020] Preferably, in any of the above solutions, the feature extraction includes extracting global features for each of the test image blocks, retrieving the TOPK nearest features in the FAISS vector feature library, and obtaining the sub-region image blocks corresponding to these features, the name of the reference image, and the position of the reference image.
[0021] Preferably, in any of the above solutions, step 4 includes the following sub-steps:
[0022] Step 41: Count all the retrieved N*TOPK features;
[0023] Step 42: For each global feature, check whether there are adjacent global features. When there are adjacent features, the count of the corresponding global feature is incremented;
[0024] Step 43: Sort according to the count of the global features. If the counts are the same, perform a secondary sort according to the order of TOPK;
[0025] Step 44: Record the test image blocks corresponding to each global feature in the sorting result.
[0026] Preferably, in any of the above solutions, step 5 includes performing RERANK processing on the test image and the reference image based on the sorting result, using LOFTR for fine image matching. When the number of matched key points exceeds a given threshold and the number of inliers after RANSAC affine transformation meets the requirements, it is regarded as a successful preliminary positioning.
[0027] Preferably, in any of the above solutions, step 6 includes, after the preliminary positioning is successful, performing a secondary LOFTR matching on the approximate matching area between the vertical image and the reference image, and obtaining the exact position of the reference image by matching the four corners of the reference image.
[0028] Preferably, in any of the above solutions, the evaluation method is:
[0029] Assume there are Y test images. For the y th test image, assume its predicted bounding box is Byp , the ground truth box is B yg . Both the predicted box and the ground truth box are determined by the coordinates of two vertices, the upper left corner and the lower right corner, and their IoU y value is B yp and B yg the intersection area of them divided by the union area of them. The calculation formula is:
[0030] IoU y = ( B yp ∩ B yg ) / ( B yp ∪ B yg )
[0031] The localization accuracy rate on Y test images is E calculated as follows: .
[0032] The present invention proposes a method for quickly locating remote sensing images at any position, which improves the effectiveness in the aspect of visual position recognition of remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a flowchart of a preferred embodiment of the method for quickly locating remote sensing images at any position according to the present invention.
[0034] Figure 2 is a flowchart of another embodiment of the visual positioning of satellite images of the method for quickly locating remote sensing images at any position according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. Embodiment 1
[0036] As Figure 1 shown, a method for quickly locating remote sensing images at any position performs step 100 to construct a database, and the database includes an offline reference image dataset, a test dataset, and a FAISS vector feature library.
[0037] The i th reference image in the offline reference image dataset T i is randomly cropped intom Sub-region image blocks are generated to form a set of sub-region image blocks T i,j ={ T i,1 , T i,2 ,…, T i,m}, where j represents the j -th sub-region, and 1 ≤ j ≤ m .
[0038] Each image in the sub-region image set T i,j is rotated, scaled, and noise-added, and stored in the test dataset. There is an overlap region with overlap pixels between each pair of the sub-region image blocks. Global features are extracted from each sub-region image block and these features are stored in the FAISS vector feature library.
[0039] Step 110 is executed to obtain the remote sensing image to be located, which is named the test image.
[0040] Step 120 is executed to preprocess the test image. The preprocessing includes calculating the minimum circumscribed quadrilateral of the effective region of the test image, and accordingly, affine-transforming the test image into a "vertical" state.
[0041] Step 130 is executed to block the preprocessed remote sensing image and perform feature extraction. The blocking includes cropping the test image into N test image blocks according to BLOCK_SIZE * BLOCK_SIZE pixels.
[0042] The feature extraction includes extracting global features from each test image block, retrieving the TOPK nearest features in the FAISS vector feature library, and obtaining the sub-region image blocks corresponding to these features, the name of the reference image, and the position of the reference image.
[0043] Step 140 is executed for feature aggregation and sorting, including the following sub-steps:
[0044] Step 141 is executed to count all the retrieved N * TOPK features;
[0045] Step 142 is executed to check, for each global feature, whether there is an adjacent global feature. When there is an adjacent feature, the count of the corresponding global feature is incremented;
[0046] Step 143 is executed to sort according to the count of the global features. If the counts are the same, secondary sorting is performed according to the order of TOPK;
[0047] Execute step 144 to record the test image blocks corresponding to each global feature in the sorting result.
[0048] Execute step 150 for re - sorting and initial positioning, including performing RERANK processing on the test image and the reference image based on the sorting result, using LOFTR for fine image matching. When the number of matched key points exceeds a given threshold and the number of inliers after RANSAC affine transformation meets the requirements, it is regarded as successful in initial positioning.
[0049] Execute step 160 for secondary sorting and precise positioning, including after successful initial positioning, performing secondary LOFTR matching on the approximate matching area between the vertical image and the reference image, and obtaining the precise position of the reference image by matching the four corners of the reference image.
[0050] Execute step 170 to evaluate the precise positioning result. The evaluation method is as follows:
[0051] Assume there are Y test images. For the y th test image, assume its predicted bounding box is B yp , and the ground - truth bounding box is B yg . Both the predicted bounding box and the ground - truth bounding box are determined by the coordinates of two vertices, the upper - left and the lower - right. Its IoU y value is the area of the intersection of B yp and B yg divided by the area of their union. The calculation formula is:
[0052] IoU = y = ( B yp ∩ B yg ) / ( B yp ∪ B yg )
[0053] The positioning accuracy rate Y on E test images is calculated as follows: .
[0054] Execute step 180. When the positioning accuracy rate E is greater than the accuracy threshold, output the final positioning position of the test image. Example Two
[0055] In the existing positioning methods, the entire remote sensing image is compared one by one with all the images in the database. This approach is slow on the one hand and occupies a large amount of computing resources on the other hand, resulting in a slower recognition and positioning speed.
[0056] The present invention creatively proposes a two-stage VPR strategy, aiming to improve the positioning accuracy of remote sensing images through efficient global feature retrieval and fine local feature reordering. In the first stage, we use the SELAVPR model for global feature search. SELAVPR realizes the hybrid extraction of global and local features by inserting a lightweight adapter into the pre-trained model, where the global feature dimension is only 1024 dimensions, which is beneficial to saving storage space and is especially suitable for large-scale remote sensing image databases.
[0057] Among numerous VPR models, SELAVPR is selected due to its compact global feature representation. This model not only retains a powerful feature extraction ability but also significantly reduces the storage requirements by optimizing the feature dimension, which is crucial for remote sensing image positioning applications that require offline storage of a large number of global features. We only use the global features of SELAVPR to build a feature database to support efficient online retrieval.
[0058] For the global retrieval results, in the second stage, LOFTR (a Transformer-based local feature matcher) is used for fine local feature matching and reordering. LOFTR abandons the traditional feature detection step and directly processes the dense feature map, using the self-attention and cross-attention mechanisms of the Transformer to capture global context information, achieving the ability to generate high-quality matching items even in low-texture, blurred, or repetitive patterns.
[0059] To further improve the positioning accuracy, we propose a spatial mutual nearest neighbor global TOPK sorting method. This method effectively compensates for the problem of low TOP1 and TOP5 accuracy in remote sensing images for TOPK retrieval based on the VPR model, and realizes a more accurate position estimation by comprehensively considering the spatial distribution and feature similarity of candidate images. Example Three
[0060] A method for rapid positioning of remote sensing images at any position, including:
[0061] 1. Offline stage: Database construction
[0062] In the implementation of VPR technology, the core task in the offline stage is to build a comprehensive and efficient database. This database focuses on storing a series of pre-acquired images that widely cover different locations and perspectives of the target environment, providing a solid data foundation for subsequent online processing. For a VPR system based on remote sensing satellite images, this stage is particularly important because it requires not only capturing a vast geographical view but also precisely handling image scaling ratios and overlaps to ensure that the database can reflect both the macroscopic features of the terrain and the accuracy of details.
[0063] 2. Online Stage: Image Retrieval and Image Registration
[0064] Entering the online stage, the system needs to have real-time or near-real-time image processing capabilities to handle dynamic query requirements. This stage consists of two core links:
[0065] Image Retrieval: Facing a new query image, the system first activates the image retrieval mechanism to quickly search for candidate images in the already constructed database that best match the query image. This process relies on efficient feature matching algorithms or deep learning techniques, aiming to rapidly narrow down the search scope, improve the recognition efficiency, and ensure that the possible target locations can be quickly located in the vast amount of data.
[0066] Image Registration: Once the candidate image is selected, the system enters the image registration stage. This link deeply analyzes the subtle differences between the query image and the candidate image, and through advanced techniques such as feature point matching and affine transformation estimation, precisely calculates the spatial correspondence relationship between the two, thereby achieving the precise positioning of the query image in the target environment.
[0067] 3. Dataset
[0068] To support the verification and optimization of the above technical processes, this paper uses the "Tianzhi Cup Arbitrary Region Optical Image Rapid Positioning Dataset" as the offline benchmark image data. This dataset consists of 980 high-resolution (1 meter) optical satellite images, and each image is an RGB three-channel image with 5000×5000 pixels, and the format is uniformly TIF. This dataset is not only large in scale but also rich in content, providing valuable experimental resources for the research and application of VPR technology.
[0069] The test dataset in this paper consists of two major parts:
[0070] Online Unpublished Test Set: This part contains 200 specially processed data sets, which are derived from the benchmark data set but have undergone a series of complex transformations to enhance their diversity and challenge. Specifically, the images in the benchmark data set are randomly cropped into 200 sub-region images, and the original size of each sub-region ranges from 300×300 pixels to 1000×1000 pixels. After cropping, these sub-region images further undergo various image processing operations such as rotation, scaling, and noise addition, and are randomly saved in JPG, PNG, or TIF formats. It should be noted that since these data sets are unpublished online, they are only used for accuracy verification through algorithm mirroring to evaluate the generalization ability of the model on unknown data.
[0071] Internally Generated Simulation Test Set: To supplement the online test set and increase the diversity of test samples, we randomly generated N test data sets ourselves according to the processing method of the online test data set. These simulation data sets strictly follow the preprocessing process of the online test set, including cropping, rotation, scaling, noise addition, and random format conversion, aiming to simulate complex scenarios and noise conditions in the real world, so as to more comprehensively evaluate the performance and robustness of the model.
[0072] By combining the use of these two parts of test data sets, we can more comprehensively and accurately evaluate the performance of the proposed algorithm under different conditions, providing strong support for the optimization and improvement of the model.
[0073] 4. Evaluation Method
[0074] In terms of evaluation metrics, the localization accuracy and localization efficiency are comprehensively used to measure the algorithm performance. The localization accuracy is obtained by averaging the localization accuracies of each test image, and the localization accuracy of each test image is evaluated by the IOU (Intersection over Union) metric.
[0075] Suppose there are Y test images. For the y th test image, assume its predicted bounding box is B yp , and the ground truth bounding box is B yg . Both the predicted bounding box and the ground truth bounding box are determined by the coordinates of the two vertices of the upper left corner and the lower right corner. Its IoU y value is the area of the intersection of B yp and B yg divided by the area of their union. The calculation formula is:
[0076] IoU y = ( B yp∩ B yg ) / ( B yp ∪ B yg )
[0077] At Y the positioning accuracy on the test images E The calculation formula is as follows: .
[0078] Example 4
[0079] As Figure 2 shown, the experimental process is mainly divided into two stages: offline and online.
[0080] Offline stage
[0081] In the offline stage, first, the reference image is segmented into image blocks of size BLOCK_SIZE×BLOCK_SIZE pixels, and there is an overlap area of overlap pixels between adjacent image blocks. Subsequently, the SELAVPR technology is used to extract global features from each image block, and these features are stored in the FAISS vector feature library for subsequent fast retrieval.
[0082] Online stage
[0083] The online stage involves the processing and positioning of the test image, and the specific steps are as follows:
[0084] Preprocessing: Since the test image may have experienced rotation and asymmetric stretching of width and height, resulting in a large invalid area. Therefore, first calculate the minimum circumscribed quadrilateral of the valid area, and accordingly perform an affine transformation on the test image to the "vertical" state to maximize the information utilization rate.
[0085] Chunking and feature extraction: The processed test image is cropped according to the BLOCK_SIZE size and considering a certain overlap degree to generate N test image blocks. Then, use SELAVPR to extract global features from each test image block, retrieve the TOPK nearest features in the FAISS database, and at the same time obtain the reference image name and its image position (IMAGENAME_LEFT_TOP) corresponding to these features.
[0086] Feature Aggregation and Sorting: Count all the retrieved N * TOPK features. For each reference feature (a reference feature is the feature of each record in the reference image dataset, which is a global feature and is associated with the image), check if there are adjacent reference features (based on IMAGENAME consistency and proximity). If adjacent features exist, the count of the corresponding reference feature is incremented; otherwise, the count is set to 0. Then, sort according to the counts. If the counts are the same, perform a secondary sort based on the order of TOPK. Meanwhile, record the test images corresponding to each reference feature.
[0087] Re - sorting and Initial Positioning: Based on the sorting results, perform RERANK processing on the test images and reference images, and use LOFTR for fine image matching. If the number of matched key points exceeds a given threshold and the number of inliers after RANSAC affine transformation meets the requirements, it is considered that the initial positioning is successful.
[0088] Secondary Matching and Precise Positioning: After the initial positioning is successful, perform secondary LOFTR matching on the approximate matching area between the "vertical image" and the reference image to further narrow down the positioning range and achieve more accurate position recognition.
[0089] Conduct Experiment 1 to evaluate the performance of the image matching and position recognition algorithms. First, the first experiment focuses on the internal test set, which has been pre - processed to remove images lacking significant features (such as images with full vegetation coverage) and uniformly adjusted to the "vertical" state. The experiment uses BLOCK_SIZE = 256 for image block division to generate N test image blocks. Subsequently, use SELAVPR to extract features, and through the feature aggregation and sorting strategy, select the global features with a sorting count greater than 1 and in the top 60% to reduce the computational load. On this basis, perform RERANK processing and LOFTR fine matching on the test images and reference images. If the number of matched key points and the number of inliers after RANSAC affine transformation meet the preset conditions, it is considered that the initial positioning is successful. Subsequently, perform secondary LOFTR matching on the approximate matching area between the "vertical" image and the reference image to further improve the positioning accuracy. This experiment generates a total of 5 groups, with 200 test data in each group. As shown in the statistical results of Table 1, the success rates and processing times of TOP1, TOP5, TOP10, and TOP128 are all excellent. Especially for TOP128, almost all images are detected, and the average processing time per single image is less than 2 seconds, with an average accuracy of 0.99638, reflecting high efficiency and accuracy.
[0090] Table 1 Optical Satellite Positioning Evaluation Results for the Internal Test Set
[0091]
[0092] Experiment 2 was conducted on the online test set with the aim of comprehensively locating all images. The experimental process was similar to the first experiment, but a loop processing step was added to meet the competition requirements, and two adjustments were made to BLOCK_SIZE (256 and 386) to adapt to different image features. First, the vast majority of images (197) were successfully located through the TOP128 strategy, and only 3 complex images could not be recognized. Subsequently, for these 3 difficult images, the TOP384 strategy was successfully used for location, as shown in the statistical results of Table 2.
[0093] Table 2 Optical Satellite Location Evaluation Results of the Online Test Set
[0094]
[0095] In summary, the experimental process effectively improved the efficiency and accuracy of visual position recognition. By recording the initial positions of the four corner points of the test images and the relative positions after transformation, the accuracy of location was ensured. In particular, by preferentially processing the features with the highest count and a sorted count greater than 0, the speed of local matching was significantly accelerated, and the success rate of location was increased. In the Tianzhi Cup competition, the accuracy of this solution reached 0.995846 based on the online unpublished test set, ranking first in the Tianzhi Cup competition, demonstrating its high efficiency and precision. After measurement, the matching accuracy of each image reached the sub-pixel level (0.6 - 0.7 pixels), further verifying its excellent performance.
[0096] To better understand the present invention, the above has been described in detail in combination with specific embodiments of the present invention, but it is not a limitation of the present invention. Any simple modification made to the above embodiments based on the technical essence of the present invention still belongs to the scope of the technical solution of the present invention. Each embodiment in this specification focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment.
Claims
1. A method for rapid positioning of remote sensing images at any position, comprising building a database, characterized in that: The following steps are also included: Step 1: Obtain the remote sensing image that needs to be located and name it as a test image; Step 2: preprocessing the test image; Step 3: Divide the preprocessed remote sensing image into blocks and perform feature extraction, wherein the block division includes cutting the test image according to BLOCK_SIZE*BLOCK_SIZE pixels to generate N test image blocks; the feature extraction includes extracting global features for each of the test image blocks, retrieving TOPK nearest features in the FAISS vector feature library, and obtaining the sub-region image blocks corresponding to these features, the names of the benchmark images, and the positions of the benchmark images; Step 4: Perform feature aggregation and sorting, including the following sub-steps: Step 41: Count all retrieved N*TOPK features; Step 42: for each reference feature, checking whether there is an adjacent reference feature, when there is an adjacent feature, the count of the corresponding reference feature is increased; Step 43: sorting according to the counts of the benchmark features, and if the counts are the same, performing secondary sorting according to the order of TOPK; Step 44: Record the test image block corresponding to each benchmark feature in the sorting result; Step 5: re-ranking and initial positioning, including performing RERANK processing on the test image and the reference image based on the ranking result, and using LOFTR to perform fine image matching. When the number of matched key points exceeds a given threshold and the number of internal points after RANSAC affine transformation meets the requirements, the initial positioning is considered to be successful; Step 6: performing secondary sorting and precise positioning, including performing secondary LOFTR matching for the matching area of the vertical image and the reference image after the initial positioning is successful, and obtaining the precise position of the reference image by matching the four corners of the reference image; Step 7: Evaluate the precise positioning results; Step 8: When the positioning accuracy E When the accuracy is greater than the threshold, the final positioning position of the test image is output.
2. The method for rapid positioning of remote sensing images at any position as claimed in claim 1, characterized in that: The database includes an offline benchmark image data set, a test data set and a FAISS vector feature library.
3. The method for rapid positioning of remote sensing images at any position as claimed in claim 2, characterized in that: The first i Baseline images T i Randomly crop according to BLOCK_SIZE*BLOCK_SIZE pixels m sub-region image blocks, generate a sub-region image block set T i,j ={ T i,1 , T i,2 , …, T i,m },in, j Indicates j sub-regions, 1≤ j ≤ m .
4. The method for rapid positioning of remote sensing images at any position as claimed in claim 3, characterized in that: Sub-region image collection T i,j Each image in is rotated, scaled and denoised and stored in the test dataset.
5. The method for rapid positioning of remote sensing images at any position as claimed in claim 4, characterized in that: An overlapping area of overlap pixels is set between each of the sub-region image blocks, global features are extracted from each sub-region image block, and the features are stored in the FAISS vector feature library.
6. The method for rapid positioning of remote sensing images at any position as claimed in claim 5, characterized in that: The preprocessing includes calculating the minimum circumscribed quadrilateral of the effective area of the test image, and affine transforming the test image into a vertical state accordingly.
7. The method for rapid positioning of remote sensing images at any position as claimed in claim 6, characterized in that: The step 7 comprises: Assume that there is Y Test image, for y A test image, assuming its prediction box is B yp , the real frame is B yg , the predicted box and the real box are determined by the coordinates of the two vertices of the upper left corner and the lower right corner. IoU y Value is B yp and B yg The intersection area divided by B yp and B yg The calculation formula for the union area is: IoU y = ( B yp ∩ B yg ) / ( B yp ∪ B yg ) exist Y The positioning accuracy on the test images E The calculation formula is as follows: 。
Citation Information
Patent Citations
Global rapid retrieval positioning method for passive remote sensing image, storage medium and equipment
CN115630236A
Remote sensing image positioning precision evaluation method based on reference base map
CN111144350A
Remote sensing image retrieval positioning method
CN114860974A
Accurate positioning method for aerial inclined remote sensing image target
CN118570294A