Image staged matching and positioning method
By employing a phased matching and localization method, and utilizing global feature optimization and local mismatch filtering of the ViT network, the problem of matching and localizing UAV images and satellite images under large-viewpoint changes was solved, achieving high-precision image matching and localization.
Patent Information
- Application Number
- CN202511247714.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-12-23
AI Technical Summary
Matching and positioning between UAV images and satellite reference images is affected by changes in flight altitude, viewing angle, and environmental interference. Traditional methods have limited matching accuracy under large viewing angle changes and it is difficult to establish robust feature associations.
A phased matching and localization method is adopted. First, semantic association is established using the ViT network through a global feature optimization strategy. Then, mismatched point pairs are eliminated by combining local feature mismatch filtering to achieve accurate mapping.
It improves the accuracy and efficiency of image matching and localization, breaks through the limitations of traditional methods under large changes in viewing angle, and enhances the accuracy and reliability of feature matching.
Smart Images

Figure CN121190792A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image analysis, in particular to a phased matching and positioning method of images. BACKGROUND
[0002] With the rapid development of unmanned aerial vehicle technology, unmanned aerial vehicles are increasingly widely used in aerial monitoring, disaster rescue, military reconnaissance and other fields. However, in the process of shooting, due to factors such as flight height, change of viewing angle and environmental interference, images of the same scene often have significant geometric distortion, target occlusion and scale change, which brings great challenges to the matching and positioning between unmanned aerial vehicle images and satellite reference images.
[0003] Traditional image matching methods based on local features (such as SIFT, ORB, etc.) often have low feature point repeatability and high false matching rate, making it difficult to achieve stable matching in cases of large viewing angle changes. In addition, due to the significant differences in imaging conditions between satellite images and unmanned aerial vehicle images, relying solely on local feature matching cannot establish a robust feature association.
[0004] Traditional image matching methods based on convolutional neural networks (CNN) mainly rely on local receptive fields to extract features. Although they perform well under small scale changes and slight viewing angle differences, in large viewing angle change scenarios, CNNs have difficulty effectively modeling long-range dependencies between images, resulting in limited matching accuracy and affecting the stable matching and positioning of images. SUMMARY
[0005] The purpose of the present application is to provide a phased matching and positioning method for images to solve the problems raised in the background art.
[0006] To achieve the above purpose, the present application is implemented by the following technical solution: a phased matching and positioning method for images, comprising the following specific steps:
[0007] Step 1: Construct a model architecture, inputting an unmanned aerial vehicle image u;
[0008] Step 2: First stage image matching based on global feature optimization strategy, using an optimized global feature extraction and metric learning model to retrieve the most similar satellite image S i from a satellite image database, and filtering candidate images {u, S i} through global semantic feature matching;
[0009] Step 3: Second stage false matching screening based on local features, screening false matching of candidate images {u, S iThe system performs a matching operation based on local features, extracts key points and descriptive words from the image, and establishes an initial set of matching point pairs. Combining geometric consistency constraints and spatial distribution characteristics, it introduces a mismatch filtering mechanism, performs homography transformation on the geographic annotation information in the satellite image, accurately maps it to the UAV image coordinate system, and determines the location of ground targets from the UAV's perspective.
[0010] Preferably, in step two, a visual ViT network is introduced as the backbone network for feature extraction to capture long-range dependencies between image patches, perform local feature extraction and local semantic information analysis on the images extracted by the model, and input UAV or satellite images χ∈R H*W*C Divide it into N image blocks of a fixed size. Where H, W, and C represent the height, width, and number of channels of the image, respectively, and positional information is integrated into each image patch P through learnable positional embedding. i middle.
[0011] Preferably, in step two, each image block P i Through linear projection, the vectors are sequentially flattened into D-dimensional vectors, which, after processing by an encoder, form a sequence of local features of the image. Retaining effective local features, integrating the spatial information of the image, and combining them into robust initial global features G, the output sequence Z of the model feature extraction module is represented as:
[0012] Z = [G,L0,L1,...,L] N ].
[0013] Preferably, the initial global feature G is updated by introducing a gating mechanism, which calculates the linear transformation λ0 of the initial hidden state of the global feature G as the update gate. The calculation method is expressed by the following formula:
[0014] λ0 = W0G + b0;
[0015] Where W0 is the weight matrix, b0 is the bias vector, and for local features L i Its corresponding reset gate λ i The calculation formula is:
[0016] λ i =W i L i +b i ,i∈[1,N];
[0017] Among them, W i It is the weight matrix, b i It is a bias vector, and the memory weight λ is applied to the dynamic update of the global feature G, forgetting unnecessary information while retaining key feature information. Its calculation formula is:
[0018]
[0019] The update process of global feature G can be expressed by the following formula:
[0020]
[0021] Preferably, based on a global feature update strategy, the model dynamically adjusts its focus on key features under limited computational resource consumption. The loss function used for model training is the InfoNCE loss of metric learning, which is calculated by measuring the similarity between the UAV query image and the satellite reference image. The calculation formula is as follows:
[0022]
[0023] Where u represents the encoded UAV query image, s is a set of encoded satellite images, and r is a single positive sample. + Matching u, τ is a fixed temperature parameter.
[0024] Preferably, in step three, the satellite image is labeled and mapped to the UAV query image. In the pixel-level matching stage, image registration is performed using feature point matching. Based on the results, the homography matrix H between the UAV query image u and its corresponding satellite image s is calculated. The calculation formula can be expressed as u = Hs. The specific formula for the mismatch filtering method based on local features is as follows:
[0025]
[0026] Among them, (x u ,y u (x) corresponds to the coordinates of feature points in the UAV query image u. s ,y s ) represents the coordinates of the feature point in the satellite image s that matches it.
[0027] Preferably, the steps of the mismatch filtering method based on local features are as follows:
[0028] S1: Input pre-matched feature point set M, minimum number of samples n required for sampling, maximum number of iterations T, maximum allowable error σ;
[0029] S2: Calculate the local feature similarity of all combinations in M, and directly remove combinations that are below the filtering threshold;
[0030] S3: Calculate the sampling probability of the remaining combinations based on local feature similarity;
[0031] S4: Select n pairs of combinations based on the sampling probability to fit the transformation model;
[0032] S5: For the remaining combinations in M, calculate their error with the current fitted model. If it is less than σ, mark it as an interior point.
[0033] S6: Determine the number of inliers. If it is greater than the number of inliers previously retained as the optimal model, record the current fitting result.
[0034] S7: Repeat steps S3-S6 until the maximum number of iterations T or a satisfactory result is found;
[0035] S8: Output the fitted homography matrix H.
[0036] Preferably, the local feature similarity of the feature point combination in S2 is calculated by extracting the local features of the image patch where the feature point is located using an image-level matching model, and then measuring the similarity by calculating the cosine similarity. The specific formula is as follows:
[0037] ω i =cosine(L s_i, L u_i ), i∈[1,N]
[0038] Where N is the number of pre-matched feature point combinations, ω i Let s_i and u_i represent the feature similarity of the i-th pair of feature points, respectively, and let L represent the image patch indices of the i-th pair of feature points in their respective images. si and L ui The local features extracted from the image patch are represented, and outlier combinations below the filtering threshold are directly removed based on the similarity calculation results.
[0039] Preferably, in step S3, the local feature similarity of the remaining combinations is normalized to obtain the sampling probability of each pair of feature point combinations. The specific formula for the sampling probability is as follows:
[0040]
[0041] Preferably, in the transformation model calculation process in S3, the minimum combination of feature points required for fitting the model is randomly sampled according to the sampling probability, and the geometric transformation model is iteratively updated using the mismatch filtering algorithm. Finally, the homography transformation matrix is fitted to complete the annotation mapping from satellite image to UAV image.
[0042] The technical effects and advantages of this invention are as follows:
[0043] (1) This invention achieves accurate positioning of ground targets from the perspective of UAV by combining global feature optimization and local mismatch filtering. In the global matching stage, a deep feature extraction network based on ViT is used to establish semantic association between UAV query images and satellite reference images to quickly filter out the best matching candidates. In the local matching stage, high-precision feature point matching and geometric consistency verification are combined to eliminate mismatched point pairs. The geographic annotation information in the satellite image is accurately mapped to the UAV image coordinate system through homography transformation, which effectively solves the problem of image matching under large perspective changes and improves the accuracy and efficiency of image matching and positioning.
[0044] (2) This invention introduces Vision ViT as the backbone network for feature extraction. Its attention mechanism enables the model to have global perception capabilities in each layer, effectively capturing long-range dependencies between image patches, thereby breaking through the inherent limitations of CNN, improving image matching performance, introducing a mismatch filtering mechanism to effectively remove outlier point pairs, further estimating the homography matrix between images, and performing homography transformation on the geographic annotation information in satellite images accordingly, accurately mapping it to the UAV image coordinate system, thereby realizing ground target localization from the perspective of the UAV. By calculating the local feature similarity to filter feature point combinations, it can effectively remove mismatched feature point combinations, improve the accuracy and reliability of feature matching, provide more accurate input data for the calculation of the homography matrix, and improve the efficiency of matching. Attached Figure Description
[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 This is a flowchart illustrating the steps of the image phased matching and localization method of the present invention;
[0047] Figure 2 This is a flowchart of the two-stage scene matching method for multi-view images according to the present invention;
[0048] Figure 3 This is a flowchart of the image-level matching stage of the present invention;
[0049] Figure 4 This is a flowchart of the global feature optimization strategy of the present invention. Detailed Implementation
[0050] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0051] This invention provides, for example Figures 1-4 This illustrates a phased matching and localization method for images.
[0052] The specific steps include the following:
[0053] Step 1: Build the model architecture, with the input drone image u.
[0054] Step Two: The first stage involves image matching based on a global feature optimization strategy. This utilizes an optimized global feature extraction and metric learning model within the satellite image database {s1,s2,...,s...}. R} Retrieve the satellite image S that is most similar to it i Where R represents the number of satellite images, and i is the index number of the best matching image. Candidate images {u,S} are filtered through global semantic feature matching. i This ensures the accuracy and efficiency of the matching process.
[0055] To address the significant visual differences caused by variations in geometric distortion, target occlusion, and scale changes resulting from changes in imaging angle during drone image capture of the same area, a global feature optimization strategy-based image matching method in the first stage introduces ViT (Visual Transformer) as the backbone network for feature extraction. Its attention mechanism enables the model to possess global perception capabilities at each layer, capturing long-range dependencies between image patches, thereby overcoming the inherent limitations of CNNs and improving image matching performance.
[0056] The model extracts local features and analyzes local semantic information from the image. The input is a UAV or satellite image χ∈R. H*W*C Divide it into N image blocks of a fixed size. Here, H, W, and C represent the image height, width, and number of channels, respectively, enabling the model to efficiently extract local features and accurately analyze local semantic information. Learnable positional embeddings integrate positional information into each image patch P. i In this process, the model's ability to perceive spatial location information and its ability to capture spatial context are further enhanced.
[0057] Each image block P i Through linear projection, the vectors are sequentially flattened into D-dimensional vectors, which, after processing by an encoder, form a sequence of local features of the image. Retaining effective local features, integrating the spatial information of the image, and combining them into robust initial global features G, the output sequence Z of the model feature extraction module is represented as:
[0058] Z = [G,L0,L1,...,L] N ];
[0059] The model feature extraction is inspired by the time-series feature modeling mechanism in gated recurrent single-mode. Its core is to selectively remember and forget feature information at different positions in the image feature sequence, so as to capture long-range dependencies more efficiently. By introducing a gating mechanism (update gate and reset gate), the global feature G can retain image detail features while selectively updating and forgetting non-critical information to enhance its feature representation ability.
[0060] The initial global feature G is updated by introducing a gating mechanism, where the linear transformation λ0 of the initial hidden state of the global feature G is calculated as the update gate. The calculation method is expressed by the following formula:
[0061] λ0 = W0G + b0;
[0062] Where W0 is the weight matrix, b0 is the bias vector, and for local features L i Its corresponding reset gate λ i The calculation formula is:
[0063] λ i =W i L i +b i ,i∈[1,N];
[0064] Among them, W i It is the weight matrix, b i It is a bias vector, and the memory weight λ is applied to the dynamic update of the global feature G, forgetting unnecessary information while retaining key feature information. Its calculation formula is:
[0065]
[0066] The update process of global feature G can be expressed by the following formula:
[0067]
[0068] Based on a global feature update strategy, the model dynamically adjusts its focus on key features with limited computational resources. The loss function used in model training is the InfoNCE loss from metric learning, which is calculated by measuring the similarity between the UAV query image and the satellite reference image. The formula for the loss function is as follows:
[0069]
[0070] Where u represents the encoded UAV query image, s is a set of encoded satellite images, and r is a single positive sample. + Matching u, τ is a fixed temperature parameter.
[0071] Step 3: The second stage is based on local feature-based mismatch filtering, which filters candidate images {u,S}. i The system performs a matching operation based on local features, extracts key points and descriptive words from the image, and establishes an initial set of matching point pairs. Combining geometric consistency constraints and spatial distribution characteristics, a mismatch filtering mechanism is introduced to effectively eliminate outlier point pairs. The homography matrix of the image is further estimated, and the geographic annotation information in the satellite image is subjected to homography transformation to accurately map it to the UAV image coordinate system. This determines the ground target location from the UAV's perspective, thus achieving ground target location for the UAV.
[0072] By mapping satellite image annotations to UAV query images, accurate localization of ground targets from the UAV's perspective is achieved. In the pixel-level matching stage, image registration is performed using feature point matching methods (such as SIFT, SuperPoint, etc.). Based on the results, the homography matrix H between the UAV query image u and its corresponding satellite image s is calculated, expressed as u = Hs. The specific formula for the mismatch filtering method based on local features is as follows:
[0073]
[0074] Among them, (x u ,y u (x) corresponds to the coordinates of feature points in the UAV query image u. s ,y s ) represents the coordinates of the feature points in the satellite image s that match it. For example, since the homography matrix H has 8 degrees of freedom, at least 4 pairs of matching feature points are needed to solve it.
[0075] Existing feature point matching algorithms mainly employ resampling elimination methods, such as RANSAC, to screen outlier combinations. These resampling elimination methods are based on hypothesis and verification strategies. They randomly sample the minimum size subset required to fit the model from the pre-matched feature point set M to calculate the geometric transformation model. Then, they verify the deviation of all combinations in the pre-matched set from the transformation model, eliminating outlier combinations (also known as outliers) that exceed the maximum permissible error σ, and counting the number of remaining reliable combinations (also known as inliers). If the number of inliers in the current fitted model is greater than the best result of the previous iteration, the fitted model is updated, and the sampling and calculation are repeated.
[0076] Due to the significant viewpoint differences between UAV query images and satellite images, the outlier rate of the pre-matched feature point set is high. This means that the minimum subset obtained by the filtering algorithm may contain many mismatches, making it difficult to accurately fit the transformation model and thus affecting the accuracy of the homography transformation calculation between images. The high outlier rate of existing feature point sets is because the pixel-level matching methods used do not consider the spatial distribution relationship of features. When there are significant viewpoint differences between input image pairs, although the extracted feature descriptors have high similarity, the feature point combinations may not correspond in location. Therefore, it is necessary to filter feature point combinations by incorporating location information to obtain a more robust and accurate feature point set to viewpoint differences. Since an effective feature association between UAV images and satellite images has been established through metric learning during the training of the image-level matching model, and image patches retain relevant location information through location embedding during local feature extraction, pre-filtering can be performed by calculating the local feature similarity of feature point combinations.
[0077] The steps of the mismatch filtering method based on local features are as follows:
[0078] S1: Input pre-matched feature point set M, minimum number of samples n required for sampling, maximum number of iterations T, maximum allowable error σ;
[0079] S2: Calculate the local feature similarity of all combinations in M, and directly remove combinations below the filtering threshold. The local feature similarity of feature point combinations is calculated by extracting the local features of the image patch where the feature point is located using an image-level matching model, and then using cosine similarity to measure similarity. The specific formula is as follows:
[0080] ω i =cosine(L s_i, L u_i ), i∈[1,N];
[0081] Where N is the number of pre-matched feature point combinations, ω i Let s_i and u_i represent the feature similarity of the i-th pair of feature points, respectively, and let L represent the image patch indices of the i-th pair of feature points in their respective images. si and L ui The local features extracted from the image patch are represented, and outlier combinations below the filtering threshold are directly removed based on the similarity calculation results.
[0082] S3: Calculate the sampling probability of the remaining combinations based on local feature similarity. Specifically, normalize the local feature similarity of the remaining combinations to obtain the sampling probability of each pair of feature point combinations. The specific formula for the sampling probability is as follows:
[0083]
[0084] S4: Select n pairs of combinations based on the sampling probability to fit the transformation model, and randomly sample the minimum combination of feature points required to fit the model using the sampling probability.
[0085] S5: For the remaining combinations in M, calculate their error with the current fitted model. If it is less than σ, mark it as an interior point to facilitate the selection of interior points.
[0086] S6: Determine the number of inliers. If it is greater than the number of inliers that were previously retained as the optimal model, record the current fitting result. If the proportion of inliers is ≥95%, or if there is no improvement after 50 consecutive iterations, the selection of the optimal model can be terminated early.
[0087] S7: Repeat steps S3-S6 until the maximum number of iterations T or a satisfactory result is found. The maximum number of iterations T in RANSAC is 1000 by default. The satisfactory result is the optimal solution that meets the threshold for the number of interior points.
[0088] S8: Output the fitted homography matrix H, use the mismatch filtering algorithm to iteratively update the geometric transformation model, and finally fit the homography transformation matrix to complete the annotation mapping from satellite image to UAV image.
[0089] In this process, by calculating the local feature similarity to filter feature point combinations, mismatched feature point combinations can be effectively eliminated, improving the accuracy and reliability of feature matching, providing more accurate input data for the calculation of the homography matrix, and achieving the goal of accurate matching and localization of images in two stages.
[0090] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A phased matching and localization method for images, characterized in that, The specific steps include the following: Step 1: Construct the model architecture, with the input drone image u; Step Two: The first stage involves image matching based on a global feature optimization strategy. Using an optimized global feature extraction and metric learning model, the most similar satellite image S is retrieved from the satellite image database. i Candidate images {u,S} are filtered through global semantic feature matching. i }; Step 3: The second stage is based on local feature-based mismatch filtering, which filters candidate images {u,S}. i The system performs a matching operation based on local features, extracts key points and descriptive words from the image, and establishes an initial set of matching point pairs. Combining geometric consistency constraints and spatial distribution characteristics, it introduces a mismatch filtering mechanism, performs homography transformation on the geographic annotation information in the satellite image, accurately maps it to the UAV image coordinate system, and determines the location of ground targets from the UAV's perspective.
2. The method for phased matching and localization of images according to claim 1, characterized in that, In step two, a visual ViT network is introduced as the backbone network for feature extraction. This network captures long-range dependencies between image patches and performs local feature extraction and local semantic information analysis on the extracted images. The input UAV or satellite image χ∈R is used as the input. H*W*C Divide it into N image blocks of a fixed size. Where H, W, and C represent the height, width, and number of channels of the image, respectively, and positional information is integrated into each image patch P through learnable positional embedding. i middle.
3. The method for phased matching and localization of images according to claim 2, characterized in that, In step two, each image block P is... i Through linear projection, the vectors are sequentially flattened into D-dimensional vectors, which, after processing by an encoder, form a sequence of local features of the image. Retaining effective local features, integrating the spatial information of the image, and combining them into robust initial global features G, the output sequence Z of the model feature extraction module is represented as: Z=[G,L0,L1,...,L N ]。 4. The method for phased matching and localization of images according to claim 3, characterized in that, The initial global feature G is updated by introducing a gating mechanism, which calculates the linear transformation λ0 of the initial hidden state of the global feature G as the update gate. The calculation method is expressed by the following formula: λ0 = W0G + b0; Where W0 is the weight matrix, b0 is the bias vector, and for local features L i Its corresponding reset gate λ i The calculation formula is: λ i =W i L i +b i ,i∈[1,N]; Among them, W i It is the weight matrix, b i It is a bias vector, and the memory weight λ is applied to the dynamic update of the global feature G, forgetting unnecessary information while retaining key feature information. Its calculation formula is: The update process of global feature G can be expressed by the following formula:
5. The method for phased matching and localization of images according to claim 4, characterized in that, Based on a global feature update strategy, the model dynamically adjusts its focus on key features with limited computational resources. The loss function used in model training is the InfoNCE loss of metric learning, which is calculated by measuring the similarity between the UAV query image and the satellite reference image. The calculation formula is as follows: Where u represents the encoded UAV query image, s is a set of encoded satellite images, and r is a single positive sample. + Matching u, τ is a fixed temperature parameter.
6. The method for phased matching and localization of images according to claim 1, characterized in that, In step three, satellite image annotations are mapped to the UAV query image. In the pixel-level matching stage, image registration is performed using feature point matching. Based on the results, the homography matrix H between the UAV query image u and its corresponding satellite image s is calculated. The calculation formula can be expressed as u = Hs. The specific formula for the mismatch filtering method based on local features is as follows: Among them, (x u ,y u (x) corresponds to the coordinates of feature points in the UAV query image u. s ,y s ) represents the coordinates of the feature point in the satellite image s that matches it.
7. The method for phased matching and localization of images according to claim 6, characterized in that, The steps of the mismatch filtering method based on local features are as follows: S1: Input pre-matched feature point set M, minimum number of samples n required for sampling, maximum number of iterations T, maximum allowable error σ; S2: Calculate the local feature similarity of all combinations in M, and directly remove combinations that are below the filtering threshold; S3: Calculate the sampling probability of the remaining combinations based on local feature similarity; S4: Select n pairs of combinations based on the sampling probability to fit the transformation model; S5: For the remaining combinations in M, calculate their error with the current fitted model. If it is less than σ, mark it as an interior point. S6: Determine the number of inliers. If it is greater than the number of inliers previously retained as the optimal model, record the current fitting result. S7: Repeat steps S3-S6 until the maximum number of iterations T or a satisfactory result is found; S8: Output the fitted homography matrix H.
8. The method for phased matching and localization of images according to claim 7, characterized in that, The local feature similarity of the feature point combination in S2 is calculated by extracting the local features of the image patch where the feature point is located using an image-level matching model, and then measuring the similarity by calculating the cosine similarity. The specific formula is as follows: ω i <cosine(L s_i, L u_i ),i∈[1,N]; Where N is the number of pre-matched feature point combinations, ω i Let s_i and u_i represent the feature similarity of the i-th pair of feature points, respectively, and let L represent the image patch indices of the i-th pair of feature points in their respective images. si and L ui The local features extracted from the image patch are represented, and outlier combinations below the filtering threshold are directly removed based on the similarity calculation results.
9. The method for phased matching and localization of images according to claim 8, characterized in that, In step S3, the local feature similarity of the remaining combinations is normalized to obtain the sampling probability of each pair of feature point combinations. The specific formula for the sampling probability is as follows:
10. The method for phased matching and localization of images according to claim 9, characterized in that, In the transformation model calculation process in S3, the minimum combination of feature points required for fitting the model is randomly sampled according to the sampling probability. The geometric transformation model is iteratively updated using the mismatch filtering algorithm, and finally the homography transformation matrix is fitted to complete the annotation mapping from satellite image to UAV image.
Citation Information
Cited By
Automatic correction method for angles of wine bottles
CN122066770A