Multi-strategy fusion robust aerial image feature matching method and device
By introducing adaptive shadow area enhancement, multi-view simulation constraints and Affine-LightGlue algorithms in aerial image feature matching, the difficulty of matching aerial images under large-view angle changes and lighting differences is solved, and more robust and efficient feature matching is achieved.
Patent Information
- Application Number
- CN202510309187.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-13
AI Technical Summary
When processing aerial images, it is difficult to obtain high-quality feature matching results under large perspective changes, significant light differences and complex shadow conditions.
A robust aerial image feature matching method with multi-strategy fusion is adopted to improve the robustness and efficiency of feature matching through technologies such as adaptive shadow area enhancement, multi-view simulation constraints and Affine-LightGlue algorithm.
It significantly improves the feature matching ability under complex imaging conditions, enhances the robustness and accuracy of image matching, and provides a more sufficient data basis for subsequent task processing.
Smart Images

Figure CN120147673A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly to a robust aerial image feature matching method and device with multi-strategy fusion. Background Art
[0002] Aerial images have the characteristics of large imaging range and high imaging resolution, and have become an important data source in the fields of surveying and mapping and remote sensing. Image matching, as an upstream task, is the basis and prerequisite for downstream tasks such as image stitching, structure from motion (SfM), and dense scene three-dimensional reconstruction. Improving the matching quality and robustness of images plays an extremely important role in the above tasks. Aerial image feature matching faces multiple challenges in practical applications: on the one hand, the inclination difference between flight lines caused by the attitude change of the aerial photography platform makes it difficult for traditional feature-based matching algorithms to obtain robust matching results; on the other hand, the changes in light intensity and direction caused by different shooting times, cloud occlusion, or ground object reflection characteristics during the imaging process will lead to problems such as local contrast imbalance in the image and loss of texture information in shadow areas.
[0003] Feature matching is mainly divided into a feature extraction-based matching algorithm and a network-based matching algorithm according to the implementation principle. The feature extraction-based matching algorithm mainly relies on manually designed feature descriptors, extracts local features in the image, and uses measures such as sum of absolute differences and normalized correlation coefficient for matching. The network-based matching algorithm uses deep learning methods to directly learn the features and matching probabilities of images through neural networks, and obtains the matching relationship between feature points.
[0004] Early feature extraction-based matching algorithms mainly utilized the grayscale information of images, such as the Moravec operator, Harris operator, etc. These methods lack robustness in terms of scale, rotation transformation, and perspective changes, and are relatively sensitive to illumination changes. To overcome the above limitations, Lowe proposed the Scale-invariant Feature Transform (SIFT) algorithm. SIFT constructs a scale space and retrieves extreme points to generate high-dimensional descriptor information, effectively solving the problems of perspective and illumination differences. Bay et al. proposed the Speeded Up Robust Features (SURF) algorithm based on SIFT. SURF constructs a scale space by adjusting the template size of the box filter and uses Harr wavelet features to generate 64-dimensional descriptors, improving the detection speed of feature points. In terms of feature detection, the FAST algorithm proposed by Alahi et al. determines corner points by analyzing the grayscale value relationship between the central pixel point and its neighboring pixel points. Rublee et al. combined FAST feature detection with the binary BRIEF feature descriptor to propose the ORB algorithm. This algorithm uses the moment estimation method to determine the direction of FAST feature points and realizes the rotation invariance of BRIEF features by rotating the generated BRIEF features to the direction of the feature points. It should be noted that traditional SIFT and its improved algorithms extract feature points by constructing a linear scale space, which may lead to the loss of detailed information and edge blurring in the image.
[0005] With the rapid development of deep learning, network-based matching methods have flourished. According to the processing process, they can be mainly divided into two categories: end-to-end matching methods and non-end-to-end matching methods. End-to-end matching methods can effectively capture repetitive key points in image pairs by constructing a dense grid on the image and directly extracting descriptors. However, for aerial images, there are a large number of repetitive textures, weak texture areas, and complex textures in the images, and end-to-end matching methods have problems of slow calculation speed and large computational volume. In contrast, non-end-to-end matching methods separate feature extraction and matching into two independent modules. Using a separately designed feature extraction network can obtain more robust descriptors, and the separated structure has higher computational efficiency. Based on network-based feature extraction, although neural networks can automatically capture various complex local features, they lack the robustness and consistency of hand-designed descriptor methods when dealing with large-scale affine transformations, and their generalization ability is poor, and they may completely fail for different scenarios.
[0006] According to the above feature matching methods, in the face of large perspective changes, significant illumination differences, and complex shadow conditions, it is often difficult for existing technologies to obtain sufficient and high-quality matching results. Summary of the Invention
[0007] In view of the above problems, the present invention proposes a robust aerial image feature matching method and device that integrates multiple strategies. By introducing strategies such as adaptive shadow region enhancement and multi-view simulation constraints of the matching network, the feature matching ability under complex imaging conditions is effectively improved.
[0008] In a first aspect, a robust aerial image feature matching method integrating multiple strategies proposed by the present invention includes:
[0009] Step 1: Perform adaptive shadow region enhancement on two aerial images to be matched to obtain a first image and a second image;
[0010] Step 2: Perform affine transformation on the enhanced first image and second image respectively to obtain multi-view simulation diagrams; wherein the multi-view simulation diagrams include a third image and a fourth image;
[0011] Step 3: Extract feature points of the first image and the second image based on the multi-view simulation diagrams, the first image, and the second image, and perform feature matching to obtain a preliminary feature matching result;
[0012] Step 4: Perform K-Means clustering on the preliminary feature matching result, filter outlier points based on the clustering result, and then perform matching optimization using the RANSAC algorithm to obtain a final feature matching result.
[0013] Further, in Step 1, the adaptive shadow region enhancement specifically includes:
[0014] Extract the shadow region on the aerial image to obtain a shadow region mask, and perform feathering processing on the shadow region mask;
[0015] Perform adaptive brightness enhancement on different shadow regions on the aerial image by calculating a brightness enhancement factor;
[0016] Perform Gaussian filtering on the aerial image after brightness enhancement.
[0017] Further, the extraction of the shadow region on the aerial image specifically includes:
[0018] Calculate the color value of the normalized color space of the aerial image, and the color value is calculated by the following formula:
[0019] r = R / (R + G + B)
[0020] g = G / (R + G + B)
[0021] b = B / (R + G + B)
[0022] Wherein, r represents the red value in the normalized color space, g represents the green value in the normalized color space, b represents the blue value in the normalized color space, R represents the value of the red channel, G represents the value of the green channel, and B represents the value of the blue channel;
[0023] The preliminary decision value is obtained by calculating the absolute value difference between the color space and the color value of the normalized color space, and is calculated by the following formula:
[0024]
[0025] Wherein, D 1 represents the preliminary decision value;
[0026] The shadow decision value is calculated based on the brightness of different color sensitivities and the preliminary decision value:
[0027] D = αD 1 -βD 2
[0028] Wherein, D represents the shadow decision value, α and β represent proportionality coefficients, and D 2 represents the brightness of different color sensitivities, and the calculation method is as follows:
[0029] D 2 = 0.46R + 0.50G + 0.04B;
[0030] Set the shadow threshold T, and determine the area where the shadow decision value is greater than T as the shadow area to obtain the shadow area mask.
[0031] Furthermore, the calculation formula of the brightness enhancement factor is as follows:
[0032] B = B f *(1 - G / 255.0)
[0033] Wherein, G represents the gray value of the shadow area, and B f represents the brightness factor of fixed enhancement.
[0034] Furthermore, in step 2, the affine transformation is calculated by the following formula:
[0035] I (1) = AI (7)
[0036] Wherein, I represents the aerial image, and I (1) represents the aerial image after affine transformation, and A represents the affine matrix, which is calculated by the following formula:
[0037]
[0038] Among them, α represents the rotation angle of the image; x and y represent the image information, that is, the offset in the horizontal and vertical directions, and τ represents the tilt coefficient. The calculation method is as follows:
[0039]
[0040] Among them, θ represents the tilt angle.
[0041] Furthermore, the SIFT algorithm is used to extract feature points.
[0042] Furthermore, Affine-LightGlue is used for feature matching; based on LightGlue, Affine-LightGlue deletes the depth adaptation and pruning strategies;
[0043] Correspondingly, the feature matching based on the extracted feature points specifically includes:
[0044] Determine the feature point d of the first image 0 and the feature point d of the second image 1 , and perform position encoding based on the coordinates of the feature points to obtain and
[0045] Through the self-attention mechanism, perform feature enhancement on and to obtain f s 0 and f s 1 ;
[0046] Through the cross-attention mechanism, perform feature interaction enhancement on f s 0 and f s 1 to obtain the optimized features and
[0047] Repeat the self-attention mechanism feature enhancement and the cross-attention mechanism feature interaction enhancement n times;
[0048] Calculate the pairwise score matrix S of the enhanced feature points between the two images ij and the matchability score σ of each point i , and combine the score matrix S ij and the matchability score σ i into the assignment matrix P, specifically as follows
[0049]
[0050] σ i= Sigmoid(Linear(x i )) ∈ [0, 1] (19)
[0051] Wherein, Linear(·) represents a learnable linear transformation, A represents the first image, B represents the second image, and P represents an assignment matrix composed of all feature points between the first image and the second image. represents the matchability score of the i-th feature point in the first image. represents the matchability score of the j-th feature point on the second image. represents the i-th feature point on the first image. represents the j-th feature point on the second image;
[0052] Feature matching points between the first image and the second image are obtained according to the assignment matrix P.
[0053] Further, step 4 specifically includes:
[0054] Step 4.1: Perform K-Means clustering on the feature matching points of the first image and the second image in the preliminary feature matching results respectively to obtain C P and C Q ;
[0055] Step 4.2: Use C P to perform clustering on the corresponding feature matching points of Q to obtain Use C Q to perform clustering on the corresponding feature matching points of P to obtain
[0056] Compare C P and C Q and Select the matching points that simultaneously satisfy the clustering constraints of the first image and the second image, and eliminate the outliers to obtain the matching point clustering results C L and C R ;
[0057] Step 4.3: Use the RANSAC algorithm to estimate the homography relationship between the cluster matching points in C L and C R to obtain the homography matrix H i (i = 1, 2, …, k) and the inliers corresponding to each cluster and
[0058] By calculating the difference between the pixel value after being transformed by the corresponding homography matrix H i and the pixel value of The formula is as follows:
[0059]
[0060] Among them, I(·) represents the pixel value of this point, represents the point after transformation by the corresponding homography matrix H i ;
[0061] For any point in any cluster, if △I > ε, then this point is considered an obvious wrong matching point, and this point is deleted from the matching points to obtain the updated matching points;
[0062] Step 4.4: Re - conduct RANSAC estimation based on the updated matching points to obtain a new homography matrix and inliers, and calculate the average pixel error of all points:
[0063]
[0064] Among them, I ave represents the average pixel error, and X i represents the updated matching points;
[0065] Step 4.5: Re - conduct K - Means clustering on the updated matching points, repeat the loop steps 4.1 - 4.4, and increase the clustering number in turn. Take the clustering number corresponding to the minimum I ave as the final K value;
[0066] Use the new K value to re - conduct K - Means clustering and outlier screening on the feature matching result after outlier removal to obtain the final feature matching result.
[0067] In the second aspect, a robust aerial image feature matching device with multi - strategy fusion proposed by the present invention includes:
[0068] An adaptive shadow area enhancement module, which is used to perform adaptive shadow area enhancement on two aerial images to be matched to obtain a first image and a second image;
[0069] A multi - perspective simulated image generation module, which is used to perform affine transformation on the enhanced first image and second image respectively to obtain multi - perspective simulation diagrams; among them, the multi - perspective simulation diagrams include a third image and a fourth image;
[0070] A feature extraction and matching module, which extracts feature points of the first image and the second image based on the multi - perspective simulation diagrams, the first image and the second image, and performs feature matching to obtain a preliminary feature matching result;
[0071] A matching result optimization module is used to perform K-Means clustering on the preliminary feature matching result. After screening out outliers based on the clustering result, the RANSAC algorithm is used for matching optimization to obtain the final feature matching result.
[0072] The beneficial effects of the present invention are as follows:
[0073] The present invention proposes a robust aerial image feature matching method with multi-strategy fusion, which effectively solves the problem of few and unevenly distributed matching points during matching due to shadow differences and perspective differences during image imaging. First, aiming at the problem of insufficient ground object details in the shadow area, by analyzing the imaging characteristics of the shadow, the shadow area is extracted and enhanced, significantly restoring the ground object detail information in the shadow area. Secondly, aiming at the problem of difficult matching caused by perspective differences, by simulating the attitude change during camera imaging, affine simulation transformation is performed on the image to obtain richer and more robust feature points. Then, efficient feature point matching is carried out through the Affine-LightGlue algorithm. Finally, the clustering RANSAC optimization strategy is used to optimize the matching points. By clustering the matching points and combining the RANSAC algorithm for local optimization of each cluster, not only the pixel distance error of the matching points is reduced, but also more inlier numbers are obtained, providing a more sufficient data basis for subsequent task processing. Description of the Drawings
[0074] Figure 1 It is a schematic flowchart of a robust aerial image feature matching method with multi-strategy fusion provided by an embodiment of the present invention;
[0075] Figure 2 It is a schematic flowchart of shadow enhancement processing provided by an embodiment of the present invention;
[0076] Figure 3 It is a schematic simulation diagram of affine transformation provided by an embodiment of the present invention;
[0077] Figure 4 It is a schematic architecture diagram of the Affine-LightGlue model provided by an embodiment of the present invention;
[0078] Figure 5 It is a schematic process diagram of outlier removal provided by an embodiment of the present invention;
[0079] Figure 6 It is a schematic process diagram of homography transformation provided by an embodiment of the present invention. Detailed Embodiment
[0080] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0081] As Figure 1 shown, an embodiment of the present invention provides a robust aviation image feature matching method integrating multiple strategies, including:
[0082] Step 1: Perform adaptive shadow area enhancement on two aviation images to be matched to obtain a first image and a second image;
[0083] Step 2: Perform affine transformation on the enhanced first image and second image respectively to obtain multi-view simulation images; where the multi-view simulation images include a third image and a fourth image;
[0084] Step 3: Extract feature points of the first image and the second image based on the multi-view simulation images, the first image, and the second image, and perform feature matching to obtain a preliminary feature matching result;
[0085] Step 4: Perform K-Means clustering on the preliminary feature matching result, screen out outliers based on the clustering result, and then use the RANSAC algorithm for matching optimization to obtain a final feature matching result.
[0086] In the embodiment of the present invention, the shadow area is first extracted based on the original information of the image and subjected to adaptive enhancement processing to fully restore the ground object detail information in the occluded area; subsequently, combined with the attitude information during image imaging, multi-view simulation images are constructed through affine transformation, thereby enhancing the adaptability of the matching network to view changes to obtain a more robust feature matching result; then, feature extraction is performed on the multi-view simulation images and aviation images, and the optimized LightGlue algorithm is used for feature matching to obtain an initial matching result. Finally, the K-Means algorithm is used to cluster the matching results, and the RANSAC algorithm is used to optimize each cluster of matching points, effectively removing false matches while retaining correct matching points, and finally obtaining a high-precision and sufficient number of matching results.
[0087] Based on the above embodiments, this embodiment provides an adaptive shadow area enhancement method, as Figure 2 shown, specifically including:
[0088] Extract the shadow area on the aviation image to obtain a shadow area mask, and perform feathering processing on the shadow area mask;
[0089] Specifically, calculate the color values in the normalized color space of the aerial image, and the color values are calculated by the following formula:
[0090] r = R / (R + G + B)
[0091] g = G / (R + G + B) (1)
[0092] b = B / (R + G + B)
[0093] Among them, r represents the red value in the normalized color space, g represents the green value in the normalized color space, b represents the blue value in the normalized color space, R represents the value of the red channel, G represents the value of the green channel, and B represents the value of the blue channel. It can be understood that the reason for the generation of shadows is that there is an object blocking between the light source and the illuminated area, so that the light cannot directly irradiate this area, resulting in insufficient or missing illumination in the area. Therefore, the present invention first processes the image in the normalized color space, and realizes the accurate characterization of the shadow area by analyzing the difference between the normalized color space and the original color space.
[0094] Obtain the preliminary decision value by calculating the absolute value difference between the color values of the color space and the normalized color space, and calculate it by the following formula:
[0095]
[0096] Among them, D 1 represents the preliminary decision value;
[0097] Calculate the shadow decision value based on the brightness with different color sensitivities and the preliminary decision value:
[0098] D = αD 1 -βD 2 (3)
[0099] Among them, D represents the shadow decision value, α and β represent proportionality coefficients, and D 2 represents the brightness with different color sensitivities. Due to the influence of occlusion, the brightness of the shadow area is usually lower than that of the non-shadow area, and the shadow area and the non-shadow area are distinguished based on the human eye's perception ability of brightness. In view of the differences in the sensitivity of the human eye to the three colors of red, green, and blue, the present invention designs a brightness based on different color sensitivities, and the calculation method is as follows:
[0100] D 2 = 0.46R + 0.50G + 0.04B (4);
[0101] Set the shadow threshold T, and determine the area where the shadow decision value is greater than T as the shadow area to obtain the shadow area mask M s .
[0102]
[0103] If direct enhancement processing is carried out, it will cause an obvious demarcation line to appear between the shadow area and the non-shadow area, thereby introducing additional features into the original image and affecting the subsequent matching effect of the image. Therefore, the present invention performs a feathering process on the shadow area mask to weaken the transition of the boundary line and make it more natural and smooth.
[0104] Adaptive brightness enhancement is performed on different shadow areas of the aerial image by calculating the brightness enhancement factor;
[0105] Specifically, since the brightness distribution of the shadow area in the aerial image varies at different positions, when performing brightness enhancement, the brightness is adaptively adjusted according to the brightness difference at each position in the shadow area to achieve a more natural enhancement effect. The calculation method of the brightness enhancement factor is as follows:
[0106] B = B f *(1 - G / 255.0) (6)
[0107] where G represents the gray value of the shadow area, and B f represents the brightness factor of fixed enhancement, which is set to 5 in the embodiment of the present invention.
[0108] Finally, Gaussian filtering processing is performed on the brightness-enhanced aerial image to eliminate the aerial image noise and further smooth the transition between the shadow area and the non-shadow area.
[0109] During the shooting process of the aerial image, due to the change of the lens angle during flight and the relative position relationship between the aircraft and the object to be photographed, different degrees of shadows usually occur in the image, resulting in the weakening of the ground object features in the shadow area, thereby affecting the accuracy and efficiency of feature matching. The adaptive shadow area enhancement algorithm proposed in the embodiment of the present invention can effectively enhance the ground object features in the shadow area and improve the efficiency and accuracy of matching.
[0110] On the basis of the above embodiment, the present embodiment provides a method for generating a multi-view simulation diagram. By simulating all possible imaging situations caused by the change of the camera principal optical axis from the orthographic direction, different affine transformations are performed on each image, such as Figure 3 shown, including the rotation angle α in the direction perpendicular to the principal optical axis plane and the angle θ between the tilted camera principal optical axis and the normal of the phase plane. Specifically, the first image and the second image are subjected to affine transformation through formula (7) to generate a multi-view simulation diagram;
[0111] I (1) = AI (7)
[0112] where I represents the aerial image, and I (1)Denote the aerial image after affine transformation as \(\widetilde{I}\), and \(A\) represents the affine matrix, which consists of rotation transformation and skew transformation. For image rotation, to ensure that the rotated image is not cropped, calculate the coordinates of the four corners of the image after rotation through the rotation matrix, find the smallest rectangle enclosing these points, determine the offset \(x\) and \(y\) to be translated, and form the preliminary affine matrix \(A\) with the rotation matrix \(R\). 0 , which is calculated by the following formula:
[0113]
[0114] For skew transformation, following the processing of the ASIFT algorithm, first perform Gaussian blurring on the image to reduce the distortion caused by scaling. By performing a scaling operation on the image, the scaling factor is f y = 1.0, and finally update the affine matrix \(A\):
[0115]
[0116] where \(\alpha\) represents the rotation angle, which is set to 180 degrees in this embodiment to simulate the rotation change between images with overlapping regions in adjacent flight lines; \(\tau\) represents the skew coefficient, and the calculation method is as follows:
[0117]
[0118] where \(\theta\) represents the skew angle, which is determined by the swing range of the camera.
[0119] During the flight of the UAV, due to changes in the shooting angle, altitude, and attitude, the same ground object may present different forms in different images, resulting in incorrect matching results and low robustness. In this embodiment, multi-view simulation is performed to generate multi-view simulation diagrams, accurately reflecting these changes to constrain the matching network and effectively optimizing the matching results.
[0120] Based on the above embodiments, the embodiments of the present invention use the SIFT algorithm to extract feature points. The SIFT feature extraction constructs a feature description with certain scale and rotation invariance by simulating the feature response in a multi-scale space and combining the design of the gradient direction. The SIFT algorithm provided by the embodiments of the present invention specifically includes:
[0121] By simulating and constructing a Gaussian pyramid to simulate the characteristics of the human eye's multi-scale observation, for the input image \(I(x,y)\), use Gaussian convolution kernels with different standard deviations \(\sigma\) to construct a scale space:
[0122] \(L(x,y,\sigma)=G(x,y,\sigma)*I(x,y)\) (11)
[0123]
[0124] Among them, G(x, y, σ) represents the Gaussian kernel function, and (x, y) represents the coordinates of each pixel in the image;
[0125] To effectively extract stable key points in the image, a Difference of Gaussian (DoG) pyramid is further constructed:
[0126] D(x, y, σ) = [G(x, y, kσ) - G(x, y, σ)] * I(x, y) (13)
[0127] Among them, k represents the scale multiplication factor; by searching for local extrema in the DoG pyramid space, the positions and scales of candidate key points are initially determined.
[0128] Regarding the possible errors in the extreme value coordinates obtained from the discrete sampling space, the SIFT algorithm adopts a three-dimensional quadratic function fitting strategy. By performing a Taylor expansion on the DoG function, sub-pixel coordinate localization is achieved. The Taylor expansion fitting formula of the D0G function is as follows:
[0129]
[0130] Among them, X = (x, y, σ). By solving the offset with a derivative of zero, the sub-pixel key point coordinates can be obtained. Further, unstable points are filtered using a contrast threshold and edge response rejection (based on the eigenvalue analysis of the Hessian matrix).
[0131] To ensure the rotation invariance of the features, SIFT is achieved by statistically analyzing the gradient direction distribution of pixels in the neighborhood of the key point. A window of size 16σ × 16σ is taken with the key point as the center, and the gradient magnitude m(x, y) and direction θ(x, y) of each pixel are calculated:
[0132]
[0133] The 360° direction range is divided into 36 intervals (each interval is 10°), and the gradient directions of the pixels in the neighborhood of the key point are assigned to the corresponding intervals to generate a gradient direction histogram. The highest peak in the histogram is selected as the main direction of the key point. At the same time, if the auxiliary peak exceeds 80% of the main peak, multiple directions are created for this key point.
[0134] In the coordinate system of the main direction of the key point, the 16σ × 16σ neighborhood is divided into 4 × 4 sub-regions. In each sub-region, a gradient direction histogram of 8 directions is calculated, and finally a 128-dimensional (4 × 4 × 8) feature vector is formed.
[0135] Correspondingly, based on the multi-view simulation diagram, the first image and the second image, the feature points of the first image and the second image are extracted, specifically including:
[0136] Extract the feature points of the third image, the fourth image, the first image, and the second image respectively;
[0137] Inverse affine transform the feature points of the third image back onto the first image, and jointly form the feature points of the first image with the feature points of the first image;
[0138] Inverse affine transform the feature points of the fourth image back onto the second image, and jointly form the feature points of the second image with the feature points of the second image.
[0139] LightGlue creates a graph neural network (GNN) with feature points as nodes, and exchanges global visual and geometric information between nodes through self-attention layers and cross-attention layers to achieve the purpose of feature aggregation. Then LightGlue uses a lightweight assignment matrix to predict the matching results. To reduce the training difficulty and complexity, LightGlue proposes a confidence classifier to achieve network depth adaptation and a pruning strategy to achieve network width adaptation. Since the depth adaptation mechanism may cause loss of some key point features, and the pruning strategy may delete some matching points, affecting the matching accuracy. To address this problem, the embodiment of the present invention proposes the Affine-LightGlue method based on the LightGlue method. By deleting the depth adaptation and pruning strategies, more feature point information is retained, and the robustness of feature matching is improved. The model architecture of Affine-LightGlue is as Figure 4 shown.
[0140] Correspondingly, feature matching is performed based on the extracted feature points, specifically including:
[0141] Determine the feature points d 0 of the first image and the feature points d 1 of the second image, and perform position encoding based on the coordinates of the feature points to obtain and
[0142] Perform feature enhancement on and through the self-attention mechanism to obtain f s 0 and f s 1 ;
[0143] Perform feature interaction enhancement on f s 0 and f s 1 through the cross-attention mechanism to obtain the optimized features and
[0144] Repeat the self-attention mechanism feature enhancement and cross-attention mechanism feature interaction enhancement n times;
[0145] Calculate the pairwise score matrix S of the enhanced feature points between the two images ij and the matchability score σ of each point i , and combine the score matrix S ij and the matchability score σ i into the assignment matrix P, which is specifically as follows:
[0146]
[0147] σ i = Sigmoid(Linear(x i )) ∈ [0, 1] (19)
[0148] where Linear(·) represents a learnable linear transformation, A represents the first image, B represents the second image, and P represents the assignment matrix composed of all feature points between the first image and the second image; represents the matchability score of the i-th feature point in the first image, indicating the possibility that the i-th feature point in the first image has a corresponding matching point in the second image; represents the matchability score of the j-th feature point on the second image, indicating the possibility that the j-th feature point in the second image has a corresponding matching point in the first image; represents the i-th feature point on the first image, represents the j-th feature point on the second image.
[0149] Obtain the feature matching points between the first image and the second image according to the assignment matrix P.
[0150] Specifically, obtaining the assignment matrix P reflects the matching probability between all feature points in the first image and the second image. When the matching probability of a pair of points is greater than the threshold and it is the maximum among any other matching probabilities along its row and column, then this pair of points is predicted to be mutually matched.
[0151] Based on the above embodiments, step 4 specifically includes:
[0152] Step 4.1: Perform K-Means clustering on the feature matching points of the first image and the second image in the preliminary feature matching result respectively to obtain C P and C Q .
[0153] Specifically, the K-Means clustering method is implemented as follows:
[0154] Determine the feature matching points P = {p of the first image and the second image in the preliminary feature matching result1 , p 2 , …, p n}, and Q = {q 1 , q 2 , …, q n}, determine the search range of the K value according to the number of feature matching points and the ground object features of the image, formulate a k value within the search range, and randomly select k centroids C = {c 1 , c 2 , …, c k};
[0155] For each feature point p i , calculate its distances to all centroids c j . If the distance from p i to the centroid c j is the smallest compared to the distances to other centroids, then set the indicator variable a ij to 1, indicating that p i is assigned to the cluster corresponding to c j as the center; otherwise, set it to 0, indicating not to be assigned to this cluster; the formula is as follows:
[0156]
[0157] After each assignment of feature points, calculate the mean of all points in each cluster to update the centroid:
[0158]
[0159] Repeat the feature point assignment using the updated centroid, and stop the iteration when the change of all centroids is less than the threshold ε or reaches the maximum number of iterations, that is Obtain and
[0160] Step 4.2: Use C P to perform clustering on the corresponding feature matching points of Q, and obtain Use C Q to perform clustering on the corresponding feature matching points of P, and obtain
[0161] Compare C P and C Q and respectively, screen out the matching points that satisfy the clustering constraints of both the first image and the second image, and remove the outliers to obtain the matching point clustering results C L and C R ;
[0162] As Figure 5 shown, compare C Pand C Q and Filter the matching points that simultaneously satisfy the clustering constraints of the left and right images, remove the outliers, and obtain the clustering result C of the matching points L and C R .
[0163] Step 4.3: Use the RANSAC algorithm to estimate the homography relationship between the matching points in C L and C R to obtain the homography matrix H i (i = 1, 2, …, k) and the inliers corresponding to each cluster and
[0164] By calculating the difference between the pixel value after transformation by the corresponding homography matrix H i and the pixel value of as shown in Figure 6 The formula is as follows:
[0165]
[0166] where I(·) represents the pixel value of this point, represents the point after transformation by the corresponding homography matrix H i ; Since the coordinates of the matching points are usually not integers, in the embodiments of the present invention, bilinear interpolation is used to calculate the pixel value of each matching point. Specifically, the pixel values of the four nearest neighbor integer coordinate points around each matching point are weighted and averaged to obtain the accurate pixel value of the non-integer coordinate point;
[0167] For any point in any cluster, if △I > ε, where ε represents the set threshold, then it is considered that this point is an obvious incorrect matching point, and this point is deleted from the matching points to obtain the updated matching points, and RANSAC estimation is performed again to obtain a new homography matrix and inliers, and the average pixel error of all points is calculated:
[0168]
[0169] where, I ave represents the average pixel error, and X i represents the updated matching points.
[0170] Step 4.5: Recluster the updated matching points by K-Means, repeat the loop steps 4.1~4.4, and sequentially increase the number of clusters. Take the number of clusters corresponding to the minimum I ave as the final K value;
[0171] Re - perform K - Means clustering on the feature matching results after outlier removal using the new K value and screen outliers to obtain the final feature matching results.
[0172] An embodiment of the present invention also provides a robust aviation image feature matching device with multi - strategy fusion, including:
[0173] An adaptive shadow area enhancement module, used to perform adaptive shadow area enhancement on two aviation images to be matched, obtaining a first image and a second image;
[0174] A multi - perspective simulation image generation module, used to perform affine transformation on the enhanced first image and second image respectively to obtain multi - perspective simulation images; where the multi - perspective simulation images include a third image and a fourth image;
[0175] A feature extraction and matching module, used to extract feature points of the first image and the second image based on the multi - perspective simulation images, the first image and the second image, and perform feature matching to obtain a preliminary feature matching result;
[0176] A matching result optimization module, used to perform K - Means clustering on the preliminary feature matching result, screen outliers based on the clustering result, and then perform matching optimization using the RANSAC algorithm to obtain the final feature matching result.
[0177] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-strategy fusion robust aerial image feature matching method, characterized in that: include: Step 1: Adaptively enhance the shadow areas of the two aerial images to be matched to obtain a first image and a second image; Step 2: performing affine transformation on the enhanced first image and the second image respectively to obtain a multi-view simulation image; wherein the multi-view simulation image includes the third image and the fourth image; Step 3: extracting feature points of the first image and the second image based on the multi-view simulation image, the first image and the second image, and the first image and the second image, and performing feature matching to obtain a preliminary feature matching result; Step 4: K-Means clustering is performed on the preliminary feature matching results. After outliers are screened based on the clustering results, the RANSAC algorithm is used to optimize the matching to obtain the final feature matching results.
2. According to claim 1, a multi-strategy fusion robust aerial image feature matching method is characterized in that: In step 1, the adaptive shadow area enhancement specifically includes: Extracting the shadow area on the aerial image to obtain a shadow area mask, and performing feathering processing on the shadow area mask; Adaptively enhancing the brightness of different shadow areas on the aerial image by calculating a brightness enhancement factor; Gaussian filtering is performed on the aerial image after brightness enhancement.
3. According to claim 2, a multi-strategy fusion robust aerial image feature matching method is characterized in that: The extracting of the shadow area on the aerial image specifically includes: The color value of the normalized color space of the aerial image is calculated, and the color value is calculated by the following formula: Wherein, r represents the red value of the normalized color space, g represents the green value of the normalized color space, b represents the blue value of the normalized color space, R represents the value of the red channel, G represents the value of the green channel, and B represents the value of the blue channel; The initial decision value is obtained by calculating the absolute value difference between the color space and the color value of the normalized color space, and is calculated by the following formula: Among them, D1 represents the initial decision value; Calculate the shadow decision value based on the brightness of the different color sensitivities and the preliminary decision value: D=αD1-βD2 (3) Where D represents the shadow decision value, α and β represent the proportional coefficients, and D2 represents the brightness of different color sensitivities. The calculation method is as follows: D2=0.46R+0.50G+0.04B (4); A shadow threshold T is set, and the area where the shadow decision value is greater than T is determined as a shadow area, thereby obtaining a shadow area mask.
4. The multi-strategy fusion robust aerial image feature matching method according to claim 2 is characterized in that: The brightness enhancement factor calculation formula is as follows: B=B f *(1-G / 255.0) (6) Among them, G represents the gray value of the shadow area, B f Indicates a fixed enhancement brightness factor.
5. The multi-strategy fusion robust aerial image feature matching method according to claim 1, characterized in that: In step 2, the affine transformation is calculated by the following formula: and (1) =AI (7) Where I represents aerial image, I (1) represents the aerial image after affine transformation, A represents the affine matrix, which is calculated by the following formula: Among them, α represents the rotation angle of the image; x and y represent the image information, that is, the offset in the horizontal and vertical directions; τ represents the tilt coefficient, which is calculated as follows: Here, θ represents the tilt angle.
6. The multi-strategy fusion robust aerial image feature matching method according to claim 1, characterized in that: The SIFT algorithm is used to extract feature points.
7. The multi-strategy fusion robust aerial image feature matching method according to claim 1, characterized in that: Affine-LightGlue is used for feature matching; based on LightGlue, Affine-LightGlue removes the depth adaptation and pruning strategies; Correspondingly, the feature matching based on the extracted feature points specifically includes: Determine the feature point d of the first image 0 and the feature point d of the second image 1 , and the position encoding is performed based on the coordinates of the feature points and Through the self-attention mechanism and Perform feature enhancement to obtain and Through the cross attention mechanism and Perform feature interaction enhancement to obtain optimized features and Repeat the self-attention mechanism feature enhancement and cross-attention mechanism feature interaction enhancement n times; Calculate the pairwise score matrix S of the enhanced feature points between the two images ij And the matching score σ of each point i , the score matrix S ij and the compatibility score σ i Combined into a distribution matrix P, as shown below: σ i =Sigmoid(Linear(x i ))∈[0,1] (19) Where Linear(·) represents a learnable linear transformation, A represents the first image, B represents the second image, and P represents the distribution matrix composed of all feature points between the first image and the second image. represents the matchability score of the i-th feature point in the first image, represents the matchability score of the jth feature point on the second image, represents the i-th feature point on the first image, represents the jth feature point on the second image; According to the allocation matrix P, feature matching points between the first image and the second image are obtained.
8. The multi-strategy fusion robust aerial image feature matching method according to claim 1, characterized in that: The step 4 specifically includes: Step 4.1: Perform K-Means clustering on the feature matching points of the first image and the second image in the preliminary feature matching results to obtain C P and C Q ; Step 4.2: Using C P Cluster the feature matching points corresponding to Q and get Using C Q Cluster the feature matching points corresponding to P and get Compare C P and C Q and Filter the matching points that satisfy the clustering constraints of the first image and the second image, remove the outliers, and obtain the matching point clustering result C L and C R ; Step 4.3: Estimate C using the RANSAC algorithm L and C R The homography relationship between the matching points of each cluster in is obtained to obtain the homography matrix H i (i=1,2,…,k) and the corresponding internal points of each cluster and By calculation Through the corresponding homography matrix H i The pixel value after transformation is The difference between the pixel values is calculated as follows: Among them, I(·) represents the pixel value of the point, express Through the corresponding homography matrix H i The transformed point; For any point in any cluster, if △I>ε, the point is considered to be an obvious wrong matching point, and the point is deleted from the matching points to obtain the updated matching point; Step 4.4: Based on the updated matching points, re-perform RANSAC estimation to obtain a new homography matrix and inliers, and calculate the average pixel error of all points: Among them, I ave represents the average pixel error, X i Represents the updated matching point; Step 4.5: Re-cluster the updated matching points using K-Means, and repeat steps 4.1 to 4.4, increasing the number of clusters in turn. ave The number of clusters corresponding to the minimum is taken as the final K value; The new K value is used to re-perform K-Means clustering and filter outliers on the feature matching results after outliers are removed to obtain the final feature matching results.
9. A multi-strategy fusion robust aerial image feature matching device, characterized in that: include: An adaptive shadow area enhancement module is used to adaptively enhance the shadow areas of two aerial images to be matched to obtain a first image and a second image; A multi-view simulated image generation module, used to perform affine transformation on the enhanced first image and the second image respectively to obtain a multi-view simulated image; wherein the multi-view simulated image includes a third image and a fourth image; A feature extraction and matching module extracts feature points of the first image and the second image based on the multi-view simulation image, the first image and the second image, and performs feature matching to obtain a preliminary feature matching result; The matching result optimization module is used to perform K-Means clustering on the preliminary feature matching results, filter outliers based on the clustering results, and then use the RANSAC algorithm to perform matching optimization to obtain the final feature matching results.