Real-time image feature matching method based on spatial distribution consistency
By analyzing the spatial distribution characteristics of motion vectors of image matching point pairs, adopting a "coarse-fine dual-stage screening + K-nearest neighbor binary classification" mismatch removal mechanism, combined with CPU-GPU heterogeneous parallel matrix calculation, the problems of high mismatch rate and high computational complexity in existing technologies are solved, and high-precision and efficient image feature matching is achieved.
Patent Information
- Application Number
- CN202510825705.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-26
AI Technical Summary
Existing image matching methods have high mismatch rates and high computational complexity in image scenes with complex backgrounds, large-scale changes, and parallax interference, making it difficult to meet real-time requirements.
By analyzing the spatial distribution characteristics of motion vectors of image matching point pairs, a "coarse-fine dual-stage screening + K-nearest neighbor binary classification" mismatch removal mechanism is adopted, combined with a CPU-GPU heterogeneous parallel matrix computing framework to achieve efficient image feature matching.
The matching accuracy and robustness in complex scenarios have been significantly improved, with the matching accuracy increased to 89.33% and the computing efficiency increased by 24.28 times, meeting the needs of real-time applications.
Smart Images

Figure CN120707884A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision image processing, and in particular to a real-time image feature matching method based on spatial distribution consistency. Background Art
[0002] Image matching is a core foundational technology in computer vision. As one of the core foundations of computer vision, image matching technology plays a vital role in a variety of key application scenarios, particularly in systems such as 3D reconstruction, visual localization, object tracking, and augmented reality. Its primary task is to accurately extract local or global feature points from images acquired from different perspectives, scales, and even sensor types, and to establish precise correspondences between these feature points. This process not only affects the stability and accuracy of subsequent image processing modules but also directly impacts the performance and intelligence of the entire visual system. By establishing robust matching relationships, image matching provides reliable prior support for key processes such as 3D structure recovery, camera pose estimation, and object recognition and tracking, making it a prerequisite for achieving high-precision visual perception. With the continuous evolution of algorithmic theory and the continuous increase in computing resources, image matching technology is continuously developing towards higher matching accuracy, greater robustness, and faster processing speed.
[0003] However, due to the limited discriminative power of feature descriptors and the influence of geometric distortion in complex scenes, a high proportion of false matches is often present in the initial matching process, severely limiting the robustness and accuracy of subsequent applications. Traditional methods such as Random Sample Consensus (RANSAC) rely primarily on parameterized models to estimate geometric transformations to filter out anomalous matching pairs, but the performance of this approach is highly dependent on the accuracy of the model assumptions. In scenes with drastic viewpoint changes, non-rigid deformations, or difficult-to-parameterize geometric constraints, such methods are prone to ineffectively removing false matches due to model assumptions that deviate from reality. Furthermore, existing matching methods are generally computationally complex, making it difficult to meet the demands of real-time processing as camera image resolution continues to increase.
[0004] The Chinese invention patent "A Method for Image Feature Extraction and Matching" (CN201810291396.8) proposes an image matching method that combines salient region detection with a multi-scale sliding window strategy. This method first performs edge and corner detection on the input two-dimensional image, and identifies salient regions within the image by fusing corner and edge information. Subsequently, the image is segmented and scanned using a multi-scale sliding window approach, and histogram of oriented gradients (HOG) features are extracted within each window. A smaller sliding step size is used within salient regions to improve the resolution and accuracy of feature extraction, while a larger step size is used in non-salient regions to enhance efficiency. Next, the similarity between each sliding window of the query image and the database image is calculated. If the distance between the windows meets a set threshold, the windows are identified as similar, and a series of optimal matching results are further selected. Finally, by eliminating erroneous matches due to inconsistent scales or large spatial position deviations, matching pairs with high consistency are retained, achieving accurate extraction of similar regions. This method balances local details and overall structure of the image during feature extraction and matching, resulting in strong matching capabilities and particularly suitable for image matching tasks in complex backgrounds. However, its performance depends largely on the accuracy of salient region detection and the reasonable setting of sliding window parameters, and is easily affected by noise interference, which may lead to insufficient matching accuracy and stability in some application scenarios.
[0005] The journal article "Fast and Robust Feature Matching Optimization Based on Spatial Constraints" (Journal of Electronics and Information Technology, 2014, 36(11):2571–2577) proposed a spatially constrained SURF feature point matching optimization algorithm - SC-SURF. The method first uses the SURF algorithm to extract image feature points and perform preliminary matching. Then, the matching points are sorted based on the nearest neighbor ratio principle, and a new reference coordinate system is constructed with the optimal matching points. The matching constraints are strengthened by encoding the spatial position relationship between the matching points. On this basis, a small number of optimal matching points are selected to construct a representative data set to simplify the computational burden of the RANSAC algorithm in estimating the target projection transformation matrix. Finally, the spatial constraint matrix is combined with the simplified RANSAC to perform geometric verification on the initial matching results, effectively eliminating incorrect matching point pairs. Experimental results show that the SC-SURF algorithm performs better in both matching accuracy and speed than the traditional SURF, RANSAC-SURF and SIFT algorithms in dealing with typical interference conditions such as image scale change, rotation, blur, illumination change, and compression. In particular, while maintaining the original computation time of the SURF algorithm, the overall matching performance was significantly improved. However, the matching effect of this algorithm is highly dependent on the screening results of the optimal matching points. Therefore, in practical applications, the optimal point selection mechanism still needs to be further optimized to improve the stability and versatility of the algorithm.
[0006] The journal article "Fast False Match Removal Algorithm Based on Motion Smoothness Constraint" (Computer Applications, 2018, 38(9):2678–2682+2741) proposes an efficient false match removal algorithm that integrates motion smoothness constraints. The method first uses ORB feature extraction and Hamming distance for preliminary matching, then introduces motion smoothness constraints and roughly eliminates the initial matching results based on the neighborhood support estimator to eliminate most of the significant false matches. Then, the structural similarity index is used to fine-tune the remaining matching points to further improve the matching accuracy. After the false match points are eliminated, the algorithm uses a group sorting sampling method to estimate the homography matrix between images and achieves image fusion through a weighted average strategy. Experimental results show that compared with the traditional RANSAC algorithm and its various improved versions, this method has significant advantages in false match removal accuracy: compared with the optimization strategy of reducing the total sampling amount, the false match removal rate is improved by 75.6%; compared with the method based on adaptive threshold, it is improved by 24%. In addition, while maintaining high matching accuracy, this method effectively reduces the computational complexity and achieves fast and high-quality image stitching. However, the algorithm is somewhat dependent on the number of feature points. When the image overlapping area is small, there is still room for optimization in the false matching elimination effect and image fusion efficiency. Subsequent research can further improve the adaptability and robustness of the algorithm in sparse feature environments.
[0007] In summary, most existing image feature matching optimization methods are based on traditional feature extraction and matching processes. While these methods can achieve high matching accuracy in low-resolution and general application scenarios, they are limited by the discriminative power of feature descriptors and the interference of geometric distortion in complex environments. When dealing with sparsely distributed feature points, drastic geometric changes, or high-resolution, large-scale image matching tasks, they often struggle to achieve both real-time and robust performance. In particular, significant room for improvement remains in terms of matching speed and false match rejection accuracy. More efficient and robust matching strategies are urgently needed to meet the practical demands of complex visual perception tasks.
[0008] The existing technology mainly has the following two shortcomings:
[0009] (1) Existing image matching methods generally have a high mismatching rate when dealing with image scenes with complex backgrounds, large-scale changes, and significant parallax interference. Existing image matching methods generally have a high mismatching rate when dealing with image scenes with complex backgrounds, large-scale changes, and significant parallax interference. Traditional feature matching optimization methods, such as those based on random sampling consensus (RANSAC) and its improved algorithms, usually rely on preset geometric models (such as affine transformation, homography matrix, or basic matrix) to perform inlier screening and outlier removal on matching points. Their matching effect depends to a large extent on the accuracy of the geometric model assumptions. However, in practical applications, there are often non-rigid deformations, complex local perturbations, and various non-parametric geometric changes between images, making it difficult for traditional models to fully describe the real spatial relationship between matching points, resulting in difficulty in effectively removing mismatches. Especially in complex scenes such as drastic dynamic changes, significant image distortion, or cross-sensor matching, traditional methods show obvious deficiencies in matching accuracy and robustness, making it difficult to meet the practical application requirements of high accuracy and strong robustness.
[0010] (2) Existing image matching methods generally have the problems of high computational complexity and long time consumption, which makes it difficult to meet the actual needs of real-time online applications. The current mainstream image feature matching methods are mostly based on local feature extraction and descriptor matching, which usually include multiple processing steps such as feature point detection, descriptor generation, initial matching, false matching elimination and geometric consistency verification. The overall process is relatively complex and the computational overhead is high. In practical applications, factors such as the number of feature points, image resolution and scene complexity will significantly affect the execution efficiency of the matching process. Especially when processing high-resolution images, large-scale image sets or multi-image stitching tasks, the computing time increases dramatically, which seriously restricts the deployment and application of the algorithm in real-time scenarios.
[0011] Therefore, the present invention aims to propose a method that can achieve real-time and accurate matching of image feature points, to solve the problem that existing feature matching methods for high-resolution images generally have a high mismatch rate when processing image scenes with complex backgrounds, large-scale changes and parallax interference; at the same time, due to its high computational complexity, the matching process is time-consuming, and it is difficult to meet the application requirements with high real-time requirements, so as to improve the speed, accuracy and robustness of image matching and solve the problems of high mismatch rate and insufficient real-time performance in complex scenes. Summary of the Invention
[0012] To address the aforementioned deficiencies in the prior art, the present invention aims to provide a real-time image feature matching method based on spatial distribution consistency. By analyzing the distribution characteristics of the motion vectors of matching point pairs during the image matching process, it is found that the motion vectors of correctly matched point pairs exhibit significant spatial consistency, while the motion vectors of incorrectly matched point pairs are distributed disorderly. A real-time image matching method with spatial distribution consistency is implemented through a "coarse-fine dual-stage screening + K-nearest neighbor binary classification" mismatch removal mechanism. A CPU-GPU heterogeneous parallel matrix computation framework is constructed, which converts the Euclidean distance calculation between high-dimensional descriptors into matrix multiplication, significantly improving matching efficiency and meeting the practical needs of real-time online applications.
[0013] The present invention is achieved through the following technical solutions.
[0014] The present invention provides a real-time image feature matching method based on spatial distribution consistency, comprising:
[0015] Obtaining high-resolution images to be matched with overlapping areas;
[0016] Perform SIFT feature point detection on the image and extract descriptor vectors corresponding to the SIFT feature points;
[0017] Calculate the Euclidean distance between the descriptors corresponding to SIFT feature points;
[0018] Match the descriptors corresponding to the SIFT feature points to obtain the initial matching point pairs;
[0019] Based on the initial matching point pairs, construct two-dimensional sample points;
[0020] Cluster the constructed two-dimensional sample points to separate the correct matching point pairs from the incorrect matching point pairs, and achieve rough screening of matching point pairs;
[0021] Based on the remaining matching point pairs after coarse screening, a four-dimensional sample vector is constructed;
[0022] The constructed four-dimensional sample vectors are clustered to separate the correct matching point pairs from the incorrect matching point pairs, thus achieving precise screening of the matching point pairs.
[0023] Preferably, the Euclidean distance between the descriptors corresponding to the SIFT feature points is calculated, including converting the Euclidean distance calculation between the SIFT feature point descriptors into a matrix multiplication form, and using GPU parallel acceleration technology to calculate the Euclidean distance between the SIFT feature point descriptors in the two images extracted in the second step.
[0024] Preferably, calculating the Euclidean distance between the descriptors corresponding to the SIFT feature points includes:
[0025] Construct SIFT feature point descriptor matrix;
[0026] Calculate the Euclidean distance between SIFT feature point descriptors;
[0027] Square and accumulate each row of the feature point descriptor matrix to obtain a square sum matrix;
[0028] Construct the length N A With N B All-1 matrix;
[0029] Calculate Image I A The SIFT feature point descriptor vector and image I B The Euclidean distance between SIFT feature point descriptor vectors in .
[0030] Preferably, matching the descriptors corresponding to the SIFT feature points to obtain initial matching point pairs includes:
[0031] The nearest neighbor ratio test and the mutual nearest neighbor matching algorithm are used to match the descriptors corresponding to the SIFT feature points. and Then the two SIFT feature point descriptor vectors and They are each other's nearest neighbors and meet the ratio test conditions to form the initial matching point pair.
[0032] Preferably, constructing two-dimensional sample points based on the initial matching point pairs includes:
[0033] Construct a series of two-dimensional sample points based on the two-dimensional coordinates of the initial matching point pairs;
[0034] Normalize each two-dimensional sample point.
[0035] Preferably, clustering is performed on the constructed normalized two-dimensional sample points to separate correct matching point pairs from incorrect matching point pairs, thereby achieving a rough screening of matching point pairs, including:
[0036] Construct a sample point distance matrix, calculate the Euclidean distance between sample points, and obtain the sample point distance matrix;
[0037] Construct a weight matrix to perform weighted processing on the sample point distance matrix;
[0038] Set the matching threshold, sort each row of the sample point distance matrix in ascending order to obtain the ordered sample point distance matrix, and calculate the neighborhood radius based on the ordered sample point distance matrix and the threshold;
[0039] For the rough screening of matching point pairs, the constructed normalized two-dimensional sample points are divided into inliers and outliers using the K-nearest neighbor binary classification algorithm; if the normalized two-dimensional sample point is classified as an inlier, its corresponding initial matching point pair is marked as a correct match; otherwise, if the normalized two-dimensional sample point is classified as an outlier, its corresponding initial matching point pair is marked as a mismatch and deleted from the initial matching point pair set.
[0040] As an optimal method, constructing a sample point distance matrix includes:
[0041] Arrange the constructed normalized two-dimensional sample points in rows to construct an N×2-dimensional sample point matrix;
[0042] Calculate the sum of squares of each sample point and construct an N×1 dimensional sample point square sum matrix;
[0043] Calculate the Euclidean distance between sample points and construct the sample point distance matrix.
[0044] Preferably, a weight matrix is constructed to perform weighted processing on the sample point distance matrix, including:
[0045] Normalize the coordinates of the initial matching points:
[0046] Use the normalized matching point pairs to construct two N×2 dimensional matrices respectively;
[0047] Construct two N×1 dimensional square sum matrices;
[0048] Construct the Euclidean distance matrix between feature points;
[0049] Construct the weight matrix.
[0050] As a preference, constructing a four-dimensional sample vector includes:
[0051] Based on the remaining matching point pairs after coarse screening, a four-dimensional sample vector containing the spatial position and direction information of the matching point pairs is constructed, and the four-dimensional sample vector is normalized.
[0052] Preferably, the fine screening of matching point pairs includes:
[0053] Arrange the constructed normalized four-dimensional sample vectors in rows to construct an M×4-dimensional sample vector matrix;
[0054] Calculate the Euclidean distance between four-dimensional sample vectors and construct a four-dimensional sample vector distance matrix;
[0055] Set the matching threshold;
[0056] Sort each row of the four-dimensional sample vector distance matrix in ascending order to obtain an ordered four-dimensional sample vector distance matrix, and calculate the neighborhood radius based on the ordered distance matrix and the threshold;
[0057] The constructed normalized four-dimensional sample vector is divided into inliers and outliers using the K-nearest neighbor binary classification algorithm;
[0058] If the normalized four-dimensional sample vector is classified as an internal point, its corresponding matching point pair is marked as a correct match; otherwise, if the normalized four-dimensional sample vector is classified as an external point, its corresponding matching point pair is marked as a wrong match and deleted from the matching point pair set.
[0059] The present invention adopts the above technical solution, which has the following beneficial effects:
[0060] The new method's two-stage screening strategy effectively eliminates outliers, increasing the final matching accuracy to 89.33%. Using CUDA parallel optimization on the GPU side, for a 5-megapixel resolution image, matching 4,000 feature descriptors takes only about 4ms, which is 24.28 times faster than the pure CPU version, meeting real-time (≈250fps) application requirements.
[0061] The method of the present invention has the following advantages:
[0062] 1. Significantly improve the matching accuracy and robustness in complex scenarios. Based on the distribution characteristics of feature point matching pairs in the motion vector space, the present invention designs a two-stage screening mechanism including coarse screening and fine screening, which effectively eliminates a large number of mismatched points. Compared with traditional matching optimization methods that rely on fixed geometric models (such as homography matrix and basic matrix), the present invention can adapt to complex application scenarios such as large perspective changes, non-rigid deformation and non-parametric geometric constraints, thereby significantly improving matching accuracy and stability. Experimental tests show that the matching accuracy can reach 89.33%.
[0063] 2. Greatly improve the real-time performance and computational efficiency of image feature matching. The present invention converts the Euclidean distance calculation of high-dimensional feature descriptors into matrix multiplication operations that can be executed in parallel on the GPU, and introduces a reasonable matrix partitioning strategy to optimize the memory access mode, giving full play to the parallel computing capabilities of the GPU. Experimental results show that in the matching task of 4000 feature point descriptors, the average matching time is reduced to about 4 milliseconds, which is about 24.28 times faster than the traditional CPU-based serial implementation solution. The overall system performance can achieve 250FPS, which is significantly better than the existing technology level. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of this application, and do not constitute an improper limitation of the present invention. In the drawings:
[0065] Figure 1 Flowchart of the specific operation steps of the present invention;
[0066] Figure 2(a) shows one image in the pair of images to be matched;
[0067] Figure 2(b) shows another image in the pair of images to be matched;
[0068] Figure 3 is the initial matching point pair detected in the input image pair to be matched;
[0069] Figure 4 is a two-dimensional sample point constructed based on the initial matching point pair;
[0070] Figure 5 is the matching point pair after coarse screening;
[0071] Figure 6 is the four-dimensional sample vector constructed based on the matching point pairs after coarse screening;
[0072] Figure 7 are the matching point pairs after fine screening. DETAILED DESCRIPTION
[0073] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The exemplary embodiments and descriptions of the present invention are used to explain the present invention but are not intended to limit the present invention.
[0074] The present invention proposes a real-time image matching method based on spatial distribution consistency classification, the characteristics and implementation process of which are as follows: Figure 1 As shown, the following steps are included:
[0075] Step 1: Input the image pair to be matched.
[0076] Two high-resolution images with a certain overlap are captured locally or controlled by an industrial camera for matching. If the input images are color images, they must be converted to grayscale. Figures 2(a) and 2(b) show two input image pairs to be matched.
[0077] Step 2: SIFT feature point detection.
[0078] Perform SFIT feature point detection on the images to be matched input in the first step. Assume that the two images to be matched are denoted as I A and I BFor each image, the GPU-accelerated SIFT feature detection method published by Acharya et al. (Acharya KA, Venkatesh Babu R, Vadhiyar SS. A real-time implementation of SIFT using GPU. Journal of Real-Time Image Processing. 2018, 14: 267-77) is applied to quickly extract the 128-dimensional descriptor vector corresponding to the SIFT feature points.
[0079] Step 3: The Euclidean distances between SIFT feature point descriptors are calculated in parallel.
[0080] The Euclidean distance calculation between SIFT feature point descriptors is converted into matrix multiplication form, and GPU parallel acceleration technology is used to calculate the Euclidean distance between the SIFT feature point descriptors in the two images extracted in the second step. The parallel calculation process of the Euclidean distance between SIFT feature point descriptors is as follows:
[0081] 1) SIFT feature point descriptor matrix construction. Assume d i,j Represents image I A The descriptor component of the i-th feature point in the k-th dimension, 1≤i≤N A ,1≤k≤128,N A Represents the image I extracted in the second step A The number of SIFT feature points in , then the descriptor matrix of all SIFT feature points is constructed as follows:
[0082]
[0083] Among them, D A Represents image I A The feature point descriptor matrix of image I can be constructed in the same way. B The feature point descriptor matrix D B Assume d j,k Represents image I B The descriptor component of the j-th feature point in the k-th dimension, 1≤j≤N B ,1≤k≤128,N B Represents the image I extracted in the second step B The number of SIFT feature points in , then the descriptor matrix of all feature points is constructed as follows:
[0084]
[0085] 2) Calculation of Euclidean distance between SIFT feature point descriptors. First, the matrix D AWith matrix D B Each row of is squared and accumulated to obtain the square sum matrix and The calculation formula is as follows:
[0086]
[0087] Where, and They are image I A , I B The matrix formed by the accumulation of 128-dimensional vectors of feature point descriptors extracted from .
[0088] Secondly, construct A With N B All-1 matrix and
[0089]
[0090] Where, is a row of length N A , a matrix with a column length of 1; It is a row length of 1 and column length of N B The matrix of .
[0091] Finally, calculate the image I A The SIFT feature point descriptor vector and image I B The Euclidean distance between SIFT feature point descriptor vectors in :
[0092]
[0093] After converting the Euclidean distance calculation between SIFT feature point descriptors into matrix multiplication, the matrix multiplication GPU parallel acceleration algorithm proposed by Bell et al. (Bell, N. and Garland, M., 2008. Efficient sparsematrix-vector multiplication on CUDA (Vol. 2, No. 5). Nvidia Technical Report NVR-2008-004, Nvidia Corporation.) is used to quickly calculate the Euclidean distance matrix D d .
[0094] Step 4: Get the initial matching point pair.
[0095] Combining the nearest neighbor ratio test with the mutual nearest neighbor matching algorithm, the SIFT feature point descriptors in the two images are matched to obtain the initial matching point pairs. This method measures the similarity by calculating the distance between the SIFT feature point descriptor vectors to determine the optimal matching relationship. The calculation formula is as follows:
[0096]
[0097] Among them, RM() represents the nearest neighbor ratio test function; and Represents image I A The i-th and n-th 128-dimensional SIFT feature point descriptor vectors, 1≤i≤N A , 1≤n≤N A ; and Represents image I B The jth and mth 128-dimensional SIFT feature point descriptor vectors, 1≤j≤N h ,1≤m≤N B ; and yes The nearest neighbor and next nearest neighbor of and yes The nearest neighbor and the next nearest neighbor; η is the matching threshold, and its value range is [0,1]; the symbol || ||2 represents the Euclidean distance between vectors, which can be calculated by the distance matrix D in the third step. d get.
[0098] when and Then the two SIFT feature point descriptor vectors and They are each other's nearest neighbors and meet the ratio test conditions to form the initial matching point pair.
[0099] Figure 3 The two input images to be matched as shown in FIG. 2( a ) and FIG. 2( b ) are processed by the method of the present invention to obtain the initial matching point pairs.
[0100] Step 5: Construct two-dimensional sample points.
[0101] The initial matching point pairs obtained in step 4 contain many mismatches. Therefore, the present invention constructs a series of two-dimensional sample points based on the two-dimensional coordinates of the initial matching point pairs for rough screening. Assume that a pair of initial matching point pairs is in image I A With image I B The two-dimensional coordinates in are (x i,A ,y i,A ) and (x i,B,y i,B ), 1≤i≤N, N represents the number of initial matching point pairs, which is also the number of constructed two-dimensional sample points. The two-dimensional sample points are constructed as follows:
[0102] v i =(v i,x ,v i,y )=(x i,A -x i,B ,y i,A -y i,B )
[0103] Among them, v i Represents the constructed two-dimensional sample points, v i,x With v i,y is its weight. i,A ,y i,A ) is the starting point, (x i,B ,y i,B ) is the motion vector corresponding to the matching point pair, so v i,x With v i,y In fact, it is the coordinate components of the motion vector amplitude in the x-axis and y-axis directions. In order to eliminate the dimensionality effect of the motion vector amplitude, each sample point is normalized. The normalization formula is:
[0104]
[0105] Among them, min() and max() represent the minimum and maximum value functions respectively. and Respectively represent the coordinates of the i-th sample point after normalization. All sample points after normalization The coordinates are all distributed in the interval [0,1], which is conducive to unified measurement and improved screening accuracy. Figure 4 Shown is based on Figure 3 The two-dimensional sample points are constructed from the initial matching point pairs obtained in .
[0106] Step 6: Coarse screening of matching point pairs.
[0107] The K-nearest neighbor binary classification algorithm is used to cluster the normalized two-dimensional sample points constructed in the fifth step, separating the correct matching point pairs from the incorrect matching point pairs, and achieving a rough screening of the matching point pairs. The specific process is as follows:
[0108] 1) Construct a sample point matrix. Arrange the normalized two-dimensional sample points constructed in step 6 in rows to construct an N×2-dimensional sample point matrix:
[0109]
[0110] To facilitate subsequent matrix operations, calculate the sum of squares of each sample point and construct an N×1 dimensional sample point square sum matrix:
[0111]
[0112] Then, calculate the Euclidean distance between sample points and construct the sample point distance matrix. The formula is as follows:
[0113]
[0114] Among them, 1 N×1 with 1 1×N is a matrix of all 1s:
[0115]
[0116] Sample point distance matrix D V The Euclidean distance between two-dimensional sample points is stored, which provides the basis for subsequent neighborhood calculations. The application of matrix multiplication allows the distance calculation between all two-dimensional sample points to be completed in the same matrix operation, greatly improving the computational efficiency.
[0117] 2) Constructing a weight matrix. To further enhance the consistency of the spatial distribution of sample points, the present invention introduces a weight matrix W to construct a weight matrix D for the sample point distance matrix D. V Perform weighted processing to adjust the influence between different sample points and improve the accuracy of screening. First, the coordinates of the initial matching point pair (x i,A ,y i,A ) and (x i,B ,y i,B ) for normalization:
[0118]
[0119] Then, two N×2 dimensional matrices S are constructed using the normalized matching point pairs. A and S B :
[0120]
[0121] Next, construct two N×1 dimensional square sum matrices and
[0122]
[0123] Then, calculate the Euclidean distance between the feature points and construct the Euclidean distance matrix D between the feature points A With D B , the calculation formula is as follows:
[0124]
[0125] Finally, calculate the weight matrix W:
[0126]
[0127] Among them, γ is a positive real number in the interval [0,50], which is used to adjust the size of the weight; min{D A ,D B} means comparing two distance matrices D element by element A With D B The corresponding elements of , take out the minimum value of each pair of elements to form a minimum value matrix; It means that the exponential operation is performed element by element on the minimum value matrix with e as the base.
[0128] The weight matrix W is the distance matrix D of the sample points constructed in step 1) by the Hadamard product operation. V Weighted, that is:
[0129]
[0130] The symbol ⊙ represents the Hadamard product (element-wise multiplication). represents a weighted distance matrix.
[0131] 3) Matching threshold setting. Set the neighborhood radius ε and the minimum neighborhood minPts value to ensure the robustness and adaptability of the binary classification. The minimum neighborhood is calculated as follows:
[0132] minPts=min(max(N×ρ,B min ),B max )
[0133] Among them, N is the number of two-dimensional sample points constructed in the fifth step, B min With B max are the upper and lower limits of the minimum neighborhood minPts, ρ is a set constant used to control the number of lowest internal points; min() and max() represent the minimum and maximum value functions.
[0134] The setting of the neighborhood radius ε needs to take into account the distance matrix And the value of the minimum number of neighborhood points minPts. First, the distance matrix Each row of is sorted in ascending order to obtain a distance matrix in which the elements are arranged in ascending order. Next, based on the ordered distance matrix And the threshold minPts, calculate the neighborhood radius ε, the specific formula is as follows:
[0135]
[0136] Where, 1≤j≤N, Represents an ordered distance matrix In the minPts-th row and j-th column, the parameter μ is a custom constant with a value range of [0, 1]. By introducing the constant μ, the neighborhood radius ε can be flexibly adjusted in different image scenarios to adapt to different data distribution characteristics.
[0137] 4) Rough screening of matching point pairs. Use the K-nearest neighbor binary classification algorithm to divide the normalized two-dimensional sample points constructed in the fifth step into inliers and outliers. The specific formula is as follows:
[0138]
[0139] in, Indicates that the two-dimensional sample points will be normalized Marked as interior point, Indicates that the two-dimensional sample points will be normalized Marked as an outlier, 1≤i≤N, Represents the normalized two-dimensional sample points Normalized two-dimensional sample points The ε-neighborhood of is defined as:
[0140]
[0141] in, Represents the weighted distance matrix The element in row i and column j is also a normalized two-dimensional sample point and The Euclidean distance between Also represents the normalized two-dimensional sample point and The Euclidean distance between .
[0142] If the normalized two-dimensional sample points Divided into inliers, the corresponding initial matching point pair (x i,A ,y i,A ) and (x i,B ,y i,B ) is marked as a correct match; otherwise, if the normalized two-dimensional sample point is divided into outliers, then the corresponding initial matching point pair (x i,A ,y i,A ) and (x i,B ,y i,B ) are marked as mismatches and deleted from the initial set of matching point pairs. Figure 5 Shown Figure 3 The matching point pairs remaining after the initial matching point pairs in
[0143] Step 7, construct a four-dimensional sample vector.
[0144] Among the matching point pairs remaining after the rough screening in Step 6, there are still a small number of mis-matches. Therefore, the present invention constructs a series of four-dimensional sample vectors using the matching point pairs remaining after the rough screening for fine screening. Assume that the number of matching point pairs remaining after the rough screening is M, where M < N. Combine the matching point pair (x i,A , y i,A ) and (x i,B , y i,B ) into a four-dimensional sample vector:
[0145]
[0146] In the formula, represents the constructed four-dimensional sample vector, 1 ≤ i ≤ M, h i,1 , h i,2 , h i,3 , h i,4 represent the four components of the four-dimensional sample vector , h i,1 = x i,A , h i,2 = y i,A , h i,3 = x i,B , h i,4 = y i,B . The constructed four-dimensional sample vector not only contains the spatial positions of the matching point pairs but also contains direction information, providing stronger support for further screening.
[0147] To eliminate the influence of the dimension of the amplitude, normalize the four-dimensional sample vector :
[0148]
[0149] In the formula and <000052o>respectively represent the k-th component of the four-dimensional sample vector before and after normalization; min() and max() represent the functions of taking the minimum value and the maximum value respectively. Figure 6 As shown in Figure 5 are the four-dimensional sample vectors constructed from the matching point pairs after rough screening in
[0150] Step 8, fine screening of matching point pairs.
[0151] The K-nearest neighbor binary classification algorithm is used to cluster the normalized four-dimensional sample vectors constructed in the seventh step, separating the correct matching point pairs from the incorrect matching point pairs, and realizing the precise screening of matching point pairs. The specific process is as follows:
[0152] 1) Arrange the normalized four-dimensional sample vectors constructed in step 7 in rows to construct an M×4 dimensional sample vector matrix:
[0153]
[0154] Next, construct an N×1 dimensional square sum matrix
[0155]
[0156] Then, calculate the Euclidean distance between the four-dimensional sample vectors and construct the four-dimensional sample vector matrix D H , the calculation formula is as follows:
[0157]
[0158] Where, 1 1×M with 1 M×1 is a matrix of all 1s:
[0159]
[0160] 2) Matching threshold setting. The setting method of neighborhood radius ε and minimum neighborhood minPts is the same as in step 6. The minimum neighborhood is calculated as follows:
[0161] minPts=min(max(M×ρ,B min ),B max )
[0162] Among them, M is the number of matching point pairs remaining after coarse screening, and parameter B min , B max , ρ, and the functions min() and max() are the same as those described in step 6.
[0163] The setting of the neighborhood radius ε is also consistent with the sixth step, and the distance matrix D needs to be considered comprehensively. H And the value of the minimum number of neighborhood points minPts. First, the distance matrix D H Each row of is sorted in ascending order to obtain a distance matrix in which the elements are arranged in ascending order. Next, based on the ordered distance matrix And the threshold minPts, calculate the neighborhood radius ε, the specific formula is as follows:
[0164]
[0165] Where, 1≤j≤M, Represents an ordered distance matrix The element in the minPts-th row and j-th column in , the parameter μ is the same as described in step 6.
[0166] 3) Fine screening of matching point pairs. Use the K-nearest neighbor binary classification algorithm to divide the normalized four-dimensional sample vector constructed in the seventh step into inliers and outliers. The specific formula is as follows:
[0167]
[0168] in, Indicates that the four-dimensional sample vector will be normalized Marked as interior point, Indicates that the four-dimensional sample points will be normalized Marked as an outlier, 1≤i≤M, Represents the normalized four-dimensional sample vector Normalized four-dimensional sample points The ε-neighborhood of is defined as:
[0169]
[0170] Among them, D H (i,m) represents the weighted distance matrix D H The element in the i-th row and the m-th column is also a normalized four-dimensional sample vector and The Euclidean distance between Also represents the normalized four-dimensional sample vector and The Euclidean distance between .
[0171] If the normalized four-dimensional sample vector Divided into inliers, the corresponding matching point pairs (x i,A ,y i,A ) and (x i,B ,y i,B ) is marked as a correct match; otherwise, if the normalized four-dimensional sample vector is divided into outliers, then the corresponding matching point pair (x i,A ,y i,A ) and (x i,B ,y i,B ) is marked as a mismatch and removed from the set of matched point pairs.
[0172] Step 9: Output matching point pairs.
[0173] Output the matching point pairs marked as correctly matched after fine screening in step 8, and the matching process ends. Figure 7 The figure shows the matching point pairs after fine screening, which are also the final output matching point pairs. Table 1 shows the matching efficiency when the resolution of the image to be matched is 2592×2048 and the number of SIFT feature points is different.
[0174] Table 1 Matching efficiency of the image pairs to be matched (resolution 2592×2048) when they contain different numbers of SIFT feature points
[0175]
[0176] Table 1 shows that in the matching task of 4,000 feature points, the average matching time was reduced to approximately 4 milliseconds, achieving a 24.28-fold acceleration compared to traditional CPU-based solutions. The overall matching rate can reach 250 FPS, a significant improvement.
[0177] The present invention is not limited to the above-mentioned embodiments. On the basis of the technical solutions disclosed in the present invention, those skilled in the art can make some substitutions and modifications to some of the technical features therein according to the disclosed technical content without creative labor, and these substitutions and modifications are all within the protection scope of the present invention.
Claims
1. A real-time image feature matching method based on spatial distribution consistency, characterized in that: include: Obtaining high-resolution images to be matched with overlapping areas; Perform SIFT feature point detection on the image and extract descriptor vectors corresponding to the SIFT feature points; Calculate the Euclidean distance between the descriptors corresponding to SIFT feature points; Match the descriptors corresponding to the SIFT feature points to obtain the initial matching point pairs; Based on the initial matching point pairs, construct two-dimensional sample points; Cluster the constructed two-dimensional sample points to separate the correct matching point pairs from the incorrect matching point pairs, and achieve rough screening of matching point pairs; Based on the remaining matching point pairs after coarse screening, a four-dimensional sample vector is constructed; The constructed four-dimensional sample vectors are clustered to separate the correct matching point pairs from the incorrect matching point pairs, thus achieving precise screening of the matching point pairs.
2. The real-time image feature matching method based on spatial distribution consistency according to claim 1, characterized in that: Calculate the Euclidean distance between the descriptors corresponding to the SIFT feature points, including converting the Euclidean distance calculation between the SIFT feature point descriptors into matrix multiplication form, and using GPU parallel acceleration technology to calculate the Euclidean distance between the SIFT feature point descriptors in the two images extracted in the second step.
3. The real-time image feature matching method based on spatial distribution consistency according to claim 2, characterized in that: Calculate the Euclidean distance between the descriptors corresponding to SIFT feature points, including: Construct SIFT feature point descriptor matrix: Where D A 、D B Represents image I A , I B The feature point descriptor matrix, d i,j Represents image I A and image I B The j-th dimension eigenvalue of the i-th feature point detected; Calculate the Euclidean distance between SIFT feature point descriptors: For the matrix D A With matrix D B Each row of is squared and accumulated to obtain the square sum matrix and Where, and They are image I A , I B The matrix formed by the accumulation of 128-dimensional vectors of feature point descriptors extracted from ; Construct the length N A With N B All-1 matrix and Where, is a row of length N A , a matrix with a column length of 1; It is a row length of 1 and column length of N B Matrix of Calculate Image I A The SIFT feature point descriptor vector and image I B The Euclidean distance between SIFT feature point descriptor vectors in :
4. The real-time image feature matching method based on spatial distribution consistency according to claim 1, characterized in that: Match the descriptors corresponding to the SIFT feature points to obtain the initial matching point pairs, including: The nearest neighbor ratio test and the mutual nearest neighbor matching algorithm are used to match the descriptors corresponding to the SIFT feature points. and Then the two SIFT feature point descriptor vectors and They are each other's nearest neighbors and meet the ratio test conditions to form the initial matching point pair.
5. The real-time image feature matching method based on spatial distribution consistency according to claim 1, characterized in that: Based on the initial matching point pairs, construct two-dimensional sample points, including: Construct a series of two-dimensional sample points based on the two-dimensional coordinates of the initial matching point pairs: v i =(v i,x ,v i,y )=(x i,A -x i,B ,y i,A -y i,B ) Among them, v i Represents the constructed two-dimensional sample points, v i,x With v i,y is its weight, (x i,A ,y i,A )、(x i,B ,y i,B ) are the starting point and end point of the motion vector corresponding to the matching point pair respectively; Normalize each two-dimensional sample point; Among them, min() and max() represent the minimum and maximum value functions respectively. and They represent the normalized coordinates of the i-th two-dimensional sample point.
6. The real-time image feature matching method based on spatial distribution consistency according to claim 1, characterized in that: Cluster the constructed two-dimensional sample points to separate the correct matching point pairs from the incorrect matching point pairs, and implement rough screening of matching point pairs, including: Construct a sample point distance matrix, calculate the Euclidean distance between sample points, and obtain the sample point distance matrix; Construct a weight matrix to perform weighted processing on the sample point distance matrix; Set the matching threshold, sort each row of the sample point distance matrix in ascending order to obtain the ordered sample point distance matrix, and calculate the neighborhood radius based on the ordered sample point distance matrix and the threshold; For the rough screening of matching point pairs, the constructed normalized two-dimensional sample points are divided into inliers and outliers using the K-nearest neighbor binary classification algorithm; if the normalized two-dimensional sample point is classified as an inlier, its corresponding initial matching point pair is marked as a correct match; otherwise, if the normalized two-dimensional sample point is classified as an outlier, its corresponding initial matching point pair is marked as a mismatch and deleted from the initial matching point pair set.
7. The real-time image feature matching method based on spatial distribution consistency according to claim 6, characterized in that: Construct a sample point distance matrix, including: Arrange the constructed normalized two-dimensional sample points in rows to construct an N×2-dimensional sample point matrix; Calculate the sum of squares of each sample point and construct an N×1 dimensional sample point square sum matrix; Calculate the Euclidean distance between sample points and construct the sample point distance matrix: Among them, 1 N×1 with 1 1×N is a matrix of all 1s.
8. The real-time image feature matching method based on spatial distribution consistency according to claim 6, characterized in that: Construct a weight matrix to perform weighted processing on the sample point distance matrix, including: Normalize the coordinates of the initial matching points: Use the normalized matching point pairs to construct two N×2 dimensional matrices respectively; Construct two N×1 dimensional square sum matrices; Construct the Euclidean distance matrix between feature points; Construct the weight matrix.
9. The real-time image feature matching method based on spatial distribution consistency according to claim 1, characterized in that: Construct a four-dimensional sample vector, including: Based on the remaining matching point pairs after coarse screening, a four-dimensional sample vector containing the spatial position and direction information of the matching point pairs is constructed, and the four-dimensional sample vector is normalized.
10. The real-time image feature matching method based on spatial distribution consistency according to claim 1, characterized in that: Fine screening of matching point pairs, including: Arrange the constructed normalized four-dimensional sample vectors in rows to construct an M×4-dimensional sample vector matrix; Calculate the Euclidean distance between four-dimensional sample vectors and construct a four-dimensional sample vector distance matrix D H : Where, 1 1×M with 1 M×1 is a matrix of all 1s, H m is a four-dimensional sample vector matrix, which records the normalized four-dimensional sample vector composed of matching point pairs, and is defined as: Where, Represents the normalized four-dimensional sample vector; Set the matching threshold; Sort each row of the four-dimensional sample vector distance matrix in ascending order to obtain an ordered four-dimensional sample vector distance matrix, and calculate the neighborhood radius based on the ordered distance matrix and the threshold; The constructed normalized four-dimensional sample vector is divided into inliers and outliers using the K-nearest neighbor binary classification algorithm; If the normalized four-dimensional sample vector is classified as an internal point, its corresponding matching point pair is marked as a correct match; otherwise, if the normalized four-dimensional sample vector is classified as an external point, its corresponding matching point pair is marked as a wrong match and deleted from the matching point pair set.
Citation Information
Patent Citations
Image feature extraction and matching method
CN108830279A
Cited By
An unmanned aerial vehicle image matching method based on spatio-temporal change and neighborhood category outlier
CN122453886A