Image feature matching method based on overlapping region
By improving the estimation method of overlapping regions and introducing sparse matching, the combination of dense matching and overlapping regions is optimized, and the rationality and insufficient accuracy in the prior art are solved, and more efficient image feature matching is achieved.
Patent Information
- Application Number
- CN202510636444.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-17
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, the problem of the combination of dense matching and overlapping region extraction is the problem of insufficient accuracy of OETR when overlapping regions are small or boundary blurred.
The image feature matching method based on overlapping regions is adopted, and a more refined region extraction mechanism is introduced by turning to sparse matching. The estimation of overlapping regions is optimized by using Gaussian process regression and cosine coordinate embedding, and the latest SoTA sparse feature matching method LightGlue and its supporting feature extractor Superpoint are combined for local feature matching.
The accuracy and matching speed of overlapping areas are improved, invalid feature points interference is reduced, and efficient image feature matching is achieved in complex scenarios.
Smart Images

Figure CN120472194A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to an image feature matching method based on overlapping areas. Background Art
[0002] The existing feature matching results based on overlapping areas are mainly OAMatcher and OETR.
[0003] As a dense matching method, OAMatcher has the following drawbacks: OAMatcher combines overlap extraction and dense matching, but dense matching itself proceeds from coarse to fine, with the coarse matching portion equivalent to extracting overlapped regions. Therefore, combining the two raises architectural design issues. Compared to OAMatcher, OETR's overlap estimation is relatively coarse, especially when the overlap is small or has blurred boundaries. OETR estimates overlap through a two-step process of feature correlation and overlap regression. While this method effectively captures most of the overlap in an image pair, it may not be as accurate as OAMatcher when processing more complex image pairs.
[0004] Therefore, the problems that the existing technology needs to solve include: the rationality of the combination of dense matching and overlapping area extraction: due to the overlap in the functions of dense matching and overlapping areas, the rationality of the current architecture design needs to be investigated, and the existing OETR method is not accurate enough when extracting overlapping areas, especially when the overlapping areas are small or the boundaries are blurred. Summary of the Invention
[0005] The purpose of the present invention is to provide an image feature matching method based on overlapping areas. By turning to sparse matching, the speed and key point representativeness are improved while making the structure more reasonable, and the estimation method of overlapping areas is improved, a more refined region extraction mechanism is introduced, and the accuracy of overlapping areas is improved.
[0006] The present invention discloses an image feature matching method based on overlapping areas, and the technical solution adopted is:
[0007] An image feature matching method based on overlapping areas comprises the following steps:
[0008] First, the overlapping area of the two images is extracted by the algorithm to obtain the matching score of the two images, and based on this, the overlapping area in the image is extracted, and then the overlapping area is cropped and resized;
[0009] Input the processed cropped image into the feature matching network for subsequent matching;
[0010] The cropped overlapping area is subjected to local feature matching through the existing feature extractor and feature matcher.
[0011] As a preferred solution, the mapping relationship between finding overlapping areas and finding points between images is the same, that is, a point in image A can be mapped to image B, indicating that the point is in the overlapping area of the two images. The overlapping area extraction method can be defined as estimating image I A ,I B The mapping relationship between pixels W A→B And its corresponding matching score p θ , where the overlapping area can be extracted only by the matching score without using the mapping relationship, so the solution is the matching score p between the pixel pairs. θ ;
[0012] The overlapping area extraction algorithm is mainly divided into three parts: First, the feature map F of the two images to be matched is obtained through the encoder. A and F B , and then the position embedding of image B is obtained for the subsequent Gaussian process module; the Gaussian module produces a predicted posterior; then, the embedding decoder finds the matching score of the image pair;
[0013] The specific model solution can be expressed by the following formula:
[0014]
[0015] Among them, G θ Represents a global matcher, consisting of an encoder E θ and a decoder D θ The encoder uses Gaussian process regression and the decoder uses Transformer.
[0016] As a preferred solution, the output of the GP posterior is unimodal, so multimodal matching will reduce its performance. In image matching, a point in the image may correspond to multiple candidate points, and the cosine coordinate embedding method is adopted:
[0017] Emb F (x; W, b) = cos(8πWx)
[0018] in, represents the image coordinates,
[0019] As a preferred solution, the specific approach of Gaussian process regression is: in the standard assumption, the observation value It is obtained under independent and identically distributed noise, and a probability model is used to describe the distribution of observations. Based on this, for the feature map The analytical formula for the posterior condition is as follows:
[0020]
[0021] in, represents the coordinates after cosine embedding, K represents the kernel matrix, μ is the posterior mean function, σ n =0.1 is the standard deviation of the measurement noise, Σ is the posterior covariance, which represents the uncertainty. The formula in the lower part describes the uncertainty when φ is known. B After that, it is relative to φ A How big is the uncertainty if φ A and φ B Very similar, i.e. is large, then the uncertainty will decrease, but if φ A Distance φ B Very far, that is is approximately 0, then the variance of the prediction will be close to The uncertainty at this time will be very great.
[0022] As a preferred solution, the decoder uses classification regression to discretize the output space. The specific formula is as follows:
[0023] Where K is the number of discrete anchor points, π k express The probability of predicting the k-th anchor point, represents a 2D uniform basis distribution, m k Represents the position coordinates of the kth anchor point. The output shape of this block is B×H×W×(K+1), where K is the number of anchor points and the extra dimension is the matching score p. The last dimension is required to extract the overlapping area.
[0024] As a preferred solution, given a set of images, denoted as (I a ,I b ), where each image is accompanied by its associated keypoint set (P a ,P b ), their respective score vectors (S a ,S b ) and feature descriptors (F a ,F b ), it is worth noting that each pixel coordinate is associated with the corresponding position confidence, so the image I a and I b Two different sets of feature descriptors are generated, respectively denoted as and The matching of image features is defined as identifying the optimal one-to-one correspondence matrix As shown in the formula:
[0025]
[0026] The goal is to minimize the relative pose error of the image pair by utilizing all matching pairs. Each feature descriptor d is allowed to match at most once for pose estimation. Key points or descriptors that still do not match are considered outliers, and an alignment matrix is derived. It represents the pairwise correspondence between the features of two images.
[0027] As a preferred solution, the specific process of cutting out the overlapping area is as follows:
[0028] Normalize the matching score p within each sample so that the sum of each sample is 1:
[0029]
[0030] Among them, p norm Represents the normalized matching score matrix, and the cumulative sum of the confidence is calculated in the width W direction and the height H direction respectively. The formula is as follows:
[0031]
[0032] The clipping boundaries are then determined based on this:
[0033]
[0034] Among them, arg max returns the first index that meets the conditions.
[0035] After obtaining the area to be cropped, it is necessary to resize the area to reduce the impact of scale differences on the matching effect. The scale adjustment ratio is calculated as follows:
[0036]
[0037] At this point, the preprocessing of the overlapping area of the image pair to be matched is completed, and the processed cropped image is input into the feature matching network for subsequent matching.
[0038] As a preferred solution, the processed overlapping area images are subjected to local feature matching using the latest open source SoTA sparse feature matching method LightGlue and its supporting feature extractor Superpoint.
[0039] The beneficial effect of the image feature matching method based on overlapping areas disclosed by the present invention is to improve the rationality of the combination of dense matching and overlapping area extraction. Since the functions of dense matching and overlapping areas overlap, it is possible to consider switching to sparse matching, which improves the speed and representativeness of key points while making the structure more reasonable. By improving the estimation method of overlapping areas and introducing a more refined area extraction mechanism, the accuracy of overlapping areas is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 The present invention is a flowchart of an image feature matching method based on overlapping areas. DETAILED DESCRIPTION
[0041] The present invention will be further described and explained below in conjunction with specific embodiments and accompanying drawings:
[0042] Please refer to Figure 1 , an image feature matching method based on overlapping areas, comprising the following steps:
[0043] First, the overlapping area of the two images is extracted by the algorithm to obtain the matching score of the two images, and based on this, the overlapping area in the image is extracted, and then the overlapping area is cropped and resized;
[0044] Input the processed cropped image into the feature matching network for subsequent matching;
[0045] The cropped overlapping area is subjected to local feature matching through the existing feature extractor and feature matcher.
[0046] The mapping relationship between finding overlapping areas and finding points between images is the same, that is, a point in image A can be mapped to image B, indicating that the point is in the overlapping area of the two images. The overlapping area extraction method can be defined as estimating image I A ,I B The mapping relationship between pixels W A→B And its corresponding matching score p θ , where the overlapping area can be extracted only by the matching score, so the solution is the matching score p between the pixel pairs θ ;
[0047] That is, the extraction algorithm of overlapping areas is mainly divided into three parts: first, the feature map F of the two images to be matched is obtained through the encoder A and F B , and then the position embedding of image B is obtained for the subsequent Gaussian process module; the Gaussian module produces a predicted posterior; then, the embedding decoder finds the matching score of the image pair;
[0048] The specific model solution can be expressed by the following formula:
[0049]
[0050] Among them, G θ Represents a global matcher, consisting of an encoder E θ and a decoder D θ The encoder uses Gaussian process regression and the decoder uses Transformer.
[0051] The output of the GP posterior is unimodal, so multimodal matching will reduce its performance. In image matching, one point in the image may correspond to multiple candidate points, so how to better handle multimodality becomes a problem that coordinate regression needs to deal with. To address this problem, the cosine coordinate embedding method is adopted:
[0052] Emb F (x; W, b) = cos(8πWx)
[0053] in, represents the image coordinates,
[0054] First, in Gaussian process regression, the output is considered as a collection of random variables. It is necessary to assume that these variables are jointly Gaussian, because this assumption ensures that Gaussian process regression can model the correlation of data in a probabilistic way. Secondly, since the Gaussian process is uniquely defined by its kernel, and the role of the kernel function is to define the covariance matrix between the outputs, it must be positive definite. Specifically, if the kernel function is non-positive definite, then the covariance matrix of this Gaussian distribution may be singular, since the observation value f follows a multivariate normal distribution. Therefore, a non-positive kernel function will make it impossible to calculate the probability density and perform Bayesian updates (because the inverse of the covariance matrix is needed to update the posterior). Furthermore, when the dimension of the kernel matrix is large, if it is directly diagonalized, the amount of calculation will be very large, which is not feasible on large-scale data. Therefore, in order to reduce the computational complexity, the assumption that the coordinate embedding dimension is unrelated is used here to make the kernel block diagonalization feasible. In this way, the kernel matrix can be split into multiple smaller diagonal blocks, thereby realizing block processing to achieve the purpose of reducing the amount of calculation. The specific formula is as follows:
[0055] After making the above assumptions, it is important to choose a suitable kernel function. Here, the exponential cosine similarity kernel is selected. First, the cosine similarity is calculated. The specific formula is as follows:
[0056]
[0057] Among them, x i ,y i Represents the eigenvectors of feature maps A and B. ε is used for numerical stability. Then the exponential cosine similarity kernel is calculated. The formula is as follows:
[0058]
[0059] Where T is the temperature hyperparameter, which is used to control the sensitivity of the kernel function.
[0060] After understanding the basic assumptions and the choice of kernel function, the specific practice of Gaussian process regression is: In the standard assumption, the observation It is obtained under independent and identically distributed noise, and a probability model is used to describe the distribution of observations. Based on this, for the feature map The analytical formula for the posterior condition is as follows:
[0061]
[0062] in, represents the coordinates after cosine embedding, K represents the kernel matrix, μ is the posterior mean function, σ n =0.1 is the standard deviation of the measurement noise, Σ is the posterior covariance, which represents the uncertainty. The formula in the lower part describes the uncertainty when φ is known. B After that, it is relative to φ A How big is the uncertainty if φ A and φ B Very similar, i.e. is large, then the uncertainty will decrease, but if φ A Distance φ B Very far, that is is approximately 0, then the variance of the prediction will be close to The uncertainty at this time will be very great.
[0063] The decoder uses classification regression to discretize the output space. The specific formula is as follows:
[0064]
[0065] Where K is the number of discrete anchor points, π k express The probability of predicting the k-th anchor point, represents a 2D uniform basis distribution, m k The output shape of this block is B×H×W×(K+1), where K is the number of anchor points and the extra dimension is the matching score p. The last dimension is required to extract the overlapping area.
[0066] In image analysis, given a set of images, denoted as (Ia ,I b ), where each image is accompanied by its associated keypoint set (P a ,P b ), their respective score vectors (S a ,S b ) and feature descriptors (F a ,F b ). It is worth noting that each pixel coordinate is associated with a corresponding position confidence. Therefore, image I a and I b Two different sets of feature descriptors are generated, respectively denoted as and The matching of image features is defined as identifying the optimal one-to-one correspondence matrix As shown in the formula:
[0067]
[0068] The goal of this work is to minimize the relative pose error of image pairs by utilizing all matching pairs. Each feature descriptor d is allowed to match at most once for pose estimation, and keypoints or descriptors that still do not match are considered outliers. Therefore, an alignment matrix is derived It represents the pairwise correspondence between the features of two images.
[0069] The specific process of cutting out the overlapping area is as follows:
[0070] Normalize the matching score p within each sample so that the sum of each sample is 1:
[0071]
[0072] Among them, p norm Represents the normalized matching score matrix. Then, the cumulative sum of confidence is calculated in the W direction (width) and H direction (height), respectively, as follows:
[0073]
[0074] The clipping boundaries are then determined based on this:
[0075]
[0076] Among them, arg max returns the first index that meets the conditions.
[0077] After obtaining the area to be cropped, it is necessary to resize the area to reduce the impact of scale differences on the matching effect. The scale adjustment ratio is calculated as follows:
[0078]
[0079] At this point, the preprocessing of the overlapping area of the image pair to be matched is completed, and the processed cropped image is input into the feature matching network for subsequent matching.
[0080] Next, the cropped overlapping areas are subjected to local feature matching using existing suitable feature extractors and feature matchers. The processed overlapping area images are then subjected to local feature matching using the latest open-source SoTA sparse feature matching method, LightGlue, and its accompanying feature extractor, Superpoint. Finally, post-processing restores the key points back to the original image and performs downstream tasks. The reason for choosing a detector-based image feature matching method here is that it is based on feature point matching rather than pixel-level matching, which can save computing time and resources. At the same time, because the image preprocessing for overlapping area extraction has been performed, the noise background is largely eliminated. Even sparse matching methods can achieve good matching results in complex scenes (such as weak textures and large-scale changes) without failing due to noise interference.
[0081] (1) The images extracted by this technical solution are more refined in the overlapping areas and can eliminate more interference from invalid feature points to achieve better matching results.
[0082] (2) This technical solution combines the advantages of existing sparse and dense matching methods, which can not only retain representative feature points, but also achieve better image feature matching under extreme viewing angles and scale changes.
[0083] (3) This technical solution has the characteristics of plug-and-play, and can flexibly transform the Detector and Matcher frameworks to integrate the latest research progress and achieve a matching effect that keeps pace with the times. The present invention provides an image feature matching method based on overlapping areas.
[0084] The present invention discloses an image feature matching method based on overlapping areas, which improves the rationality of combining dense matching with overlapping area extraction. Since the functions of dense matching and overlapping areas overlap, it can be considered to switch to sparse matching, which improves the speed and representativeness of key points while making the structure more reasonable. By improving the estimation method of overlapping areas and introducing a more refined area extraction mechanism, the accuracy of overlapping areas is improved.
[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the scope of protection of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. A method for image feature matching based on overlapping areas, characterized in that: The following steps are involved: First, the overlapping area of the two images is extracted by the algorithm to obtain the matching score of the two images, and based on this, the overlapping area in the image is extracted, and then the overlapping area is cropped and resized; Input the processed cropped image into the feature matching network for subsequent matching; The cropped overlapping area is subjected to local feature matching through the existing feature extractor and feature matcher.
2. The image feature matching method based on overlapping areas according to claim 1, characterized in that: The mapping relationship between finding overlapping areas and finding points between images is the same, that is, a point in image A can be mapped to image B, indicating that the point is in the overlapping area of the two images. The overlapping area extraction method can be defined as estimating image I A ,I B The mapping relationship between pixels W A→B And its corresponding matching score p θ , where the overlapping area can be extracted only by the matching score without using the mapping relationship, and the solution is the matching score p between the pixel pairs θ ; The overlapping area extraction algorithm is mainly divided into three parts: First, the feature map F of the two images to be matched is obtained through the encoder. A and F B , and then the position embedding of image B is obtained for the subsequent Gaussian process module; the Gaussian module produces a predicted posterior; then, the embedding decoder finds the matching score of the image pair; The specific model solution can be expressed by the following formula: Among them, G θ Represents a global matcher, consisting of an encoder E θ and a decoder D θ The encoder uses Gaussian process regression and the decoder uses Transformer.
3. The image feature matching method based on overlapping areas according to claim 2, characterized in that: The output of the GP posterior is unimodal, so multimodal matching will reduce its performance. In image matching, a point in the image may correspond to multiple candidate points, and the cosine coordinate embedding method is used: Emb F (x;W,b)=cos(8πWx) in, represents the image coordinates, 4. The image feature matching method based on overlapping areas according to claim 3, characterized in that: The specific approach of Gaussian process regression: In the standard assumption, the observations It is obtained under independent and identically distributed noise, and a probability model is used to describe the distribution of observations. Based on this, for the feature map The analytical formula for the posterior condition is as follows: in, represents the coordinates after cosine embedding, K represents the kernel matrix, μ is the posterior mean function, σ n =0.1 is the standard deviation of the measurement noise, Σ is the posterior covariance, which represents the uncertainty. The formula in the lower part describes the uncertainty when φ is known. B After that, it is relative to φ A How big is the uncertainty if φ A and φ B Very similar, i.e. is large, then the uncertainty will decrease, but if φ A Distance φ B Very far, that is is approximately 0, then the variance of the prediction will be close to The uncertainty at this time will be very great.
5. The image feature matching method based on overlapping areas according to claim 4, characterized in that: The decoder uses classification regression to discretize the output space. The specific formula is as follows: Where K is the number of discrete anchor points, π k express The probability of predicting the k-th anchor point, represents a 2D uniform basis distribution, m k Represents the position coordinates of the kth anchor point, and the output shape is B×H×W×(K+1), where K is the number of anchor points and the extra dimension is the matching score p. The last dimension is required to extract the overlapping area.
6. The image feature matching method based on overlapping areas according to claim 5, characterized in that: In image analysis, given a set of images, denoted as (I a ,I b ), where each image is accompanied by its associated keypoint set (P a ,P b ), their respective score vectors (S a ,S b ) and feature descriptors (F a ,F b ), it is worth noting that each pixel coordinate is associated with the corresponding position confidence, so the image I a and I b Two different sets of feature descriptors are generated, respectively denoted as and The matching of image features is defined as identifying the optimal one-to-one correspondence matrix As shown in the formula: The goal is to minimize the relative pose error of the image pair by utilizing all matching pairs. Each feature descriptor d is allowed to match at most once for pose estimation. Key points or descriptors that still do not match are considered outliers, and an alignment matrix is derived. Represents the pairwise correspondence between the features of two images.
7. The image feature matching method based on overlapping areas according to claim 1, characterized in that: The specific process of cutting out the overlapping area is as follows: Normalize the matching score p within each sample so that the sum of each sample is 1: Among them, p norm Represents the normalized matching score matrix, and the cumulative sum of confidence is calculated in the width W direction and the height H direction respectively. The formula is as follows: The clipping boundaries are then determined based on this: Among them, arg max returns the first index that meets the conditions. After obtaining the area to be cropped, it is necessary to resize the area to reduce the impact of scale differences on the matching effect. The scale adjustment ratio is calculated as follows: At this point, the preprocessing of the overlapping area of the image pair to be matched is completed, and the processed cropped image is input into the feature matching network for subsequent matching.
8. The image feature matching method based on overlapping areas according to claim 1, characterized in that: The processed overlapping area images are subjected to local feature matching through the latest open source SoTA sparse feature matching method LightGlue and its supporting feature extractor Superpoint.