An Image Feature Point Matching Method Based on Affine Transformation and Deep Learning

The integration of affine transformations and deep learning with SuperPoint for feature extraction significantly improves image matching accuracy and robustness to large geometric transformations.

CN116452839BActive Publication Date: 2025-07-15ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310373428.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-10
Publication Date
2025-07-15
Estimated Expiration
2043-04-10

AI Technical Summary

Technical Problem

Existing image matching algorithms are difficult to accurately match feature points when viewing angles and scales change greatly, especially traditional methods and some deep learning methods do not perform well in complex scenarios and occlusions.

Method used

The SuperPoint algorithm, which combines affine transformation and deep learning, uses a convolutional neural network to extract feature points, and uses inverse projection of the affine transformation matrix and Euclidean distance matching feature points, and combines the random sampling consistency algorithm to select the correct matching pairs.

Benefits of technology

It improves the success rate and accuracy of feature point matching, especially when the perspective angle changes are large, it can effectively deal with complex scenes and occlusion situations, and improves the number and internal point rate of matching point pairs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116452839B_ABST
    Figure CN116452839B_ABST
Patent Text Reader

Abstract

The present invention discloses an image feature point matching method based on affine transformation and deep learning. The present invention generates multiple different images under affine transformation, uses a deep learning network to extract the feature points of the image pair, and generates corresponding descriptors. By calculating the inverse of the affine transformation matrix, the feature points are back-projected onto the original image, and the Euclidean distance of the descriptors is used to match the feature points. Finally, the correct matching pairs are screened out by the random sample consensus algorithm. The present invention combines affine transformation and deep learning feature points, extracts feature points at multiple transformation scales, improves the accuracy and robustness of image feature point matching, enables the finally obtained feature points to have strong scale and deformation invariance, thereby improving the success rate of feature point matching in large-scale and perspective deformation scenarios. The present invention has achieved a significant improvement in the matching success rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image feature point matching method in the field of image matching technology, and particularly to an image feature point matching method based on affine transformation and deep learning. Background Art

[0002] Image matching is an important research direction in the field of computer vision, mainly referring to the process of finding similar points or identical regions between two or more images, which mainly includes the following aspects: feature extraction, feature description, feature matching, and outlier filtering. Traditional image matching algorithms are divided into two categories. One is the feature point matching method, which first needs to extract and describe the feature points in two images, and then matches the feature points in the two images by comparing the distances of feature vectors. Commonly used feature point matching algorithms include SIFT, SURF, ORB, etc. The other is the gray-level matching method, such as optical flow matching, template matching, etc. Among them, the template matching method first needs to select a template image in one image, and then search for the region most similar to the template image in another image. Commonly used template matching algorithms include mean absolute error matching, squared error matching, etc. The optical flow matching algorithm is a matching method based on motion information. This method analyzes the pixel changes between two consecutive frames, calculates the motion direction and speed of each pixel point, and then matches the two images according to the motion information. Commonly used optical flow matching algorithms include block-based optical flow matching, globally optimized optical flow matching, etc.

[0003] In recent years, with the development of deep learning technology, more and more image matching algorithms based on neural networks have been proposed and achieved better results than traditional methods. Convolutional neural network is one of the most commonly used deep learning algorithms, which can automatically learn features from the original image. While retaining the original structural information of the image, CNN can effectively extract and represent the features in the image, making it an important tool in image matching. The feature description algorithm describes the extracted features, and the commonly used method is to transform them into some vectors. Well-known deep learning feature points include Superpoint, etc. Through big data learning and training, they can have a certain adaptability to a small amount of deformation and scale transformation. However, the effect is not good when the scale and deformation are large. Summary of the Invention

[0004] The present invention aims to solve the problems existing in the background art and provides an image feature point matching method based on affine transformation and deep learning, so as to greatly improve the image feature matching effect. The present invention proposes a method of using the SuperPoint algorithm invariant to affine transformation to extract features, aiming to solve the problem that it is difficult to accurately match between images with large changes in viewpoints and scales. The proposed image matching method combines affine transformation and the deep learning SuperPoint feature point extraction method, effectively increasing the number of matching point pairs and the inlier rate of the matching.

[0005] The present invention proposes an image feature point matching method based on affine transformation and deep learning. The present invention uses affine transformation to generate multiple different images and retains the corresponding affine transformation matrices. At the same time, a convolutional neural network is adopted to automatically learn the feature points in the images and generate descriptors, calculate the inverse of the affine transformation matrix, back-project the feature points to the original images, and use the Euclidean distance of the descriptors to match the feature points in the images. Finally, the correct matching pairs are screened out by the Random Sample Consensus (RANSAC) algorithm.

[0006] The present invention can effectively improve the success rate of feature point matching, especially in the case of large viewpoint transformation, and has a good matching effect. Its feature lies in using the SuperPoint algorithm for feature point extraction and combining affine transformation to achieve more accurate matching.

[0007] The technical solution adopted by the present invention is as follows:

[0008] 1) After preprocessing the pair of images to be matched, the first preprocessed image and the second preprocessed image are obtained;

[0009] 2) After respectively performing affine transformation and interpolation on the first preprocessed image and the second preprocessed image by using multiple affine transformation matrices, the corresponding multiple affine-transformed images are respectively obtained;

[0010] 3) Respectively extract and back-project the feature points of the first preprocessed image and the second preprocessed image corresponding to the multiple affine-transformed images, thereby constructing a first feature point set and a second feature point set;

[0011] 4) Use the brute-force matching method to match the feature points before affine transformation in the first feature point set and the second feature point set to obtain an initial matching pair set;

[0012] 5) Use the Random Sample Consensus algorithm to screen the initial matching pair set to obtain the final matching pair set.

[0013] The specific content of 1) is as follows:

[0014] The two images in the image pair to be matched are respectively subjected to grayscale and normalization processing, and the corresponding grayscale-normalized images are respectively obtained and denoted as the first preprocessed image and the second preprocessed image.

[0015] 3.1) Use a convolutional neural network to respectively extract the feature points of the corresponding multiple affine-transformed images of the first preprocessed image and the second preprocessed image, so as to obtain a first initial feature point set and a second initial feature point set;

[0016] 3.2) According to the principle of non-maximum suppression, respectively remove the feature points with lower confidence in the first initial feature point set and the second initial feature point set, and respectively obtain the corresponding first denoised feature point set and second denoised feature point set;

[0017] 3.3) According to multiple affine transformation matrices, respectively perform back-projection on the first denoised feature point set and the second denoised feature point set, and respectively obtain a first feature point set before affine transformation and a second feature point set before affine transformation;

[0018] 3.4) Respectively remove the feature points outside the boundaries of the corresponding preprocessed images in the first feature point set before affine transformation and the second feature point set before affine transformation, so as to obtain the final first feature point set and second feature point set.

[0019] The convolutional neural network in the above 3.1) is a SuperPoint network, that is, a SuperPoint network.

[0020] In the above 4), for each feature point i before affine transformation in the first feature point set, calculate the Euclidean distance between the feature point i before affine transformation and the descriptors corresponding to all feature points before affine transformation in the second feature point set, and use the feature point before affine transformation with the smallest Euclidean distance in the second feature point set as the matching point of the feature point i before affine transformation and form a matching pair. If the similarity between the descriptors of the current matching pair is less than a preset threshold, then remove the matching pair, otherwise retain the matching pair.

[0021] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0022] 1. Compared with the original SuperPoint algorithm, the present invention introduces affine invariance into the extraction of feature points and the generation of descriptors, making it have better robustness to affine transformations such as rotation and scaling of images.

[0023] 2. The present invention can handle various affine transformations, and also has good performance in the extraction of feature points in complex scenes and occlusion situations. At the same time, the descriptors obtained by the present invention are unique, making it show good performance in both matching effect and computational efficiency. Description of the Drawings

[0024] Figure 1 This is the method flow chart of the present invention.

[0025] Figure 2 This is the matching result diagram of the present invention. Detailed implementation manners

[0026] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0027] As Figure 1 shown, the present invention includes the following steps:

[0028] 1) After preprocessing the image pairs to be matched, a first preprocessed image and a second preprocessed image are obtained;

[0029] 1) Specifically:

[0030] The two images in the image pairs to be matched have the same size. In specific implementation, the smaller the size of the image, the higher the processing efficiency. The two images in the image pairs to be matched are respectively subjected to grayscale and normalization processing, and the corresponding grayscale normalized images are respectively obtained and denoted as the first preprocessed image and the second preprocessed image. Among them, first, the image is subjected to grayscale conversion, and then the grayscale image is divided by 255 to obtain the normalized image.

[0031] 2) After respectively performing affine transformation and interpolation on the first preprocessed image and the second preprocessed image by using a plurality of affine transformation matrices, a plurality of corresponding affine transformed images are respectively obtained;

[0032] In the process of generating the affine transformation image, it includes three parameters: rotation, scaling, and translation. Among them, the settings of rotation and scaling need to be controlled by some parameters. The image is shrunk or enlarged in the scale space, and the scaled image is rotated. Through these operations, a group of images with different rotations and scalings are generated, and the affine transformation matrix at this time is recorded. The present invention selects different rotation angles and scaling factors for image transformation. Here, the scaling factor t is set to Here is a trade-off between accuracy and the number of images, that is, in specific implementation, the affine transformation matrix of each image is 6. The rotation is based on the scaling value, in the interval [0, 180°], and the rotation angle interval phi is set to 72 / t. When the scaling parameter t = 1 and the rotation parameter phi = 0, no change is made. Finally, the perspective scaling change between the two images can reach 32, and the covered area can reach 13.5 times that of the original image.

[0033] The affine transformation matrix A can be expressed by the following formula:

[0034]

[0035] Among them, a 11 、a 12 、a 21 、a 22 respectively represent the first to fourth parameters after the combined effects of image scaling, rotation, and shearing, and t x and t y represent the first and second translation parameters.

[0036] After generating the image, since the transformed coordinates may not be integers, it is necessary to interpolate the value of each pixel to obtain the pixel value in the final image.

[0037] 3) Respectively extract and back-project the feature points of the first preprocessed image and the second preprocessed image corresponding to multiple affine-transformed images, so as to construct a first feature point set and a second feature point set;

[0038] 3) Specifically:

[0039] 3.1) Use a convolutional neural network (CNN) to respectively extract the feature points of the first preprocessed image and the second preprocessed image corresponding to multiple affine-transformed images, so as to obtain a first initial feature point set and a second initial feature point set; the convolutional neural network is the SuperPoint network.

[0040] 3.2) Remove the feature points with lower confidence in the first initial feature point set and the second initial feature point set respectively according to the principle of non-maximum suppression, and obtain the corresponding first denoised feature point set and second denoised feature point set respectively;

[0041] 3.3) Back-project the first denoised feature point set and the second denoised feature point set respectively according to multiple affine transformation matrices, and obtain a first set of feature points before affine transformation and a second set of feature points before affine transformation respectively; specifically, solve the inverse matrix of the affine transformation matrix A, and use the inverse matrix to back-project the corresponding feature points to obtain the corresponding feature points before affine transformation.

[0042] 3.4) Remove the feature points outside the boundaries of the corresponding preprocessed images in the first set of feature points before affine transformation and the second set of feature points before affine transformation respectively, so as to obtain the final first feature point set and second feature point set.

[0043] CNN is a powerful image processing tool that can automatically extract high-level features of images. In the present invention, the SuperPoint algorithm is used as the feature point detector, which is based on the CNN network and can extract rotation-invariant and scale-invariant feature points.

[0044] 4) Use the brute-force matching method to match the feature points before affine transformation in the first set of feature points and the second set of feature points, and obtain the initial set of matching pairs;

[0045] In 4), for each feature point i before affine transformation in the first set of feature points, calculate the Euclidean distance between the feature point i before affine transformation and the descriptors corresponding to all feature points before affine transformation in the second set of feature points. Take the feature point before affine transformation with the smallest Euclidean distance in the second set of feature points as the matching point of the feature point i before affine transformation and form a matching pair. If the similarity between the descriptors of the current matching pair is less than the preset threshold, remove the matching pair; otherwise, retain the matching pair.

[0046] Among them, the calculation formula of the Euclidean distance is:

[0047]

[0048] Among them, m m,j represents the matching degree between the i-th feature point in the first preprocessed image and the j-th feature point in the second preprocessed image, that is, the Euclidean distance.

[0049] 5) Since there may be problems such as occlusion and noise in the image, the matching may not be accurate. Use the Random Sample Consensus (RANSAC) algorithm to screen the initial set of matching pairs to obtain the final set of matching pairs. Calculate the H matrix based on these matching points, and then calculate the reprojection error of the matching points after being mapped by the H matrix. Classify the points with an error less than a threshold as inliers, and the points with a distance greater than the threshold as outliers. Next, repeat this process multiple times and select the H with the most inliers as the optimal model.

[0050] Specifically:

[0051] Randomly select a set of samples from the matching pairs, and estimate the H matrix based on this set of samples. Calculate the reprojection error of the matching pairs after being mapped by the H matrix. Classify the points with a distance less than a threshold as inliers, and the points with a distance greater than the threshold as outliers. If the number of inliers exceeds the preset threshold, use the inliers to re-estimate the H matrix. Repeat this process multiple times and select the model with the most inliers as the optimal model.

[0052] The present invention combines affine transformation and deep learning methods. It uses the SuperPoint algorithm to extract feature points and descriptors, improving the accuracy and robustness of image feature point matching. Under affine transformation, multiple different images are generated. The deep learning network is used to extract the feature points of each image and generate corresponding descriptors. By calculating the inverse of the affine transformation matrix, the feature points are back-projected onto the original image, and the Euclidean distance of the descriptors is used to match the feature points. Finally, the correct matching pairs are selected through the Random Sample Consensus (RANSAC) algorithm. Through this method, the present invention can effectively process images with large perspective changes, increasing the number of matching point pairs and the inlier rate of the matching.

[0053] To evaluate the effect of the present invention, the present invention uses a publicly available dataset for experiments and obtains the following quantitative analysis results: namely, the number of matched feature points and the inlier rate of the matching after outlier filtering. In the case of large perspective changes, the number of feature points and the inlier rate of the present invention are significantly better than traditional methods. Therefore, the present invention can draw the conclusion that through the way of affine transformation, the inlier rate of the matching and the number of feature points are greatly improved in the case of large perspective changes. The qualitative results can be seen in Table 1, and the qualitative analysis results can be seen Figure 2 , Figure 2 (a)-(j) of are the image matching diagrams corresponding to 10 kinds of image pairs. In each image matching diagram, from left to right are the image matching diagrams obtained after processing with SIFT, SuperPoint, and Affine SuperPoint (i.e., the present invention). It can be seen from this that the number of feature points and the inlier rate extracted by the present invention on 10 images are the highest. Although the inlier rate is only about 1% higher than that of SuperPoint, the number of feature point matching pairs of the present invention is larger, and thus the corresponding number of inlier matching is also larger, proving the feasibility and effectiveness of the present invention.

[0054] Table 1 Number of image matching feature points and inlier rate

[0055]

[0056] The above embodiments are used to explain the present invention, rather than limit the present invention. Any modification and change made to the present invention within the spirit and scope of the protection of the claims of the present invention fall within the protection scope of the present invention.

Claims

1. An image feature point matching method based on affine transformation and deep learning, characterized in that, It includes the following steps: 1) After preprocessing the image pairs to be matched, obtain the first preprocessed image and the second preprocessed image; 2) Use multiple affine transformation matrices to perform affine transformation and interpolation on the first preprocessed image and the second preprocessed image respectively, and obtain the corresponding multiple affine-transformed images respectively; 3) Extract and back-project the feature points of the corresponding multiple affine-transformed images of the first preprocessed image and the second preprocessed image respectively, so as to construct the first feature point set and the second feature point set; The specific content of 3) is: 3.1) Use a convolutional neural network to extract the feature points of the corresponding multiple affine-transformed images of the first preprocessed image and the second preprocessed image respectively, so as to obtain the first initial feature point set and the second initial feature point set; 3.2) Remove the feature points with relatively low confidence in the first initial feature point set and the second initial feature point set respectively according to the principle of non-maximum suppression, and obtain the corresponding first denoised feature point set and the second denoised feature point set respectively; 3.3) Back-project the first denoised feature point set and the second denoised feature point set respectively according to multiple affine transformation matrices, and obtain the first set of feature points before affine transformation and the second set of feature points before affine transformation respectively; 3.4) Remove the feature points outside the boundaries of the corresponding preprocessed images in the first set of feature points before affine transformation and the second set of feature points before affine transformation respectively, so as to obtain the final first feature point set and the second feature point set; 4) Use the brute-force matching method to match the feature points before affine transformation in the first feature point set and the second feature point set, and obtain the initial matching pair set; 5) Use the random sample consensus algorithm to screen the initial matching pair set, and obtain the final matching pair set.

2. The image feature point matching method based on affine transformation and deep learning according to claim 1, characterized in that, The specific content of 1) is: Perform gray-scale and normalization processing on the two images in the image pairs to be matched respectively, and obtain the corresponding gray-scale normalized images, which are denoted as the first preprocessed image and the second preprocessed image respectively.

3. A method for matching image feature points based on affine transformation and deep learning according to claim 1, characterized in that The convolutional neural network in 3.1) is the SuperPoint network.

4. A method for image feature point matching based on affine transformation and deep learning according to claim 1, characterized in that, In 4), for each feature point i before affine transformation in the first feature point set, calculate the Euclidean distance between the feature point i before affine transformation and the descriptors corresponding to all feature points before affine transformation in the second feature point set, and use the feature point before affine transformation with the smallest Euclidean distance in the second feature point set as the matching point of the feature point i before affine transformation and form a matching pair. If the similarity between the descriptors of the current matching pair is less than the preset threshold, remove the matching pair, otherwise retain the matching pair.

Citation Information

Patent Citations

  • Image classification network based on SuperPoint feature

    CN110263868A

  • Image feature matching method based on affine transformation and ORB

    CN115689873A