A method for constructing regional feature descriptors adapting to high affine transformation

By combining simulation affine transformation and MSER algorithm, regional feature descriptors adapted to high affine transformation are generated, which solves the descriptor invariance and difference problems and improves the accuracy of feature matching.

CN116503617BActive Publication Date: 2025-06-06JIANGNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310352872.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-04
Publication Date
2025-06-06
Estimated Expiration
2043-04-04

AI Technical Summary

Technical Problem

Under high affine transformation conditions, it is difficult for the prior art to maintain the invariance and difference of feature descriptors, resulting in a decrease in feature matching accuracy.

Method used

Through classification simulation affine transformation, the degree of affine of the image to be matched is judged, and SIFT feature points are described on the simulated image set. Image segmentation is performed by combining the MSER algorithm, and area description information is generated using grayscale histograms and grayscale center of mass method, and fused with feature point descriptors.

Benefits of technology

Improve the accuracy of feature matching, enhance the descriptor difference under high affine transformation conditions, and ensure the stability and accuracy of feature point matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116503617B_ABST
    Figure CN116503617B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for constructing a regional feature descriptor that is adaptable to high affine transformation, and belongs to the field of computer vision technology. The present invention simulates affine transformation for image classification with different affine degrees to generate a new image set, and then calculates the neighborhood information of feature points on the new image set, and finally combines the grayscale histogram of the maximum stable extreme value region to which the feature point belongs, and the normalized position of the feature point relative to the grayscale centroid of the region to generate a new descriptor. By comparing the feature matching indicators under the affine transformation scenario, it is proved that the descriptor of the present invention has higher precision and robustness than the existing classical descriptors, and has the robustness of fusion with other descriptors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a method for constructing a regional feature descriptor adapting to high affine transformation, and belongs to the technical field of computer vision. Background Art

[0002] Feature descriptors are used to distinguish the differences between different feature points in the same image and to maintain the same information of the same feature point in different images. They are used in target detection, image matching, image retrieval, 3D reconstruction, and other fields. However, when an image undergoes an affine transformation, especially when the inclination angle between the optical axis and the scene is greater than 40 degrees, the neighborhood grayscale information used to describe the feature point will change significantly, making it difficult for the descriptor to maintain the difference and invariance of the feature point information.

[0003] In order to maintain the invariance of descriptors, the effective solution strategy in the current methods is to maintain the invariance of descriptors by simulating image transformations under different degrees of affine. However, the existing simulation transformation strategies all use simulation transformations of fixed data. For example, the famous ASIFT (Affine Scale Invariant Feature Transform) algorithm only simulates the magnified affine transformation of the image, which results in a larger affine gap between the two sets of affine images generated when facing a larger affine transformation, and the invariance of the same descriptor cannot be guaranteed. In addition, in order to maintain the difference of descriptors, the effective solution strategy in the current methods is to use the neighborhood information supplement method, but the limited neighborhood information often makes it impossible to distinguish the difference of descriptors. If the description area is directly expanded, although it can effectively alleviate the similarity problem of descriptors in different large areas, it will also cause the overlap of regional information corresponding to similar descriptors in the same large area, which seriously affects the difference of descriptors in the same large area, thereby reducing the accuracy of feature point matching based on such descriptors. Summary of the invention

[0004] In order to solve the problem that it is difficult to maintain the identity and difference of descriptors due to changes in regional information caused by a large degree of affine transformation, so as to improve the accuracy of subsequent feature matching, the present invention provides a method for constructing regional feature descriptors that adapt to high affine transformation. The technical solution is as follows:

[0005] The first object of the present invention is to provide a method for constructing an image region feature descriptor, comprising:

[0006] Step 1: In the classified simulated affine transformation, the affine degree of the image to be matched is judged, and a relatively high / low degree of affine transformation is obtained, thereby respectively reducing / increasing the affine degree of the simulated image, and the existing SIFT feature point description is performed on the obtained simulated image set;

[0007] Step 2: During the adding of the descriptor region information, image enhancement is performed on the image to be matched;

[0008] Step 3: performing region segmentation on the image enhanced in step 2 based on the MSER algorithm;

[0009] Step 4: Use the grayscale histogram to describe the segmented image area and obtain the area description information;

[0010] On the basis of normalization of regional coordinates, the grayscale centroid method is used to rotate the main direction of the region to be consistent with the positive direction of the x-axis, and the relative position of the feature point and the normalized centroid is calculated;

[0011] Step 5: assign weights to the region description information and the relative position information respectively, and fuse the weighted region description information and the relative position information with the feature point descriptor obtained in step 1.

[0012] Optionally, the process of determining the degree of affine in step 1 includes:

[0013] Simulate adding the input images a and b to be matched respectively times the affine amount, generate images A and B respectively, and match them through SIFT feature extraction and description algorithm. The number of feature matches between image A and image B is M aB , the number of feature matches between image A and image b is M Ab , if M aB >M Ab , then image a is an image with a lower degree of affine, otherwise image b is an image with a lower degree of affine.

[0014] Optionally, the process of reducing / increasing the affine degree of the simulated image in step 1 includes:

[0015] For the image with a lower affine degree in the input image, the simulated enlarged affine amount t_1 is used, and t_1∈{√2,2,2√2}; for the image with a higher affine degree in the input image, the simulated reduced affine amount t_2 is used, and t_2∈{√2 / 2,1 / 2,√2 / 4}. The tilt angle φ of the two images is simulated from -45 degrees to 45 degrees, with a step size of 15 degrees.

[0016] Optionally, in step 2, after enhancing the image contrast using the CLAHE algorithm, a bilateral filter is used to smooth the image area details.

[0017] Optionally, in step 2, the bilateral filter adds a weight to the pixel gray value on the basis of the Gaussian filter, thereby reducing the regional details while strengthening the retention of edge information, as shown in the following formula:

[0018]

[0019]

[0020]

[0021]

[0022] Where h(x) is the output grayscale value, k(x) is the amount of weight normalization, x is the input pixel, S is the filter window corresponding to the x pixel, f(x) is the grayscale value of pixel x, c(ξ,x) is the spatial weight, s(f(ξ),f(x)) is the grayscale weight, σ d and σ r is the standard deviation of spatial proximity and grayscale similarity.

[0023] Optionally, step 4 includes:

[0024] Step 4-1: Normalize the coordinates of the region, and regard the region as the integral of several parallel line segments. The coordinate normalization formula is:

[0025]

[0026]

[0027] Where (x, y) is the input pixel coordinate, (norm_x, norm_y) is the normalized coordinate, min_x, min_y, max_x and max_y are the minimum and maximum x and y coordinate values ​​of all pixel coordinates in the region;

[0028] Step 4-2: When an affine transformation occurs in the same area, the grayscale of the pixel becomes extremely small and the line segment ratio remains unchanged. The grayscale centroid of the area will correspond to the same pixel point. The coordinates of the area centroid C and the main direction angle θ of the area can be calculated by the following formula:

[0029] m pq =∑ x,y∈S x p y q I(x,y) p,q={0,1}

[0030]

[0031]

[0032] Among them, m pq is the moment of region S, m 00 Represents the total grayscale of pixels in area S, m 10 is the grayscale moment of region S in the x direction, m 01is the grayscale moment of region S in the y direction, and the influence of rotation on the coordinate value in affine transformation is eliminated by turning the positive direction of the x-axis of the region to the main direction;

[0033] Step 4-3: Based on the coordinate normalization and the area centroid C, the relative position of the feature point and the centroid is calculated.

[0034] The second object of the present invention is to provide an image matching method, which first constructs a region feature descriptor for an input image to be matched using the above-mentioned image region feature descriptor construction method, and then performs feature matching based on the constructed feature descriptor.

[0035] The third object of the present invention is to provide an image matching device, comprising: a processor and a memory, wherein the memory stores instructions executed by the processor, and when the instructions are executed by the processor, the image matching device implements the above-mentioned image matching method.

[0036] A fourth object of the present invention is to provide a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the above-mentioned image matching method is implemented.

[0037] The beneficial effects of the present invention are:

[0038] The present invention uses a strategy of classified simulated affine transformation to solve the problem of difficulty in maintaining the invariance of the same descriptor due to obvious regional changes when a large affine transformation occurs in the image to be matched, and generates a new descriptor based on the existing SIFT descriptor combined with regional information that is invariant under affine transformation, thereby enhancing the difference between different descriptors when distinguishing similar small regions in affine changes. Simulation results under different data sets prove that the descriptor proposed by the present invention can more accurately improve the accuracy of feature matching. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0040] Figure 1 It is an overall flow chart of a regional feature descriptor adapted to high affine transformation according to an embodiment of the present invention.

[0041] Figure 2 Schematic diagram of the affine degree of the original image in an embodiment of the present invention.

[0042] Figure 31 is an image enhancement result diagram in an embodiment of the present invention, wherein (a) is the original image, (b) is the effect diagram of the CLAHE algorithm, and (c) is the effect diagram of the enhancement by the CLAHE algorithm and the bilateral filtering algorithm.

[0043] Figure 4 1 is an image segmentation result diagram in an embodiment of the present invention, wherein (a) is the image to be segmented, and (b) is the image after segmentation.

[0044] Figure 5 Result diagrams of two standard image sets, Graffiti (a) and Wall (b), in an embodiment of the present invention. DETAILED DESCRIPTION

[0045] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0046] Embodiment 1:

[0047] This embodiment provides a method for constructing an image region feature descriptor, including:

[0048] Step 1: In the classified simulated affine transformation, the affine degree of the image to be matched is judged, and a relatively high / low degree of affine transformation is obtained, thereby respectively reducing / increasing the affine degree of the simulated image, and the existing SIFT feature point description is performed on the obtained simulated image set;

[0049] Step 2: During the adding of the descriptor region information, image enhancement is performed on the image to be matched;

[0050] Step 3: performing region segmentation on the image enhanced in step 2 based on the MSER algorithm;

[0051] Step 4: Use the grayscale histogram to describe the segmented image area and obtain the area description information;

[0052] On the basis of normalization of regional coordinates, the grayscale centroid method is used to rotate the main direction of the region to be consistent with the positive direction of the x-axis, and the relative position of the feature point and the normalized centroid is calculated;

[0053] Step 5: assign weights to the region description information and the relative position information respectively, and fuse the weighted region description information and the relative position information with the feature point descriptor obtained in step 1.

[0054] Embodiment 2:

[0055] This embodiment provides a method for constructing a regional feature descriptor that is adaptable to high affine transformation, see Figure 1 , including the following steps:

[0056] Step 1: In the classified simulated affine transformation, the affine degree of the original image is judged, and a relatively high / low degree of affine transformation is obtained, thereby respectively reducing / increasing the affine degree of the simulated image, and the existing SIFT feature point description is performed on the obtained simulated image set;

[0057] The method for judging the radiometric degree of the image to be matched is: simulate the increase of the input images a and b respectively. times the affine amount, generate images A and B, and match them through SIFT feature extraction and description algorithm. The number of feature matches between image A and image B is M aB , the number of feature matches between image A and image b is M Ab , if M aB >M Ab , then image a is an image with a lower degree of affine, otherwise image b is an image with a lower degree of affine.

[0058] like Figure 2 As shown, plane u is the plane where the shooting scene is located, where θ is the angle between the optical axis and the normal vector of plane u, that is, the inclination angle of the optical axis. In this embodiment, it is considered that the angle θ corresponding to the input image has a certain initial value, so θ can be reduced for images with a high degree of affine to reduce the degree of image affine, and vice versa, θ can be increased to increase the degree of image affine, so that the simulated image set is closer to the purpose.

[0059] Reducing the affine degree difference between images is conducive to increasing the number of feature matches, and vice versa. Therefore, cross-feature matching can be performed by comparing the number of feature matches between the two input images to be matched and the simulated image with the amplified affine degree. The steps of reducing / increasing the affine degree of the simulated image include:

[0060] Step 1-1: For the image with a lower affine degree in the input image, the affine amount t of simulated enlargement is used 1 ,and For the image with a higher degree of affine in the input image, the affine amount t of simulated reduction is used. 2 ,and

[0061] Step 1-2: Both simulations for the tilt angle φ are from -45 degrees to 45 degrees, with a step size of 15 degrees.

[0062] In order to conveniently illustrate the advantages of the simulation strategy of this embodiment, equations (1) and (2) are introduced here.

[0063]

[0064] Among them, max_affine represents the simulated maximum affine amount, and the max() and min() functions correspond to the maximum and minimum affine amounts in the sampling set respectively.

[0065]

[0066] Among them, average_differ is the overall average affine amount, a means that the affine amount of the high affine image is a times that of the low affine image, and the count() function is used to calculate the number of elements in the sampling set.

[0067] Step 2: In the addition of descriptor region information, the two input images to be matched are enhanced in advance. The specific steps are as follows:

[0068] Step 2-1: Contrast Limited Adaptive Histogram Equalization (CLAHE), by dividing the image into several sub-block areas, using formula (3) to perform histogram equalization on a single area, and then intercepting the pixels in the histogram of a single sub-area that exceed the limit value and evenly distributing them to the pixels corresponding to each grayscale value level, as shown in formula (3).

[0069]

[0070] Among them, r is the input grayscale value, s is the output grayscale value, c x is the number of pixels with gray value x, P(x) is the distribution probability of gray value x, M and N are the length and width of the divided sub-block.

[0071] At this time, the new histogram can be used to obtain the corresponding new grayscale value using formula (4).

[0072]

[0073] Among them, new_c x is the number of pixels corresponding to the gray level x after redistribution, clip is the clipping face value, the number of pixels increased for each gray level is bin_incr=clip_sum / 256, and clip_sum is the total number of clipped pixels.

[0074] Step 2-2: At this time, the differences between the sub-blocks are obvious, and they need to be smoothed using bilinear interpolation. Divide the image into the same sub-blocks as above, and each sub-block is divided into 4×4 small areas. Use formula (5) to calculate the pixel values ​​in the corresponding area.

[0075]

[0076] The gray value f(x, y) is the gray value of the pixel to be sought, and the four vertices of the rectangular area where the pixel point (x, y) is located are denoted as A(x 1 ,y 1 )、B(x 2 ,y 1 )、C(x 2 ,y 1 )、D(x 2 ,y 2 ), the gray values ​​are T A , T B , T C , T D .

[0077] Step 2-3: After the image contrast is enhanced by the CLAHE algorithm, a bilateral filter is used to smooth the image area details. The bilateral filter increases the weight of the pixel gray value on the basis of Gaussian filtering, which reduces the area details while strengthening the retention of edge information, as shown in formula (6).

[0078]

[0079] Where h(x) is the output grayscale value, k(x) is the amount of weight normalization, x is the input pixel, S is the filter window corresponding to the x pixel, f(x) is the grayscale value of pixel x, c(ξ, x) is the spatial weight, s(f(ξ), f(x)) is the grayscale weight, σ d and σ r is the standard deviation of spatial proximity and grayscale similarity.

[0080] Step 3: Based on the MSER (Maximally Stable Extremal Regions) algorithm, the enhanced image in step 2 is segmented into regions, which specifically includes the following steps:

[0081] Step 3-1: The region segmentation of this embodiment uses the maximum stable extreme region algorithm MSER implemented by the Nister David improved watershed method. The watershed method compares the pixel grayscale value of the image to the altitude, sets a grayscale threshold as the horizontal plane, and regards the connected region boundary formed by the pixels under the grayscale threshold as the watershed. In the process of grayscale threshold from low to high, several connected regions can be obtained.

[0082] Step 3-2: Obtain the maximum stable extreme value region through the change rate constraint of the same region. The core operation is to start from a certain point in the image, establish a set of pixel points related to the gray threshold of the current point through a 4-neighborhood search, and perform corresponding operations on the point set according to the threshold change of the search point during the search process to obtain connected areas with different gray thresholds. Then, the maximum stable extreme value region is selected through formula (7) and the overlapping regions are merged.

[0083]

[0084] Among them, i is the gray value threshold, Q i is a connected area when the threshold is i, delta is a small change in the gray threshold, and q(i) is the area Q when the threshold is i i When the rate of change is less than the set maximum rate of change, the connected region is considered to be the maximum stable extreme value region.

[0085] Step 4: Use the grayscale histogram to describe the region after image segmentation in step 3, and then use the grayscale centroid method to rotate the main direction of the region to the positive direction of the x-axis based on the normalization of the region coordinates, and calculate the relative position of the feature point and the normalized centroid, including:

[0086] Step 4-1: Calculate the relative centroid position of the feature point. It is necessary to consider the feature point position offset caused by the affine transformation. Therefore, the coordinates of the region are normalized first. The region is regarded as the integral of several parallel line segments. Under coordinate normalization, the influence of the scale and affine amount on the coordinate value caused by the affine transformation can be greatly alleviated. Coordinate normalization is shown in formula (8).

[0087]

[0088] Among them, (x, y) is the input pixel coordinates, (norm_x, norm_y) is the normalized coordinates, min_x, min_y, max_x and max_y are the minimum and maximum x, y of all pixel coordinates in the area.

[0089] Step 4-2: When the same region undergoes affine transformation, the pixel grayscale becomes extremely small and the line segment ratio remains unchanged, then the grayscale centroid of the region will correspond to the same pixel point. The coordinates of the region centroid C and the main direction angle θ of the region are calculated by equation (9).

[0090] m pq =∑ x,y∈S x p y q I(x,y) p,q={0,1}

[0091]

[0092]

[0093] Step 4-3: Based on the above coordinate normalization and grayscale centroid method, the relative position of the feature point and the centroid is calculated. This relative position can greatly distinguish similar descriptors in the same area and ensure the invariance of the relative position value of the feature point in the affine transformation.

[0094] Step 5: Assign weights α to the histogram distribution probability of the area corresponding to the feature point and the relative centroid position of the feature point 1 =300 and α 2 =300, and add the two to the descriptor obtained in step 1 to obtain the new area information descriptor of this embodiment.

[0095] Based on the above specific implementation, the effect of the present invention is verified in combination with specific experiments below:

[0096] The proposed method is compared and evaluated with SIFT, ASIFT and SuperPoint methods on the Graffiti image set, Wall image set and three sets of simulated image sets. Figure 5 As shown in the figure, it can be seen from the two standard image sets of Graffiti and Wall that the accuracy of feature matching and subsequent homography matrix of the method of the present invention is higher than that of other methods, and it shows strong accuracy stability.

[0097] This example was completed in PyCharm Professional 2020.1 software, with the main dependent libraries being Opencv-python library version 4.5.3 and Numpy library version 1.19.2. The hardware environment was a laptop with a 3.20GHz i7 processor and 16GB of running memory, and the experimental process was relatively stable.

[0098] Some steps in the embodiments of the present invention may be implemented using software, and the corresponding software program may be stored in a readable storage medium, such as a CD or a hard disk.

[0099] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for constructing an image region feature descriptor, It is characterized in that The method comprises: Step 1: In the classified simulated affine transformation, the affine degree of the image to be matched is judged, and a relatively high / low degree of affine transformation is obtained, thereby respectively reducing / increasing the affine degree of the simulated image, and the existing SIFT feature point description is performed on the obtained simulated image set; Step 2: During the adding of the descriptor region information, image enhancement is performed on the image to be matched; Step 3: performing region segmentation on the image enhanced in step 2 based on the MSER algorithm; Step 4: Use the grayscale histogram to describe the segmented image area and obtain the area description information; On the basis of normalization of regional coordinates, the grayscale centroid method is used to rotate the main direction of the region to be consistent with the positive direction of the x-axis, and the relative position of the feature point and the normalized centroid is calculated; Step 5: assign weights to the region description information and the relative position information respectively, and fuse the region description information and the relative position information assigned with the weights with the feature point descriptor obtained in step 1.

2. The method for constructing an image region feature descriptor according to claim 1, It is characterized in that The process of determining the degree of affine in step 1 includes: Simulate adding the input images a and b to be matched respectively times the affine amount, generate images A and B respectively, and match them through SIFT feature extraction and description algorithm. The number of feature matches between image A and image B is M aB , the number of feature matches between image A and image b is M Ab , if M aB >M Ab , then image a is an image with a lower degree of affine, otherwise image b is an image with a lower degree of affine.

3. The method for constructing an image region feature descriptor according to claim 1, It is characterized in that The process of reducing / increasing the degree of affine of the simulated image in step 1 comprises: For the image with lower affine degree in the input image, the affine amount t of simulated enlargement is used 1 ,and For the image with a higher degree of affine in the input image, the affine amount t of simulated reduction is used. 2 ,and The tilt angle φ of both images is simulated from -45° to 45° with a step size of 15°.

4. The method for constructing an image region feature descriptor according to claim 1, It is characterized in that In step 2, after enhancing the image contrast using the CLAHE algorithm, a bilateral filter is used to smooth the image area details.

5. The method for constructing an image region feature descriptor according to claim 4, It is characterized in that In step 2, the bilateral filter increases the weight of the pixel gray value on the basis of Gaussian filtering, thereby reducing the regional details while strengthening the retention of edge information, as shown in the following formula: Where h(x) is the output grayscale value, k(x) is the amount of weight normalization, x is the input pixel, S is the filter window corresponding to the x pixel, f(x) is the grayscale value of pixel x, c(ξ,x) is the spatial weight, s(f(ξ),f(x)) is the grayscale weight, σ d and σ r is the standard deviation of spatial proximity and grayscale similarity.

6. The method for constructing an image region feature descriptor according to claim 1, It is characterized in that The step 4 comprises: Step 4-1: Normalize the coordinates of the region, and regard the region as the integral of several parallel line segments. The coordinate normalization formula is: Where (x, y) is the input pixel coordinate, (norm_x, norm_y) is the normalized coordinate, min_x, min_y, max_x and max_y are the minimum and maximum x and y coordinate values ​​of all pixel coordinates in the region; Step 4-2: When an affine transformation occurs in the same area, the grayscale of the pixel becomes extremely small and the line segment ratio remains unchanged. The grayscale centroid of the area will correspond to the same pixel point. The coordinates of the area centroid C and the main direction angle θ of the area can be calculated by the following formula: m pq =∑ x,y∈S x p y q I(x,y) p,q={0,1} Among them, m pq is the moment of region S, m 00 Represents the total grayscale of pixels in area S, m 10 is the grayscale moment of region S in the x direction, m 01 is the grayscale moment of region S in the y direction, and the influence of rotation on the coordinate value in affine transformation is eliminated by turning the positive direction of the x-axis of the region to the main direction; Step 4-3: Based on the coordinate normalization and the area centroid C, the relative position of the feature point and the centroid is calculated.

7. An image matching method, It is characterized in that Firstly, the method for constructing an image region feature descriptor according to any one of claims 1 to 6 is used to construct a region feature descriptor for an input image to be matched, and then feature matching is performed based on the constructed feature descriptor.

8. An image matching device, It is characterized in that include: A processor and a memory, wherein the memory stores instructions executed by the processor, and when the instructions are executed by the processor, the image matching device implements the image matching method according to claim 7. 9 . A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the image matching method according to claim 7 .

Citation Information

Patent Citations

  • Multimodal image feature extraction and matching method based on ASIFT (affine scale invariant feature transform)

    CN102231191A

  • Method for extracting feature points with invariable affine sizes

    CN103186899A