Octagonal subdomain scale invariant feature conversion image registration method based on FPGA
By using the octagonal subdomain SIFT algorithm on the FPGA platform, replacing the traditional square subdomain and increasing the number of feature directions, the problems of high computational complexity and high hardware resource requirements in remote sensing image registration are solved, and efficient and accurate image registration is achieved.
Patent Information
- Application Number
- CN202510114191.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-03
AI Technical Summary
Traditional SIFT algorithms have high computational complexity and high hardware resource requirements in remote sensing image registration, resulting in insufficient feature matching accuracy and speed.
Using the FPGA-based octagonal subdomain scale invariant feature conversion image registration method, 64-dimensional feature descriptors are generated by replacing 4 concentric regular octagons with the traditional 16 square subdomains and increasing the statistical direction in the feature subdomain from 8 to 16.
Fast and high-precision remote sensing agricultural image registration is achieved, improving the robustness and stability of the feature extraction process, and reducing the computing complexity and hardware resource requirements.
Smart Images

Figure CN120088299A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image registration, and particularly relates to an octagon sub-region scale-invariant feature transform image registration method based on FPGA. Background Art
[0002] During the growth cycle of rice, lodging will inevitably occur due to natural factors such as typhoons and heavy rains, or human factors such as violent trampling, which has a serious impact on the rice yield. Therefore, agricultural monitoring and early warning has become the primary issue to ensure stable rice yields.
[0003] The registration of remote sensing images is a very widely used technology in the field of image registration and occupies an important position in the field of image registration. Due to external natural weather factors such as clouds, lighting, and hurricanes, or various self-factors such as sensors, lenses, and shooting angles, the UAV equipment may generate blurred or multiple images that need to be stitched; these problems can be corrected using advanced image processing algorithms or during the image reconstruction process. Therefore, the main goal of remote sensing image registration technology is to accurately align the corresponding structures or features in multiple agricultural images so that the agricultural department can comprehensively utilize the image information from multiple angles to accurately monitor and judge agricultural disasters and conduct post-disaster rescue.
[0004] Currently, the commonly used remote sensing image registration methods are roughly divided into feature-based methods and region-based methods. Feature-based image registrations such as SIFT and SURF usually have good robustness to remote sensing images and can handle a certain degree of image changes; moreover, only a small number of feature points or feature descriptors need to be processed, and the calculation speed is relatively fast, which is suitable for real-time or near-real-time application scenarios; however, the extraction and matching of feature points may be affected by noise, local deformation, or missing, resulting in unstable corresponding relationships. As for region-based methods such as mutual information and cross-correlation, they can use the information of the entire image region for registration, can handle a large range of deformations or transformations; often can provide stable registration results and are not sensitive to local changes; but this registration method involves traversing all pixels in the image region, usually has a high computational complexity, and requires a longer processing time.
[0005] In recent years, with the increasing application of FPGA technology in the field of image registration, the SIFT image registration technology accelerated by FPGA has become one of the research hotspots. However, there are still many challenges in implementing SIFT for remote sensing image registration using FPGA. Firstly, although compared with methods such as mutual information and cross-correlation, SIFT image registration only needs to obtain a small number of feature points and generate feature descriptors for registration, and the amount of calculation has been reduced a lot. However, due to the multi-scale processing, feature point detection, descriptor calculation, and high-precision requirements of SIFT, there are still multiple image convolution operations, feature descriptor calculation, and storage bandwidth bottlenecks, etc., which all make the implementation of the SIFT algorithm on FPGA still pose great challenges. Secondly, FPGA is good at fixed-point operations, while many feature points and feature descriptors generated by traditional SIFT registration involve floating-point operations. However, the floating-point operation speed is slow, which affects the parallel computing advantage of FPGA; and it will affect the accuracy of image registration; at the same time, it also consumes more hardware resources. Therefore, in order to achieve efficient SIFT remote sensing agricultural image registration, the present invention optimizes the algorithm for generating feature descriptors.
[0006] The traditional SIFT algorithm generates feature descriptors by dividing the area around the feature point into 16 square sub-regions and calculating the gradients and modulus values at 8 equally spaced angles in each sub-region. These descriptors can effectively cover most of the features of the image, and the coordinate axes are rotated by finding the main direction, and finally a 128-dimensional feature vector is formed to describe the feature point. However, before generating the sub-region descriptors, the original SIFT algorithm needs to rotate the coordinates of the feature point neighborhood to the main direction and perform image interpolation operations, which poses relatively high resource requirements for hardware implementation. Summary of the Invention
[0007] Aiming at the technical problems that the traditional SIFT has high computational complexity and high hardware resource requirements when generating feature descriptors in remote sensing image registration, which may lead to insufficient feature matching accuracy and speed, this technical solution provides an octagon sub-domain scale-invariant feature transform image registration method based on FPGA. Based on FPGA technology, an optimized feature descriptor generation algorithm is proposed, the number of feature directions is changed, and an optimized algorithm for matching the combination of the number of feature directions is used to reduce the dimension of the feature descriptor; to achieve a fast and high-precision remote sensing agricultural image registration system; it can effectively solve the above problems.
[0008] The present invention is realized through the following technical solutions:
[0009] An octagon sub - domain scale - invariant feature transform image registration method based on FPGA. The image registration method is completed based on an image registration module. The image registration module consists of three sub - modules: a feature point detection module, a feature descriptor generation module, and a feature matching module. The feature point detection module detects extreme points. The feature descriptor generation module calculates the local region feature gradient through the coordinates of the extreme points. Calculating the local region feature gradient includes calculating the feature modulus and constructing a gradient histogram. And use the histogram to statistically calculate the main direction of the extreme points in this region. Then, replace the original 16 squares with 4 concentric regular octagons as the feature sub - domains of this region, and increase the statistical directions in the feature sub - domains from 8 to 16. Then, determine the feature main direction of each small sub - domain by judging the maximum values of the 16 directions in the 4 - layer sub - domains, and then rotate the coordinate axes to the main direction of the extreme points to achieve rotational invariance. Finally, generate a 64 - dimensional feature descriptor of 4 * 16, so that each extreme point realizes two characteristics of scale invariance and rotational invariance. Finally, the feature matching module matches the set of extreme points in the two images to complete the registration of remote sensing images.
[0010] Furthermore, in the calculation of the feature modulus, in order to make the descriptor rotation - invariant, it is necessary to use the local features of the image to assign a reference direction to each extreme point. The method of image gradient is used to obtain the stable direction of the local structure.
[0011] For the extreme points detected in the DOG pyramid, collect the gradient and direction distribution characteristics of the pixels within the 3σ neighborhood window of the Gaussian pyramid image where they are located. The gradient modulus is:
[0012]
[0013] The gradient direction is:
[0014]
[0015] In the above formula, L(x, y) is the scale - space value of different feature points. According to the suggestion of scholar Lowe, the gradient modulus m(x, y) is added according to the Gaussian distribution of σ = 1.5σ_oct, that is, the gradient modulus around the feature points at a specific scale in the image is weighted, m(x, y)=(m(x, y)*G(x, y,), 1.5σ_oct); here, σ_oct refers to the change in the standard deviation between adjacent two layers in each scale layer. The fixed ratio is 1.5.
[0016] Furthermore, in the construction of the gradient histogram, after the gradient calculation of the extreme points is completed, use the histogram to statistically calculate the gradient and direction of the pixels in the neighborhood. The gradient histogram divides the 0 - 360 - degree direction range into 36 bins, with each bin being 10 degrees.
[0017] The peak of the histogram represents the direction of the neighborhood gradient at the extreme point, and the maximum value in the histogram is taken as the main direction of the extreme point; to enhance the robustness of the matching, only the directions with peaks greater than 80% of the peak of the main direction are retained as the secondary directions of the extreme point;
[0018] To prevent a sudden change in the angle of a certain gradient direction due to noise interference, it is necessary to smooth the gradient direction histogram, and the smoothing formula is:
[0019]
[0020] where \(i\in[0,35]\) is the angle index, and \(h\) and \(H\) represent the histogram before and after smoothing respectively.
[0021] Furthermore, the operation method of generating a 64-dimensional feature descriptor of \(4\times16\) is as follows: after completing the gradient magnitude statistics, the SIFT algorithm obtains three pieces of information for each key point: the coordinates of the extreme point, the image scale, and the main direction of the extreme point. According to these three pieces of information, a feature descriptor can be generated so that it is not affected by changes in illumination and viewing angle;
[0022] Rotate the coordinate axis to the main direction of the feature point shown in the histogram, and the coordinate axis rotation formula is:
[0023]
[0024] In the above formula is the coordinate of the feature point, is the rotated coordinate, \(\theta\) is the rotation angle and also the rotation direction.
[0025] Furthermore, the side lengths of the 4 octagons are set to 9, 8, 5, and 2 respectively; the interval angles of the 16 statistical directions in the feature sub-region are the same, which is \(22.5^{\circ}\).
[0026] Furthermore, the operation method of the feature point detection module for detecting extreme points is as follows: input the reference remote sensing image data and the remote sensing image data to be registered into the feature point detection module respectively. This module first constructs a difference-of-Gaussians pyramid. For each pixel point in each layer, with this pixel as the center, it is compared with 8 adjacent pixels of the same scale and 9 pixels of the adjacent upper and lower scales, a total of 26 points. An extreme point is detected in both the scale space and the two-dimensional image space; multiple extreme points in multiple spaces are obtained by using the above method.
[0027] Furthermore, the operation method of the feature point detection module for detecting extreme points includes the following steps:
[0028] Step 1: Construct a Gaussian pyramid:
[0029] The Gaussian pyramid of an image is obtained by blurring and downsampling the image using the Gaussian function. The calculation formula for the number of pyramid groups is:
[0030] O = [logmin(M, N)] - 2 (1)
[0031] where M and N are the number of rows and columns of the original image;
[0032] Each image in the pyramid is represented by L(x, y, σ), and its calculation formula is:
[0033]
[0034]
[0035] where I(x, y) represents the image, G(x, y, σ) represents the Gaussian function, represents convolution, e is a constant, called the base of the natural logarithm, and its value is approximately 2.71828; σ represents the image scale parameter, that is, the blur coefficient. The calculation formula for the Gaussian blur coefficient is as follows:
[0036]
[0037] where o is the group index number, r is the layer index number, s is the number of layers in each group of the pyramid, and σ o is the initial value of Gaussian blur. According to the suggestion of scholar Lowe, it is set to 1.6, but the actual initial value can be set to 1.52;
[0038] The convolution of the Gaussian function with scale σ and the image is used as the first layer. The scale of each upper layer image is k times that of the lower layer, that is, the nth layer image is the convolution of the Gaussian function with scale K n-1 σ and the image;
[0039] Step 2: Construct the Difference of Gaussian (DOG) pyramid:
[0040] The DOG pyramid is operated on the basis of the Gaussian pyramid. Its construction process is: in each group of the Gaussian pyramid, subtract the adjacent two layers, subtract the upper layer from the lower layer to generate the Gaussian difference pyramid; the expression of each image in the generated Gaussian difference image pyramid is:
[0041] D(x, y, σ) = L(x, y, kσ) - L(x, y, σ) (5)
[0042] where D(x, y, σ) is the difference image, and L(x, y, kσ) and L(x, y, σ) represent images of different scales respectively;
[0043] Step 3: Detect spatial extreme points:
[0044] The preliminary exploration of the extreme points is completed by comparing the adjacent two layers of images of each DOG within the same group; to find the extreme points of the DOG function, each pixel point needs to be compared with all its adjacent points to see if it is larger or smaller than its adjacent points in the image domain and scale domain; the middle pixel point is compared with 8 adjacent points of the same scale and 9×2 points corresponding to the upper and lower adjacent scales, a total of 26 points. If the currently detected pixel point is the maximum or minimum among them, it is retained as an extreme point, otherwise, it is not an extreme point and the detection of other points is carried out.
[0045] Furthermore, the specific operation steps for the feature matching module to match the extreme point sets in the two images include:
[0046] Euclidean distance is used to achieve feature matching. Key point descriptor sets are established for the template image and the real-time image respectively, and the recognition of the target is completed by comparing the key point descriptors within the two point sets;
[0047] The formula for the traditional SIFT algorithm to achieve feature matching using Euclidean distance is:
[0048]
[0049] where, R i and S i respectively represent the descriptors of two feature points, that is, two 128-dimensional feature vectors, here they are 64-dimensional; r ij and s ij respectively represent the values of the j-th dimension in the descriptors R i and S i .
[0050] Beneficial effects
[0051] An octagon sub-domain scale-invariant feature transform image registration method based on FPGA proposed by the present invention has the following beneficial effects compared with the prior art:
[0052] (1) Based on FPGA technology, the present invention proposes an optimized algorithm for generating feature descriptors, changes the number of feature directions, and improves the matching with the optimized algorithm for the number of feature direction combinations, reducing the dimension of the feature descriptor; to achieve a fast and high-precision remote sensing agricultural image registration system; in the original SIFT algorithm, the original 4×4 square feature sub-domain is replaced by 4 concentric regular octagons, the side lengths of the four concentric octagons are set to 9, 8, 5, and 2 respectively, and the statistical direction of the feature points within each sub-domain is increased from 8 to 16. The design of the 4 concentric octagon regions, with 16 directions set in each octagon region, makes the dimension of the feature descriptor 4×16 = 64 dimensions. In the new feature region, the rotation direction is calculated to achieve rotation invariance.
[0053] (2) The present invention uses four concentric octagons to replace the traditional 16 square sub-domains, which helps to improve the robustness of the feature extraction process under transformations such as image rotation and scale change. Octagons have better symmetry than squares during image rotation, such as the issue of the size of the rotation radius. And they can better adapt to the morphological and size changes of local features during scale change, thereby enhancing the stability and distinctiveness of the descriptor. This optimization enables the image feature extraction to maintain consistency and reliability when facing different perspectives and scale changes.
[0054] (3) The present invention increases the traditional 8 feature directions to 16 feature directions, improving the ability to capture the detailed direction distribution of local image regions. This enables OS-SIFT to more accurately reflect the direction features of key points in the image, enhancing the stability and accuracy of image matching. Description of the Drawings
[0055] Figure 1 It is the overall framework diagram of the OS-SIFT registration system in Embodiment 1.
[0056] Figure 2 It is the flowchart of feature point detection of the present invention.
[0057] Figure 3 It is the flowchart of generating the feature descriptor of the present invention.
[0058] Figure 4 It is the flowchart of feature matching of the present invention.
[0059] Figure 5 It is the schematic diagram of the Gaussian pyramid structure of the present invention.
[0060] Figure 6 It is the schematic diagram of the Difference-of-Gaussians pyramid structure of the present invention.
[0061] Figure 7 It is the schematic diagram of detecting spatial extreme points of the present invention.
[0062] Figure 8 It is the schematic diagram of the gradient direction of feature points in the scale image of the present invention.
[0063] Figure 9 It is the schematic diagram of the gradient direction histogram of the present invention.
[0064] Figure 10 It is the schematic diagram of the radius change caused by rotating the picture in the traditional SIFT algorithm.
[0065] Figure 11 It is the schematic diagram of rotating the coordinate axes to the main direction by rotating the picture in the traditional SIFT algorithm.
[0066] Figure 12Schematic diagram of the 16-square sub-region of the traditional SIFT algorithm.
[0067] Figure 13 Schematic diagram of the 8-direction feature gradients of the traditional SIFT algorithm.
[0068] Figure 14 Schematic diagram of the 4 concentric regular octagon sub-regions in the improved SIFT algorithm of the present invention.
[0069] Figure 15 Schematic diagram of the 16-direction feature gradients in the improved SIFT algorithm of the present invention.
[0070] Figure 16 Scale effect diagram after the hardware in the present invention completes the matching of the picture.
[0071] Figure 17 Scale effect diagram after the software matching of the original SIFT algorithm.
[0072] Figure 18 Rotation effect diagram after the hardware in the present invention completes the matching of the picture.
[0073] Figure 19 Rotation effect diagram after the software matching of the original SIFT algorithm. Detailed implementation manners
[0074] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Without departing from the design concept of the present invention, various variations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope of the present invention.
[0075] Embodiment 1:
[0076] An octagon sub-region scale-invariant feature transform image registration method based on FPGA. The image registration method is completed based on an image registration module. The image registration module is set in the FPGA image registration system, and the system architecture is as Figure 1 shown, including a data reception module, a preprocessing module, a data storage module (DDR3), an image registration module, and an HDMI output module. When the entire system receives image data, it will first perform preprocessing operations such as grayscale processing and noise suppression, and transmit the preprocessed two pieces of image data to the image registration module through DDR3. This system can effectively reduce the amount of computation, improve the registration speed and accuracy, and thus achieve more efficient image registration processing.
[0077] The described image registration module consists of three sub-modules, including a feature point detection module, a feature descriptor generation module, and a feature matching module. The specific process of each sub-module is as follows:
[0078] Feature point detection module: The operation method is as Figure 2 shown. Specifically: The reference remote sensing image data and the remote sensing image data to be registered are respectively input into the feature point detection module. This module first constructs a Difference of Gaussian pyramid (DoG pyramid), and then for each pixel point in each layer, with this pixel as the center, it is compared with 8 adjacent pixels of the same scale and 9 pixels of the adjacent upper and lower scales (a total of 18 points in the upper and lower two layers), and a total of 26 points are compared. To ensure that an extreme point (feature point) is detected in both the scale space and the two-dimensional image space; in this way, we obtain many spatial extreme points (feature points).
[0079] Feature descriptor generation module: The operation method is as Figure 3 shown. Specifically: In the feature descriptor generation module, first calculate the local region feature gradient through the above extreme point (feature point) coordinates, that is, the feature modulus and the feature direction, and use a histogram to statistically calculate the main direction of the feature points in this region. The peak direction of the histogram represents the main direction of the feature point, and the direction greater than 80% of the main direction peak is used as the secondary direction of the feature point.
[0080] Then use four concentric octagons to replace the original 16 squares as the feature sub-region of this region. The lengths of the octagons are set to 9, 8, 5, and 2 respectively; and the statistical directions in the feature sub-region are increased from 8 to 16 (one direction every 22.5°). Then determine the main feature direction of each small sub-region by judging the maximum value of the 16 directions in the 4-layer sub-region, and then rotate the coordinate axis to the main direction of the feature point to achieve rotation invariance; finally, generate a 64-dimensional feature descriptor of 4 * 16.
[0081] So far, each extreme point (feature point) realizes two characteristics: scale invariance and rotation invariance.
[0082] Feature matching module: Next, match the feature point sets of the two images to complete the registration of the remote sensing images.
[0083] The above is the complete process of the entire OS-SIFT (Octagonal Subregion-SIFT) remote sensing image registration system. The specific situation is as Figure 4 shown:
[0084] Step 1: Use the feature point detection module to detect feature points;
[0085] Step 1.1: Construct a Gaussian pyramid;
[0086] The Gaussian Pyramid of an image is obtained by blurring and downsampling the image using a Gaussian function. In this embodiment, taking 1 group of 4 layers as an example, specifically as Figure 5 shown.
[0087] Calculation formula for the number of pyramid groups:
[0088] O = [log min(M, N)] - 2 (1)
[0089] where M and N are the number of rows and columns of the original image respectively.
[0090] Each image in the pyramid is represented by L(x, y, σ), and its calculation formula is:
[0091]
[0092]
[0093] where I(x, y) represents the image, and G(x, y, σ) represents the Gaussian function; represents convolution; e is a constant, called the base of the natural logarithm, and its value is approximately 2.71828; σ represents the image scale parameter, that is, the blur coefficient, and the calculation formula for the Gaussian blur coefficient is as follows:
[0094]
[0095] where o is the group index number, r is the layer index number, s is the number of layers in each group of the pyramid, and σ o is the initial value of Gaussian blur (set to 1.6 according to the suggestion of scholar Lowe, but the actual initial value can be set to 1.52).
[0096] Taking the convolution of the Gaussian function with scale σ and the image as the first layer, the scale of the image corresponding to each upper layer is k times that of the lower layer. That is, the nth layer image is the convolution of the Gaussian function with scale K n-1 σ and the image.
[0097] Step 1.2: Construct the Difference of Gaussian (DOG) pyramid;
[0098] The Difference of Gaussian pyramid is operated on the basis of the Gaussian pyramid. Its establishment process is: subtracting adjacent two layers in each group of the Gaussian pyramid (subtracting the upper layer from the lower layer) generates the Difference of Gaussian pyramid. Specifically as Figure 6 shown.
[0099] The expression for each image in the generated Difference of Gaussian image pyramid (DOG scale space) is:
[0100] D(x, y, σ) = L(x, y, kσ) - L(x, y, σ) (5)
[0101] Among them, D(x, y, σ) is the difference image, and L(x, y, kσ) and L(x, y, σ) represent images of different scales respectively.
[0102] Step 1.3: Detect spatial extreme points;
[0103] The preliminary exploration of extreme points is completed by comparing between adjacent two layers of images within the same group of DOG. To find the extreme points of the DOG function, each pixel point needs to be compared with all its adjacent points to see if it is larger or smaller than its adjacent points in the image domain and scale domain. As Figure 7 shown, the middle pixel point is compared with 8 adjacent points of the same scale and 9×2 points corresponding to the upper and lower adjacent scales, a total of 26 points. If the currently detected pixel point is the maximum or minimum among them, it is retained as an extreme point; otherwise, it is not an extreme point, and other points are detected.
[0104] Step Two: Use the generated feature descriptor module to generate feature descriptors;
[0105] Step 2.1: Calculate the gradient value of the feature point:
[0106] To make the descriptor rotation invariant, it is necessary to use the local features of the image to assign a reference direction to each feature point. Therefore, the method of image gradient is used to obtain the stable direction of the local structure.
[0107] For the feature points detected in the DOG pyramid, collect the gradient and direction distribution features of the pixels within the 3σ neighborhood window of the Gaussian pyramid image where they are located.
[0108] Gradient magnitude:
[0109]
[0110] Gradient direction:
[0111]
[0112] L(x, y) is the scale space value where different feature points are located. According to the suggestion of scholar Lowe, the gradient magnitude m(x, y) is added according to the Gaussian distribution of σ = 1.5σ_oct, that is, the gradient magnitude around the feature points at a specific scale in the image is weighted, m(x, y) = (m(x, y) * G, y,), 1.5σ_oct). σ_oct in this implementation refers to the change amount of the standard deviation between adjacent two layers in each scale layer (octave). The fixed ratio is 1.5.
[0113] Step 2.2: Construct a gradient histogram:
[0114] After the gradient calculation of the feature points is completed, the gradient and direction of the pixels in the neighborhood are statistically analyzed using a histogram. The gradient histogram divides the 0 to 360-degree direction range into 36 bins, with each bin covering 10 degrees. Specifically, as shown in Figure 8 and Figure 9 (For simplicity, only 8 directions are drawn in the figure).
[0115] The peak of the direction histogram represents the direction of the neighborhood gradient at that feature point, and the maximum value in the histogram is taken as the main direction of the key point. To enhance the robustness of the matching, only the directions with a peak greater than 80% of the main direction peak are retained as the secondary directions of the key point.
[0116] To prevent a sudden change in a certain gradient direction angle due to noise interference, we also need to smooth the gradient direction histogram.
[0117]
[0118] where \(i\in[0,35]\) is the angle index, and \(h\) and \(H\) represent the histogram before and after smoothing respectively.
[0119] Step 2.3: Generate feature descriptors:
[0120] After the gradient magnitude statistics, the SIFT algorithm obtains three pieces of information for each key point: the key point coordinates, the image scale, and the main direction of the key point. At this time, the feature descriptors can be generated to make them invariant to various changes (such as illumination changes and perspective changes).
[0121] The traditional SIFT algorithm divides the neighborhood of the feature point into a \(d\times d\) square region (according to Professor Lowe's suggestion, \(d\) is set to 4), and only 8 directions are statistically analyzed in each sub-region. The side length of the image window in each region is \(3\times3\times\sigma\); but considering the rotation change, the actual radius is:
[0122]
[0123] The specific rotation-induced radius change is as shown in Figure 10 When the rotation direction points to the four sides of the square, each square sub-region uses half of the square side length \(r\) as the rotation radius. However, when the rotation direction points from the four sides of the square to the 90° inner angle of the square, the rotation radius changes from \(r\) to
[0124] At this time, the coordinate axes need to be rotated to the main direction of the feature point shown in the histogram to ensure rotation invariance. Specifically, as shown in Figure 11 In the example, the main direction in the square sub-region points to the 45° direction in the third quadrant. At this time, with the Y-axis as the reference, rotating 225° clockwise can turn the coordinate axes to the main direction of the feature point, successfully achieving rotation invariance.
[0125] The coordinate rotation formula is as follows:
[0126]
[0127] In the above formula is the coordinate of the feature point, is the rotated coordinate, θ is the rotation angle and also the rotation direction. After the rotation in the main direction is completed, the quantities in 8 directions within a 4×4 square sub-region can be counted, each direction being 45°, and a gradient histogram is generated, thus generating a 128-dimensional SIFT feature descriptor of 4*4*8. Specifically, as shown in Figure 12 and 13 For example: within a 4*4 16-square sub-region, each small square sub-region has a certain number of feature directions. The original algorithm would count the total number of 8 feature directions within the sub-region and then select the maximum value as the main direction for convenient coordinate axis rotation output.
[0128] However, this method has too high requirements for hardware. Considering the convenience of use, this embodiment proposes to use 4 concentric octagons to replace the original 16 square sub-regions, and increase the gradient direction of the feature descriptor from 8 to 16. Then the dimension of the new descriptor is 64, which not only reduces the computational complexity but also retains sufficient feature information, helping to optimize the performance in hardware implementation.
[0129] Specifically, as shown in Figure 14 and 15 : First, select 4 regular octagons of different sizes. After they are mutually contained, set them as the feature sub-regions. Considering that there will also be a certain number of overlapping feature directions within the 4 layers of regular octagon sub-regions of different sizes, after each time the feature directions within the smaller octagon are counted, the larger octagon only needs to count the non-overlapping regions of the two.
[0130] In order to obtain more accurate details of the sub-region direction, the feature direction is increased from 8 directions to 16 directions, and it is counted once every 22.5°. From the perspective of the clock, the extreme value points (feature points) of the main direction among the 16 feature directions can be judged in 4 clocks. That is, the feature descriptor finally generated for each extreme value point (feature point) is 4*16, a 64-dimensional descriptor, and the amount of calculated data is half less than that of the original algorithm.
[0131] Step 3: Adopt a feature matching module to perform feature matching on the reference image and the registered image;
[0132] Image registration is essentially to detect the similarity of the image feature descriptors. When a large number of regions with the same descriptors are detected in two images, the possibility of successful matching of that region is greater. The traditional SIFT algorithm uses the Euclidean distance to achieve feature matching, and key point descriptor sets are established for the template image and the real-time image respectively. The recognition of the target is completed by comparing the key point descriptors within the two point sets.
[0133] The traditional SIFT algorithm uses the Euclidean distance to achieve feature matching, and the specific formula is as follows:
[0134]
[0135] Where R i and S i respectively represent the descriptors of two feature points (that is, two 128-dimensional feature vectors, which are 64-dimensional in this design); r ij and s ij respectively represent the values of the j-th dimension in the descriptors R i and S i During the matching process, the smaller the distance between descriptors, the more similar they are. Finally, the registration of two remote sensing images is achieved and the result is output through the HDMI interface.
[0136] The specific implementation effect is as Figures 16 to 19 shown, and the test image is a remote sensing agricultural image. Figure 17 and Figure 19 are the software matching effect diagrams of the original SIFT algorithm. In the diagrams, only 3 feature points are matched in both the scale invariance test and the rotation invariance test, and the matching positions of 2 of the feature points are repeated, with the correct rate being only 66.6%. Figure 16 and Figure 18 are the effect diagrams after hardware matching after the optimization of this embodiment. In the diagrams, at least more than 18 feature points are matched in both the scale invariance test and the rotation invariance test, and the correct rate reaches 94.5%; scale invariance and rotation invariance can be achieved.
[0137] Through the above comparison, the conclusion can be drawn that the SIFT algorithm of the octagonal subdomain implemented based on hardware (FPGA) is superior to the original SIFT algorithm implemented by software (MATLAB) in terms of both the number of feature point matches and the feature point matching accuracy when applied in the field of remote sensing. This also indicates that the present invention is more suitable for application in real life.
[0138] The above embodiments are only used to illustrate the technical concept and features of the present invention, and their purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly, and cannot be used to limit the protection scope of the present invention. Any equivalent transformation or modification made according to the spirit of the present invention should be covered by the present invention.
Claims
1. An octagonal subdomain scale-invariant feature conversion image registration method based on FPGA, wherein the image registration method is completed based on an image registration module; characterized in that: The image registration module is composed of three submodules: a feature point detection module, a feature description generation submodule and a feature matching module; the feature point detection module detects extreme points, and the feature description generation submodule calculates local area feature gradients through extreme point coordinates, and the calculation of local area feature gradients includes calculating feature modulus values and constructing gradient histograms; and the histogram is used to count the main directions of the extreme points in the area; then 4 concentric regular octagons replace the original 16 squares as the feature subdomain of the area, and the statistical directions in the feature subdomain are increased from 8 to 16; then the feature main direction of each small subdomain is determined by judging the maximum value of 16 directions in 4 layers of subdomains, and then the coordinate axis is rotated to the main direction of the extreme point to achieve rotation invariance; finally, a 4*16 64-dimensional feature descriptor is generated, so that each extreme point achieves two characteristics of scale invariance and rotation invariance; finally, the extreme point sets in the two images are matched by the feature matching module to complete the registration of remote sensing images.
2. The FPGA-based octagonal subdomain scale-invariant feature conversion image registration method according to claim 1, characterized in that: In the calculation of the characteristic modulus value, in order to make the descriptor rotationally invariant, it is necessary to use the local features of the image to assign a reference direction to each extreme point; the image gradient method is used to obtain the stable direction of the local structure; The extreme points detected in the DOG pyramid are used to collect the gradient and directional distribution characteristics of the pixels in the 3σ domain window of the Gaussian pyramid image; the gradient modulus is: The gradient direction is: In the above formula, L(x,y) is the scale space value of different feature points. According to the suggestion of scholar Lowe, the gradient modulus m(x,y) is added according to the Gaussian distribution of σ=1.5σ_oct, that is, the gradient modulus around the feature points at a specific scale in the image is weighted m(x,y)=(m(x,y)*G(x,y,0,1.5σ_oct); here σ_oct refers to the change in the standard deviation between two adjacent layers in each scale layer; the fixed ratio is 1.
5.
3. The FPGA-based octagonal subdomain scale-invariant feature conversion image registration method according to claim 1, characterized in that: In the construction of the gradient histogram, after completing the gradient calculation of the extreme point, the gradient and direction of the pixels in the field are counted using the histogram; the gradient histogram divides the direction range of 0 to 360 degrees into 36 bins, each of which is 10 degrees; The peak of the histogram represents the direction of the neighborhood gradient at the extreme point, and the maximum value in the histogram is taken as the main direction of the extreme point. In order to enhance the robustness of the matching, only the direction with a peak value greater than 80% of the peak value of the main direction is retained as the auxiliary direction of the extreme point. In order to prevent a certain gradient direction angle from changing suddenly due to noise interference, the gradient direction histogram needs to be smoothed. The smoothing formula is: Among them, i∈[0,35] is the angle index, h and H represent the histogram before and after smoothing, respectively.
4. The FPGA-based octagonal subdomain scale-invariant feature conversion image registration method according to claim 3, characterized in that: The operation method of generating a 4*16 64-dimensional feature descriptor is as follows: after completing the gradient amplitude statistics, the SIFT algorithm obtains three information for each key point: the extreme point coordinates, the image scale, and the extreme point main direction. Based on these three information, a feature descriptor can be generated so that it is not affected by changes in illumination and viewing angle.
5. The FPGA-based octagonal subdomain scale-invariant feature transformation image registration method according to claim 1 or 3, characterized in that: The side lengths of the four octagons are set to 9, 8, 5, and 2 respectively; the interval angles of the 16 statistical directions in the characteristic subdomain are the same, which is 22.5°.
6. The FPGA-based octagonal subdomain scale-invariant feature conversion image registration method according to claim 1, characterized in that: The operation mode of the feature point detection module to detect the extreme point is as follows: the reference remote sensing image data and the remote sensing image data to be registered are respectively input into the feature point detection module, and this module first constructs a Gaussian difference pyramid. For each pixel point in each layer, the pixel is taken as the center, and the 8 adjacent pixels of the same scale and the 9 pixels of the upper and lower adjacent scales are compared, a total of 26 points, and an extreme point is detected in both the scale space and the two-dimensional image space; the extreme points of multiple spaces are obtained by using the above method.
7. The FPGA-based octagonal subdomain scale-invariant feature conversion image registration method according to claim 1 or 6, characterized in that: The operation method of the feature point detection module to detect and obtain extreme points includes the following steps: Step 1: Construct a Gaussian pyramid: The image Gaussian pyramid is obtained by blurring and downsampling the image using a Gaussian function. The formula for calculating the number of pyramid groups is: O=[logmin(M,N)]-2 (1) Where M and N are the number of rows and columns of the original image; Each image in the pyramid is represented by L(x,y,σ), which is calculated as: Among them, I(x,y) represents the image, G(x,y,σ) represents the Gaussian function, represents convolution, e is a constant, called the base of the natural logarithm, and its value is approximately 2.71828; σ represents the image scale parameter, that is, the blur coefficient, where the Gaussian blur coefficient calculation formula is as follows: Among them, o is the group index number, r is the layer index number, s is the number of layers in each group of pyramid, σ o is the initial value of Gaussian blur, which is set to 1.6 according to the suggestion of scholar Lowe, but the actual initial value can be set to 1.52; The convolution of the Gaussian function with a scale of σ and the image is used as the first layer. The scale of each previous layer of images is k times that of the next layer, that is, the nth layer of images is composed of images with a scale of K. n-1 Convolution of the Gaussian function of σ with the image; Step 2: Construct a difference pyramid DOG: The difference pyramid is based on the Gaussian pyramid. The establishment process is: in each group of the Gaussian pyramid, two adjacent layers are subtracted, and the next layer is subtracted from the previous layer to generate a Gaussian difference pyramid; the expression for generating each image in the Gaussian difference image pyramid is: D(x,y,σ)=L(x,y,kσ)-L(x,y,σ) (5) Among them, D(x,y,σ) is the difference image, L(x,y,kσ) and L(x,y,σ) represent images of different scales respectively; Step 3: Detect spatial extreme points: The preliminary detection of extreme points is completed by comparing the two adjacent layers of images of each DOG in the same group; in order to find the extreme points of the DOG function, each pixel point is compared with all its adjacent points to see whether it is larger or smaller than its adjacent points in the image domain and scale domain; the middle pixel point is compared with its 8 adjacent points of the same scale and 9×2 points corresponding to the upper and lower adjacent scales, a total of 26 points. If the currently detected pixel point is the maximum or minimum value among them, it is retained as an extreme point, otherwise, it is not an extreme point and other points are detected.
8. The FPGA-based octagonal subdomain scale-invariant feature conversion image registration method according to claim 1, characterized in that: The specific operation steps of the feature matching module for matching the extreme point sets in the two images include: The Euclidean distance is used to achieve feature matching. Key point descriptor sets are established for the template image and the real-time image respectively. The target is identified by comparing the key point descriptors in the two point sets. The traditional SIFT algorithm uses the Euclidean distance to achieve feature matching formula: Among them, R i and S i Respectively represent the descriptors of two feature points, that is, two 128-dimensional feature vectors, here 64-dimensional; r ij and ij Respectively represent the descriptor R i and S i The value of the j-th dimension in .