Initial value search algorithm for SIFT feature point matching based on GPU acceleration
The GPU-accelerated SIFT feature point matching algorithm optimizes the strain measurement of the synchronous automatic clutch, solving the real-time measurement problem of the core synchronization mechanism of the synchronous automatic clutch, improving the measurement speed and accuracy, extending the service life and improving reliability.
Patent Information
- Application Number
- CN202510522077.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-04-24
AI Technical Summary
In the prior art, the core synchronization mechanism of the synchronous automatic clutch takes a long time in strain measurement, resulting in real-time full-field strain measurement, which affects its operating efficiency and reliability.
The SIFT feature point matching method based on GPU acceleration is adopted to optimize the initial value search process through parallel calculations, including allocating and aligning the image memory, building Gaussian pyramids and differential pyramids, computing and matching of feature point descriptors, and solving the initial matching value using the least squares method.
Significantly improves the speed and accuracy of the synchronous automatic clutch synchronous mechanism strain measurement, meets real-time measurement needs, optimizes structure and materials, extends service life, and predicts damage locations for improved reliability and safety.
Smart Images

Figure CN120495700A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and image processing, and in particular relates to an initial value search algorithm for SIFT feature point matching based on GPU acceleration. Background Art
[0002] The synchronous automatic clutch is a core component of the combined power transmission system. The stability and safety of its core synchronization mechanism directly impacts its operating efficiency and lifespan. As a critical component, the core synchronization mechanism is subjected to high-frequency alternating loads for a long time, making it prone to deformation and fracture. This poses a challenge to the safety and reliability of the synchronous automatic clutch. To gain a deeper understanding of the deformation of the core synchronization mechanism during operation, the three-dimensional digital image correlation method, a feature of binocular vision technology, can be used to obtain accurate strain data of the core synchronization mechanism during operation. This data can help optimize the structure and materials of the core synchronization mechanism, improve its load-bearing capacity and fatigue resistance, and thus extend its service life. Furthermore, the strain response and fracture process of the core synchronization mechanism under stable resonant excitation can be studied to predict possible damage locations and develop effective maintenance strategies, further improving the reliability and safety of the synchronous automatic clutch.
[0003] Although the digital image correlation method can obtain the full-field strain, it is impossible to realize real-time full-field strain measurement in the working structure of the synchronous mechanism of the synchronous automatic clutch due to the large amount of data, the cumbersome matching calculation process and the long time consumption. Among them, the SIFT corner point search and matching algorithm used in the initial value search is time-consuming. In order to solve the limitations of the traditional method in measurement speed and efficiency, the present invention adopts a SIFT feature point matching method based on GPU acceleration, optimizes the initial value search process through parallel computing, and significantly improves the measurement efficiency. Specifically, GPU has outstanding performance in the acceleration of three-dimensional digital image correlation algorithms, especially in the SIFT corner point search algorithm, which has injected new vitality into the innovation of high-frequency fatigue testing technology and provided more reliable data support for its design, manufacturing and operation. In summary, the initial value search algorithm based on GPU-accelerated SIFT feature point matching not only improves the measurement speed of the digital image correlation method, but also provides a new path and possibility for improving the performance and safety of the synchronous automatic clutch. Summary of the Invention
[0004] The purpose of this section is to summarize some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the abstract and title of this application to avoid obscuring the purpose of this section, the abstract and the title of the invention, and such simplifications or omissions should not be used to limit the scope of the present invention.
[0005] In view of the above or existing problems of the initial value search algorithm for SIFT feature point matching based on GPU acceleration, the present invention is proposed.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] An embodiment of the present invention provides an initial value search algorithm for SIFT feature point matching based on GPU acceleration, including: allocating and aligning image memory; constructing a multi-scale space of adjacent multi-frame Gaussian pyramids based on the GPU to obtain a complete Gaussian pyramid and a Gaussian difference pyramid, and performing GPU-accelerated calculation of Gaussian blur; calculating feature point descriptors; matching SIFT feature points; and constructing a control point matching equation based on the calculation of SIFT feature points.
[0008] As a preferred solution of the initial value search algorithm for SIFT feature point matching based on GPU acceleration described in the present invention, the memory allocation and alignment of the image includes: allocating memory of width w × height h for the image, and the memory size is pitch × h × size, wherein pitch is the smallest multiple of 128 greater than w, and size is the number of data bytes for each pixel, ensuring that the starting address of each row of data meets the memory alignment requirements of the GPU.
[0009] As a preferred solution of the initial value search algorithm for SIFT feature point matching based on GPU acceleration described in the present invention, wherein: the GPU-based Gaussian pyramid construction includes:
[0010] The Gaussian kernel function is stored in the constant memory of the GPU in the form of an array. Gaussian blur is performed on each group of images. A two-dimensional Gaussian kernel function is used to convolve with the lowest layer image. The convolution kernel size is (2R+1)×(2R+1). R is the Gaussian blur convolution kernel window radius. The Gaussian blur convolution kernel window radius is set to 9. The two-dimensional Gaussian blur convolution kernel is shown below:
[0011]
[0012] According to the separability of the two-dimensional Gaussian function, the two-dimensional Gaussian blur convolution kernel is expressed as the product of two one-dimensional Gaussian functions, as shown in the following formula:
[0013]
[0014] Where x and y are the horizontal and vertical position coordinates, σ is the standard deviation of the Gaussian distribution, and S is the normalization constant of the two-dimensional Gaussian function. and are the one-dimensional Gaussian functions along the horizontal and vertical directions respectively, S´ is the normalization constant of the one-dimensional Gaussian function;
[0015] After reducing the dimensionality of the two-dimensional Gaussian blur convolution kernel, for images with a resolution of w×h, the time complexity of the algorithm is reduced from O(R×R×w×h) to O(R×R×w)+O(R×R×h), reducing the amount of computation.
[0016] As a preferred solution of the initial value search algorithm for SIFT feature point matching based on GPU acceleration described in the present invention, the GPU accelerated calculation of Gaussian blur includes: allocating GPU threads, in the GPU architecture, a thread bundle is a group of 32 threads executed in parallel; setting each thread block to have 32×32 threads, and each thread block is responsible for calculating the final convolution value of a 32×24 area.
[0017] As a preferred solution of the initial value search algorithm for SIFT feature point matching based on GPU acceleration described in the present invention, the GPU accelerated calculation of Gaussian blur also includes: performing convolution calculation on GPU threads, each thread first calculates a row convolution, and for a thread block, a row convolution calculation result of a 32×32 area is obtained; the calculation result is stored in the shared memory of the thread block, and then each thread located in the middle area of 32×24 of the thread block calculates a column convolution on the row convolution calculation result in the shared memory to obtain the final convolution value within the 32×24 area. After all threads are completed, the complete image after Gaussian blur can be obtained.
[0018] As a preferred solution of the initial value search algorithm for GPU-accelerated SIFT feature point matching described in the present invention, the construction of a Gaussian difference pyramid includes: performing Gaussian blurring on the original image seven times in sequence to obtain a group of eight images, performing interval sampling on the third-to-last image as the 0th layer image of the next group, and then performing Gaussian blurring on the image seven times in sequence to obtain the next group of images, and so on. Each interval sampling reduces the length and width of the image to half of the original length and width until the length or width of the image is only one pixel. Finally, the difference between two adjacent images in each group is taken to generate a Gaussian difference pyramid.
[0019] As a preferred solution of the initial value search algorithm for GPU-accelerated SIFT feature point matching described in the present invention, the construction of a Gaussian difference pyramid further includes: after constructing the Gaussian difference pyramid, performing local extreme point detection on the middle five layers of the seven-layer difference pyramid in each group, each thread block containing 32 threads, completing extreme value detection for 30 pixel points. When detecting each point, the maximum and minimum values of eight points above, below, left, right, and diagonally, as well as nine points at corresponding positions in each of the upper and lower layers, are calculated, and a determination is made as to whether the point is between the minimum and maximum values. If not, the point is a local extreme point and is stored as a feature point.
[0020] As a preferred solution of the initial value search algorithm for GPU-accelerated SIFT feature point matching described in the present invention, the calculation of the feature point descriptor includes: first classifying the feature points according to the number of Laplacian pyramid levels in which they are located, and then calculating the main directions and descriptors of the feature points in the same level in parallel; the feature descriptor is calculated as follows: the thread block is set to a two-dimensional thread set containing 16×16, one thread block is responsible for completing the calculation of all pixels within a circular domain with a radius R of a feature point and obtaining a corresponding 128-dimensional feature vector, each thread is responsible for completing the calculation of a pixel within a circular domain with a radius R of the feature point in each layer of the image, obtaining the contribution of its corresponding seed point in 8 directions, accumulating it to the descriptor, placing the domain information of each feature point in global memory, and returning it to the CPU to complete the normalization step of the generated SIFT descriptor.
[0021] As a preferred solution of the initial value search algorithm for SIFT feature point matching based on GPU acceleration described in the present invention, wherein: the matching of SIFT feature points includes: in the SIFT feature point matching of two images, a thread block calculates the correlation score of 32 feature points in image one and all feature points in image two, a thread block is divided into 32×8 threads, 8 threads jointly calculate the correlation score of the same feature point with all feature points in image two, each thread calculates the correlation score of the 16-dimensional feature point descriptor of the same feature point with all feature points in image two, and finally the total score is obtained through an atomic addition operation, and the point with the largest correlation score is the matching point of the point.
[0022] As a preferred solution of the initial value search algorithm for SIFT feature point matching based on GPU acceleration described in the present invention, the control point matching equation constructed based on the calculation of SIFT feature points includes:
[0023] After obtaining the matched feature point pairs, the calculation steps for the initial matching value of the detection point are as follows: In the reference image, select the n points closest to the detection point as control points (n ≥ 3) for initial value solution. In this algorithm, n = 4 is taken. After obtaining the coordinates of the control point pairs, the first-order displacement model is used to establish the equation system for each pair of control points as follows:
[0024]
[0025] in, is the coordinate of the detection point in the reference image, is the coordinate of the control point under the target image, is the relative coordinate of the control point and the reference point in the reference image, is the first-order shape function corresponding to the detection point to be determined; u and v are the displacements of the detection point in the x and y directions of the image, respectively. and are the rates of change of displacement u in the x and y directions, and are the rates of change of displacement v in the x-direction and y-direction respectively; the equations of the four control points are connected and solved by the least squares method to obtain the first-order shape function, and then the matching initial value is calculated .
[0026] The beneficial effects of the present invention are as follows: the present invention performs GPU-based parallel optimization on the initial value search algorithm based on SIFT feature point matching in the strain measurement process of the synchronous automatic clutch synchronization mechanism, mainly optimizing the scale space construction and descriptor calculation in the SIFT feature point detection algorithm, thereby improving the measurement speed of the strain measurement of the synchronous automatic clutch synchronization mechanism; to adapt to the initial value search under the complex working conditions of the synchronous automatic clutch core synchronization mechanism, by finding the four feature point pairs with the closest matching degree to the detection point, a first-order displacement model equation group is established, allowing the target subset to translate, rotate, uniformly shear and stretch, and obtain the initial matching value through the least squares method, thereby improving the accuracy of the initial value search; meeting the real-time measurement of the full-field strain data of the synchronous automatic clutch core synchronization mechanism, helping to optimize the structure and materials of the core synchronization mechanism, improve its load-bearing capacity and fatigue resistance, and thus extend its service life. In addition, the strain response and fracture process of the core synchronization mechanism under stable resonant excitation can also be studied to predict possible damage locations and formulate effective maintenance strategies to further improve the reliability and safety of the synchronous automatic clutch. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0028] Figure 1 A flowchart of an initial value search algorithm for SIFT feature point matching based on GPU acceleration is provided in an embodiment of the present invention.
[0029] Figure 2 A schematic diagram of a GPU-accelerated Gaussian blur algorithm provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0030] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0031] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0032] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0033] Example 1
[0034] Reference Figure 1~Figure 2 , is an embodiment of the present invention, which provides an initial value search algorithm for SIFT feature point matching based on GPU acceleration, including:
[0035] S1: Allocate and align the memory of the image.
[0036] Preferably, memory of width w × height h is allocated for the image, and the memory size is pitch × h × size, where pitch is the smallest multiple of 128 greater than w, and size is the number of data bytes per pixel, ensuring that the starting address of each row of data meets the memory alignment requirements of the GPU.
[0037] Furthermore, assume that the image width w = 1024 pixels, the image height h = 768 pixels, and the data size of each pixel size = 4 bytes;
[0038] Calculate the number of bytes required for each row when it is not aligned: bytes_per_row = w × size = 1024 × 4 = 4096 bytes;
[0039] In this example, 4096 itself is a multiple of 128, so pitch = 4096;
[0040] Allocate GPU memory of size pitch × h × size: total_memory = pitch × h = 4096 × 768 = 3145728 bytes. Use the cudaMallocPitch function to allocate memory for the image to automatically ensure the correct alignment of the pitch.
[0041] S2: Get the complete Gaussian pyramid.
[0042] Preferably, the Gaussian kernel function is stored in the constant memory of the GPU in the form of an array, and Gaussian blur is performed on each group of images. A two-dimensional Gaussian kernel function is used to convolve with the lowest layer image. The convolution kernel size is (2R+1)×(2R+1), and R is the Gaussian blur convolution kernel window radius. The Gaussian blur convolution kernel window radius is set to 9. The two-dimensional Gaussian blur convolution kernel is as follows:
[0043]
[0044] According to the separability of the two-dimensional Gaussian function, the two-dimensional Gaussian blur convolution kernel is expressed as the product of two one-dimensional Gaussian functions, as shown in the following formula:
[0045]
[0046] Where x and y are the horizontal and vertical position coordinates, σ is the standard deviation of the Gaussian distribution, and S is the normalization constant of the two-dimensional Gaussian function. and are the one-dimensional Gaussian functions along the horizontal and vertical directions respectively, S´ is the normalization constant of the one-dimensional Gaussian function;
[0047] After reducing the dimensionality of the two-dimensional Gaussian blur convolution kernel, for images with a resolution of w×h, the time complexity of the algorithm is reduced from O(R×R×w×h) to O(R×R×w)+O(R×R×h), reducing the amount of computation.
[0048] Furthermore, assume that the Gaussian blur convolution kernel window radius R = 9 (meaning the convolution kernel size is (2R + 1) × (2R + 1) = 19 × 19), and the standard deviation σ = R / 3;
[0049] Calculate the normalization constant S′:
[0050]
[0051] Construct one-dimensional Gaussian kernels Gx and Gy of length (2R+1)(2R + 1)(2R+1), where x,y∈[-R,R];
[0052] Perform horizontal convolution on each row of the image, and calculate each pixel as follows:
[0053]
[0054] After completing the horizontal blur, perform vertical convolution:
[0055]
[0056] Starting from the lowest level, the next level of the pyramid image is obtained by Gaussian blurring. The above blurring operation is repeated, and the image is scaled (for example, the resolution of each level is reduced to half of the previous level) until the complete pyramid is constructed.
[0057] S3: GPU accelerated calculation of Gaussian blur.
[0058] Preferably, GPU threads are allocated. In the GPU architecture, a warp is a group of 32 threads executed in parallel. If the number of threads in each thread block is a multiple of 32, then all warps are fully filled and there will be no unused threads. Each thread block is set to have 32×32 threads, and each thread block is responsible for calculating the final convolution value of a 32×24 area.
[0059] Preferably, convolution calculation is performed on GPU threads, and each thread first calculates a row convolution. For a thread block, the row convolution calculation result of a 32×32 area is obtained;
[0060] The calculation result is stored in the shared memory of the thread block. Then, each thread in the middle area of 32×24 in the thread block performs a column convolution on the row convolution result in the shared memory to obtain the final convolution value in the 32×24 area. After all threads are completed, the complete image after Gaussian blur can be obtained.
[0061] Furthermore, the value of the one-dimensional Gaussian kernel function is calculated (the kernel radius is 9), stored in an array, and loaded into the constant memory of the GPU for fast access during the convolution process.
[0062] The entire image is divided into several 32×24 small areas. Each small area is processed by a thread block to ensure that the divided small blocks can completely cover the image. If the image size is not an integer multiple of 32×24, some extra pixels can be padded at the edge of the image to make it fit the size.
[0063] For each small block, all threads first perform row-wise Gaussian convolution on the small area they are responsible for and store the results in shared memory so that subsequent operations can quickly access them.
[0064] After the row convolution is completed, the threads in the middle 32×24 area of the thread block read the row convolution results in the shared memory again and perform Gaussian convolution in the column direction. Since the shared memory can be read very quickly, the time required for the convolution operation can be greatly reduced.
[0065] Merge small blocks to generate the final blurred image. Merge the blurred results of the 32×24 small areas calculated by each thread block to generate a complete blurred image.
[0066] S4: Get the complete Gaussian difference pyramid.
[0067] Preferably, the original image is Gaussian blurred seven times in sequence to obtain a group of 8 images. The third-to-last image is sampled as the 0th layer image of the next group. The image is then Gaussian blurred seven times in sequence to obtain the next group of images. And so on. Each time the image is sampled, the length and width of the image are reduced to half of the original length or width until the length or width of the image is only one pixel. Finally, the difference between the two adjacent images in each group is taken to generate a Gaussian difference pyramid.
[0068] Preferably, after constructing the Gaussian difference pyramid, local extreme point detection is performed on the middle 5 layers of the 7-layer difference pyramid in each group. Each thread block contains 32 threads, and extreme value detection of 30 pixel points is completed. When each point is detected, the maximum and minimum values of 8 points above, below, left, right and diagonal lines and 9 points at corresponding positions in each layer of the upper and lower layers are calculated to determine whether the point is between the minimum and maximum values. If not, the point is a local extreme point and is stored as a feature point.
[0069] Furthermore, starting from an original image, the image is Gaussian blurred seven times in sequence to obtain eight layers of images (including the original image).
[0070] Each Gaussian blur uses a different standard deviation σ, and the standard deviation of each blur increases successively to ensure that the image smoothness at different scales is different and to ensure that the Gaussian kernel function is stored in the constant memory of the GPU to improve access efficiency.
[0071] Take the third-to-last blurred image and use it as the starting layer for the next set of images. Downsample this image to reduce its resolution to half (the length and width are halved respectively). Then Gaussian blur it seven times to get the next set of images, and repeat the above process. Each downsampling process gradually reduces the size of the image until the length or width of the image is only one pixel.
[0072] For each set of images in the Gaussian pyramid, the image difference between two adjacent layers is calculated to generate seven differential images. The differential operation is achieved by pixel-wise subtraction of adjacent Gaussian blurred images to obtain a Gaussian difference pyramid. GPU threads perform this operation in parallel to reduce processing time.
[0073] After constructing the Gaussian Difference pyramid, the five middle layers of the pyramid are used to detect local extreme points for feature point extraction. In the Gaussian Difference pyramid, each pixel to be detected is compared with its eight adjacent points on the upper, lower, left, and right diagonals of the image, as well as nine points at the same location in the upper and lower layers. If the pixel value of a point is greater or less than the values of all points in the detection area, it is determined to be a local extreme point and retained as a feature point.
[0074] Each GPU thread block contains 32 threads. Through proper allocation, each thread block processes a 30-pixel area, ensuring maximum utilization of the GPU's parallel computing capabilities. This block division allows threads within each thread block to process pixels in different areas in parallel. Each thread reads the value of the pixel it is responsible for, obtains the values of the eight adjacent pixels, and the corresponding nine pixels above and below it for comparison.
[0075] S5: Calculate the feature point descriptor.
[0076] Preferably, the feature points are first classified according to the number of Laplacian pyramid levels they are in, and then the main directions and descriptors of the feature points in the same level are calculated in parallel;
[0077] The feature descriptor is calculated as follows: the thread block is set to a two-dimensional thread set containing 16×16. One thread block is responsible for completing the calculation of all pixel points within a circular domain with a radius of R of a feature point and obtaining the corresponding 128-dimensional feature vector. Each thread is responsible for completing the calculation of a pixel point within a circular domain with a radius of R of a feature point in each layer of the image, obtaining the contribution of its corresponding seed point in 8 directions, accumulating it to the descriptor, placing the domain information of each feature point in the global memory, and returning it to the CPU to complete the normalization step of the generated SIFT descriptor.
[0078] Furthermore, assuming that the input data image resolution is 512×512 pixels, the number of feature points is 64, the circular radius R of each feature point is 8 pixels, the size of each thread block is 16×16 threads, and the dimension of the descriptor is 128 dimensions, assuming that the position of one of the feature points is (256,256), the area around it will be divided into a circular area with a radius of 8 for calculation.
[0079] Calculate the main direction of the feature point. Input the feature point coordinates: (256, 256). Circle pixels: includes all pixels within an 8-pixel radius, about 225 pixels.
[0080] Calculate the gradient for (256, 254), assuming the horizontal gradient is 5 and the vertical gradient is 3, the gradient amplitude is equal to 52+32=34≈5.83 and the direction angle is equal to atan(3 / 5)≈30∘;
[0081] The direction of each pixel is accumulated to the corresponding position of the 36 direction histograms. Assuming that the (256, 254) direction angle is 30°, it will be accumulated to the histogram interval of 0° to 10°. Assuming that the maximum value of the histogram in the 90° direction is 40, then this direction is the main direction, and the main direction is determined to be 90°.
[0082] Gradient direction contribution calculation: Each thread calculates the gradient information of the assigned pixel. Assume that a pixel has a gradient of 7 and a direction of 110°. In the direction histogram, it is classified as "0°-45°" and has a weighted contribution of 7.
[0083] All threads in the thread block accumulate the results in shared memory to reduce global memory access, and accumulate the information of each direction of the 4×4 small area to obtain a 128-dimensional feature descriptor.
[0084] S6: Match SIFT feature points.
[0085] Preferably, in the SIFT feature point matching of two images, one thread block calculates the correlation scores of 32 feature points in image one and all feature points in image two, one thread block is divided into 32×8 threads, 8 threads jointly calculate the correlation scores of the same feature point and all feature points in image two, each thread calculates the correlation scores of the 16-dimensional feature point descriptors of the same feature point and all feature points in image two, and finally the total score is obtained through atomic addition operation, and the point with the largest correlation score is the matching point of that point.
[0086] Furthermore, the example has 4 thread blocks, each processing 128 feature points in parallel:
[0087] Thread block 1: Process feature points 1-32 in image 1;
[0088] Thread block 2: Process feature points 33-64 in image 1;
[0089] Thread block 3: Processing feature points 65-96 in image 1;
[0090] Thread block 4: Processing feature points 97-128 in image 1;
[0091] After each thread block completes the calculation, the best matching feature point result for each feature point in image 1 is saved to shared memory. Finally, it is transferred back to the CPU, where all matching pairs of feature points are summarized to form complete image matching data.
[0092] Thanks to the 32×8 thread configuration, the calculation process of each feature point is fully parallelized, significantly improving performance. The use of shared memory reduces the number of global memory accesses and increases data processing speed.
[0093] Assuming that a single-threaded feature point matching calculation takes 0.005 seconds, the time can be shortened to less than 0.1 seconds after GPU parallel calculation.
[0094] S7: Construct control point matching equation based on the calculation of SIFT feature points.
[0095] Preferably, after obtaining the matched feature point pairs, the calculation steps for the initial matching value of the detection point are as follows: in the reference image, select n points closest to the detection point as control points (n≥3) for solving the initial value. In this algorithm, n=4 is taken. After obtaining the coordinates of the control point pairs, the first-order displacement model is used to establish the following equation system for each pair of control points:
[0096]
[0097] in, is the coordinate of the detection point in the reference image, is the coordinate of the control point under the target image, is the relative coordinate of the control point and the reference point in the reference image, is the first-order shape function corresponding to the detection point to be determined; u and v are the displacements of the detection point in the x and y directions of the image, respectively. and are the rates of change of displacement u in the x and y directions, and are the rates of change of displacement v in the x and y directions, respectively;
[0098] The equations of the four control points are connected and solved by the least square method to obtain the first-order shape function, and then the matching initial value is calculated. .
[0099] Furthermore, assuming image A and target image B, find the matching position in image B corresponding to the detection point in image A.
[0100] Assume that a feature point P in image A has coordinates (x p ,y p ), using the SIFT matching algorithm, assuming that 100 pairs of feature points are obtained, 4 of which are selected as control points.
[0101] Control point coordinates (reference image A):
[0102] C1=(x c1 ,y c1 ), C2=(x c2 ,y c2 ), C3=(x c3 ,y c3 ), C4=(x c4 ,y c4);
[0103] Assume that the matching points of these four control points in the target image B are:
[0104] C 1′ =(x c1′ ,y c1′ ),C 2′ =(x c2′ ,y c2′ ), C 3′ =(x c3′ ,y c3′ ), C 4′ =(x c4′ ,y c4′ );
[0105] Assume that the matching point corresponding to the detection point P is P′=(x p′ ,y p′ ), the displacement pattern can be estimated by the control points. In order to obtain the initial value of P′, it is necessary to construct the relative displacement equation between the control points.
[0106] Representation of the displacement equation:
[0107] x p′ =a1⋅x p +a2⋅y p +b1,y p′ =a3⋅x p +a4⋅y p +b2;
[0108] Use 4 control points to construct a system of equations, organize these equations into a matrix form AX=B, and solve the X matrix using the least squares method, that is:
[0109]
[0110] Get the parameters of the first-order shape function a1, a2, a3, a4, b1, b2. Assume that after the solution, a1=0.98, a2=0.03, b1=5.1, a3=-0.04, a4=1.02, b2=3.5.
[0111] Calculate the matching initial value: P=(50,30)→P′=(x p′ ,y p′ )=(54.7,35.4).
[0112] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. An initial value search algorithm for SIFT feature point matching based on GPU acceleration, characterized in that: include: Allocate and align the memory of the image; Construct a multi-scale space of adjacent multi-frame Gaussian pyramids based on GPU, obtain a complete Gaussian pyramid and Gaussian difference pyramid, and perform GPU-accelerated calculation of Gaussian blur; Calculate the feature point descriptor; Match SIFT feature points; The control point matching equation is constructed based on the calculation of SIFT feature points.
2. The initial value search algorithm for SIFT feature point matching based on GPU acceleration according to claim 1, characterized in that: The memory allocation and alignment of the image includes: Allocate memory of width w × height h for the image. The memory size is pitch × h × size, where pitch is the smallest multiple of 128 greater than w, and size is the number of data bytes per pixel. Ensure that the starting address of each row of data meets the GPU memory alignment requirements.
3. The initial value search algorithm for SIFT feature point matching based on GPU acceleration as claimed in claim 1, characterized in that: The complete Gaussian pyramid is obtained, including: The Gaussian kernel function is stored in the constant memory of the GPU in the form of an array. Gaussian blur is performed on each group of images. A two-dimensional Gaussian kernel function is used to convolve with the lowest layer image. The convolution kernel size is (2R+1)×(2R+1). R is the Gaussian blur convolution kernel window radius. The Gaussian blur convolution kernel window radius is set to 9. The two-dimensional Gaussian blur convolution kernel is shown below: ; According to the separability of the two-dimensional Gaussian function, the two-dimensional Gaussian blur convolution kernel is expressed as the product of two one-dimensional Gaussian functions, as shown in the following formula: ; Where x and y are the horizontal and vertical position coordinates, σ is the standard deviation of the Gaussian distribution, and S is the normalization constant of the two-dimensional Gaussian function. and are the one-dimensional Gaussian functions along the horizontal and vertical directions respectively, S´ is the normalization constant of the one-dimensional Gaussian function; After reducing the dimensionality of the two-dimensional Gaussian blur convolution kernel, for images with a resolution of w×h, the time complexity of the algorithm is reduced from O(R×R×w×h) to O(R×R×w)+O(R×R×h), reducing the amount of computation.
4. The initial value search algorithm for SIFT feature point matching based on GPU acceleration according to claim 1, characterized in that: The GPU-accelerated calculation of Gaussian blur includes: Allocate GPU threads. In the GPU architecture, a warp is a group of 32 threads that execute in parallel. Each thread block is set to have 32×32 threads, and each thread block is responsible for calculating the final convolution value of a 32×24 area.
5. The initial value search algorithm for SIFT feature point matching based on GPU acceleration according to claim 1 or 4, characterized in that: The GPU-accelerated calculation of Gaussian blur also includes: Convolution calculation is performed on GPU threads. Each thread first calculates a row convolution. For a thread block, the row convolution calculation result of a 32×32 area is obtained. The calculation result is stored in the shared memory of the thread block. Then, each thread in the middle area of 32×24 in the thread block performs a column convolution on the row convolution result in the shared memory to obtain the final convolution value in the 32×24 area. After all threads are completed, the complete image after Gaussian blur can be obtained.
6. The initial value search algorithm for SIFT feature point matching based on GPU acceleration as claimed in claim 1, characterized in that: The complete Gaussian difference pyramid is obtained, including: The original image is Gaussian blurred seven times in sequence to obtain a group of 8 images. The third-to-last image is sampled as the 0th layer image of the next group. The image is then Gaussian blurred seven times in sequence to obtain the next group of images. And so on. Each time the length and width of the image are reduced to half of the original, until the length or width of the image is only one pixel. Finally, the difference between the two adjacent images in each group is taken to generate a Gaussian difference pyramid.
7. The initial value search algorithm for SIFT feature point matching based on GPU acceleration according to claim 1 or 6, characterized in that: The method of obtaining a complete Gaussian difference pyramid further includes: After constructing the Gaussian difference pyramid, local extreme point detection is performed on the middle 5 layers of the 7-layer difference pyramid in each group. Each thread block contains 32 threads and completes extreme value detection of 30 pixel points. When detecting each point, the maximum and minimum values of the 8 points above, below, left, right and diagonal lines and the 9 points at the corresponding positions in each layer of the upper and lower layers are calculated to determine whether the point is between the minimum and maximum values. If not, the point is a local extreme point and is stored as a feature point.
8. The initial value search algorithm for SIFT feature point matching based on GPU acceleration according to claim 1, characterized in that: The calculating of the feature point descriptor includes: The feature points are first classified according to the number of Laplacian pyramid layers they are in, and then the main directions and descriptors of the feature points in the same layer are calculated in parallel; The feature descriptor is calculated as follows: the thread block is set to a two-dimensional thread set containing 16×16. One thread block is responsible for completing the calculation of all pixel points within a circular domain with a radius of R of a feature point and obtaining the corresponding 128-dimensional feature vector. Each thread is responsible for completing the calculation of a pixel point within a circular domain with a radius of R of a feature point in each layer of the image, obtaining the contribution of its corresponding seed point in 8 directions, accumulating it to the descriptor, placing the domain information of each feature point in the global memory, and returning it to the CPU to complete the normalization step of the generated SIFT descriptor.
9. The initial value search algorithm for SIFT feature point matching based on GPU acceleration according to claim 1, characterized in that: The matching of SIFT feature points includes: In the SIFT feature point matching of two images, one thread block calculates the correlation score between the 32 feature points in image one and all the feature points in image two. One thread block is divided into 32×8 threads. The eight threads jointly calculate the correlation score between the same feature point and all the feature points in image two. Each thread calculates the correlation score of the 16-dimensional feature point descriptor of the same feature point and all the feature points in image two. Finally, the total score is obtained through the atomic addition operation. The point with the largest correlation score is the matching point of that point.
10. The initial value search algorithm for SIFT feature point matching based on GPU acceleration according to claim 1, characterized in that: The calculation based on SIFT feature points to construct a control point matching equation includes: After obtaining the matched feature point pairs, the calculation steps for the initial matching value of the detection point are as follows: In the reference image, select the n points closest to the detection point as control points (n ≥ 3) for initial value solution. In this algorithm, n = 4 is taken. After obtaining the coordinates of the control point pairs, the first-order displacement model is used to establish the equation system for each pair of control points as follows: ; in, is the coordinate of the detection point in the reference image, is the coordinate of the control point under the target image, is the relative coordinate of the control point and the reference point in the reference image, is the first-order shape function corresponding to the detection point to be determined; u and v are the displacements of the detection point in the x and y directions of the image, respectively. and are the rates of change of displacement u in the x and y directions, and are the rates of change of displacement v in the x and y directions, respectively; The equations of the four control points are connected and solved by the least square method to obtain the first-order shape function, and then the matching initial value is calculated. .
Citation Information
Patent Citations
SIFT feature matching method based on OpenCL parallel acceleration
CN104732221A
Binocular matching method based on GPU acceleration
CN114937159A
Method for accelerating a CDVS extraction process based on a gpgpu platform
US20190139186A1