Infrared and visible light image registration method based on improved SIFT algorithm
By improving the SIFT algorithm, Gaussian differential pyramids, local adaptive FAST key point extraction and main direction allocation, combined with the optimized RANSAC algorithm, the problem of incomplete detection of key point and redundancy in infrared and visible image registration is solved, and more efficient and accurate image registration is achieved.
Patent Information
- Application Number
- CN202510098789.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-13
AI Technical Summary
In the process of infrared and visible light image registration, the SIFT algorithm has problems such as incomplete detection of key points and incomplete redundancy and matching processes for gradient direction and mode value calculation, which makes it difficult to take into account both registration accuracy and real-timeness.
By constructing a Gaussian differential pyramid of infrared and visible light images, local adaptive FAST key point extraction, main direction allocation and feature descriptor generation, combined with the optimized RANSAC algorithm, mismatch point removal is achieved, efficient registration of infrared and visible light images.
It improves the robustness and effectiveness of key points, enhances the accuracy and efficiency of feature matching, and overcomes the problem of poor feature extraction effect in low-contrast or strong noise images.
Smart Images

Figure CN120147375A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of digital image processing, and in particular relates to an infrared and visible light image registration method based on an improved SIFT algorithm. Background Art
[0002] Image registration is a classic problem and technical difficulty in the field of image processing. Its goal is to establish the optimal mapping relationship between images acquired under different conditions (such as different acquisition devices, shooting time or angle). This technology has important practical value in computer vision, medical image processing, remote sensing and material mechanics. However, due to the diversity of image types and application scenarios, it is currently impossible to develop a general optimization method that meets all needs. Among many registration tasks, the registration of infrared images and visible light images plays an irreplaceable role because it can fuse two complementary information and is widely used in practical scenarios such as target detection, monitoring and navigation.
[0003] The SIFT algorithm plays an important role in the field of classical image processing. Compared with other registration algorithms, SIFT is robust to scale changes, and its descriptor is invariant to image rotation, and has a wide range of applications. Compared with the currently emerging deep learning methods, SIFT does not require labeled data and directly extracts and matches features of the input image. At the same time, the entire SIFT process is based on a clear mathematical model, has high interpretability, and does not require a large amount of computing resources and high-performance hardware platforms. For objects to be processed such as infrared and visible light images that have significant spectral differences, SIFT's gradient information extraction does not depend on pixel intensity values, so it has advantages for this cross-spectral registration.
[0004] Although SIFT is highly competitive in the field of image registration, it still has problems such as poor performance in applications with high real-time requirements and poor feature extraction in low-contrast or strong noise images due to its high computational complexity and reliance on gradients. The present invention proposes an infrared and visible light image registration method based on an improved SIFT algorithm, which overcomes the shortcomings of the existing SIFT algorithm that image registration and accuracy cannot be taken into account at the same time. Summary of the invention
[0005] Purpose of the invention: The purpose of the present invention is to solve the problems of incomplete key point detection in the SIFT algorithm during the registration process, redundant calculation of gradient direction and modulus value in the process of key point determination of the main direction, and imperfect matching process, and propose an infrared and visible light image registration method based on an improved SIFT algorithm.
[0006] Technical solution: The infrared and visible light image registration method based on the improved SIFT algorithm described in the present invention specifically includes the following steps:
[0007] (1) Construct the Gaussian difference pyramid of infrared and visible images to effectively capture the multi-scale features existing in the images;
[0008] (2) Perform local adaptive FAST key point extraction on each layer of the pyramid to detect the significant feature points of the images at the multi-scale level;
[0009] (3) Based on the gradient direction distribution of the images within the neighborhood of the key points, assign the main direction to the key points so that the key points can still maintain consistency under different rotation angles;
[0010] (4) Encode the key points based on the gradient information features and generate the corresponding feature descriptors for subsequent registration operations;
[0011] (5) For the feature points extracted from the infrared and visible images, use the similarity between the descriptors for initial registration, and use the optimized RANSAC algorithm to verify and optimize the initial registration result to eliminate the mismatched points.
[0012] Furthermore, the specific steps of step (1) are as follows:
[0013] Input the original infrared image I ir (x, y) and the visible image I vi (x, y). On the basis of downsampling the two-modal original images and adding Gaussian filtering, obtain images with different scales and sizes from large to small and from bottom to top; Subtract the adjacent two layers of Gaussian pyramids to obtain the Gaussian difference pyramid, which is convenient for effectively extracting stable key points;
[0014] Among them, the Gaussian scale space is defined as: L(x, y, σ) = G(x, y, σ) * I(x, y), where I(x, y) represents the pixel value of the input image at (x, y), G(x, y, σ) represents the Gaussian kernel function, and the variance is σ 2 , and L(x, y, σ) represents the scale image of the original image after Gaussian filtering; The Gaussian difference pyramid is the difference between the adjacent two layers of scale images, which is defined as:
[0015] D(x, y, σ) = [G(x, y, kσ) - G(x, y, σ)] * I(x, y) = L(x, y, kσ) - L(x, y, σ)
[0016] Among them, k is the scale change coefficient.
[0017] Furthermore, the specific steps of step (2) are as follows:
[0018] Perform local adaptive FSAT corner detection on the scaled image: Extract the neighborhood for each pixel point, calculate the statistical characteristics of the pixel brightness within the neighborhood, and based on the neighborhood statistical characteristics, calculate the dynamic threshold T = t(T max -T min ), where t is the adjustment coefficient, and T max and T min are the maximum and minimum mean values of the pixels defined within the neighborhood; compare the brightness of the neighborhood pixels. If at least J consecutive pixels have an absolute brightness difference greater than T from the central pixel, then this pixel point is marked as a key point.
[0019] Furthermore, the neighborhood pixels are 16 pixels on a circular window with a diameter of 7.
[0020] Furthermore, the specific steps of step (3) are as follows:
[0021] Taking the key point as the center, select the 3σ window size of its pyramid image as the neighborhood. Starting from a certain starting pixel point in the neighborhood, gradually mark the pixels with an absolute pixel difference within the threshold M as the same region, and record this region as the connected region; for the points within the neighborhood, only calculate the gradient direction and gradient magnitude of all pixel points outside the connected region and the outermost circle of pixel points within the connected region;
[0022] After completing the gradient calculation of the key points, use the gradient histogram to count the gradients and directions of the pixels within the neighborhood to determine the main direction of the key points; the peak direction of the histogram represents the main direction of the key points.
[0023] Furthermore, the gradient histogram divides the 0 - 360° direction into 10 bins, each bin representing 36°, and the peak direction of the histogram represents the main direction of the key points.
[0024] Furthermore, the specific steps of step (4) are as follows:
[0025] Divide the neighborhood near the key point into 4 * 4 sub - regions. Each sub - region is used as a seed point. Each seed point takes the main direction of the feature point as the reference direction, calculates the angle between its gradient direction and the reference direction, and evenly distributes it into 8 directions at intervals of 45° within the range of 0 - 360°. When calculating the gradient direction and gradient magnitude of the pixel points; to highlight that the pixels at the center of the sub - region contribute more to the gradient description, use a Gaussian kernel to weight the gradient amplitude of each pixel. The Gaussian weight is:
[0026]
[0027] where (x c ,y c) is the center of the sub-region, and σ controls the distribution range of the weighting; the weighted gradient magnitude is: G′ = G·w(x,y). After obtaining the optimized SIFT feature descriptor, to reduce the influence of illumination and enhance the robustness, normalization processing is performed on it.
[0028] Further, the specific step (5) is as follows:
[0029] The key point descriptor R i extracted from the infrared image i1 =(r i2 , r i128 ),…, r i ), and the key point descriptor V i1 extracted from the visible light image i2 =(v i128 ),…, v
[0030]
[0031] Then the Euclidean distance between the two points is: 1 For a descriptor r 1 in a certain infrared image 2 , find two closest descriptors v 1 and v 2 in the visible light image, and the corresponding distances are d Calculate the ratio 1 If r is less than a certain set threshold, then it is considered that r 1 matches v
[0032] Otherwise, discard it;
[0033] Beneficial effects: Compared with the prior art, the beneficial effects of the present invention are:
[0034] Based on local adaptive FAST corners, the present invention generates a large number of preselected key points with uniform distribution while reducing the calculation cost, preventing the omission of potential feature points. At the same time, it realizes the dynamic adjustment of the detection threshold in regions with different brightness or contrast, effectively avoiding the problem of key point loss in low-contrast regions and improving the robustness of key points. By detecting connected regions, the present invention greatly reduces unnecessary gradient calculations when calculating the main direction of key points, excludes low-texture regions (i.e., connected regions), focuses on the parts with stronger structural information in the image, and improves the effectiveness and matching efficiency of feature points. The present invention uses the improved RANSAC algorithm to speed up the iteration speed when removing mismatched points and improve the matching efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a flowchart of the present invention;
[0036] Figure 2 is a comparison diagram of infrared image feature point detection provided in an embodiment of the present invention;
[0037] Figure 3 is a schematic diagram of binary image connected region retrieval provided in an embodiment of the present invention;
[0038] Figure 4 is a schematic diagram of 4-neighbor connected region retrieval of an infrared image in an actual scenario provided by the present invention;
[0039] Figure 5 is a result diagram of image registration using the traditional SIFT algorithm provided by the present invention;
[0040] Figure 6 is a result diagram of image registration based on the improved SIFT algorithm provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The present invention will be further described in detail below with reference to the accompanying drawings.
[0042] As Figure 1 shown, the present invention discloses an infrared and visible light image registration method based on an improved SIFT algorithm, including the following steps:
[0043] Step 1: Construct a Gaussian difference pyramid of the image.
[0044] Input the original image I ir (x,y) and I vi (x,y). On the basis of downsampling the original images of the two modalities and adding Gaussian filtering, images with different scales and sizes from large to small and from bottom to top are obtained; the difference between adjacent layers of the Gaussian pyramid is taken to obtain a Gaussian difference pyramid, which is convenient for effectively extracting stable key points.
[0045] To ensure scale invariance, features need to be extracted in a multi-scale space. First, the input original image I(x, y) is convolved with the multi-scale Gaussian kernel G(x, y, σ) to construct the Gaussian multi-scale space L(x, y, σ) = G(x, y, σ) * I(x, y); I(x, y) represents the pixel value of the input image at (x, y), G(x, y, σ) represents the Gaussian kernel function with variance σ 2 , and L(x, y, σ) represents the scale image of the original image after Gaussian filtering.
[0046] Construct a Gaussian difference pyramid. Take the difference of the obtained multi-scale images to get the Gaussian difference image D(x, y, σ) = [G(x, y, kσ) - G(x, y, σ)] * I(x, y) = L(x, y, kσ) - L(x, y, σ), where k is the scale change coefficient. Thus, the blurring effect in the image processing process is weakened, and features such as edges and corners are more prominent in the difference image, facilitating subsequent detection.
[0047] Step 2: Perform fast key point extraction on each layer of the pyramid.
[0048] Divide the neighborhood of the detected pixel points. Usually, a circular window with a diameter of 7 centered on the pixel point is selected.
[0049] Set the dynamic threshold within this area. Compared with the traditional extreme value detection method, the local custom FAST corner detection is more sensitive to changes in noise and illumination, and the key points selected contain less information and have low stability. To find a more generalizable threshold, perform a gray-level histogram statistics on the divided neighborhood window, g min is the minimum pixel gray value in the histogram, g max is the maximum pixel gray value in the histogram, g 1 is located on the right side of g min is located on the left side of g 2 is located on the right side of g min is located on the left side of g
[0050]
[0051] where MN is the number of pixel points within this area range, n i is the number of pixel points with pixel value equal to g i ; the adaptive threshold T = t(T max - T min ), where t is the adjustment coefficient.
[0052] Select the key point positions. Subtract the gray value of the central pixel from the gray values of the pixels on the edge of the circular neighborhood, and compare the absolute value of the difference with T. If the absolute values of the differences between the central pixel and J consecutive pixels are all greater than T, then this point is considered a key point. As Figure 2 shown, Figure 2 in (a) are the key points extracted by the original SIFT algorithm, Figure 2 in (b) are the key points extracted by using the FAST corner detection method, Figure 2 and in (c) are the key points located by using the adaptive threshold FAST corner detection method in this step.
[0053] Step 3: Assign the main direction to the key points.
[0054] This step mainly aims at the problems such as the long calculation process caused by a large number of trigonometric calculations and modulus calculations involved in the SIFT key point direction assignment process. The specific steps are as follows:
[0055] Taking the key point as the center, select the 3σ window size of its image scale as the neighborhood. Starting from a certain starting pixel point in the neighborhood, considering the 4-connected method, perform region growing operations in the local neighborhood. Select a starting pixel as the seed point. For the current pixel, use depth-first search. Starting from the starting pixel point, visit this pixel point and mark it as visited, and recursively visit the unvisited adjacent pixel points of this pixel point. If all the adjacent pixel points of a certain pixel point have been visited, then backtrack to the upper-level pixel point.
[0056] Repeat the above steps until all nodes are visited. Check the four adjacent pixels of its up, down, left, and right. If the absolute value of the gray difference is less than the threshold, add it to the current connected region, and add the four adjacent pixels to the queue to be detected. Repeat this process until all eligible pixels are marked as part of the connected region. Figure 3 is a schematic diagram of the 4-connected region of a binary image, Figure 4 and is a schematic diagram of a certain connected region of an infrared image in the actual scene.
[0057] For the points in the neighborhood of the key point, only calculate the gradient directions and gradient moduli of all pixels outside the connected region and the outermost circle of pixels inside the connected region.
[0058] After completing the calculation of the relevant pixel points, count the gradients and directions of the pixels in the gradient histogram; the gradient histogram divides the 0-360° direction into 10 bins, each bin representing 36°, and the peak direction of the histogram represents the main direction of the key point.
[0059] Step 4: Describe the features of the key points.
[0060] The neighborhood near the key point is divided into 4*4 sub-regions. The size of each sub-region is the same as that when the key point direction is assigned. Each sub-region serves as a seed point. Taking the main direction of the feature point as the reference direction, the included angle between its gradient direction and the reference direction is calculated and evenly distributed over 8 directions at intervals of 45° within the range of 0-360°. When calculating the gradient direction and gradient magnitude of a pixel point, the pixel selection rule is the same as that described in step 3.
[0061] To highlight that the pixel at the center of the sub-region contributes more to the gradient description, a Gaussian kernel is used to weight the gradient amplitude of each pixel. The Gaussian weight is: (x c ,y c ) is the center of the sub-region, and σ controls the distribution range of the weighting; the weighted gradient amplitude is: G′ = G·w(x,y), where G is the original gradient amplitude of this point.
[0062] Generate a SIFT descriptor, which is a total of 16×4×4 = 128 dimensions. Let the obtained descriptor be H = (h 1 ,h 2 ,…,h 128 ). To remove the influence of illumination changes, they need to be normalized to shift the overall image gray value. Then the normalized feature vector L = (l 1 ,l 2 ,…,l 128 ).
[0063] Step 5: For the feature points extracted from the infrared and visible light images, perform initial registration and use the optimized RANSAC algorithm to remove mis-matched points.
[0064] The key point descriptor R i =(r i1 ,r i2 ,…,r i128 ) extracted from the infrared image, and the key point descriptor V i =(v i1 ,v i2 ,…,v i128 ) extracted from the visible light image. Then the Euclidean distance between two points For a descriptor r 1 in a certain infrared image, find two closest descriptors v 1 and v 2 in the visible light image, with corresponding distances d 1 and d 2 . Calculate the ratio If r is less than a certain set threshold, then r 1 is considered to be corresponding to v 1Match, otherwise discard.
[0065] Given a set of matching point pairs \(\{(x i , x' i )\}, where \(x i \) and \(x' i \) respectively represent the matching feature points in the infrared and visible light images, and the Euclidean distance between the matching points is \(d(x i , x' i ) = ||x i - x' i ||. Arrange the Euclidean distances of all matching point pairs in ascending order, and randomly select four points from the top N points with the highest scores as inliers to construct a transformation model. Let the affine transformation matrix \(T affine \) be the coordinate system for the transformation from the infrared image to the visible light image. The matrix can be calculated by \(X' = T affine X\), where \(x = [x, y, 1] T \) is the coordinate in the infrared image, and \(x' = [x', y', 1] T \) is the coordinate in the visible light image. The transformation matrix \(T affine \) can be solved by constructing a system of equations through the selected four point pairs.
[0066] Randomly select a point pair \((x j , x' j )\) from the remaining matching point pairs, substitute it into the transformation matrix \(T affine \) and calculate the reprojection error of this point. If the error is less than the set threshold \(\tau\), then this point pair is considered an inlier; if the error is greater than \(\tau\), then the matching of this point pair is considered a false match and needs to be excluded. For each calculated transformation model, record the number of inliers. By repeating the above steps, continuously update the inlier set. After multiple iterations, select the transformation model with the largest number of inliers as the final estimated transformation model.
[0067] The inventor uses the scene images of vehicles driving on the street in the Roadscene dataset for experimental analysis, and conducts a comparative experiment on feature extraction, registration, and optimization of the experimental images. The algorithm evaluation is carried out from multiple aspects such as the number of extracted feature points, the number of matching feature point pairs, the root mean square error (RMSE), the accuracy rate, and the time consumed for matching. Figure 5 is the matching result of the traditional SIFT algorithm, Figure 6 is the matching result achieved by the improved SIFT algorithm of the present invention.
[0068] Table 1 shows the comparative experimental results of registering infrared and visible light images using the traditional SIFT algorithm and the improved SIFT algorithm of the present invention.
[0069] Table 1 Comparison results of infrared and visible light image registration
[0070]
[0071] Infrared and visible light images with a resolution of 500*329 were selected for result comparison. The results in Table 1 show that in terms of key point localization, the method of the present invention can detect more feature points compared with traditional algorithms. Combining Figure 5 , Figure 6 and the visualization results, it can be seen that the optimized key point localization algorithm can locate corners, edges, etc. with more local features, and the key points located at some textures are also more comprehensive than those of the traditional extreme point localization algorithm. The number of feature point pairs extracted and matched by the method of the present invention has increased significantly, and good results have also been achieved in terms of RMSE and accuracy. At the same time, when the quantity and quality of the matching point pairs have been significantly improved, the method of the present invention runs faster. Based on the above analysis, the improved method not only enhances the quantity and quality of the extracted features, improves the robustness of the features, but also effectively improves the matching efficiency.
[0072] The above embodiments are only used to illustrate the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any modification made on the basis of the technical solution according to the technical idea proposed by the present invention shall fall within the protection scope of the present invention.
Claims
1. A method for infrared and visible light image registration based on an improved SIFT algorithm, characterized in that: The following steps are involved: (1) Construct a Gaussian difference pyramid of infrared and visible light images to effectively capture the multi-scale features in the images; (2) Perform local adaptive FAST key point extraction on each layer of the pyramid to detect the salient feature points of the image at multiple scales; (3) Based on the gradient direction distribution of the image in the neighborhood of the key point, the main direction is assigned to the key point so that the key point can remain consistent under different rotation angles; (4) Encode the key points based on the gradient information features and generate corresponding feature descriptors for application in subsequent registration operations; (5) For the feature points extracted from the infrared and visible light images, the similarity between the descriptors is used for initial registration, and the optimized RANSAC algorithm is used to verify and optimize the initial registration results to eliminate mismatched points.
2. The infrared and visible light image registration method based on the improved SIFT algorithm according to claim 1, characterized in that: The step (1) is specifically: Input original infrared image I ir (x,y) and visible light image I vi (x, y), the original images of the two modes are downsampled and Gaussian filtered to obtain images of different scales and sizes from large to small and from bottom to top; the two adjacent layers of Gaussian pyramid are subtracted to obtain Gaussian difference pyramid, which is convenient for effectively extracting stable key points; The Gaussian scale space is defined as: L(x,y,σ)=G(x,y,σ)*I(x,y), where I(x,y) represents the pixel value of the input image at (x,y), G(x,y,σ) represents the Gaussian kernel function, and the variance is σ 2 , L(x,y,σ) represents the scaled image of the original image after Gaussian filtering; Gaussian difference pyramid is the difference between two adjacent scaled images, defined as: D(x,y,σ)=[G(x,y,kσ)-G(x,y,σ)]*I(x,y)=L(x,y,kσ)-L(x,y,σ) Where k is the scale variation coefficient.
3. The infrared and visible light image registration method based on the improved SIFT algorithm according to claim 1, characterized in that: The step (2) is specifically: For scaled images, local adaptive FSAT corner detection is performed: extract the neighborhood of each pixel, calculate the statistical characteristics of the pixel brightness in the neighborhood, and calculate the dynamic threshold T = t (T max -T min ), where t is the adjustment coefficient, T max With T min It is the self-defined maximum and minimum mean of pixels in the neighborhood; the brightness of the neighborhood pixels is compared. If the absolute value of the brightness difference between at least J consecutive pixels and the central point pixel is greater than T, the pixel is marked as a key point.
4. The infrared and visible light image registration method based on the improved SIFT algorithm according to claim 3, characterized in that: The neighborhood pixels are 16 pixels on a circular window with a diameter of 7.
5. The infrared and visible light image registration method based on the improved SIFT algorithm according to claim 1, characterized in that: The step (3) is specifically: Taking the key point as the center, the 3σ window size of the pyramid image where it is located is selected as the neighborhood. Starting from a certain starting pixel point in the neighborhood, the connected pixels whose absolute value of pixel difference is within the threshold M are gradually marked as the same area, and the area is recorded as the connected area. For the points in the neighborhood, only the gradient direction and gradient modulus of all pixel points outside the connected area and the outermost circle of pixel points in the connected area are calculated; After completing the gradient calculation of the key point, the gradient histogram is used to count the gradient and direction of the pixels in the field to determine the main direction of the key point; the peak direction of the histogram represents the main direction of the key point.
6. The infrared and visible light image registration method based on the improved SIFT algorithm according to claim 5, characterized in that: The gradient histogram divides the direction of 0-360° into 10 columns, each column represents 36°, and the peak direction of the histogram represents the main direction of the key point.
7. The infrared and visible light image registration method based on the improved SIFT algorithm according to claim 1, characterized in that: The step (4) is specifically: The neighborhood near the key point is divided into 4*4 sub-regions, each sub-region is used as a seed point, and each seed point takes the main direction of the feature point as the reference direction. The angle between its gradient direction and the reference direction is calculated and evenly distributed to 8 directions with an interval of 45° within the range of 0-360°. When calculating the gradient direction and gradient modulus of the pixel point; in order to highlight that the pixel at the center of the sub-region contributes more to the gradient description, the Gaussian kernel is used to weight the gradient amplitude of each pixel. The Gaussian weight is: Among them, (x c ,y c ) is the sub-region center, σ controls the weighted distribution range; the weighted gradient amplitude is: G′=G·w(x,y). After obtaining the optimized SIFT feature descriptor, it is normalized to reduce the influence of illumination and enhance robustness.
8. The infrared and visible light image registration method based on the improved SIFT algorithm according to claim 1, characterized in that: The step (5) is specifically as follows: Key point descriptor R extracted from infrared image i =(r i1 ,r i2 ,…,r i128 ), the key point descriptor V extracted from the visible light image i =(v i1 ,v i2 ,…,v i128 ), then the Euclidean distance between two points is: For a descriptor r1 in an infrared image, find the two closest descriptors v1 and v2 in the visible light image, with corresponding distances d1 and d2, and calculate the ratio If r is less than a certain threshold, r1 is considered to match v1, otherwise it is discarded; After the preliminary matching, the wrong matching points are removed: the Euclidean distances of all matching point pairs are arranged in ascending order, four of the N points with the highest scores are randomly selected as inliers to build the transformation model, and the remaining matching point pairs are selected to calculate the error. Matching points less than the threshold are regarded as inliers, otherwise they are regarded as outliers and removed; For each calculated transformation model, the number of inliers is recorded, and the inlier set is continuously updated by repeating the above steps. After multiple iterations, the transformation model with the largest number of inliers is selected as the final estimated transformation model.
Citation Information
Cited By
Purple sand ware feature acquisition method and device based on image processing, and electronic equipment
CN120852797A
Power distribution cabinet safety monitoring method and system based on image recognition
CN121033052A
Feature extraction description method and system for fire scene image
CN121686000A
A steel tower precision positioning method and system
CN122550706A