Infrared image and visible light image matching method and device, medium and equipment
The method enhances image alignment by using deep learning and polynomial fitting to uniformly distribute feature points, addressing noise and distortion issues in infrared and visible light image matching for accurate fault detection in power line inspections.
Patent Information
- Application Number
- CN202510394585.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-15
AI Technical Summary
The prior art cannot accurately match the abnormal points in infrared images into visible light images, and there are problems such as noise, influence of light changes, difficulty in processing complex geometric distortions, unsupported traditional feature descriptors, sensitive modal changes, high cost and slow deep learning.
The RANSAC fitting technology and preset thresholds are used to filter the point pairs of the same name, combined with infrared image size and grid division method, distribution adjustment is performed, and then the mapping relationship is established through bidirectional polynomial fitting and adjacent interpolation methods to achieve high-precision matching of infrared and visible images.
It improves the accuracy and robustness of feature matching, handles complex nonlinear deformation, significantly improves the accuracy and reliability of image registration, and can more accurately locate and diagnose equipment failures.
Smart Images

Figure CN120318278A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of infrared image and visible light image matching, and particularly to a method, device, medium and equipment for infrared image and visible light image matching. Background Technique
[0002] Cross-modal image registration technology is crucial in scenarios such as power drone inspection. By aligning infrared images with visible light images, it enables precise positioning of temperature anomaly points and improves the accuracy of fault diagnosis. However, there are many challenges in the existing technology. Region-based registration algorithms are vulnerable to noise and illumination changes and are difficult to handle complex geometric distortions. Feature-based registration algorithms can handle large-scale geometric transformations, but in cross-modal matching of infrared and visible light images, traditional methods such as SIFT, SURF and other feature descriptors are not robust and are prone to failure, while fast algorithms such as ORB, BRISK are sensitive to modal changes and have a low feature matching rate. Deep learning-based registration algorithms have strong robustness, but require a large amount of labeled data and computing resources, with high costs and slow speeds. In addition, common linear transformation models cannot handle non-linear deformations, resulting in insufficient registration accuracy. These problems cause the existing technology to be unable to accurately match the anomaly points in the infrared image to the visible light image. Summary of the Invention
[0003] The present invention provides a method, device, medium and equipment for infrared image and visible light image matching to solve the problem in the existing technology that the anomaly points in the infrared image cannot be accurately matched to the visible light image.
[0004] In a first aspect, the present application provides a method for infrared image and visible light image matching, including:
[0005] Obtaining a preset infrared image and a preset visible light image;
[0006] Performing feature matching on the infrared image and the visible light image to obtain respective first corresponding point pairs in which each corresponding point in the infrared image and the visible light image is matched;
[0007] According to the RANSAC fitting technique and a preset threshold, screening the respective first corresponding point pairs to obtain respective second corresponding point pairs;
[0008] According to the size of the infrared image and a preset grid division method, adjusting the distribution of each corresponding point of each second corresponding point pair in the infrared image to obtain respective evenly distributed corresponding points in the infrared image;
[0009] Establish a mapping relationship between the infrared image and the visible light image based on the bidirectional polynomial fitting, the respective uniformly distributed corresponding points of the same name, and the respective second corresponding point pairs of the same name, and obtain the polynomial coefficients of the mapping relationship;
[0010] Map the edge box in the visible light image to the infrared image containing the respective uniformly distributed corresponding points of the same name according to the mapping relationship, the polynomial coefficients, and the preset neighborhood interpolation method to complete the matching.
[0011] This application obtains a preset infrared image and a visible light image, inputs them into a preset feature matching network, uses deep learning technology to extract and match the feature points in the image pair, and generates the first corresponding point pairs of the same name. Then, these corresponding point pairs of the same name are screened by the RANSAC fitting technology and a preset threshold to remove the mis-matched points and obtain more accurate second corresponding point pairs of the same name. Next, according to the size of the infrared image and the preset grid division method, these corresponding point pairs of the same name are adjusted in distribution to make them evenly distributed in the image and reduce the deformation caused by local optimization. After that, a mapping relationship between the infrared image and the visible light image is established by using bidirectional polynomial fitting and these uniformly distributed corresponding points of the same name, and the polynomial coefficients of the mapping relationship are obtained. Finally, through the preset neighborhood interpolation method, the edge box in the visible light image is mapped to the infrared image to generate a visible light image matched with the infrared image. This process not only improves the accuracy and robustness of feature matching, but also processes complex non-linear deformations through polynomial fitting, significantly improving the accuracy and reliability of image registration. Therefore, in application scenarios such as power drone patrol, equipment faults can be more accurately located and diagnosed. This application solves the problem in the prior art that the abnormal points in the infrared image cannot be accurately matched to the visible light image.
[0012] As a preferred embodiment of the first aspect, the screening of the respective first corresponding point pairs of the same name according to the RANSAC fitting technology and the preset threshold to obtain the respective second corresponding point pairs of the same name is specifically:
[0013] Screen the respective first corresponding point pairs of the same name according to the RANSAC fitting technology and the preset maximum number of iterations to obtain the respective confidence values of the respective first corresponding point pairs of the same name;
[0014] Retain the corresponding point pairs of the same name whose confidence values of the respective first corresponding point pairs of the same name are greater than the preset threshold to obtain the respective second corresponding point pairs of the same name.
[0015] In this preferred embodiment, the present application screens the initial corresponding point pairs by adopting the RANSAC fitting technique and a preset threshold, effectively improving the accuracy and reliability of the corresponding point pairs. Specifically, the RANSAC fitting technique is used in combination with a preset maximum number of iterations to evaluate all the corresponding point pairs, and the confidence value of each corresponding point pair is calculated. This process ensures, through random sampling and consistency testing, that the selected corresponding point pairs are highly reliable statistically. Subsequently, by retaining the corresponding point pairs with confidence values greater than the preset threshold, mis-matched point pairs are further eliminated, and more accurate second corresponding point pairs are obtained. This screening mechanism not only reduces the impact of mis-matches on the subsequent registration accuracy but also improves the robustness and efficiency of the entire image registration process, laying a solid foundation for achieving high-precision infrared and visible light image registration.
[0016] As a preferred embodiment of the first aspect, according to the size of the infrared image and a preset grid division method, the distribution of each corresponding point of each second corresponding point pair in the infrared image is adjusted to obtain evenly distributed corresponding points in the infrared image, specifically:
[0017] According to the size of the infrared image, a first space of the same size as the infrared image is divided in a preset space with a grid.
[0018] Each corresponding point in the infrared image is placed into the first space, and the number of corresponding points in each grid in the first space is calculated to obtain the average number of corresponding points in each grid in the first space.
[0019] According to the average number of corresponding points, the distribution of each corresponding point in the infrared image is adjusted to obtain evenly distributed corresponding points in the infrared image.
[0020] In this preferred embodiment, the present application realizes the even distribution of corresponding points in the infrared image by adjusting the distribution of corresponding points according to the size of the infrared image and a preset grid division method. First, according to the size of the infrared image, a first space of the same size as it is divided in a preset grid space. Then, the corresponding points in the infrared image are placed in this space, and the number of corresponding points in each grid is calculated to obtain the average value. Based on this average value, the corresponding points are redistributed to ensure that they are evenly spread in the infrared image. This process effectively avoids the problem of local optimization caused by the excessive concentration of corresponding points in a certain area, thereby improving the globality and accuracy of the subsequent polynomial fitting and enhancing the overall effect and reliability of image registration.
[0021] As a preferred embodiment of the first aspect, establishing a mapping relationship between the infrared image and the visible light image based on the bidirectional polynomial fitting, the respective evenly distributed corresponding points, and the respective second corresponding point pairs to obtain polynomial coefficients with a mapping relationship, specifically:
[0022] Calculate the mapping relationship from the visible light image to the infrared image and the mapping relationship from the infrared image to the visible light image based on the bidirectional polynomial fitting, the respective evenly distributed corresponding points, and the respective second corresponding point pairs to obtain two sets of polynomial coefficients;
[0023] Optimize the two sets of polynomial coefficients according to the least squares method to obtain polynomial coefficients with a mapping relationship.
[0024] In this preferred embodiment, the present application calculates the mapping relationships from the visible light image to the infrared image and from the infrared image to the visible light image through the bidirectional polynomial fitting technology, in combination with the evenly distributed corresponding points and the selected second corresponding point pairs, thereby obtaining two sets of polynomial coefficients. Further, the least squares method is used to optimize these two sets of polynomial coefficients to minimize the sum of squared residuals between the corresponding point pairs, ensuring that the obtained polynomial coefficients have the best fitting effect. This process not only improves the accuracy and robustness of the mapping relationship but also effectively handles the complex non-linear deformation between the infrared and visible light images, thereby significantly enhancing the accuracy and reliability of image registration and providing a solid technical guarantee for achieving high-precision cross-modal image registration.
[0025] As a preferred embodiment of the first aspect, mapping the edge frame in the visible light image to the infrared image containing the respective evenly distributed corresponding points according to the mapping relationship, the polynomial coefficients, and a preset neighbor interpolation method to complete the matching, specifically:
[0026] Map the first coordinates of each pixel point in the visible light image to the infrared image containing the respective evenly distributed corresponding points according to the mapping relationship and the polynomial coefficients;
[0027] Convert the respective first coordinates to respective integer coordinates according to the preset neighbor interpolation method;
[0028] Screen out the mapped coordinates within a preset range in the infrared image containing the respective evenly distributed corresponding points according to the respective integer coordinates;
[0029] Assign the pixel values at the mapped coordinates in the infrared image containing the respective evenly distributed corresponding points to each pixel point in the visible light image to generate a visible light image matched with the infrared image.
[0030] In this preferred embodiment, the present application achieves a high-precision mapping from the edge boxes in the visible light image to the infrared image through an accurate mapping relationship and polynomial coefficients, in combination with a preset neighborhood interpolation method. First, using the mapping relationship and polynomial coefficients, the first coordinates of each pixel point in the visible light image are mapped to the infrared image. Then, through the preset neighborhood interpolation method, these floating-point coordinates are converted into integer coordinates, ensuring the validity of the coordinates in the image grid. Next, the mapped coordinates within the preset range of the infrared image are filtered out, further improving the accuracy of the mapping. Finally, the pixel values in the infrared image corresponding to these valid mapped coordinates are assigned to the corresponding pixel points in the visible light image, generating a visible light image matched with the infrared image. This process not only improves the accuracy of image registration but also ensures the robustness and reliability of the mapping through filtering and interpolation methods, thus significantly enhancing the overall effect of cross-modal image registration and providing more accurate fault diagnosis support for application scenarios such as power drone inspections.
[0031] In a second aspect, the present application provides a matching device for infrared images and visible light images. The matching device for infrared images and visible light images includes an acquisition module, an input / output module, a filtering module, an adjustment module, and a mapping module;
[0032] The acquisition module is used to acquire a preset infrared image and a preset visible light image;
[0033] The input / output module is used to perform feature matching on the infrared image and the visible light image to obtain each first corresponding point pair in which the corresponding points between the infrared image and the visible light image match;
[0034] The filtering module is used to filter each of the first corresponding point pairs according to the RANSAC fitting technique and a preset threshold to obtain each second corresponding point pair;
[0035] The adjustment module is used to adjust the distribution of each corresponding point in each second corresponding point pair in the infrared image according to the size of the infrared image and a preset grid division method to obtain each evenly distributed corresponding point in the infrared image;
[0036] The mapping module is used to establish a mapping relationship between the infrared image and the visible light image according to the bi-directional polynomial fitting, each of the evenly distributed corresponding points, and each of the second corresponding point pairs, and obtain polynomial coefficients with a mapping relationship;
[0037] According to the mapping relationship, the polynomial coefficients, and a preset neighborhood interpolation method, map the edge boxes in the visible light image to the infrared image containing each evenly distributed corresponding point to complete the matching.
[0038] The device uses five modules to divide the work and coordinate with each other, which can more accurately match the infrared image and the visible light image. This application obtains the preset infrared image and visible light image, and inputs them into the preset feature matching network. Using deep learning technology, it extracts and matches the feature points in the image pair to generate the first corresponding point pair. Then, through the RANSAC fitting technology and the preset threshold, these corresponding point pairs are screened to remove the mis-matched points, and more accurate second corresponding point pairs are obtained. Then, according to the size of the infrared image and the preset grid division method, these corresponding point pairs are adjusted in distribution to make them evenly distributed in the image, reducing the deformation caused by local optimization. After that, using the bidirectional polynomial fitting and these evenly distributed corresponding points, the mapping relationship between the infrared image and the visible light image is established, and the polynomial coefficients of the mapping relationship are obtained. Finally, through the preset neighborhood interpolation method, the edge box in the visible light image is mapped into the infrared image to generate the visible light image matched with the infrared image. This process not only improves the accuracy and robustness of feature matching, but also processes complex non-linear deformations through polynomial fitting, significantly improving the accuracy and reliability of image registration. Thus, in application scenarios such as power drone inspection, it can more accurately locate and diagnose equipment failures. This application solves the problem in the prior art that the abnormal points in the infrared image cannot be accurately matched to the visible light image.
[0039] As a preferred embodiment of the second aspect, the step of screening each of the first corresponding point pairs according to the RANSAC fitting technology and the preset threshold to obtain each of the second corresponding point pairs is specifically as follows:
[0040] According to the RANSAC fitting technology and the preset maximum number of iterations, each of the first corresponding point pairs is screened to obtain the confidence values of each of the first corresponding point pairs;
[0041] Each of the corresponding point pairs with a confidence value greater than the preset threshold among each of the first corresponding point pairs is retained to obtain each of the second corresponding point pairs.
[0042] In this preferred embodiment, the present application screens the initial corresponding point pairs by adopting the RANSAC fitting technique and a preset threshold, effectively improving the accuracy and reliability of the corresponding point pairs. Specifically, the RANSAC fitting technique is used in combination with a preset maximum number of iterations to evaluate all the corresponding point pairs, and the confidence value of each corresponding point pair is calculated. This process ensures, through random sampling and consistency testing, that the selected corresponding point pairs are highly reliable statistically. Subsequently, by retaining the corresponding point pairs with confidence values greater than the preset threshold, mis-matched point pairs are further eliminated, and more accurate second corresponding point pairs are obtained. This screening mechanism not only reduces the impact of mis-matches on the subsequent registration accuracy but also improves the robustness and efficiency of the entire image registration process, laying a solid foundation for achieving high-precision infrared and visible light image registration.
[0043] As a preferred embodiment of the second aspect, according to the size of the infrared image and a preset grid division method, the distribution of each corresponding point of each second corresponding point pair in the infrared image is adjusted to obtain evenly distributed corresponding points in the infrared image, specifically as follows:
[0044] According to the size of the infrared image, a first space of the same size as the infrared image is divided in a preset space with grids.
[0045] Each corresponding point in the infrared image is placed into the first space, and the number of corresponding points in each grid in the first space is calculated to obtain the average number of corresponding points in each grid in the first space.
[0046] According to the average number of corresponding points, the distribution of each corresponding point in the infrared image is adjusted to obtain evenly distributed corresponding points in the infrared image.
[0047] In this preferred embodiment, the present application realizes the even distribution of corresponding points in the infrared image by adjusting the distribution of corresponding points according to the size of the infrared image and a preset grid division method. First, according to the size of the infrared image, a first space of the same size is divided in a preset grid-based space. Then, the corresponding points in the infrared image are placed in this space, and the number of corresponding points in each grid is calculated to obtain the average value. Based on this average value, the corresponding points are redistributed to ensure their even dispersion in the infrared image. This process effectively avoids the problem of local optimization caused by the excessive concentration of corresponding points in a certain area, thereby improving the globality and accuracy of the subsequent polynomial fitting and enhancing the overall effect and reliability of image registration.
[0048] In a third aspect, the present application provides a computer-readable storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the infrared image and visible light image matching method as described above. Its beneficial effects are the same as those of the infrared image and visible light image matching method provided in the first aspect of the present application.
[0049] In a fourth aspect, the present application provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements any one of the infrared image and visible light image matching methods as described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 : It is a schematic flowchart of an embodiment of the infrared image and visible light image matching method provided by the present application;
[0051] Figure 2 : It is a schematic structural diagram of an embodiment of the feature matching network provided by the present application;
[0052] Figure 3 : It is a schematic structural diagram of an embodiment of the infrared image and visible light image matching process provided by the present application;
[0053] Figure 4 : It is a schematic structural diagram of an embodiment of the infrared image and visible light image matching device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0055] Embodiment 1
[0056] Please refer to Figure 1 , which is an infrared image and visible light image matching method provided by an embodiment of the present invention.
[0057] In this embodiment, the process of the infrared image and visible light image matching method in the present application is described in detail through steps S01 - S06.
[0058] S01: Obtain a preset infrared image and a preset visible light image.
[0059] S02: Perform feature matching on the infrared image and the visible light image to obtain each first corresponding point pair where corresponding points in the infrared image and the visible light image match each other.
[0060] As a preferred embodiment of Embodiment 1, the performing feature matching on the infrared image and the visible light image to obtain each first corresponding point pair where corresponding points in the infrared image and the visible light image match each other is specifically as follows:
[0061] As Figure 2 shown, input the infrared and visible light image sets, and through a convolutional network, downsample to scales of 1 / 8, 1 / 4, and 1 / 2; then upsample through a feature pyramid to feature maps at scales of 1 / 2, 1 / 4, and 1 / 8. Coarse registration is performed on the feature maps at the 1 / 8 scale. First, expand it into 2D, and through a self-attention layer and a cross-attention layer, obtain the feature map Then, calculate the similarity matrix S using the following formula:
[0062]
[0063] where Linear(·) is a linear layer, <·,·> represents the inner product. τ is a scaling factor used to control the similarity score.
[0064] Next, perform a softmax operation on each row (i) of the similarity matrix S to obtain the probability value related to i (i.e., among all possible j, softmax selects the most likely match):
[0065] P 0 (i,j) = Softmax(S(i,·)) j
[0066] Similarly, perform a softmax operation on each column (j) of the similarity matrix S to obtain the probability value related to j:
[0067] P 1 (i,j) = Softmax(S(·,j)) i
[0068] The above is the introduced one-to-many and many-to-one matching mechanisms, which improve the robustness of the matching and obtain the matching probabilities between all feature points of the infrared image and the visible light image.
[0069] Since the infrared image and the visible light image perceive different spectra, the feature points show obvious differences in the two images. This operation increases the flexibility and robustness of the matching by allowing multiple potential matching points.
[0070] Next, screen out the final coarse matching pairs from these two coarse matching probability matrices Obtain the rough matching matrix M c . For the matrix P 0 , it is required to be greater than the threshold θ c and be the maximum value in row i; for the matrix P 1 ,
[0071]
[0072] it is required to be greater than the threshold θ c and be the maximum value in column j:
[0074] The next step is fine matching. Based on the rough matching pairs provided by the rough matching matrix M c , crop out local windows of scales 1×1, 3×3, and 5×5 on the feature maps F 1 / 4 , F 1 / 2 respectively. These windows pass through convolutional layers respectively, and then perform self-attention and cross-attention operations at the local window level to obtain local feature maps and Based on the local feature maps, find the similarity matrix for each rough matching pair , and obtain the fine matching probability matrix through double
[0075]
[0076] Softmax:
[0077] Then, screen out the fine matching pairs with matching probabilities greater than the threshold θ f from the fine matching probability matrix to obtain the fine matching matrix M f :
[0078]
[0079] The last step is to splice together the feature vectors corresponding to the coordinates of the fine matching matrix M and based on the local feature maps f , and regress four sub-pixel coordinates through the MLP network and the Tanh activation function These coordinates represent fine offset amounts. Finally, by adding the coordinates of the fine matching to the regressed sub-pixel offset amounts, the positions of more accurate homologous points are obtained.
[0080] This feature matching module adapts to the conversion between infrared and visible light modalities during the deep learning process and has a global perception field of view, and thus performs well in cross-modal matching tasks.
[0081] As a preferred embodiment of Embodiment 1, before inputting the infrared image and the visible light image into the preset feature matching network, the following steps are further included:
[0082] Before inputting the image pair into the feature matching network, mask occlusion pre-training is first performed. This pre-training process uses the same network structure as the feature matching network and applies a dual-channel mask strategy. 50% of the patch mask blocks are randomly generated and applied respectively on the input infrared and visible light images. This allows the model to perform better feature learning and reconstruction through the information complementarity of the two modalities.
[0083] For the different-scale feature maps generated after passing through the CNN feature extractor, the patch masks are respectively replaced with trainable mask tokens, enabling them to be gradually optimized during the training process to learn the information of the masked areas.
[0084] The above operations use mask tokens of different scales, which can promote the model to learn features at different levels, capture local details and global context, and are very important for processing images of different resolutions.
[0085] For the mask tokens after rough matching, the same method as in the fine matching stage is used to obtain the local feature maps and Then, through linear mapping, it is enlarged to an image resolution of 10×10 to correspond to the resolution of the original image and filled back into the original pixel space. In this way, the masked areas of the entire image can be reconstructed to form the final complete image.
[0086] During the learning process, the mean square error (MSE) is used as the loss function to measure the difference between the reconstructed image and the original image.
[0087] Generally speaking, mask occlusion pre-training can help the model better familiarize itself with the non-linear intensity differences between visible light and thermal infrared images, and learn how to find complementary relationships between features in multi-modal inputs. By directly using the hyperparameters and weights learned in the pre-training stage for the feature matching network, these capabilities are directly transferred to the refinement task, improving the feature matching performance of the model.
[0088] S03: According to the RANSAC fitting technique and the preset threshold, screen the respective first corresponding point pairs to obtain respective second corresponding point pairs.
[0089] As a preferred embodiment of Embodiment 1, the step of screening the respective first corresponding point pairs according to the RANSAC fitting technique and the preset threshold to obtain respective second corresponding point pairs is specifically as follows:
[0090] First, based on the matched corresponding points, the fundamental matrix F is fitted using RANSAC, satisfying:
[0091]
[0092] Ensure that the obtained fundamental matrix is globally optimal through the maximum number of iterations N (default value is 10000) and high confidence (default value is 0.9999).
[0093] Based on the reprojection error threshold T (default is 0.5), set the points with reprojection error e less than T as inliers, and remove the rest as outliers:
[0094]
[0095] Next, sort the matching confidence of the corresponding point pairs and select the confidence conf_thres at the 80th percentile (default parameter in_conf = 0.2):
[0096] mconf sorted = sort(mconf)
[0097]
[0098] conf_thres = min(mconf sorted [-index:])
[0099] If it is less than the set threshold conf_thres_set (default value 0.5), the value becomes conf_thres_set, otherwise it remains unchanged. Filter out all matching points with matching confidence less than conf_thres.
[0100] Finally, filter out the image pairs with the number of corresponding point pairs less than the threshold (default parameter min_match_num is 20).
[0101] In this stage, the corresponding points obtained by matching are screened multiple times, removing the wrong point pairs, and effectively improving the accuracy of subsequent fitting.
[0102] In this preferred embodiment, the present application screens the initial corresponding point pairs by using the RANSAC fitting technique and a preset threshold, effectively improving the accuracy and reliability of the corresponding point pairs. Specifically, the RANSAC fitting technique is used in combination with a preset maximum number of iterations to evaluate all the corresponding point pairs, and the confidence value of each corresponding point pair is calculated. This process ensures, through random sampling and consistency testing, that the selected corresponding point pairs are highly reliable statistically. Subsequently, by retaining the corresponding point pairs with confidence values greater than the preset threshold, mis-matched point pairs are further eliminated, and more accurate second corresponding point pairs are obtained. This screening mechanism not only reduces the impact of mis-matches on the subsequent registration accuracy but also improves the robustness and efficiency of the entire image registration process, laying a solid foundation for achieving high-precision infrared and visible light image registration.
[0103] S04: According to the size of the infrared image and a preset grid division method, adjust the distribution of each corresponding point of each second corresponding point pair in the infrared image to obtain evenly distributed corresponding points in the infrared image.
[0104] As a preferred embodiment of Embodiment 1, the adjusting the distribution of each corresponding point of each second corresponding point pair in the infrared image according to the size of the infrared image and a preset grid division method to obtain evenly distributed corresponding points in the infrared image is specifically as follows:
[0105] In an image space with the same size as the thermal infrared image, divide the space into square grids, and the total number of grids is y. Calculate the number x of all the selected corresponding point pairs, and place all the corresponding points in the thermal infrared image into this space. Calculate the average number of corresponding points in each grid For the grids containing more corresponding points than the average randomly delete the corresponding points therein until the number of corresponding points therein is equal to
[0106] In this way, the distribution of corresponding points in the image is balanced, thereby reducing local deformation in the polynomial fitting process and improving the registration accuracy.
[0107] In this preferred embodiment, the present application adjusts the distribution of corresponding points according to the size of the infrared image and a preset grid division method, achieving a uniform distribution of corresponding points in the infrared image. First, according to the size of the infrared image, a first space of the same size is divided in the preset grid space. Then, the corresponding points in the infrared image are placed in this space, and the number of corresponding points in each grid is calculated to obtain an average value. Based on this average value, the corresponding points are redistributed to ensure their uniform dispersion in the infrared image. This process effectively avoids the problem of local optimization caused by the over-concentration of corresponding points in a certain area, thereby improving the global nature and accuracy of subsequent polynomial fitting and enhancing the overall effect and reliability of image registration.
[0108] S05: Establish a mapping relationship between the infrared image and the visible light image based on the bidirectional polynomial fitting, the uniformly distributed corresponding points, and the pairs of second corresponding points, and obtain the polynomial coefficients with the mapping relationship.
[0109] S06: Map the edge frame in the visible light image to the infrared image containing the uniformly distributed corresponding points according to the mapping relationship, the polynomial coefficients, and a preset neighborhood interpolation method to complete the matching.
[0110] As a preferred embodiment of Embodiment 1, the establishment of the mapping relationship between the infrared image and the visible light image based on the bidirectional polynomial fitting, the uniformly distributed corresponding points, and the pairs of second corresponding points, obtaining the polynomial coefficients with the mapping relationship, and mapping the edge frame in the visible light image to the infrared image containing the uniformly distributed corresponding points according to the mapping relationship, the polynomial coefficients, and a preset neighborhood interpolation method to complete the matching are specifically as follows:
[0111] Perform bidirectional polynomial fitting based on the matched and filtered corresponding points to find the mapping relationship from visible light coordinates to thermal infrared coordinates and the mapping relationship from thermal infrared coordinates to visible light coordinates, and obtain two corresponding sets of polynomial coefficients.
[0112] According to the fitting coefficients for converting thermal infrared coordinates to the visible light coordinate system, obtain the coordinates of each thermal infrared grid coordinate mapped to the visible light coordinate system; then perform resampling, and replace the pixel values of the corresponding visible light coordinates into the thermal infrared grid to obtain the image of the visible light image resampled to the thermal infrared coordinate system.
[0113] Perform a reverse coordinate remapping once. According to the fitting coefficients for converting visible light coordinates to the thermal infrared coordinate system, obtain the coordinates of the contour points based on the visible light coordinate system mapped to the thermal infrared coordinate system.
[0114] Take the mapping from infrared to visible light images as an example. First, perform polynomial fitting. Assume that the mapping relationship from visible light images to infrared images is a third-order polynomial. Substitute the coordinates of homologous points \(\{(x,y)\}\) in the visible light image into the following formula:
[0115]
[0116]
[0117] where are the estimated coordinates of homologous points in the infrared image obtained by mapping, and \(a_0,a_1,\cdots,a_9\) and \(b_0,b_1,\cdots,b_9\) are the coefficients to be determined.
[0118] Next, use the least squares method to solve this linear system. The following is the formula for the sum of squared residuals:
[0119]
[0120] where \(u\) i , \(v\) i are the coordinates of the corresponding homologous points in the real infrared image.
[0121] By minimizing the sum of squared residuals, find the optimal third-order polynomial fitting coefficients.
[0122] Next, perform resampling. Based on the fitting coefficients and the fitting polynomial, map the coordinates of each pixel point in the visible light image to the corresponding coordinates in the infrared image, and then obtain integer coordinates through nearest-neighbor interpolation. Then filter out the mapped coordinates that are not within the valid range of the infrared image. Finally, assign the pixels at the corresponding coordinates in the infrared image to the corresponding coordinates in the visible light coordinate system, that is, map the infrared image to the coordinate system of the visible light image.
[0123] To obtain the corresponding coordinates of the contour points based on the visible light coordinate system mapped to the infrared coordinate system, reverse polynomial fitting and coordinate remapping are also performed, and the logic is the same as above.
[0124] Since infrared and visible light images have different imaging characteristics, the mapping relationship between them may be non-linear. Using non-linear polynomial fitting can better capture these complex relationships, thereby improving the registration accuracy.
[0125] In this preferred embodiment, the present application uses a bidirectional polynomial fitting technique to calculate the mapping relationships from the visible light image to the infrared image and from the infrared image to the visible light image by combining evenly distributed corresponding points of the same name and the selected second corresponding point pairs of the same name, thereby obtaining two sets of polynomial coefficients. Further, the least squares method is used to optimize these two sets of polynomial coefficients to minimize the sum of the squared residuals between the corresponding point pairs, ensuring that the obtained polynomial coefficients have the best fitting effect. This process not only improves the accuracy and robustness of the mapping relationship, but also effectively handles the complex non-linear deformation between the infrared and visible light images, thus significantly improving the accuracy and reliability of image registration and providing a solid technical guarantee for achieving high-precision cross-modal image registration.
[0126] The flow of the entire technical solution of the present application is as Figure 3 shown. The present application acquires a preset infrared image and a visible light image and inputs them into a preset feature matching network. Deep learning technology is used to extract and match the feature points in the image pair to generate the first corresponding point pairs of the same name. Then, these corresponding point pairs of the same name are screened by the RANSAC fitting technique and a preset threshold to remove the mismatched points and obtain more accurate second corresponding point pairs of the same name. Next, according to the size of the infrared image and a preset grid division method, these corresponding point pairs of the same name are redistributed to make them evenly distributed in the image, reducing the deformation caused by local optimization. After that, a mapping relationship between the infrared image and the visible light image is established using bidirectional polynomial fitting and these evenly distributed corresponding points of the same name to obtain the polynomial coefficients of the mapping relationship. Finally, through a preset neighborhood interpolation method, the edge box in the visible light image is mapped into the infrared image to generate a visible light image that matches the infrared image. This process not only improves the accuracy and robustness of feature matching, but also processes complex non-linear deformations through polynomial fitting, significantly improving the accuracy and reliability of image registration. Thus, in application scenarios such as power drone inspections, equipment faults can be more accurately located and diagnosed. The present application solves the problem in the prior art that the abnormal points in the infrared image cannot be accurately matched to the visible light image.
[0127] Embodiment 2
[0128] Please refer to Figure 4 , which is a matching device for infrared images and visible light images provided by an embodiment of the present application.
[0129] In this embodiment, the matching device for infrared images and visible light images includes an acquisition module 10, an input / output module 20, a screening module 30, an adjustment module 40, and a mapping module 50.
[0130] The acquisition module 10 is used to acquire a preset infrared image and a preset visible light image.
[0131] The input / output module 20 is used to perform feature matching on the infrared image and the visible light image to obtain each first corresponding point pair in which the corresponding points between the infrared image and the visible light image are matched.
[0132] As a preferred embodiment of the second embodiment, the performing feature matching on the infrared image and the visible light image to obtain each first corresponding point pair in which the corresponding points between the infrared image and the visible light image are matched is specifically as follows:
[0133] As Figure 2 shown, the infrared and visible light image sets are input, passed through a convolutional network, and downsampled to 1 / 8, 1 / 4, and 1 / 2 scales; then upsampled through a feature pyramid to feature maps of 1 / 2, 1 / 4, and 1 / 8 scales. Coarse registration is performed on the feature maps of the 1 / 8 scale. First, it is unfolded into 2D, and through the self-attention layer and the cross-attention layer, the feature map Then, the similarity matrix S is calculated using the following formula:
[0134]
[0135] where Linear(·) is a linear layer, <·,·> represents the inner product. τ is a scaling factor used to control the similarity score.
[0136] Next, a softmax operation is performed on each row (i) of the similarity matrix S to obtain the probability value related to i (i.e., among all possible j, softmax selects the most likely match):
[0137] P 0 (i,j) = Softmax(S(i,·)) j
[0138] Similarly, a softmax operation is performed on each column (j) of the similarity matrix S to obtain the probability value related to j:
[0139] P 1 (i,j) = Softmax(S(·,j)) i
[0140] The above is the introduced one-to-many and many-to-one matching mechanisms, which improve the robustness of the matching and obtain the matching probabilities between all the feature points of the infrared image and the visible light image.
[0141] Since the infrared image and the visible light image perceive different spectra, the feature points show obvious differences in the two images. This operation increases the flexibility and robustness of the matching by allowing multiple potential matching points.
[0142] Next, the final coarse matching pairs are screened out from these two coarse matching probability matrices Obtain the rough matching matrix M c For matrix P 0 it is required to be greater than the threshold θ c and be the maximum value in row i; for matrix P 1
[0143]
[0144] it is required to be greater than the threshold θ c and be the maximum value in column j:
[0145] The next step is fine matching. Based on the rough matching pairs provided by the rough matching matrix M c crop local windows of scales 1×1, 3×3, and 5×5 on the feature maps F 1 / 4 and F 1 / 2 respectively. These windows pass through convolutional layers and then perform self-attention and cross-attention operations at the local window level to obtain local feature maps and Based on the local feature maps, find the similarity matrix for each rough matching pair and obtain the fine matching probability matrix through double Softmax:
[0146]
[0147] Then, screen out the fine matching pairs with matching probabilities greater than the threshold θ f from the fine matching probability matrix to obtain the fine matching matrix M f :
[0148]
[0149] The last step is to splice together the feature vectors corresponding to the coordinates of the fine matching matrix M and based on the local feature maps f and regress four sub-pixel coordinates through an MLP network and a Tanh activation function These coordinates represent fine offsets. Finally, by adding the coordinates of the fine matching to the regressed sub-pixel offsets, the position of the more accurate homologous points is obtained.
[0150] This feature matching module adapts to the conversion between infrared and visible light modalities during the deep learning process and has a global perception field of view, and thus performs well in cross-modal matching tasks.
[0151] As a preferred embodiment of the second embodiment, before inputting the infrared image and the visible light image into the preset feature matching network, it further includes:
[0152] Before inputting the image pair into the feature matching network, mask occlusion pre-training is first carried out. This pre-training process uses the same network structure as the feature matching network and applies a dual-channel mask strategy. 50% of the patch mask blocks are randomly generated and applied to the input infrared and visible light images respectively. This allows the model to perform better feature learning and reconstruction through the complementary information of the two modalities.
[0153] For the different-scale feature maps generated after passing through the CNN feature extractor, the patch masks are respectively replaced with trainable mask tokens, enabling them to be gradually optimized during the training process to learn the information of the masked regions.
[0154] The above operations use mask tokens of different scales, which can promote the model to learn features at different levels, capture local details and global context, and are very important for processing images of different resolutions.
[0155] For the mask tokens after coarse matching, the local feature maps are obtained using the same method as in the fine matching stage. and Then, through linear mapping, it is enlarged to an image resolution of 10×10 to correspond to the resolution of the original image and filled back into the original pixel space. In this way, the masked regions of the entire image can be reconstructed to form the final complete image.
[0156] During the learning process, the mean squared error (MSE) is used as the loss function to measure the difference between the reconstructed image and the original image.
[0157] Generally speaking, mask occlusion pre-training can help the model better familiarize itself with the non-linear intensity differences between visible light and thermal infrared images, and learn how to find complementary relationships between features in multi-modal inputs. By directly using the hyperparameters and weights learned in the pre-training stage for the feature matching network, these capabilities are directly transferred to the refinement task, improving the feature matching performance of the model.
[0158] The screening module 30 is used to screen the respective first corresponding point pairs according to the RANSAC fitting technique and a preset threshold to obtain respective second corresponding point pairs.
[0159] As a preferred embodiment of the second embodiment, the screening of the respective first corresponding point pairs according to the RANSAC fitting technique and the preset threshold to obtain respective second corresponding point pairs is specifically as follows:
[0160] First, based on the matched corresponding points, the fundamental matrix F is fitted by RANSAC, satisfying:
[0161]
[0162] Ensure that the obtained fundamental matrix is globally optimal through the maximum number of iterations N (default value is 10,000) and high confidence (default value is 0.9999).
[0163] Based on the reprojection error threshold T (default is 0.5), set the points with reprojection error e less than T as inliers, and remove the rest as outliers:
[0164]
[0165] Next, sort the matching confidence of corresponding point pairs and select the confidence conf_thres at the 80th percentile (default parameter in_conf = 0.2):
[0166] mconf sorted = sort(mconf)
[0167]
[0168] conf_thres = min(mconf sorted [-index:])
[0169] If it is less than the set threshold conf_thres_set (default value 0.5), the value becomes conf_thres_set, otherwise it remains unchanged. Filter out all matching points with matching confidence less than conf_thres.
[0170] Finally, filter out the image pairs with the number of corresponding point pairs less than the threshold (default parameter min_match_num is 20).
[0171] In this stage, the corresponding points obtained by matching are screened multiple times, removing the wrong point pairs, and effectively improving the accuracy of subsequent fitting.
[0172] In this preferred embodiment, the present application screens the initial corresponding point pairs by adopting the RANSAC fitting technique and preset thresholds, effectively improving the accuracy and reliability of the corresponding point pairs. Specifically, using the RANSAC fitting technique combined with the preset maximum number of iterations, all corresponding point pairs are evaluated, and the confidence value of each corresponding point pair is calculated. This process ensures that the selected corresponding point pairs are highly reliable statistically through random sampling and consistency checking. Subsequently, by retaining the corresponding point pairs with confidence values greater than the preset threshold, the mis-matched point pairs are further removed, obtaining more accurate second corresponding point pairs. This screening mechanism not only reduces the impact of mis-matches on the subsequent registration accuracy, but also improves the robustness and efficiency of the entire image registration process, laying a solid foundation for achieving high-precision infrared and visible light image registration.
[0173] The adjustment module 40 is configured to adjust the distribution of each corresponding point in each second corresponding point pair in the infrared image according to the size of the infrared image and a preset grid division method, so as to obtain each evenly distributed corresponding point in the infrared image.
[0174] As a preferred embodiment of the second embodiment, the adjusting the distribution of each corresponding point in each second corresponding point pair in the infrared image according to the size of the infrared image and a preset grid division method to obtain each evenly distributed corresponding point in the infrared image is specifically as follows:
[0175] In an image space having the same size as the thermal infrared image, the space is divided into square grids, and the total number of grids is y. Calculate the number x of all selected corresponding point pairs, and put the corresponding points in all thermal infrared images into this space. Calculate the average number of corresponding points in each grid For the grids whose number of corresponding points included exceeds the average value randomly delete the corresponding points therein until the number of corresponding points therein is equal to
[0176] In this way, the distribution of corresponding points in the image is balanced, thereby reducing local deformation in the polynomial fitting process and improving the registration accuracy.
[0177] In this preferred embodiment, the present application realizes the uniform distribution of corresponding points in the infrared image by adjusting the distribution of corresponding points according to the size of the infrared image and a preset grid division method. First, according to the size of the infrared image, a first space of the same size is divided in a preset grid space. Then, the corresponding points in the infrared image are placed in this space, and the number of corresponding points in each grid is calculated to obtain the average value. Based on this average value, the corresponding points are redistributed to ensure that they are evenly scattered in the infrared image. This process effectively avoids the problem of local optimization caused by the over-concentration of corresponding points in a certain area, thereby improving the globality and accuracy of subsequent polynomial fitting, and enhancing the overall effect and reliability of image registration.
[0178] The mapping module 50 is configured to establish a mapping relationship between the infrared image and the visible light image according to the bidirectional polynomial fitting, the evenly distributed corresponding points, and the second corresponding point pairs, and obtain the polynomial coefficients with the mapping relationship.
[0179] The mapping module 50 is further configured to map the edge frame in the visible light image to the infrared image containing the evenly distributed corresponding points according to the mapping relationship, the polynomial coefficients, and a preset neighborhood interpolation method, so as to complete the matching.
[0180] As a preferred embodiment of the second embodiment, the mapping relationship between the infrared image and the visible light image is established based on the bidirectional polynomial fitting, the evenly distributed corresponding points of the same name, and the pairs of the second corresponding points of the same name, and the polynomial coefficients with the mapping relationship are obtained. According to the mapping relationship, the polynomial coefficients, and the preset neighborhood interpolation method, the edge frame in the visible light image is mapped into the infrared image containing the evenly distributed corresponding points of the same name to complete the matching. Specifically:
[0181] Perform bidirectional polynomial fitting based on the corresponding points of the same name after matching and screening, find the mapping relationship from the visible light coordinates to the thermal infrared coordinates and the mapping relationship from the thermal infrared coordinates to the visible light coordinates, and obtain two corresponding sets of polynomial coefficients.
[0182] According to the fitting coefficients from the thermal infrared coordinates to the visible light coordinate system, obtain the coordinates of each thermal infrared grid coordinate mapped to the visible light coordinate system; then perform resampling, and replace the pixel values of the corresponding visible light coordinates into the thermal infrared grid to obtain the image of the visible light image resampled to the thermal infrared coordinate system.
[0183] Perform a coordinate remapping in the reverse direction once. According to the fitting coefficients from the visible light coordinates to the thermal infrared coordinate system, obtain the coordinates of the contour point coordinates based on the visible light coordinate system mapped to the thermal infrared coordinate system.
[0184] Take the mapping from the infrared image to the visible light image as an example. First, perform polynomial fitting. Assume that the mapping relationship from the visible light image to the infrared image is a third-order polynomial, and substitute the corresponding point coordinates {(x,y)} in the visible light image into the following formula:
[0185]
[0186] where is the estimated coordinate of the corresponding point in the infrared image obtained by mapping, and a0, a1, …, a9 and b0, b1, …, b9 are the coefficients to be determined.
[0187] Next, use the least squares method to solve this linear system. The following is the formula for the sum of squared residuals:
[0188]
[0189] where u i , v i are the coordinates of the corresponding point in the real infrared image.
[0190] By minimizing the sum of squared residuals, find the optimal third-order polynomial fitting coefficients.
[0191] Next, resampling is performed. Based on the fitting coefficients and the fitting polynomial, the coordinates of each pixel in the visible light image are mapped to the corresponding coordinates in the infrared image, and then integer coordinates are obtained through nearest neighbor interpolation. Then, the mapped coordinates outside the valid range of the infrared image are filtered out. Finally, the pixels at the corresponding coordinates in the infrared image are assigned to the corresponding coordinates in the visible light coordinate system, that is, the infrared image is mapped into the coordinate system of the visible light image.
[0192] In order to obtain the corresponding coordinates of the contour points based on the visible light coordinate system mapped to the infrared coordinate system, reverse polynomial fitting and coordinate remapping are also performed, and the logic is the same as above.
[0193] Since the infrared and visible light images have different imaging characteristics, the mapping relationship between them may be non-linear. Using non-linear polynomial fitting can better capture these complex relationships, thereby improving the registration accuracy.
[0194] In this preferred embodiment, the present application calculates the mapping relationship from the visible light image to the infrared image and from the infrared image to the visible light image through a two-way polynomial fitting technique, in combination with evenly distributed corresponding points and the selected second corresponding point pairs, thereby obtaining two sets of polynomial coefficients. Further, the least squares method is used to optimize these two sets of polynomial coefficients to minimize the sum of the squared residuals between the corresponding point pairs, ensuring that the obtained polynomial coefficients have the best fitting effect. This process not only improves the accuracy and robustness of the mapping relationship, but also effectively handles the complex non-linear deformation between the infrared and visible light images, thereby significantly improving the accuracy and reliability of image registration, providing a solid technical guarantee for achieving high-precision cross-modal image registration.
[0195] The flow of the entire technical solution of the present application is as Figure 3As shown, the present application obtains a preset infrared image and a visible light image, and inputs them into a preset feature matching network. Using deep learning technology, it extracts and matches the feature points in the image pair to generate a first pair of corresponding points. Then, through the RANSAC fitting technology and a preset threshold, these pairs of corresponding points are screened to remove mis-matched points, obtaining a more accurate second pair of corresponding points. Next, according to the size of the infrared image and a preset grid division method, the distribution of these pairs of corresponding points is adjusted to make them evenly distributed in the image, reducing the distortion caused by local optimization. After that, using bidirectional polynomial fitting and these evenly distributed pairs of corresponding points, the mapping relationship between the infrared image and the visible light image is established, obtaining the polynomial coefficients of the mapping relationship. Finally, through a preset neighborhood interpolation method, the edge box in the visible light image is mapped into the infrared image, generating a visible light image that matches the infrared image. This process not only improves the accuracy and robustness of feature matching, but also processes complex non-linear deformations through polynomial fitting, significantly improving the accuracy and reliability of image registration. Thus, in application scenarios such as power drone inspections, it can more accurately locate and diagnose equipment failures. The present application solves the problem in the prior art that abnormal points in the infrared image cannot be accurately matched to the visible light image.
[0196] Embodiment 3:
[0197] The embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program. Among them, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the method for matching an infrared image and a visible light image described above;
[0198] Among them, for the method for matching an infrared image and a visible light image, if it is implemented in the form of a software functional unit and used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above embodiment methods of the present invention, it can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0199] Embodiment 4
[0200] The present application provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements any one of the infrared image and visible light image matching methods described in Embodiment 1.
[0201] For the specific embodiments described above, the objectives, technical solutions, and beneficial effects of the present invention have been further described in detail. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. In particular, for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for matching infrared images and visible light images, characterized in that, Including: Obtain a preset infrared image and a preset visible light image; Perform feature matching on the infrared image and the visible light image to obtain respective first corresponding point pairs in which corresponding points of the infrared image and the visible light image match each other; According to the RANSAC fitting technique and a preset threshold, screen the respective first corresponding point pairs to obtain respective second corresponding point pairs; According to the size of the infrared image and a preset grid division method, adjust the distribution of corresponding points of the respective second corresponding point pairs in the infrared image to obtain evenly distributed corresponding points in the infrared image; According to the bilinear polynomial fitting, the respective evenly distributed corresponding points, and the respective second corresponding point pairs, establish a mapping relationship between the infrared image and the visible light image, and obtain polynomial coefficients with the mapping relationship; According to the mapping relationship, the polynomial coefficients, and a preset adjacent interpolation method, map the edge frame in the visible light image to the infrared image containing the respective evenly distributed corresponding points to complete the matching.
2. The matching method for infrared images and visible light images according to claim 1, characterized in that The step of screening the respective first corresponding point pairs according to the RANSAC fitting technique and a preset threshold to obtain respective second corresponding point pairs specifically is: According to the RANSAC fitting technique and a preset maximum number of iterations, screen the respective first corresponding point pairs to obtain respective confidence values of the respective first corresponding point pairs; Retain the corresponding point pairs in which the confidence values of the respective first corresponding point pairs are greater than the preset threshold to obtain respective second corresponding point pairs.
3. The matching method for infrared images and visible light images according to claim 1, characterized in that The step of adjusting the distribution of corresponding points of the respective second corresponding point pairs in the infrared image according to the size of the infrared image and a preset grid division method to obtain evenly distributed corresponding points in the infrared image specifically is: According to the size of the infrared image, divide a first space with the same size as the infrared image in a preset space with grids; Put the corresponding points in the infrared image into the first space, calculate the number of corresponding points in each grid in the first space, and obtain the average number of corresponding points in each grid in the first space; According to the average number of corresponding points, adjust the distribution of the corresponding points in the infrared image to obtain evenly distributed corresponding points in the infrared image.
4. The infrared image and visible light image matching method according to claim 1, wherein The step of establishing a mapping relationship between the infrared image and the visible light image according to the bilinear polynomial fitting, the respective evenly distributed corresponding points, and the respective second corresponding point pairs to obtain polynomial coefficients with the mapping relationship specifically is: According to the bilinear polynomial fitting, the respective evenly distributed corresponding points, and the respective second corresponding point pairs, calculate the mapping relationship from the visible light image to the infrared image and the mapping relationship from the infrared image to the visible light image to obtain two sets of polynomial coefficients; Optimize the two sets of polynomial coefficients according to the least squares method to obtain polynomial coefficients with the mapping relationship.
5. The method for matching an infrared image and a visible light image according to any one of claims 1-4, characterized in that The step of mapping the edge frame in the visible light image to the infrared image containing the respective evenly distributed corresponding points according to the mapping relationship, the polynomial coefficients, and a preset adjacent interpolation method to complete the matching specifically is: Map the first coordinates of each pixel point in the visible light image to the infrared image containing uniformly distributed corresponding points according to the mapping relationship and the polynomial coefficients; Convert the first coordinates into integer coordinates according to a preset neighborhood interpolation method; Filter out the mapping coordinates within a preset range of the infrared image containing uniformly distributed corresponding points according to the integer coordinates; Assign the pixel values at the mapping coordinates in the infrared image containing uniformly distributed corresponding points to each pixel point in the visible light image to generate a visible light image matched with the infrared image.
6. An infrared image and visible light image matching device, characterized in that, It includes: An acquisition module, an input / output module, a filtering module, an adjustment module, and a mapping module; The acquisition module is used to acquire a preset infrared image and a preset visible light image; The input / output module is used to perform feature matching on the infrared image and the visible light image to obtain each first corresponding point pair where corresponding points in the infrared image and the visible light image match; The filtering module is used to filter each first corresponding point pair according to the RANSAC fitting technique and a preset threshold to obtain each second corresponding point pair; The adjustment module is used to perform distribution adjustment on the corresponding points of each second corresponding point pair in the infrared image according to the size of the infrared image and a preset grid division method to obtain uniformly distributed corresponding points in the infrared image; The mapping module is used to establish a mapping relationship between the infrared image and the visible light image according to bidirectional polynomial fitting, the uniformly distributed corresponding points, and the second corresponding point pairs, and obtain polynomial coefficients with a mapping relationship; Map the edge frame in the visible light image to the infrared image containing uniformly distributed corresponding points according to the mapping relationship, the polynomial coefficients, and a preset neighborhood interpolation method to complete the matching.
7. The matching device for infrared images and visible light images according to claim 6, characterized in that, The step of filtering each first corresponding point pair according to the RANSAC fitting technique and a preset threshold to obtain each second corresponding point pair is specifically: Filter each first corresponding point pair according to the RANSAC fitting technique and a preset maximum number of iterations to obtain the confidence values of each first corresponding point pair; Retain the corresponding point pairs with confidence values of each first corresponding point pair greater than the preset threshold to obtain each second corresponding point pair.
8. The matching device for infrared images and visible light images according to claim 6, characterized in that, The step of performing distribution adjustment on the corresponding points of each second corresponding point pair in the infrared image according to the size of the infrared image and a preset grid division method to obtain uniformly distributed corresponding points in the infrared image is specifically: Divide a first space with the same size as the infrared image in a preset space with a grid according to the size of the infrared image; Put the corresponding points in the infrared image into the first space, calculate the number of corresponding points in each grid in the first space, and obtain the average number of corresponding points in each grid in the first space; Perform distribution adjustment on the corresponding points in the infrared image according to the average number of corresponding points to obtain uniformly distributed corresponding points in the infrared image.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the method for matching infrared images and visible light images according to any one of claims 1 to 5.
10. A terminal device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the method for matching infrared images and visible light images according to any one of claims 1 to 5.
Citation Information
Cited By
Photoelectric pod multi-modal image adaptive registration method in non-feature scene
CN121120720A