Low-cost high-precision visible light and infrared ray fusion method, medium and equipment
Through residual block encoder, attention mechanism and jump connection image fusion method, the problem of information retention and artifacts in infrared and visible light image fusion is solved, and a high-precision image fusion effect is achieved.
Patent Information
- Application Number
- CN202510483223.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art is difficult to effectively retain the complementary information of the image in the fusion of infrared and visible light images, and there are problems of artifacts and information loss.
An encoder fusion framework with residual blocks is adopted, combining attention mechanism and jump connection, deep features of infrared and visible light images are extracted, and feature point matching pairing is obtained through geometric constraints to optimize the feature fusion process.
High-precision image fusion is achieved, the target information of infrared light images and the detailed information of visible light images are retained, artifacts are reduced, and the clarity and objective indicator performance of the fused image are improved.
Smart Images

Figure CN120388260A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a low-cost and high-precision visible light and infrared fusion method, medium and device. Background Art
[0002] The infrared and visible light image fusion technology has wide applications in the fields of computer vision, remote sensing, medical imaging, military detection, etc. The superiority of fusing infrared and visible light pictures lies in that: the visible light image contains rich texture details, while the infrared image can highlight the thermal target even under poor lighting conditions or severe occlusion. Due to their strong complementarity, an image containing thermal targets and rich background information can be obtained after fusion. In order to present the complementary information transmitted by different source images in one image, an image fusion method based on residual blocks is proposed, which can effectively extract features without generating artifacts, and can largely retain the information contained in the original image after fusion. The research work specifically includes the following parts: (1) Introduce the residual module into the network encoder. Add three residual blocks and skip connections during the encoding process, so that the encoder can generate richer features.
[0003] (2) Introduce the model based on the attention mechanism into the feature fusion stage. Different from the previous methods of fusing by weighted average or summation superposition, the attention mechanism model is adopted, which can better fuse the information in the infrared and visible light images.
[0004] (3) Introduce the skip connection into the network decoder. Transfer the feature maps obtained in the first and second layers to the corresponding deconvolution layers through the skip connection for processing, and better obtain the detail information of the original image, and finally obtain a fusion image with better effect. Summary of the Invention
[0005] A low-cost and high-precision visible light and infrared fusion method, device and storage medium proposed by the present invention can at least solve one of the technical problems in the background art.
[0006] To achieve the above object, the present invention adopts the following technical solutions: A low-cost and high-precision visible light and infrared fusion method, which is executed by a computer device through the following steps: S100. Obtain the feature point matching pairs of the infrared and visible light images by using the geometric constraint method; S200. Based on the feature point matching pairs, extract the deep features of the infrared and visible light images through the encoder fusion framework with residual blocks; S300. Obtain the attention map by adding the attention mechanism, and fuse the attention map with the deep features of the infrared and visible light images to obtain the feature map; The S400 passes the feature map through skip connections to the corresponding transposed convolution layer for processing to obtain a fused image.
[0007] Furthermore, the method for obtaining the feature point matching pairs of the infrared and visible light images by using the geometric constraint method in step S100 of the present invention includes: Design of geometric constraint algorithm In the algorithm, find the optimal similarity of two similar triangles and set the expression for quantifying the similarity of the two triangles:
[0008] First, calculate the ratio of the lengths of the corresponding line segments of the triangles. The length of each line segment is the Euclidean distance length in the pixel coordinate system. When designing the algorithm, set the threshold of the line segment ratio to ensure that the images are similar. The formula is as follows:
[0009] In the program, the designed threshold is 0.02. Only when both are less than the threshold and meet the conditions of the above two formulas can it be proved that the ratios of the corresponding three line segments are close to the same. If it is greater than the threshold, the triangles are judged to be dissimilar. Under this condition of ensuring similarity, use the following conditions to select the most similar similar triangles:
[0010] dist The smaller it is, the closer the ratios of the corresponding line segments of the two triangles are, and at the same time, it also means that the two triangles are more similar. Among the rough matching pairs of feature points preliminarily screened by the method of the ratio of the nearest Euclidean distance to the second nearest Euclidean distance, arbitrarily select three pairs of feature point matching pairs. The three pairs of feature point matching pairs will form two triangles. Calculate the lengths of each line segment according to the coordinates of the three feature points, then calculate the ratio of the lengths of the corresponding line segments of the triangles to obtain three ratios, and finally find the most similar three pairs of feature point matching pairs according to the above two conditions.
[0011] Furthermore, in step S200 of the present invention, the encoder fusion framework for extracting the deep features of the infrared and visible light images based on the feature point matching pairs includes: source image input, encoder, feature fusion, decoder, fused image, and loss function. Among them, the training stage only includes the encoder and decoder parts; Encoder part: It consists of three convolutional layers and three residual blocks. The size of the input training data is 256×256. The first convolutional layer does not change the size of the input image, and the second and third convolutional layers are half of the input size. Through the residual network, reuse the previous features to make up for the image details lost during the convolution process, and add three residual blocks after the last convolutional layer of the residual network. All the convolutional operations are used as feature extractors, retaining the texture and structural information of the source image. The output of the encoder has 256 intermediate features with a size of 64×64, retaining more original structural details; Decoder part: The starting and ending requirements of the skip connection are that the number of feature channels and the size should be consistent. Therefore, the output sizes of the corresponding convolutional layers in the decoding and encoding parts are the same. For the input depth features, first perform a transposed convolution to obtain a feature layer with the same size as the shallow features of the second layer of the encoder, and then superimpose them. After the superposition, enter the next transposed convolution and continue to superimpose with the shallow features of the first layer. Finally, after another transposed convolution, the fused image is obtained. Among them, the kernel size of the transposed convolution layer is the same as that of the convolutional layer, both set to 3×3.
[0012] Furthermore, the skip connection of the present invention includes: In the training network, connect the first and second convolutional layers of the encoder to the corresponding layers of the decoder through skip connections. At the same time, also use skip connections in the second and third residual blocks to reuse the results of the third convolutional layer and the first residual block, that is, a "large residual block" structure is newly constructed, and a residual block is constructed on the outside.
[0013] Furthermore, step S300 of the present invention, obtaining an attention map by adding an attention mechanism and fusing the attention map with deep features to obtain a feature map includes: In the feature fusion framework of the attention mechanism, the output of the residual block is a series of deep features ; Among them, res represents the output of the residual block, k The value of k is 1, 2. When A =1, it represents the source image k =2 represents the source image B ; i is the number of intermediate features of the input image. In order to accurately reflect the significant features of the source image, an attention map needs to be created from these feature maps; each feature map has its own weight, which is calculated by the softmax function; is the probability weight calculated by the L 1-norm and the softmax function. The specific weight is calculated by the following formula:
[0014] Among them represents the L1 norm, (x,y) represents the deep features of two source images at the corresponding position in the probability weight , is the depth feature vector at the point (x,y). All intermediate features are multiplied by the corresponding probability weights to generate the attention map of the source image. Its calculation formula is:
[0015] Among them, Represents the enhanced deep features, reflecting the attention map of the source image weight. Before feature-level fusion, the attention map is used to optimize the intermediate features, and then the softmax Calculate the probability weights and get is the optimal weight map of the source image features, calculated as:
[0016] Finally, we get The fused features that will be sent to the decoder are calculated as:
[0017] in, i =1,…,256, subscript 1 is the source image A , 2 is the source image B , the final output feature layer size is the same as the input.
[0018] Furthermore, in S400 of the present invention, the feature map is skip-connected and transferred to the corresponding deconvolution layer for processing, and the method for obtaining the fused image includes: Pixel loss L pixel and SSIM loss SSIML The sum of the total loss function is calculated L total The total loss function can be expressed as the following formula, where L total Indicates the total loss, L pixel represents pixel loss and L SSIM express SSIM loss:
[0019] use L 2 norm to calculate pixel loss L pixel The formula is:
[0020] in express L 2 norm; O and I are the output and input images respectively; By definition, the structural similarity of the entire graph MSSIM、 The structural similarity loss is shown as follows:
[0021]
[0022] in O and Iare the output and input images respectively, which is calculated by the above formula, M 、 N are the lengths and widths of different input pictures, and are the corresponding position values in the two pictures.
[0023] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to execute the steps of the above method.
[0024] On yet another aspect, the present invention also discloses a computer device including a memory and a processor, where the memory stores a computer program, and when the computer program is executed by the processor, it causes the processor to execute the steps of the above method.
[0025] As can be seen from the above technical solutions, this part mainly proposes an encoder fusion framework with residual blocks to extract deep features in infrared and visible light images; in order to better fuse the target information of infrared light images and the detailed information contained in visible light images, an attention mechanism is added in the feature fusion stage to obtain an attention map for fusing deep features; finally, the feature maps obtained in the first and second layers, that is, shallow features, are passed to the corresponding deconvolution layers through skip connections for processing to obtain a fused image. Experimental results show that the fusion result of this method is clearer subjectively and achieves better results than existing methods in objective indicators such as average gradient, structural similarity, peak signal-to-noise ratio, and mutual information. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is the rough matching pair diagram of the feature points of the visible light image and the infrared image of the present invention; Figure 2 is the flowchart of the RANSAC algorithm of the present invention; Figure 3 are similar triangles; Figure 4 is the acquisition process of geometric constraint feature point matching pairs; Figure 5 are the visible light feature points obtained by the traditional RANSAC algorithm; Figure 6 are the corresponding feature points of the infrared image obtained by the traditional RANSAC algorithm; Figure 7 are the visible light image feature points obtained by geometric constraints; Figure 8 are the infrared image feature points obtained by geometric constraints; Figure 9 is the schematic diagram of the fusion framework based on the residual module; Figure 10 Schematic diagram of the overall training structure of the present invention; Figure 11 Schematic diagram of the residual block and skip connection combination module of the present invention; Figure 12 Schematic diagram of the feature fusion framework based on the attention mechanism; Figure 13 Schematic diagram of the infrared and visible light image fusion result; Figure 14 Schematic diagram of the infrared and visible light image fusion result; Figure 15 Schematic diagram of the fusion effect of the present invention in other scenarios. Specific implementation manners
[0027] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention.
[0028] As Figure 1 shown, the low-cost and high-precision visible light and infrared fusion method described in this embodiment is executed by a computer device through the following steps: S100. Obtain the feature point matching pairs of the infrared and visible light images by using the geometric constraint method; S200. Based on the feature point matching pairs, extract the deep features of the infrared and visible light images through an encoder fusion framework with residual blocks; S300. Obtain an attention map by adding an attention mechanism, and fuse the attention map with the deep features of the infrared and visible light images to obtain a feature map; S400. Pass the feature map through skip connections to the corresponding deconvolution layer for processing to obtain a fused image.
[0029] The following is a detailed description of each step: S100. Obtain the feature point matching pairs of the infrared and visible light images by using the geometric constraint method; S110. Obtain the rough feature point matching pairs; After detecting the feature points of the image by using the SURF algorithm, it is necessary to match the feature points of the images to be registered to form feature point matching pairs. Image registration essentially is to find the correct feature point matching pairs of the images to be registered, and then through the geometric position relationship of the feature point matching pairs, solve the specific transformation parameters of the corresponding geometric transformation model, and perform geometric registration on the images to be registered.
[0030] Because the physical environment during imaging can vary, different images of the same scene may appear "identical" to the human eye, but they may appear vastly different to a computer. Furthermore, each image has numerous feature points, so during feature point matching, incorrectly matched pairs often form. These incorrectly matched pairs are extremely detrimental to image registration. If incorrectly matched pairs are used to find geometric transformations, the resulting geometric transformation parameters will not be compatible with the images being registered. Therefore, finding the correct feature point matching pairs for the images being registered is crucial for registration algorithms.
[0031] The main function of the feature point descriptor vector is to quantify the features of the image feature points into digital representations. This means that the description of features is no longer limited to visual descriptions but can also be quantified into mathematical descriptions. It is even possible to compare the similarities of similar features at the mathematical level. To describe the similarity of features, the Euclidean distance is generally used for judgment, and the formula is as follows:
[0032] in A 、 B Represents two feature points, a i Representing feature points A The feature point descriptor vector i Component values, b i Same reason. D(A,B) The smaller the value, the smaller the overall difference between the feature points. The two feature points can be considered similar and can be considered as a feature point matching pair. On the contrary, the greater the difference, the feature points do not form a matching pair.
[0033] If we simply consider the Euclidean distance between different feature points as the only criterion for forming feature point matching pairs, the following two problems will arise: (1) The minimum value of the Euclidean distance between feature points must exist, but the graph A The feature points in the figure B There may not necessarily be matching points in , so the feature point matching pair formed by the minimum Euclidean distance criterion is not necessarily the correct feature point matching pair.
[0034] (2) Figure A The feature points in the figure B There may be multiple extremely similar feature points in a dataset, and it is not reliable to simply exclude other feature points based on the minimum Euclidean distance criterion.
[0035] Regarding the above two feature point matching pair problems, the first problem always objectively exists, which is also one of the main reasons for the source of incorrect feature point matching pairs. For the second problem, the method of using the ratio of the nearest Euclidean distance to the second-nearest Euclidean distance can be used to mitigate the harm brought by the second problem. Calculate the Euclidean distances between a feature point in A Figure and all feature points in the image B to be registered. It is certain that the nearest Euclidean distance D min and the second-nearest Euclidean distance D smin can be obtained, and the ratio of the two is calculated:
[0036] If the ratio k is less than the set threshold, a matching pair is formed with the feature point of this nearest Euclidean distance; otherwise, no matching pair is formed. The ideal value of the threshold is generally set between 0.5 and 0.7, and the middle value 0.6 is generally taken. It can be known from the formula analysis that k the smaller it is, the more obvious the nearest Euclidean distance is less than the second-nearest Euclidean distance, and a feature point matching pair can be formed. If k the value is large, the second-nearest Euclidean distance is relatively close to the nearest Euclidean distance, and no obvious advantage can be formed, so no feature point matching pair can be formed. The method of using the ratio of the nearest Euclidean distance to the second-nearest Euclidean distance cannot completely solve the above two problems. It can only form a rough matching pair for the feature points between the images to be registered. There will still be incorrect matches in the feature point matching pairs preliminarily screened by this method. The situation of incorrect feature point matching pairs objectively exists, and other restrictive methods must be used to eliminate the incorrect feature point matching pairs.
[0037] According to the research on the principle of the SURF algorithm, the feature points extracted by using the SURF algorithm ( Figure 1 as shown), and then according to the method of using the ratio of the nearest Euclidean distance to the second-nearest Euclidean distance of the image feature point descriptor vectors, the rough matching pairs of the feature points preliminarily screened are shown in the following figure: Such as Figure 1As shown, there are approximately 40 pairs of formed feature point matches. Moreover, most of the feature point matches formed by the initial screening of the infrared image and the visible light image are basically mis-matches. The main reason for the large number of mis-matches here lies in the different imaging mechanisms of the two heterogeneous image sensors. The visible light imaging sensor forms an image based on the brightness of the light reflected by the object, while the intensity value of the pixel in the infrared image sensor is related to the heat radiated by the object. The higher the temperature of the object, the more heat it radiates, and the larger the intensity value of the pixel in the infrared imaging sensor. The lower the temperature of the object, the less heat it radiates, and the smaller the intensity value of the pixel in the infrared sensor. When using the SURF algorithm to extract the image blob features for these two types of images, the feature points extracted from the visible light image are blobs of brightness, while for the infrared image, the extracted blob features are blobs of heat. It is precisely due to the different imaging mechanisms that the properties of the image feature points extracted by the algorithm are different. If the feature points in two images are to form a correct match pair, then it is required that this point be not only a blob at the brightness level but also a blob at the heat level. It is very demanding to simultaneously meet the above two conditions to form a feature point match pair. Therefore, most of the feature point match pairs are mis-matches, and only a few correct feature point match pairs exist.
[0038] S120, Classical RANSAC Algorithm By the method of the ratio of the nearest Euclidean distance to the second-nearest Euclidean distance, relatively accurate feature point match pairs can be preliminarily screened out. These feature point match pairs have a higher accuracy rate compared to the feature point match pairs directly using the minimum Euclidean distance criterion for matching. However, even so, the problem of feature point mis-match pairs still objectively exists. The currently popular method is to use the RANSAC (Random Sample Consensus) algorithm to eliminate the mis-feature point match pairs. This algorithm was jointly proposed by Fischler and Bolles in 1981. This algorithm for eliminating mis-match pairs mainly uses the known data to estimate the parameters of the unknown mathematical model, and this method can eliminate the obvious incorrect data. In the SURF registration algorithm, the data is the rough match pairs of the feature points between the images to be registered preliminarily screened by the nearest neighbor ratio to the second-nearest neighbor method, and the mathematical model is the geometric transformation model for registering the images to be registered. In this method, the selected one is the affine transformation model.
[0039] First, a set of data in the known dataset is used to estimate the parameters of the mathematical model. After obtaining the parameters of the mathematical model, a specific model is obtained. Then, this model is used to verify other data in the dataset. If the verified data conforms to the model, the confidence level of the model is incremented by one. After verifying all the data, the total confidence level of the model is obtained. Then, another set of data in the dataset is used to obtain the model parameters of this set of data, and the overall confidence level of the model is calculated. This step is repeated until all the data in the dataset is traversed. According to this iterative method, the confidence levels of all possible mathematical models can be obtained, and finally, the model parameters with the highest confidence level are selected as the final model.
[0040] As Figure 2 shown: When using the SURF feature extraction algorithm for image registration, a set of feature point matching pair data is selected from the data of the rough feature point matching pairs in the dataset. The specific parameters of the model are solved according to the geometric transformation model selected by the registration algorithm. Then, this model is used to verify other feature point matching pairs in the data of the rough feature point matching pairs. If the feature point matching pair passes the verification of the geometric transformation model, the confidence level of the geometric transformation model is incremented by one. After verifying all the feature point matching pairs in the data of the rough feature point matching pairs, the total confidence level of the geometric transformation model can be obtained. Then, another set of feature point matching pair data is selected from the dataset of feature point matching pairs to obtain the geometric transformation model parameters determined by this set of feature point matching pairs, and the above process of calculating the model confidence level is repeated. Finally, after iterating through all the data in the dataset, the confidence levels of all possible geometric transformation models are obtained. The one with the highest confidence level is selected as the final geometric transformation model.
[0041] In essence, the RANSAC algorithm uses the data in the dataset to estimate the total confidence level of the mathematical model generated by this data with respect to all the data in the dataset, and obtains the model with the most votes, that is, the highest confidence level. The main advantage of this algorithm is its relatively high robustness, but its disadvantages are also relatively obvious. The iterative calculation amount is relatively large. Because the mathematical model generated by each set of data has to poll all the data in the entire dataset for calculation, if the dataset is large, it consumes a lot of computing resources. In addition to this problem, it also has certain requirements for the data in the dataset. The premise for the RANSAC algorithm to be effective is that the correct and valid data in the dataset is more than the abnormal data. If the number of correct matching pairs in the feature point matching pairs is less than the number of feature point mis-matching pairs, then the votes of the incorrect data in the algorithm are more than the votes of the correct data. Then, the total confidence level of the model calculated based on the dataset model is also necessarily untrustworthy.
[0042] S130. Geometric constraint method By analyzing, the prerequisite for the normal operation of the RANSAC algorithm is that the number of correct and effective matching pairs in the feature point matching pair dataset should be more than the number of wrong matching pairs. For homologous image registration, since the imaging principles of the sensors for obtaining images are the same, the image features generally do not differ much, and it is relatively easy to obtain more correct feature point matching pairs. However, for heterologous images, due to the different physical mechanisms for obtaining images, the feature points detected by using image feature algorithms may be very different in nature. For this reason, the number of wrong matching pairs of feature points that may be extracted may be greater than the number of correct matching pairs. In the face of this situation, directly using the RANSAC algorithm may not have good results, and a method for extracting feature point matching pairs based on geometric constraints is proposed.
[0043] S131. Principle of eliminating wrong matching pairs by geometric constraints The method of obtaining feature point matching pairs of images based on geometric constraints is essentially to make the conditional requirements for combining the feature points of two images to form feature point matching pairs more stringent. It is not only necessary to meet the requirement of the ratio of the nearest neighbor Euclidean distance to the second nearest neighbor Euclidean distance, but also the positions of the feature point matching pairs should meet the geometric constraint conditions. Only the corresponding two feature points that meet the above two conditions can form a feature point matching pair.
[0044] After camera calibration, the imaging of the imaging sensor can be considered to be without distortion. Excluding the possible non-linear transformation caused by image distortion, this means that there is a large part of similar information in the two images of the same scene, and there is only a linear geometric transformation relationship in this part of the image. Then the geometric positions of the image feature points extracted from the similar parts of the two images should also be similar. That is, if three feature points are extracted from the visible light image to form a triangle, then the corresponding feature points in the infrared image should also form a similar triangle, and there is only a relationship of translation, rotation and scaling between these two triangles. Instead, think in reverse. If there are three or more pairs of correct feature point matching pairs in the two images, then the two triangles formed by these three pairs of feature point matching pairs must also be the most similar among all the matching pairs. Based on this assumption, the wrong matching pairs can be eliminated by adding the constraint condition of triangle similarity to the feature point matching pairs, and three correct pairs of feature point matching pairs can be obtained, and the affine transformation model parameters between the images to be registered can be solved.
[0045] S132. Design of geometric constraint algorithm To determine the similarity of two triangles, the ratio of the lengths of corresponding line segments can be used as a criterion, or the angles of the two triangles can be used as the similarity criterion. If two of the three corresponding angles are the same, then the two triangles are similar. Generally, the cosine theorem is used to calculate the angles of a triangle, and the cosine theorem also requires calculating the lengths of the line segments of the triangle. Therefore, it is better to directly select the ratio of the line segment lengths as the similarity criterion. In this method program, this ratio of the lengths of corresponding sides is also selected as the similarity criterion.
[0046] As Figure 3 shown, for two similar triangles, how to find the optimal similarity in the algorithm? This requires finding an expression to quantify the similarity of the two triangles. It can be designed like this in the program:
[0047] First, calculate the ratio of the lengths of the corresponding line segments of the triangle. Here, the length of each line segment is the Euclidean distance length in the pixel coordinate system. Theoretically, the ratio of the lengths of the corresponding line segments of similar triangles should be equal. However, in the program, the equality cannot be used as the condition to determine the similarity of the triangles. When designing the algorithm, a threshold for the line segment ratio can be set to ensure that the images are similar. The formula is as follows:
[0048] In the program, the threshold is designed to be 0.02. Only when both are less than the threshold and meet the conditions of the above two formulas can it be proved that the ratios of the three corresponding line segments are close to being the same. If it is greater than the threshold, the triangle is judged to be not similar. Under this condition of ensuring similarity, the most similar similar triangles are selected using the following conditions:
[0049] As Figure 4 shown: dist The smaller it is, the closer the ratios of the corresponding line segments of the two triangles are, and at the same time, it also means that the two triangles are more similar. Among the rough matches of feature points preliminarily screened by the method of the nearest Euclidean distance ratio to the second-nearest Euclidean distance, any three pairs of feature point matches are selected. The three pairs of feature point matches will form two triangles. According to the coordinates of the three feature points, the lengths of each line segment can be obtained, and then the ratio of the lengths of the corresponding line segments of the triangle is calculated to obtain three ratios. Finally, the most similar three pairs of feature point matches are found according to the above two conditions.
[0050] S140. Comparative Experimental Analysis of Traditional RANSAC Algorithm and Geometric Constraint Algorithm For the obtained rough matches of feature points, the experimental results obtained by using the RANSAC algorithm and the proposed geometric constraint method respectively are as Figure 5 、 6 、7, 8 shown: It can be found from the above four figures that among the three pairs of feature point matching pairs directly obtained by using the traditional RANSAC algorithm for the rough matching of feature points, two of the matching pairs belong to false matches. This also exactly proves that the premise for the effectiveness of the RANSAC algorithm is that the number of correct matching pairs in the feature point matching pair dataset is more than the number of false matching pairs. The feature point matching pairs selected by using the geometric constraint method are significantly better than those selected by the RANSAC algorithm. As long as there are three or more correct feature point matching pairs in the feature point matching pairs, the method of using geometric constraints to obtain feature point matching pairs can surely find the correct feature point matching pairs.
[0051] In this part, the RANSAC algorithm and its principle of removing false matching pairs are mainly introduced, and its existing problems are also analyzed. In the registration of infrared images and visible light images, due to the different imaging mechanisms, the number of false matching pairs in the feature point matching pairs screened by the ratio of the nearest Euclidean distance to the second nearest Euclidean distance is much more than that of the correct matching pairs. At this time, directly using the RANSAC algorithm to remove the false matching pairs of feature points does not have a very good effect. Then a method of adopting geometric constraints to obtain feature point matching pairs is proposed, the logical feasibility is discussed, and the control experiment results are given. The experiment shows that two of the three pairs of feature point matching pairs obtained by the traditional RANSAC are false matching pairs, but all three pairs of feature point matching pairs obtained based on geometric constraints are correct matching pairs.
[0052] The experimental results show that the fusion result of this method is clearer in subjective feeling and achieves better results than the existing methods in objective indexes such as average gradient, structural similarity, peak signal-to-noise ratio, and mutual information.
[0053] S200. Based on the feature point matching pairs, obtain the deep features of infrared and visible light images through the encoder fusion framework with residual blocks; As Figure 9 shown, the encoder fusion framework consists of the following six parts: source image input, encoder, feature fusion, decoder, fused image, and loss function. This encoder-decoder structure has good reconstruction characteristics without supervised learning.
[0054] Specifically, a publicly available dataset is used to train an encoder-decoder network with residual blocks, and the trained model is used to encode two source images respectively. The feature maps obtained from the first and second convolutional layers carry different texture details of the source images, that is, shallow features, while the result obtained after the third convolutional layer is intermediate features. Three residual blocks are added after the third layer, and skip connections are used in the second and third residual blocks to reuse the results of the third convolutional layer and the first residual block. In this way, while deepening the network, the information of the intermediate features is also well utilized to obtain deep features. Feature fusion consists of two parts. First, the deep features extracted are passed through an attention mechanism to obtain two attention maps, and then the two attention maps are fused. The fusion result obtained by fusing these deep features can be used as a target information. The decoder part consists of three transposed convolutional layers. Since many detailed parts are obtained from the first and second convolutional layers, but a lot of information will inevitably be lost during the deepening of the network, shallow features are also important for obtaining the fused image. Therefore, skip connections are also added to the decoder part to utilize the shallow features during the decoding process, and finally the fused image is obtained.
[0055] As Figure 10 shown, the training stage only includes the encoder and decoder parts and does not include the fusion part. The main purpose of the training stage is to train the encoder part to better extract features, and then enable the decoder to accurately reconstruct the original dataset while minimizing the reconstruction error. The smaller the reconstruction error, the more representative the extracted features are. The training network consists of encoding, decoding, and two skip connections. The functions of each module are introduced below.
[0056] (1) Encoder part: It consists of three convolutional layers and three residual blocks. The size of the input training data is 256×256. The first convolutional layer does not change the size of the input image, and the second and third convolutional layers are half of the input size. To make up for the image details lost during the convolution process, through the residual network, the previous features are further reused. And in this network, three residual blocks are added after the last convolutional layer. All convolutional operations serve as feature extractors, fully retaining the texture and structural information of the source image. The output of the encoder has 256 intermediate features with a size of 64×64, retaining more original structural details.
[0057] (2) Residual blocks: Three residual blocks are added to the encoder part, which have two purposes: one is to ensure the optimal convergence during the training of the deep network; the other is to better utilize the intermediate features generated by the third convolutional layer. At the same time, as Figure 11 shown, through skip connections, the input and output of the first residual block are respectively connected to the input and output of the third residual block, so that the features of the intermediate layer can be better used and the feature information can be better extracted.
[0058] (3) Decoder part: Since the start and end of the skip connection require the number of feature channels and the size to be consistent, the output sizes of the corresponding convolutional layers in the decoding and encoding parts are the same. For the input depth features, first perform a transposed convolution to obtain a feature layer with the same size as the shallow features of the second layer of the encoder for superposition. After superposition, enter the next transposed convolution and continue to superpose with the shallow features of the first layer. Finally, after another transposed convolution, the fused image is obtained. Among them, in order to facilitate superposition, the kernel size of the transposed convolution layer is set to 3×3, the same as the kernel size of the convolutional layer.
[0059] (4) Skip connection: In the training network, the first and second layer convolutions of the encoder are connected to the corresponding decoder layers through skip connections. At the same time, skip connections are also used in the second and third residual blocks to reuse the results of the third layer convolution and the first residual block. Substantially, it is equivalent to newly constructing a "large residual block" structure, and then constructing a residual block on the outside. This structure can further promote the interaction and fusion of the shallow features and deep features of the source image, obtain more useful features during the transposed convolution process, and ensure better fusion of the images.
[0060] S300. Obtain the attention map by adding the attention mechanism, and fuse the attention map with the deep features of the infrared and visible light images to obtain the feature map; As Figure 12 shown, it is Figure 9 's feature fusion framework based on the attention mechanism. Input the feature map obtained in the encoding network into the feature fusion framework, and the significant objects in the source image are given higher attention values. The result output after feature fusion is superposed and sent to the decoder to reconstruct the fused image. However, in this process, a large amount of detailed information is lost in the intermediate features. Therefore, in order to retain the detailed information in the fusion result, the shallow features obtained after the first and second layer convolutions in the encoder are retained using skip connections.
[0061] In the feature fusion framework of the attention mechanism, the output of the residual block is a series of deep features . Among them, res represents the output of the residual block, k 's value is 1, 2. When k = 1, it represents the source image A , k = 2 represents the source image B . i is the number of intermediate features of the input image. In order to accurately reflect the significant features of the source image, it is necessary to create an attention map from these feature maps. Each feature map has its own weight, and its weight can be calculated by the softmax function. is composed of L 1-norm and softmaxThe probability weights calculated by the function. The specific weights are calculated by the following formula:
[0062] where represents the L1 norm, and (x, y) represents the deep features of two source images At the corresponding position in the probability weight is the depth feature vector at the point (x, y). All intermediate features are multiplied by the corresponding probability weights to generate the source image attention map. Its calculation formula is:
[0063] where represents the enhanced depth feature, reflecting the attention map of the source image weights. Before feature-level fusion, the intermediate features are optimized using the attention map, and continue to use softmax to calculate the probability weights, and the obtained is the optimal weight map of the source image features, calculated as:
[0064] Finally, the obtained will be sent to the fused features in the decoder part, and the calculation formula is:
[0065] where i = 1,..., 256, the subscript 1 is the source image A , 2 is the source image B , and the size of the finally output feature layer is the same as the input.
[0066] S400. Through the skip connection feature map, it is passed to the corresponding deconvolution layer for processing to obtain the fused image.
[0067] In this method, the pixel loss L pixel and SSIM loss SSIML are used to calculate the total loss function L total . The total loss function can be expressed as the following formula, where L total , L pixel and L SSIM represent the total loss, pixel loss, and SSIM loss respectively:
[0068] Use L the 2-norm to calculate the pixel loss L pixel , as shown in the following formula, where represents L 2-norm; O and I are the output and input images respectively:
[0069] According to the definition, the structural similarity of the whole image MSSIM、 The structural similarity loss is shown as the following formula:
[0070]
[0071] where O and I are the output and input images respectively, which are calculated by the above formula, M , N are the length and width of the input different pictures, and are the corresponding position values in the two pictures.
[0072] Such as Figure 13 , 14 , shown in 15, demonstrates the fusion effects of the present method in other different scenarios. From the Figure 13 results, it can be seen that the present method makes full use of and retains the detail layer information of the image, and well shows the target area of the image, and also enhances some edges that are relatively blurred in visible light. The fusion results are highly robust to images with different sharpness.
[0073] An encoder fusion framework based on residual blocks is proposed. Three residual blocks and three convolutional layers are combined to design a feature extraction network, which effectively extracts the features of the active image; and an attention mechanism is introduced for feature fusion; in the decoding part, the shallow features are also reused through skip connections, better retaining the detail information. The features generated by this method are unique, with low redundancy, avoiding the possible distortion and inconsistency in the fused image, and solving the artifact problem while retaining more detail information and target information. Compared with a variety of classical methods, the results show that the proposed fusion method has achieved good results.
[0074] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to execute the steps of the above method.
[0075] On yet another aspect, the present invention also discloses a computer device including a memory and a processor, where the memory stores a computer program, and when the computer program is executed by the processor, it causes the processor to execute the steps of the above method.
[0076] In another embodiment provided by the present application, a computer program product including instructions is further provided. When it runs on a computer, it causes the computer to execute any one of the low-cost and high-precision visible light and infrared fusion methods in the above embodiments.
[0077] It can be understood that the systems, devices, and storage media provided by the embodiments of the present invention correspond to the methods provided by the embodiments of the present invention. For the explanations, examples, and beneficial effects of the relevant content, reference can be made to the corresponding parts in the above methods.
[0078] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)).
[0079] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.
[0080] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiment.
[0081] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A low-cost and high-precision visible light and infrared fusion method, characterized in that, The following steps are performed by a computer device: S100. Obtain the feature point matching pairs of infrared and visible light images by using the geometric constraint method; S200. Based on the feature point matching pairs, extract the deep features of infrared and visible light images through an encoder fusion framework with residual blocks; S300. Obtain an attention map by adding an attention mechanism, and fuse the attention map with the deep features of infrared and visible light images to obtain a feature map; S400. Pass the feature map through a skip connection to the corresponding deconvolution layer for processing to obtain a fused image.
2. The low-cost and high-precision visible light and infrared fusion method according to claim 1, wherein The method for obtaining the feature point matching pairs of infrared and visible light images by using the geometric constraint method in step S100 includes: Geometric constraint algorithm design Find the optimal similarity of two similar triangles in the algorithm and set the expression for quantifying the similarity of the two triangles: First, calculate the ratio of the lengths of the corresponding line segments of the triangle. The length of each line segment is the Euclidean distance length in the pixel coordinate system. When designing the algorithm, set the threshold of the line segment ratio to ensure that the images are similar. The formula is as follows: The threshold is designed to be 0.02 in the program. Only when both of them are less than the threshold can the conditions of the above two formulas be satisfied to prove that the corresponding ratios of the three line segments are approximately the same; if it is greater than the threshold, the triangle is judged to be non-similar; under this condition that guarantees similarity, the most similar similar triangles are selected using the following conditions: dist The smaller the ratio, the closer the corresponding line segments of the two triangles are, and it also indicates that the two triangles are more similar. Among the rough matches of feature points preliminarily screened by the method of the ratio of the nearest Euclidean distance to the second-nearest Euclidean distance, randomly select three pairs of feature point matches. The three pairs of feature point matches will form two triangles. Calculate the lengths of each line segment based on the coordinates of the three feature points, then calculate the ratio of the lengths of the corresponding line segments of the triangles to obtain three ratios. Finally, find the three pairs of feature point matches that are the most similar according to the above two conditions.
3. The low-cost and high-precision visible light and infrared fusion method according to claim 1, characterized in that, In step S200, the encoder fusion framework for extracting the deep features of infrared and visible light images based on the feature point matching pairs includes: source image input, encoder, feature fusion, decoder, fused image, and loss function; among them, the training stage only includes the encoder and decoder parts; Encoder part: It consists of three convolutional layers and three residual blocks. The size of the input training data is 256×256. The first convolutional layer does not change the size of the input image. The second and third convolutional layers are half of the input size; through the residual network, the previous features are reused to make up for the image details lost during the convolution process, and three residual blocks are added after the last convolutional layer of the residual network. All convolutional operations are used as feature extractors, retaining the texture and structure information of the source image. The output of the encoder has 256 intermediate features, with a size of 64×64, retaining more original structural details; Decoder part: The start and end of the skip connection require the feature channel number and size to be consistent, so the output sizes of the corresponding convolutional layers in the decoding and encoding parts are the same; for the input deep features, first perform a deconvolution to obtain a feature layer with the same size as the shallow features of the second layer of the encoder for superposition. After superposition, enter the next deconvolution and continue to superpose with the shallow features of the first layer. Finally, after another deconvolution, obtain the fused image; among them, the kernel size of the deconvolution layer is the same as that of the convolutional layer, both set to 3×3.
4. The low-cost and high-precision visible light and infrared fusion method according to claim 3, wherein The skip connection includes: In the training network, connect the first and second convolutional layers of the encoder to the corresponding layers of the decoder through skip connections, and at the same time use skip connections in the second and third residual blocks to reuse the results of the third convolutional layer and the first residual block, that is, a "large residual block" structure is newly constructed, and a residual block is constructed on the outside.
5. The low-cost and high-precision visible light and infrared fusion method according to claim 1, characterized in that In step S300, the method for obtaining an attention map by adding an attention mechanism and fusing the attention map with the deep features to obtain a feature map includes: In the feature fusion framework of the attention mechanism, the output of the residual block is a series of deep features ; Among them, res represents the output of the residual block, k The value of is 1, 2, when k = 1 represents the source image A , k = 2 represents the source image B , i is the number of intermediate features of the input image; in order to accurately reflect the significant features of the source image, it is necessary to create an attention map from these feature maps; each feature map has its own weight, which is calculated by the softmax function; is L the 1-norm and softmax the probability weight calculated by the function, and the specific weight is calculated by the following formula: Among them represents the L1 norm, and (x, y) represents the deep features of two source images at the corresponding position in the probability weight ; is the depth feature vector of the point (x, y); all intermediate features are multiplied by the corresponding probability weights to generate the source image attention map, and its calculation formula is: Among them, represents the enhanced depth feature, which reflects the attention map of the source image weights; before feature-level fusion, the intermediate features are optimized using the attention map, and continue to use softmax to calculate the probability weights, and the obtained is the optimal weight map of the source image features, calculated as: Finally, the obtained fused features to be sent to the decoder part are calculated as follows: Among them, i = 1, …, 256, where the subscript 1 is for the source image A , 2 is for the source image B , and the size of the finally output feature layer is the same as the input.
6. The low-cost and high-precision visible light and infrared fusion method according to claim 1, characterized in that The method in S400 for passing the feature map through a skip connection to the corresponding deconvolution layer for processing to obtain a fused image includes: Pixel loss L pixel and SSIM loss SSIML to calculate the total loss function L total , the total loss function is expressed as the following formula, where L total represents the total loss, L pixel represents the sum of pixel losses L SSIM represents SSIM loss: Use L the 2-norm to calculate the pixel loss L pixel The formula is: Among them denotes L the 2-norm; O and I are the output and input images, respectively; According to the definition, the structural similarity of the whole picture MSSIM The structural similarity loss is shown as follows: where O and I are the output and input images respectively, is calculated by the above formula, M and N are the lengths and widths of different input images, and are the corresponding position values in the two images.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the processor is caused to execute the method according to any one of claims 1 to 6.
8. A computer device, comprising a memory and a processor, characterized in that, The memory stores a computer program, which, when executed by the processor, causes the processor to execute the method according to any one of claims 1 to 6.