A registration method for registering infrared light images to visible light images
Through the VGG16 model, the features of infrared and visible images are extracted and fused, and combined with correlation solution and regression network, the problem of infrared and visible images is solved, achieving a higher accuracy registration effect.
Patent Information
- Application Number
- CN202111399384.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-19
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-11-19
AI Technical Summary
It is difficult for the prior art to accurately establish a reliable matching model to achieve accurate registration of infrared and visible images, especially when the mode and scale differences are large.
The VGG16 pre-trained model is used to extract the features of infrared and visible light images, and the shallow and deep features are fused through the feature fusion module to generate a fusion feature map. Then, the affine transformation parameters are predicted by correlation solution and regression network model to achieve registration of infrared light images.
It effectively improves the accuracy of infrared and visible images registration, reduces feature differences, and improves the registration effect.
Smart Images

Figure CN114066955B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a registration method, and particularly to a registration method for registering an infrared light image to a visible light image. Background Art
[0002] Infrared and visible light images are two common images with complementary information. Using both infrared and visible light images simultaneously has great advantages in various fields of computer vision. Therefore, an accurate and efficient image registration method is an important prerequisite. An infrared image is obtained by measuring the heat radiated by an object, which can intuitively reflect the energy magnitude of the object's thermal radiation, but it will lose appearance information such as the texture and structure of the object. While a visible light image contains relatively rich texture structure and color information, but is greatly affected by light and occlusion. Due to the obvious modal differences between infrared and visible light images, small correlation, and lack of consistent features, the registration is difficult.
[0003] Current registration methods are mainly divided into region-based methods and feature-based methods. Region-based registration methods use different types of similarity metrics, including mutual information, normalized mutual information, FFT-based phase correlation method, etc. However, these methods may lead to local extrema when processing images with large modal differences such as infrared and visible light images, and the computational load is large. Compared with region-based registration techniques, feature-based registration methods mainly solve the rotation and scale transformation between images by establishing reliable feature matches. Currently, many scholars have carried out research on extracting similar features from heterologous images. Some scholars have made the feature descriptors such as SIFT and SURF adapt to the obvious differences in pixel gray values between images obtained by different sensors by improving traditional feature matching algorithms. In addition, some scholars use edge information to overcome the differences between different modal images, but this type of method cannot guarantee that a large number of repeated and consistent feature points can be extracted from heterologous images, and there are a large number of outliers in the process of feature matching, and accurate registration effects cannot be obtained when there are large gray differences between images due to different shooting times, spectra, and sensors. Summary of the Invention
[0004] Aiming at the technical problem that for the registration of infrared and visible light images with large modal and scale differences, it is difficult for traditional heterologous image registration algorithms to accurately establish a reliable matching model to achieve image registration, the present invention provides a registration method for registering an infrared light image to a visible light image.
[0005] The present invention is implemented by the following technical solutions: A registration method for registering an infrared light image to a visible light image, which includes the following steps:
[0006] Step 1: Extract the shallow infrared features and deep infrared features from an infrared image, and extract the shallow visible features and deep visible features from a visible image; the infrared image and the visible image are two different modality images of the same scene.
[0007] Step 2: Perform feature fusion on the shallow infrared features and the deep infrared features to form an infrared fusion feature map with infrared fusion features, and perform feature fusion on the shallow visible features and the deep visible features to form a visible fusion feature map with visible fusion features.
[0008] Step 3: Solve the correlation between each infrared fusion feature in the infrared fusion feature map and each visible fusion feature in the visible fusion feature map to obtain a correlation map with corresponding correlations.
[0009] Step 4: Input the correlation map into a regression network model to predict the affine transformation parameters for registering the infrared image to the visible image.
[0010] Step 5: Register the infrared image to be registered to the corresponding visible image through the affine transformation parameters.
[0011] As a further improvement of the above solution, use the deep convolutional network model, namely the VGG16 model, as the feature extraction model. The full English name of VGG is Visual Geometry Group, and the Chinese translation is Visual Geometry Group Network.
[0012] Further, extract the shallow infrared features and deep infrared features from the infrared image through convolution operations, and extract the shallow visible features and deep visible features from the visible image.
[0013] Further, the ways of feature fusion are all:
[0014] F = Cat(Down(F1) 8 , Down(F2) 4 , Down(F3) 2 , F4)
[0015] Among them, F represents the fusion feature map, referring to the infrared fusion feature map or the visible fusion feature map, Cat(·) represents the splicing of the fusion feature map in the channel dimension, Down(·) represents the downsampling operation on the shallow features in the image, and its subscript is the multiple of downsampling. F1, F2, F3, and F4 are the feature maps of the 1st, 3rd, 5th, and 8th layer convolutions of the feature extraction model respectively.
[0016] As a further improvement of the above solution, the feature extraction model is trained multiple times, and the trained feature extraction model is used as the feature extraction model adopted in step one.
[0017] Further, during the training process of the feature extraction model, three sets of images need to be input. The first set is the infrared light image I to be registered Inf , the second set is the visible light image I with the same content as the infrared light image to be registered pos , and the third set is the target visible light image I vis . First, the features of the three sets of images are respectively extracted by the feature extraction model, and the 1000-dimensional feature vector generated by the fully connected layer FC of the feature extraction model is used as the input of the triplet loss, further narrowing the feature expression between the I Inf and I pos images and reducing the feature difference between the infrared light image and the visible light image.
[0018] Further, the specific formula of the triplet loss TripletLoss is:
[0019]
[0020] where ||·|| is the Euclidean distance. Therefore represents the feature distance metric between the visible light image F(I pos ) with the same content as the infrared light image and the corresponding infrared light image F(I Inf ), represents the feature distance metric between the target visible light image F(I vis ) and the corresponding infrared light image F(I Inf ); a represents the minimum margin, and + means that when the value in [·] is greater than zero, the value is taken as the loss, and when the value is less than zero, the loss is zero, that is if the value in [·] is greater than zero, a loss will be generated. When the loss is zero; i represents the i-th coordinate point on the infrared light image, i = 1 - N, and N is a positive integer representing the number of selected coordinate points.
[0021] As a further improvement of the above solution, the correlation C AB at the k-th channel and the (i, j) coordinate position in the correlation graph is solved in the following way:
[0022] C AB (i, j, k) = F vis (i, j) T F Inf (i k , j k )
[0023] where k represents the k-th channel in the correlation graph, (i, j) represents the coordinate position on the infrared light fusion feature map, and (ik, jk) represents the coordinate position on the infrared light fusion feature map corresponding to (i, j); C AB (i, j, k) represents the pixel F at the k-th channel and the (i, j) coordinate position on the infrared light fusion feature map vis (i, j) to the pixel F at the k-th channel and the (i k ,j k ) coordinate position on the visible light fusion feature map Inf (i k ,j k ) correlation
[0024] As a further improvement of the above solution, the design of the regression network model:
[0025] The regression network includes a convolutional layer, Batch Normalization (Batch Normalization in Chinese), a ReLU (Rectified linear unit in English full name, Rectified linear unit in Chinese) activation function, and a fully connected layer FC (fully connected layers in English full name, fully connected layer in Chinese). The convolutional kernel sizes of the first four convolutional layers are all 7×7, and the convolutional kernel size of the last layer is 6×6. After passing through the fully connected layer FC, the output is a 6-dimensional vector
[0026] As a further improvement of the above solution, train the regression network model multiple times, and use the trained regression network model as the regression network model adopted in step four
[0027] The training of the regression network model inputs two sets of images. The first set is the infrared light image to be registered, and the second set is the target visible light image. The affine transformation parameters between the infrared light image and the visible light image are obtained through the regression network model, and the regression network model is tuned and optimized through the grid loss. The specific formula of the grid loss is:
[0028]
[0029] where i represents the i-th coordinate point on the infrared light image, N represents the number of selected coordinate points, gi represents the coordinate of the selected point on the infrared light image, that is, the i-th coordinate point, and T θ (g i ) represents the coordinate point corresponding to the predicted affine transformation of the coordinate point represents the coordinate point corresponding to the true affine transformation of the coordinate point
[0030] Compared with the existing methods, the present invention uses the VGG16 pre-trained model to extract the features of infrared and visible light images. At the same time, in order to better extract the consistent features of infrared and visible light images, the present invention fine-tunes the parameters of the pre-trained model by designing a triplet loss, further reducing the feature difference between infrared and visible light images, and improving the ability of the VGG16 pre-trained model to extract the consistent features of infrared and visible light. At the same time, through the feature fusion module, the multi-layer feature maps of infrared and visible light are fused respectively, making full use of the shallow and deep features extracted by the model, which can better express the features of infrared and visible light, and enabling the relevant layers to calculate the feature correlation between infrared and visible light more accurately. Finally, the inverse consistency loss is used to further constrain the regression model to generate affine transformation parameters, effectively improving the accuracy of infrared and visible light image registration. Description of the Drawings
[0031] Figure 1 It is a flowchart of the registration method for registering an infrared light image to a visible light image in Embodiment 1 of the present invention.
[0032] Figure 2 It is a registration network structure diagram adopted by the registration method for registering an infrared light image to a visible light image in Embodiment 2 of the present invention.
[0033] Figure 3 It is a feature fusion module diagram adopted by the registration method for registering an infrared light image to a visible light image in Embodiment 2 of the present invention.
[0034] Figure 4 It is a schematic diagram of the inverse consistency loss adopted by the registration method for registering an infrared light image to a visible light image in Embodiment 2 of the present invention. Detailed Embodiments
[0035] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0036] The registration method for registering an infrared light image to a visible light image of the present invention is based on an end-to-end network of deep learning, and is used to solve the problem of unsatisfactory registration caused by the large spectral difference between infrared and visible light images. Technically, the present invention overcomes the problem that the prior art cannot obtain accurate registration when facing images with large spectral differences, effectively improving the accuracy of infrared and visible light image registration, and obtaining good registration results on some data sets.
[0037] Embodiment 1
[0038] As Figure 1 shown, the registration method for registering an infrared light image to a visible light image includes the following steps.
[0039] Step 1: Extract the shallow infrared features and deep infrared features from an infrared light image, and extract the shallow visible features and deep visible features from a visible light image through a feature extraction model; the infrared light image and the visible light image are two different modality images of the same scene.
[0040] The present invention uses a deep convolutional network model, namely the VGG16 model, as the feature extraction model. The VGG16 pre-trained model is used to extract features from infrared and visible light images respectively. Since the pre-trained model is trained on visible light images, it cannot obtain good results in the feature extraction of infrared images. Therefore, the triplet loss is used during the network training process to better extract the consistent features between infrared and visible light images.
[0041] Step 2: Perform feature fusion on the shallow infrared features and the deep infrared features to form an infrared light fusion feature map with infrared light fusion features, and perform feature fusion on the shallow visible features and the deep visible features to form a visible light fusion feature map with visible light fusion features.
[0042] The present invention uses the VGG16 network structure as the backbone network of the feature extraction module (i.e., the feature extraction model), and initializes the feature extraction model parameters by loading the pre-trained VGG16 model; at the same time, the shallow features and deep features are fused through the feature fusion module. Specifically, the scale of the shallow feature map extracted by the feature extraction network is larger, while the scale of the deep feature map is smaller. Therefore, the shallow feature map is downsampled and then fused with the deep feature map.
[0043] During the model training process, three images are input, where two are the corresponding infrared and visible light images, and the other is an image obtained by randomly affine-transforming the visible light image. During the training process, the feature extraction network is further optimized through the triplet loss, and the regression network model is optimized through the grid loss and the inverse consistency loss; during the testing process, the two infrared and visible light images to be registered are input into the trained network model, and the affine transformation parameters for registering the infrared image to the visible light image can be obtained.
[0044] Step 3: Solve the correlation between each infrared light fusion feature in the infrared light fusion feature map and each visible light fusion feature in the visible light fusion feature map to obtain a correlation map with corresponding correlations.
[0045] Design of the correlation layer module and correlation calculation: From Step 1, the feature maps fA and fB of the infrared and visible light images can be generated respectively, where fA, fB ∈ Rh×w×d, that is, the sizes of the feature maps are both h×w×d. The correlation layer module takes the two feature maps as inputs, calculates the correlation between all pixels in the two feature maps, and generates the correlation map C AB . Calculate the correlation between the infrared and visible light features through the correlation layer module.
[0046] Step 4: Input the correlation map into a regression network model to predict the affine transformation parameters for registering the infrared light image to the visible light image.
[0047] Among them, the regression network can be composed of five convolutional layers, Batch Normalization, ReLU activation function, and a fully connected layer FC; input the correlation map C calculated in Step 2 AB into the regression network, and the affine transformation parameters corresponding to the infrared and visible light images can be predicted. Estimate the affine transformation parameters between the infrared and visible light images based on the correlation and the regression model.
[0048] Input the correlation between the infrared and visible light features into the regression network, and estimate the transformation parameters for mapping the infrared image to the visible light image and the transformation parameters for mapping the visible light image to the visible light image respectively. At the same time, through the inverse consistency loss, the regression network can predict more accurate affine transformation parameters, and finally register the infrared image to the visible light image according to the affine transformation parameters.
[0049] Step 5: Register the infrared light image to be registered to the corresponding visible light image through the affine transformation parameters. Align the infrared image with the visible light image according to the affine transformation result.
[0050] Embodiment 2
[0051] Please refer to Figure 2 、 Figure 3 、 Figure 4 . The registration method for registering the infrared light image to the visible light image in this embodiment includes the following steps.
[0052] Step 1.1, Design of the feature extraction module: Use the network structure of the deep convolutional network VGG16 as the feature extraction module.
[0053] Step 1.2, Feature extraction and feature fusion: Input the image into the feature extraction module, and extract the features of the infrared and visible light images respectively through convolutional operations; perform feature fusion using the feature maps of different sizes generated by different convolutional layers. Therefore, it is necessary to downsample the larger feature maps in the shallow layer and then splice them with the deep feature maps to generate the final fused feature map F; the feature fusion operation is expressed as:
[0054] F = Cat(Down(F1) 8 , Down(F2) 4 , Down(F3) 2 , F4)
[0055] Where F represents the fused feature map, Cat represents the concatenation of feature maps in the channel dimension, Down(·) represents the downsampling operation, and its subscript is the multiple of downsampling. F1, F2, F3, and F4 are the feature maps of the 1st, 3rd, 5th, and 8th convolutional layers of the feature extraction network respectively.
[0056] Step 2.1, Calculation of infrared and visible light feature correlation: Calculate the correlation of the fused feature maps F of infrared and visible light obtained in Step 1.2 Inf 、F vis to generate a correlation map C AB , where the correlation calculation formula is expressed as:
[0057] C AB (i, j, k) = F vis (i, j) T F Inf (i k , j k )
[0058] k = h(j k - 1) + i k
[0059] Where k represents the kth channel in the correlation map, and (i, j) and (i k , j k ) represent the position indices on the visible light feature map and the infrared feature map respectively.
[0060] Step 3.1, Design of the regression network model: The regression network consists of five convolutional layers, Batch Normalization, ReLU activation function, and a fully connected layer FC. The convolutional kernel sizes of the first four convolutional layers are all 7×7, the convolutional kernel size of the last layer is 6×6, and after passing through the fully connected layer FC, the output is a 6-dimensional vector.
[0061] Step 3.2, Input the correlation map C obtained in Step 2.1 AB into the regression network model, and the 6-degree-of-freedom affine transformation parameters corresponding to the infrared and visible light images can be predicted.
[0062] Step 4.1, Training of the feature extraction module: During the training of the feature extraction model, three groups of images need to be input. The first group is the infrared image I to be registered Inf , and the second group is the visible light image I with the same content as the infrared image pos, the third group is the target visible light image I vis , first, the features of the three groups of images are extracted by the feature extraction network respectively, and the 1000-dimensional feature vector generated by the fully connected layer of the feature extraction network is used as the input of the triplet loss, which can further narrow the feature expression of the feature extraction network for I Inf and I pos images, and reduce the feature difference between infrared and visible light images; the specific formula of the triplet loss is:
[0063]
[0064] where ||·|| is the Euclidean distance, so represents the feature distance metric between the visible light image with the same content as the infrared image and the infrared image, represents the feature distance metric between the target visible light image and the infrared image; a represents the minimum interval, and + means that when the value in [-] is greater than zero, the value is taken as the loss, and when the value is less than zero, the loss is zero, that is the value in will be greater than zero, and a loss will be generated. When the loss is zero.
[0065] Step 4.2, training of the regression network model: The training of the regression model requires two groups of images as input. The first group is the infrared image to be registered, and the second group is the target visible light image. The affine transformation parameters between the infrared and visible light images are obtained through the regression network model, and the regression network model is tuned and optimized through the grid loss. The specific formula of the grid loss is:
[0066]
[0067] where N represents the number of selected coordinate points, and g i represents the coordinates of the selected points on the infrared image, and T θ (g i ) represents the coordinate points corresponding to the predicted affine transformation of the coordinate points, represents the coordinate points corresponding to the true affine transformation of the coordinate points.
[0068] At the same time, in order to make the regression model more accurately predict the affine transformation between the infrared and visible light images, an inverse consistency loss is designed to enhance the constraint on the regression model. First, the affine transformation θ XY from the infrared image X mapped to the visible light image and the affine transformation θ YX from the visible light image mapped to the infrared image are obtained through the registration network respectively. Then, ideally, the product between θ XY and θ YX is the identity affine transformation. The specific formula of the inverse consistency loss ICLoss is:
[0069] ICLoss = GridLoss(θ XY *θ YX , ε)
[0070] Step 4.3, Testing of the network model: The testing of the registration model requires two sets of images as input. The first set is the infrared image to be registered, and the second set is the target visible light image. The affine transformation parameters between the infrared and visible light images can be obtained through the registration network, and the registration effect of the infrared and visible light images is evaluated by PCK and RMSE metrics. Among them, the PCK metric randomly selects 100 point coordinates and sets a threshold. If the error between the corresponding coordinates generated by the affine transformation predicted by the model and the coordinates generated by the true affine transformation does not exceed the threshold, it is considered a correct match. PCK represents the ratio of the number of correct match point pairs to the total number of match point pairs; the RMSE metric represents the average error of all corresponding point coordinates.
[0071] In summary, the registration method for registering an infrared light image to a visible light image according to the present invention is used to solve the problem of unsatisfactory registration caused by the large spectral difference between infrared and visible light images. The steps of this method include: Step 1, feature extraction is respectively performed on the infrared and visible light images through the VGG16 pre-trained model. Since the pre-trained model is trained on visible light images and cannot obtain good results in the feature extraction of infrared images, the triplet loss is used during the network training process to better extract the consistency features between the infrared and visible light images; Step 2, calculate the correlation between the infrared and visible light features through the relevant layer module; Step 3, estimate the affine transformation parameters between the infrared and visible light images based on the correlation and the regression model. Step 4, align the infrared image with the visible light image according to the affine transformation result. The present invention technically overcomes the problem in the prior art that accurate registration cannot be obtained due to the large spectral difference, effectively improving the accuracy of the registration of infrared and visible light images, and obtaining good registration effects on some data sets.
[0072] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A registration method for registering an infrared light image to a visible light image, characterized in that, it includes the following steps: Step 1: Extract infrared light shallow features and infrared light deep features in an infrared light image, and extract visible light shallow features and visible light deep features in a visible light image through a feature extraction model; the infrared light image and the visible light image are two different modality images in the same scene; Step 2: Perform feature fusion on the infrared light shallow features and the infrared light deep features to form an infrared light fusion feature map with infrared light fusion features, and perform feature fusion on the visible light shallow features and the visible light deep features to form a visible light fusion feature map with visible light fusion features; Step 3: Solve the correlation between each infrared light fusion feature in the infrared light fusion feature map and each visible light fusion feature in the visible light fusion feature map to obtain a correlation map with corresponding correlations; Step 4: Input the correlation map into a regression network model to predict the affine transformation parameters for registering the infrared light image to the visible light image; Step 5: The infrared light image to be registered is registered to the corresponding visible light image through the affine transformation parameters; Among them, three groups of images need to be input during the training of the feature extraction model. The first group is the infrared image I to be registered Inf , the second group is the visible light image I with the same content as the infrared image pos , and the third group is the target visible light image I vis . First, the features of the three groups of images are extracted through the feature extraction network respectively, and the 1000-dimensional feature vectors generated by the fully connected layer of the feature extraction network are used as the input of the triplet loss. The specific formula of the triplet loss is as follows: where ||·|| is the Euclidean distance, so represents the feature distance metric between the visible light image F(I pos ) that is the same as the content of the infrared light image and the corresponding infrared light image F(I Inf ); represents the feature distance metric between the target visible light image F(I vis ) and the corresponding infrared light image F(I Inf ); a represents the minimum interval, + means that when the value in [·] is greater than zero, take this value as the loss, and when the value is less than zero, the loss is zero, that is if the value in [·] is greater than zero, there will be a loss, when the loss is zero; i represents the i-th coordinate point on the infrared light image, i = 1 - N, and N is a positive integer representing the number of selected coordinate points.
2. The registration method for registering an infrared light image to a visible light image as described in claim 1, characterized in that, using the deep convolutional network model, namely the VGG16 model, as the feature extraction model.
3. The registration method for registering an infrared light image to a visible light image as described in claim 2, characterized in that, Extract infrared light shallow features and infrared light deep features from the infrared light image through convolution operations, and extract visible light shallow features and visible light deep features from the visible light image.
4. The registration method for registering an infrared light image to a visible light image as described in claim 3, characterized in that, The two feature fusion methods in Step 2 are both: F = Cat(Down(F1) 8 , Down(F2) 4 , Down(F3) 2 , F4) Among them, F represents the fusion feature map, referring to the infrared light fusion feature map or the visible light fusion feature map, Cat(·) represents the splicing of the fusion feature map in the channel dimension, Down(·) represents the downsampling operation on the shallow features in the image, and its subscript is the downsampling multiple. F1, F2, F3, and F4 are the convolutional feature maps of the 1st, 3rd, 5th, and 8th layers of the feature extraction model respectively.
5. The registration method for registering an infrared light image to a visible light image as described in any one of claims 1 to 4, characterized in that, Train the feature extraction model multiple times, and use the trained feature extraction model as the feature extraction model adopted in Step 1.
6. The registration method for registering an infrared light image to a visible light image as described in claim 1, characterized in that, The correlation C at the (i, j) coordinate position in the k-th channel of the correlation graph AB (i, j, k) is solved in the following manner: C AB (i, j, k) = F vis (i, j) T F Inf (i k , j k ) where k represents the k-th channel in the correlation graph, and (i, j) represents the coordinate position on the infrared light fusion feature map, (i k , j k ) represents the coordinate position on the infrared light fusion feature map corresponding to (i, j); C AB (i, j, k) represents the pixel F at the k-th channel and the (i, j) coordinate position on the infrared light fusion feature map vis (i, j) to the pixel F at the k-th channel and the (i k , j k ) coordinate position on the visible light fusion feature map Inf (i k , j k ) correlation.
7. The registration method for registering an infrared light image to a visible light image as described in claim 1, characterized in that, The design of the regression network model: The regression network includes a convolutional layer, Batch Normalization, ReLU activation function, and a fully connected layer FC. Among them, the convolutional kernel sizes of the first four convolutional layers are all 7×7, the convolutional kernel size of the last layer is 6×6, and after passing through the fully connected layer FC, the output is a 6-dimensional vector.
8. The registration method for registering an infrared light image to a visible light image according to claim 1, characterized in that, the regression network model is trained multiple times, and the trained regression network model is used as the regression network model adopted in step four; two sets of images are input for the training of the regression network model. The first set is the infrared light image to be registered, and the second set is the target visible light image. The affine transformation parameters between the infrared light image and the visible light image are obtained through the regression network model, and the regression network model is adjusted and optimized through the grid loss. The specific formula of the grid loss is: Among them, i represents the i-th coordinate point on the infrared light image, N represents the number of selected coordinate points, and g i represents the coordinates of the selected point on the infrared light image, that is, the coordinates of the i-th coordinate point, and T θ (g i ) represents the coordinate point corresponding to the predicted affine transformation of the coordinate point, and T θGT (g i ) represents the coordinate point corresponding to the true affine transformation of the coordinate point.
Citation Information
Patent Citations
Infrared image and visible image registration method based on visual attention
CN103714548A
Infrared and visible light image fusion method based on phase consistency and target enhancement
CN111462028A