Infrared and visible light image fusion method and system based on multi-scale information global integration and salient region preservation
By adopting a neural network with multi-scale information global integration and significant region maintenance in infrared and visible light images, combined with the synergistic effect of convolutional neural network and Transformer, the problems of global information loss and insufficient retention of significant region in the prior art are solved, and a higher quality image fusion effect is achieved.
Patent Information
- Application Number
- CN202510249011.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The existing infrared and visible image fusion algorithm based on autoencoder has limitations in feature extraction, which may cause the loss of some global information and insufficient retention of scene position information.
A neural network based on global integration of multi-scale information and significant region maintenance is adopted. Through the synergy between convolutional neural network and Transformer, a multi-scale perception module and a feature fusion module for significant region maintenance is constructed. Combined with the loss function of significance similarity, the multi-scale information integration of the image and the effective retention of significant regions are achieved.
It effectively makes up for the shortcomings of global information extraction, retains effective information of prominent areas in the source image, and improves the quality and visual effect of image fusion.
Smart Images

Figure CN120163718A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing, and particularly relates to an infrared and visible light image fusion method and system based on global integration of multi-scale information and preservation of significant regions. Background Art
[0002] Infrared and visible light image fusion refers to effectively combining the information of infrared images and visible light images to overcome their respective deficiencies, thereby obtaining a richer and more comprehensive image representation. Infrared and visible light image fusion is usually applied in low-light environments, night monitoring, bad weather conditions, etc., and has important practical significance. Specifically, infrared images can provide significant information about heat sources or targets in environments with insufficient visible light, while visible light images can provide higher spatial resolution and texture details. The fused image can not only clearly display infrared targets but also provide richer texture information, significantly improving the overall visual quality and providing an effective approach for tasks such as target tracking, monitoring, and target detection.
[0003] Due to differences in factors such as lighting, weather, and environment in real imaging scenarios or external and internal changes of imaging targets, the infrared and visible light image fusion task often faces problems of image information inconsistency and complementarity. To overcome these problems, considering the differences in spatial resolution, texture information, and other features between infrared and visible light images, the research on fusion algorithms has received extensive attention. Currently, infrared and visible light image fusion methods based on deep learning can generally be divided into algorithms based on convolutional neural networks, methods based on generative adversarial networks, and methods based on autoencoders. Methods based on convolutional neural networks perform poorly in the absence of real labels because their performance heavily depends on loss functions and network designs. Methods based on generative adversarial networks provide feasible solutions for image fusion, but they face challenges such as high training complexity, instability during the training process, and low resource utilization efficiency. The autoencoder-based image fusion method consists of an encoder, a fusion module, and a decoder, and has flexible fusion strategies, strong adaptability, and robust performance. However, existing autoencoder-based algorithms have limitations in feature extraction, which may cause loss of some global information and deficiencies in retaining scene location information.
[0004] The autoencoder algorithm that combines convolutional neural networks and Transformers can make up for the deficiency of global information extraction by extracting features from source images. However, there is a situation of local information loss in its extraction method based on independent double branches of source images. Currently, there are few algorithms based on global integration of multi-scale information that can exert the superiority of the autoencoder algorithm that combines convolutional neural networks and Transformers. Summary of the Invention
[0005] To solve the problems in the prior art, the present invention proposes an infrared and visible light image fusion method and system based on global integration of multi-scale information and preservation of significant regions.
[0006] The technical solution adopted by the present invention is as follows:
[0007] In a first aspect, the present invention discloses an infrared and visible light image fusion method based on global integration of multi-scale information and preservation of significant regions, including the following steps:
[0008] Step 1): Obtain the gray-scale images of the registered and aligned infrared and visible light images, and obtain the blue and red concentration offset maps of the visible light images;
[0009] Step 2): Construct a neural network based on global integration of multi-scale information and preservation of significant regions. The neural network takes the registered and aligned infrared and visible light images as inputs and finally outputs the fused gray-scale image. The neural network includes an encoder module for global integration of multi-scale information, a feature fusion module based on preservation of significant regions, and a feature decoder module;
[0010] Step 3): Construct a loss function based on saliency similarity to train the neural network based on global integration of multi-scale information and preservation of significant regions, and obtain a trained neural network;
[0011] Step 4): For the infrared and visible light images to be fused, use the trained neural network to obtain the fused gray-scale image, and construct an image post-processing module to perform image mode transformation on the fused gray-scale image output by the neural network to obtain the fused color image; realizing the fusion of infrared and visible light images.
[0012] In a second aspect, the present invention discloses an infrared and visible light image fusion system based on global integration of multi-scale information and preservation of significant regions for the above method, including:
[0013] An image preprocessing module, which is used to obtain the gray-scale images of the registered and aligned infrared and visible light images, and obtain the blue and red concentration offset maps of the visible light images;
[0014] A neural network construction module, which is used to construct a neural network based on global integration of multi-scale information and preservation of significant regions. The neural network is used to obtain the fused image based on the gray-scale images of the infrared and visible light images. The neural network includes an encoder module for global integration of multi-scale information, a feature fusion module based on preservation of significant regions, and a feature decoder module;
[0015] An image post-processing module, which is used to perform image mode transformation on the fused gray-scale image to obtain the fused color image;
[0016] A neural network training module, which is used to construct a loss function module based on saliency similarity, train a neural network based on global integration of multi-scale information and saliency region preservation, and obtain a trained neural network;
[0017] An infrared and visible light image fusion module, which is used to use the trained neural network to obtain the fusion image result of the infrared and visible light images to be fused, and realize the fusion of the infrared and visible light images.
[0018] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0019] 1) The present invention proposes an encoder module based on global integration of multi-scale information, effectively combines a convolutional neural network and a Transformer, gives full play to their respective advantages, and extracts multi-scale local features and global correlations at different levels.
[0020] 2) The present invention proposes a fusion module based on saliency region preservation, dynamically adjusts the fusion weight at the feature level through a spatial selection mechanism, generates a saliency attention map, and thus retains the effective information in the salient regions of the source images.
[0021] 3) The present invention proposes a loss function based on saliency similarity, adaptively calculates the global structural correlation of an image based on the extracted global features, and calculates the similarity loss of the salient regions using the attention map in the fusion module, retaining the effective information in the important regions. Description of the Drawings
[0022] Figure 1 It is a basic step flowchart of an embodiment of the infrared and visible light image fusion method of the present invention;
[0023] Figure 2 It is a schematic structural diagram of the infrared and visible light image fusion device of the present invention;
[0024] Figure 3 It is a schematic diagram of the multi-scale perception module of the present invention;
[0025] Figure 4 It is a schematic diagram of the feature fusion module based on saliency region preservation of the present invention;
[0026] Figure 5 It is a fusion result diagram of the MSRS infrared and visible light image dataset fused by the embodiment of the present invention and different methods. Detailed Embodiments
[0027] The present invention will be further described and explained below in conjunction with specific embodiments. The described embodiments are merely illustrative of the present disclosure and do not delimit the scope of limitation. The technical features of each embodiment of the present invention can be combined accordingly without conflict.
[0028] As Figure 1 shown, it is a basic step flowchart of an embodiment of the infrared and visible light image fusion method of the present invention, mainly including:
[0029] Step 1): Image preprocessing
[0030] For the infrared image and visible light image to be processed, in this embodiment, the method of mapping the RBG image mode to the YCbCr image mode is adopted, and the visible light image in the RBG mode is converted into the picture format in the YCbCr mode. W represents the width of the image, and H represents the height of the image. The calculation formula is as follows:
[0031] I VI-Y ,I VI-Cb ,I VI-Cr = RTY(I VI-RGB )
[0032] where, I VI-Y ,I VI-Cb ,I VI-Cr represent the luminance component, blue and red concentration offset component maps of the visible light picture in the YCbCr mode, and RTY(·) represents the operation of mapping the image from the RBG mode to the YCbCr mode. The specific calculation formula is as follows:
[0033] Y = 0.299·R + 0.587·G + 0.114·B
[0034] Cb = -0.1687·R - 0.3313·G + 0.5·B
[0035] Cr = 0.5·R - 0.4600·G - 0.0402·B
[0036] where, Y, Cb, and Cr are respectively the luminance component, blue and red concentration offset component maps of the image in the YCbCr mode, and R, G, and B are the three color channel values of the image in the RBG mode; then, for the registered and aligned infrared and visible light images, the infrared image I IR ∈R H×W×1 and the luminance component map I VI-Y of the visible light image are used as the inputs of the subsequent neural network.
[0037] Step 2): Construct a neural network based on global integration of multi-scale information and preservation of significant regions
[0038] The neural network is used to obtain a fused image based on the grayscale images of infrared and visible light images. The neural network includes an encoder module for globally integrating multi-scale information, a feature fusion module for preserving significant regions, and a feature decoder module. Among them, the encoder module for globally integrating multi-scale information is used to extract multi-source features containing rich multi-scale information and global correlation information from the source images. The feature fusion module for preserving significant regions obtains an adaptive significant region mask for the multi-source features and weights the multi-source features to obtain the final fused features. The feature decoder module processes the fused features and reconstructs the image to generate the fused image. In this embodiment, random parameters are used as the initial weights of the network.
[0039] The specific working processes of each module in the neural network based on global integration of multi-scale information and preservation of significant regions are as follows:
[0040] Step 21): The encoder module for globally integrating multi-scale information includes B serial multi-scale perception modules and a global feature extraction module based on Transformer. The structure of the multi-scale perception module is as Figure 3 shown. The multi-scale perception module first constructs multi-level multi-scale features based on the input preprocessed infrared image I IR and the luminance component map I VI-Y of the visible light image. The calculation formula is as follows:
[0041] F (1) = MP(I)
[0042] F (i) = MP(F (i-1) )
[0043] where MP(·) represents the multi-scale perception module, I ∈ {I IR , I VI-Y}, i = 1, 2,..., B. In this embodiment, B = 5. The output of the (i - 1)-th multi-scale perception module will continue to be fed into the i-th multi-scale perception module. The multi-level outputs are concatenated to obtain the final multi-scale information features. The calculation formula is as follows:
[0044] Φ L = Concat(F (1) , F (2) ,..., F (B) )
[0045] where represents the multi-level multi-scale information features, and Concat(·) represents the channel concatenation operation; is the multi-level multi-scale information features of the infrared image, is the multi-level multi-scale information features of the visible light image.
[0046] Then, for the extracted multi-level and multi-scale information features To further establish the global correlation between information at different scales, a global feature extraction module based on Transformer is constructed, and its structure is as Figure 1 shown. First, the multi-level and multi-scale information features are downsampled and embedded into the hidden layer, and the dense feature blocks are encoded to prepare for global context extraction. The global feature extraction module based on Transformer contains L Transformer layers, and each layer consists of a normalization layer, a multi-head attention mechanism, and a multi-layer perceptron. When the encoded features pass through the L Transformer layers, the features are gradually upsampled to restore the spatial dimension, and the output features at different scales are concatenated with the features in the downsampling stage to output an infrared image feature tensor containing rich local context information and a visible light image feature tensor Therefore, the encoder outputs a total of 4 features:
[0047] Step 22): The features output by the encoder module include features with multi-scale information and the global features after global integration of local information The present invention designs a feature fusion module based on significant region preservation, which is used to generate a set of adaptive weights for each feature, and the fused feature is obtained by weighting based on the weights. Its structure is as Figure 4 shown. The feature fusion module based on significant region preservation first concatenates the input four features, and then performs operations on the concatenated feature Φ using an average pooling layer and a max pooling layer. The formula is as follows:
[0048]
[0049] where, and represent the average pooling operation and the max pooling operation respectively, and PM avg , PM max are their corresponding saliency feature descriptors respectively. To promote the interaction of significant information between different source features, the saliency feature descriptors PM avg and PM max are aggregated into a single saliency feature descriptor. Then, through a convolution operation it is converted into the corresponding N saliency attention maps, and the expression is as follows:
[0050]
[0051] where, represents the saliency attention map, and in this embodiment, N = 4. For each saliency attention map Apply the Sigmoid activation function to obtain the corresponding saliency selection mask for extracting features from infrared and visible light images. The calculation formula is as follows:
[0052]
[0053] Among them, represents the saliency selection mask, and ρ(·) represents the Sigmoid activation function for generating dynamic fusion weights. Then, the features extracted from infrared and visible light images are weighted and fused through the saliency selection mask to obtain the fused feature tensor:
[0054]
[0055] The feature fusion module based on saliency region preservation finally outputs the fused features containing saliency information from multiple source images
[0056] Step 23): The feature decoder module aims to restore the fully fused features to the original size of the initial input image and output the fused grayscale image. The whole process is implemented based on a multi-layer neural network and is expressed by the formula:
[0057]
[0058] Among them, is the fused grayscale image generated by the model. LeakyReLU and Softmax represent the Leaky ReLU non-linear activation function and the Softmax non-linear activation function respectively, and Conv represents the two-dimensional convolutional layer.
[0059] Step 3: Neural network model training
[0060] Construct a loss function based on saliency similarity, which is used to guide the neural network. The loss function based on saliency similarity consists of three parts: gradient loss, mean square error loss, and saliency similarity loss. By combining these loss functions, it can help the model learn local and global information, retain salient features, and at the same time maintain the structural integrity of the fused image. The calculation formula of the loss function is as follows:
[0061] L = λ1L grad + λ2L MSE + λ3L S-SIM
[0062] Among them, L grad represents the gradient loss, L MSE represents the mean square error loss, L S-SIMIndicates the significant similarity loss. λ1, λ2, λ3 are the balance weights of each loss part. In this embodiment, λ1 = 50, λ2 =
[0063] 5, λ3 = 2.
[0064] The gradient loss part is used to preserve the details and texture information of the fused image. It calculates the gradient difference at the corresponding positions by comparing the gradients of the fused image and the source image, squares them and averages over all pixels. The calculation formula is as follows:
[0065]
[0066] Where, represents the gradient of the image at pixel i, N is the total number of pixels in the image, and max represents the operation of taking the larger value.
[0067] The mean square error loss part evaluates the fusion effect by calculating the square difference between the fused image and the source image at each pixel position. This loss helps to make the fused image as close as possible to the source image. The calculation formula is as follows:
[0068]
[0069] Where, α1 and α2 are parameters used to balance the mean square error loss part between the infrared image and the visible image. In this embodiment, α1 = α2 = 1.
[0070] The significant similarity loss part includes two main components, the significant region similarity loss and the global structure similarity loss. The expression is as follows:
[0071] L S-SIM = β1L SRS + β2L GSS
[0072] Where, L SRS represents the significant region similarity loss, and L GSS represents the global structure similarity loss. β1 and β2 are the corresponding weights. In this embodiment, β1 = 1, β2 = 5.
[0073] The significant region similarity loss aims to retain the effective information in the significant regions of the image. By using the feature selection mask of the feature fusion module based on significant region preservation to adaptively fuse the intensity information of the image, the calculation formula is as follows:
[0074]
[0075] Where, δ represents the step function, and W1 and W2 are the weights generated based on the significant feature selection masks of the visible image and the infrared image.
[0076] The global structural similarity loss is used to preserve the global structure of the image, and the structural similarity is calculated by comparing the global features obtained by the Transformer-based feature extraction module. First, the mean μ of the global features IR and μ VI are used to calculate the weights W3 and W4 for the global structural similarity. The specific formulas are as follows:
[0077]
[0078] where MEAN(·) represents the operation of calculating the mean. Based on the calculated weights W3 and W4, the structural similarity between the fused image and the source image is further calculated.
[0079] Finally, the global structural similarity loss is calculated by weighting based on the weight map of the global features. The calculation formula is as follows:
[0080]
[0081] where represents the structural similarity between the infrared image I IR and the fused grayscale image I FUSE-Y , represents the structural similarity between the luminance component map I VI-Y of the visible light image and the fused grayscale image I FUSE-Y .
[0082] Using the entire infrared and visible light images as training samples, the network adopts an unsupervised training method. The samples are input into the network model. Based on the constructed saliency similarity loss function module, the network weights are updated using the Adam gradient descent method with an adaptive learning rate adjustment. In this embodiment, the training is iterated 60 times in total.
[0083] Step 4: Based on the trained neural network model, the pre-processed infrared and visible light images are used as the input of the network, and the output of the neural network based on the global integration of multi-scale information and the preservation of significant regions is post-processed. The image post-processing steps are as follows:
[0084] For the fused grayscale image, in this embodiment, the method of mapping the YCbCr image mode to the RGB image mode is adopted, and the fused grayscale image I Fuse-Y in the YCbCr mode is converted into the picture format in the RGB mode. The calculation formula is as follows:
[0085] I FUSE = YTR(I FUSE-Y , I VI-Cb , I VI-Cr )
[0086] Among them, is the fused color image, I VI-Cb , I VI-Cr represents the blue concentration offset component map and the red concentration offset component map of the visible light image in the YCbCr color space. YTR(·) represents the method of mapping the image from the YCbCr color space to the RGB color space. The specific calculation formula is as follows:
[0087] R = Y + 1.402·(Cr - 128)
[0088] G = Y - 0.344136·(Cb - 128) - 0.714136·(Cr - 128)
[0089] B = Y + 1.772·(Cb - 128)
[0090] Finally, through the image post - processing module, a color infrared and visible light fused image with three channels is obtained, realizing the result of infrared and visible light image fusion.
[0091] Figure 2 Shown in Figure 2 is an infrared and visible light image fusion system based on global integration of multi - scale information and preservation of salient regions according to the embodiment. The system includes:
[0092] An image pre - processing module, which is used to obtain the grayscale images of the registered and aligned infrared and visible light images, and obtain the blue and red concentration offset maps of the visible light image;
[0093] A neural network construction module, which is used to construct a neural network based on global integration of multi - scale information and preservation of salient regions. The neural network is used to obtain the fused image based on the grayscale images of the infrared and visible light images. The neural network includes an encoder module for global integration of multi - scale information, a feature fusion module for preserving salient regions, and a feature decoder module;
[0094] An image post - processing module, which is used to perform image mode transformation on the fused grayscale image to obtain the fused color image;
[0095] A neural network training module, which is used to construct a loss function module based on saliency similarity, train the neural network based on global integration of multi - scale information and preservation of salient regions, and obtain the trained neural network;
[0096] An infrared and visible light image fusion module, which is used to use the trained neural network to obtain the fusion image result of the infrared and visible light images to be fused, realizing the fusion of the infrared and visible light images.
[0097] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0098] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions of the method embodiments. The device embodiments described above are only illustrative. For example, the image preprocessing module can be a logical function division, and in actual implementation, there can be other division methods. For example, multiple modules can be combined or integrated into another unit. Another point is that the connections between the displayed or discussed modules can be communication connections through some interfaces, which can be electrical or other forms. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present application. Those of ordinary skill in the art can understand and implement it without creative efforts. The following takes real infrared and visible light images as examples to illustrate the specific implementation manners to reflect the technical effects of the present invention, and the specific steps in the embodiments will not be repeated.
[0099] Embodiment
[0100] Next, taking the MSRS infrared and visible light image dataset as the research object, the fusion method of the present invention is verified. In order to comprehensively compare the fusion performance and display the fusion results from the perspectives of visualization and quantification, the fusion result graph and evaluation indicators: the mean square error (MSE), peak signal-to-noise ratio (PSNR), artifact-based metric N AB / F , correlation coefficient (CC), and multi-scale structural similarity (MS-SSIM) are used to evaluate the proposed fusion method; the formulas are as follows:
[0101]
[0102] Among them, MSE X,F represents the calculation of the mean square error between the source image X and the fused image F, where A represents the visible light image and B represents the infrared image. A smaller mean square error value indicates a smaller difference between the fused image and the source image, that is, a better fusion effect.
[0103] The peak signal-to-noise ratio quantifies the ratio of the maximum signal intensity to the noise level in the fused image. It indicates the degree of distortion introduced at the pixel level during the fusion process. The specific calculation formula is as follows:
[0104]
[0105] Among them, r represents the peak of the fused image. A larger peak signal-to-noise ratio indicates less distortion during the fusion process, resulting in better fusion quality.
[0106] NAB / F This metric is used to measure the degree of artifacts introduced during the fusion process. The specific calculation formula is as follows
[0107]
[0108] where measures the localization of artifacts in the fused image. Specifically, if the edge intensity at a certain position in the fused image is greater than that at the corresponding position in the source image, then that position is regarded as an artifact. and respectively represent the edge intensities of the fused image and the source image at position i. A lower N AB / F value indicates fewer artifacts introduced during the fusion process.
[0109] The correlation coefficient is used to evaluate the linear relationship between the fused image and the original image. The specific calculation formula is as follows:
[0110]
[0111] where A represents the visible light image and B represents the infrared image. X i and F i respectively represent the pixel values of the source image and the fused image at position i, and respectively represent the means of the source image and the fused image. A higher CC value indicates a higher similarity between the fused image and the source image.
[0112] As a metric for multi-scale analysis of image structure, multi-scale structural similarity can capture the structural features of an image more comprehensively. The calculation process of multi-scale structural similarity is as follows: First, perform operations such as downsampling on the image to obtain image pairs of different sizes. Then, calculate the structural similarity values of the image pairs at each scale. Finally, perform weighted calculation on the structural similarity values at different scales to obtain the multi-scale structural similarity. The specific calculation formula is as follows:
[0113] MS-SSIM X,F = 1 - MSSSIM(F, X)
[0114] MS-SSIM = MS-SSIM A,F + MS-SSIM B,F
[0115] where MSSSIM(·) represents the calculation of multi-scale structural similarity. A higher multi-scale structural similarity value indicates a higher structural consistency between the fused image and the source image.
[0116] The MSRS dataset contains a total of 1444 groups of infrared and visible light images, and each image contains 480×640 pixels.
[0117] Figure 5 The fusion result graph obtained by the embodiments of the present invention and different fusion algorithms for the MSRS dataset.
[0118] Table 1 Evaluation metrics for the fusion results of the MSRS infrared and visible image datasets
[0119]
[0120] The comparison method MDLatLRR is from: H. Li, X.-J. Wu, and J. Kittler, “Mdlatlrr: A novel decomposition method for infrared and visible image fusion,” IEEE Transactions on Image Processing, vol. 29, pp. 4733–4746, 2020.
[0121] The comparison method FusionDN is from: H. Xu, J. Ma, Z. Le, J. Jiang, and X. Guo, “Fusiondn: A unified densely connected network for image fusion,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, no. 07, pp. 12 484–12 491, 2020.
[0122] The comparison method ReCoNet is from: Z. Huang, J. Liu, X. Fan, R. Liu, W. Zhong, and Z. Luo, “Reconet: Recurrent correction network for fast and efficient multi-modality image fusion,” in Computer Vision–ECCV 2022, 2022, vol. 13678, pp. 539–555.
[0123] The comparison method SeAFusion is from: L. Tang, J. Yuan, and J. Ma, “Image fusion in the loop of high-level vision tasks: A semantic-aware real-time infrared and visible image fusion network,” Information Fusion, vol. 82, pp. 28–42, 2022.
[0124] The comparison method PIAFusion is from: L. Tang, J. Yuan, H. Zhang, X. Jiang, and J. Ma
[0125] “Piafusion: A progressive infrared and visible image fusion network based on illumination aware,” Information Fusion, vol. 83–84, pp. 79–92, 2022.
[0126] The comparison method TarDAL is from: J. Liu, X. Fan, Z. Huang, G. Wu, R. Liu, W. Zhong, and Z. Luo, “Target-aware dual adversarial learning and a multi-scenario multi-modality benchmark to fuse infrared and visible for object detection,” in 2022 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 5792–5801.
[0127] The comparison method SwinFusion is from: J. Ma, L. Tang, F. Fan, J. Huang, X. Mei, and Y. Ma, “Swinfusion: Cross-domain long-range learning for general image fusion via swin transformer,” IEEE / CAA Journal of Automatica Sinica, vol. 9, no. 7, pp. 1200–1217, 2022.
[0128] The comparison method CDDFuse is from: Z. Zhao, H. Bai, J. Zhang, Y. Zhang, S. Xu, Z. Lin, R. Timofte, and L. Van Gool, “Cddfuse: Correlation-driven dual-branch feature decomposition for multi-modality image fusion,” 2023.
[0129] The comparison method NestFuse is from: H. Li, X.-J. Wu, and T. Durrani, “Nestfuse: An infrared and visible image fusion architecture based on nest connection and spatial / channel attention models,” IEEE Transactions on Instrumentation and Measurement, vol. 69, no. 12, pp. 9645–9656, 2020.
[0130] The comparison method DIDFuse is from: Z. Zhao, S. Xu, C. Zhang, J. Liu, P. Li, and J. Zhang,
[0131] “Didfuse: Deep image decomposition for infrared and visible image fusion,” in Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, 2020, pp. 970–976.
[0132] The comparison method DeFusion is from: P. Liang, J. Jiang, X. Liu, and J. Ma, “Fusion from decomposition: A self-supervised decomposition approach for image fusion,” in Computer Vision–ECCV 2022, 2022, vol. 13678, pp. 719–735.
[0133] The comparison method CrossFuse is from: H. Li and X.-J. Wu, “Crossfuse: A novel cross-attention mechanism based infrared and visible image fusion approach,” Information Fusion, vol. 103, pp. 102 147–102 147, 2024.
[0134] The fusion results of the MSRS dataset are as Figure 5 shown, and the quantization results are shown in Table 1. In the quantization results, the present invention achieved the best performance and had a good suppression effect on the noise and artifacts existing in the source images. This algorithm extracts features based on a neural network for global integration of multi-scale information and preservation of significant regions, and has a good expression in terms of scene, location, and detailed texture when fusing images; in addition, the feature fusion module based on the preservation of significant regions and the loss function module based on significant similarity can provide effective information retention and a richer fusion result.
[0135] The above-described embodiments merely represent several implementation manners of the present invention, and the description thereof is relatively specific and detailed, but should not be construed as a limitation on the scope of the present invention. For those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. A method for fusion of infrared and visible light images based on global integration of multi-scale information and preservation of salient regions, characterized in that: The following steps are involved: Step 1): obtaining a grayscale image of infrared and visible light images, and obtaining a blue and red concentration offset map of the visible light image; Step 2): constructing a neural network based on global integration of multi-scale information and preservation of salient regions, wherein the neural network takes the registered and aligned infrared and visible light images as input and finally outputs a fused grayscale image, and the neural network includes an encoder module for global integration of multi-scale information, a feature fusion module based on preservation of salient regions, and a feature decoder module; Step 3): construct a loss function based on saliency similarity, train the neural network based on global integration of multi-scale information and preservation of salient regions, and obtain a trained neural network; Step 4): For the fused infrared and visible light images, a fused grayscale image is obtained using a trained neural network, and an image post-processing module is constructed to perform image mode transformation on the fused grayscale image output by the neural network to obtain a fused color image; thus, the fusion of infrared and visible light images is realized.
2. The infrared and visible light image fusion method based on multi-scale information global integration and salient region preservation according to claim 1 is characterized in that: The step 1) comprises: Based on the input aligned infrared image and visible light image, the visible light image in RBG mode is Convert to the image format in YCbCr mode, W represents the width of the image, H represents the height of the image, and the calculation formula is as follows: I VI-Y ,I VI-Cb ,I VI-Cr =RTY(I VI-RGB ) Among them, I VI-Y ,I VI-Cb ,I VI-Cr Represents the brightness component, blue density offset component map and red density offset component map of the visible light image in YCbCr mode. RTY(·) represents the operation of mapping the image from RBG mode to YCbCr mode. Then, for the aligned infrared image and visible light image, the infrared image I IR ∈R H×W×1 And the brightness component map of the visible light image I VI-Y As the input of the subsequent neural network.
3. The infrared and visible light image fusion method based on multi-scale information global integration and salient region preservation according to claim 1, characterized in that: The encoder module based on global integration of multi-scale information in step 2) is used to mine multi-source features containing rich multi-scale information and global correlation information from the source image, wherein the source image is a brightness component map of an infrared image and a visible light image; the feature fusion module based on salient region preservation obtains an adaptive salient region mask for the multi-source features, and weights the features to obtain the final fusion feature; The feature decoder module processes the fused features and reconstructs the image to generate a fused image.
4. The infrared and visible light image fusion method based on multi-scale information global integration and salient region preservation according to claim 3 is characterized in that: The encoder module based on global integration of multi-scale information includes B serial multi-scale perception modules and a Transformer-based global feature extraction module; The multi-scale perception module firstly inputs the pre-processed infrared image I IR Brightness component diagram of visible light image I VI-Y , construct multi-level and multi-scale features, and the calculation formula is as follows: F (1) =MP(I) F (i) =MP(F (i-1) ) Among them, MP(·) represents the multi-scale perception module, I∈{I IR ,I VI-Y }, i = 1, 2, ..., B; the output of the i-1th multi-scale perception module will continue to be sent to the i-th multi-scale perception module; the multi-level outputs are spliced to obtain the final multi-scale information features, and the calculation formula is as follows: Φ 1 =Concat(F (1) ,F (2) ,…,F (B) ) in, It expresses multi-level and multi-scale information features, and Concat(·) represents the channel concatenation operation; is the multi-level and multi-scale information feature of the infrared image. It is the multi-level and multi-scale information feature of visible light images; Then, for the extracted multi-level and multi-scale information features and Construct a Transformer-based global feature extraction module; first, downsample the multi-level and multi-scale information features and embed them into the hidden layer, encode the dense feature blocks, and prepare for global context extraction; the Transformer-based global feature extraction module contains L Transformer layers, each of which consists of a normalization layer, a multi-head attention mechanism, and a multi-layer perceptron; after the encoded features pass through L Transformer layers, the features are gradually upsampled to restore the spatial dimension, and the output features of different scales are spliced with the features of the downsampling stage to output the infrared image feature tensor containing rich local context information and the visible light image feature tensor Therefore, the encoder outputs a total of 4 features:
5. The infrared and visible light image fusion method based on multi-scale information global integration and salient region preservation according to claim 4, characterized in that: The feature fusion module based on salient region preservation first integrates the four input features After splicing, the average pooling layer and the maximum pooling layer are performed on the spliced feature Φ. The formula is as follows: in, and Represents average pooling operation and maximum pooling operation, PM avg ,PM max are the corresponding salient feature descriptors; in order to promote the interaction of salient information between different source features, the salient feature descriptor PM avg and PM max are aggregated into a single salient feature descriptor; then, after the convolution operation Converted into the corresponding 4 saliency attention maps, the expression is as follows: in, Represents 4 salient attention maps; for each salient attention map The Sigmoid activation function is applied to obtain the corresponding saliency selection mask, which is used to extract the features of infrared images and visible light images. The calculation formula is as follows: in, represents the saliency selection mask, ρ(·) represents the Sigmoid activation function, which is used to generate dynamic fusion weights. Then, the features extracted from the infrared and visible light images are weightedly fused by the saliency selection mask to obtain the fused feature tensor: The feature fusion module based on salient region preservation finally outputs fused features containing salient information from multiple source images.
6. The infrared and visible light image fusion method based on multi-scale information global integration and salient region preservation according to claim 5, characterized in that: The feature decoder module is implemented based on a multi-layer neural network and outputs a fused grayscale image, which is expressed as: in, The fused grayscale image generated by the model, LeakyReLU and Softmax represent the Leaky ReLU non-linear activation function and the Softmax non-linear activation function respectively, and Conv represents the two-dimensional convolutional layer.
7. The infrared and visible light image fusion method based on multi-scale information global integration and salient region preservation according to claim 1, characterized in that: In the step 4), The fused grayscale image I in YCbCr mode Fuse-Y Convert to the image format in RGB mode, the calculation formula is as follows: I FUSE =YTR(I FUSE-Y ,I VI-Cb ,I VI-Cr ) in, is the fused color image, I VI-Cb ,I VI-Cr Represents the blue density offset component map and the red density offset component map of the visible light image in the YCbCr mode. YTR(·) represents the operation of mapping the image from the YCbCr mode to the RBG mode. Finally, after the image post-processing module, a three-channel color infrared and visible light fusion image is obtained.
8. The infrared and visible light image fusion method based on multi-scale information global integration and salient region preservation according to claim 6, characterized in that: In the step 3), The loss function L based on saliency similarity consists of three components: gradient loss L grad , mean square error loss L MSE and the saliency similarity loss L 2-SIM ; The calculation formula of the loss function is as follows: L=λ1L grad +λ2L MIE +λ3L S-SIM Among them, λ1,λ2,λ3 are the balance weights of each loss part; The calculation formula of the gradient loss part is as follows: in, Represents the gradient of the image at pixel i, N is the total number of pixels in the image, and max represents the larger value operation; The mean square error loss calculation formula is as follows: Among them, α1 and α2 are parameters used to balance the mean square error loss between infrared images and visible images; The saliency similarity loss part includes saliency region similarity loss and global structure similarity loss; the expression is as follows: L S-SIM =β1L SRS +β2L GSS Among them, L SRS represents the significant region similarity loss, L GSS represents the global structural similarity loss, β1 and β2 are the corresponding weights; The salient region similarity loss is achieved by using a feature selection mask based on the feature fusion module based on salient region preservation. The intensity information of the image is adaptively fused, and the calculation formula is as follows: Where δ represents a step function, W1 and W2 are weights for mask generation based on the salient features of visible and infrared images; The global structural similarity loss is used to preserve the global structure of the image. The structural similarity is calculated by comparing the global features obtained by the Transformer-based feature extraction module. First, the mean μ of the global features is used to calculate the structural similarity. IR and μ VI Calculate the weights W3 and W4 of global structural similarity. The specific formula is as follows: Wherein, MEAN(·) represents the mean operation; based on the calculated weights W3 and W4, the structural similarity between the fused image and the source image is further calculated; Finally, the global structural similarity loss is weighted by using a weight map based on global features. The calculation formula is as follows: in, Represents infrared image I IR And the fused grayscale image I FUSE-Y The structural similarities between Represents the brightness component image I of the visible light image VI-Y And the fused grayscale image I FUSE-Y The structural similarities between The loss function based on the significant similarity is used as the loss function, and the weight of the neural network is updated based on the gradient descent method with adaptively adjusted learning rate.
9. An infrared and visible light image fusion system based on multi-scale information global integration and salient region preservation for implementing the method of claim 1, characterized in that: include: An image preprocessing module is used to obtain a grayscale image of the infrared image and the visible light image registration and alignment, and obtain a blue and red concentration offset map of the visible light image; A neural network building module, which is used to construct a neural network based on global integration of multi-scale information and preservation of salient regions, wherein the neural network is used to obtain a fused image based on grayscale images of infrared and visible light images, and the neural network includes an encoder module for global integration of multi-scale information, a feature fusion module based on preservation of salient regions, and a feature decoder module; An image post-processing module is used to perform image mode transformation on the fused grayscale image to obtain a fused color image; A neural network training module is used to construct a loss function module based on saliency similarity, train a neural network based on global integration of multi-scale information and preservation of salient regions, and obtain a trained neural network; The infrared and visible light image fusion module is used to obtain the fusion image result of the infrared and visible light images to be fused by using the trained neural network, so as to realize the fusion of the infrared and visible light images.
Citation Information
Patent Citations
Infrared and visible light image fusion method based on multi-mode features
CN114639002A
Fusion network construction method for multispectral images and corresponding fusion method
CN115909000A
Multi-modal medical image fusion method based on multi-scale transformer
CN115984257A
Saliency Prioritization for Image Processing
US20220198258A1