Deep Learning-Based Infrared Image and Visible Light Image Fusion Method and System
By constructing a U-shaped network and two-stage training image fusion method, using wavelet convolution and Kolmogolov-Arnold feature extraction module, the problems of inaccurate feature extraction and insufficient loss function in infrared images and visible light images are solved, and high-quality image fusion effect is achieved.
Patent Information
- Application Number
- CN202510585967.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-08
AI Technical Summary
The existing infrared image and visible image fusion technology have problems such as inaccurate feature extraction and insufficient design of fixed activation functions for network expression capabilities and loss function, resulting in poor fusion quality in complex scenarios.
The image fusion network is constructed, the U-shaped network structure is adopted, and the encoder and fusion module are optimized through two-stage training strategies, and the multi-scale features are extracted using wavelet convolution and the Kolmogorov-Arnold feature extraction module, and the adaptive loss function is designed to improve the fusion quality.
The visual effect and objective evaluation indicators of the fusion image are improved, which is better than the existing methods, and better preserve texture details and information in complex scenarios.
Smart Images

Figure CN120107089B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly relates to a method and system for fusing infrared images and visible light images based on deep learning. Background Art
[0002] Image fusion is an image processing technology that combines multi-modal images of the same scene to generate an image that can accurately and comprehensively describe the scene. Infrared images can work in low light, dark or harsh environments. By capturing the thermal radiation information of objects, they provide temperature distribution, heat source detection and concealed target recognition. However, they cannot accurately capture the details and color information of objects, and are easily affected by changes in environmental temperature and humidity. Visible light images, on the other hand, can provide rich detail and color information and are suitable for most daily scenes, but perform poorly in low light or dark environments and are limited by lighting conditions. Therefore, fusing infrared images and visible light images can make full use of the advantages of both, enabling stronger perception capabilities in complex low light, night or variable environments. This technology is widely used in military reconnaissance, target tracking, intelligent driving, security monitoring and other fields.
[0003] Currently, infrared image and visible light image fusion technologies can be roughly divided into traditional fusion methods and deep learning-based fusion methods. Traditional fusion methods include methods based on multi-scale transformation, sparse representation, subspace, saliency, total variation, etc. However, these traditional fusion methods have limitations. Their fusion strategies rely too much on manual design and cannot adapt to increasingly complex fusion scenarios, resulting in limited fusion performance. Deep learning-based methods include fusion methods based on autoencoders, generative adversarial networks, and convolutional neural networks. Due to the strong feature extraction ability of convolutional operations and attention mechanisms and the drive of big data, deep learning-based fusion methods have achieved satisfactory fusion effects. However, there are still some deficiencies in existing deep learning-based fusion algorithms. First, in terms of feature extraction, existing methods all use ordinary convolutions for feature extraction, and the extracted features contain information with mixed frequencies, making it difficult for the model to distinguish features of different frequencies, resulting in limited fusion quality. Second, in traditional deep learning networks, a fixed activation function is usually used to introduce non-linearity to help the network learn complex mapping relationships. However, the choice of this fixed activation function limits the expressive power of the network, making it have certain limitations in processing complex features. Third, there are certain defects in the design of the loss function in existing infrared and visible light image fusion methods, resulting in weak ability to preserve image texture details in complex scenes and easy loss of important information. Summary of the Invention
[0004] The object of the present invention is to provide a method and system for fusing infrared images and visible light images based on deep learning, which can improve the quality of the fused images.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] On the one hand, an embodiment of the present invention provides a method for fusing infrared images and visible light images based on deep learning, the method comprising the following steps:
[0007] Construct an image fusion network, the image fusion network comprising an encoder, a fusion module and a decoder; the encoder comprises an infrared encoder and a visible light encoder;
[0008] Form a U-shaped network with the encoder and the decoder, and train the U-shaped network based on a first-stage loss function to obtain a trained encoder;
[0009] Fix the parameters of the encoder, and train the image fusion network based on a second-stage loss function to obtain a trained image fusion network;
[0010] Use the trained image fusion network to fuse the input infrared image and visible light image to generate a fused image.
[0011] Preferably, the step of using the trained image fusion network to fuse the input infrared image and visible light image to generate a fused image comprises:
[0012] Obtain the input infrared image and visible light image;
[0013] Input the infrared image and the visible light image into the encoder to output multi-layer feature maps; the feature maps include infrared feature maps and visible light feature maps;
[0014] Concatenate the infrared feature maps and the visible light feature maps in each layer of the feature maps in the channel direction to obtain multi-layer concatenated maps;
[0015] Input the multi-layer concatenated maps into the fusion module for processing to obtain fused feature maps;
[0016] Input the fused feature maps into the decoder for processing to obtain a fused image.
[0017] Preferably, the encoder comprises an infrared encoder and a visible light encoder, and both the infrared encoder and the visible light encoder comprise 3 wavelet convolution modules for the encoder and 2 block-based Kolmogorov-Arnold feature extraction modules cascaded in sequence; the wavelet convolution module for the encoder is successively composed of a common convolution layer, a reflection padding layer, a wavelet convolution layer, a batch normalization layer and a rectified linear unit;
[0018] The calculation formula of the infrared feature map is as follows:
[0019] ;
[0020] ;
[0021] The calculation formula of the visible light feature map is as follows:
[0022] ;
[0023] ;
[0024] Among them, and respectively represent the feature maps of the infrared encoder and the visible light encoder at the th layer; and respectively represent the input infrared image and visible light image, , , , respectively represent the number of channels, height and width of the image; and respectively represent the wavelet convolution modules for the encoder in the th layer of the infrared encoder and the visible light encoder; and respectively represent the block-based Kolmogorov-Arnold feature extraction modules in the th layer of the infrared encoder and the visible light encoder.
[0025] Preferably, the fusion module includes a channel attention mechanism, a large kernel attention mechanism and a convolution module, and the convolution module includes a common convolution layer, a reflection padding layer, a batch normalization layer and a rectified linear unit;
[0026] The process of inputting the multi-layer spliced image into the fusion module to obtain a fusion feature map includes:
[0027] Inputting the multi-layer spliced image into the corresponding channel attention mechanism respectively to obtain a multi-layer channel spliced image; The formula of the channel spliced image is expressed as follows:
[0028] ;
[0029] Among them, represents the channel attention feature map of the th layer, represents global average pooling, Denotes a fully connected layer; Denotes the sigmoid activation function; the symbol Denotes matrix multiplication;
[0030] Input the multi - layer channel splicing map into the large - kernel attention mechanism to obtain a multi - layer attention splicing map; the formula of the attention splicing map is as follows:
[0031] ;
[0032] ;
[0033] Among them, Denotes the th layer of large - kernel convolutional attention feature map, Denotes depth - wise convolution; Denotes depth - wise dilated convolution; Denotes Convolution operation of;
[0034] The convolutional module processes the multi - layer attention splicing map to obtain a fused feature map; the convolutional module is expressed by the formula as follows:
[0035] ;
[0036] Among them, Denotes the th layer of fused feature map, Denotes a convolution operation with a convolution kernel of 3 and a stride of 1; Denotes a reflection padding operation with a padding number of 1; Denotes a batch normalization operation; Denotes a rectified linear unit.
[0037] Preferably, the decoder includes 3 decoder wavelet convolutional modules cascaded in sequence and 2 block - based Kolmogorov - Arnold feature extraction modules;
[0038] The calculation formula of the fused image is:
[0039] ;
[0040] );
[0041] );
[0042] ;
[0043] ;
[0044] Among them, represents the fused image, represents the fused feature map after upsampling in the th layer; and respectively represent the wavelet convolution module and the block-based Kolmogorov - Arnold feature extraction module for the decoder in the th layer; represents the upsampling operation; the symbol
[0045] Preferably, the formula of the first-stage loss function is:
[0046] ;
[0047] where represents the first-stage loss function, 、 and respectively represent the pixel loss, the structural similarity loss, and the first-stage gradient loss function;
[0048] The formula of the pixel loss function is:
[0049] ;
[0050] where and respectively represent the L1 norm and the L2 norm; and respectively represent the output image and the input image; and respectively represent the height and width of the image;
[0051] The formula of the structural similarity loss function is:
[0052] ;
[0053] where represents the structural similarity index;
[0054] The formula of the first-stage gradient loss function is:
[0055] ;
[0056] ;
[0057] ;
[0058] where represents the pixel value of the image at the position ; represents the horizontal gradient, reflecting the pixel changes in the direction of the image; represents the vertical gradient, reflecting the pixel changes in the direction of the image.
[0059] Preferably, the formula of the second-stage loss function is:
[0060] ;
[0061] ;
[0062] ;
[0063] ;
[0064] wherein, represents the second-stage loss function, , and respectively represent the intensity loss, the detail texture loss, and the second-stage gradient loss function; and respectively represent the height and width of the image; , and respectively represent the fused image, the infrared image, and the visible light image; The symbol represents the L1 norm; The symbol represents the element-wise maximum selection; The symbol represents the Sobel gradient operator.
[0065] On the other hand, an embodiment of the present invention provides an infrared image and visible light image fusion system based on deep learning, and the system includes:
[0066] At least one processor;
[0067] At least one memory for storing at least one program;
[0068] When the at least one program is executed by the at least one processor, the at least one processor implements the method described in any one of the above.
[0069] On the other hand, an embodiment of the present invention provides a computer-readable storage medium, in which a processor-executable program is stored, and characterized in that the processor-executable program is used to execute the method described in any one of the above when executed by the processor.
[0070] The beneficial effects of the present invention are as follows: First, the present invention constructs an image fusion network, which includes an encoder, a fusion module, and a decoder, and adopts a two-stage training strategy, respectively designing loss functions suitable for each stage; in the first-stage training, the fusion module is removed, and only the encoder and the decoder are trained to enhance the feature extraction capabilities of the infrared encoder and the visible-light encoder; subsequently, the parameters of the trained encoder are fixed, and in the second-stage training, the fusion module and the decoder are optimized to improve the fusion quality; after the training is completed, an infrared image and a visible-light image are input to generate a high-quality fused image. The method of the present invention is compared with the existing fusion methods, and shows good fusion performance in terms of visual effects and objective evaluation indicators, and has high application value. Description of the Drawings
[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0072] Figure 1 It is a schematic flowchart of the method for fusing infrared images and visible-light images based on deep learning in the embodiments of the present invention;
[0073] Figure 2 It is a network structure diagram of the method for fusing infrared images and visible-light images based on deep learning in the embodiments of the present invention;
[0074] Figure 3 It is a network structure diagram of the wavelet convolution module EWTB for the encoder in the embodiments of the present invention;
[0075] Figure 4 It is a network structure diagram of the wavelet convolution module DWTB for the decoder in the embodiments of the present invention;
[0076] Figure 5 It is a network structure diagram of the block-based Kolmogorov-Arnold feature extraction module (Tok-KAN) in the embodiments of the present invention;
[0077] Figure 6 It is a network structure diagram of the fusion module in the embodiments of the present invention;
[0078] Figure 7 It is a structural diagram of the implementation process of the first-stage training strategy in the embodiments of the present invention;
[0079] Figure 8 It is a structural diagram of the implementation process of the second-stage training strategy in the embodiments of the present invention;
[0080] Figure 9 It is a schematic structural diagram of an infrared image and visible light image fusion system based on deep learning in an embodiment of the present invention. Specific embodiments
[0081] The following will clearly and completely describe the concept, specific structure, and technical effects generated by the present invention in combination with embodiments and drawings to fully understand the purpose, solution, and effects of the present invention. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0082] Refer to Figure 1 and Figure 2 The present invention provides a method for fusing infrared images and visible light images based on deep learning. The method includes the following steps:
[0083] S100. Construct an image fusion network, where the image fusion network includes an encoder, a fusion module, and a decoder; the encoder includes an infrared encoder and a visible light encoder;
[0084] S200. Form a U-shaped network with the encoder and the decoder, and train the U-shaped network based on a first-stage loss function to obtain a trained encoder.
[0085] In this example, the first stage is mainly used to train the infrared encoder and the visible light encoder. In this stage, the fusion module in the image fusion network is removed, and the encoder and the decoder together form a U-shaped network. To enable the network to effectively extract image features, this stage can be regarded as an image reconstruction task, that is, through optimization, the output image is made as close as possible to the input image, thereby improving the feature extraction ability of the encoder.
[0086] S300. Fix the parameters of the encoder, and train the image fusion network based on a second-stage loss function to obtain a trained image fusion network;
[0087] In this example, the task of the second stage is to train the fusion module and the decoder. In this stage, the parameters of the infrared encoder and the visible light encoder are kept frozen, and only the fusion module and the decoder are optimized to ensure that the fusion network can make full use of multi-scale features to achieve high-quality fusion image reconstruction.
[0088] S400. Use the trained image fusion network to fuse the input infrared image and visible light image to generate a fusion image.
[0089] The present invention proposes an infrared image and visible light image fusion method and system based on deep learning. The method first constructs an image fusion network, which includes an encoder, a fusion module, and a decoder, and adopts a two-stage training strategy, respectively designing loss functions suitable for each stage; in the first-stage training, the fusion module is removed, and only the encoder and decoder are trained to enhance the feature extraction capabilities of the infrared encoder and the visible light encoder; subsequently, the parameters of the trained encoder are fixed, and in the second-stage training, the fusion module and the decoder are optimized to improve the fusion quality; after the training is completed, the infrared and visible light images are input to generate a high-quality fused image. The method of the present invention is compared with existing fusion methods, and shows good fusion performance in terms of visual effects and objective evaluation indicators, and has high application value.
[0090] In some embodiments, in S400, fusing the input infrared image and visible light image by using the trained image fusion network to generate a fused image includes:
[0091] S410, obtaining the input infrared image and visible light image;
[0092] S420, inputting the infrared image and visible light image into the encoder to output multi-layer feature maps; the feature maps include infrared feature maps and visible light feature maps;
[0093] Specifically, inputting the infrared image as a training sample into the infrared encoder to output multi-layer infrared feature maps; inputting the visible light image as a training sample into the visible light encoder to output multi-layer visible light feature maps;
[0094] S430, respectively splicing the infrared feature maps and visible light feature maps in each layer of feature maps in the channel direction to obtain multi-layer spliced maps;
[0095] S440, inputting the multi-layer spliced maps into the fusion module for processing to obtain fused feature maps;
[0096] S450, inputting the fused feature maps into the decoder for processing to obtain a fused image.
[0097] In some embodiments, both the infrared encoder and the visible light encoder are composed of 3 wavelet convolutional modules for the encoder (EWTB, Encoder Wavelet Transform Block) and 2 block-based Kolmogorov-Arnold feature extraction modules (Tok-KAN).
[0098] The wavelet convolution module for the encoder consists of a normal convolution layer, a reflection padding layer, a wavelet convolution layer, a batch normalization layer, and a rectified linear unit (ReLU) in sequence; among them, the convolution kernel size of the normal convolution layer is 3, the stride is 1, the padding number of the reflection padding layer is 1, and the normal convolution layer and the reflection padding layer are used to adjust the number of channels of the feature map; the convolution kernel size of the wavelet convolution layer is 5, and the stride is 2; by performing downsampling operation, the length and width of the feature map are reduced to half of the original.
[0099] Given a pair of infrared image and visible light image, for the infrared encoder, the calculation formula of the infrared feature map is:
[0100] ;
[0101] ;
[0102] For the visible light encoder, the calculation formula of the visible light feature map is:
[0103] ;
[0104] ;
[0105] Among them, and respectively represent the feature maps of the infrared encoder and the visible light encoder at the th layer; and respectively represent the input infrared image and visible light image, , , , respectively represent the number of channels, height and width of the image; and respectively represent the wavelet convolution modules for the encoder in the th layer of the infrared encoder and the visible light encoder; and respectively represent the block-based Kolmogorov-Arnold feature extraction modules in the th layer of the infrared encoder and the visible light encoder.
[0106] Specifically, the encoder is respectively used to extract multi-scale features of infrared and visible light images, and is divided into five layers in total. The number of channels of the feature maps of each layer is 1, 32, 64, 128, 256, and 512 in sequence, and the spatial resolution decreases layer by layer, and is respectively , , , and 。
[0107] The encoder consists of 3 wavelet convolutional modules for the encoder (EWTB) and 2 block-based Kolmogorov-Arnold feature extraction modules (Tok-KAN). As Figure 3 shown, the wavelet convolutional module for the encoder (EWTB) consists of a common convolutional layer, a reflection padding layer, a wavelet convolutional layer, a batch normalization layer, and a rectified linear unit. Among them, the convolutional kernel size of the common convolutional layer is 3, the stride is 1, and the padding number of the reflection padding layer is 1. These two layers are used to adjust the number of channels of the feature map; the convolutional kernel size of the wavelet convolutional layer is 5, the stride is 2, and it performs downsampling operations to reduce the length and width of the feature map to half of the original. As Figure 4 shown, the block-based Kolmogorov-Arnold feature extraction module (Tok-KAN) consists of a block embedding layer, layer normalization, a KAN layer, and layer normalization in sequence. Among them, the block embedding layer is implemented by a common convolution, and its convolutional kernel size is 3 and the stride is 1.
[0108] In some embodiments, processing the multi-layer splicing map by inputting it into the fusion module to obtain a fused feature map includes:
[0109] Constructing the fusion module: The fusion module is located at the skip connection of the network, corresponding to Figure 2 the dotted lines in, a total of 5 places, aiming to adjust and balance the contributions of the multi-scale feature maps from the infrared and visible light encoders to improve the feature fusion effect.
[0110] As Figure 6 shown, the fusion module consists of a channel attention mechanism, a large kernel attention mechanism, and a convolutional module. The convolutional module consists of a common convolutional layer, a reflection padding layer, a batch normalization layer, and a rectified linear unit (ReLU) in sequence. Among them, the convolutional kernel size of the common convolution is 3, the stride is 1, the padding number of the reflection padding layer is 1, and the output channels are half of the input channels, which is used to halve the number of channels of the feature map.
[0111] Given the infrared feature map and the visible light feature map of the infrared encoder and the visible light encoder at the layer , where , , , respectively represent the number of channels, height, width, and layer number of the feature map. We first splice the feature maps in the channel direction to obtain a spliced map ; subsequently, we use the spliced map Input into the channel attention mechanism to adjust the contribution of each feature source, obtaining the channel attention feature map; the formula of the channel attention mechanism is as follows:
[0112] ;
[0113] Among them, represents the channel attention feature map of the th layer, represents global average pooling, represents a fully connected layer; represents the sigmoid activation function; the symbol represents matrix multiplication.
[0114] Subsequently, we input the channel concatenation map of the th layer into the large kernel attention mechanism, obtaining the large kernel attention feature map of the th layer; the large kernel attention mechanism is used to extract more extensive context information and optimize the feature representation, and the formula of the large kernel attention mechanism is as follows:
[0115] ;
[0116] ;
[0117] Among them, represents the large kernel convolution attention feature map of the th layer, represents depthwise convolution; represents depthwise dilated convolution; represents convolution operation;
[0118] Finally, the convolution module processes the large kernel convolution attention feature map to halve the number of channels of the attention concatenation map, reducing the computational complexity, and obtaining the fused feature map; this process is represented by the formula as follows:
[0119] ;
[0120] Among them, represents the fused feature map of the th layer, represents the convolution operation, and its convolution kernel is 3 and the stride is 1; represents the reflection padding operation, and its padding number is 1; represents the batch normalization operation; represents the rectified linear unit.
[0121] In some embodiments, the decoder and the encoder are symmetric in structure, and the two together form a U-shaped architecture. AsFigure 4 As shown in Figure 4 , the decoder consists of 3 wavelet convolution modules for the decoder (DWTB, Decoder Wavelet Transform Block) and 2 block-based Kolmogorov-Arnold feature extraction modules (Tok-KAN). There are some differences between the wavelet convolution module for the decoder (DWTB) and the wavelet convolution module for the encoder (EWTB). The wavelet convolution module for the decoder (DWTB) consists of a wavelet convolution layer, an ordinary convolution layer, a reflection padding layer, a batch normalization layer, and a rectified linear unit (ReLU) in sequence. Among them, the convolution kernel size of the wavelet convolution layer is 5, the stride is 1, and no downsampling operation is performed; the convolution kernel size of the ordinary convolution layer is 3, the stride is 1, and the padding number of the reflection padding layer is 1. These two layers are used to adjust the number of channels of the feature map.
[0122] The decoding process can be expressed by the following formula:
[0123] ;
[0124] );
[0125] );
[0126] ;
[0127] ;
[0128] Among them, represents the fused image, represents the th upsampled fused feature map after the and respectively represent the wavelet convolution module for the decoder and the block-based Kolmogorov-Arnold feature extraction module at the th layer; represents the upsampling operation; the symbol represents the concatenation operation.
[0129] In some embodiments, the loss function in the first stage is designed as follows: The goal of the first stage of training is to enhance the feature extraction capabilities of the infrared encoder and the visible light encoder respectively, and ensure that they can accurately extract infrared target information and visible light texture details. For this purpose, the following loss functions are mainly used in the first stage of training:
[0130] Pixel Loss: Pixel Loss is used to measure the difference between the reconstructed image and the input image at the pixel level to ensure that the output image is numerically as close as possible to the input image.
[0131] The formula for the loss function in the first stage is as follows:
[0132] ;
[0133] Among them, represents the loss function in the first stage, 、 and respectively represent pixel loss, structural similarity loss, and the gradient loss function in the first stage;
[0134] The formula for the pixel loss function is:
[0135] ;
[0136] Among them, and respectively represent the L1 norm and the L2 norm; and respectively represent the output image and the input image; and respectively represent the height and width of the image;
[0137] Structural Similarity Loss (SSIM Loss): Although pixel loss can effectively measure the numerical difference between the reconstructed image and the input image, it cannot accurately reflect the perceptual quality of the human eye. To address this limitation, we further introduce structural similarity loss, which uses brightness, contrast, and structural information as evaluation metrics to measure the perceptual similarity between the output image and the input image. The formula for the structural similarity loss function is:
[0138] ;
[0139] Among them, represents the structural similarity index;
[0140] Gradient Loss (Grad Loss): To further improve the quality of the reconstructed image, we introduce gradient loss. This loss function can more effectively capture the structural information in the image, especially in terms of preserving edge and texture details. The formula for the gradient loss function in the first stage is:
[0141] ;
[0142] ;
[0143] ;
[0144] Among them, represents the pixel value of the image at position ; Represents the horizontal gradient, reflecting the pixel change of the image in the x direction; Represents the vertical gradient, reflecting the pixel change of the image in the y direction.
[0145] It should be noted that in the first stage, the fusion module is removed. As Figure 7 shown, we respectively constructed an infrared autoencoder and a visible light autoencoder, and each autoencoder consists of an encoder and a decoder to independently learn and extract infrared target information and visible light texture details.
[0146] As Figure 8 shown, the training process in the second stage aims to make full use of the multi-scale features extracted by the encoder and integrate information through the fusion module to generate a high-quality fused image.
[0147] Design of the loss function in the second stage: In this example, the task in the second stage is to train the fusion module and the decoder. At this stage, the parameters of the infrared encoder and the visible light encoder are kept frozen, and only the fusion module and the decoder are optimized to ensure that the fusion network can make full use of the multi-scale features to achieve high-quality fused image reconstruction.
[0148] In some embodiments, the formula of the second-stage loss function is:
[0149] ;
[0150] ;
[0151] ;
[0152] ;
[0153] Among them, represents the second-stage loss function, , and respectively represent the intensity loss, the detail texture loss, and the second-stage gradient loss function; and respectively represent the height and width of the image; , and respectively represent the fused image, the infrared image, and the visible light image; The symbol represents the L1 norm; The symbol represents the element-wise maximum selection; The symbol represents the Sobel gradient operator.
[0154] To further verify the performance of this method, comparative experiments were conducted on three datasets: TNO, M3FD, and LLVIP. Tables 1, 2, and 3 respectively show the experimental results of the fusion method proposed in the present invention on the three datasets, and a comparative analysis was carried out with existing methods to objectively evaluate its performance and effectiveness.
[0155] Table 1: Comparison of quantitative evaluation indicators between the present invention and existing fusion methods on the TNO dataset.
[0156]
[0157] Table 2: Comparison of quantitative evaluation indicators between the present invention and existing fusion methods on the M3FD dataset.
[0158]
[0159] Table 3: Comparison of quantitative evaluation indicators between the present invention and existing fusion methods on the LLVIP dataset.
[0160]
[0161] The relevant literature in Tables 1, 2, and 3 is as follows:
[0162] [1] XYDEAS C S, PETROVIĆ V. Objective image fusion performancemeasure[J / OL]. Electronics Letters, 2000, 36(4): 308-309[2025-04-13]. http: / / digital-library.theiet.org / doi / 10.1049 / el%3A20000267. DOI:10.1049 / el:20000267.
[0163] [2] RAO Y J. In-fibre bragg grating sensors[J / OL]. DOI:10.1088 / 0957-0233 / 8 / 4 / 002.
[0164] [3] ROBERTS J W, VAN AARDT J, AHMED F. Assessment of image fusionprocedures using entropy, image quality, and multispectral classification[J].
[0165] [4] CHEN H, VARSHNEY P K. A human perception inspired quality metric for image fusion based on regional information[J / OL]. Information Fusion,2007, 8(2): 193-207[2025-04-13]. https: / / linkinghub.elsevier.com / retrieve / pii / S1566253505000898. DOI:10.1016 / j.inffus.2005.10.001.
[0166] [5] CHEN Y, BLUM R S. A new automated quality assessment algorithm for image fusion[J / OL]. DOI:10.1016 / j.imavis.2007.12.002.
[0167] [6] ESKICIOGLU A M, FISHER P S. Image quality measures and their performance[J / OL]. IEEE Transactions on Communications, 1995, 43(12): 2959-2965[2025-04-13]. http: / / ieeexplore.ieee.org / document / 477498 / . DOI:10.1109 / 26.477498.
[0168] [7] SHEIKH H R, BOVIK A C. Image information and visual quality[J].
[0169] [8] CUI G. Detail preserved fusion of visible and infrared images using regional saliency extraction and multi-scale image decomposition[J / OL]. DOI:10.1016 / j.optcom.2014.12.032.
[0170] [9] NAIDU V P S. Image fusion technique using multi - resolution singular value decomposition[J / OL]. Defence Science Journal, 2011, 61(5): 479[2025 - 04 - 13]. http: / / publications.drdo.gov.in / ojs / index.php / dsj / article / view / 705. DOI:10.14429 / dsj.61.705.
[0171]
[10] GAO Z. Texture clear multi - modal image fusion with joint sparsity model[J / OL]. DOI:10.1016 / j.ijleo.2016.09.126.
[0172]
[11] LIU Y, CHEN X, WARD R K, et al. Image fusion with convolutional sparse representation[J / OL]. IEEE Signal Processing Letters, 2016, 23(12):1882 - 1886[2025 - 04 - 13]. https: / / ieeexplore.ieee.org / document / 7593316 / . DOI:10.1109 / LSP.2016.2618776.
[0173]
[12] LI H, WU X J, KITTLER J. MDLatLRR: A novel decomposition method for infrared and visible image fusion[J / OL]. IEEE Transactions on Image Processing, 2020, 29: 4733 - 4746[2025 - 04 - 13]. https: / / ieeexplore.ieee.org / document / 9018389 / . DOI:10.1109 / TIP.2020.2975984.
[0174]
[13] LI H, WU X J, KITTLER J. Infrared and visible image fusion usinga deep learning framework[C / OL] / / 2018 24th International Conference onPattern Recognition (ICPR). Beijing: IEEE, 2018: 2705-2710[2025-04-13].https: / / ieeexplore.ieee.org / document / 8546006 / . DOI:10.1109 / ICPR.2018.8546006.
[0175]
[14] LI H, WU X jun, DURRANI T S. Infrared and visible image fusionwith ResNet and zero-phase component analysis[J / OL]. Infrared Physics&Technology, 2019, 102: 103039[2025-04-13]. https: / / linkinghub.elsevier.com / retrieve / pii / S1350449519301525. DOI:10.1016 / j.infrared.2019.103039.
[0176]
[15] LI H, WU X J. DenseFuse: A fusion approach to infrared andvisible images[J / OL]. IEEE Transactions on Image Processing, 2019, 28(5):2614-2623[2024-09-21]. http: / / arxiv.org / abs / 1804.08361. DOI:10.1109 / TIP.2018.2887342.
[0177]
[16] MA J, YU W, LIANG P, et al. FusionGAN: A generative adversarial network for infrared and visible image fusion[J / OL]. Information Fusion, 2019, 48: 11-26[2024-09-29]. https: / / linkinghub.elsevier.com / retrieve / pii / S1566253518301143. DOI:10.1016 / j.inffus.2018.09.004.
[0178]
[17] MA J, ZHANG H, SHAO Z, et al. GANMcC: A Generative Adversarial Network With Multiclassification Constraints for Infrared and Visible Image Fusion[J / OL]. IEEE Transactions on Instrumentation and Measurement, 2021, 70: 1-14[2024-09-29]. https: / / ieeexplore.ieee.org / document / 9274337 / . DOI:10.1109 / TIM.2020.3038013.
[0179]
[18] LI H, WU X J, KITTLER J. RFN-Nest: An end-to-end residual fusion network for infrared and visible images[J / OL]. Information Fusion, 2021, 73: 72-86[2024-09-21]. https: / / linkinghub.elsevier.com / retrieve / pii / S1566253521000440. DOI:10.1016 / j.inffus.2021.02.023.
[0180]
[19] XU H, WANG X, MA J. DRF: Disentangled representation for visibleand infrared image fusion[J]. IEEE TRANSACTIONS ON INSTRUMENTATION ANDMEASUREMENT, 2021,70.
[0181]
[20] JUNG H, KIM Y, JANG H, et al. Unsupervised deep image fusionwith structure tensor representations[J / OL]. IEEE Transactions on ImageProcessing, 2020, 29: 3845-3858[2025-04-13]. https: / / ieeexplore.ieee.org / document / 8962327 / . DOI:10.1109 / TIP.2020.2966075.
[0182]
[21] LIU J, FAN X, HUANG Z, et al. Target-aware Dual AdversarialLearning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared andVisible for Object Detection[A / OL]. arXiv, 2022[2024-09-29]. http: / / arxiv.org / abs / 2203.16220. DOI:10.48550 / arXiv.2203.16220.
[0183]
[22] LI H, WU X J. CrossFuse: A novel cross attention mechanism based infrared and visible image fusion approach[J / OL]. Information Fusion, 2024, 103: 102147[2024-11-15]. https: / / linkinghub.elsevier.com / retrieve / pii / S1566253523004633. DOI:10.1016 / j.inffus.2023.102147.
[0184] Compared with the prior art, the present invention has the following beneficial effects:
[0185] 1. The designed wavelet convolution modules (EWTB and DWTB) of the present invention can more effectively separate feature information of different frequencies. Compared with the existing methods that only use ordinary convolution for feature extraction, this method can more accurately capture the key features of infrared images and visible light images, avoid the interference caused by mixed frequency information, thereby enhancing the feature representation ability and further improving the quality of the fused image.
[0186] 2. The introduction of the Kolmogorov - Arnold representation theorem in the present invention can help the network learn more complex mapping relationships, breaking through the problem of weak feature expression ability due to the use of fixed activation functions in traditional deep learning networks.
[0187] 3. The fusion module proposed by the present invention can adaptively adjust the contribution ratio of each feature source (infrared and visible light features), thereby achieving a more balanced and efficient fusion process, improving the quality of the fused image and the information expression ability.
[0188] 4. The designed loss function of the present invention optimizes the deficiencies of the existing infrared and visible light image fusion methods in terms of the loss function, effectively improving the quality of the fused image and better retaining its texture details.
[0189] 5. Through comparison on the TNO, M3FD and LLVIP datasets (a total of 186 pairs of infrared and visible light images), the algorithm proposed by the present invention is superior to 14 existing fusion algorithms and achieves state - of - the - art performance in terms of visual quality and objective evaluation metrics.
[0190] Corresponding to Figure 1 the method of Figure 9 , an embodiment of the present invention provides a deep - learning - based infrared image and visible light image fusion system, including:
[0191] At least one processor;
[0192] At least one memory for storing at least one program;
[0193] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.
[0194] It can be seen that the content in the above method embodiments is applicable to the system embodiments herein. The functions specifically implemented by the system embodiments herein are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0195] In addition, an embodiment of the present invention also discloses a computer program product or a computer program. The computer program product or the computer program is stored in a computer-readable storage medium. The processor of the computer device can read the computer program from the computer-readable storage medium, and the processor executes the computer program so that the computer device executes the above method. Similarly, the content in the above method embodiments is applicable to the storage medium embodiments herein. The functions specifically implemented by the storage medium embodiments herein are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0196] Those of ordinary skill in the art can understand that all or some of the methods and systems disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or a non-transitory medium) and a communication medium (or a transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disc (DVD), or other optical disc storage, magnetic cassette, tape, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium generally includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0197] The foregoing has described the preferred embodiments of the present disclosure in detail, but the present disclosure is not limited to the above-described embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present disclosure, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present disclosure.
Claims
1. An infrared image and visible light image fusion method based on deep learning, characterized in that, The method includes the following steps: Construct an image fusion network, which includes an encoder, a fusion module, and a decoder; the encoder includes an infrared encoder and a visible light encoder; Form a U-shaped network with the encoder and the decoder, and train the U-shaped network based on the first loss function to obtain a trained encoder; Fix the parameters of the encoder, and train the image fusion network based on the second loss function to obtain a trained image fusion network; Use the trained image fusion network to fuse the input infrared image and visible light image to generate a fused image; The step of using the trained image fusion network to fuse the input infrared image and visible light image to generate a fused image includes: Obtain the input infrared image and visible light image; Input the infrared image and visible light image into the encoder to output multi-layer feature maps; the feature maps include infrared feature maps and visible light feature maps; In the channel direction, splice the infrared feature maps and visible light feature maps in each layer of the feature maps respectively to obtain multi-layer spliced maps; Input the multi-layer spliced maps into the fusion module for processing to obtain fused feature maps; Input the fused feature maps into the decoder for processing to obtain a fused image; Both the infrared encoder and the visible light encoder include 3 wavelet convolution modules for the encoder and 2 block-based Kolmogorov-Arnold feature extraction modules cascaded in sequence; the wavelet convolution module for the encoder is composed of a common convolution layer, a reflection padding layer, a wavelet convolution layer, a batch normalization layer, and a rectified linear unit in sequence; The fusion module includes a channel attention mechanism, a large kernel attention mechanism, and a convolution module, and the convolution module includes a common convolution layer, a reflection padding layer, a batch normalization layer, and a rectified linear unit; The step of inputting the multi-layer spliced maps into the fusion module for processing to obtain fused feature maps includes: Input the multi-layer spliced maps into the corresponding channel attention mechanism respectively to obtain channel attention feature maps; the channel attention mechanism is represented by the following formula: ; Among them, represents the channel attention feature map of the th layer, represents global average pooling, represents a fully connected layer; represents the sigmoid activation function; the symbol represents matrix multiplication; is the concatenated graph of the th layer; Input the channel attention feature maps into the large kernel attention mechanism to obtain large kernel convolution attention feature maps; the large kernel attention mechanism is represented by the following formula: ; ; Among them, represents the layer large kernel convolutional attention feature map, represents depthwise convolution; represents depthwise dilated convolution; represents convolution operation of Process the multi-layer attention spliced maps by the convolution module to obtain fused feature maps; the convolution module is represented by the following formula: ; Among them, represents the layer fusion feature map, represents a convolution operation with a convolution kernel of 3 and a stride of 1; represents a reflection padding operation with a padding number of 1; represents a batch normalization operation; represents a rectified linear unit.
2. The method according to claim 1, wherein The calculation formula for the infrared feature map is: ; ; The calculation formula for the visible light feature map is: ; ; Among them, and respectively represent the feature maps of the infrared encoder and the visible light encoder in the th layer; and respectively represent the input infrared image and visible light image, , , , respectively represent the number of channels, height, and width of the image; and respectively represent the wavelet convolution modules for the encoder in the th layer of the infrared encoder and the visible light encoder; and respectively represent the block-based Kolmogorov - Arnold feature extraction modules in the th layer of the infrared encoder and the visible light encoder.
3. The method according to claim 2, wherein The decoder includes 3 wavelet convolution modules for the decoder and 2 block-based Kolmogorov-Arnold feature extraction modules cascaded in sequence; The calculation formula for the fused image is: ; ; ; ; ; Among them, represents the fused image, represents the fused feature map after upsampling in the th layer; and respectively represent the wavelet convolution for the decoder and the block-based Kolmogorov-Arnold feature extraction module in the th layer; represents the upsampling operation; the symbol 4. The method according to claim 1, characterized in that, The formula for the first loss function is: ; Among them, represents the first loss function, , and respectively represent pixel loss, structural similarity loss, and the first gradient loss function; The formula for the pixel loss function is: ; Among them, and represent the L1 norm and the L2 norm respectively; and represent the output image and the input image respectively; and represent the height and width of the image respectively; The formula for the structural similarity loss function is: ; Among them, represents the structural similarity index; The formula for the first gradient loss function is: ; ; ; Among them, represents the pixel value of the image at the position ; represents the horizontal gradient, reflecting the pixel change of the image in the direction; represents the vertical gradient, reflecting the pixel change of the image in the direction.
5. The method according to claim 1, characterized in that The formula for the second loss function is: ; ; ; ; Among them, represents the second loss function, , and represent the intensity loss, the detail texture loss, and the second gradient loss function respectively; and represent the height and width of the image respectively; , and represent the fused image, the infrared image, and the visible light image respectively; The symbol represents the L1 norm; The symbol represents the element-wise maximum selection; The symbol represents the Sobel gradient operator.
6. An infrared image and visible light image fusion system based on deep learning, characterized in that, The system includes: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 5.
7. A computer-readable storage medium storing a program executable by a processor, characterized in that, The program executable by the processor, when executed by the processor, is used to perform the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Infrared and visible light image fusion method based on visual enhancement
CN116363036A
Image reconstruction algorithm, image processing method and device, vehicle and storage medium
CN119579713A