Infrared image and visible light image fusion method and system based on deep learning
By building an image fusion network and adopting a two-stage training strategy, the limitations of the infrared image and visible image fusion methods in the prior art in feature extraction and activation function selection are solved, and high-quality image fusion is achieved, especially in complex scenarios, the retention ability of texture details is significantly improved.
Patent Information
- Application Number
- CN202510585967.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-08
AI Technical Summary
The existing deep learning-based infrared image and visible image fusion methods have limitations in feature extraction and activation function selection, resulting in limited fusion quality, especially in complex scenarios, with weak ability to retain image texture details.
Build an image fusion network, including an encoder, a fusion module and a decoder, and adopts a two-stage training strategy. The first stage removes the fusion module and trains only the encoder and decoder to enhance feature extraction capabilities. The second stage is fixed encoder parameters, and the fusion module and decoder are optimized to improve the fusion quality.
Through the two-stage training strategy and the loss function of the design, the quality of the fusion image is significantly improved, and the ability to retain image texture details in complex scenes is enhanced, which is better than 14 existing fusion algorithms.
Smart Images

Figure CN120107089A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a method and system for fusing infrared images and visible light images based on deep learning. Background Art
[0002] Image fusion is an image processing technology that combines multimodal images of the same scene to generate an image that can accurately and comprehensively describe the scene. Infrared images can work in low-light, dark or harsh environments. By capturing the thermal radiation information of objects, they provide temperature distribution, heat source detection and hidden target identification. However, they cannot accurately capture the details and color information of objects, and are easily affected by changes in ambient temperature and humidity. Visible light images can provide rich details and color information and are suitable for most daily scenes, but they perform poorly in low-light or dark environments and are limited by lighting conditions. Therefore, fusing infrared images and visible light images can make full use of the advantages of both, enabling them to provide stronger perception capabilities in complex low-light, nighttime or changing environments. This technology is widely used in military reconnaissance, target tracking, intelligent driving, security monitoring and other fields.
[0003] At present, infrared image and visible light image fusion technology can be roughly divided into traditional fusion methods and deep learning-based fusion methods. Traditional fusion methods include methods based on multi-scale transformation, sparse representation, subspace, saliency, total variation, etc. However, these traditional fusion methods have limitations. Their fusion strategies rely too much on manual design and cannot adapt to increasingly complex fusion scenarios, resulting in limited fusion performance. Deep learning-based methods include fusion methods based on autoencoders, generative adversarial networks, and convolutional neural networks. Due to the strong ability of convolution operations and attention mechanisms to extract features and driven by big data, deep learning-based fusion methods have achieved satisfactory fusion results. However, there are still some shortcomings in the existing deep learning-based fusion algorithm methods. First, in terms of feature extraction, the existing methods all use ordinary convolution for feature extraction. The extracted features contain information about mixed frequencies, making it difficult for the model to distinguish features of different frequencies, resulting in limited fusion quality. Second, in traditional deep learning networks, fixed activation functions are usually used to introduce nonlinearity to help the network learn complex mapping relationships. However, the choice of this fixed activation function limits the network's expression ability, making it have certain limitations when processing complex features. Third, the existing infrared and visible light image fusion methods have certain defects in the loss function design, which leads to weak ability to preserve image texture details in complex scenes and easily causes the loss of important information. Summary of the invention
[0004] The purpose of the present invention is to provide a method and system for fusing infrared images and visible light images based on deep learning, which can improve the quality of the fused image.
[0005] In order to achieve the above object, the present invention provides the following technical solutions: On the one hand, an embodiment of the present invention provides a method for fusing infrared images and visible light images based on deep learning, the method comprising the following steps: Constructing an image fusion network, the image fusion network comprising an encoder, a fusion module and a decoder; the encoder comprising an infrared encoder and a visible light encoder; The encoder and the decoder form a U-shaped network, and the U-shaped network is trained based on the first-stage loss function to obtain a trained encoder; Fixing the parameters of the encoder, and training the image fusion network based on the second-stage loss function to obtain a trained image fusion network; The trained image fusion network is used to fuse the input infrared image and visible light image to generate a fused image.
[0006] Preferably, the step of using a trained image fusion network to fuse the input infrared image and the visible light image to generate a fused image includes: Obtain input infrared images and visible light images; Inputting the infrared image and the visible light image into an encoder and outputting a multi-layer feature map; the feature map includes an infrared feature map and a visible light feature map; The infrared feature map and the visible light feature map in each layer of feature map are spliced in the channel direction to obtain a multi-layer spliced map; The multi-layer spliced image is input into the fusion module for processing to obtain a fused feature map; The fused feature map is input into the decoder for processing to obtain a fused image.
[0007] Preferably, the encoder includes an infrared encoder and a visible light encoder, and the infrared encoder and the visible light encoder each include three wavelet convolution modules for encoders and two Kolmogorov-Arnold feature extraction modules based on block cascading in sequence; the wavelet convolution module for the encoder is composed of a normal convolution layer, a reflection filling layer, a wavelet convolution layer, a batch normalization layer and a rectified linear unit in sequence; The calculation formula of the infrared characteristic map is: ; ; The calculation formula of the visible light characteristic graph is: ; ; in, and Indicates the infrared encoder and visible light encoder in the Feature map of the layer; and Represent the input infrared image and visible light image respectively, , , , Respectively represent the number of channels, height and width of the image; and Respectively represent the infrared encoder and visible light encoder The wavelet convolution module of the layer is used for the encoder; and Respectively represent the infrared encoder and visible light encoder A block-based Kolmogorov-Arnold feature extraction module for the layer.
[0008] Preferably, the fusion module includes a channel attention mechanism, a large kernel attention mechanism and a convolution module, and the convolution module includes a common convolution layer, a reflection padding layer, a batch normalization layer and a rectified linear unit; The multi-layer spliced image is input into the fusion module for processing to obtain a fusion feature map, including: The multi-layer splicing graph is input into the corresponding channel attention mechanism respectively to obtain a multi-layer channel splicing graph; the formula of the channel splicing graph is as follows: ; in, Indicates Layer channel attention feature map, represents global average pooling, represents a fully connected layer; Represents the sigmoid activation function; symbol Represents matrix multiplication; The multi-layer channel splicing graph is input into the large core attention mechanism to obtain a multi-layer attention splicing graph; the formula of the attention splicing graph is as follows: ; ; in, Indicates Layer large kernel convolution attention feature map, represents depthwise convolution; represents a depth-dilated convolution; express Convolution operation; The multi-layer attention splicing map is processed by the convolution module to obtain a fused feature map; the convolution module is expressed by the following formula: ; in, Indicates Layer fusion feature map, Represents a convolution operation, with a convolution kernel of 3 and a step size of 1; Indicates a reflection fill operation, and its fill number is 1; Represents batch normalization operation; represents a rectified linear unit.
[0009] Preferably, the decoder comprises three wavelet convolution modules for the decoder and two Kolmogorov-Arnold feature extraction modules based on block connection in sequence; The calculation formula of the fused image is: ; ); ); ; ; in, represents the fused image, Indicates The fused feature map of the layer after upsampling; and Respectively expressed in The layer is used for the wavelet convolution module and the block-based Kolmogorov-Arnold feature extraction module of the decoder; Represents an upsampling operation; the symbol Represents a concatenation operation.
[0010] Preferably, the formula of the first stage loss function is: ; in, represents the first stage loss function, 、 and Represent pixel loss, structural similarity loss and first-stage gradient loss function respectively; The formula of pixel loss function is: ; in, and Represent the L1 norm and L2 norm respectively; and Represent the output image and input image respectively; and Respectively represent the height and width of the image; The formula of the structural similarity loss function is: ; in, represents the structural similarity index; The formula of the gradient loss function in the first stage is: ; ; ;
[0011] in, Indicates that the image is at position The pixel value at ; Represents the horizontal gradient, reflecting the image Pixel changes in direction; Represents the vertical gradient, reflecting the image The pixel change in direction.
[0012] Preferably, the formula of the second stage loss function is: ; ; ; ; in, represents the second stage loss function, , and Represent intensity loss, detail texture loss, and second-stage gradient loss function respectively; and Represent the height and width of the image respectively; , and Represent the fused image, infrared image and visible light image respectively; The symbol represents the L1 norm; The symbol indicates element-wise maximum selection; The symbol represents the Sobel gradient operator.
[0013] On the other hand, an embodiment of the present invention provides an infrared image and visible light image fusion system based on deep learning, the system comprising: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements any one of the methods described above.
[0014] On the other hand, an embodiment of the present invention provides a computer-readable storage medium, which stores a program executable by a processor, wherein the program executable by the processor is used to execute any of the methods described above when executed by the processor.
[0015] The beneficial effects of the present invention are as follows: the present invention first constructs an image fusion network, which includes an encoder, a fusion module and a decoder, and adopts a two-stage training strategy to design loss functions suitable for each stage respectively; in the first stage of training, the fusion module is removed, and only the encoder and decoder are trained to enhance the feature extraction capabilities of the infrared encoder and the visible light encoder; then, the trained encoder parameters are fixed, and in the second stage of training, the fusion module and the decoder are optimized to improve the fusion quality; after the training is completed, the infrared image and the visible light image are input to generate a high-quality fused image. Compared with the existing fusion method, the method of the present invention shows good fusion performance in terms of visual effects and objective evaluation indicators, and has high application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0017] Figure 1 is a flowchart of a method for fusing infrared images and visible light images based on deep learning in an embodiment of the present invention; Figure 2 is a network structure diagram of the infrared image and visible light image fusion method based on deep learning in an embodiment of the present invention; Figure 3 is a network structure diagram of a wavelet convolution module EWTB used for an encoder in an embodiment of the present invention; Figure 4 is a network structure diagram of a wavelet convolution module DWTB used in a decoder in an embodiment of the present invention; Figure 5 is a network structure diagram of a block-based Kolmogorov-Arnold feature extraction module (Tok-KAN) in an embodiment of the present invention; Figure 6 is a network structure diagram of a fusion module in an embodiment of the present invention; Figure 7 is a structural diagram of the implementation process of the first stage training strategy in an embodiment of the present invention; Figure 8 is a structural diagram of the implementation process of the second stage training strategy in an embodiment of the present invention; Fig. 9 It is a structural schematic diagram of an infrared image and visible light image fusion system based on deep learning in an embodiment of the present invention. DETAILED DESCRIPTION
[0018] The following will be combined with the embodiments and drawings to clearly and completely describe the concept, specific structure and technical effects of the present invention, so as to fully understand the purpose, scheme and effect of the present invention. It should be noted that the embodiments and features in the embodiments of the present invention can be combined with each other without conflict.
[0019] See also Figure 1 and Figure 2 The present invention provides a method for fusing infrared images and visible light images based on deep learning, the method comprising the following steps: S100, constructing an image fusion network, wherein the image fusion network includes an encoder, a fusion module, and a decoder; the encoder includes an infrared encoder and a visible light encoder; S200, the encoder and the decoder form a U-shaped network, and the U-shaped network is trained based on the first-stage loss function to obtain a trained encoder.
[0020] In this example, the first stage is mainly used to train the infrared encoder and the visible light encoder. In this stage, the fusion module in the image fusion network is removed, and the encoder and decoder together form a U-shaped network. In order to enable the network to effectively extract image features, this stage can be regarded as an image reconstruction task, that is, through optimization, the output image is made as close as possible to the input image, thereby improving the feature extraction ability of the encoder.
[0021] S300, fixing the parameters of the encoder, and training the image fusion network based on the second stage loss function to obtain a trained image fusion network; In this example, the task of the second stage is to train the fusion module and decoder. In this stage, the parameters of the infrared encoder and the visible light encoder remain frozen, and only the fusion module and the decoder are optimized to ensure that the fusion network can fully utilize multi-scale features to achieve high-quality fusion image reconstruction.
[0022] S400 uses the trained image fusion network to fuse the input infrared image and visible light image to generate a fused image.
[0023] The present invention proposes a method and system for fusion of infrared images and visible light images based on deep learning. The method first constructs an image fusion network, which includes an encoder, a fusion module and a decoder, and adopts a two-stage training strategy to design loss functions suitable for each stage respectively; in the first stage of training, the fusion module is removed, and only the encoder and decoder are trained to enhance the feature extraction capabilities of the infrared encoder and the visible light encoder; then, the trained encoder parameters are fixed, and in the second stage of training, the fusion module and the decoder are optimized to improve the fusion quality; after the training is completed, the infrared and visible light images are input to generate high-quality fused images. Compared with the existing fusion methods, the method of the present invention shows good fusion performance in terms of visual effects and objective evaluation indicators, and has high application value.
[0024] In some embodiments, in S400, the step of fusing the input infrared image and the visible light image using a trained image fusion network to generate a fused image includes: S410, acquiring an input infrared image and a visible light image; S420, inputting the infrared image and the visible light image into an encoder, and outputting a multi-layer feature map; the feature map includes an infrared feature map and a visible light feature map; Specifically, an infrared image as a training sample is input into an infrared encoder, and a multi-layer infrared feature map is output; a visible light image as a training sample is input into a visible light encoder, and a multi-layer visible light feature map is output; S430, splicing the infrared feature map and the visible light feature map in each layer of feature map in the channel direction to obtain a multi-layer spliced map; S440, inputting the multi-layer spliced image into a fusion module for processing to obtain a fusion feature map; S450, input the fused feature map into the decoder for processing to obtain a fused image.
[0025] In some embodiments, both the infrared encoder and the visible light encoder are composed of three wavelet convolution modules (EWTB, Encoder Wavelet Transform Block) for encoders and two block-based Kolmogorov-Arnold feature extraction modules (Tok-KAN).
[0026] The wavelet convolution module used for the encoder is composed of a normal convolution layer, a reflection padding layer, a wavelet convolution layer, a batch normalization layer, and a rectified linear unit (ReLU) in sequence; the convolution kernel size of the normal convolution layer is 3, the step size is 1, the padding number of the reflection padding layer is 1, and the normal convolution layer and the reflection padding layer are used to adjust the number of channels of the feature map; the convolution kernel size of the wavelet convolution layer is 5, and the step size is 2; by performing a downsampling operation, the length and width of the feature map are reduced to half of the original.
[0027] Given a pair of infrared images and visible light images, for the infrared encoder, the calculation formula of the infrared feature map is: ; ; For a visible light encoder, the calculation formula of the visible light characteristic map is: ; ; in, and Indicates the infrared encoder and visible light encoder in the Feature map of the layer; and Represent the input infrared image and visible light image respectively, , , , Respectively represent the number of channels, height and width of the image; and Respectively represent the infrared encoder and visible light encoder Layer's wavelet convolution module for encoder; and Respectively represent the infrared encoder and visible light encoder A block-based Kolmogorov-Arnold feature extraction module for the layer.
[0028] Specifically, the encoder is used to extract multi-scale features of infrared and visible light images, respectively, and is divided into five layers. The number of feature map channels in each layer is 1, 32, 64, 128, 256, and 512, respectively, and the spatial resolution decreases layer by layer, which are the original image , , , and .
[0029] The encoder consists of three wavelet convolution modules for encoder (EWTB) and two block-based Kolmogorov-Arnold feature extraction modules (Tok-KAN). Figure 3 As shown in the figure, the wavelet convolution module (EWTB) for the encoder consists of a normal convolution layer, a reflection padding layer, a wavelet convolution layer, a batch normalization layer, and a rectified linear unit. The convolution kernel size of the normal convolution layer is 3, the step size is 1, and the padding number of the reflection padding layer is 1. These two layers are used to adjust the number of channels of the feature map; the convolution kernel size of the wavelet convolution layer is 5, the step size is 2, and the downsampling operation is performed to reduce the length and width of the feature map to half of the original. Figure 4 As shown in Figure 1, the block-based Kolmogorov-Arnold feature extraction module (Tok-KAN) consists of a block embedding layer, layer normalization, a KAN layer, and layer normalization. The block embedding layer is implemented by ordinary convolution, with a convolution kernel size of 3 and a stride of 1.
[0030] In some embodiments, the step of inputting the multi-layer spliced image into a fusion module for processing to obtain a fused feature image includes: Constructing fusion module: The fusion module is located at the jump connection of the network, corresponding to Figure 2 There are 5 dotted lines in the figure, which are intended to adjust and balance the contributions of multi-scale feature maps from the infrared and visible light encoders to improve the feature fusion effect.
[0031] like Figure 6 As shown in the figure, the fusion module consists of a channel attention mechanism, a large kernel attention mechanism, and a convolution module. The convolution module consists of a normal convolution layer, a reflection padding layer, a batch normalization layer, and a rectified linear unit (ReLU) in sequence. Among them, the convolution kernel size of the normal convolution is 3, the stride is 1, the padding number of the reflection padding layer is 1, and the output channel is half of the input channel, which is used to halve the number of channels of the feature map.
[0032] Given an infrared encoder and a visible light encoder in Infrared and visible light signatures of the layers ,in , , , Represent the number of channels, height, width and layer number of the feature map respectively. We first transform the feature map in the channel direction Perform stitching to obtain a stitching map ; Then, we will stitch the graph Input into the channel attention mechanism to adjust the contribution of each feature source and obtain the channel attention feature map; the formula of the channel attention mechanism is as follows: ; in, Indicates The channel attention feature map of the layer, represents global average pooling, represents a fully connected layer; Represents the sigmoid activation function; symbol Represents matrix multiplication.
[0033] Afterwards, we will Layer channel mosaic Input into the large core attention mechanism to get the Layer large core attention feature map; the large core attention mechanism is used to extract more extensive context information and optimize feature representation. The formula of the large core attention mechanism is as follows: ; ; in, Indicates Layer large kernel convolution attention feature map, represents depthwise convolution; represents a depth-dilated convolution; express Convolution operation; Finally, the convolution module convolves the attention feature map with a large kernel Processing is performed to halve the number of channels of the attention splicing map, reduce the computational complexity, and obtain a fused feature map; the process is expressed by the following formula: ; in, Indicates Layer fusion feature map, Represents a convolution operation, with a convolution kernel of 3 and a step size of 1; Indicates a reflection fill operation, and its fill number is 1; Represents batch normalization operation; represents a rectified linear unit.
[0034] In some embodiments, the decoder and the encoder are symmetrical in structure, and the two together form a U-shaped architecture. Figure 4As shown in the figure, the decoder consists of three wavelet convolution modules (DWTB, Decoder WaveletTransform Block) for the decoder and two block-based Kolmogorov-Arnold feature extraction modules (Tok-KAN). There are some differences between the wavelet convolution module (DWTB) for the decoder and the wavelet convolution module (EWTB) for the encoder. The wavelet convolution module (DWTB) for the decoder consists of a wavelet convolution layer, a normal convolution layer, a reflection padding layer, a batch normalization layer, and a rectified linear unit (ReLU) in sequence. Among them, the convolution kernel size of the wavelet convolution layer is 5, the step size is 1, and no downsampling operation is performed; the convolution kernel size of the normal convolution layer is 3, the step size is 1, and the padding number of the reflection padding layer is 1. These two layers are used to adjust the number of channels of the feature map; The decoding process can be expressed by the following formula: ; ); ); ; ; in, represents the fused image, Indicates The fused feature map of the layer after upsampling; and Respectively expressed in The layer is used for the wavelet convolution module and the block-based Kolmogorov-Arnold feature extraction module of the decoder; Represents an upsampling operation; the symbol Represents a concatenation operation.
[0035] In some embodiments, the loss function of the first stage is designed as follows: The goal of the first stage training is to enhance the feature extraction capabilities of the infrared encoder and the visible light encoder respectively, ensuring that they can accurately extract infrared target information and visible light texture details. To this end, the first stage training mainly uses the following loss function: Pixel Loss: Pixel loss is used to measure the difference between the reconstructed image and the input image at the pixel level to ensure that the output image is as close to the input image as possible in terms of value.
[0036] The formula of the first stage loss function is: ; in, represents the first stage loss function, 、 and Represent pixel loss, structural similarity loss and first-stage gradient loss function respectively; The formula of pixel loss function is: ; in, and Represent the L1 norm and L2 norm respectively; and Represent the output image and input image respectively; and Respectively represent the height and width of the image; Structural Similarity Loss (SSIM Loss): Although pixel loss can effectively measure the numerical difference between the reconstructed image and the input image, it cannot accurately reflect the perceived quality of the human eye. To address this limitation, we further introduce structural similarity loss, which uses brightness, contrast, and structural information as evaluation indicators to measure the perceptual similarity between the output image and the input image. The formula of the structural similarity loss function is: ; in, represents the structural similarity index; Gradient Loss: To further improve the quality of the reconstructed image, we introduced gradient loss. This loss function can more effectively capture the structural information in the image, especially in terms of retaining edge and texture details. The formula of the first stage gradient loss function is: ; ; ; in, Indicates that the image is at position The pixel value at ; Represents the horizontal gradient, reflecting the pixel change of the image in the x direction; Represents the vertical gradient, reflecting the pixel changes of the image in the y direction.
[0037] It should be noted that in the first stage, the fusion module is removed. Figure 7 As shown in the figure, we construct an infrared autoencoder and a visible light autoencoder respectively, each of which consists of an encoder and a decoder to independently learn and extract infrared target information and visible light texture details.
[0038] like Figure 8 As shown, the training process in the second stage aims to make full use of the multi-scale features extracted by the encoder and integrate the information through the fusion module to generate a high-quality fused image.
[0039] Loss function design in the second stage: In this example, the task of the second stage is to train the fusion module and decoder. In this stage, the parameters of the infrared encoder and the visible light encoder remain frozen, and only the fusion module and the decoder are optimized to ensure that the fusion network can fully utilize multi-scale features to achieve high-quality fusion image reconstruction.
[0040] In some embodiments, the formula of the second stage loss function is: ; ; ; ; in, represents the second stage loss function, , and Represent intensity loss, detail texture loss, and second-stage gradient loss function respectively; and Represent the height and width of the image respectively; , and Represent the fused image, infrared image and visible light image respectively; The symbol represents the L1 norm; The symbol indicates element-wise maximum selection; The symbol represents the Sobel gradient operator.
[0041] In order to further verify the performance of this method, comparative experiments were conducted on three datasets: TNO, M3FD, and LLVIP. Tables 1, 2, and 3 show the experimental results of the fusion method proposed in this paper on the three datasets, and compared and analyzed them with the existing methods to objectively evaluate its performance and effect.
[0042] Table 1: Comparison of quantitative evaluation indicators between the present invention and the existing fusion methods on the TNO dataset.
[0043]
[0044] Table 2: Comparison of quantitative evaluation indicators between the present invention and the existing fusion method on the M3FD dataset.
[0045]
[0046] Table 3: Comparison of quantitative evaluation indicators between the present invention and the existing fusion methods on the LLVIP dataset.
[0047]
[0048] Check 1, Check 2 and Check 3 for a snowstorm: [1] XYDEAS CS, PETROVIĆ V. Objective image fusion performancemeasure[J / OL]. Electronics Letters, 2000, 36(4):308-309[2025-04-13]. http: / / digitallibrary.theiet.org / doi / 10.1049 / el%3A20000267. DOI:10.1049 / el:20000267. [2] RAO Y J. In-fiber bragg grating sensors[J / OL]. DOI:10.1088 / 0957-0233 / 8 / 4 / 002. [3] ROBERTS JW, VAN AARDT J, AHMED F. Assessment of image fusionprocedures using entropy, image quality, and multispectral classification[J]. [4] CHEN H, VARSHNEY P K. A human perception inspired quality metricfor image fusion based on regional information[J / OL]. Information Fusion,2007, 8(2):193-207[2025-04-13]. https: / / linkinghub.elsevier.com / retrieve / pii / S1566253505000898. DOI:10.1016 / j.inffus.2005.10.001. [5] CHEN Y, BLUM RS.A new automated quality assessment algorithmfor image fusion[J / OL]. DOI:10.1016 / j.imavis.2007.12.002. [6] ESKICIOGLU A M, FISHER P S. Image quality measures and theirperformance[J / OL]. IEEE Transactions on Communications, 1995, 43(12): 2959-2965[2025-04-13]. http: / / ieeexplore.ieee.org / document / 477498 / . DOI:10.1109 / 26.477498. [7] SHEIKH H R, BOVIK A C. Image information and visual quality[J]. [8] CUI G. Detail preserved fusion of visible and infrared imagesusing regional saliency extraction and multi-scale image decomposition[J / OL].DOI:10.1016 / j.optcom.2014.12.032. [9] NAIDU V P S. Image fusion technique using multi-resolutionsingular value decomposition[J / OL]. Defence Science Journal, 2011, 61(5): 479[2025-04-13]. http: / / publications.drdo.gov.in / ojs / index.php / dsj / article / view / 705. DOI:10.14429 / dsj.61.705.
[10] GAO Z. Texture clear multi-modal image fusion with jointsparsity model[J / OL]. DOI:10.1016 / j.ijleo.2016.09.126.
[11] LIU Y, CHEN X, WARD R K, et al. Image fusion with convolutionalsparse representation[J / OL]. IEEE Signal Processing Letters, 2016, 23(12):1882-1886[2025-04-13]. https: / / ieeexplore.ieee.org / document / 7593316 / . DOI:10.1109 / LSP.2016.2618776.
[12] LI H, WU X J, KITTLER J. MDLatLRR: A novel decomposition methodfor infrared and visible image fusion[J / OL]. IEEE Transactions on ImageProcessing, 2020, 29: 4733-4746[2025-04-13]. https: / / ieeexplore.ieee.org / document / 9018389 / . DOI:10.1109 / TIP.2020.2975984.
[13] LI H, WU X J, KITTLER J. Infrared and visible image fusion usinga deep learning framework[C / OL] / / 2018 24th International Conference onPattern Recognition (ICPR). Beijing: IEEE, 2018: 2705-2710[2025-04-13].https: / / ieeexplore.ieee.org / document / 8546006 / . DOI:10.1109 / ICPR.2018.8546006.
[14] LI H, WU X jun, DURRANI T S. Infrared and visible image fusionwith ResNet and zero-phase component analysis[J / OL]. Infrared Physics&Technology, 2019, 102: 103039[2025-04-13]. https: / / linkinghub.elsevier.com / retrieve / pii / S1350449519301525. DOI:10.1016 / j.infrared.2019.103039.
[15] LI H, WU X J. DenseFuse: A fusion approach to infrared andvisible images[J / OL]. IEEE Transactions on Image Processing, 2019, 28(5):2614-2623[2024-09-21]. http: / / arxiv.org / abs / 1804.08361. DOI:10.1109 / TIP.2018.2887342.
[16] MA J, YU W, LIANG P, et al. FusionGAN: A generative adversarialnetwork for infrared and visible image fusion[J / OL]. Information Fusion,2019, 48: 11-26[2024-09-29]. https: / / linkinghub.elsevier.com / retrieve / pii / S1566253518301143. DOI:10.1016 / j.inffus.2018.09.004.
[17] MA J, ZHANG H, SHAO Z, et al. GANMcC: A Generative AdversarialNetwork With Multiclassification Constraints for Infrared and Visible ImageFusion[J / OL]. IEEE Transactions on Instrumentation and Measurement, 2021, 70:1-14[2024-09-29]. https: / / ieeexplore.ieee.org / document / 9274337 / . DOI:10.1109 / TIM.2020.3038013.
[18] LI H, WU X J, KITTLER J. RFN-Nest: An end-to-end residual fusionnetwork for infrared and visible images[J / OL]. Information Fusion, 2021, 73:72-86[2024-09-21]. https: / / linkinghub.elsevier.com / retrieve / pii / S1566253521000440. DOI:10.1016 / j.inffus.2021.02.023.
[19] XU H, WANG X, MA J. DRF: Disentangled representation for visibleand infrared image fusion[J]. IEEE TRANSACTIONS ON INSTRUMENTATION ANDMEASUREMENT, 2021,70.
[20] JUNG H, KIM Y, JANG H, et al. Unsupervised deep image fusion with structure tensor representations[J / OL]. IEEE Transactions on ImageProcessing, 2020, 29: 3845-3858[2025-04-13]. https: / / ieeexplore.ieee.org / document / 8962327 / . DOI:10.1109 / TIP.2020.2966075.
[21] LIU J, FAN DOI:10.48550 / arXiv.2203.16220.
[22] LI H, WU DOI:10.1016 / j.inffus.2023.102147. Compared with the prior art, the present invention has the following beneficial effects: 1. The wavelet convolution modules (EWTB and DWTB) designed in this invention can more effectively separate the feature information of different frequencies. Compared with the existing method that only uses ordinary convolution for feature extraction, this method can more accurately capture the key features of infrared images and visible light images, avoid the interference caused by mixed frequency information, thereby improving the feature representation ability and further improving the quality of the fused image.
[0049] 2. The present invention introduces the Kolmogorov-Arnold representation theorem to help the network learn more complex mapping relationships, breaking through the problem of weak feature expression ability of using fixed activation functions in traditional deep learning networks.
[0050] 3. The fusion module proposed in the present invention can adaptively adjust the contribution ratio of each feature source (infrared and visible light features), thereby achieving a more balanced and efficient fusion process and improving the quality and information expression capability of the fused image.
[0051] 4. The loss function designed in the present invention optimizes the shortcomings of the existing infrared and visible light image fusion method in terms of loss function, effectively improves the quality of the fused image, and better preserves its texture details.
[0052] 5. Through comparison on the TNO, M3FD and LLVIP datasets (a total of 186 pairs of infrared and visible light images), the algorithm proposed in this paper outperforms 14 existing fusion algorithms and achieves state-of-the-art performance in terms of visual quality and objective evaluation indicators.
[0053] and Figure 1 Corresponding to the method, refer to Fig. 9 , an embodiment of the present invention provides an infrared image and visible light image fusion system based on deep learning, comprising: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the above method.
[0054] It can be seen that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0055] In addition, an embodiment of the present invention further discloses a computer program product or a computer program, which is stored in a computer-readable storage medium. A processor of a computer device can read the computer program from a computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the above method. Similarly, the contents of the above method embodiment are all applicable to the storage medium embodiment, and the functions specifically implemented by the storage medium embodiment are the same as those of the above method embodiment, and the beneficial effects achieved are also the same as those achieved by the above method embodiment.
[0056] It will be appreciated by those skilled in the art that all or some of the methods disclosed above and the system may be implemented as software, firmware, hardware and appropriate combinations thereof. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or as hardware, or as an integrated circuit, such as an application specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or transient medium). As known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0057] The above is a specific description of the preferred implementation of the present disclosure, but the present disclosure is not limited to the above-mentioned implementation mode. Technical personnel familiar with the field can also make various equivalent deformations or substitutions without violating the spirit of the present disclosure. These equivalent deformations or substitutions are all included in the scope defined by the claims of the present disclosure.
Claims
1. A method for fusion of infrared images and visible light images based on deep learning, characterized in that: The method comprises the following steps: Constructing an image fusion network, the image fusion network comprising an encoder, a fusion module and a decoder; the encoder comprising an infrared encoder and a visible light encoder; The encoder and the decoder form a U-shaped network, and the U-shaped network is trained based on the first-stage loss function to obtain a trained encoder; Fixing the parameters of the encoder, and training the image fusion network based on the second-stage loss function to obtain a trained image fusion network; The trained image fusion network is used to fuse the input infrared image and visible light image to generate a fused image.
2. The method according to claim 1, characterized in that The method of using a trained image fusion network to fuse the input infrared image and the visible light image to generate a fused image includes: Obtain input infrared images and visible light images; Inputting the infrared image and the visible light image into an encoder and outputting a multi-layer feature map; the feature map includes an infrared feature map and a visible light feature map; The infrared feature map and the visible light feature map in each layer of feature map are spliced in the channel direction to obtain a multi-layer spliced map; The multi-layer spliced image is input into the fusion module for processing to obtain a fused feature map; The fused feature map is input into the decoder for processing to obtain a fused image.
3. The method according to claim 2, characterized in that The encoder includes an infrared encoder and a visible light encoder, and the infrared encoder and the visible light encoder each include three wavelet convolution modules for encoders and two Kolmogorov-Arnold feature extraction modules based on block cascading in sequence; the wavelet convolution module for the encoder is composed of a common convolution layer, a reflection filling layer, a wavelet convolution layer, a batch normalization layer and a rectified linear unit in sequence; The calculation formula of the infrared characteristic map is: ; ; The calculation formula of the visible light characteristic graph is: ; ; in, and Indicates the infrared encoder and visible light encoder in the Feature map of the layer; and Represent the input infrared image and visible light image respectively, , , , Respectively represent the number of channels, height and width of the image; and Respectively represent the infrared encoder and visible light encoder The wavelet convolution module of the layer is used for the encoder; and Respectively represent the infrared encoder and visible light encoder A block-based Kolmogorov-Arnold feature extraction module for the layer.
4. The method according to claim 3, characterized in that The fusion module includes a channel attention mechanism, a large kernel attention mechanism and a convolution module, and the convolution module includes a common convolution layer, a reflection padding layer, a batch normalization layer and a rectified linear unit; The multi-layer spliced image is input into the fusion module for processing to obtain a fusion feature map, including: The multi-layer splicing graph is input into the corresponding channel attention mechanism respectively to obtain the channel attention feature map; the channel attention mechanism is expressed by the formula as follows: ; in, Indicates The channel attention feature map of the layer, represents global average pooling, represents a fully connected layer; Represents the sigmoid activation function; symbol Represents matrix multiplication; The channel attention feature map is input into the large core attention mechanism to obtain the large core convolution attention feature map; the large core attention mechanism is expressed by the following formula: ; ; in, Indicates Layer large kernel convolution attention feature map, represents depthwise convolution; represents a depth-dilated convolution; express Convolution operation; The multi-layer attention splicing map is processed by the convolution module to obtain a fused feature map; the convolution module is expressed by the following formula: ; in, Indicates Layer fusion feature map, Represents a convolution operation, with a convolution kernel of 3 and a step size of 1; Indicates a reflection fill operation, and its fill number is 1; Represents batch normalization operation; represents a rectified linear unit.
5. The method according to claim 4, characterized in that The decoder includes three wavelet convolution modules for the decoder and two Kolmogorov-Arnold feature extraction modules based on block connection, which are cascaded in sequence; The calculation formula of the fused image is: ; ); ); ; ; in, represents the fused image, Indicates The fused feature map of the layer after upsampling; and Respectively expressed in The wavelet convolution and block-based Kolmogorov-Arnold feature extraction modules for the decoder are implemented in layers; Represents an upsampling operation; the symbol Represents a concatenation operation.
6. The method according to claim 1, characterized in that The formula of the first stage loss function is: ; in, represents the first stage loss function, 、 and Represent pixel loss, structural similarity loss and first-stage gradient loss function respectively; The formula of pixel loss function is: ; in, and Represent the L1 norm and L2 norm respectively; and Represent the output image and input image respectively; and Respectively represent the height and width of the image; The formula of the structural similarity loss function is: ; in, represents the structural similarity index; The formula of the gradient loss function in the first stage is: ; ; ; in, Indicates that the image is at position The pixel value at ; Represents the horizontal gradient, reflecting the image Pixel changes in direction; Represents the vertical gradient, reflecting the image The pixel change in direction.
7. The method according to claim 1, characterized in that The second stage Loss Function The formula is: ; ; ; ; in, represents the second stage loss function, , and Represent intensity loss, detail texture loss, and second-stage gradient loss function respectively; and Respectively represent the height and width of the image; , and Represent the fused image, infrared image and visible light image respectively; The symbol represents the L1 norm; The symbol indicates element-wise maximum selection; The symbol represents the Sobel gradient operator.
8. A deep learning-based infrared image and visible light image fusion system, characterized in that: The system comprises: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 7.
9. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program is used to perform the method according to any one of claims 1 to 7 when executed by the processor.
Citation Information
Patent Citations
Infrared and visible light image fusion method based on visual enhancement
CN116363036A
Infrared-visible light image fusion method based on attention mechanism
CN117197624A
Image reconstruction algorithm, image processing method and device, vehicle and storage medium
CN119579713A
Method and apparatus for computer vision
US20210125338A1
Cited By
Unified image fusion method and system based on adaptive distribution difference perception
CN120410893A
A unified image fusion method and system based on adaptive distribution difference perception
CN120410893B
Infrared light and visible light image fusion method, system and device based on double-branch cross-domain feature fusion and medium
CN121032817A
Image fusion method and system based on depth estimation and dual-module attention
CN121437293A