Different-source image fusion method based on attention mechanism and dense connection
Through the method based on attention mechanism and dense connection, the problem of difference in the feature description of infrared and visible light images is solved, and the stable extraction and full fusion of heterologous images are achieved, which improves the comprehensiveness and quality of image information acquisition.
Patent Information
- Application Number
- CN202510201406.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, infrared images and visible light images have different image feature descriptions due to different sensor imaging methods. Most algorithms cannot fully explore multimodal source image features, especially in harsh environments, and the existing methods lack the perception of significant areas.
A heterologous image fusion method based on attention mechanism and dense connection is adopted, features are extracted through the encoder, features are enhanced by spatial and channel attention mechanisms, and images are reconstructed in combination with the decoder to achieve complementary fusion of infrared and visible images.
It realizes stable extraction and full fusion of infrared and visible image features, and can obtain more comprehensive image information in harsh environments, improving the generalization ability and effect of image fusion.
Smart Images

Figure CN120339763A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a heterogeneous image fusion method based on an attention mechanism and dense connections. Background Art
[0002] Image fusion can integrate and process image information of the same scene from different sensors, and is one of the sub-fields of information processing. In some harsh environments, such as low light, strong light, smoke occlusion and other environments, a single image sensor can only collect partial information of the scene, and even the performance degrades and it is difficult to obtain effective information. However, using a multi-sensor system that combines several sensors to collect image information of the same scene can obtain richer and more reliable image information and achieve a more comprehensive perception of the scene. However, due to different imaging methods of the sensors, there are huge differences in the image feature description methods between infrared images and visible light images. At the same time, parameter differences of the sensors also lead to differences in gray scale, viewing angle, resolution, image quality, etc. between infrared images and visible light images. The image fusion results of most existing algorithms do not sufficiently perceive the typical features of multi-modal source images. Due to the difference in the imaging principles of infrared and visible light sensors, most deep learning-based methods extract features by simply stacking convolutional layers, resulting in insufficient mining of feature information of multi-modal source images and insufficient information interaction in the fused image, especially the inability to well perceive the significant regions of multi-modal images. Summary of the Invention
[0003] Object of the Invention: In order to solve the problems existing in the above-mentioned prior art, the present invention provides a heterogeneous image fusion method based on an attention mechanism and dense connections.
[0004] Technical Solution: The present invention provides a heterogeneous image fusion method based on an attention mechanism and dense connections, which specifically includes the following steps:
[0005] Step 1: Input two heterogeneous images A and B into an encoder respectively for feature extraction;
[0006] Step 2: Input the features of the two heterogeneous images into a fusion module based on an attention mechanism to obtain a fused feature image;
[0007] Step 3: Input the fused image into a decoder module for image reconstruction to obtain the final fused image.
[0008] Further, in step 1, the encoder includes a first convolutional layer and first to fourth densely connected layers; the first convolutional layer is connected to the first densely connected layer, and the first to fourth densely connected layers are connected in sequence. Each densely connected layer takes the outputs of all previous layers as the input of this layer. The convolutional kernel size of the first convolutional layer is 3×3, the number of input channels is 1, and the number of output channels is 12. The convolutional kernel sizes of the densely connected layers are all 3×3, and the number of output channels are all 12. Each densely connected layer is activated using the ReLU activation function.
[0009] Further, the fusion module based on the attention mechanism includes two spatial attention modules, an addition module, and a channel attention module; the features of two heterogeneous images are respectively input into the corresponding spatial attention modules, and then the outputs of the two spatial attention modules are input into the channel attention module through the addition module.
[0010] Further, the spatial attention module extracts the channel maximum value and the channel average value of the input feature to obtain the maximum value feature and the average value feature, performs per-channel convolution on the maximum value feature and the average value feature, and then activates it through the Sigmoid activation function to obtain the spatial attention weight output by the Sigmoid activation function. The original image input into the spatial attention module is multiplied by the spatial attention weight to obtain the enhanced image feature.
[0011] Further, the channel attention module uses a global pooling layer to compress the input feature to obtain the compressed channel attention weight, uses a first fully connected layer to reduce the dimension of the compressed channel attention weight, and activates it through the ReLU activation function. A second fully connected layer is used to increase the dimension of the output of the ReLU activation function, and then it is activated through the Sigmoid activation function. The output of the Sigmoid activation function is multiplied by the original image input into the channel attention module to obtain the fused feature image.
[0012] Further, the decoder module includes five convolutional layers connected in sequence.
[0013] Further, when training the encoder and the decoder, the loss function Loss is as follows:
[0014] Loss = Loss_pixel + λLoss_ssim
[0015] Loss_pixel = MSE(O, I)
[0016] Loss_ssim = 1 - SSIM(O, I)
[0017] Among them, I represents the input image, O represents the output image, λ is a coefficient, Loss_pixel is the structural loss, Loss_ssim is the pixel loss, SE(.) represents the mean square error between the input image and the output image, MSE(.) is the structural similarity function, and the expression of SSIM(O, I) is as follows:
[0018]
[0019] Among them, μ I 、μ O respectively represent the means of the input image and the output image, σ I 、σ O respectively represent the standard deviations of the input image and the output image, σ IO represents the covariance between the input image and the output image, and c1, c2 are two non-zero constants.
[0020] An electronic device / system for heterologous image fusion includes a processor and a memory. The memory stores the execution instructions of the processor, and the processor is configured to execute the execution instructions to implement the heterologous image fusion method.
[0021] A computer-readable storage medium is used to store a program, and executing the program realizes the heterologous image fusion method.
[0022] Beneficial effects: The present invention constructs an encoder through dense connections, which can effectively extract image features and reuse shallow features, making feature extraction more stable and sufficient. The present invention uses a fusion module that combines a spatial attention mechanism and a channel attention mechanism to fuse infrared / visible light heterologous image features. Feature enhancement based on spatial attention is performed on different dominant regions of infrared / visible light features, and different weights are assigned to each channel of the fused features based on channel attention, realizing targeted complementary fusion of infrared / visible light image features. The present invention is an unsupervised end-to-end deep learning method that can perform autonomous extraction and fusion of heterologous image features without the need for artificial design of feature extraction and fusion strategies, and has strong generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is the heterologous image fusion network designed by the present invention.
[0024] Figure 2 is a schematic diagram of the encoder structure based on dense connections.
[0025] Figure 3 is a schematic diagram of the fusion module structure based on the attention mechanism.
[0026] Figure 4 is a schematic diagram of the decoder module structure based on convolution.
[0027] Figure 5 It is a schematic diagram of the spatial attention module in the fusion module.
[0028] Figure 6 It is the schematic structure of the channel attention module in the fusion module.
[0029] Figure 7 It is the fusion effect of infrared / visible light image fusion using the present invention, where (a) is the original visible light image, (b) is the original infrared light image, and (c) is the fused image. Detailed implementation manners
[0030] The accompanying drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0031] The core idea of the present invention is to construct a heterogeneous image fusion network based on dense connection and attention mechanism, and propose an infrared / visible light image fusion method including an encoder, a fusion module and a decoder. The specific network structure is as Figure 1 shown, and the process of the present invention will be specifically described below.
[0032] Step 1: The encoder module for feature extraction is composed of a common convolutional layer and 4 dense connection layers, as Figure 2 shown. The convolutional layer has a convolutional kernel size of 3×3, an input channel number of 1, and an output channel number of 12; the convolutional kernel sizes of the dense connection layers are all 3×3, and the output channel numbers are all 12. Each dense connection layer receives the outputs of all the previous layers as inputs, and multiplexes the shallow features through dense connection to realize the feature extraction of infrared / visible light source images.
[0033] The convolutional layer performs convolutional processing on the input source image to initially extract its features. The input of each dense connection layer is a combination of the outputs of all the previous dense connection layers, and this combination is carried out through channel-level connection. At the same time, each dense connection layer uses the ReLU activation function for activation. This encoder network has strong stability, and the use of the dense connection layer and the ReLU activation function can enable the encoder to more effectively extract and utilize the source image features.
[0034] Step 2: The fusion module is composed of a spatial attention mechanism and a channel attention mechanism, as Figure 3 shown. The infrared / visible light source image features are respectively input into two spatial attention modules, and for different regions of the features, the regions with significant intensity are enhanced. After adding the enhanced infrared / visible light image features, the features are enhanced based on the channel attention mechanism to obtain the fused image features.
[0035] As Figure 4 shown, the spatial attention module is mainly composed of a convolutional layer and a Sigmoid activation function.
[0036] The source image features of H×W×C (where H is the height, W is the width, and C is the number of channels) are input into the spatial attention module. The channel maximum value and the channel average value are calculated respectively to obtain the maximum value feature of H×W×1 and the average value feature of H×W×1.
[0037] The maximum value feature and the average value feature are convolved channel by channel and activated through the Sigmoid activation function. The dimension of the convolutional layer is 1×1×2, and the spatial attention weight of H×W×1 is output. The source image features of H×W×C are multiplied by the spatial attention weight of H×W×1 to obtain the enhanced image features of dimension H×W×C. The above spatial attention enhancement processing is performed on the infrared image and the visible light image features respectively to obtain the image feature 1 and the image feature 2 of dimension H×W×C. The two features are added to obtain the image feature 3 of dimension H×W×C.
[0038] As Figure 5 shown, the channel attention module is mainly composed of a global pooling layer, a fully connected layer, a ReLU activation function, and a Sigmoid activation function. It mainly includes three parts: Squeeze (compression), Excitation (excitation), and Reweight (reweighting).
[0039] The image feature 3 is input into the channel attention module. The compression part performs global pooling, and the image feature 3 of H×W×C is compressed into the channel attention weight 1 of 1×1×C; the excitation part reduces the dimension of the weight 1 through the fully connected layer 1 and activates it through the ReLU activation function to obtain the weight 2; the dimension of the weight 2 is increased through the fully connected layer 2 and activated through the Sigmoid activation function to obtain the weight 3 of 1×1×C; the reweighting part multiplies the weight 3 by the input feature 3 to obtain the fused image feature of H×W×C.
[0040] Step 3: The decoder module for image reconstruction is composed of five convolutional layers, as Figure 6 shown. The features of multiple channels are reconstructed into a single channel to obtain the final infrared / visible light fused image.
[0041] Before performing heterologous image fusion, the above network is first trained. During network training, only the feature extraction and image reconstruction parts of the network, as well as the encoder and decoder modules, are trained, and the fusion module is not added. The COCO2014 dataset is selected for network training. The Loss function includes the structural loss Loss_pixel and the pixel loss Loss_ssim, where λ is a coefficient, taken as 100 here. The formula is as follows:
[0042] Loss = Loss_pixel + λLoss_ssim
[0043] Loss_pixel = MSE(O, I)
[0044] Loss_ssim = 1 - SSIM(O, I)
[0045] Among them, I represents the training source image, O represents the network output image, and MSE (Mean Squared Error) is the mean square error between the input image and the output image.
[0046] Loss_pixel measures the pixel gray-scale change and image intensity change of the fusion result by calculating the pixel similarity between the fusion image and the input infrared image and visible light image. Here, the mean square error calculation method is adopted. However, since this calculation method only calculates through the differences of single pixel points, it is more inclined to capture the local changes of the image and cannot effectively reflect the global similarity between the input image and the output image. In addition, the presence of noise also affects Loss_pixel.
[0047] Loss_ssim measures the similarity in aspects such as contrast, brightness, and structure between the original image and the fusion image. The contour and structural features of the image are included in the evaluation scope. Therefore, it has stronger anti-noise ability compared to Loss_pixel. Among them, the SSIM (Structural Similarity Index Measure) function is a structural similarity function that can measure the distortion degree of the output image. In the fusion task of infrared images and visible light images, it can evaluate the structural similarity between the fusion image and the original image. Its calculation method is as follows:
[0048]
[0049] Among them, μ I 、μ O respectively represent the means of the input image and the output image, σ I 、σ O represent the standard deviations, and σ IOis the covariance between the input image and the output image, and c1, c2 are two non-zero constants, usually taken as 0.012 and 0.032, to prevent the denominator from being zero.
[0050] The fusion effect of infrared / visible light image fusion using the present invention is as Figure 7 shown. It can be seen that the present invention has a good fusion effect and can extract most features of the original image.
[0051] In addition, it should be noted that, in the above specific embodiments, the various specific technical features described can be combined in any appropriate way without contradiction. To avoid unnecessary repetition, the present invention will not describe various possible combination methods separately.
Claims
1. A heterogeneous image fusion method based on attention mechanism and dense connection, characterized in that Specifically, it includes the following steps: Step 1: Input two heterologous images A and B into the encoder for feature extraction respectively; Step 2: Input the features of the two heterologous images into the fusion module based on the attention mechanism to obtain the fused feature image; Step 3: Input the fused image into the decoder module for image reconstruction to obtain the final fused image.
2. The method for fusing heterogeneous images based on an attention mechanism and dense connection according to claim 1, wherein In Step 1, the encoder includes a first convolutional layer and the first to fourth dense connection layers; the first convolutional layer is connected to the first dense connection layer, and the first to fourth dense connection layers are connected in sequence. Each dense connection layer takes the outputs of all previous layers as the input of this layer. The convolutional kernel size of the first convolutional layer is 3×3, the number of input channels is 1, and the number of output channels is 12. The convolutional kernel sizes of the dense connection layers are all 3×3, and the number of output channels are all 12; each dense connection layer is activated using the ReLU activation function.
3. A method for fusing heterogeneous images based on an attention mechanism and dense connection according to claim 1, characterized in that The fusion module based on the attention mechanism includes two spatial attention modules, an addition module, and a channel attention module; the features of the two heterologous images are input into the corresponding spatial attention modules respectively, and then the outputs of the two spatial attention modules are input into the channel attention module through the addition module.
4. A heterogeneous image fusion method based on an attention mechanism and dense connection according to claim 3, characterized in that The spatial attention module extracts the channel maximum value and the channel average value of the input feature to obtain the maximum value feature and the average value feature, performs per-channel convolution on the maximum value feature and the average value feature, and then activates it through the Sigmoid activation function to obtain the spatial attention weight output by the Sigmoid activation function. Multiply the original image input into the spatial attention module by the spatial attention weight to obtain the enhanced image feature.
5. The method for fusing heterogeneous images based on attention mechanism and dense connection according to claim 3, wherein The channel attention module uses a global pooling layer to compress the input feature to obtain the compressed channel attention weight, uses a first fully connected layer to reduce the dimension of the compressed channel attention weight, and activates it through the ReLU activation function. Use a second fully connected layer to increase the dimension of the output of the ReLU activation function, and then activate it through the Sigmoid activation function. Multiply the output of the Sigmoid activation function by the original image input into the channel attention module to obtain the fused feature image.
6. A method for fusing heterogeneous images based on an attention mechanism and dense connection according to claim 1, characterized in that, The decoder module includes five convolutional layers connected in sequence.
7. A method for fusing heterogeneous images based on an attention mechanism and dense connections according to claim 1, characterized in that, When training the encoder and the decoder, the loss function Loss is as follows: Loss = Loss_pixel + λLoss_ssim Loss_pixel = MSE(O, I) Loss_ssim = 1 - SSIM(O, I) Where, I represents the input image, O represents the output image, λ is a coefficient, Loss_pixel is the structural loss, Loss_ssim is the pixel loss, SE(.) represents the mean square error between the input image and the output image, MSE(.) is the structural similarity function, and the expression of SSIM(O, I) is as follows: Among them, μ I and μ O represent the means of the input image and the output image respectively, σ I and σ O represent the standard deviations of the input image and the output image respectively, σ IO represents the covariance between the input image and the output image, and c1, c2 are two non-zero constants.
8. An electronic device / system for heterologous image fusion, characterized in that, It includes a processor and a memory. The memory stores the execution instructions of the processor, and the processor is configured to execute the execution instructions to implement the heterologous image fusion method according to any one of claims 1-7.
9. A computer-readable storage medium for storing a program, characterized in that, Execute the program to implement the heterologous image fusion method described in any one of claims 1-7.