Infrared visible light image fusion method driven by triple attention generative adversarial network
Through the triple attention generation adversarial network, combined with nested connection generators and dual discriminators, the loss function is optimized, and the efficient deep fusion of infrared and visible images is achieved, solving the problem of insufficient image quality under low light conditions. The generated fusion images are outstanding in visual and information retention, and are suitable for scenes such as drone night patrols.
Patent Information
- Application Number
- CN202510616849.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-05-14
AI Technical Summary
In the existing infrared and visible image fusion methods, insufficient feature interaction, rigid fusion strategy and one-sided loss function lead to limited image quality and it is difficult to effectively retain infrared target significance and visible light texture details under low light conditions.
A triple attention generation adversarial network is adopted, including channel, space and point attention mechanisms, combined with nested connection generators and dual discriminators, and the deep fusion of infrared and visible images is achieved through multimodal loss function optimization.
The generated fusion images visually present the details of infrared targets and visible scenes more clearly, improve information richness and structural similarity, significantly improve image quality under low-light conditions, and support applications such as drone night patrols.
Smart Images

Figure CN120543392A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing and computer vision technology, and in particular to an infrared and visible light image fusion method driven by a triple attention generative adversarial network. Background Art
[0002] In drone-based nighttime campus patrols, visible light images suffer from low contrast and blurred details due to low light conditions. While infrared images can capture thermal radiation information, they lack texture detail. Existing fusion strategies struggle to achieve an effective balance between the two. Traditional image fusion methods (such as multi-scale transformation) rely on manually designed rules and are prone to losing modality-specific features. Deep learning-based methods (such as FusionGAN and DDcGAN) improve fusion performance through adversarial training, but still suffer from the following shortcomings:
[0003] (1) Insufficient feature extraction:
[0004] The existing network structure does not adequately explore the interaction of channel, spatial and local detail features, resulting in incomplete retention of key information in the fused image (such as infrared target saliency and visible light texture).
[0005] (2) Single fusion strategy:
[0006] It relies on fixed fusion rules (such as simple weighting or splicing) and lacks a trainable adaptive mechanism, making it difficult to cope with modal differences in complex scenarios.
[0007] (3) Limitations of loss function design:
[0008] Only focusing on pixel intensity or single modality features without taking into account the texture details and structural information of multimodal images results in limited quality of the fused image.
[0009] Therefore, there is an urgent need for an end-to-end deep learning framework that can achieve deep fusion and efficient retention of multimodal features through a trainable attention mechanism and optimized loss function. Summary of the Invention
[0010] In view of the above-mentioned shortcomings of the prior art, the purpose of the present invention is to provide an infrared-visible light image fusion method driven by a triple attention generative adversarial network, aiming to solve the problems of insufficient feature interaction, rigid fusion strategy, and one-sided loss function in the prior art of infrared and visible light image fusion.
[0011] In order to achieve the above object, the present invention adopts the following technical solutions:
[0012] A triple attention generative adversarial network-driven infrared and visible light image fusion method, including:
[0013] Step 1: Obtain the registered infrared image and visible light image of the same scene, perform preprocessing, unify the number of channels of the visible light image and infrared image, and stitch them into a multi-channel image;
[0014] Step 2: Set up a feature extraction network including a triple attention module to extract features from multi-channel images and generate a single-channel fusion feature map;
[0015] Step 3: The single-channel fusion feature map generated in step 2 is transformed through a convolutional layer to generate a final fusion image containing infrared target saliency and visible light rich texture;
[0016] Step 4. Based on steps 1-3, the fused image is discriminated against the real infrared and visible light images respectively through a dual discriminator. The distribution difference between the generated image and the real image is calculated through the loss calculation based on the least squares method, guiding the generator to learn feature representations that conform to the bimodal distribution, ensuring that the fused image retains both the infrared target saliency and the visible light texture details.
[0017] Furthermore, in step 1, the infrared image and visible light image after registration of the same scene are obtained, and the resolution of the two images is consistent. The image sources include drone sensors, cameras and other equipment or public data sets; if the visible light image is RGB three-channel, it is converted into a single-channel grayscale image, so that both the visible light and infrared images are single-channel to unify the number of channels, and then the two are normalized, and the formula is used. The pixel value range is adjusted to [-1, 1], and finally the normalized infrared and visible light images are spliced along the channel dimension into a 2×H×W input tensor as the initial input of the generator.
[0018] Furthermore, in step 2, the triple attention module of the feature extraction network includes a channel attention mechanism, a spatial attention mechanism, and a point attention mechanism, wherein the channel attention generates channel weights through a bilinear layer, a ReLU activation function, and a Sigmoid function to dynamically enhance key feature channels; the spatial attention extracts spatial features based on average pooling and maximum pooling, generates a spatial weight tensor through a convolutional layer and a Sigmoid function, and focuses on the area of interest; the point attention analyzes the contextual relevance of each pixel in the feature map, strengthens the interaction of local detail features through residual transformation, and captures subtle texture differences. After the three are independently processed, the feature map is spliced to form a multi-perspective enhanced feature representation.
[0019] Furthermore, in step 2, the feature extraction network also includes an encoder and a decoder. The feature map processed by the triple attention module is input to an encoder consisting of 5 residual blocks, and multi-scale features are gradually extracted through downsampling. The number of channels in each layer increases, thereby realizing deep encoding of modal difference features. The high-level features output by the encoder are input to the simplified nested connection decoder of UNet++, which is connected to the corresponding layer of the encoder through upsampling, and the underlying details and high-level semantics are gradually restored, and finally a single-channel fusion feature map is generated.
[0020] Furthermore, in step 3, the feature map output in step 2 is converted into a fused image with a pixel value range of [-1, 1] through the last convolution layer. If normalization is used in preprocessing, the pixel values need to be restored to the original range of [0, 255] through denormalization, using the following processing formula:
[0021]
[0022] Among them, the left I fused Represents the pixel value of the fused image restored to the original range [0,255] after denormalization, and the right I fused Represents the pixel value output by the last convolutional layer, ranging from [-1, 1]. max(I) represents the maximum value of the original image pixel value, and min(I) represents the minimum value of the original image pixel value.
[0023] The final output contains the final fused image of infrared target saliency and visible light rich texture.
[0024] Furthermore, in step 4, when using the least squares method, the loss function of the discriminator is determined by the mean square error loss between the discriminator output and the preset target label, which ensures that the discriminator can effectively distinguish between real images and generated images. The discriminator loss function is as follows:
[0025]
[0026] Among them, MSE(·) represents the mean square error loss, D real Denotes the output prediction of the discriminator for the real image, D fake Represents the output prediction of the discriminator for the generated image, N represents the dimension of the vector, L D represents the total loss of the discriminator.
[0027] Furthermore, the generator loss function is mainly divided into two parts. The first part is the adversarial loss, which is determined based on the output of the discriminator, and the second part is the content loss, which includes pixel loss and other losses. The samples generated by the generator are input into two different branches of the discriminator, namely D i and D v, respectively calculate the adversarial loss in the visible light and infrared modes. The total adversarial loss of the generator is the sum of the losses of these two branches. The formula for the adversarial loss is as follows:
[0028] L adv =MSE(D v (fake),1)+MSE(D i (fake),1)
[0029] Among them, D v (fake) represents the discriminator’s judgment result on the fused image under the visible spectrum.
[0030] D i (fake) represents the judgment result of the discriminator on the fusion image under the infrared spectrum, MSE(·) represents the mean square error loss, L adv Represents resistance to loss.
[0031] Furthermore, the pixel loss measures the pixel difference between the generated image and the original image using the Frobenius norm, and the formula is:
[0032]
[0033] Among them, Fused represents the fused image, Input ir Represents the original infrared image, Input vis represents the original visible light image, ||·|| F represents the Frobenius norm, ω p Indicates the balance coefficient in loss, L ir pixel and L vis pixel represent pixel losses for infrared and visible light, respectively;
[0034] The gradient loss uses the L1 norm to measure the difference in the gradient domain between the generated image and the real image. The formula is:
[0035]
[0036] Among them, Fused represents the fused image, Input ir Represents the original infrared image, Input vis represents the original visible light image, Grad(·) represents the gradient calculation of the image, ||·|| L represents the L1 norm, ω g represents the balance coefficient in the loss, and represent the gradient loss of infrared and visible light respectively;
[0037] The content loss is the weighted sum of the two:
[0038] L con =L pixel +λL grad
[0039] Among them, L con Indicates content loss, L pixel Indicates pixel loss, L grad represents the gradient loss, and λ represents the balance coefficient between the two losses;
[0040] The total loss of the generator is:
[0041] L G =L adv +L con
[0042] Among them, L G Represents the total loss, L adv represents the adversarial loss generated by the generator, L con Indicates content loss.
[0043] The technical solution adopted by the present invention has the following beneficial effects:
[0044] This paper constructs an efficient infrared and visible light image fusion framework through the collaborative design of a triple attention mechanism, a nested connection generator, a dual discriminator, and a multimodal loss function. The triple attention mechanism strengthens feature representation from multiple dimensions. Channel attention dynamically adjusts the weights of each channel, allowing the model to focus more on key information channels such as infrared thermal radiation characteristics and visible light texture. Spatial attention generates weight tensors by analyzing the spatial distribution of feature maps, effectively highlighting target contours, edges, and other areas of interest. Point attention explores the subtle correlations of each pixel in the feature map and strengthens the interaction of local details. The combination of the three achieves feature complementarity and enhancement at the channel, spatial, and point levels, significantly improving the model's ability to capture complex features in multimodal images. The nested connection generator leverages the multi-scale feature extraction of residual blocks in the encoder and the dense cross-layer connections of the decoder to integrate high-level semantic features while preserving underlying detail information. This achieves progressive recovery from coarse-grained structure to fine-grained texture, avoiding the information loss problem caused by single-scale processing in traditional networks. The dual discriminators are designed for infrared and visible light modalities respectively. Through adversarial training, they force the generator to learn feature representations that conform to the true distribution of the two modalities, ensuring that the fused image retains the target saliency in the infrared image while presenting the rich texture of the visible light image, avoiding the modal bias problem that may be caused by a single discriminator. The loss function system, through the organic combination of adversarial loss, pixel loss, and gradient loss, takes into account the texture details of pixel intensity and gradient domains while constraining the visual authenticity of the generated image. This ensures that the fusion result achieves higher-quality information integration while maintaining the original features of infrared and visible light. Experiments show that the framework's fusion effect on public datasets is significantly better than that of various advanced methods. The generated images not only visually present the details of infrared targets and visible light scenes more clearly, such as the outline of heat source targets and the texture structure of the surrounding environment in night scenes, but also perform outstandingly in key indicators such as information richness and structural similarity, effectively solving the problems of poor visible light image quality under low-light conditions and incomplete feature preservation of traditional fusion methods. In practical applications, this method provides clearer fused images for drone night patrols, enabling subsequent target detection and recognition tasks to more accurately locate and analyze key information in the scene, significantly improving monitoring efficiency and safety in complex environments. At the same time, it also demonstrates strong adaptability in fields such as remote sensing imaging and medical imaging, providing an innovative solution for the deep fusion and practical application of multimodal images, with important technical value and broad engineering application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 This is a diagram of the triple attention mechanism model, including the structural design of channel attention (a), spatial attention (b), and point attention (c);
[0046] Figure 2This is the generator structure diagram;
[0047] Figure 3 It is the structure diagram of the dual discriminator;
[0048] Figure 4 Qualitative analysis results of the attention ablation experiment;
[0049] Figure 5 This is a qualitative analysis diagram of the fusion effect of the TNO dataset. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical solution and effect of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0051] A triple attention generative adversarial network-driven infrared and visible light image fusion method is implemented through the following core modules and steps:
[0052] 1) Image acquisition and preprocessing
[0053] ① Data input: Get the registered infrared image of the same scene (single channel, denoted as I ir ) and the visible light image single / three channels, denoted as I vis ), both have the same resolution, and their sources include drone sensors, cameras and other equipment or public datasets (TNO datasets).
[0054] ② Preprocessing operation: If the visible light image is RGB three-channel, convert it into a single-channel grayscale image, so that both the visible light and infrared images are single-channel to unify the number of channels, and then normalize the two. The pixel value range is adjusted to [-1, 1], and finally the normalized infrared and visible light images are spliced along the channel dimension into a 2×H×W input tensor as the initial input of the generator.
[0055] 2) Feature extraction and fusion (generator processing)
[0056] The generator structure is as follows Figure 2 As shown, the spliced multi-channel image is processed in turn by the core module of the generator:
[0057] ① Triple attention feature enhancement: The triple attention mechanism model is as follows Figure 1As shown in the figure, after independent processing of channel attention, spatial attention, and point attention, the generated feature maps are concatenated to form a multi-dimensional enhanced feature representation. Channel attention generates channel weights through bilinear layers, ReLU, and Sigmoid to dynamically enhance the key features of the infrared thermal radiation channel and the visible light texture channel; spatial attention uses average pooling and maximum pooling to extract spatial features and generate spatial weight tensors to focus on areas of interest such as target contours and edges; and point attention analyzes the contextual relevance of pixels and enhances the interaction of local details through residual transformation to capture subtle texture differences.
[0058] ② Multi-scale feature extraction of residual encoder: The feature map processed by the attention module is input to the encoder composed of 5 residual blocks. Multi-scale features are gradually extracted through downsampling (such as convolution or pooling operations). The number of channels in each layer increases (such as 64→128→256→512), realizing deep encoding of modal difference features.
[0059] ③ Dense decoder feature recovery and fusion: The high-level features output by the encoder are input into the simplified nested connection decoder of UNet++, which is connected to the corresponding layer of the encoder through upsampling to gradually recover the underlying details and high-level semantics, and finally generate a single-channel fused feature map.
[0060] 3) Fusion image generation and post-processing
[0061] ① Image reconstruction: The feature map output by the decoder is converted into a fused image I with a pixel value range of [-1, 1] through the last convolution layer (1×1 convolution) fused .
[0062] ② Denormalization: If the normalization operation is used in preprocessing, the pixel values need to be restored to the original range [0, 255] through denormalization. The formula is:
[0063] ③ Output result: Generate a final fused image containing infrared target saliency and visible light rich texture, which is used for subsequent target detection, recognition and other tasks.
[0064] 4) Dual Discriminator Adversarial Training (Training Phase)
[0065] The dual discriminator structure is as follows Figure 3 As shown, during the training phase, the dual discriminators (convolutional neural networks with the same structure) are used to train the fused image I fused Discrimination with real infrared and visible light images:
[0066] The discriminator outputs a single-channel probability value, calculates the distribution difference between the generated image and the real image through the least squares loss, and guides the generator to learn feature representations that conform to the bimodal distribution, ensuring that the fused image retains both the infrared target saliency and the visible light texture details.
[0067] Loss function optimization:
[0068] 1) Discriminator Loss Function: When the discriminator processes real samples, the probability of the output classification being true should approach 1. Therefore, the target label of the real sample is set to 1, and the mean squared error loss between the discriminator output and the target label is calculated based on this. For generated samples, the probability of the discriminator output classification being false should approach 0. Therefore, the target label of the generated sample is set to 0, and the corresponding mean squared error loss is calculated.
[0069] The main goal of the least squares loss function is to enable the discriminator to better distinguish between real samples and generated samples by minimizing the mean squared error. Compared with the traditional cross entropy loss, the advantage of this loss function is that it can provide a more stable training process and generate clearer images.
[0070] When using the least squares method, the discriminator loss function is determined by the mean square error loss between the discriminator output and the preset target label, which ensures that the discriminator can effectively distinguish between real images and generated images. The formula of the discriminator loss function is shown in (1):
[0071]
[0072] Among them, MSE(·) represents the mean square error loss, D real Denotes the output prediction of the discriminator for the real image, D fake Represents the output prediction of the discriminator for the generated image, N represents the dimension of the vector, L D represents the total loss of the discriminator.
[0073] 2) Generator loss function: It is mainly divided into two parts. The first part is the adversarial loss, which is determined based on the output of the discriminator. The second part is the content loss, which is composed of pixel loss and other losses.
[0074] In the least squares (LS) framework, the generator loss is determined by the mean squared error (MSE) between generated samples and real samples. The generator generates samples and sends them to the discriminator, which classifies and outputs probabilities. Because the generator wants the samples to be realistic, it sets the target probability to 1 and uses this to calculate the MSE loss. By minimizing the loss, the generator adjusts its parameters to produce more realistic samples.
[0075] In the actual implementation scenario, the samples generated by the generator are input into two different branches of the discriminator, namely D i and D v , respectively calculate the adversarial loss in the visible light and infrared modalities. The total adversarial loss of the generator is the sum of the losses of these two branches. The formula definition of the adversarial loss is shown in (2):
[0076] L adv=MSE(D v (fake),1)+MSE(D i (fake),1) (2)
[0077] Among them, D v (fake) represents the discriminator’s judgment result on the fused image under the visible spectrum.
[0078] D i (fake) represents the judgment result of the discriminator on the fusion image under the infrared spectrum, MSE(·) represents the mean square error loss, L adv Represents resistance to loss.
[0079] Next, we discuss the pixel loss in content loss. This loss is quantified using the Frobenius norm and is used to measure the difference between the generated image and the original image at the pixel level. Since visible light images typically have high contrast, when calculating pixel loss, the loss between the generated image and the two original images is calculated separately and assigned different weights to control the original image to contribute more pixel intensity information to the generated image. The pixel loss expression is defined in (3).
[0080]
[0081] Among them, Fused represents the fused image, Input ir Represents the original infrared image, Input vis represents the original visible light image, ||·|| F represents the Frobenius norm, ω p Indicates the balance coefficient in loss, L ir pixel and L vis pixel Indicates pixel loss for infrared and visible light respectively.
[0082] Gradient loss is used to measure the difference between the generated image and the real image in the gradient domain, reflecting the gradient-level characteristics. Given that infrared images contain texture details, when calculating the gradient loss, the gradient losses of the generated image and the two original images are calculated separately, and different weights are assigned to them. This is used to control the original image to contribute more texture details to the generated image. The gradient loss expression is shown in (4).
[0083]
[0084] Among them, Fused represents the fused image, Input ir Represents the original infrared image, Input vis represents the original visible light image, Grad(·) represents the gradient calculation of the image, ||·|| L represents the L1 norm, ωg represents the balance coefficient in the loss, and represent the gradient loss of infrared and visible light respectively.
[0085] We derive the content loss by performing a weighted summation of pixel loss and gradient loss. The expression of content loss is defined as shown in (5):
[0086] L con =L pixel +λL grad (5)
[0087] Among them, L con Indicates content loss, L pixel Indicates pixel loss, L grad represents the gradient loss, and λ represents the balance coefficient between the two losses.
[0088] In summary, the total loss of the generator is the sum of the adversarial loss and the content loss. The expression of the total loss is defined as shown in (6):
[0089] L G =L adv +L con (6)
[0090] Among them, L G Represents the total loss, L adv represents the adversarial loss generated by the generator, L con Indicates content loss.
[0091] Specifically, the present invention verifies the implementation effect of the present invention through the following examples:
[0092] (1) Experimental platform and data preparation
[0093] 1) Hardware and Framework: The experiment is implemented based on the PyTorch deep learning framework and runs on a server equipped with NVIDIA V100 GPU.
[0094] 2) Dataset construction:
[0095] ① Training data: 30,000 pairs of infrared and visible light images with a resolution of 128×128 were selected from public urban scene data, covering scenes such as low light at night and complex backgrounds.
[0096] ② Test data: The TNO public dataset contains 32 pairs of multimodal images, covering a variety of scenes such as natural scenes and man-made objects, to verify the fusion performance of the model in real complex environments.
[0097] (2) Network training parameter setting
[0098] 1) Optimization strategy: Use the Adam optimizer to synchronously update the generator and discriminator parameters. Set the initial learning rate to 1E-4, the batch size to 4, and the total number of training rounds to 50 to ensure a balance between model convergence and training efficiency.
[0099] 2) Input processing: The input infrared and visible light images are normalized and input into the generator’s triple attention module after channel splicing.
[0100] (3) Evaluation indicators and testing methods
[0101] 1) Quantitative evaluation indicators: Seven classic indicators are used to comprehensively evaluate the fusion quality, including:
[0102] ① Entropy (En): measures the richness of image information. A higher value indicates more effective information.
[0103] ② Sum of Correlated Differences (SCD): This evaluates the degree of modal difference preservation between the original image and the fused image, reflecting the balance between infrared and visible light features.
[0104] ③ Multi-scale structural similarity (MS-SSIM): measures image structural similarity from multiple spatial scales, reflecting the preservation of texture and edge details;
[0105] In addition, it also includes mutual information (MI), pixel-level fusion quality index (FMI_pixel), etc., which comprehensively evaluate from the dimensions of information sharing, feature alignment, etc.
[0106] 2) Comparison Method: Five currently advanced fusion algorithms are selected for comparison, including NestFuse, RFN-Nest, DDcGAN, SeAFusion, and U2-Fusion. All comparison models use open source code and maintain the original parameter settings.
[0107] (4) Ablation experiment design
[0108] To verify the effectiveness of the triple attention mechanism, a comparative experiment was designed: under the same training parameters, a TA-NcGAN model containing a complete triple attention module (denoted as "with attention") and a simplified model without the attention module (denoted as "without attention") were trained respectively, and the differences were analyzed through qualitative visual comparison and quantitative indicators. The qualitative analysis results are shown in Figure 2. Figure 4 The results show that the fused image without the attention model is brighter overall, with reduced contrast, blurred texture details, and reduced significance of key information such as infrared heat source areas. In terms of quantitative indicators, the attention model outperforms the non-attention model in six metrics, including En, SCD, and MS-SSIM, demonstrating the key role of the triple attention mechanism in feature enhancement and information retention. The quantitative comparison data is shown in Table 1.
[0109] Table 1. Ablation comparison of attention modules
[0110] Table.1 Quantitative Comparison of Attention Module Ablation
[0111]
[0112] (5) Comparative experimental results and analysis
[0113] 1) Qualitative analysis: Qualitative fusion effect such as Figure 5 As shown in the figure, in typical scenes from the TNO dataset, such as nighttime streets, the fused image generated by TA-NcGAN clearly preserves both the outlines of heat source targets in infrared images, such as pedestrians and vehicle engines, and texture details in visible light images, such as road ridges and building windows. Its contrast and color reproduction are superior to those of contrast methods. For example, in low-light scenes, contrast methods may suffer from blurred edges of heat source targets or overly smoothed visible light textures. However, the fusion results of TA-NcGAN accurately present the target location and surrounding environment details, resulting in a visual effect that is closer to the real scene.
[0114] 2) Quantitative Analysis: Statistical results on 32 pairs of test images show that TA-NcGAN outperforms most compared methods in the En metric, indicating that its fused images contain richer multimodal information. The MS-SSIM metric leads significantly, demonstrating its superiority in structural similarity and detail preservation. The SCD metric remains within a reasonable range, demonstrating that the model effectively balances the modal differences between infrared and visible light, avoiding the over-dominance of single modal features. Overall, TA-NcGAN ranks highly in many of the seven evaluation metrics, demonstrating overall superior performance compared to existing state-of-the-art methods. The quantitative analysis results are shown in Table 2.
[0115] Table 2 Quantitative Analysis of Fusion Results on TNO Dataset
[0116]
[0117] In summary, to address the technical challenges of infrared and visible light image fusion in low-light scenarios, this paper proposes a TA-NcGAN framework. Through the collaborative design of a triple attention mechanism, a nested connection generator, a dual discriminator, and a multimodal loss function, it achieves deep fusion and efficient preservation of multimodal features. The triple attention mechanism dynamically enhances key features from the channel, spatial, and point dimensions. The nested generator achieves progressive integration of multi-scale features. The dual discriminator ensures that the fused image conforms to the bimodal distribution. The multimodal loss function balances pixel intensity and texture detail. Experiments demonstrate that the proposed framework significantly outperforms existing methods in fusion on public datasets, generating images that combine infrared target saliency with visible light texture richness, with outstanding information preservation and structural similarity. Ablation experiments validate the core role of the attention mechanism, effectively improving target detection accuracy in scenarios such as drone night patrols. This method provides a new approach for multimodal image fusion and has broad application potential in fields such as remote sensing and medical imaging, demonstrating both theoretical innovation and practical engineering value.
[0118] Other embodiments of the present invention will readily occur to those skilled in the art after considering the specification and practicing the embodiments disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the invention being indicated by the claims.
Claims
1. A triple attention generative adversarial network driven infrared and visible light image fusion method, characterized by: include: Step 1: Obtain the registered infrared image and visible light image of the same scene, perform preprocessing, unify the number of channels of the visible light image and infrared image, and stitch them into a multi-channel image; Step 2: Set up a feature extraction network including a triple attention module to extract features from multi-channel images and generate a single-channel fusion feature map; Step 3: The single-channel fusion feature map generated in step 2 is transformed through a convolutional layer to generate a final fusion image containing infrared target saliency and visible light rich texture; Step 4. Based on steps 1-3, the fused image is discriminated against the real infrared and visible light images respectively through a dual discriminator. The distribution difference between the generated image and the real image is calculated through the loss calculation based on the least squares method, guiding the generator to learn feature representations that conform to the bimodal distribution, ensuring that the fused image retains both the infrared target saliency and the visible light texture details.
2. The infrared-visible image fusion method driven by triple attention generative adversarial network according to claim 1 is characterized in that: In step 1, the infrared image and visible light image after registration of the same scene are obtained, and the resolution of the two images is consistent. The image sources include drone sensors, cameras and other equipment or public data sets; if the visible light image is RGB three-channel, it is converted into a single-channel grayscale image, so that both the visible light and infrared images are single-channel to unify the number of channels, and then the two are normalized by the formula The pixel value range is adjusted to [-1, 1], and finally the normalized infrared and visible light images are spliced along the channel dimension into a 2×H×W input tensor as the initial input of the generator.
3. The infrared-visible image fusion method driven by triple attention generative adversarial network according to claim 1 is characterized in that: In step 2, the triple attention module of the feature extraction network includes a channel attention mechanism, a spatial attention mechanism, and a point attention mechanism, wherein the channel attention generates channel weights through a bilinear layer, a ReLU activation function, and a Sigmoid function, and dynamically enhances key feature channels; the spatial attention extracts spatial features based on average pooling and maximum pooling, generates a spatial weight tensor through a convolutional layer and a Sigmoid function, and focuses on the area of interest; the point attention analyzes the contextual relevance of each pixel in the feature map, strengthens the interaction of local detail features through residual transformation, and captures subtle texture differences. After the three are independently processed, the feature map is spliced to form a multi-view enhanced feature representation.
4. The infrared-visible image fusion method driven by triple attention generative adversarial network according to claim 1 is characterized in that: In step 2, the feature extraction network also includes an encoder and a decoder. The feature map processed by the triple attention module is input to an encoder consisting of 5 residual blocks. Multi-scale features are gradually extracted through downsampling, and the number of channels in each layer increases to achieve deep encoding of modal difference features. The high-level features output by the encoder are input to the UNet++ simplified nested connection decoder, which is connected to the corresponding layer of the encoder through upsampling to gradually restore the underlying details and high-level semantics, and finally generate a single-channel fusion feature map.
5. The infrared-visible image fusion method driven by triple attention generative adversarial network according to claim 1 is characterized in that: In step 3, the feature map output in step 2 is converted into a fused image with a pixel value range of [-1, 1] through the last convolution layer. If normalization is used in preprocessing, the pixel values need to be restored to the original range of [0, 255] through denormalization, using the following processing formula: Among them, the left I fused Represents the pixel value of the fused image restored to the original range [0, 255] after denormalization, and the right I fused Represents the pixel value output by the last convolutional layer, ranging from [-1, 1], max(I) represents the maximum value of the original image pixel value, and min(I) represents the minimum value of the original image pixel value; The final output contains the final fused image of infrared target saliency and visible light rich texture.
6. The infrared-visible image fusion method driven by triple attention generative adversarial network according to claim 1 is characterized in that: In step 4, when using the least squares method, the discriminator's loss function is determined by the mean square error loss between the discriminator output and the preset target label, which ensures that the discriminator can effectively distinguish between real images and generated images. The discriminator loss function is as follows: Among them, MSE(·) represents the mean square error loss, D real Denotes the output prediction of the discriminator for the real image, D fake Represents the output prediction of the discriminator for the generated image, N represents the dimension of the vector, L D represents the total loss of the discriminator.
7. The infrared-visible image fusion method driven by triple attention generative adversarial network according to claim 1 is characterized in that: The generator loss function is mainly divided into two parts: the first part is the adversarial loss, which is determined based on the output of the discriminator, and the second part is the content loss, including pixel loss and other losses, where the samples generated by the generator are input into two different branches of the discriminator, namely D i and D v , respectively calculate the adversarial loss in the visible light and infrared modes. The total adversarial loss of the generator is the sum of the losses of these two branches. The formula for the adversarial loss is as follows: L adv =MSE(D v (fake),1)+MSE(D i (fake),1) Among them, D v (fake) represents the discriminator’s judgment result on the fused image under the visible spectrum, D i (fake) represents the judgment result of the discriminator on the fusion image under the infrared spectrum, MSE(·) represents the mean square error loss, L adv Represents resistance to loss.
8. The infrared-visible image fusion method driven by triple attention generative adversarial network according to claim 7 is characterized in that: The pixel loss measures the pixel difference between the generated image and the original image using the Frobenius norm, and the formula is: Among them, Fused represents the fused image, Input ir Represents the original infrared image, Input vis represents the original visible light image, ||·|| F represents the Frobenius norm, ω p Indicates the balance coefficient in loss, L ir pixel and L vis pixel represent pixel losses for infrared and visible light, respectively; The gradient loss uses the L1 norm to measure the difference in the gradient domain between the generated image and the real image. The formula is: Among them, Fused represents the fused image, Input ir Represents the original infrared image, Input vis represents the original visible light image, Grad(·) represents the gradient calculation of the image, ||·|| L represents the L1 norm, ω g represents the balance coefficient in the loss, and represent the gradient loss of infrared and visible light respectively; The content loss is the weighted sum of the two: THE con =L pixel +λL grad Among them, L con Indicates content loss, L pixel Indicates pixel loss, L grad represents the gradient loss, and λ represents the balance coefficient between the two losses; The total loss of the generator is: L G =L adv +L con Among them, L G Represents the total loss, L adv represents the adversarial loss generated by the generator, L con Indicates content loss.
Citation Information
Patent Citations
Infrared and visible light image fusion method based on multi-scale attention mechanism
CN115423734A
Infrared and visible light image fusion method based on multi-discriminator generative adversarial network
CN115601282A
Image fusion method and system based on double-discriminator generative adversarial network
CN115830384A
Multi-modal medical image fusion method based on SGDD GAN
CN117475268A
4K image defogging algorithm based on weak supervision contrast learning
CN118918039A