Infrared and visible light image fusion method based on cross-modal complementary learning

By adopting a cross-modal complementary learning method in infrared and visible image fusion, using dual-branch dense residual blocks and spatial attention blocks to extract and enhance feature information, combining adversarial training and content loss function to optimize the network, the problem of poor fusion quality and visual effect in the prior art is solved, and higher quality image fusion is achieved.

CN120107082APending Publication Date: 2025-06-06NANKAI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510181813.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing infrared and visible image fusion methods cannot fully retain source image information during feature extraction, and the convolutional layer has limited long-distance dependence ability to capture, resulting in poor fusion quality and visual effect, and the instability of the training process of the generative adversarial network, which is prone to artifacts and noise.

Method used

Using a method based on cross-modal complementary learning, texture details are extracted from visible light images through double-branch dense residual blocks, contrast information is extracted from infrared images, and feature information is enhanced using spatial attention blocks and interactive complementary mining engines, combining adversarial training loss function and content loss function to optimize the network.

Benefits of technology

The maximum preservation of source image information is achieved, the fusion quality and visual effect is improved, the limitations of convolutional layers in capturing long-distance dependencies are overcome, and the stability is improved the quality of image fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107082A_ABST
    Figure CN120107082A_ABST
Patent Text Reader

Abstract

The invention provides an infrared and visible light image fusion method based on cross-modal complementary learning, and relates to the technical field of image fusion. The invention discloses an infrared and visible light image fusion method based on cross-modal complementary learning. The infrared and visible light image fusion method comprises the following steps: inputting an infrared image and a visible light image into a network model comprising a generator and two content perception discriminators; the generator extracts texture details from a visible light image and extracts contrast information from an infrared image through a double-branch dense residual block; performing spatial feature enhancement on the extracted features by using a spatial attention block; splicing the enhanced features in channel dimensions, and inputting the spliced features into an interactive complementary mining engine to mine cross-modal space information; the fused image is transmitted into a parameter-locked double-branch dense residual block module again, and the fused image is optimized by using an adversarial training loss function; the source image information is kept to the maximum extent, and the fusion quality and the visual effect are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image fusion, and in particular to a method for fusion of infrared and visible light images based on cross-modal complementary learning. Background Art

[0002] In the field of image fusion, although the existing technology of generative adversarial networks can fuse infrared and visible light images, extract fused texture and contour features through the generator, and then compare the fused image with the labeled image through the discriminator, and generate the fused image through adversarial training to achieve Nash equilibrium, there are still many shortcomings. For example, the traditional method cannot fully retain the source image information when extracting features, and the convolution layer has limited ability to capture long-distance dependencies, resulting in poor fusion quality and visual effects; the training process of the generative adversarial network is unstable and prone to artifacts, noise and other problems, which affect the quality of the fused image and the performance of subsequent visual tasks.

[0003] To address these limitations, this paper proposes Fusion, a novel infrared and visible light image fusion method based on a dual-discriminator generative adversarial network with attention mechanism. Summary of the invention

[0004] The present application aims to solve at least one of the technical problems existing in the prior art. To this end, the present application proposes a method for fusion of infrared and visible light images based on cross-modal complementary learning, comprising the following steps: S1. Input infrared images and visible light images into a network model consisting of a generator and two content-aware discriminators; S2. The generator extracts texture details from the visible light image and contrast information from the infrared image through a dual-branch dense residual block; S3. Use spatial attention blocks to enhance the extracted features; S4. After concatenating the enhanced features in the channel dimension, input them into the interactive complementary mining engine to mine cross-modal spatial information; S5. The fused image is again passed into the parameter-locked dual-branch dense residual block module, and the fused image is optimized using the adversarial training loss function; S6. Train the network according to the defined generator loss function, combination loss and content loss. Use the trained network to fuse the infrared image and the visible light image to generate a synthetic image.

[0005] In addition, the infrared and visible light image fusion method based on cross-modal complementary learning according to an embodiment of the present application also has the following additional technical features: Preferably, the two content-aware discriminators are a visible light discriminator and an infrared discriminator, which are used to evaluate the authenticity of the fused image output by the generator.

[0006] Preferably, each module in the dual-branch dense residual block has the same structure and maintains independent network parameters.

[0007] Preferably, the spatial attention block enhances pixel correlations, overcoming the limitations of convolutional layers in capturing long-range dependencies.

[0008] Preferably, the interactive complementary mining engine uses a loss function to control the mined features so that they do not interfere with each other and maintain the unique characteristics of each modality.

[0009] Preferably, the loss function of the generator is composed of an adversarial loss and a combined loss, wherein the adversarial loss is defined based on the judgment of two content-aware discriminators on the features of the fused image.

[0010] Preferably, the combined loss is used to guide the training of the generator, balancing the performance of image fusion and visual tasks through hyperparameters.

[0011] Preferably, the content loss includes intensity loss and texture loss, the intensity loss measures the difference between the fused image and the source image at the pixel level, and the texture loss is used to enhance the fine-grained texture information of the fused image.

[0012] Preferably, the intensity loss adopts a maximum selection strategy to integrate the pixel intensity distributions of the infrared and visible images and constrain the brightness distribution of the fused image.

[0013] Preferably, the texture loss measures fine-grained texture information of the image through a gradient operator.

[0014] According to an infrared and visible light image fusion method based on cross-modal complementary learning in an embodiment of the present application, the beneficial effects are: 1. The generator uses texture details extracted from visible light images and contrast information extracted from infrared images to ensure that the source image information is retained to the maximum extent; 2. The lightweight attention mechanism enhances pixel correlation and overcomes the limitations of convolutional layers in capturing long-distance dependencies, thereby improving fusion quality and visual effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the implementation methods of the present application, the drawings required for use in the implementation methods will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0016] Figure 1 According to the embodiment of the present application Fusion network framework diagram; Figure 2 is a comparison chart of fusion results according to the embodiment of the present application and the most advanced infrared and visible light image fusion method; Figure 3 It is the features and loss functions learned by different modules according to the embodiment of the present application Schematic diagram of performance on the Roadscene dataset; Figure 4 This is a schematic diagram comparing the fusion results of infrared and visible light images in a power plant environment. DETAILED DESCRIPTION

[0017] In order to make the purpose, technical solutions and advantages of the implementation methods of this application clearer, the technical solutions in the implementation methods of this application will be clearly and completely described below in conjunction with the drawings in the implementation methods of this application. Obviously, the described implementation methods are part of the implementation methods of this application, not all of the implementation methods. Based on the implementation methods in this application, all other implementation methods obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0018] The following describes in detail a method for fusion of infrared and visible light images based on cross-modal complementary learning of the present invention through specific implementation methods.

[0019] one, Fusion's network framework Fusion's fusion algorithm framework is as follows Figure 1 As shown in Figure 1, it consists of a generator and two content-aware discriminators, namely visible light discriminator and infrared discriminator. Given an infrared image and visible light images As the input of the network model, synthetic images can be generated through feature extraction and feature fusion In the generator, a dual-branch dense residual block (DRC) is constructed, and each module in the dual-branch structure has the same structure but maintains independent network parameters.

[0020] After the infrared image and the visible light image are feature extracted by the dual-branch DRC structure, the spatial feature extraction in the image is enhanced by SA1 and SA2. The features enhanced by SA1 and SA2 spatial attention are spliced ​​in the channel dimension, and the spliced ​​feature maps are input into the interactive complementary mining engine (ICME). ICME mines the cross-modal spatial information in the spliced ​​feature maps and adopts the loss function To control the mined features to not interfere with each other and maintain the unique characteristics of each modality. The superposition method can make the dual-branch feature extraction more accurate and effective.

[0021] Subsequently, the fused image is passed again to the DRC module with locked parameters. Without updating the network parameters, the adversarial training loss function is used to make the fused image closer to the visible light and infrared images, thereby allowing the features output from the dual-branch DRC to be and Learn more relevant texture information and structural information of the fused image. By utilizing spatial attention blocks and ICME, the dual-path fusion features show different response distributions to cross-modal features, thereby improving the quality of the fusion output. This makes Fusion is able to focus more on target information in typical infrared images and texture in visible light images.

[0022] The BFFC module uses two content-aware discriminators to evaluate the authenticity of the fused image output by the generator. Infrared content-aware discriminator (ICDNet) and visible light content-aware discriminator (VCDNet) are used to verify , And the fusion features output from the parameter-locked DRC and Due to the rich texture details in visible light images, the mutual iterative training process of the generator and the discriminator enables accurate target recognition in different scenarios, especially in low light conditions and when the target is occluded, while ensuring high fusion quality. The infrared discriminator and the visible light discriminator play the role of distinguishing the source image from the generated fused image. The infrared discriminator and the visible light discriminator consist of three 3×3 convolutional layers in series and one linear layer. The activation function after all 3×3 convolutional layers is LReLU. Adversarial loss of infrared discriminator and visible light discriminator and is defined as follows: in and Represents the features of the fused image derived from the parameter-locked two-branch DRC. represents infrared discriminator, represents the visible light discriminator, and N represents the number of fused images.

[0023] The loss function of the generator By adversarial loss and the combined loss Combined composition, It can be defined as: in, Represents the first coefficient.

[0024] Adversarial Loss It can be defined as: 2. Loss Function We construct a combined loss To guide the training of the generator. Define it as: in, represents the second coefficient, represents the third coefficient, which are two hyperparameters that indicate the importance of each loss to maintain a balance between image fusion and visual tasks (such as target detection), ensuring the performance of visual tasks without reducing the performance of the fusion model. is the SSIM loss function.

[0025] The instability of generative adversarial networks during training often leads to artifacts and noisy or incomprehensible output. To address this issue, we introduce a content loss , which imposes constraints on the network to fully integrate the complementary information in the source images, such as salient objects in infrared images and texture details in visible images. This content loss not only enhances the visual quality of the fused output, but also facilitates the quantification of performance metrics. The content loss consists of two components, intensity loss and texture loss The specific description of content loss is defined as follows: in, Constrain the overall appearance strength of the fused image, while The forced fused image contains finer-grained texture detail information. represents the fourth coefficient, represents the fifth coefficient, and It is used to strike a balance between intensity loss and texture loss. Intensity loss measures the difference between the fused image and the source image at the pixel level. The intensity loss for infrared and visible light images can be defined as: in and are the height and width of the image, Represents the maximum number of selected elements. A maximum selection strategy is adopted to integrate the pixel intensity distributions of infrared and visible images, and the brightness distribution of the fused image is constrained accordingly. This ensures that the fused image retains rich texture details while achieving optimal brightness. In order to address the coarse-grained constraint of intensity loss, we introduce texture loss to enhance the fused image with finer-grained texture information. The texture loss is defined as: in Represents the gradient operator, which is used to measure the fine-grained texture information of an image. Refers to absolute operation. Assuming that the optimal texture of the fused image is the largest set of infrared and visible light image textures, the results show that the DRC-based fusion network can achieve the optimal grayscale distribution and retain rich detail information under the guidance of content loss. In other words, content loss can effectively ensure that our model achieves its first goal, which is to improve the visual quality and statistical evaluation indicators of the fused image.

[0026] 3. Experimental Verification To show Fusion's fusion performance uses Fusion compares its performance with seven state-of-the-art infrared and visible image fusion methods, including four deep fusion models, namely RFN-Nest, MFEIF, IFCNN, and SeAFusion, and three latest Transformer-based fusion models, namely SwinFusion, DATFusion, and YDRT. The source code of these seven methods is either public or provided by their authors. We randomly selected 20 pairs of images from the RoadScene and M3FD datasets for testing, which include labels such as people, cars, roads, buildings, plants, and lamps. Table 1 shows the average of the quantitative results of these 20 pairs of images fused by eight different methods. The results show that Fusion obtains the best fusion indexes of MI, SD, and SF in both datasets. The significant improvement of SD value indicates that the fused image has clear contours and the features of visible light image and infrared image are well integrated. The lack of attention mechanism and discriminator in SeAFusion limits its feature extraction ability, resulting in Fusion is weaker than Semantic Outline. Although YDTR and DATFuse adopt Transformer-based structures, they lack fine-grained feature extraction design, resulting in similar performance across five indicators. In contrast, Fusion showed excellent performance, achieving the highest MI values, which reflects enhanced texture detail and visual appeal. The best SF values ​​indicate improved grayscale variation, resulting in a clearer, higher-quality fusion, while the highest SD values ​​highlight a sharper, more contrast-rich result. Overall, Fusion outperforms other methods and can generate fused images with the best visual effects.

[0027] like Figure 2 As shown, Fusion is good at preserving texture and background information in images, providing excellent visual quality. In contrast, SwinFusion cannot reconstruct salient objects and backgrounds, while RFN-Nest causes background blur and removes detail information in images, resulting in unrecognizable license plates in the Roadscene dataset. Other methods rely too much on infrared information, resulting in the loss of salient objects. In forest and dense fog scenes, Fusion effectively integrates target contours in infrared images to enhance missing visible details while retaining complete pedestrian structure information.

[0028] As shown in Table 2, the average quantitative results of the fusion of these 20 image pairs by four different methods are shown. Due to the interactive use of spatial attention and The fused images reused in Fusion retain the inherent characteristics of the features extracted by the dual-path process of infrared and visible light modalities while incorporating rich cross-modal features. The effectiveness of Fusion, we can observe from Table 2: through the spatial attention module and DRC reuse fused image, and the integration of L added value method, the metrics of improved fused image are achieved in five key aspects. Fusion achieved the best performance.

[0029] like Figure 3 As shown, the results of four groups of experiments are given. Fusion achieves the best visual performance in both background and texture details. In addition, combined with multi-spatial attention, the image is fused into the parameter-locked DRC and added The visual quality is further improved, verifying the effectiveness of our designed modules and loss functions.

[0030] like Figure 4 The figure shows infrared and visible light images collected by a quadruped robot in a power plant. The comparison of this method with two other advanced methods, TarDAL and SeAFusion, shows that in high-dynamic scenes with blurred visible images, the use of Fusion's fused images can still reconstruct high-texture and clear images. Compared with the fused images of other methods, only Fusion successfully and clearly reconstructed the pipeline image.

[0031] Table 1: Fusion compares quantitative results with seven state-of-the-art methods

[0032] Table 2: ICMEFusion ablation experiment on Roadscene dataset

[0033] The above embodiments are only used to illustrate the specific embodiments of the present invention and are not limited thereto. For those skilled in the art, various similar deformations and changes can be made according to the concept of the present invention, and these deformations and changes should be regarded as the protection scope of the present invention.

Claims

1. A method for fusion of infrared and visible light images based on cross-modal complementary learning, characterized in that: The following steps are involved: S1. Input infrared images and visible light images into a network model consisting of a generator and two content-aware discriminators; S2. The generator extracts texture details from the visible light image and contrast information from the infrared image through a dual-branch dense residual block; S3. Use the spatial attention block to enhance the extracted features; S4. After concatenating the enhanced features in the channel dimension, input them into the interactive complementary mining engine to mine cross-modal spatial information; S5. The fused image is again passed into the parameter-locked dual-branch dense residual block module, and the fused image is optimized using the adversarial training loss function; S6. Train the network according to the defined generator loss function, combination loss and content loss.

2. The infrared and visible light image fusion method based on cross-modal complementary learning as claimed in claim 1, characterized in that: The two content-aware discriminators are a visible light discriminator and an infrared discriminator, which are used to evaluate the authenticity of the fused image output by the generator.

3. The infrared and visible light image fusion method based on cross-modal complementary learning as claimed in claim 1, characterized in that: Each module in the dual-branch dense residual block has the same structure and maintains independent network parameters.

4. The infrared and visible light image fusion method based on cross-modal complementary learning as claimed in claim 1, characterized in that: The spatial attention block enhances pixel correlations and overcomes the limitation of convolutional layers in capturing long-range dependencies.

5. The infrared and visible light image fusion method based on cross-modal complementary learning as claimed in claim 1, characterized in that: The interactive complementary mining engine uses a loss function to control the mined features to prevent them from interfering with each other and maintain the unique characteristics of each modality.

6. The infrared and visible light image fusion method based on cross-modal complementary learning as claimed in claim 1, characterized in that: The loss function of the generator is composed of an adversarial loss and a combination loss, where the adversarial loss is defined based on the judgment of two content-aware discriminators on the features of the fused image.

7. The infrared and visible light image fusion method based on cross-modal complementary learning as claimed in claim 1, characterized in that: The combined loss is used to guide the training of the generator, balancing the performance of image fusion and vision tasks through hyperparameters.

8. The infrared and visible light image fusion method based on cross-modal complementary learning as claimed in claim 1, characterized in that: The content loss includes intensity loss and texture loss, wherein the intensity loss measures the difference between the fused image and the source image at the pixel level, and the texture loss is used to enhance the fine-grained texture information of the fused image.

9. The infrared and visible light image fusion method based on cross-modal complementary learning as claimed in claim 8, characterized in that: The intensity loss adopts a maximum selection strategy to integrate the pixel intensity distributions of the infrared and visible images and constrain the brightness distribution of the fused image.

10. The infrared and visible light image fusion method based on cross-modal complementary learning as claimed in claim 8, characterized in that: The texture loss measures the fine-grained texture information of the image through a gradient operator.

Citation Information

Cited By

  • Infrared and visible light image fusion method and system based on instance perception

    CN121353096A

  • An instance perception-based infrared and visible light image fusion method and system

    CN121353096B