Infrared and visible light image fusion method based on interactive complementary mining engine
By using an interactive complementary mining engine and spatial attention module in infrared and visible image fusion, the problem of poor fusion image quality in the prior art is solved, and a higher quality image fusion effect is achieved.
Patent Information
- Application Number
- CN202510181903.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is prone to overfitting when fusing infrared and visible light images, resulting in poor quality of the fusion image and failing to fully consider the characteristics of different modal images.
Using an interactive complementary mining engine method, the features of infrared and visible light images are extracted through the dual-branch dense residual connection module and the spatial attention module, and the interactive complementary mining engine and loss function are used to perform feature fusion to generate high-quality fusion images.
Through this method, the visual fidelity and contrast of the fusion image are greatly improved, which can more effectively retain the texture characteristics and structural information of infrared and visible images, and generate clearer and richer contrast fusion images.
Smart Images

Figure CN120107083A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to an infrared and visible light image fusion method based on an interactive complementary mining engine. Background Art
[0002] The fusion of infrared and visible light images has been a hot topic in the field of image processing in recent years. Infrared images and visible light images can capture different features of objects. Infrared images can be used to detect objects well in poor lighting conditions. However, due to their low spatial resolution, infrared images provide limited scene texture details. In contrast, visible light images have higher spatial resolution, but they become blurred at night or in bad weather conditions. The fusion of infrared and visible light images can generate images that contain more prominent objects and clearer backgrounds.
[0003] The convolutional neural network (CNN) in the related art effectively fuses the content and texture information of infrared and visible light images. The network is trained using blurred background and foreground images and a binary weight map is obtained. During the test, the original image is combined with the weight map to obtain a fused multi-focus image. Therefore, researchers began to integrate CNN-based methods into infrared and visible light image fusion. In the initial stage, most researchers tried to introduce CNN into traditional methods to inject rich semantic information into the fused image. Related studies use the VGG19 network to further process the details after multi-scale decomposition, thereby retaining rich texture information in the fused image. It has been found that the features extracted by CNN can reflect the proportion of the original image in the fusion process to a certain extent. Therefore, the down-sampling sequence of the convolution weight map is used as the fusion ratio map of the down-sampling sequences of the two branches, avoiding the artificial design of the fusion strategy.
[0004] There are also related studies that use zero-phase component analysis and L1 norm to obtain a weight map that reflects the proportion of the original image, overcoming the problem of information loss when the original image is downsampled. These methods use the network's powerful feature extraction capabilities to retain rich fine-grained features in the fused image. The related technology uses two channels to input the original image into two CNNs, and the obtained feature map is cascaded into the convolution layer to obtain the fused image. Although this method can complete image fusion through a simple fusion structure, it is prone to overfitting during training, resulting in poor quality of the fused image. In addition, based on the dense residual network, the related technology proposes a deep dense residual network that can increase the amount of information in the fused image. An attention module is also embedded in the network, which can selectively provide more brightness and gradient information for the fused image. The end-to-end fusion model based on CNN connects the source image at the input stage or connects the deep features at the fusion layer. However, these methods generally use simple concatenation operations and fail to fully consider different modal images. GAN consists of a generator and a discriminator. The generator extracts and fuses the texture and contour features of infrared and visible light images. Then, the discriminator compares the fused image with the labeled image. Through continuous iterative adversarial training, it is possible to reach a Nash equilibrium with the generator, thereby generating an image with good fusion effect. In order to solve these defects, this patent proposes an infrared and visible light image fusion method based on an interactive complementary mining engine. Summary of the invention
[0005] In order to make up for the above shortcomings, the present invention provides an infrared and visible light image fusion method based on an interactive complementary mining engine, aiming to improve the problem in the related art that infrared and visible light images cannot be fused into high-quality images.
[0006] The present invention is achieved in that: The present invention provides an infrared and visible light image fusion method based on an interactive complementary mining engine, comprising the following steps: S1: The given infrared image and visible light images As input to the network model; S2: In the generator, a dual-branch dense residual connection module is constructed to extract infrared images and visible light images Infrared signatures Visible light characteristics ; S3: The focused features are obtained by element-wise multiplication with the attention map output by the SA1 spatial attention module , The focused features are obtained by element-wise multiplication with the attention map output by the SA2 spatial attention module. ; S4: and After splicing in the channel dimension, the input is fed into the interactive complementary mining engine to obtain two different feature maps SAM3 and SAM4. S5: Attention map SAM3 and Perform element-wise multiplication to get ; Attention map SAM4 and Doing element-wise multiplication gives ; S6: Finally and Perform channel dimension stitching and fusion to generate images .
[0007] In one embodiment of the present invention, the generated image after fusion in the above step S6 is The double-branch dense residual connection module is passed again, using the loss function Bringing visible light and infrared images closer to fused images , allowing the dual-path dense residual connection module to output features with richer fine-grained texture features and more complete structural information and .
[0008] In one embodiment of the present invention, the loss function Control SAM3 and SAM4 contain different modal information.
[0009] In one embodiment of the present invention, the loss function Adopt gradient smoothing, MSE loss function which is sensitive to large errors.
[0010] In one embodiment of the present invention, each module in the dual-branch structure has the same structure, but maintains independent network parameters.
[0011] The beneficial effect of the present invention is that the infrared and visible light image fusion method based on the interactive complementary mining engine obtained by the present invention through the above design extracts the texture features, contour shape features and structural features of the infrared image and the visible light image by combining the spatial attention module (SA1, SA2) and the interactive complementary mining engine (ICME). The visual fidelity of the fused image is fully enhanced. Compared with other advanced infrared and visible light image fusion algorithms, the various image quality evaluation indicators after the algorithm of this patent is fused are greatly improved.
[0012] The present invention utilizes the interactive use of spatial attention blocks and ICME to make the dual-path fusion features show different response distributions to cross-modal features, thereby improving the quality of the fusion output. This enables ICMEFusion to pay more attention to typical infrared targets and visible textures.
[0013] ICME guides the model to locate and focus on valid information in different ways. It forces the texture information of different modal areas to be inconsistent, and needs to be biased towards the information of the original modality of the source image, encouraging the model to mine complementary features from visible light and infrared images. The ICME proposed in this patent aims to effectively focus on the spatial dimension information of the input bimodal fusion features. ICME imposes constraints on cross-modal features to ensure that by introducing a specially designed loss function , so that the fused features of the two modalities will not interfere with each other while retaining their unique characteristics. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0015] Figure 1 It is a framework diagram of an infrared and visible light image fusion method based on an interactive complementary mining engine provided by an embodiment of the present invention; Figure 2 The dual discriminator and ICME provided by the embodiment of the present invention and Ablation experiment diagram of loss function; Figure 3 This is a diagram of the fusion result of infrared and visible light images of the power plant environment provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0016] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0017] Example See also Figure 1The present invention provides an infrared and visible light image fusion method based on an interactive complementary mining engine, which includes a generator and two content-aware discriminators, a visible light discriminator and an infrared discriminator. Given an infrared image and visible light images As the input of the network model, a fused image can be generated through feature extraction, integration and reconstruction In the generator, a dual-branch dense residual block (DRC) is constructed, and each module in the dual-branch structure has the same structure but maintains independent network parameters.
[0018] The visible light and infrared features after DRC extraction are: and Then The focused features are obtained by element-wise multiplication with the attention map output by the SA1 spatial attention module , similarly, The focused features are obtained by element-wise multiplication with the attention map output by the SA2 spatial attention module. .
[0019] Afterwards, and After splicing the channel dimension, the images are fed into the Interactive Complementary Mining Engine (ICME). ICME effectively focuses on the spatial dimension of the image and obtains two different feature maps: one is SAM3 and the other is SAM4.
[0020] At the same time, this invention patent designs an interactive spatial attention loss function ,pass To control the different modal information contained in SAM3 and SAM4, The MSE loss function is used. The MSE loss function is simple and easy to use, has a smooth gradient, is sensitive to large errors, and can effectively reduce the overall prediction bias in regression problems.
[0021] W is the width of the feature map, and H is the height of the feature map. and Represents the spatial attention map corresponding to the pixel position (i, j). The attention map SAM3 and Perform element-by-element multiplication; the attention map SAM4 and Do element-by-element multiplication and finally get and .
[0022] Afterwards, and The channel dimension is stitched. The stitched feature map is passed through a series of fusion layers to generate an infrared and visible light fusion image. The fused image is passed to the DRC module again. Without updating the network parameters, the loss function is used to make the visible light and infrared images closer to the fusion image. , thus allowing the features output from the dual-path DRC module to and It has richer fine-grained texture features and more complete structural information.
[0023] Experimental verification: The experimental data set of this patent uses M 3 FD and RoadScene datasets are two mainstream datasets for infrared and visible light image fusion. In the experiment, 4411 pairs of infrared and visible light images were collected from the two datasets and randomly divided into three datasets for training, validation, and testing. They contain 3094, 1297, and 20 pairs of infrared and visible light images, respectively.
[0024] M 3 FD dataset: This dataset contains 4200 pairs of registered images captured in various environments, lighting conditions, seasons, and weather conditions. We use independent scenes and the resolution of the image pairs is up to 640×480 pixels.
[0025] RoadScene dataset: Contains 221 pairs of registered images with a resolution of 563×459 pixels, captured using real-world cameras, which are selected from the FLIR dataset.
[0026] Power Plant Dataset: We apply a quadruped robot (OptrisPI640) equipped with visible and infrared cameras in a power plant to capture infrared and visible images in high-dynamic scenes. The dataset consists of 5047 pairs of infrared and visible images, each with a resolution of 640×480 pixels.
[0027] Experimental implementation details: In ICMEFusion, the generator and dual discriminators are trained iteratively. In order to maintain the training stability of the two discriminators, the training time between the dual content-aware discriminators adopts a 1:1 strategy. The learning rate is fixed at 1×10 −3 , with a batchsize of 4. Both the generator and the discriminator use the Adam optimizer to optimize the deep model with 10 epochs. Hyperparameters are used to maintain the balance between the various loss components and establish the preference of the generator during training. In practice, to ensure balanced training, we empirically assign initial values to the hyperparameters, aiming to make the size of the loss components equal. These values are then refined through a combination of qualitative and quantitative experimental results.
[0028] Information loss is a catastrophic problem in image fusion tasks. Therefore, except for the 1×1 convolutional layer, the padding in our fusion network is set to “same” and the stride is set to “1”. These settings ensure that the fusion layer does not introduce any downsampling and the size of the fused image remains consistent with the source image. The visible light discriminator and the infrared discriminator are designed to distinguish the output features of the dual-branch DRC and the features of the fused image generated by the parameter-locked dual-branch DRC. There is a conflict between the guidance of the visible light discriminator to the generator to generate the fused image and the guidance of the infrared discriminator to the generator to generate the fused image. In our work, we consider the adversarial relationship between the generator and the dual discriminator. Otherwise, as the training proceeds, the strength of one discriminator will eventually lead to the inefficiency of the other discriminator. This balance is achieved by the design of the network structure and training strategy. The visible light discriminator and the infrared discriminator have the same structure, and their settings are simpler than the generator structure. The stride of all convolutional layers is set to “2”. Five indicators are used to evaluate the fusion performance: information entropy (EN), mutual information (MI), peak spatial frequency (SF), standard deviation (SD), and visual fidelity (VIF).
[0029] To demonstrate the performance of ICMEFusion, we compare the performance of our method with seven state-of-the-art infrared and visible image fusion methods, including four deep fusion models, namely RFN-Nest, MFEIF, IFCNN, and SeAFusion, and three state-of-the-art Transformer-based fusion models, namely SwinFusion, DATFusion, and YDTR. The source codes of these seven methods are either publicly available or provided by their authors. We use the source code from RoadScene and M 3 20 pairs of images are randomly selected from the FD dataset for testing, which include target objects such as people, cars, roads, buildings, plants, and street lights.
[0030] Quantitative comparisons are performed with seven state-of-the-art methods. The average quantitative test results of 20 randomly selected image pairs are shown in Table 1, and the best values are marked in bold and underlined. ICMEFusion obtains the best image fusion metrics of mutual information, information entropy, and spatial frequency in both datasets. The significant improvement in the spatial frequency value indicates that the fused image has clear contours and the features of the visible light image are well fused with the infrared image. The lack of attention mechanism and discriminator in SeAFusion limits its feature extraction ability, resulting in weaker semantic contours compared with ICMEFusion. Although YDTR and DATFuse adopt a Transformer-based structure, they lack a fine-grained feature extraction design, resulting in similar performance. Across five metrics. In contrast, ICMEFusion shows superior performance and achieves the highest mutual information value, which reflects the enhanced texture details and visual appeal. The best standard deviation value indicates that the grayscale variation of the fused image is improved, and the highest standard deviation value highlights the clearer and more contrast-rich results.
[0031] As shown in Table 2, the average quantitative results of the fusion of these 20 pairs of images by two different methods are shown. Due to the interactive use of the ICME engine and spatial attention, the features extracted by the dual-path process of infrared and visible light modalities retain their inherent characteristics while incorporating rich cross-modal features. To verify the effectiveness of ICMEFusion, we can observe from Table 2 that ICMEFusion achieves the best performance by using the spatial attention module and ICME engine, as well as the loss function The fused image achieves the best results in five key indicators. ICMEFusion in this paper achieves the best performance.
[0032] like Figure 2 As shown in Figure 2, the results of four sets of experiments are given. ICMEFusion achieves the best visual performance in both background and texture details. In addition, by combining multi-space attention and adding loss function The visual quality of the fused image is further improved, verifying the effectiveness of the modules and loss functions we designed.
[0033] like Figure 3 The figure shows the effect of fusion of infrared and visible light images collected by a quadruped robot in a power plant. The visualization of the experimental results shows that in high-dynamic scenes with blurred visible images, the fused images using ICMEFusion can still reconstruct high-texture and clear multimodal fusion images. The multimodal complementary information in the fused image is obvious, and the infrared contour and visible light texture information are rich, achieving an excellent fusion effect.
[0034] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. The infrared and visible light image fusion method based on interactive complementary mining engine is characterized by: The steps include: S1: The given infrared image and visible light images As input to the network model; S2: In the generator, a dual-branch dense residual connection module is constructed to extract infrared images and visible light images Infrared signatures Visible light characteristics ; S3: The focused features are obtained by element-wise multiplication with the attention map output by the SA1 spatial attention module , The focused features are obtained by element-wise multiplication with the attention map output by the SA2 spatial attention module. ; S4: and After splicing in the channel dimension, the input is fed into the interactive complementary mining engine to obtain two different feature maps SAM3 and SAM4. S5: Attention map SAM3 and Perform element-wise multiplication to get ; Attention map SAM4 and Doing element-wise multiplication gives ; S6: Finally and Perform channel dimension stitching and fusion to generate images .
2. The infrared and visible light image fusion method based on the interactive complementary mining engine according to claim 1 is characterized in that: The generated image after fusion in the above step S6 The double-branch dense residual connection module is passed again, using the loss function Bringing visible light and infrared images closer to fused images , allowing the dual-path dense residual connection module to output features with richer fine-grained texture features and more complete structural information and .
3. The infrared and visible light image fusion method based on interactive complementary mining engine according to claim 2 is characterized in that: The loss function Control SAM3 and SAM4 contain different modal information.
4. The infrared and visible light image fusion method based on interactive complementary mining engine according to claim 3 is characterized in that: The loss function Adopt gradient smoothing, MSE loss function which is sensitive to large errors.
5. The infrared and visible light image fusion method based on interactive complementary mining engine according to claim 1, characterized in that: Each module in the dual-branch structure has the same structure but maintains its own independent network parameters.
Citation Information
Cited By
Infrared-visible light image fusion method based on pseudo twin network
CN121121382A
An Infrared-Visible Image Fusion Method Based on Pseudo-Twin Networks
CN121121382B