Method for fusing infrared light and visible light images

By adopting CNN network, gradient aggregation residual dense blocks and spatial channel attention mechanisms in the fusion of infrared light and visible light images, the shortcomings of existing methods in integrating texture features and suppressing interference are solved, and high-quality image fusion and semantic information enhancement are achieved.

WO2025103079A1PCT designated stage expired Publication Date: 2025-05-22CHONGQING UNIV OF TECH

Patent Information

Application Number
PCT/CN2024/126006
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-17
Filing Date
2024-10-21
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

The existing infrared light and visible light image fusion method ignores the needs of advanced visual tasks, is difficult to effectively integrate thick and fine texture features, and is disturbed by thermal radiation, affecting image quality.

Method used

Using an end-to-end fusion network based on CNN, combining gradient aggregation residual dense blocks and spatial channel attention mechanism, retaining features through split-transformation-merging and dense connections, amplifying useful information and suppressing interference from useless information.

Benefits of technology

Improve the accuracy of image fusion, enhance the semantic and spatial information of the fused image, promote the application of advanced visual tasks, and perform better than the most advanced methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024126006_22052025_PF_FP_ABST
    Figure CN2024126006_22052025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of image fusion. Disclosed is a method for fusing infrared light and visible light images. According to the present invention, design is carried out on the basis of a spatial channel attention mechanism and a gradient aggregation residual dense block, a strong texture and weak texture of a feature are retained by integrating Sobel and Laplacian operators while combining the advantages of ResNeXt and DenseNet, and channel information and spatial information of a feature map are refined by means of introducing a spatial and channel attention mechanism, thus the information capture capability of the feature map is increased, and refined spatial and channel feature maps are fused by using a pooled fusion block, to obtain a high-quality fusion feature.
Need to check novelty before this filing date? Find Prior Art

Description

A method for fusion of infrared and visible light images Technical Field

[0001] The present invention relates to the technical field of image fusion, and in particular to a method for fusing infrared light and visible light images. Background Art

[0002] [Corrected 07 / 11 / 2024 according to Rule 26] Image fusion is an important image enhancement technique that extracts meaningful information from different source images and fuses them into a new image. The fused image is typically robust, information-rich, and capable of representing more complex and detailed scenes. This not only reduces data redundancy but also facilitates the development of subsequent applications and decision-making. Fusion has been widely used as a preprocessing module for advanced vision tasks and holds great promise for applications in object detection, object tracking, and semantic segmentation.

[0003] Due to the practicality of infrared and visible light images, a number of image fusion techniques have emerged in recent years, which can be roughly divided into two categories: traditional methods and deep learning-based methods. Based on the different mathematical transformations, traditional methods can be further divided into methods based on multi-scale transformations, such as the discrete wavelet transform (DWT); methods based on representation learning, such as sparse representation (SR) and joint sparse representation (JSR); subspace-based methods, saliency-based methods, and hybrid models. Deep learning-based methods, on the other hand, can be divided into three categories based on the network architecture: models based on autoencoders (AEs), models based on convolutional neural networks (CNNs), and models based on generative adversarial networks (GANs).

[0004] Existing deep learning models are generally built on CNN or GAN networks. Although these methods have achieved good results in the field of image fusion, they emphasize the quality of image fusion while ignoring the needs of advanced visual tasks. In 2022, Tang et al. first proposed SeAFusion, an image fusion framework that combines advanced visual tasks. They did not make significant innovations in network architecture or learning paradigms, but they examined the image fusion task from a new perspective, namely using advanced visual tasks to drive image fusion. The emergence of SeAFusion has proposed new possibilities for image fusion.

[0005] In order to better promote the performance of infrared and visible light image fusion, the use of spatial and channel attention mechanisms can amplify useful information in the image and suppress the interference of harmful information by assigning different weights, while enriching the semantic and spatial information of the image to promote the application of subsequent high-level visual tasks. The NestFuse network is a pioneering work in the field of image fusion. For multi-scale deep feature fusion, they proposed a fusion strategy based on spatial and channel attention models. In the spatial attention model, the deep features are calculated by l1 norm and soft-max operator to obtain enhanced deep features. In the channel attention model, the channel fusion features are obtained by global pooling and soft-max operator, and finally the spatial features and channel features are added to obtain the final fusion feature reference:

[0006] Deficiencies of existing technology:

[0007] 1) Traditional methods use the same transformation to extract features from different source images. However, this operation fails to account for the differences in the source image features, potentially resulting in poorly expressed features. Furthermore, the complexity of the fusion strategy and its design in traditional methods limits performance. These manually designed fusion strategies are not only incapable of learning but also introduce artifacts into the fusion results.

[0008] 2) Before fusion of infrared and visible light images, it is necessary to extract useful feature information from them, such as texture and edges. However, existing network architectures cannot effectively extract detailed features of both coarse and fine granularity and are easily interfered with by thermal radiation, thus affecting the quality of the fused image. Further breakthroughs are needed to effectively integrate coarse and fine texture features while suppressing interference from useless information.

[0009] 3) Existing fusion methods tend to pursue better visual quality and higher evaluation metrics, rarely systematically considering whether the fused image can facilitate high-level visual tasks. While some studies have introduced perceptual losses at the feature level to constrain the fused and source images, perceptual losses are not effective in enhancing semantic information in the fused image. Furthermore, other researchers have used segmentation masks to guide the image fusion process, but these masks only segment prominent objects, which is limited in enhancing semantic information.

[0010] Therefore, a new solution to the above problems needs to be proposed. Summary of the Invention

[0011] The purpose of the present invention is to provide a method for fusing infrared light and visible light images to solve the technical problems raised in the background technology.

[0012] To achieve the above-mentioned object, the present invention provides the following technical solution: a method for fusing infrared and visible light images, characterized in that it comprises at least the following steps:

[0013] S1: Build a fusion network, using the end-to-end CNN fusion network as the basic framework. The end-to-end CNN network has powerful feature extraction and learning capabilities;

[0014] S2: Build a gradient aggregation residual dense block. The gradient aggregation residual dense block combines the advantages of ResNext and DenseNet networks. It retains shallow and deep features through split-conversion-merging and dense connections, and enhances the network's feature extraction capabilities.

[0015] S3: Build a spatial channel attention module, which is used to amplify useful information and suppress the interference of useless information, while promoting the development of semantic segmentation tasks and processing the deep features of the input;

[0016] S4: Build a separation network to fully enhance the semantic information of the fused image.

[0017] Preferably, the application of the fusion network in S1 at least includes the following steps:

[0018] By sending the infrared and visible light images to the feature extraction module respectively, the depth features of each are extracted;

[0019] The extracted deep features are then concatenated and fed into the spatial channel attention mechanism to further extract features and suppress the interference of useless information;

[0020] Finally, the fused image is generated through the feature reconstruction module.

[0021] Preferably, the branches in the gradient aggregated residual dense block in S2 perform a set of transformations, each on a low-dimensional embedding, whose outputs are aggregated by summing.

[0022] Preferably, the gradient aggregated residual dense block in S2 sets the cardinality of the aggregate transformation to 2, and is composed of at least one residual block and one residual dense block. Each gradient aggregated residual dense block contains three branches to improve the diversity of extracted features, so that it can fully utilize the deep features extracted by each convolutional layer in the block;

[0023] Since the source image contains rich texture details, the gradient aggregation residual dense block also integrates the Laplacian operator and the Sobel operator to retain more coarse textures and fine textures in the image.

[0024] Preferably, the spatial channel attention module in S3 includes at least multi-scale feature extraction, spatial channel attention mechanism and pooling fusion block.

[0025] Preferably, the application of the spatial channel attention module in S3 comprises at least the following steps:

[0026] Four convolutional blocks with different kernel sizes are used to capture multi-scale and multi-receptive field deep features, which are then fed into the spatial and channel attention mechanisms respectively;

[0027] Generate spatial and channel attention masks through self-attention functions to improve the network's information capture ability, thereby obtaining more accurate spatial and semantic information;

[0028] After the refinement of the attention branch, the channel feature map is upsampled to restore it to its original size. In addition, the present invention uses convolution operation to control the number of channels of the spatial feature map;

[0029] The generated spatial and channel attention feature maps are fused using the pooling fusion block to generate a fused feature map that satisfies both intra-class similarity and inter-class difference requirements.

[0030] Preferably, the application of the S4 separation network includes at least the following steps:

[0031] A real-time segmentation model, Bilateral Attention Decoder, is introduced to segment the fused image. The segmentation network outputs the segmentation result and auxiliary segmentation result.

[0032] The gap between the segmentation result and the semantic label reflects the richness of the semantic information contained in the fused image;

[0033] Exploit this gap to construct semantic loss;

[0034] Semantic loss is used to guide the training of the fusion network through back-propagation, forcing the fused image to contain more semantic information.

[0035] Compared with the prior art, the present invention has the following beneficial effects:

[0036] The present invention is designed based on the spatial channel attention mechanism and the gradient aggregation residual dense block. It combines the advantages of ResNeXt and DenseNet while integrating Sobel and Laplacian operators to retain the strong and weak textures of the features. By introducing the spatial and channel attention mechanism, the channel information and spatial information of the feature map are refined to improve the information capture ability of the feature map. A pooling fusion block is used to fuse the refined spatial and channel feature maps to obtain high-quality fused features. The experimental results of the present invention show that the proposed method performs better than the most advanced methods on the MSRS dataset, highlighting the potential future research direction to improve the accuracy of the fused image while promoting the development of advanced vision tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0038] Figure 1 shows the spatial attention model under the existing technology;

[0039] Figure 2 shows the channel attention model under the existing technology;

[0040] FIG3 is a schematic diagram of the overall network framework of SCGRFuse of the present invention;

[0041] FIG4 is a schematic diagram of a fusion network of the present invention;

[0042] FIG5 is a schematic diagram of a GRXDB module of the present invention;

[0043] FIG6 is a schematic diagram of a multi-scale spatial attention module of the present invention;

[0044] FIG7 is a schematic diagram showing a qualitative comparison of the SCGRFuse of the present invention and nine state-of-the-art methods on the MSRS dataset;

[0045] FIG8 is a schematic diagram of the process of generating channel attention and spatial attention masks according to the present invention;

[0046] FIG9 is a schematic diagram of the overall structure of the pooling fusion block of the present invention. DETAILED DESCRIPTION

[0047] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0048] Referring to Figures 1 and 2, we can see the spatial attention model and channel attention model under the existing technology.

[0049] Example 1:

[0050] This embodiment discloses a method for fusing infrared and visible light images based on a spatial channel attention mechanism and a gradient aggregation residual dense block.

[0051] SCGRFuse combines image fusion and semantic segmentation tasks. It first uses a fusion network to generate a fused image from the input infrared and visible light images. A content loss is used to guide the training of the fusion network. The resulting fused image is then fed into a semantic segmentation network. The semantic loss generated by the segmentation network integrates more semantic information into the fused image, thereby optimizing the overall fused image. Figure 3 shows the overall network architecture.

[0052] 1) Fusion Network

[0053] This method uses an end-to-end fusion network based on CNN as its basic framework. This end-to-end CNN network possesses powerful feature extraction and learning capabilities. The infrared and visible light images are fed into a feature extraction module to extract their respective deep features. These deep features are then concatenated and fed into a spatial channel attention mechanism to further extract features and suppress interference from useless information. Finally, a feature reconstruction module generates a fused image. The network structure is shown in Figure 4.

[0054] 2) Gradient Aggregation Residual Dense Block (GRXDB)

[0055] Gradient Aggregation Residual Dense Blocks (GRXDB) combine the strengths of ResNext and DenseNet networks, preserving both shallow and deep features through a split-transform-merge approach and dense connections. The branches within the GRXDB module perform a set of transformations, each on a low-dimensional embedding, whose outputs are aggregated by summation. The present invention sets the cardinality of the aggregation transformation to 2, with the module consisting of a residual block and a residual dense block. More specifically, each GRXDB contains three branches to increase the diversity of extracted features, enabling it to fully utilize the deep features extracted by each convolutional layer within the block. Because the source image contains rich texture details, GRXDB also integrates the Laplacian and Sobel operators to preserve more coarse and fine textures in the image. GRXDB significantly enhances the network's feature extraction capabilities. Its structure is shown in Figure 5.

[0056] 3) Spatial Channel Attention Module (SCAM)

[0057] To amplify useful information, suppress the interference of useless information, and promote the semantic segmentation task, this paper proposes a multi-scale spatial-channel attention mechanism to process input deep features. SCAM consists of three main components: multi-scale feature extraction, spatial-channel attention mechanism, and pooling fusion block. Because same-scale convolutions hinder the extraction of multi-scale information from different features, before calculating attention weights, to obtain a more comprehensive representation of the source image, this paper uses four convolution blocks with different kernel sizes to capture multi-scale and multi-receptive field deep features. These captured deep features are then fed into the spatial and channel attention mechanisms, respectively. To improve the network's information capture capability and thereby obtain more accurate spatial and semantic information, this paper uses a self-attention function to generate spatial and channel attention masks. After refinement by the attention branch, this paper upsamples the channel feature maps to restore them to their original size. Furthermore, this paper uses convolution operations to control the number of channels in the spatial feature maps. To generate feature maps that meet both intra-class similarity and inter-class difference requirements, this paper uses a pooling fusion block (PFB) to fuse the generated spatial and channel attention feature maps, resulting in a fused feature map that meets the requirements. The structure of the spatial channel attention module is shown in Figure 6.

[0058] 4) Segmentation Network

[0059] To fully enhance the semantic information of the fused image, this paper introduces a real-time segmentation model, the Bilateral Attention Decoder, to segment the fused image. The segmentation network outputs both the segmentation result and the auxiliary segmentation result. The gap between the segmentation result and the semantic label reflects the richness of the semantic information contained in the fused image. Therefore, this gap can be used to construct a semantic loss. Finally, the semantic loss is used to guide the training of the fusion network through backpropagation, forcing the fused image to contain more semantic information.

[0060] Example 2:

[0061] This embodiment further discloses an application of performance evaluation based on the above embodiment 1;

[0062] To comprehensively evaluate the proposed method, we conduct extensive quantitative and qualitative evaluations on the MSRS dataset. We also compare the proposed method with nine state-of-the-art methods, including one traditional method (GTF), one ae-based method (DenseFuse), two GAN-based methods (FusionGAN and GANMcC), four CNN-based methods (IFCNN, SDNet, U2Fusion, and SeAFusion), and one image decomposition-based method (DeFusion). The implementations of these nine methods are all public, and the parameters set in the present invention are consistent with those in the original paper. Among them, DenseFuse and IFCNN adopt element-wise addition and element-wise maximum fusion strategies to fuse deep features, respectively.

[0063] In terms of quantitative evaluation, this paper uses seven metrics, namely, EN, MI, VIF, SF, SD, SCD, and Qabf, to objectively evaluate the fusion effect. Furthermore, this paper uses IoU to quantify segmentation performance. Larger values ​​for these metrics indicate better fusion performance. The quantitative results of these seven statistical metrics for 361 image pairs are shown in Table 1. The segmentation results are shown in Table 2.

[0064] Table 1 quantitatively compares 7 indicators on 361 pairs of images from the MSRS dataset.

[0065] MethodENSDSFMISCD.VIFQabfGTF5.4719.568.481.670.760.510.40DenseFuse5.9423.576.032.651.250.690.37Fus ionGAN5.4417.074.421.870.980.440.14IFCNN6.2831.5010.762.821.380.770.60GANMcC5.9122.844.922.531.240. 590.25SDNet5.2517.358.671.650.990.450.38U2Fusion5.5627.719.241.961.260.550.42DeFusion6.4637.638.60 2.161.350.770.54SeAFusion6.6541.8411.114.041.690.940.67SCGRFuse(Ours)6.6842.7611.285.081.671.040.69

[0066] As shown in Table 1, the present invention demonstrates significant advantages in SD, SF, and MI. A higher SD indicates the highest contrast in the fused image. A higher SF value indicates a clearer fused image and better quality. A higher MI value indicates that the fused image conveys more information. Furthermore, the present invention exhibits the best VIF, demonstrating that the fused image is more consistent with the human visual system. The present invention also achieves the best Qabf, indicating that the fused image retains more edge information. Furthermore, the present invention exhibits the highest EN, indicating that the fused image contains the most information. The present invention trails SeAFusion only slightly in the SCD metric. Qualitative results of SCGRFuse on the MSRS dataset are shown in Figure 7. The present invention displays the fusion results for two scenes, primarily day and night, respectively. A red frame is used to magnify an area to illustrate the varying degrees of spectral contamination of texture details. Furthermore, the present invention uses a green frame to highlight the problem of useless information weakening the highlighted object. In daytime scenes, neither GTF nor FusionGAN can effectively preserve the texture details of visible light images, and other methods are inevitably affected by useless information. For nighttime scenes, we can see that all methods fuse the complementary information between infrared and visible light images to some extent, but most introduce useless information into the fused image, manifesting as a weakening of salient objects and contamination of texture and background details. Only our method and SeAFusion can preserve rich texture details and highlight objects, and our method's fused image has higher contrast than SeAFusion.

[0067] Table 2 Segmentation performance (miou) of visible light, infrared and fusion images on the MSRS dataset.

[0068]

[0069] The segmentation results are shown in Table 2. It can be seen that the present invention is generally in a leading position in all types of IoU and ranks first in MIoU. This is due to two advantages of the present invention. On the one hand, the present invention effectively integrates the complementary information of infrared and visible light images, which helps the segmentation model fully understand the imaging scene. On the other hand, SCGRFuse improves the information capture capability under the action of spatial and channel attention mechanisms. Under the guidance of semantic loss, it enhances spatial and semantic information, enabling the segmentation network to more accurately describe the imaging scene. Example 3:

[0070] This embodiment is used to disclose the specific operation process under the above embodiment.

[0071] The goal of image fusion is to preserve the strengths of the different images in the fused image. However, existing fusion methods are complex and neglect the impact of attention mechanisms on deep features. To address these issues, benfam proposes a method called SCGRFuse for fusion of infrared and visible light images. First, we construct a Gradient Aggregation Residual Dense Block (GRXDB), which combines the strengths of ResNeXt and DenseNet while integrating Sobel and Laplacian operators to preserve both strong and weak textures in features. We then introduce spatial and channel-wise attention mechanisms to refine the channel-wise and spatial information of feature maps, improving their information capture capabilities. A pooling fusion block is then used to fuse the refined spatial and channel-wise feature maps to produce high-quality fused features. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods on the MSRS dataset, highlighting potential future research directions to improve the accuracy of fused images and promote the development of advanced vision tasks.

[0072] Among them, the gradient aggregation residual dense block (GRXDB) and the spatial channel attention module (SCAM) are the key technologies of this invention.

[0073] 1) Gradient Aggregation Residual Dense Block

[0074] The overall operation flow of the Gradient Aggregation Residual Dense Block (as shown in Figure 5) is as follows: Our GRXDB consists of three branches. The main branch deploys three 3x3 convolutional blocks and two 1x1 convolutional blocks. To more fully utilize the various convolutional layers for feature extraction, we introduce dense connections into the main branch and use 1x1 convolutional blocks to balance the differences between channel dimensions. Furthermore, the main branch also incorporates the Laplacian operator to further extract weak texture information from features. Furthermore, in addition to two 3x3 convolutional blocks and one 1x1 convolutional block, the residual branch also integrates the Sobel operator to preserve strong texture information from features. The third branch, referred to as the residual branch, performs no processing and preserves input feature information. Finally, the outputs of the main, residual, and residual branches are summed via element-wise addition to integrate deep features. It is worth noting that the Relu activation function is equivalent to directly discarding negative activations. While this approach may be effective for classification tasks, it is not suitable for image fusion tasks. To better meet the requirements of image fusion, our GRXDB activation function is set to Leaky Relu, which preserves negative activation information.

[0075] 2) Spatial Channel Attention Module

[0076] The operation flow of the spatial channel attention module (as shown in Figure 6) is as follows: This invention uses four convolutional blocks with different kernel sizes to capture multi-scale and multi-receptive field deep features. The convolutional block sizes are 1x1, 3x3, 5x5, and 7x7, respectively. The activation function for the convolutional blocks is still Leaky ReLU. The obtained deep features are then concatenated channel-wise. Finally, a 1x1 convolutional block with a different number of channels is used to send the deep features to the spatial and channel attention layers. In the attention mechanism, generating the attention masks Ms and Mc is the most important step, improving the network's information capture capability and thus obtaining more accurate spatial and semantic information. This invention uses a self-attention function to generate spatial and channel attention masks, the main process of which is shown in Figure 8. To generate the channel attention mask, this invention first uses global average pooling to average all pixels in each channel map, resulting in a new 1x1 channel map. The feature map values ​​are then adjusted using the Tanh activation function and a 1x1 convolutional block. Notably, the calculated value of the Tanh activation function is divided by 2 and added with 0.5 to map it to the range [0, 1]. Finally, the present invention multiplies the calculated channel map with the input feature map to generate the output of the channel attention branch. Similarly, to obtain the spatial attention mask, the present invention first uses a 1x1 convolution block to reduce the number of channels in the input feature map. Then, max pooling and average pooling are used to generate two feature maps. Both feature maps have only one channel and are the same size as the input feature map. Next, the two feature maps are concatenated and the number of channels is reduced to 1 using a 1x1 convolution block. The activation function of the convolution block is still tanh, consistent with the channel attention mechanism. Finally, the generated spatial attention mask is multiplied with the input feature map to obtain the final output of the spatial attention branch. After refinement by the attention branch, the present invention uses an upsampling operation to restore the channel feature map to its original size. In addition, convolution operations are used to control the number of channels in the spatial feature map. To generate a feature map that meets both intra-class similarity and inter-class difference requirements, the present invention uses a Pooling Fusion Block (PFB) to fuse the generated spatial and channel attention feature maps. The structure of the PFB is shown in Figure 9. The present invention uses 3x3 average pooling to smooth the upsampled channel feature map and uses reflection padding to construct the true boundary and reduce boundary artifacts. Finally, it is spliced ​​with the spatial feature map and the boundary is corrected using a 1x1 convolution block and a leaky ReLU activation function to generate a fused feature map that meets the requirements.

[0077] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

Claims

1. A method for fusing infrared light and visible light images, characterized in that: At least the following steps are included: S1: Build a fusion network, using CNN's end-to-end fusion network as the basic framework. The end-to-end CNN network has powerful feature extraction and learning capabilities; S2: Build a gradient aggregation residual dense block, which combines the advantages of ResNext and DenseNet networks, retains shallow and deep features through split-conversion-merge and dense connection, and enhances the feature extraction ability of the network; S3: Build a spatial channel attention module, which is used to amplify useful information, suppress the interference of useless information, and promote the development of semantic segmentation tasks to process the deep features of the input; S4: Build a separation network to fully enhance the semantic information of the fused image.

2. The infrared light and visible light image fusion method according to claim 1, characterized in that: The application of the fusion network in S1 at least includes the following steps: By sending the infrared light and visible light images to the feature extraction module respectively, the respective depth features are extracted; The extracted deep features are then concatenated and sent to the spatial channel attention mechanism to further extract features and suppress the interference of useless information; Finally, the fused image is generated through the feature reconstruction module.

3. The infrared light and visible light image fusion method according to claim 1, characterized in that: The branches in the gradient aggregation residual dense block in S2 perform a set of transformations, each on a low-dimensional embedding, whose outputs are aggregated by summing.

4. The infrared light and visible light image fusion method according to claim 3, characterized in that: The gradient aggregated residual dense block in S2 sets the cardinality of the aggregated transformation to 2, and is composed of at least one residual block and one residual dense block. Each gradient aggregated residual dense block contains three branches to improve the diversity of extracted features, so that it can make full use of the deep features extracted by each convolutional layer in the block; Since the source image contains rich texture details, the gradient aggregation residual dense block also integrates the Laplacian operator and the Sobel operator to retain more coarse textures and fine textures in the image.

5. The infrared light and visible light image fusion method according to claim 1, characterized in that: The spatial channel attention module in S3 at least includes multi-scale feature extraction, spatial channel attention mechanism and pooling fusion block.

6. The method for fusing infrared light and visible light images according to claim 5, characterized in that: The application of the spatial channel attention module in S3 at least comprises the following steps: Four convolution blocks with different kernel sizes are used to capture multi-scale and multi-receptive field deep features, which are then fed into the spatial and channel attention mechanisms respectively; Generate spatial and channel attention masks through self-attention functions to improve the network's information capture ability, thereby obtaining more accurate spatial and semantic information; After the refinement of the attention branch, the channel feature map is upsampled to restore it to its original size. In addition, the present invention uses a convolution operation to control the number of channels of the spatial feature map; The generated spatial and channel attention feature maps are fused using the pooling fusion block to generate a fused feature map that satisfies both the intra-class similarity and inter-class difference requirements.

7. The infrared light and visible light image fusion method according to claim 1, characterized in that: The application of the S4 separation network includes at least the following steps: A real-time segmentation model Bilateral attention decoder is introduced to segment the fused image, and the segmentation network outputs the segmentation result and auxiliary segmentation result; The gap between the segmentation result and the semantic label reflects the richness of the semantic information contained in the fused image; This gap is used to construct semantic loss; Semantic loss is used to guide the training of the fusion network through back-propagation, forcing the fused image to contain more semantic information.

Citation Information

Patent Citations

  • Visible light and infrared image fusion enhancement method and system in low-light environment

    CN115063329A

  • Substation unmanned aerial vehicle inspection method based on infrared and visible light image fusion

    CN115471723A

  • Infrared light and visible light image fusion method

    CN117576543A

  • Progressive image fusion

    US11113802B1

Cited By

  • Method and device for multi-scale fusion of hyperspectral image and multispectral image

    CN120259100A

  • A method and device for multi-scale fusion of hyperspectral images and multispectral images

    CN120259100B

  • Isovariant consistency image fusion method and system based on task driving

    CN120339092A

  • Multispectral and visible light remote sensing image fusion segmentation method, system and medium

    CN120472333A

  • Photovoltaic hot spot detection method and system based on visible light and thermal infrared image fusion

    CN120495302A