Image fusion method and system based on depth adaptive channel-space attention
By introducing a deep adaptive channel-spatial attention mechanism into the image fusion technology, the problem of unsatisfactory image fusion effect under extreme lighting conditions is solved, and a higher quality and robust image fusion effect is achieved.
Patent Information
- Application Number
- CN202510056920.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-16
AI Technical Summary
Existing infrared and visible image fusion technologies are difficult to effectively fuse images under extreme lighting conditions, resulting in unsatisfactory fusion effects and loss of important information, especially in low-light environments.
The image fusion method based on depth adaptive channel-spatial attention is adopted, and the initial features of infrared and visible light images are extracted through a dual-branch interactive encoder, and the weighted fusion is used to use the depth adaptive score module and the adaptive channel-spatial attention module for weighted fusion, and the fusion weight is dynamically adjusted to adapt to different lighting conditions.
It improves the quality and robustness of image fusion, enhances the processing capability of infrared images in low-light environments, and ensures better fusion effect under different lighting conditions.
Smart Images

Figure CN120013776A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing technology, and in particular relates to an image fusion method and system based on deep adaptive channel-spatial attention. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] In the current intelligent security, autonomous driving and other fields, infrared and visible light image fusion technology has a wide range of application needs. Infrared sensors capture images through thermal radiation and are not affected by weather and lighting conditions. However, infrared images often lack background details and structural textures. On the other hand, images captured by visible light sensors can provide rich color details and texture information, but are easily affected by lighting conditions, which makes it challenging to obtain high-quality images in adverse environments. Since the imaging mechanisms of infrared and visible light images are obviously complementary, designing algorithms that can fuse these two imaging modes can not only generate images that conform to human visual perception, but also retain key information in the environment, thereby improving the accuracy and robustness of visual tasks. Through the fused image, the outline of the target object is clearer, not only can infrared invisible objects be detected in dark or complex environments, but also the recognition accuracy can be enhanced through the details of visible light. In addition, in scenes with dense crowds, insufficient light or obstructions, the fused image is more conducive to pedestrian identification and improves the recognition ability of public safety and monitoring systems.
[0004] The main challenge of fusing infrared and visible light images is how to effectively minimize the noise interference of each mode and how to cleverly integrate complementary features to improve the quality of the generated image. To solve this problem, the algorithms proposed so far can be roughly divided into two categories: traditional methods and deep learning-based methods.
[0005] Traditional infrared and visible light image fusion methods can be further divided into several categories, including multi-scale transform-based methods, sparse representation methods, saliency methods, subspace techniques, and hybrid methods. Usually, these traditional methods extract features from images through specific transformations, manually design fusion rules to integrate these features, and then reconstruct the image through inverse transformation. However, traditional methods have inherent limitations. These methods rely on complex manually designed fusion rules, which are less flexible and difficult to adapt to the characteristic differences of images of different modalities. Therefore, such rules are difficult to be widely applied when dealing with specific tasks.
[0006] In contrast, deep learning methods have shown significant advantages in overcoming the limitations of traditional methods. These methods can automatically learn the hierarchical representation and fusion rules of data through the adaptability of neural networks. Deep learning methods have performed outstandingly in solving the defects of traditional methods and improving the performance of various applications. Its fusion algorithm automatically completes feature extraction, feature fusion and image reconstruction through neural networks, without the need to manually design complex rules, so it has higher versatility in different fusion scenarios. Image fusion technology based on deep learning covers a variety of types, including algorithms based on convolutional neural networks (CNN), generative adversarial networks (GAN), autoencoders, and Transformers.
[0007] Although deep learning methods have achieved remarkable results in image fusion, they usually ignore the specific problems faced by infrared and visible light images under different lighting conditions. Existing studies often assume that the input images have the same quality, but this assumption does not always hold true in real applications, especially under extreme lighting conditions. This may lead to suboptimal fusion results and loss of important information, especially for infrared images taken in low-light environments. This gap needs to be bridged by more sophisticated methods to adapt to different image modes and environmental conditions to ensure better fusion performance. Summary of the invention
[0008] In order to overcome the shortcomings of the above-mentioned prior art, the present invention provides an image fusion method and system based on deep adaptive channel-spatial attention, which interactively processes the input features of the two modal images and dynamically adjusts the fusion weights according to different lighting conditions. Specifically, a novel interactive extraction structure is introduced in the feature extraction stage of infrared light and visible light, so that more complementary information can be captured. The quality of the input features is evaluated using a deep adaptive fusion module, and weighted fusion is performed through a dual attention mechanism of channels and space, thereby achieving a better fusion effect.
[0009] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:
[0010] A first aspect of the present invention provides an image fusion method based on deep adaptive channel-spatial attention;
[0011] Image fusion method based on deep adaptive channel-spatial attention, including:
[0012] Acquire infrared light images and visible light images;
[0013] Input the acquired infrared image and visible light image into a dual-branch interactive encoder to extract initial features of the infrared image and the visible light image;
[0014] The weights of the initial features of the infrared image and the visible light image are obtained by using the deep adaptive score module to generate weighted features; the weighted features are merged and input into the adaptive channel-spatial attention module to realize the fusion of infrared and visible light features;
[0015] The fused features are input into the decoder for dimensionality reduction and reconstruction, and the fused results are mapped back to the original representation.
[0016] As a further technical solution, the process of inputting the acquired infrared image and visible light image into the dual-branch interactive encoder to extract initial features of the infrared image and the visible light image includes:
[0017] First, the infrared image and the visible light image are processed by the 1x1 convolutional layer of the dual-branch interactive encoder to reduce the modal difference between the infrared image and the visible light image.
[0018] Secondly, the 3x3 convolutional layer and SE layer of the dual-branch interactive encoder are used to extract the initial features of the infrared image and visible light image, and the CBAEM module is used to gradually integrate the complementary information.
[0019] As a further technical solution, the process of gradually integrating complementary information using the CBAEM module is as follows: the extracted initial features of the infrared image and the visible light image are input into the CBAEM module to complete channel-level feature enhancement and spatial-level enhancement;
[0020] The channel-level feature enhancement is as follows: the initial features of the extracted infrared image and visible light image are processed using global average pooling and global maximum pooling to extract the global features of each channel; the global features are processed through a two-layer convolutional network, and the weight of each channel is generated through a Sigmoid function; the channel weight is multiplied by the original input feature to complete the channel-level feature enhancement;
[0021] The spatial level enhancement is as follows: perform average pooling and maximum pooling on the channel dimension of the feature map to generate an average feature map and a maximum response map; concatenate the average feature map and the maximum response map along the channel dimension, and use convolution operations to generate a spatial weight map; multiply the spatial weight map with the input feature to complete the spatial level feature enhancement.
[0022] As a further technical solution, the process of using the depth adaptive score module to obtain the weights of the initial features of the infrared image and the visible light image and generate weighted features is as follows:
[0023] Extract the maximum channel, average channel, and standard deviation channel of the input features and connect them;
[0024] Integrate feature information through multi-scale convolutional layers; perform adaptive maximum pooling and adaptive average pooling on the convolution results to optimize feature expression from different angles;
[0025] The deep adaptive score is obtained through the Sigmoid function for dynamic weighted features;
[0026] The formula is as follows:
[0027] f1=Concat(Max(F input ),Std(F input ),Mean(F input ));
[0028] f2=Add(3×3conv(f1),5×5conv(f1),7×7conv(f1));
[0029] S=Sigmoid(Add(AdaptiveMaxPool2d(f2),AdaptiveAveragePooli2d(f2)));
[0030] Among them, Conv represents the convolution layer, Concat represents the connection layer, Sigmoid represents the Sigmoid activation function, Add represents the element addition operation, AdaptiveMaxPool2d represents the adaptive maximum pooling operation, and AdaptiveAveragePool2d represents the adaptive average pooling operation; f1 represents the feature vector after dimensionality reduction, f2 represents the feature vector of f1 after integration through 3×3, 5×5 and 7×7 multi-scale convolution layers, and S represents the weighted score of the input feature.
[0031] As a further technical solution, the process of implementing infrared and visible light feature fusion in the adaptive channel-spatial attention module is as follows:
[0032] In the channel attention module, adaptive average pooling and maximum pooling are first used to obtain the average pooling and maximum pooling of the input features respectively; channel-level attention is generated through two fully connected layers, and the ReLU activation function is used for nonlinear transformation. Finally, the channel weight is generated through the Sigmoid function; the channel weight is multiplied by the input feature channel by channel to enhance the expression of related features.
[0033] In the spatial attention module, the average and maximum values of the channel dimension of the feature map are first calculated to form a spatial saliency map. The local spatial information is further captured through convolution, and the spatial attention map is generated through Sigmoid activation. The input features are weighted pixel by pixel to highlight the important areas in space. After the output features are weighted by channel attention, they are further processed through spatial attention to enhance the representation ability of the features.
[0034] The spatial information is aggregated through adaptive average pooling to generate compact channel information and further processed by the fully connected layer; the output weights are applied to the feature map again to ensure that the final output features have been weighted by adaptive channels and spatial attention to generate pixel-level feedback features; the specific workflow is as follows:
[0035] f3=F concat ·CA(F concat );
[0036] f4=F concat ·SA(F concat );
[0037] F fusion =F concat ·Sigmoid(FC(AdaptiveAveragePooling(Add(f1,f2))));
[0038] Among them, CA stands for adaptive channel attention module, SA stands for adaptive spatial attention module, and F concat is the input of the adaptive channel-spatial attention module, f3 represents the output features after the adaptive channel attention module, f4 represents the output features after the adaptive spatial attention module, FC represents the fully connected layer, and F fusion It is the fusion output of the entire deep adaptive fusion module.
[0039] As a further technical solution, the decoder includes four 3×3 convolutional layers with output channel numbers of 256, 128, 64 and 32 respectively. Through the four 3×3 convolutional layers, the output feature map of the fusion layer is reconstructed by dimensionality reduction to restore a representation similar to the original image.
[0040] As a further technical solution, the image fusion method based on deep adaptive channel-spatial attention also includes training the fusion results of infrared light image and visible light image through loss function; wherein the loss function is composed of content loss L content and the gradient loss L gradient The content loss L content for:
[0041]
[0042] L content =α1*L en +α2*L me ;
[0043] Where L enrepresents the MEA loss applied to the visible light and infrared light fusion image pair; after calculating the image entropy weights of visible light and infrared light respectively, the fusion results will be used to calculate the MEA of the two modes of the source image; similarly, L me represents the MEA loss of visible light and infrared light under the median weighted influence; the overall L content By L en and L me Composition, where α1 and α2 are hyperparameters;
[0044] The gradient loss L gradient for:
[0045]
[0046] in, Represents the gradient operator of image texture. “│·│” indicates an absolute operation.
[0047] A second aspect of the present invention provides an image fusion system based on deep adaptive channel-spatial attention.
[0048] Image fusion system based on deep adaptive channel-spatial attention, including:
[0049] The image acquisition module is configured to: acquire infrared light images and visible light images;
[0050] The initial feature extraction module is configured to: input the acquired infrared light image and visible light image into the dual-branch interactive encoder to extract initial features of the infrared light image and the visible light image;
[0051] The image fusion module is configured to: use the deep adaptive score module to obtain the weights of the initial features of the infrared light image and the visible light image to generate weighted features; merge the weighted features and input them into the adaptive channel-spatial attention module to realize the fusion of infrared and visible light features;
[0052] The image reconstruction module is configured to: input the fused features into the decoder for dimensionality reduction reconstruction, and map the fused results back to the original representation.
[0053] The third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps in the image fusion method based on deep adaptive channel-spatial attention as described in the first aspect of the present invention.
[0054] The fourth aspect of the present invention provides an electronic device, comprising a memory, a processor, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the image fusion method based on deep adaptive channel-spatial attention as described in the first aspect of the present invention are implemented.
[0055] One or more of the above technical solutions have the following beneficial effects:
[0056] The present invention adopts a dual-branch interactive encoder structure to achieve the interaction of complementary features of two modalities in the feature extraction process, and enhances the capture of shallow texture features. The deep adaptive fusion module is used to adaptively assign weights to different features by calculating weight scores to achieve efficient feature fusion. The introduction of channel and spatial attention further improves the aggregation effect of features. Through the loss function, the training process is constrained based on the entropy ratio of the input image, and the median weighting mechanism is combined to reduce noise interference.
[0057] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0059] Figure 1 This is a flow chart of the method of the first embodiment.
[0060] Figure 2 It is a schematic diagram of the overall architecture of the first embodiment.
[0061] Figure 3 Schematic diagram of the structure of the CBAEM enhancement module in the first embodiment.
[0062] Figure 4 It is a schematic diagram of the structure of the CBA module in the first embodiment.
[0063] Figure 5 Schematic diagram of the channel attention module structure in the CBA module in the first embodiment.
[0064] Figure 6 Schematic diagram of the structure of the spatial attention module in the CBA module in the first embodiment.
[0065] Figure 7 FIG. 1 is a schematic diagram of the overall architecture of the CSAF module in the first embodiment.
[0066] Figure 8 Schematic diagram of the overall architecture of the deep adaptive scoring module in the first embodiment.
[0067] Fig. 9 Schematic diagram of the overall structure of the adaptive channel-spatial attention module in the first embodiment.
[0068] Fig.10 Schematic diagram of the specific structure of the channel attention module in the adaptive channel-spatial attention module in the first embodiment.
[0069] Fig.11 Schematic diagram of the specific structure of the spatial attention module in the adaptive channel-spatial attention module in the first embodiment.
[0070] Fig.12 It is a system structure diagram of the second embodiment. DETAILED DESCRIPTION
[0071] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0072] It should be noted that the terms used herein are for describing specific embodiments only and are not intended to be limiting of exemplary embodiments according to the present invention.
[0073] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.
[0074] The present invention proposes an image fusion method and system based on deep adaptive channel-spatial attention, which includes three parts: feature extraction, feature fusion and image construction. In the feature extraction part, initial feature extraction is performed on infrared images and visible light images through a dual-branch interactive encoder. Subsequently, the extracted initial features are input into the CSAF fusion module to deeply fuse the features of the two modalities. The CSAF module obtains the weights of the two modes, fuses the infrared and visible light features through the deep adaptive scoring (DAS) module, and then fuses the infrared and visible light features through the adaptive channel-spatial joint attention (ACSA) module. Finally, the fused features are input into the image construction part, and the fusion results are mapped back to the original representation.
[0075] Embodiment 1
[0076] This embodiment discloses an image fusion method based on deep adaptive channel-spatial attention;
[0077] Combination Figure 1 and Figure 2 , an image fusion method based on deep adaptive channel-spatial attention, including:
[0078] Step S1, acquiring an infrared light image and a visible light image;
[0079] Step S2, inputting the acquired infrared light image and visible light image into a dual-branch interactive encoder to extract initial features of the infrared light image and the visible light image;
[0080] Step S3, using the depth adaptive score module to obtain the weights of the initial features of the infrared light image and the visible light image to generate weighted features; merging the weighted features and inputting them into the adaptive channel-spatial attention module to realize infrared and visible light feature fusion;
[0081] In step S4, the fused features are input into the decoder for dimensionality reduction reconstruction, and the fused results are mapped back to the original representation.
[0082] In step S2, a dual-branch interactive encoder is used for feature extraction. The dual-branch interactive encoder has five convolutional layers, which is designed to fully extract complementary features and common features. First, the 1×1 convolutional layer is designed to reduce the modality difference between infrared images and visible light images. Then, a feature extraction module consisting of four weight-sharing convolutional layers and SE layers is used to extract deep features of infrared images and visible light images.
[0083] Table 1 lists more details of all convolutional layers in the feature extractor, such as kernel size, output channels, and activation functions. Except for the first layer, the kernel size of all convolutional layers is 3. All layers of the feature extractor use Leaky ReLU as the activation function.
[0084] Table 1 Kernel sizes, output channels, and activation functions of all convolutional layers in the infrared and visible image fusion model network based on deep adaptive channel-spatial attention.
[0085]
[0086] In addition, the differential information is also utilized through the CBAEM enhancement module in step 2. Specifically, the outputs of the 2nd, 3rd, and 4th layers of the dual-branch interactive encoder are followed up by the CBAEM module to exchange modal complementary features. The CBAEM module is able to gradually integrate complementary information in the feature extraction stage, ensuring that the common and complementary features in the infrared and visible light images are fully extracted.
[0087] Combination Figure 3 ,In the CBAEM enhancement module, a design similar to the differential ,amplifier circuit and a CBA module are included. The idea of the differential ,amplifier circuit is used to perform differential operations on the ,features of visible light and infrared light images to extract the ,difference information between the two.
[0088] The differential amplifier circuit is an electronic circuit whose core function is to amplify the difference between the input signals. Its basic structure includes two input terminals V1 and V2, an output terminal and a gain-controlled operational amplifier. The output signal is the difference between the two input signals V out =A×(V1-V2), where A is the gain.
[0089] The idea of differential amplifier circuit is applied to image fusion. In the image fusion task, the goal is to effectively fuse images from different modalities (such as visible light images and infrared images). These images usually have similar features in some areas, but significant differences in other areas. Using the idea of differential amplifier circuit, the difference information between the two can be extracted, so as to better fuse them. Specifically, in this embodiment, V1 is the processed visible light image feature; V2 is the processed infrared image feature; V out The difference information between the two.
[0090] The CBAM module is added to the differential features, and the differential features are refined in multiple dimensions through the channel and spatial attention mechanism. Figure 4 , in the CBA module, it mainly includes the channel attention module and the spatial attention module. Figure 5 As shown in Figure 1, the channel attention module uses global average pooling and global maximum pooling to process the feature map and extract the global features of each channel. These features are then processed through a two-layer convolutional network (1x1 convolution) and finally the weight of each channel is generated through the Sigmoid function. Finally, the channel weight is multiplied by the original input feature to complete the channel-level feature enhancement. Figure 6 As shown in the figure, the spatial attention module first performs average pooling and maximum pooling on the channel dimension of the feature map to generate an average feature map and a maximum response map. Then, the two feature maps are concatenated along the channel dimension and a convolution operation is used to generate a spatial weight map. Finally, this spatial weight map is multiplied with the input feature to complete the spatial level feature enhancement.
[0091] Step 21, extract features from the visible light image and the infrared image respectively to obtain the feature F of the visible light image vi and the infrared image feature F ir , whose formula is:
[0092] F vi =FeatureExtractor(I vi )
[0093] F ir =FeatureExtractor(I ir )
[0094] In the formula, Ivi ,I ir are the input visible light image and infrared light image respectively.
[0095] Step S22, after initial feature extraction of infrared and visible light image inputs, both modes are enhanced with CBAEM. In subsequent feature extraction layers, both visible and infrared image features depend on the differential features of the two modes in the previous layer. This ensures that in each layer, the features of the two modes can interact with each other, rather than each mode having a separate model weight. The design of the complementary enhancement module is as follows:
[0096]
[0097] F vi ′=CBA(F vi_complement ),F ir ′=CBA(F ir_complement );
[0098] Among them, F vi_complement and F ir_complement Represent the differential features of visible light image and infrared image respectively; F vi and F ir represents the visible light and infrared features input into the CBAEM enhancement module, CBA represents the CBA layer used to further enhance the complementary information, and F vi ′ and F ir ′ represents the module output. Using the concept of differential amplifier circuit, the differential features of visible light and infrared light can be obtained. These features are then enhanced through the channel-space dual attention mechanism in the CBA module and the features are recalibrated to obtain the output of this layer. This output is then fed back to the respective branches to further extract features.
[0099] In the dual-branch interactive encoder, by introducing three rounds of feature interaction before merging the features of the two modalities, this integration ensures the comprehensive integration of shallow feature information. The feature extraction of each modality is no longer isolated, because the feature information of each modality in each layer will affect the feature extraction weights of the other modality in the next layer.
[0100] Step S3, using the depth adaptive score module to obtain the weights of the initial features of the infrared light image and the visible light image to generate weighted features; merging the weighted features and inputting them into the adaptive channel-spatial attention module to realize infrared and visible light feature fusion;
[0101] Since the quality of visible and infrared images varies in different scenarios, even if feature extraction can effectively obtain information from both inputs, it may not be able to solve this problem. For example, in extremely dark environments, the quality of visible images is extremely low, while infrared images contain rich information. In other words, in dark conditions, the weight of infrared features should be greater than that of visible light features. Therefore, how to cleverly fuse the features of the two modes is a key step in image fusion algorithms.
[0102] In step S3, a fusion module based on channel and spatial attention mechanism, called deep adaptive fusion, is introduced. This module learns complementary information between different modalities by dynamically adjusting the weight of each channel. Specifically, it combines multi-scale convolutional features and channel-space adaptive weight allocation to enhance important targets from visible light and infrared light, thereby achieving a more robust fusion effect. This method can effectively improve the quality of image fusion.
[0103] like Figure 7 As shown in Figure 1, the deep adaptive fusion module takes the output of the feature extraction module as input. First, the deep adaptive score module generates independent weight scores for the initial features of the visible light and infrared images respectively. These weight scores are multiplied by the respective features channel by channel to generate weighted features. Then, the weighted features are merged and input into the adaptive channel-spatial attention module for further fusion processing. The final output achieves accurate multimodal information fusion by applying adaptive channel-spatial weights to the merged features.
[0104] Furthermore, in this embodiment, by combining maximum pooling and average pooling, the channel of feature standard deviation is further increased. The specific implementation is as follows Figure 8 As shown. In the pooling layer design of the fusion module, an adaptive pooling layer is used to enhance the adaptability of the model. First, the maximum channel, average channel, and standard deviation channel of the input feature are extracted, and they are connected to integrate the feature information through 3×3, 5×5, and 7×7 multi-scale convolution layers. Then, the convolution results are processed by adaptive maximum pooling and adaptive average pooling to optimize the feature expression from different angles. Finally, the deep adaptive score is obtained through the Sigmoid function for dynamic weighted features. The specific calculation formula is as follows:
[0105] f1=Concat(Max(F input ),Std(F input ),Mean(F input ));
[0106] f2=Add(3×3conv(f1),5×5conv(f1),7×7conv(f1));
[0107] S=Sigmoid(Add(AdaptiveMaxPool2d(f2),AdaptiveveragePooli2d(f2)));
[0108] Among them, Conv represents the convolution layer, Concat represents the connection layer, Sigmoid represents the Sigmoid activation function, Add represents the element addition operation, AdaptiveMaxPool2d represents the adaptive maximum pooling operation, and AdaptiveAveragePool2d represents the adaptive average pooling operation. f1 represents the feature vector after dimensionality reduction, the symbol f2 represents the feature vector after f1 is integrated through the multi-scale convolution layers of 3×3, 5×5 and 7×7, and S represents the weighted score of the input feature.
[0109] Furthermore, after obtaining the weights of visible light and infrared light respectively through the deep adaptive scoring module, the combined features are input into the adaptive channel-spatial attention module (such as Fig. 9 In the adaptive channel-spatial attention module, the channel and spatial attention mechanisms work together to enhance the expressiveness of features, highlight key information, and suppress irrelevant features.
[0110] like Fig.10 As shown in the figure, in the channel attention module, adaptive average pooling and maximum pooling are first used to obtain the average pooling and maximum pooling of the input features respectively to capture the importance of different channels. Then, the generation of channel-level attention is realized through two fully connected layers, and the ReLU activation function is used for nonlinear transformation. Finally, the channel weights are generated through the sigmoid function. These channel weights will be multiplied by the input features channel by channel to highlight the channels that are important for feature performance, thereby enhancing the expression of related features.
[0111] like Fig.11 As shown in the figure, in the spatial attention module, the average and maximum values of the channel dimension of the feature map are first calculated to form a spatial saliency map. Local spatial information is further captured through convolution, and a spatial attention map is generated through sigmoid activation. The input features are weighted pixel by pixel to highlight important areas in space. After the output features are weighted by channel attention, they are further processed by spatial attention so that important information of both channels and space is retained, thereby effectively enhancing the representation ability of the features. Then, the spatial information is aggregated through adaptive average pooling to generate compact channel information and further processed through the fully connected layer. The output weights are applied to the feature map again to ensure that the final output features have been weighted by adaptive channel and spatial attention to generate pixel-level feedback features. The specific workflow is as follows:
[0112] f3=F concat ·CA(F concat);
[0113] f4=F concat ·SA(F concat );
[0114] F fusi0n =F concat ·Sigmoid(FC(AdaptiveAveragePooling(Add(f1,f2))));
[0115] Among them, CA stands for adaptive channel attention module, SA stands for adaptive spatial attention module, and F concat is the input of the adaptive channel-spatial attention module, f3 represents the output features after the adaptive channel attention module, f4 represents the output features after the adaptive spatial attention module, FC represents the fully connected layer, and F fusion It is the fusion output of the entire deep adaptive fusion module.
[0116] Under nighttime lighting conditions, it is reasonable to infer that the infrared mode has a higher weight, as infrared images can effectively capture nighttime details. Under daytime lighting conditions, the contribution of color images is greater, so the visible light mode should have a higher weight. A classification network is used to identify whether the input image corresponds to day or night, and this information is used to dynamically adjust the weights during training. According to the proposed fusion module, a deep adaptive scoring mechanism is used to weight the features of the two modes through channel and spatial attention mechanisms. Under good lighting conditions, the visible light image has a higher weight score, while under poor lighting conditions, the infrared image has a higher weight score. This design ensures optimal feature fusion under different lighting conditions.
[0117] In step S4, the fused features are input into the decoder for dimensionality reduction reconstruction, and the fused results are mapped back to the original representation.
[0118] In the image reconstruction part, four 3×3 convolutional layers are used, and the number of output channels is 256, 128, 64, and 32 respectively. Through these layers, the output feature map of the fusion layer is reconstructed by dimensionality reduction to restore a representation similar to the original image. This process not only retains the structure, texture, and detail information of the fused image, but also reduces the feature dimension and redundancy. The final generated image minimizes information loss, thereby maintaining visual consistency and natural effects in terms of color, brightness, contrast, etc., ensuring the realism of the fused image.
[0119] In this embodiment, the fusion result of the infrared image and the visible light image is also trained by a loss function. The loss function is composed of the content loss L content and the gradient loss L gradient composition.
[0120] Although infrared images and visible light images differ in features, the instance structures of both types of images are stable. Image quality can be evaluated by utilizing parts with stable but inconsistent structures. In this case, using image entropy (EN) to measure quality can reflect parts with inconsistent structures. However, the EN value of an image is greatly affected by image noise. Therefore, median weighting is introduced to mitigate the impact of source image noise on the fusion result. Content loss L content for:
[0121]
[0122] L content =α1*L en +α2*L me ;
[0123] Where L en represents the MEA loss applied to the visible and infrared fused image pair. After calculating the image entropy weights for visible and infrared light respectively, the fusion results are used to calculate the MEA of the two modes of the source image. Similarly, L me represents the MEA loss of visible and infrared light under the median weighted influence. Overall L content By L en and L me where α1 and α2 are hyperparameters.
[0124] In addition, by introducing L gradient Gradient to ensure that the fusion result retains rich texture details and maintains the intensity distribution of the image. By maximizing the use of texture sets in visible and infrared images, the fusion result is forced to obtain more texture feature information:
[0125]
[0126] in, Represents the gradient operator of image texture. “│·│” indicates an absolute operation.
[0127] The final loss is:
[0128] L total =λ1·L content +λ2·L gradient ;
[0129] By L content Image entropy and median weighting are introduced to constrain the weights of the fusion results of the two modes, thereby obtaining more intensity information. gradient It is also used to constrain the fusion result to retain more texture information. Together, they constitute L total , where λ1 and λ2 are hyperparameters.
[0130] Experimental verification
[0131] The MSRS dataset was selected as the training set, which includes 427 daytime source images and 376 nighttime source images, so that the model can learn the fusion ability under different lighting conditions. These images were originally 640×480 in size and then cropped to 64×64, resulting in a total of 26,112 infrared-visible image pairs. This preprocessing step increases the size of the dataset, enhances the model's ability to fuse small objects, reduces overfitting, improves training performance, and speeds up training.
[0132] Three datasets are selected as test sets: TNO, LLVIP, and M3FD. The TNO dataset is one of the most common datasets in the field of image fusion, containing infrared and visible images in military-related scenes. M3FD contains high-resolution infrared-visible images from various scenes such as campuses, resorts, and roads. The LLVIP dataset consists of data pairs that are strictly calibrated in both time and space, so it is suitable for detecting pedestrians using infrared and visible light under low-light conditions. The performance of the model is evaluated on these three test sets representing different scenes and lighting conditions. Table 2 lists the quantitative results of our proposed deep adaptive channel-spatial attention infrared and visible image fusion model and other 11 algorithms on the TNO, M3FD, and LLVIP datasets. As can be seen from Table 2, the scheme in the present invention has achieved excellent results in many indicators of the three datasets.
[0133] Table 2 Quantitative analysis results of the TNO dataset (the first column of data for each indicator), the M3FD dataset (the second column of data for each indicator), and the LLVIP dataset (the third column of data for each indicator) (red indicates the best result, and blue indicates the second best result).
[0134]
[0135] Embodiment 2
[0136] This embodiment discloses an image fusion system based on deep adaptive channel-spatial attention;
[0137] like Fig.12 As shown, the image fusion system based on deep adaptive channel-spatial attention includes:
[0138] The image acquisition module is configured to: acquire infrared light images and visible light images;
[0139] The initial feature extraction module is configured to: input the acquired infrared light image and visible light image into the dual-branch interactive encoder to extract initial features of the infrared light image and the visible light image;
[0140] The image fusion module is configured to: use the deep adaptive score module to obtain the weights of the initial features of the infrared light image and the visible light image to generate weighted features; merge the weighted features and input them into the adaptive channel-spatial attention module to realize the fusion of infrared and visible light features;
[0141] The image reconstruction module is configured to: input the fused features into the decoder for dimensionality reduction reconstruction, and map the fused results back to the original representation.
[0142] Embodiment 3
[0143] The purpose of this embodiment is to provide a computer-readable storage medium.
[0144] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the image fusion method based on deep adaptive channel-spatial attention as described in Example 1.
[0145] Embodiment 4
[0146] The purpose of this embodiment is to provide an electronic device.
[0147] An electronic device comprises a memory, a processor and a program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the image fusion method based on deep adaptive channel-spatial attention as described in Example 1 are implemented.
[0148] The steps involved in the apparatuses of the above embodiments 2, 3 and 4 correspond to the method embodiment 1, and the specific implementation methods can refer to the relevant description part of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0149] Those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0150] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.
Claims
1. Image fusion method based on deep adaptive channel-spatial attention, characterized in that: include: Acquire infrared light images and visible light images; Input the acquired infrared image and visible light image into a dual-branch interactive encoder to extract initial features of the infrared image and the visible light image; The weights of the initial features of the infrared image and the visible light image are obtained by using the deep adaptive score module to generate weighted features; the weighted features are merged and input into the adaptive channel-spatial attention module to realize the fusion of infrared and visible light features; The fused features are input into the decoder for dimensionality reduction and reconstruction, and the fused results are mapped back to the original representation.
2. The image fusion method based on deep adaptive channel-spatial attention as claimed in claim 1, characterized in that: The process of inputting the acquired infrared light image and visible light image into the dual-branch interactive encoder to extract initial features of the infrared light image and the visible light image includes: First, the infrared image and the visible light image are processed by the 1x1 convolutional layer of the dual-branch interactive encoder to reduce the modal difference between the infrared image and the visible light image. Secondly, the 3x3 convolutional layer and SE layer of the dual-branch interactive encoder are used to extract the initial features of the infrared image and visible light image, and the CBAEM module is used to gradually integrate the complementary information.
3. The image fusion method based on deep adaptive channel-spatial attention as claimed in claim 2, characterized in that: The process of gradually integrating complementary information using the CBAEM module is as follows: inputting the extracted initial features of the infrared image and the visible light image into the CBAEM module to complete channel-level feature enhancement and spatial-level enhancement; The channel-level feature enhancement is as follows: the initial features of the extracted infrared image and visible light image are processed using global average pooling and global maximum pooling to extract the global features of each channel; the global features are processed through a two-layer convolutional network, and the weight of each channel is generated through a Sigmoid function; the channel weight is multiplied by the original input feature to complete the channel-level feature enhancement; The spatial level enhancement is as follows: perform average pooling and maximum pooling on the channel dimension of the feature map to generate an average feature map and a maximum response map; concatenate the average feature map and the maximum response map along the channel dimension, and use convolution operations to generate a spatial weight map; multiply the spatial weight map with the input feature to complete the spatial level feature enhancement.
4. The image fusion method based on deep adaptive channel-spatial attention as claimed in claim 1, characterized in that: The process of using the depth adaptive score module to obtain the weights of the initial features of the infrared image and the visible light image and generate weighted features is as follows: Extract the maximum channel, average channel, and standard deviation channel of the input features and connect them; Integrate feature information through multi-scale convolutional layers; perform adaptive maximum pooling and adaptive average pooling on the convolution results to optimize feature expression from different angles; The deep adaptive score is obtained through the Sigmoid function for dynamic weighted features; The formula is as follows: f1=Concat(Max(F input ),Std(F input ),Mean(F input )); f2=Add(3×3conv(f1), 5×5conv(f1), 7×7conv(f1)); S=Sigmoid(Add(AdaptiveMaxPool2d(f2), AdaptiveAveragePooli2d(f2))); Among them, Conv represents the convolution layer, Concat represents the connection layer, Sigmoid represents the Sigmoid activation function, Add represents the element addition operation, AdaptiveMaxPool2d represents the adaptive maximum pooling operation, and AdaptiveAveragePool2d represents the adaptive average pooling operation; f1 represents the feature vector after dimensionality reduction, f2 represents the feature vector of f1 after integration through 3×3, 5×5 and 7×7 multi-scale convolution layers, and S represents the weighted score of the input feature.
5. The image fusion method based on deep adaptive channel-spatial attention as claimed in claim 1, characterized in that: The process of implementing infrared and visible light feature fusion in the adaptive channel-spatial attention module is as follows: In the channel attention module, adaptive average pooling and maximum pooling are first used to obtain the average pooling and maximum pooling of the input features respectively; channel-level attention is generated through two fully connected layers, and the ReLU activation function is used for nonlinear transformation. Finally, the channel weight is generated through the Sigmoid function; the channel weight is multiplied by the input feature channel by channel to enhance the expression of related features; In the spatial attention module, the average and maximum values of the channel dimension of the feature map are first calculated to form a spatial saliency map. The local spatial information is further captured through convolution, and the spatial attention map is generated through Sigmoid activation to weight the input features pixel by pixel to highlight the important areas in space. After the output features are weighted by channel attention, they are further processed by spatial attention to enhance the representation ability of the features; The spatial information is aggregated through adaptive average pooling to generate compact channel information and further processed by the fully connected layer; the output weights are applied to the feature map again to ensure that the final output features have been weighted by adaptive channels and spatial attention to generate pixel-level feedback features; the specific workflow is as follows: <h2 style=";text-align:left;direction:ltr">f3=F<h2 style=";text-align:left;direction:ltr"> concat <h2 style=";text-align:left;direction:ltr"> CA(F<h2 style=";text-align:left;direction:ltr"> concat <h2 style=";text-align:left;direction:ltr"> ); f4=F concat ·SA(F concat ); F fusion =F concat ·Sigmoid(FC(AdaptiveAveragePooling(Add(f1,f2)))); Among them, CA stands for adaptive channel attention module, SA stands for adaptive spatial attention module, and F concat is the input of the adaptive channel-spatial attention module, f3 represents the output features after the adaptive channel attention module, f4 represents the output features after the adaptive spatial attention module, FC represents the fully connected layer, and F fusion It is the fusion output of the entire deep adaptive fusion module.
6. The image fusion method based on deep adaptive channel-spatial attention as claimed in claim 1, characterized in that: The decoder includes four 3×3 convolutional layers, and the number of output channels is 256, 128, 64 and 32 respectively. The output feature map of the fusion layer is reconstructed by reducing the dimension through the four 3×3 convolutional layers to restore the representation form similar to the original image.
7. The image fusion method based on deep adaptive channel-spatial attention as claimed in claim 1, characterized in that: The image fusion method based on deep adaptive channel-spatial attention also includes training the fusion results of infrared light image and visible light image through loss function; wherein the loss function is composed of content loss L content and the gradient loss L gradient The content loss L content for: L content =α1*L en +α2*L me ; Where L en represents the MEA loss applied to the visible light and infrared light fusion image pair; after calculating the image entropy weights of visible light and infrared light respectively, the fusion results will be used to calculate the MEA of the two modes of the source image; similarly, L me represents the MEA loss of visible light and infrared light under the median weighted influence; the overall L content By L en and L me Composition, where α1 and α2 are hyperparameters; The gradient loss L gradient for: in, Represents the gradient operator of image texture, "|·|" represents absolute operation.
8. Image fusion system based on deep adaptive channel-spatial attention, characterized by: include: The image acquisition module is configured to: acquire infrared light images and visible light images; The initial feature extraction module is configured to: input the acquired infrared light image and visible light image into the dual-branch interactive encoder to extract initial features of the infrared light image and the visible light image; The image fusion module is configured to: use the deep adaptive score module to obtain the weights of the initial features of the infrared light image and the visible light image to generate weighted features; merge the weighted features and input them into the adaptive channel-spatial attention module to realize the fusion of infrared and visible light features; The image reconstruction module is configured to: input the fused features into the decoder for dimensionality reduction reconstruction, and map the fused results back to the original representation.
9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps in the image fusion method based on deep adaptive channel-spatial attention as described in any one of claims 1 to 7 are implemented.
10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the image fusion method based on deep adaptive channel-spatial attention as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Highway pavement disease detection method and system
CN120182625A
Target detection method and system based on visible light and infrared light feature fusion
CN120219408A
Multispectral and visible light remote sensing image fusion segmentation method, system and medium
CN120472333A
Multispectral and visible light remote sensing image fusion segmentation method, system and medium
CN120472333B
Infrared light and visible light image fusion method, system and device based on double-branch cross-domain feature fusion and medium
CN121032817A