Multispectral image fusion method for dark light environment
Through the multi-layer cascade feature extraction and differential feature interaction mechanism of the DFDFuse model, the problems of unclear details and blurred target contours in multispectral image fusion in dark light environments are solved, and high-quality image fusion in low-brightness environments is achieved.
Patent Information
- Application Number
- CN202510583008.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-10-10
Smart Images

Figure CN120766070A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of image fusion, and particularly relates to a multispectral image fusion method for a dark light environment BACKGROUND
[0002] With the development of multi-modal perception technology, multispectral image fusion is widely used in night vision monitoring, automatic driving, medical diagnosis and other fields. Among them, visible light images have rich texture details but are limited by lighting conditions, and infrared light images can effectively penetrate smoke, haze and other interference through thermal radiation imaging, but lack of detailed features. In the extreme working environment of coal mine underground (dark, humid, narrow, high temperature), the imaging system faces many challenges: the roadway with a depth of more than 2000 meters completely relies on explosion-proof LED lighting, resulting in a sharp decrease in effective illumination from 10 lux at the light source to 0.3 lux 50 meters away, plus the ±95% illumination fluctuation caused by the dynamic shielding of the coal mining machine; the multi-phase medium interference is manifested as the coupling scattering effect of coal dust (particle size 2-50 μm, concentration 800-1500 mg / m 3 ) and water mist (RH≥95%), making the scattering coefficient of 550 nm visible light reach 1.2×10 -3 m -1 , and causing the change of lens refractive index; at the same time, 10-200Hz mechanical vibration causes the PSF half-width of the image to expand to 5 pixels, and the optical axis offset (0.13mm / 10℃) caused by the temperature gradient (20℃ / 10m) further deteriorates the imaging quality, which seriously restricts the reliability of the intelligent perception system of the mine. How to effectively fuse two types of modal information has become the key to improving the environmental perception ability of the vision system.
[0003] Existing methods of visible light and infrared light image fusion: traditional multi-scale decomposition method relies on artificial rules and has poor adaptability in complex scenes; in the deep learning framework, the dense connection architecture represented by DenseFuse in the convolutional network (CNN) maintains the continuity of features, but deep convolution easily leads to loss of weak target details, and the dynamic convolution kernel strategy alleviates this problem, but still lacks the ability to model global dependence across modalities; Transformer enhances cross-modal interaction by long-range feature association, but faces the bottleneck of high computational complexity and insufficient modeling of local feature heterogeneous association; GAN network is unstable during training and is prone to texture distortion.
[0004] The multiscale spatial decomposition and fusion method decomposes visible and infrared images at multiple scales, then performs feature selection at different scales. This extracts and processes different details, reducing information redundancy. While this method may perform well in specific scenarios, it lacks adaptability to complex scenarios or diverse data. The DenseFuse model employs an encoding-decoding approach, similar to residual connections, concatenating the feature maps output by each layer and feeding them into the next layer to extract more information. However, its over-reliance on convolutional computations results in low overall fusion quality. The SDCFusion model utilizes deep semantic information to drive image fusion, improving the performance of image fusion in downstream high-level vision tasks such as semantic segmentation and object detection. The SwimFuse model introduces the Transformer into the fusion network, addressing the issue of CNNs losing relevant information when extracting information. The DRMF model employs the Diffusion method to address image quality degradation under complex degradation conditions. The DDcGAN model uses a GAN approach for self-supervised generation of fused images, but training suffers from numerous instabilities that can lead to network failure.
[0005] Therefore, it is necessary to solve three issues: the lack of feature decoupling mechanism, insufficient directional feature extraction and blurred edge information. Summary of the Invention
[0006] The purpose of the embodiments of the present invention is to provide a multispectral image fusion method for a dark light environment, so as to solve the problem that the fused image details are unclear and the target outline is blurred in a dark light environment in the existing fusion model.
[0007] To solve the above technical problems, the technical solution adopted by the present invention is a multispectral image fusion method for a dark light environment, comprising the following steps:
[0008] S1: Build paired visible and infrared light datasets;
[0009] Select the registered visible light and infrared images, and then randomly divide the image pairs into training and test sets;
[0010] S2: Build the fusion model DFDFuse;
[0011] S21: Build the Base Avgformer Encoder module to decompose and extract image features;
[0012] S22: Construct a Fusion Layer module to fuse cross-modal features and construct feature maps;
[0013] S23: Construct a Decoder module to process the input feature map and output a fused image or the original image;
[0014] S3: training model;
[0015] The model training is divided into two steps. The first step is to train the encoder and decoder, and the output result is a pair of visible light and infrared light images. The second step is to train the encoder, fusion layer and decoder, and the output result is a fused image.
[0016] S4: Test model.
[0017] Furthermore, the S21 includes:
[0018] S211: Construct the bottleneck layer Bottle-Neck module; Bottle-Neck consists of three 3×3 convolutions and one GeLu activation function in sequence;
[0019] S212: Construct Block module; input F0 is processed by two branches. The first branch obtains feature map F1 after a 3×3 convolution and an LReLu activation function, and then is divided into two equal parts in the channel direction. The first half obtains feature map F3 and the second half obtains F4. Feature map F3 is normalized to obtain feature map F5. Feature map F5 is processed by a Sigmoid activation function and then multiplied with feature map F4 to obtain feature map F6. Then, feature map F6 and feature map F5 are spliced in the channel direction and sent to a 3×3 convolution and an LReLu activation function to obtain feature map F7; the second branch obtains feature map F8 after a grouped convolution GConv and a 1×1 convolution. Feature map F8 is added to feature map F7 to obtain the output of the Block module;
[0020] The Block feature extraction process is as follows:
[0021] F1=LReLu(Conv 3×3 (F0)) (1)
[0022] F2,F3=separate(F1) (2)
[0023] F4=Norm(F3); F5=F3·(Sigmoid(F4)) (3)
[0024] F6=LReLu(Conv 3×3 (Concat(F4,F5))) (4)
[0025] F7=Conv 1×1 (Conv 3×3 (F0)),Output=F7+F6 (5)
[0026] S213: constructing a deep feature extraction module Skip-Net; the module Skip-Net is sequentially composed of a first Block module, a second Block module and a third Block module, and there is a ReLu activation function after every 1 Block; wherein the output of the first Block module, the output of the second Block module and the input of the module Skip-Net are spliced in the channel as the input of the third Block module, and the output of the third Block module and the input of the module Skip-Net are added as the output of the module Skip-Net;
[0027] S214: constructing a multi-scale feature correlation module Self-Conv; in the module Self-Conv, the input feature map Z0 enters an adaptive average pooling module AdaptiveAvgPool of 3 parallel branches, the first branch obtains a 3x3 feature map Z1, the second branch obtains a 2x2 feature map Z2, and the third branch obtains a 6x6 feature map Z3, then the feature maps Z1, Z2 and Z3 are expanded in one dimension and spliced in order to send a Linear module to obtain a feature map Z4, the feature map Z4 is sent to two parallel branches, the first branch passes through a Linear module to obtain a feature map Z5, the second branch passes through a Sigmoid activation function to obtain a feature map Z6, then the feature maps Z5 and Z6 are added to obtain a 3x3 convolution kernel Z7, the input feature map Z0 is convolved with the convolution Z7 to obtain the output of the module Self-Conv;
[0028] The feature processing process of Self-Conv is as follows:
[0029] Z1=AdaptiveAvgPool(Z0),Z2=AdaptiveAvgPool(Z0),
[0030] Z3=AdaptiveAvgPool(Z0) (6)
[0031] Z4=Linear(Concat(Z1,Z2,Z3)) (7)
[0032] Z5=Linear(Z4),Z6=Sigmoid(Z4) (8)
[0033] Z7=Z5+Z6,Output=Conv(Z0,Z7) (9)
[0034] S215: Construct an edge perception module EPM; process the input feature map A0 through the module Self-Conv to obtain a feature map A1, and process the feature map A1 through the two parallel branches of directional feature extraction convolution d_conv and h_conv to obtain two results, which are added to the input feature map A0 to obtain a feature map A2, and then enter the dual-branch spatial attention mechanism module. The result of the feature map A2 obtained by the first branch spatial_mean and the result of the feature map A2 obtained by the second branch spatial_max are spliced on the channel and sent to 1 1×1 convolution and 1 Sigmoid activation function to obtain a feature map A3. The feature map A2 and the feature map A3 are multiplied to obtain a feature map A4, and then the feature map A4 is sent to the dual-branch channel attention mechanism module. The first branch is processed by channel_avg and MLP to obtain a feature map A5, and the second branch is processed by channel_max and MLP to obtain a feature map A6. The feature map A5 and the feature map A6 are added, and the result of multiplication with the feature map A4 is added to the feature map A2 to obtain the output of the module EPM;
[0035] Feature processing of module EPM:
[0036] A1=Self-Conv(A0) (10)
[0037] A2=d_conv(A1)+h_conv(A1)+A0 (11)
[0038] A3=Sigmoid(Conv 1×1 (Concat(spatial_mean(A2),spatial_max(A2)))) (12)
[0039] A4=A2×A3,A5=MLP(channel_avg(A4)) (13)
[0040] A6=MLP(channel_max(A4)) (14)
[0041] Output=((A5+A6)×A4)+A2 (15)
[0042] S216: constructing an average attention module Avg-Attention; first, input feature map B0 is passed through a 1*1 convolution and a 3*3 depth separable convolution DWConv to obtain feature map B1, then feature map B1 is equally divided into three feature maps in the channel direction, denoted as feature Q, feature K, and feature V, feature Q is subjected to average pooling Avg_pooling in the H direction to obtain feature map B2, feature K is subjected to average pooling Avg_pooling in the H direction to obtain feature map B3, the product of feature map B2 and feature map B3 is multiplied by feature V to obtain the output of the average attention module Avg-Attention;
[0043] S217: constructing a background feature extraction module BAN; in the module BAN, the result obtained by passing input feature map C0 through a Norm module and Avg-Attention is added to input feature map C0 to obtain feature map C1, and the result obtained by processing feature map C1 through a Norm module and an MLP is added to feature map C1 to obtain the output of the module BAN;
[0044] S218: constructing a module Mix; the module Mix is composed of two parallel branch convolution modules, and each convolution is followed by a ReLu6 activation function, the first branch is sequentially composed of a 3*3 convolution, a 3*3 depth separable convolution, and a 3*3 convolution, the second branch is a 3*3 convolution, and the outputs of the two branches are added after being multiplied by a learnable parameter to obtain the output of the module Mix;
[0045] S219: constructing a detail feature extraction module DCE; input feature map is divided into two equal parts X1 and X2 in the channel direction, X1 is processed by the module Mix and then added to X2 to obtain X3, X3 is processed by the module Mix and a Sigmoid activation function and then multiplied by X1 to obtain X4, X3 is processed by the module Mix and then added to X4 to obtain X5, and X5 and X3 are spliced in the channel direction to obtain the output, and three identical modules are sequentially connected to form the module DCE;
[0046] In the encoder Base Avgformer Encoder module, the input image is subjected to feature decomposition by a Bottle-Neck, a Skip-Net, and an EPM module, and then the output of the EPM module is input into the parallel BAN and DCE to obtain background feature D and detail feature E. The visible light and infrared light images share the modules and the parameters are shared.
[0047] Further, the S22 includes:
[0048] S221: constructing a difference feature module DFE; taking the processing of the background feature as an example, the visible light background feature D obtained in the encoder is passed through a 1*1 convolution, a 3*3 depth separable convolution, a 3*3 convolution, and a 1*1 convolution to obtain feature map D1, the product of feature map D1 and the output of the module DFE is added to the output of the module DFE to obtain the output of the module DFE.vi and infrared light background feature D ir As input, D vi and D ir respectively pass through 1x1 convolution to obtain D vi1 and D ir1 , then D vi1 and D ir1 after difference pass through Sigmoid activation function and multiply D vi0 to obtain D2, D ir1 and D vi1 after difference pass through Sigmoid activation function and multiply D ir to obtain D3, then D2, D3, D vi , D ir add to obtain the output F of the module DFE 背景 ;
[0049] Feature processing process of the module DFE:
[0050] D vi1 =Conv 1×1 (D vi ), D ir =Conv 1×1 (D ir1 ) (16)
[0051] D2=Sigmoid(D vi1 -D ir1 )xD vi , D3=Sigmoid(D ir1 -D vi1 )xD ir (17)
[0052] F 背景 =D2+D3+D vi +D ir (18)
[0053] S222: build cross-attention mechanism module Cross-Atten; visible light and infrared light images are processed by the encoder and the DFE to obtain features F 背景 and F 细节 , cross-attention calculation is performed with F 背景 as Q, F 细节 as K and V to obtain fusion feature F;
[0054] S223: build fusion feature enhancement module FFE; feature processing process of the FFE module:
[0055] Output=c-attention(Conv_c1(Input))*a+c-attention(Conv_c2(Input))*(1-a) (19)
[0056] In the FFE module, four special convolution kernels are added and spliced on the channel to form Conv_c1 and Conv_c2. The FFE input enters two parallel branches. The first branch passes through 1 Conv_c1, 1 channel attention mechanism c-attention and multiplies the result with the learnable parameter a. The second branch passes through 1 Conv_c2, 1 channel attention mechanism c-attention and multiplies the result with (1-a). The results of the two branches are added to the FFE input to obtain the output of the FFE module.
[0057] In the Fusion Layer module, the visible light background feature D vi0 and detail features E vi0 , infrared light background characteristics D ir0 and detail features E vi0 As the input of DFE, it then passes through the Cross-Atten module and the FFE module. The outputs of the FFE module and the Cross-Atten module are spliced on the channel as the output of the Fusion Layer module.
[0058] Furthermore, the S23 includes:
[0059] The output of the Fusion Layer module serves as the input of the Decoder module. The Decoder module consists of a 1×1 convolution, a feature enhancement module Feature-enchancement, and a res-conv module in sequence. The feature enhancement module consists of a 3×3 convolution and three Transformers in sequence. The res-conv module consists of a 3×3 convolution, a GELU activation function, and a 3×3 convolution in sequence. The output of the res-conv module is added to the visible light image and then activated by the Sigmoid function to obtain the output image of the Decoder module.
[0060] The beneficial effects of the present invention are:
[0061] To address the shortcomings of existing encoders in feature extraction, decomposition, and extraction, this paper proposes a multi-layer cascaded feature extraction architecture: a layered cascade structure consisting of a Skip-Net, an Edge Perception Module (EPM), and a dual-branch High-Frequency Detail and Low-Frequency Background Processing Module (DCE and BAN) is designed to enhance the internal correlation and detail processing capabilities of feature blocks. Furthermore, Avgformer average pooling is used to reduce computational complexity while strengthening the information connection between features. A multi-scale feature correlation module (a channel-feature driven parameter sharing mechanism) is designed to enhance the expressive power of features.
[0062] To address the blurring of object edges and loss of detail in existing fusion methods, this paper proposes a differentially driven cross-modal interaction mechanism: a differential feature extraction module (DFE) that enhances the differential features between modalities. Cross-Attention (Cross-Atten) enhances intermodal detail differences, preserving detail in visible and infrared images even in low-light environments and optimizing the processing of brightness and contrast differences. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0064] Figure 1 It is a structural diagram of the DFDFuse model of the method of the present invention.
[0065] Figure 2 It is a structural diagram of the deep feature extraction module Skip-Net of the method of the present invention.
[0066] Figure 3 (a) and (b) are schematic diagrams of the structures of the edge perception module and the multi-scale feature correlation module of the method of the present invention, respectively.
[0067] Figure 4 (a) and (b) are schematic diagrams of the structures of the background feature extraction module and the detail feature extraction module of the method of the present invention respectively.
[0068] Figure 5 (a) and (b) are schematic diagrams of the structures of the differential feature module and the fusion feature enhancement module of the method of the present invention respectively.
[0069] Figure 6 The method of the present invention is combined with ReCoNet, SuperFusion, YDTR, CrossFuse 0,CDDFuse 0 , comparison chart of fusion results of DATFuse, LRRNet, MMIF-EMMA, and PFCFuse image fusion methods. 0 DETAILED DESCRIPTION
[0070] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0071] This embodiment provides a multispectral image fusion method for a dark environment. The complete steps of the method are as follows:
[0072] S1: Dataset Construction. The dataset uses the public MSRS and LLVIP datasets. The MSRS dataset contains 1083 paired visible and infrared image pairs as the training set. The image pairs in the training set are randomly cropped into small image pairs of size 128×128, and low-contrast image pairs are removed. The test set uses 361 image pairs from MSRS and 360 image pairs from LLVIP.
[0073] S2: Build the fusion model DFDFuse. Figure 1 The fusion model DFDFuse consists of the encoding module BaseAvgformer Encoder, the fusion module Fusion Layer, and the decoder module Decoder. The specific steps are as follows:
[0074] S21: Construct the encoder module. The encoder module consists of a Bottle-Neck module, a Skip-Net deep feature extraction module, an EPM edge perception module, a BAN dual-branch parallel background feature extraction module, and a DCE detail feature extraction module. This module aims to improve the model's ability to decompose image detail features and extract directional edge features. The specific steps are as follows:
[0075] S211: Construct the Bottle-Neck module. Bottle-Neck preprocesses the input image into a feature map, consisting of three 3×3 convolutions and one GeLu activation function.
[0076] S212: Construct Block module; Figure 2As shown in the figure, the input F0 is processed by two branches. The first branch obtains the feature map F1 after a 3×3 convolution and an LReLu activation function, and then is divided into two equal parts in the channel direction. The first half obtains the feature map F3 and the second half obtains F4. The feature map F3 is normalized to obtain the feature map F5. The feature map F5 is processed by a Sigmoid activation function and then multiplied with the feature map F4 to obtain the feature map F6. Then, the feature map F6 and the feature map F5 are spliced in the channel direction and sent to a 3×3 convolution and an LReLu activation function to obtain the feature map F7; the second branch obtains the feature map F8 after a grouped convolution GConv and a 1×1 convolution. The feature map F8 is added to the feature map F7 to obtain the output of the Block module;
[0077] The Block feature extraction process is as follows:
[0078] F1=LReLu(Conv 3×3 (F0)) (1)
[0079] F2,F3=separate(F1) (2)
[0080] F4=Norm(F3); F5=F3·(Sigmoid(F4)) (3)
[0081] F6=LReLu(Conv 3×3 (Concat(F4,F5))) (4)
[0082] F7=Conv 1×1 (Conv 3×3 (F0)),Output=F7+F6 (5)
[0083] S213: Constructing the deep feature extraction module Skip-Net; Figure 2 As shown in the figure, the Skip-Net module consists of the first, second, and third blocks in sequence, with each block followed by a ReLu activation function. The output of the first and second blocks are concatenated with the input of the Skip-Net module on the channel to serve as the input of the third block module. The output of the third block module is then added to the input of the Skip-Net module to serve as the output of the Skip-Net module. This prevents loss of detail information during deep feature extraction and preserves the original input.
[0084] S214: Constructing a multi-scale feature correlation module Self-Conv; Figure 3As shown in (b), in the module Self-Conv, the input feature map Z0 enters the adaptive average pooling module AdaptiveAvgPool of three parallel branches. The first branch obtains a 3×3 feature map Z1, the second branch obtains a 2×2 feature map Z2, and the third branch obtains a 6×6 feature map Z3. Then, the feature maps Z1, Z2, and Z3 are expanded in one dimension and sequentially spliced into a Linear module to obtain a feature map Z4. The feature map Z4 is sent to two parallel branches. The first branch passes through a Linear module to obtain a feature map Z5, and the second branch passes through a Sigmoid activation function to obtain a feature map Z6. Then, the feature maps Z5 and Z6 are added to obtain a 3×3 convolution kernel Z7. The convolution kernel is constructed using the feature map itself, which is beneficial to enhance the expression of feature information and reduce the amount of calculation. The input feature map Z0 is convolved with the convolution Z7 to obtain the output of the module Self-Conv.
[0085] Self-Conv feature processing process:
[0086] Z1=AdaptiveAvgPool(Z0), Z2=AdaptiveAvgPool(Z0),
[0087] Z3=AdaptiveAvgPool (Z0) (6)
[0088] Z4=Linear (Concat (Z1, Z2, Z3)) (7)
[0089] Z5=Linear (Z4), Z6=Sigmoid (Z4) (8)
[0090] Z7=Z5+Z6, Output=Conv(Z0, Z7) (9)
[0091] S215: build edge perception module EPM; input feature map A0 is processed through module Self-Conv to obtain feature map A1, which will enhance the expression of feature map detail information and reduce the convolution low parameter quantity, feature map A1 is processed through two parallel directional feature extraction convolutions d_conv and h_conv to obtain two results, which are added to input feature map A0 to obtain feature map A2, which further improves the perception of the model to the edge details of the target, then enters the double-branch spatial attention mechanism module, the result obtained by feature map A2 through the first branch spatial_mean is spliced with the result obtained by feature map A2 in the second branch spatial_max in the channel, and then sent to a 1x1 convolution and a Sigmoid activation function to obtain feature map A3, feature map A2 and feature map A3 are multiplied to obtain feature map A4, and then feature map A4 is sent to the double-branch channel attention mechanism module, the first branch is processed through channel_avg and MLP to obtain feature map A5, the second branch is processed through channel_max and MLP to obtain feature map A6, the result of adding feature map A5 and feature map A6 is multiplied with feature map A4, and then added to feature map A2 to obtain the output of module EPM;
[0092] The feature processing process of module EPM:
[0093] A1=Self-Conv(A0) (10)
[0094] A2=d_conv(A1)+h_conv(A1)+A0 (11)
[0095] A3=Sigmoid(Conv 1×1 (Concat(spatial_mean(A2),spatial_max(A2)))) (12)
[0096] A4=A2×A3,A5=MLP(channel_avg(A4)) (13)
[0097] A6=MLP(channel_max(A4)) (14)
[0098] Output=((A5+A6)×A4)+A2 (15)
[0099] S216: Construct an average attention module Avg-Attention; while establishing long-distance feature interactions, it also reduces the amount of computation. First, the input feature map B0 is passed through a 1×1 convolution and a 3×3 depthwise separable convolution DWConv to obtain a feature map B1. The feature map B1 is then divided equally into three feature maps in the channel, represented as feature Q, feature K, and feature V. Feature Q is average pooled along the H direction in space to obtain feature map B2. Feature K is average pooled along the H direction in space to obtain feature map B3. The result of multiplying feature map B2 and feature map B3 is then multiplied by feature V to obtain the output of the average attention module Avg-Attention.
[0100] S217: Constructing background feature extraction module BAN; Figure 4 As shown in (a), the input feature map C0 is processed by a Norm module and Avg-Attention, and then added to the input feature map C0 to obtain the feature map C1. The feature map C1 is processed by a Norm module and an MLP, and then added to the feature map C1 to obtain the output of the module BAN.
[0101] S218: Construct module Mix; Module Mix consists of two parallel branches of convolutional modules, and each convolution is followed by a ReLu6 activation function. The first branch undergoes a 3×3 convolution, a 3×3 depthwise separable convolution, and a 3×3 convolution in sequence. The second branch undergoes a 3×3 convolution. The outputs of the two branches are multiplied by a learnable parameter and then added together as the output of the module Mix.
[0102] S219: Construct detail feature extraction module DCE; Figure 4 As shown in (b), the input feature map is divided into two equal parts X1 and X2 in the channel. X1 is processed by the module Mix and added to X2 to obtain X3. X3 is multiplied by X1 after the module Mix and Sigmoid activation function to obtain X4. X3 is processed by the module Mix and added to X4 to obtain X5. X5 and X3 are spliced in the channel direction to obtain the output. There are three identical modules connected in sequence to form the module DCE. Here, the output of the third module is used as the output of the module DCE.
[0103] In the encoder Base Avgformer Encoder module, the input image is decomposed through the Bottle-Neck, Skip-Net, and EPM modules, and then the output of the EPM module is sent to the dual-branch parallel BAN and DCE to obtain the background feature D vi 、D ir and detail features E vi 、Eir , visible light and infrared light images share the modules in the encoder and the parameters are shared.
[0104] S22: Construct the fusion layer. The fusion layer consists of the differential feature module DFE, the cross-attention mechanism module Cross-Atten, and the fusion feature enhancement module FFE. The specific steps are as follows:
[0105] S221: constructing a differential feature module DFE; Figure 5 As shown in (a), taking background feature processing as an example, the visible light background feature D in the encoder is vi and infrared background characteristics D ir As input, D vi With D ir After 1×1 convolution, D vi1 and D ir1 , then D vi1 and D ir1 After the difference is made, it is activated by Sigmoid function and then combined with D vi0 Multiply to get D2, D ir1 With D vi1 After the difference is made, it is activated by Sigmoid function and then combined with D ir Multiply to get D3, then D2, D3, D vi 、D ir Add together to get the output F of module DFE 背景 Similarly, F can also be obtained through detailed features 细节 ;
[0106] The feature processing process of module DFE:
[0107] D vi1 =Conv 1×1 (D vi ),D ir =Conv 1×1 (D ir1 ) (16)
[0108] D2=Sigmoid(D vi1 -D ir1 )×D vi ,D3=Sigmoid(D ir1 -D vi1 )×D ir (17)
[0109] F 背景 =D2+D3+D vi +D ir (18)
[0110] S222: Construct the Cross-Atten mechanism module; the visible light and infrared light images are processed by the encoder and DFE, and the feature F 背景 and F 细节 , F 背景 As features Q, F 细节 Perform cross attention calculation on feature K and feature V to obtain fused feature F;
[0111] S223: Constructing a fusion feature enhancement module FFE; Figure 5 As shown in (b), in the FFE module, the four special convolution kernels are added and spliced on the channel to form Conv_c1 and Conv_c2. The FFE input enters two parallel branches. The first branch passes through 1 Conv_c1, 1 channel attention mechanism c-attention and multiplies the result with the learnable parameter a. The second branch passes through 1 Conv_c2, 1 channel attention mechanism c-attention and multiplies the result with (1-a). The results of the two branches are added to the FFE input to obtain the output of the FFE module.
[0112] Feature processing of the FFE module:
[0113] Output=c-attention(Conv_c1(Input))*a+c-attention(Conv_c2(Input))*(1-a) (19)
[0114] In the Fusion Layer module, the background feature D vi 、D ir and detail features E vi 、E ir As the input of DFE, it then passes through the Cross-Atten module and the FFE module. The outputs of the FFE module and the Cross-Atten module are spliced on the channel as the output of the Fusion Layer module.
[0115] S23: Construct a decoder; the output of the Fusion Layer module is used as the input of the Decoder module. The Decoder module consists of a 1×1 convolution, a feature enhancement module Feature-enchancement, and a res-conv module in sequence. The feature enhancement module consists of a 3×3 convolution and three Transformers in sequence. The res-conv module consists of a 3×3 convolution, a GELU activation function, and a 3×3 convolution in sequence. The output of the res-conv module is added to the visible light image and then activated by the Sigmoid function to obtain the output image of the Decoder module.
[0116] S3: Train the model using the MSRS training set. Training is done in two steps: the first step trains the encoder and decoder, outputting visible and infrared image pairs; the second step trains the encoder, fusion layer, and decoder, outputting the fused image. This training approach stabilizes model training and improves the quality of generated images.
[0117] S4: Test the model using the MSRS and LLVIP test sets.
[0118] Comparison of nine classic image fusion networks: ReCoNet, SuperFusion, YDTR, CrossFuse 0 、CDDFuse 0,DATFuse, LRRNet, MMIF-EMMA, PFCFuse,The experimental results are shown in Table 1, where bold represents the best indicators. ReCoNet (see Huang Z, Liu J, Fan awareness[J].IEEE / CAA Journal of AutomaticaSinica,2022,9(12):2121-2137.), YDTR (see Wei Tang, Fazhi He, Yu Liu.YDTR: Infraredand Visible Image Fusion Via Y-Shape Dynamic Transformer[J], IEEE Transactionson Multimedia,2023,25:5413-5428.), CrossFuse (see Wang Z, Shao W, Chen Y, et al. Across-scale for details iterative attentional adversarial fusion network for infrared and visible images[J]. IEEE Transactions on Circuits and Systems for VideoTechnology, 2023, 33(8):3677-3688.), CDDFuse (see Zixiang Z, Haowen B, Jiangshe Z, Yulun Z, Shuang X, Zudi L, Radu T, Luc VG, et al.CDDFuse:Correlation-Driven Dual-Branch Feature Decomposition for Multi-Modality Image Fusion.[J],Proceedings-IEEE Computer Society Conference on Computer Vision and Pattern Recognition,2023:5906-5916.)DATFuse(详见Tang W,He F,Liu Y,et al.DATFuse:Infrared andvisible image fusion via dual attention transformer[J].IEEE Transactions onCircuits and Systems for Video Technology,2023,33(7):3159-3172.),LRRNet(详见Hui L,Tianyang X,Xiao-Jun W,Jiwen L,Josef K,et al.LRRNet:A NovelRepresentation Learning Guided Fusion Network for Infrared and VisibleImages.[J],IEEE Transactions on Pattern Analysis and Machine Intelligence,2023,45(9):11040-11052.),MMIF-EMMA(详见Zhao Z,Bai H,Zhang J,et al.Equivariantmulti-modality image fusion[C] / / Proceedings of the IEEE / CVF Conference onComputer Vision and Pattern Recognition.2024:25912-25921.),PFCFuse(详见Hu X,Liu Y,Yang F.PFCFuse:A Poolformer and CNN fusion network for Infrared-VisibleImage Fusion[J].IEEE Transactions on Instrumentation and Measurement,2024.).
[0119] Table 1 Comparison of fusion results of DFDFuse and box model on MSRS dataset
[0120]
[0121]
[0122] Table 1 shows the fusion results of each model in the MSRS test set. DFDFuse achieves the best result in six of the seven key indicators, with a 9.6% and 3.2% improvement over the second-best MMIF-EMMA model in VIF and SD, respectively, and a 3.2%, 4.5%, and 13.9% improvement over the second-best PFCFuse model in AG, SF, and MI, respectively, indicating that the complementary modeling capability of the fusion layer to visible light texture and infrared thermal radiation features is significantly better than existing methods; AG and SD are improved simultaneously, verifying the effectiveness of each module in the encoder in enhancing edge features; and the module structure design of the model is reasonable.
[0123] Figure 6 is a comparison chart of DFDFuse and each model on the LLVIP dataset. As can be seen, the fusion image of DFDFuse in low light scenes has balanced brightness and contrast, retains the background information of the infrared light image and the light source information of the visible light image as a whole, and fully restores the brightness of the infrared light image and the local detail texture of the visible light (such as the texture of the sidewalk tiles), and is closer to human visual perception; the overall background of ReCoNet, CrossFuse, and LRRNet fails to well fuse the infrared light information, and the preservation of visible light details by other models has degenerated, reflecting the full extraction and decomposition of image detail features by each module in the encoder, the differential feature DFE in the fusion layer highlighting the cross-modal difference information and highlighting the edge detail information of each modality, and the fusion feature enhancement module FFE preserving the respective significant features in visible light and infrared light. Overall, the fusion image of DFDFuse is the best.
[0124] In summary, the DFDFuse network proposed in the present application, under the feature decomposition-dynamic fusion collaborative architecture, uses multi-scale dynamic convolution to enhance edge representation, establishes a cross-modal attention mechanism to achieve adaptive feature complementarity, and ensures that image detail information and target edges are not missing while improving the structural consistency of the fusion image.
[0125] The various embodiments in the specification are described in a related manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments.
[0126] The above only describes the preferred embodiments of the present application, and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A multispectral image fusion method for dark light environment, characterized in that: include: S1: Build paired visible and infrared light datasets; Select the registered visible light and infrared images, and then randomly divide the image pairs into training and test sets; S2: Build the fusion model DFDFuse; S21: Build the Base Avgformer Encoder module to decompose and extract image features; S22: Construct a Fusion Layer module to fuse cross-modal features and construct feature maps; S23: Construct a Decoder module to process the input feature map and output a fused image or the original image; S3: training model; Model training is divided into two steps. The first step is to train the encoder and decoder, and the output result is a pair of visible light and infrared light images; the second step is to train the encoder, fusion layer and decoder, and the output result is a fused image. S4: Test model.
2. The multispectral image fusion method for a dark light environment according to claim 1, characterized in that: The S21 includes: S211: Construct the bottleneck layer Bottle-Neck module; Bottle-Neck consists of three 3×3 convolutions and one GeLu activation function in sequence; S212: Construct Block module; input F0 is processed by two branches. The first branch obtains feature map F1 after a 3×3 convolution and an LReLu activation function, and then is divided into two equal parts in the channel direction. The first half obtains feature map F2 and the second half obtains F3. Feature map F3 is normalized to obtain feature map F4. Feature map F4 is processed by a Sigmoid activation function and then multiplied with feature map F3 to obtain feature map F5. Then, feature map F5 and feature map F4 are spliced in the channel direction and sent to a 3×3 convolution and an LReLu activation function to obtain feature map F6; the second branch obtains feature map F7 after a grouped convolution GConv and a 1×1 convolution. Feature map F7 is added to feature map F6 to obtain the output of the Block module; The Block feature extraction process is as follows: F1=LReLu(Conv 3×3 (F0)) (1) F2,F3=separate(F1) (2) F4=Norm(F3); F5=F3·(Sigmoid(F4)) (3) F6=LReLu(Conv 3×3 (Concat(F4,F5))) (4) F7=Conv 1×1 (Conv 3×3 (F0)),Output=F7+F6 (5) S213: Construct a deep feature extraction module Skip-Net; the Skip-Net module is composed of the first Block module, the second Block module, and the third Block module in sequence, and each Block is followed by a ReLu activation function; the output of the first Block module, the output of the second Block module, and the input of the Skip-Net module are spliced on the channel as the input of the third Block module, and the output of the third Block module is added to the input of the Skip-Net module as the output of the Skip-Net module; S214: Construct a multi-scale feature correlation module Self-Conv; in the module Self-Conv, input feature map Z0 enters the adaptive average pooling module AdaptiveAvgPool of three parallel branches, the first branch obtains a 3×3 feature map Z1, the second branch obtains a 2×2 feature map Z2, and the third branch obtains a 6×6 feature map Z3, then the feature maps Z1, Z2, and Z3 are expanded in one dimension and sequentially spliced into a Linear module to obtain a feature map Z4, the feature map Z4 is sent to two parallel branches, the first branch passes through a Linear module to obtain a feature map Z5, and the second branch passes through a Sigmoid activation function to obtain a feature map Z6, then the feature maps Z5 and Z6 are added to obtain a 3×3 convolution kernel Z7, the input feature map Z0 is convolved with the convolution Z7 to obtain the output of the module Self-Conv; Self-Conv feature processing process: Z1=AdaptiveAvgPool(Z0), Z2=AdaptiveAvgPool(Z0), Z3=AdaptiveAvgPool (Z0) (6) Z4=Linear (Concat (Z1, Z2, Z3)) (7) Z5=Linear (Z4), Z6=Sigmoid (Z4) (8) Z7=Z5+Z6, Output=Conv(Z0, Z7) (9) S215: Construct an edge perception module EPM; process the input feature map A0 through the module Self-Conv to obtain a feature map A1, and process the feature map A1 through the two parallel branches of directional feature extraction convolution d_conv and h_conv to obtain two results, which are added to the input feature map A0 to obtain a feature map A2, and then enter the dual-branch spatial attention mechanism module. The result of the feature map A2 obtained by the first branch spatial_mean and the result of the feature map A2 obtained by the second branch spatial_max are spliced on the channel and sent to 1 1×1 convolution and 1 Sigmoid activation function to obtain a feature map A3. The feature map A2 and the feature map A3 are multiplied to obtain a feature map A4, and then the feature map A4 is sent to the dual-branch channel attention mechanism module. The first branch is processed by channel_avg and MLP to obtain a feature map A5, and the second branch is processed by channel_max and MLP to obtain a feature map A6. The feature map A5 and the feature map A6 are added, and the result of multiplication with the feature map A4 is added to the feature map A2 to obtain the output of the module EPM; Feature processing of module EPM: A1=Self-Conv (A0) (10) A2=d_conv (A1)+h_conv (A1)+A0 (11) A3=Sigmoid(Conv 1×1 (Concat(spatial_mean(A2),spatial_max(A2)))) (12) A4=A2×A3,A5=MLP(channel_avg(A4)) (13) A6=MLP(channel_max(A4)) (14) Output=((A5+A6)×A4)+A2 (15) S216: Construct an average attention module Avg-Attention; first, the input feature map B0 is passed through a 1×1 convolution and a 3×3 depthwise separable convolution DWConv to obtain a feature map B1, and then the feature map B1 is equally divided into three feature maps in the channel, represented as feature Q, feature K, and feature V. Feature Q is average pooled along the H direction in space to obtain feature map B2, and feature K is average pooled along the H direction in space to obtain feature map B3. The result of multiplying feature map B2 and feature map B3 is then multiplied by feature V to obtain the output of the average attention module Avg-Attention; S217: Construct a background feature extraction module BAN. In the module BAN, add the result obtained by passing the input feature map C0 through a Norm module and Avg-Attention to the input feature map C0 to obtain a feature map C1. Then, add the result obtained by passing the feature map C1 through a Norm module and an MLP to the feature map C1 to obtain the output of the module BAN. S218: Construct module Mix; Module Mix consists of two parallel branches of convolutional modules, and each convolution is followed by a ReLu6 activation function. The first branch undergoes a 3×3 convolution, a 3×3 depthwise separable convolution, and a 3×3 convolution in sequence. The second branch undergoes a 3×3 convolution. The outputs of the two branches are multiplied by a learnable parameter and then added together as the output of the module Mix. S219: Construct a detail feature extraction module DCE; divide the input feature map into two equal parts X1 and X2 on the channel, X1 is processed by the module Mix and added to X2 to obtain X3, X3 is multiplied with X1 after the module Mix and Sigmoid activation function to obtain X4, X3 is processed by the module Mix and added to X4 to obtain X5, X5 and X3 are spliced in the channel direction to obtain the output, and there are 3 identical modules connected in sequence to form the module DCE.
3. The multispectral image fusion method for a dark light environment according to claim 1, characterized in that: The S22 includes: S221: Construct differential feature module DFE; take background feature processing as an example, the visible light background feature D vi and infrared background characteristics D ir As input, D vi With D ir After 1×1 convolution, D vi1 and D ir1 , then D vi1 and D ir1 After the difference is made, it is activated by Sigmoid function and then combined with D vi Multiply to get D2, D ir1 With D vi1 After the difference is made, it is activated by Sigmoid function and then combined with D ir Multiply to get D3, then D2, D3, D vi 、D ir Add together to get the output F of module DFE 背景 ; The feature processing process of module DFE: D vi1 =Conv 1×1 (D vi ),D ir =Conv 1×1 (D ir1 ) (16) D2=Sigmoid(D vi1 -D ir1 )×D vi ,D3=Sigmoid(D ir1 -D vi1 )×D ir (17) F 背景 =D2+D3+D vi +D ir (18) S222: Construct the Cross-Atten mechanism module; the visible light and infrared light images are processed by the encoder and DFE, and the feature F 背景 and F 细节 , F 背景 As features Q, F 细节 Perform cross attention calculation on feature K and feature V to obtain fused feature F; S223: Constructing a fusion feature enhancement module FFE; feature processing process of the FFE module: Output=c-attention(Conv_c1(Input))*a+c-attention(Conv_c2(Input))*(1-a)(19) In the FFE module, four special convolution kernels are added and spliced on the channel to form Conv_c1 and Conv_c2. The FFE input enters two parallel branches. The first one passes through 1 Conv_c1, 1 channel attention mechanism c-attention and multiplies the result with the learnable parameter a. The second one passes through 1 Conv_c2, 1 channel attention mechanism c-attention and multiplies the result with (1-a). The results of the two branches are added to the FFE input to obtain the output of the FFE module. The output of the Cross-Atten module and the output of the FFE module are spliced on the channel to obtain the output of the FusionLayer module.
4. The multispectral image fusion method for a dark light environment according to claim 1, characterized in that: The S23 includes: The output of the FusionLayer module serves as the input of the Decoder module. The Decoder module consists of a 1×1 convolution, a feature enhancement module, and a res-conv module in sequence. The feature enhancement module consists of a 3×3 convolution and three Transformers in sequence. The res-conv module consists of a 3×3 convolution, a GELU activation function, and a 3×3 convolution in sequence. The output of the res-conv module is added to the visible light image and then activated by the Sigmoid function to obtain the output image of the Decoder module.