Multi-spectral image fusion model and fusion method based on double-branch self-attention-generative adversarial network

Through the multispectral image fusion model of dual-branch self-attention-generative adversarial network, the problems of brightness contrast imbalance and insufficient detail preservation in the fusion of visible light and infrared images are solved, and efficient and stable image fusion is achieved. It is suitable for embedded platforms and improves the real-time performance and quality of image fusion.

CN120747677APending Publication Date: 2025-10-03ANHUI UNIV OF SCI & TECH
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510577018.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing technologies in the fusion of visible light and infrared images have problems such as insufficient detail preservation, imbalance in brightness and contrast, unstable unsupervised training, and difficulty in real-time deployment.

Method used

A multispectral image fusion model based on a dual-branch self-attention-generative adversarial network is adopted, including an input preprocessing module, a dual-branch encoder module, a fusion decoder module and a dual-domain discriminator module. It is trained through a combined loss function of multi-scale structural similarity loss, pixel loss, gradient loss and adversarial loss to achieve efficient image fusion.

Benefits of technology

It achieves dual fidelity of brightness and details, improves the stability of unsupervised training and real-time reasoning capabilities, is suitable for embedded platforms, meets low-power edge reasoning needs, and demonstrates excellent fusion effects on multiple data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747677A_ABST
    Figure CN120747677A_ABST
Patent Text Reader

Abstract

The invention discloses a multispectral image fusion model based on a double-branch self-attention-generative adversarial network and a fusion method thereof, and the model sequentially comprises an input preprocessing module which is used for carrying out the same-amplitude mapping, normalization and overlapping block embedding of visible light and infrared original images, and generating a to-be-fused feature block; the double-branch encoder module captures a cross-modal long-distance dependence and overall brightness structure through multi-head deep convolution transpose attention (MDTA) and a gated deep convolution feedforward network (GDFN), and extracts high-frequency texture and edge information by using a reversible residual gating layer and detail DetailNode iteration; the fusion decoder module is used for carrying out multi-level self-attention-convolution reconstruction on the two paths of features after channel dimension splicing, and outputting a single-frame high-resolution fusion image; and the double-domain discriminator module comprises a visible light domain discriminator and an infrared domain discriminator which are respectively used for carrying out adversarial evaluation on the fused image and the corresponding modal truth value image so as to improve the detail authenticity and the thermal target contrast ratio of the fusion result. The technical problems that an existing infrared-visible light image fusion method is insufficient in detail reservation, unbalanced in brightness and contrast, poor in unsupervised training stability and the like are solved, the method can be deployed on embedded platforms needing real-time and multi-modal information enhancement such as night monitoring, unmanned driving and edge security and protection, and high-contrast and high-information-amount fusion imaging is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of multimodal image processing and computer vision, and proposes a real-time visible light-infrared image fusion model and a fusion method based on a dual-branch self-attention-generative adversarial network. Background Art

[0002] Image fusion aims to fuse multi-source image data from different sensors or sources to preserve more target information, texture details, or spectral features in a single output image. Taking visible light and infrared image fusion as an example, the fused image has important applications in security surveillance, target detection, medical imaging, and remote sensing monitoring. Traditional image fusion methods primarily rely on rule-based algorithms or transform domain processing, such as multi-scale decomposition and sparse representation. These methods offer some interpretability but are sensitive to noise and struggle to maintain stable fusion results in complex scenes (such as strong illumination and partial occlusion). With the advancement of deep learning, researchers have begun to introduce models such as convolutional neural networks (CNNs), generative adversarial networks (GANs), and transformers for image fusion. These data-driven methods can automatically learn fusion strategies, overcoming the limitations of traditional methods and improving the quality and detail fidelity of fused images.

[0003] Existing CNN-based fusion methods extract multi-scale features through multi-layer convolution, simplifying the process of manually designing rules. The resulting fused images are typically less artifact-prone and better detail-preserving. However, pure CNN methods suffer from issues such as lack of interpretability, large number of model parameters, strong dependence on training data, and inflexible fusion strategies. GAN-based fusion methods use adversarial training of the generator and discriminator to generate fusion results without the need for real fusion image pairs. However, GAN methods also suffer from training instability, difficult-to-control results, and high computational overhead. Furthermore, the emerging Transformer model, relying on a self-attention mechanism for global modeling, excels at capturing long-range dependencies and global structure. However, the Transformer network has a large number of parameters, and direct application to high-resolution images results in slow inference. Therefore, in fusion tasks, it is often necessary to combine the Transformer with convolutional or other lightweight modules to achieve a balance between global modeling capabilities and computational efficiency.

[0004] In summary, existing technologies in the field of visible light and infrared image fusion still face the following major challenges: First, how to simultaneously preserve the thermal radiation information of the infrared image and the detailed texture of the visible light image to avoid information imbalance and detail loss; second, how to enhance the clarity of subtle structures and edges during the fusion process; and third, how to improve the overall contrast and visual realism of the fused image while ensuring the stability and efficiency of model training. These issues require urgent solutions, necessitating a new technical solution to improve image fusion. Summary of the Invention

[0005] The purpose of the embodiments of the present invention is to provide a multispectral image fusion model and a fusion method based on a dual-branch self-attention-generative adversarial network to solve the problems of existing infrared-visible light image fusion technology such as insufficient detail preservation, brightness contrast imbalance, unstable unsupervised training, and difficulty in real-time deployment.

[0006] To solve the above technical problems, the present invention adopts the following technologies to design a multimodal image fusion model driven by a dual-branch self-attention-generative adversarial network, including:

[0007] The input preprocessing module is used to unify the size, normalize and embed overlapping blocks of visible light images and infrared images, and output feature blocks to be fused;

[0008] The dual-branch encoder module consists of a global structure encoder and a fine-scale enhancement unit in parallel, which extract global structural features and high-frequency detail features respectively;

[0009] The fusion decoder module is used to concatenate the two features in the channel dimension and reconstruct the output fused image through multiple layers of self-attention-convolution layers;

[0010] The dual-domain discriminator module includes a visible light domain discriminator and an infrared domain discriminator, which respectively perform authenticity discrimination on the fused image and the real image of the corresponding modality;

[0011] Furthermore, the multispectral image fusion model based on a dual-branch self-attention-generative adversarial network is characterized in that the input preprocessing module includes, in sequence: a size unification unit, which aligns and resamples the visible light and infrared images to the same resolution; a normalization unit, which implements zero mean-unit variance standardization on each channel; and an overlapping block embedding unit, which uses a 3×3, stride 1 convolution to map the spliced ​​image into a high-dimensional feature block with a local overlapping receptive field as the encoder input.

[0012] Furthermore, the multispectral image fusion model based on the dual-branch self-attention-generative adversarial network is characterized in that the dual-branch encoder includes two parallel paths: a global structure branch and a fine-scale enhancement branch: the global structure branch connects several efficient Transformer blocks (multidimensional convolution transposed attention MDTA + gated deep convolution feedforward network GDFN) in series to extract low-frequency contours and illumination; the fine-scale enhancement branch divides the input features equally and progressively enhances them through multiple levels of DetailNode, accumulating high-frequency textures in an interactive-modulation manner; the outputs of the two branches are spliced ​​in the channel dimension and merged with the shallow residual; the input features of the fine-scale enhancement branch are first evenly split into two groups of channels Z1 and Z2 through the channel split-shuffle convolution module and randomly interleaved, and then four DetailNodes are connected in series in sequence, and the DetailNode internally includes: a reversible residual gating map θ φ (·) is applied to Z1; the amplitude modulation mapping θ p (·) and the translation map θ η (·) is applied to Z2 and then iteratively updated; all DetailNode outputs are concatenated in the channel dimension and integrated through 1×1 convolution to obtain detail features.

[0013] Furthermore, the fusion decoder module includes a channel compression unit, a hierarchical reconstruction stack and an output convolution unit in sequence; the channel compression unit reduces the dimensionality of the high-dimensional tensor after splicing the global features and detail features through 1×1 convolution; the hierarchical reconstruction stack is stacked from top to bottom by several decoding layers, and each level first uses upsampling operation to restore the spatial resolution, and then integrates cross-branch and cross-scale information through the self-attention-convolution fusion unit, and obtains the details of the corresponding encoding layer through jump connection; the output convolution unit uses 3×3 convolution to map the accumulated features to a single-channel grayscale image, and finally generates a fused image.

[0014] Furthermore, the dual-domain discriminator module consists of a visible light discriminator and an infrared discriminator in parallel, both of which have the same structure and independent parameters; each discriminator sequentially contains a four-level convolution-downsampling unit, a batch normalization layer and a Leaky-ReLU activation, followed by global average pooling to compress the spatial dimension, and then outputs a single authenticity score through a fully connected layer and a Sigmoid; the convolution kernel size is set to 7×7, 5×5, 3×3, and 3×3 from top to bottom, with a stride of 2, to extract multi-scale discriminative features layer by layer and suppress artifacts.

[0015] Furthermore, a loss function of a multispectral image fusion model based on a dual-branch self-attention-generative adversarial network is characterized in that the following weighted combination loss function system is used when training the generator-discriminator adversarial network: (1) multi-scale structural similarity loss: the structural similarity between the fused image and the visible light image, and the fused image and the infrared image is calculated using three-scale SSIM, the weight of each scale is set to 0.4, 0.3, and 0.3 respectively, and the sum of the two 1-SSIM is taken as the structure preservation term; (2) pixel loss and gradient loss: first, the maximum value of the visible light Y channel and the infrared image is taken pixel by pixel in the pixel domain, and the L1 norm is used to constrain the maximum value map to keep the brightness consistent with the fused image; then the Sobel operator is used to extract the gradient amplitude of the three, the maximum value of the visible light and infrared gradient amplitudes is taken, and the L1 norm is used to constrain it to keep the same as the gradient amplitude of the fused image, and the gradient term is given a higher weight than the pixel term; (3) adversarial loss: the visible light discriminator and the infrared discriminator described in claim 5 are used to force the fused image to be judged as real according to the least squares adversarial criterion to improve cross-modal visual authenticity;

[0016] The combined loss of its model loss function and the optimization module construction is:

[0017] L total =L pixel +α·L grad +β·L ssim +η·L adv

[0018] Among them L pixel and L grad are pixel-level and gradient-level maximum constraints, L ssim is the multi-scale structural similarity loss, L adv is the least squares adversarial loss, and α, β, and η are weight coefficients.

[0019] Furthermore, the multispectral image fusion model based on the dual-branch self-attention-generative adversarial network is characterized in that its training process is as follows: first, the registered image pairs are cut into overlapping blocks of 128×128, stride 96 and contrast greater than 0.05; 8 pairs of samples are taken in each batch, the fusion image is output by the generator and the generator is reversely updated with the total generator loss defined in claim 7, and the Adam optimizer is used (the learning rate is initialized to 2×10 -4 , α=0.5, β=0.6, η=0.005) to update the generator; then freeze the generator, and update the two discriminators with the true modal graph and the fusion graph respectively; alternate iterations until the validation set converges.

[0020] The present invention also provides a multimodal image fusion method based on the model, comprising the following steps:

[0021] Step S1: convert the original visible light image The luminance component I is extracted by Y-Cb-Cr conversion v , the infrared grayscale image is recorded as I r , both are first subjected to the same geometric registration-resampling operator Unified to the resolution H×W, and then normalized by zero mean-unit variance to obtain the visible light brightness map I v and infrared image I r Normalized matrix and Then splice in the channel dimension And with 3×3 convolution Conv 3×3 (·) Projected as a feature tensor Where C is the number of convolution output channels;

[0022] Step S2: X0 is sent to both the global structure branch and the fine-scale enhancement branch. The former extracts the scene contour and low-frequency illumination through a 4-level efficient Transformer block to obtain the global feature G = f Global (X0); the latter first divides X0 by channel, and then performs reversible scale-shift transformation on the 4-level DetailNode to enhance the texture and edge, and outputs the detail feature D = f Detail (X0);

[0023] Step S3: After concatenating G and D in the channel dimension, the 1×1 convolutional compression operator ψ(·) is used to reduce the dimension and then sent to the K-level decoding Transformer stack. After 3×3 convolution and Sigmoid activation σ(·), the normalized fusion image is output.

[0024] Step S4: Input the fusion image Y into the visible light domain discriminator D vi and infrared domain discriminator D ir , and get the true probability P vi =D vi (Y), P ir =D ir (Y), used to measure the consistency of Y with the true distribution under each mode;

[0025] Step S5: Calculate the combined loss based on the loss function and the optimization module, use the Adam optimizer to backpropagate and update the generator and discriminator weights, and independently update D with the discriminator loss. vi With D ir ; The loop is iterated until the MS-SSIM and gradient consistency indicators of the fusion results in the validation set reach the preset threshold.

[0026] The beneficial effects of the present invention are:

[0027] Brightness-Detail Dual Fidelity: Pixel / gradient maximum constraints work in conjunction with a dual-branch encoder to preserve both highlighted objects and texture details. Training Stability and Unlabeled Learning: A dual-domain discriminator combined with a least-squares adversarial loss provides a flexible gradient, significantly accelerating the convergence of unsupervised training. Efficient Real-Time Inference: Cross-Channel Attention (MDTA) linearizes computational complexity and, in conjunction with a lightweight convolutional decoder, enables real-time fusion in high-resolution scenarios. Multi-Scenario Generalization: Seven objective indicators achieve optimal or suboptimal performance on the three public datasets of LLVIP, M3FD, and MSRS, as well as in real-world industrial monitoring scenarios. The mAP@0.5:0.95 ratio for downstream object detection in YOLOv7 is improved by 2.9-6.5 percentage points. Deployment-Friendly: The model boasts a symmetrical structure and scalable parameters, making it compatible with embedded platforms such as Jetson Orin NX and Rock 5 Plus, meeting low-power edge inference requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for describing the embodiments or the prior art.

[0029] Figure 1 2 is a schematic diagram of the structure of a dual-discriminator image fusion network provided by an embodiment of the present invention.

[0030] Figure 2 It is a schematic diagram of the overall structure of the fusion model provided by an embodiment of the present invention.

[0031] Figure 3 It is a schematic diagram of the generator encoder structure provided by an embodiment of the present invention.

[0032] Figure 4 It is a structural diagram of the multidimensional convolutional transposed attention (MDTA) and the gated deep convolutional feedforward network (GDFN) provided in an embodiment of the present invention.

[0033] Figure 5 2 is a schematic structural diagram of a global structure encoder module provided by an embodiment of the present invention.

[0034] Figure 6 2 is a schematic structural diagram of a fine scale enhancement unit (Fine Scale Enhancement Unit) provided in an embodiment of the present invention.

[0035] Figure 7 It is a schematic diagram of the decoder structure provided by an embodiment of the present invention.

[0036] Figure 8 2 is a schematic diagram of the discriminator structure provided by an embodiment of the present invention.

[0037] Figure 93 is a schematic diagram of subjective comparison of example images of the M3FD dataset under different fusion methods provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0038] The following is a further description of the dual-branch self-attention and generative adversarial network driven image fusion method proposed in the present invention in conjunction with the accompanying drawings and specific embodiments. It should be understood that the embodiments described herein are only used to explain the present invention and are not intended to limit the scope of protection of the present invention.

[0039] This embodiment provides an image fusion model based on a generative adversarial network, whose structure is as follows: Figure 1 and Figure 2 The model consists of a generator and two discriminators. The generator uses an improved encoder-decoder structure to fuse the input visible light image and infrared image to generate a single image. The two discriminators work independently, respectively constraining the fused image from the visible light domain and the infrared domain to be close to the real image distribution.

[0040] First, the visible light image and infrared image to be fused are merged according to the channel dimension as multi-channel input. Figure 3 As shown in Figure 2, the encoder of the generator extracts features from the input image. Specifically, the input passes through the overlapping block embedding module, which divides the original image into local blocks and maps them into a high-dimensional feature space. At the same time, local continuity across blocks is introduced through convolution to avoid the edge incoherence problem caused by simple block division. Then, the encoder stacks 4 layers of Transformer Block to extract features layer by layer. Each layer of the module contains an MDTA attention unit and a GDFN feedforward unit (such as Figure 4 MDTA (Multi-Dimensional Convolutional Transposed Attention) calculates global correlations in the channel dimension, improving computational efficiency by reducing the dimension of the attention matrix and preserving spatial neighborhood information in combination with local convolution operations. GDFN (Gated Deep Convolutional Feedforward Network) introduces a gating mechanism and deep convolutions, enabling the feedforward network to selectively pass important features and fuse information from adjacent pixels. Thanks to residual connections, each Transformer Block layer retains existing information while extracting new features, preventing gradient vanishing and achieving layer-by-layer feature enhancement.

[0041] After the encoder multi-layer extraction, the obtained feature map contains both visible light and infrared information. Subsequently, the output of the encoder is divided into two paths to achieve dual-branch feature extraction: the first branch is fed into the global structure encoder ( Figure 5 ), the second branch is fed into the fine-scale enhancement unit module ( Figure 6The global structure encoder consists of several series-connected attention-MLP units, which are used to further refine the global structural features of the image, such as the main outline of the scene, the general shape of the target, and the distribution of strong and weak light areas. The fine scale enhancement unit is used to enhance the detail information. Its internal structure is as follows: Figure 6 As shown in Figure 2, the fine-scale enhancement unit divides the input features into two parts, Z1 and Z2, in the channel dimension, and then iteratively transforms and updates Z1 and Z2 through multiple DetailNode modules: each DetailNode first concatenates Z1 and Z2 in the channel dimension and re-scatters them into two new sub-features through shuffled convolution to increase the information interaction between the two sub-features; then, the transformation rule is used to update the new Z2′ and Z1′. Among them, the mapping function θ φ (·) Extract fine details from Z1 and add them to Z2, mapping function θ p (·) and θ η (·) Extracting the scaling factor and offset from the updated Z2′ to adjust Z1. Through the cascading of multiple DetailNodes, the features of the detail branch are gradually enhanced, incorporating rich fine-grained information such as edges and textures. The features of the global branch, on the other hand, primarily preserve the image's large-scale structure and background information. The dual-branch encoder architecture enables the present invention to separately learn and preserve key features of images from different modalities: the heat source intensity of infrared images and the detailed texture of visible light images.

[0042] Next, if Figure 2 and Figure 7 As shown, the decoder of the generator fuses and reconstructs the above two features. First, the global features and detail features are spliced ​​in the channel dimension, and the number of channels is compressed to the number of channels of the original input image (for example, single-channel grayscale) through 1×1 convolution. Then, the fused features are sent to the multi-level decoding module to gradually restore the image: each level of the decoding module can correspond to a Transformer Block or convolution layer, corresponding to the level of the encoder. From high to low layers, the decoder side gradually introduces details to achieve the integration of information at different scales. In a preferred embodiment, the input of each layer of the decoder not only comes from the decoding output of the previous layer, but can also be spliced ​​with the features of the corresponding layer of the encoder through skip connection to supplement the detail information that may be lost during the reconstruction process. After layer-by-layer upsampling and convolution reconstruction, the decoder outputs the final fused image I fused Since the fusion process fully combines global and detailed features, the fused image of the present invention not only retains the highlighted targets in the infrared image but also presents the clear texture of the visible light image, thus achieving the complementary advantages of the two source image information.

[0043] While the generator completes the forward fusion, the present invention uses two discriminators to conduct adversarial discrimination training on the fusion results. Figure 1 and Figure 8 As shown, the discriminator D vi and D ir They are all composed of convolutional networks and work independently. vi The input of is the real visible light image and the generated fusion image (which appears similar to the low-light visible light image in the visible light mode), and the output is the classification result, which is used to judge whether the fusion image is indistinguishable from the real visible light image. Similarly, the discriminator D ir Input a real infrared image and a fused image, and determine whether the fused image has the thermal radiation characteristics of the real infrared image. Each discriminator contains multiple convolutional layers that gradually downsample and extract features, and finally outputs a discrimination probability value between 0 and 1 through a fully connected layer activated by Sigmoid. An output value close to 1 indicates that the discriminator believes that the input is a real image, and a value close to 0 indicates that the input is a generated fake image. During the training process, the two discriminators calculate their respective loss functions respectively. By continuously optimizing the discriminator parameters, it can accurately identify the difference between the generated image and the real image, thereby providing an improved signal to the generator during back propagation. It should be noted that the dual discriminators of the present invention do not simply exist at the same time, but each performs its own function: D vi Focusing on image details, texture, visible light brightness and other information, D ir The focus is on infrared features such as heat intensity distribution. This division of labor requires the generator to meet the requirements of both aspects during training, so that the generated fused image can appear more "realistic" and natural in both infrared and visible light domains.

[0044] In summary, the workflow of the generator-discriminator adversarial framework of the present invention is as follows: In each iteration, the training sample is input into the generator to obtain the fused image I fused ; Then I fused Send to D vi and D ir Compare and discriminate with the corresponding real image and calculate the discriminator loss; at the same time, according to the generated I fused The fusion loss, structural similarity loss and adversarial loss of the generator are calculated with the corresponding source image, and the weighted sum is used to obtain the total loss of the generator. Next, the generator parameters are fixed and the discriminator loss is minimized to update D vi and D irOnce for each parameter; then, the discriminator parameters are fixed and the generator parameters are updated once to minimize the generator loss. Through alternating training, after multiple rounds of iterations, the generator is able to produce fused images that are difficult for the two discriminators to distinguish between real and fake, while the discriminator maximizes its accuracy. When training reaches equilibrium, the fusion network model is trained. At this point, the generator has learned to fuse unseen visible light and infrared image pairs into a single output image. In inference applications, simply feed the newly input visible light and infrared images into the trained generator to quickly obtain the fused image.

[0045] To verify the effectiveness of the proposed method, we conducted comparative experiments on multiple standard datasets. The datasets include: the low-light visible-infrared dataset LLVIP, the multi-scene visible-infrared fusion dataset M3FD, and the day-night dual-scene visible-infrared dataset MSRS. The proposed method is compared with nine existing representative fusion algorithms, including: the 2024 CrossFuse method, the 2023 DATFuse method, the 2022 FLFuse method, the 2023 LRRNet method, the 2022 ReCoNet method, the 2022 SeAFusion method, the 2022 SwinFusion method, the 2024 TIMFusion method, and the 2022 YDTR method (the year in brackets indicates the year of publication). In order to objectively evaluate the fusion performance, the information fidelity (VIF), entropy (EN), average gradient (AG), spatial frequency (SF), standard deviation (SD), structural similarity difference (SCD), and Qabf evaluation indicators are used to quantitatively compare the fusion results. Among them: VIF reflects the fidelity of the fused image to the source image information. The larger the value, the more information details of the source image the fused image contains. EN represents the information entropy of the image. The higher the value, the richer the information contained in the fused image. AG represents the average gradient of the image, which is used to measure the image clarity and edge strength. SF represents the spatial frequency, which reflects the richness of the image texture details. SD is the grayscale standard deviation, which can be used to approximately measure the contrast of the image. SCD is an indicator that measures the difference in gradient distribution between the fused image and the source image. Qabf is a comprehensive indicator for evaluating the overall quality of the fused image (the value range is 0 to 1, and the closer to 1, the better the fusion quality).

[0046] Table 1 Comparison of quantitative indicators of LLVIP dataset (65 pairs of images)

[0047]

[0048] Note: Bold indicates the optimal value, and underlined indicates the suboptimal value.

[0049] Table 2 Comparison of quantitative indicators of the M3FD dataset (57 pairs of images)

[0050]

[0051]

[0052] Note: Bold indicates the optimal value, and underlined indicates the suboptimal value.

[0053] Table 3 Comparison of quantitative indicators of MSRS dataset (62 pairs of images)

[0054]

[0055] Note: Bold indicates the optimal value, and underlined indicates the suboptimal value.

[0056] Tables 1 through 3 list the average performance comparisons of the proposed method and the nine comparison algorithms on the LLVIP, M3FD, and MSRS datasets, respectively. It should be noted that bold values ​​indicate the best performance for that metric, while underlined values ​​indicate the second-best performance. This comparison demonstrates that the proposed method achieves leading overall performance on all datasets.

[0057] The LLVIP dataset (Table 1) was used to test visible light and infrared image fusion in low-light scenarios, with 65 pairs of images selected for evaluation. As shown in Table 1, the proposed method achieved the highest values ​​across all seven metrics. The VIF index reached 0.92, significantly higher than the second-place SwinFusion method's 0.84; the entropy value (EN) reached 7.56, surpassing the next-highest SwinFusion method (7.45); and detail clarity metrics such as average gradient (AG) and spatial frequency (SF) reached 6.82 and 22.13, respectively, significantly outperforming other methods. These results demonstrate that under low-light conditions, the proposed method is able to maximize the fusion of infrared and visible light information, producing images rich in detail and high in information content. Among the other methods, SwinFusion performed particularly well, achieving the second-best performance in some metrics (e.g., VIF = 0.84, AG = 6.36), significantly outperforming other traditional algorithms, but still lagging behind the proposed method overall. This demonstrates that the proposed dual-branch Transformer encoder and dual-discriminator GAN framework possesses enhanced fusion capabilities under complex conditions.

[0058] The M3FD dataset (Table 2) contains 57 pairs of visible-light and infrared images under various weather and lighting conditions. As shown in Table 2, the proposed method also achieves the highest values ​​across all metrics. In particular, the average gradient (AG) and spatial frequency (SF) metrics achieved by the proposed method reach 5.71 and 16.17, respectively, far exceeding the next highest methods (e.g., CrossFuse's AG of 5.44 and SF of 15.53). Furthermore, the proposed method achieves a Vibration Index (VIF) of 0.79, significantly ahead of the second-place SeAFusion (0.70). The entropy (EN) value reaches 7.16, the highest among all methods. This demonstrates that the proposed method can consistently produce high-quality fused images across a wide range of scenarios, including daytime and nighttime. Notably, while some comparison methods approach the proposed method in individual metrics, such as SeAFusion's VIF of 0.70 (the second highest) and CrossFuse's suboptimal performance in some metrics, no single method surpasses the proposed method across all metrics. This demonstrates the overall performance advantage of the proposed method.

[0059] The MSRS dataset (Table 3) combines two types of scenes, day and night, with a total of 62 pairs of test images. As can be seen from Table 3, most of the indicators of the method of the present invention still rank first, and are only slightly lower than the TIMFusion method in the entropy value EN. Specifically, the EN of TIMFusion reaches 7.08, and the EN of the present invention is 6.92, with a small gap. Except for the entropy value, the present invention has achieved the best or tied for the best results in the other six indicators: for example, the VIF is 1.0908 (significantly higher than the second highest SeAFusion method of 0.99), indicating that the fused image contains more effective information than the source image; the average gradient AG is 4.68, exceeding the second best value of 4.44; the standard deviation SD reaches 48.06, which is also the highest among all methods, indicating that the contrast of the fused image is significantly improved. Overall, even in the MSRS dataset with complex and changeable scenes, the present invention still demonstrates excellent robustness and generalization ability. While the TIMFusion method has a slight advantage in luminance information (EN), it lags significantly behind in other metrics such as detail clarity and structural fidelity (for example, TIMFusion's Qabf is only 0.44, while the proposed method is 0.69). Therefore, the overall fusion quality is still inferior to that of the proposed method. This shows that leading in individual metrics cannot compensate for deficiencies in other areas. The proposed method achieves balanced optimization in maintaining brightness, enhancing details, and improving contrast, thereby achieving the best overall fusion performance.

[0060] From the above quantitative results, it can be seen that the method of the present invention can achieve a better balance and higher index values ​​in terms of brightness, contrast, and detail clarity compared to various existing fusion algorithms. This is mainly due to: the dual discriminator mechanism ensures that both infrared intensity and visible light details are preserved, avoiding the information of one side from drowning out the other; the dual-branch encoder structure fully extracts multimodal features, and the detail enhancement unit improves the amount of texture and edge information; the fusion loss function provides clear optimization goals at the brightness and gradient levels, prompting the fused image to achieve the best in both subjective and objective indicators. Experimental results show that the fused image of the present invention restores the scene more naturally in human vision, important targets are easier to identify, and it is also comprehensively ahead in objective indicators, thus verifying the effectiveness and superiority of the present invention.

Claims

1. A multispectral image fusion model based on a dual-branch self-attention-generative adversarial network, characterized by: include: The input preprocessing module is used to unify the size, normalize and embed overlapping blocks of visible light images and infrared images, and output feature blocks to be fused; The dual-branch encoder module consists of a global structure encoder and a fine-scale enhancement unit in parallel, which extract global structural features and high-frequency detail features respectively; The fusion decoder module is used to concatenate the two features in the channel dimension and reconstruct the output fused image through multiple layers of self-attention-convolution layers; The dual-domain discriminator module includes a visible light domain discriminator and an infrared domain discriminator, which respectively perform authenticity discrimination on the fused image and the real image of the corresponding modality.

2. A multispectral image fusion model based on a dual-branch self-attention-generative adversarial network according to claim 1, characterized in that: The input preprocessing module includes: a size unification unit, which aligns and resamples the visible light and infrared images to the same resolution; a normalization unit, which implements zero mean and unit variance standardization on each channel; and an overlapping block embedding unit, which uses 3×3, stride 1 convolution to map the spliced ​​image into high-dimensional feature blocks with local overlapping receptive fields as encoder input.

3. A multispectral image fusion model based on a dual-branch self-attention-generative adversarial network according to claim 1, characterized in that: The dual-branch encoder includes two parallel paths: a global structure branch and a fine-scale enhancement branch; The global structure branch connects several efficient Transformer blocks (Multi-Dimensional Convolutional Transposed Attention MDTA + Gated Deep Convolutional Feedforward Network GDFN) in series to extract low-frequency contours and illumination; The fine-scale enhancement branch evenly divides the input features and progressively enhances them through multiple levels of DetailNode, accumulating high-frequency textures in an interactive-modulation manner; the outputs of the two branches are spliced ​​in the channel dimension and merged with the shallow residual; The input features of the fine-scale enhancement branch are first split into two groups of channels, Z1 and Z2, by the channel split-shuffle convolution module and randomly interleaved, and then four DetailNodes are connected in series. The DetailNode internally includes: reversible residual gating mapping θ φ (·) is applied to Z1; the amplitude modulation mapping θ p (·) and the translation map θ η (·) is applied to Z2 and then iteratively updated; all DetailNode outputs are concatenated in the channel dimension and integrated through 1×1 convolution to obtain detail features.

4. A multispectral image fusion model based on a dual-branch self-attention-generative adversarial network according to claim 1, characterized in that: The fusion decoder module includes a channel compression unit, a hierarchical reconstruction stack and an output convolution unit in sequence; The channel compression unit reduces the dimensionality of the high-dimensional tensor obtained by concatenating the global features and the detail features through 1×1 convolution; The hierarchical reconstruction stack consists of several decoding layers stacked from top to bottom. Each layer first uses upsampling operations to restore spatial resolution, then integrates cross-branch and cross-scale information through self-attention-convolution fusion units, and obtains detailed supplements from the corresponding encoding layer through skip connections; The output convolution unit uses 3×3 convolution to map the accumulated features to a single-channel grayscale image, and finally generates a fused image.

5. A multispectral image fusion model based on a dual-branch self-attention-generative adversarial network according to claim 1, characterized in that: The dual-domain discriminator module consists of a visible light discriminator and an infrared discriminator in parallel, both of which have the same structure and independent parameters. Each discriminator sequentially contains a four-stage convolution-downsampling unit, a batch normalization layer and a Leaky-ReLU activation, followed by global average pooling to compress the spatial dimension, and then outputs a single authenticity score through a fully connected layer and a Sigmoid. The convolution kernel sizes are set from top to bottom to 7×7, 5×5, 3×3, and 3×3, with a stride of 2, to extract multi-scale discriminative features layer by layer and suppress artifacts.

6. A loss function for a multispectral image fusion model based on a dual-branch self-attention-generative adversarial network, characterized by: The following weighted combination of loss functions is used when training the generator-discriminator adversarial network: (1) Multi-scale structural similarity loss: The structural similarity between the fused image and the visible light image, and between the fused image and the infrared image is calculated using three-scale SSIM. The weight of each scale is set to 0.4, 0.3, and 0.3 respectively. The sum of the two is taken as the structure preservation term. (2) Pixel loss and gradient loss: First, the maximum value of the visible light Y channel and the infrared image is taken pixel by pixel in the pixel domain, and the L1 norm is used to constrain the maximum value map to keep the brightness consistent with the fused image; then the Sobel operator is used to extract the gradient amplitude of the three, and the maximum value of the visible light and infrared gradient amplitude is taken, and the L1 norm is used to constrain it to keep the same as the gradient amplitude of the fused image. The gradient term is given a higher weight than the pixel term; (3) Adversarial loss: using the visible light discriminator and infrared discriminator described in claim 5, respectively forcing the fused image to be judged as real according to the least squares adversarial criterion, so as to improve the cross-modal visual authenticity; The combined loss of its model loss function and the optimization module construction is: L total =L pixel +α·L grad +β·L ssim +η·L adv Among them L pixel and L grad are pixel-level and gradient-level maximum constraints, L ssim is the multi-scale structural similarity loss, L adv is the least squares adversarial loss, and α, β, and η are weight coefficients.

7. A multispectral image fusion model based on a dual-branch self-attention-generative adversarial network according to any one of claims 1 to 5, characterized in that: The training process is as follows: first, the registered image pairs are cut into overlapping blocks of 128×128, stride 96 and contrast greater than 0.05; 8 pairs of samples are taken in each batch, and the fusion image is output by the generator and the total generator loss defined in claim 7 is used to update the generator in reverse, and the Adam optimizer is used (the learning rate is initialized to 2×10 -4 , α=0.5,β=0.6、η=0.005) update the generator; Then freeze the generator and update the two discriminators with the real modal image and the fusion image respectively; iterate alternately until the validation set converges.

8. A multispectral image fusion method based on a dual-branch self-attention-generative adversarial network, characterized in that: The following steps are involved: Step S1: convert the original visible light image The luminance component I is extracted by Y-Cb-Cr conversion v , the infrared grayscale image is recorded as I r , both are first subjected to the same geometric registration-resampling operator Unified to the resolution H×W, and then normalized by zero mean-unit variance to obtain the visible light brightness map I v and infrared image I r Normalized matrix and Then splice in the channel dimension And with 3×3 convolution Conv 3×3 (·) Projected as a feature tensor Where C is the number of convolution output channels; Step S2: X0 is sent to both the global structure branch and the fine-scale enhancement branch. The former extracts the scene contour and low-frequency illumination through a 4-level efficient Transformer block to obtain the global feature G = f Global (X0); the latter first divides X0 by channel, and then performs reversible scale-shift transformation on the 4-level DetailNode to enhance the texture and edge, and outputs the detail feature D = f Detail (X0); Step S3: After concatenating G and D in the channel dimension, the 1×1 convolutional compression operator ψ(·) is used to reduce the dimension and then sent to the K-level decoding Transformer stack. After 3×3 convolution and Sigmoid activation σ(·), the normalized fusion image is output. Step S4: Input the fusion image Y into the visible light domain discriminator D vi and infrared domain discriminator D ir , and get the true probability P vi =D vi (Y), P ir =D ir (Y), used to measure the consistency of Y with the true distribution under each mode; Step S5: Calculate the combined loss based on the loss function and the optimization module, use the Adam optimizer to backpropagate and update the generator and discriminator weights, and independently update D with the discriminator loss. vi With D ir ; The loop is iterated until the MS-SSIM and gradient consistency indicators of the fusion results in the validation set reach the preset threshold.

Citation Information

Cited By

  • Spectrum reconstruction method based on dual-branch heterogeneous feature fusion

    CN121074274A

  • Multi-modal image fusion method, device and equipment based on double-branch heterogeneous network

    CN121235918A

  • A multi-modal image fusion method, device and equipment based on a dual-branch heterogeneous network

    CN121235918B

  • Infrared light-visible light image data fusion method

    CN121330445A

  • Frequency modulation and wavelet sub-band guided double-domain cooperative Transform X-ray image denoising method

    CN121981912A