Image fusion method and system based on two-stage adversarial training and edge perception

By adopting two-stage adversarial training and edge perception methods in image fusion, using a two-way interactive encoder and a two-branch decoder, the defects of infrared and visible light images in the prior art are solved, and a high-quality image fusion effect is achieved.

CN120163719AActive Publication Date: 2025-06-17TAISHAN UNIV

Patent Information

Application Number
CN202510321415.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-17
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

The existing infrared and visible image fusion methods have defects such as unnatural fusion results, loss of details, insufficient contrast, and blurred edge details. GAN-based methods are difficult to take into account the quality of feature extraction and image reconstruction, and lack edge feature processing mechanisms.

Method used

Using an image fusion method based on two-stage adversarial training and edge perception, the two-way interactive encoder, dual-branch decoder and optimized discriminator structure is designed to achieve refined processing of the frequency domain and feature extraction to improve the quality of the fused image.

Benefits of technology

The quality of the fusion image is significantly improved, ensuring that the fusion image reaches a high quality level in overall structure and detail performance, and has better detail clarity and overall consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163719A_ABST
    Figure CN120163719A_ABST
Patent Text Reader

Abstract

The invention provides an image fusion method and system based on two-stage adversarial training and edge perception, and belongs to the field of image processing. The method comprises the following steps: selecting an infrared-visible light sample image pair data set, and combining paired images to obtain an initial combined image; constructing a fusion feature generation adversarial model comprising a bidirectional interaction encoder, a double-branch decoder and a discriminator, inputting the initial merged image into the bidirectional interaction encoder, and enhancing features through a multi-scale feature enhancement and high and low frequency feature decomposition module; and obtaining multi-scale fusion features, inputting the multi-scale fusion features into a double-branch decoder, and fusing and reconstructing infrared and visible light image features to form fusion features. And inputting the image into a discriminator, training the model according to a two-stage training strategy until the loss function converges, and fusing the to-be-fused and merged image based on the trained model. By designing a two-stage training strategy and introducing an edge feature enhancement mechanism, fine processing of a frequency domain is realized, and feature extraction and fusion effects and quality of a fused image are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to an image fusion method and system based on two-stage adversarial training and edge perception. Background Art

[0002] With the development of image processing technology, infrared and visible light image fusion plays a crucial role in many important fields, such as night vision monitoring, security systems, military reconnaissance, etc. Visible light images can present rich texture details and color information, while infrared images can capture the target thermal radiation information in low-light environments. The fusion of the two can obtain more comprehensive scene information, helping to improve the accuracy of target detection and scene understanding.

[0003] Current infrared and visible light image fusion methods have various defects: traditional multi-scale decomposition-based methods (such as wavelet transform, pyramid transform, etc.) have problems such as unnatural fusion results and detail loss; sparse representation-based methods are computationally complex and difficult to preserve the complete information of the source images; pulse-coupled neural network-based methods have defects such as insufficient contrast and blurred edge details in the fused images.

[0004] With the rise of deep learning technology, image fusion methods based on generative adversarial networks (GANs) have become a research hotspot, but many problems have also emerged: their single training stage is difficult to balance feature extraction and image reconstruction quality, resulting in poor fusion results in terms of details and structural integrity; the lack of a dedicated edge feature processing mechanism makes the edge contours of the fused images unclear; the handling of different frequency components is rough, and it is unable to effectively separate and enhance high and low frequency information. At the same time, the existing discriminator structure is simple, only relying on a simple convolutional network to discriminate features, lacking multi-scale perception ability, which limits the effect of generative adversarial learning; the decoding stage uses a single reconstruction path and cannot fully utilize feature information at different levels, restricting the quality of the reconstructed images. Summary of the Invention

[0005] To solve the above problems, the present invention proposes an image fusion method and system based on two-stage adversarial training and edge perception. By designing a two-stage training strategy and introducing an edge feature enhancement mechanism, fine-grained processing in the frequency domain and an optimized design of the discriminator structure are realized, improving feature extraction and fusion effects, and effectively enhancing the quality of the fused images.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] In a first aspect, the present invention provides an image fusion method based on two-stage adversarial training and edge perception, including:

[0008] Select an infrared-visible light sample image pair dataset, and perform channel merging on the infrared images and visible light images in the dataset to obtain an initial merged image;

[0009] Construct a fusion feature generation adversarial model, which includes a bidirectional interaction encoder, a two-branch decoder, and a discriminator;

[0010] Input the initial merged image into the bidirectional interaction encoder. Based on the bidirectional interaction mechanism, use a multi-scale feature enhancement module and a high-low frequency feature decomposition module to perform feature enhancement to obtain multi-scale fusion features; input the multi-scale fusion features into the two-branch decoder for fusion, and respectively reconstruct the infrared and visible light image features to obtain fusion features;

[0011] Input the fusion features into the discriminator, and train the model based on a two-stage training strategy for authenticity judgment; when the loss function converges, the training of the model is completed;

[0012] Input the merged image of the infrared image and the visible light image to be fused into the trained fusion feature generation adversarial model to obtain a fused image.

[0013] In a second aspect, the present invention provides an image fusion system based on two-stage adversarial training and edge perception, including:

[0014] A merging module, which is used to select an infrared-visible light sample image pair dataset, and perform channel merging on the infrared images and visible light images in the dataset to obtain an initial merged image;

[0015] A fusion model construction module, which is used to construct a fusion feature generation adversarial model, and the model includes a bidirectional interaction encoder, a two-branch decoder, and a discriminator;

[0016] A generator, which is used to input the initial merged image into the bidirectional interaction encoder. Based on the bidirectional interaction mechanism, use a multi-scale feature enhancement module and a high-low frequency feature decomposition module to perform feature enhancement to obtain multi-scale fusion features; input the multi-scale fusion features into the two-branch decoder for fusion, and respectively reconstruct the infrared and visible light image features to obtain fusion features;

[0017] A discriminator, which is used to obtain the fusion features and train the model based on a two-stage training strategy for authenticity judgment; when the loss function converges, the training of the model is completed;

[0018] A fusion module, which is used to input the merged image of the infrared image and the visible light image to be fused into the trained fusion feature generation adversarial model to obtain a fused image.

[0019] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in an image fusion method based on two-stage adversarial training and edge perception described in the first aspect are implemented.

[0020] In a fourth aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in an image fusion method based on two-stage adversarial training and edge perception described in the first aspect are implemented.

[0021] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0022] (1) Through an improved generative adversarial network (GAN) structure, the present invention realizes a high-quality infrared and visible light image fusion method, demonstrating significant technical advantages and effects. By adopting a two-stage alternating training mechanism, namely the adversarial training stage and the reconstruction training stage, the model gradually optimizes the feature extraction and image reconstruction capabilities, ensuring that the fused image reaches a high-quality level in both overall structure and detail performance. The generator integrates an encoder, a decoder, a high-low frequency feature decomposition module (HLFD module), and a multi-scale feature enhancement module (MEEM module). These modules work together to effectively enhance the feature representation ability, enabling the fused image to have better detail clarity and overall consistency.

[0023] (2) Through the joint optimization of multiple loss functions such as adversarial loss, identity preservation loss, edge feature loss, gradient reconstruction loss, and mean square error loss, the generated image not only maintains consistency with the target image at the pixel level but also has high quality in terms of edge details and global vision. The model is trained using a diverse dataset, including infrared-visible light image pairs and natural images. The combined use of this dataset enables the generator to have strong generalization ability and exhibit good fusion effects and detail performance in various scenarios. In addition, by using the Adam optimizer combined with the StepLR learning rate dynamic adjustment strategy, the stability of the training process and the convergence speed of the model are effectively improved. Finally, the technical effect of the present invention is to significantly improve the edge features, detail quality, and overall visual effect of the generated image, ensure the effective fusion of multi-modal information, have strong generalization ability and stability, and provide an efficient and practical solution for the field of multi-modal image fusion.

[0024] The advantages of the additional aspects of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments of the invention and their descriptions are used to explain the invention and do not constitute a limitation to the invention.

[0026] Figure 1 The main flowchart of an image fusion method based on two-stage adversarial training and edge perception provided by an embodiment of the present invention;

[0027] Figure 2 The overall flowchart of an image fusion method based on two-stage adversarial training and edge perception provided by an embodiment of the present invention;

[0028] Figure 3 The schematic diagram of a bidirectional interactive encoder provided by an embodiment of the present invention;

[0029] Figure 4 The schematic diagram of a primary feature extraction module provided by an embodiment of the present invention;

[0030] Figure 5 The schematic diagram of a multi-scale feature enhancement module provided by an embodiment of the present invention;

[0031] Figure 6 The schematic diagram of a high-low frequency feature decomposition module provided by an embodiment of the present invention;

[0032] Figure 7 The schematic diagram of a dual-branch decoder provided by an embodiment of the present invention;

[0033] Figure 8 The schematic diagram of a discriminator provided by an embodiment of the present invention;

[0034] Figure 9 The flowchart of a discriminator provided by an embodiment of the present invention;

[0035] Figure 10 The overall structural diagram of a training process provided by an embodiment of the present invention;

[0036] Figure 11 The schematic diagram of an encoder training module provided by an embodiment of the present invention;

[0037] Figure 12 The schematic diagram of a decoder training module provided by an embodiment of the present invention;

[0038] Figure 13 The schematic diagram of the effect display provided by an embodiment of the present invention. Detailed implementation manners

[0039] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0040] Embodiment 1

[0041] AsFigure 1 As shown in Figure 1 , this embodiment discloses an image fusion method based on two-stage adversarial training and edge perception, including the following steps:

[0042] S1: Select an infrared-visible light sample image pair dataset, and merge the infrared images and visible light images in the dataset to obtain an initial merged image;

[0043] S2: Construct a fusion feature generation adversarial model, where the generator in the model includes a bidirectional interactive encoder and a dual-branch decoder;

[0044] S3: Input the initial merged image into the bidirectional interactive encoder for feature enhancement to obtain multi-scale enhanced features; input the multi-scale enhanced features into the dual-branch decoder for fusion to obtain fusion features;

[0045] S4: Input the fusion features into the discriminator for authenticity judgment. Based on the discrimination results, use a loss function to alternately train the generator and the discriminator. When the loss function converges, complete the training of the model;

[0046] S5: Input the merged image of the infrared image and the visible light image to be fused into the trained fusion feature generation adversarial model to obtain the fused image.

[0047] Next, in combination with Figure 2 , a detailed description will be given of an image fusion method based on two-stage adversarial training and edge perception disclosed in this embodiment.

[0048] I. Data Input and Preprocessing

[0049] The data input and preprocessing module of the present invention aims to preprocess the original input infrared and visible light images to ensure the effectiveness of subsequent feature extraction and fusion steps. The specific steps of data input and preprocessing are as follows:

[0050] (I) Input Image Data

[0051] Obtain infrared images and visible light images respectively. The infrared images have unique thermal information and can effectively display the contours and thermal radiation of objects when visible light conditions are insufficient. The visible light images are used to provide detailed texture and color information of the target object. Both types of images come from different spectral acquisitions with the same field of view and may have different spatial resolutions when input.

[0052] (II) Data Preprocessing Steps

[0053] 1. Grayscale Processing

[0054] First, for the visible light image I VISPerform grayscale processing to convert it from a three-channel RGB image to a single-channel grayscale image. The pixel values of the grayscale image are calculated by the following formula:

[0055]

[0056] Among them, respectively represent the red, green, and blue channel values at position (x, y) in the visible light image, represents the pixel value of the grayscale image. The processed grayscale image is denoted as

[0057] 2. Size adjustment

[0058] Since the sizes of the input images may be different, in order to unify the sizes and adapt to subsequent feature extraction operations, it is necessary to adjust the sizes of the infrared image and the grayscale visible light image.

[0059] Unify the infrared image I IR and the grayscale visible light image to the size of H×W. In this embodiment, H = 128 and W = 128 are set.

[0060]

[0061] The size adjustment is achieved through bilinear interpolation to ensure that the details of the image are retained as much as possible during the scaling process.

[0062] 3. Normalization processing

[0063] The infrared image and the visible light image that have undergone grayscale processing and size adjustment need to be further normalized to improve the stability of the input data and the convergence speed of training. Normalize the pixel values of the image to the interval [-1, 1]. The normalization formula is:

[0064]

[0065] Among them, I I ′ R , I′ VIS ∈[-1, 1] H×W×1 . This can make the mean value of the input image close to 0 and avoid training instability caused by too large a range of different input data.

[0066] 4. Data alignment and merging

[0067] To ensure that the infrared and visible light images have consistent spatial information, the normalized infrared image and visible light image are merged in the channel dimension to obtain the initial merged image I input of the model input data:

[0068] Iinput = Concat(I I ′ R , I′ VIS );

[0069] Among them, the merged input data contains two channels: one is infrared information and the other is visible light information, and B is the number of batches; this data will be input into the encoder for feature extraction.

[0070] II. Bidirectional Interactive Encoder

[0071] As Figure 3 shown, the main purpose of the bidirectional interactive encoder is to extract high-quality multi-scale features from the input merged data (infrared and visible light images), enhance feature details and texture information, so as to support the generation of high-quality fused images by the decoder.

[0072] To achieve this goal, the bidirectional interactive encoder includes a primary feature extraction module, a high-low frequency feature decomposition module (HLFD module) and a multi-scale feature enhancement module (MEEM module). Through the close cooperation between these modules, the encoder can gradually extract and enhance image features.

[0073] (1) Overall Architecture of the Encoder

[0074] The bidirectional interactive encoder described in this embodiment adopts an innovative bidirectional interactive design, specifically the interaction between the HLFD module and the MEEM module, aiming to perform multi-scale feature extraction and enhancement on the input infrared and visible light images. Through the cooperation and bidirectional interaction between the modules, the encoder forms a closed-loop optimization path, thereby improving the feature expression ability and the quality of the fused image.

[0075] 1. Primary Feature Extraction Module

[0076] First of all, the input features are subjected to preliminary convolution processing through the primary feature extraction module to obtain a preliminary multi-scale feature representation.

[0077] The primary feature extraction module is used to extract preliminary features from the input initial merged image, and extracts feature representations of different scales and levels through gradual downsampling and feature enhancement.

[0078] As a specific implementation method, the primary feature extraction module includes three first convolution blocks, second convolution blocks and third convolution blocks that are the same in structure and connected in sequence. The structure includes a convolution layer with a convolution kernel size of 3×3, a stride of 2, and a padding of 1, a batch normalization layer and a non-linear activation layer, as Figure 4 shown.

[0079] First of all, the initial merged image Enter the first convolutional block, and perform preliminary feature extraction and downsampling operations through the convolutional layer.

[0080]

[0081] Among them, the output of the convolutional operation The number of output channels is 128, and the size of the feature map is reduced to half of the original. Next, apply the batch normalization (Batch Normalization) operation to the convolutional output to standardize the feature values, reduce the internal covariate shift during training, and improve the stability of the model. The output of batch normalization is:

[0082]

[0083] Use the Leaky ReLU activation function to perform non-linear mapping on the normalized features. The negative slope of the activation function is 0.2, which is used to retain some negative value feature information to avoid sparsity. The activated features are represented as:

[0084]

[0085] The size of the output feature map is Containing preliminary spatial and texture features.

[0086] Take the output of the first convolutional block As the input, perform further feature extraction and downsampling through the convolutional layer of the second convolutional block:

[0087]

[0088] Output features The number of channels is 64, and the size of the feature map is one-fourth of the input. Use the batch normalization layer to standardize the output features:

[0089]

[0090] Perform non-linear mapping on the normalized features using the Leaky ReLU activation function:

[0091]

[0092] The size of the output feature map is

[0093] Take the output of the second convolutional block As the input, perform further downsampling and deep feature extraction through the convolutional layer of the third convolutional block:

[0094]

[0095] Output features The number of channels is reduced to 32, and the spatial dimension of the feature map is further reduced to one-eighth of the original. Apply batch normalization to the output features:

[0096]

[0097] Use the Leaky ReLU activation function to perform non-linear mapping on the features:

[0098]

[0099] The size of the output feature map is of the initial multi-scale features.

[0100] The primary feature extraction module of this embodiment realizes multi-level abstraction and enhancement of the input features through step-by-step downsampling, non-linear mapping, and feature channel changes. In each convolutional block, use a convolutional operation with a stride of 2 to downsample the feature map, effectively reducing the feature size and capturing a larger range of context information. After convolution, introduce the Leaky ReLU activation function (negative slope of 0.2) for non-linear enhancement, retain negative value feature information to avoid sparsification, and ensure that more feature information can be passed to the next layer. In terms of feature channels, layer-by-layer feature extraction gradually increases and then decreases the number of channels. Such a design fully captures local details in the initial stage and gradually compresses features in the subsequent stage to extract high-level semantic information, thus effectively balancing the expression of feature details and overall structure.

[0101] 2. Multi-scale Feature Enhancement Module (MEEM Module)

[0102] Input the initial multi-scale features into the MEEM module for multi-scale edge enhancement.

[0103] As Figure 5 shown, the MEEM module performs convolutional operations on the feature map in parallel with different convolutional kernel sizes (such as 3×3, 5×5, 7×7) to extract edge information of different scales respectively. Then, through a fusion mechanism, fuse these edge features of different scales to obtain multi-scale edge enhanced features Generally expressed as:

[0104]

[0105] The goal of the Multi-Scale Edge Enhancement Module (MEEM) is to perform multi-stage and multi-scale enhancement on the fused features to further enhance the edge and detail information of the features. This module adopts a progressive multi-stage feature enhancement structure to capture the multi-scale attributes of the input features and enhance the edge information.

[0106] As a specific implementation, the enhancement process of this module consists of feature enhancement units with multi-stage progression. Each unit performs multi-scale processing and enhancement on the input features through convolution operations, pooling operations, and edge enhancement operations. The specific steps are as follows:

[0107] The input features of the multi-scale feature enhancement module come from the features extracted by the primary module, denoted as

[0108] First, perform initial feature transformation. The feature transformation unit with a 1×1 convolutional layer structure compresses the channels of the input features to reduce the dimension and computational complexity of the features. The convolution operation formula is:

[0109]

[0110] where C2 is the number of compressed channels. This represents the features after preliminary convolution and is passed as input to the subsequent multi-scale processing stage.

[0111] After that, perform cascaded enhancement on the transformed features. The cascaded enhancement unit contains L cascaded feature enhancement stages. The specific steps of each stage are as follows:

[0112] (1) Feature pooling: Perform average pooling operation on the input transformed features to extract the low-frequency components in the features to capture more context information. The feature pooling in the first stage is expressed as:

[0113]

[0114] Output feature

[0115] If there are high-frequency features with HLFD feedback then

[0116]

[0117] The output feature remains and the subsequent process remains unchanged.

[0118] Specifically, in this embodiment, high-frequency feature adaptation and adaptive adjustment are introduced. Before performing multi-scale feature enhancement, based on the two-way interaction mechanism, if there are high-frequency enhancement features with HLFD feedback then calculate the weight coefficient G through the Gate Network MEEM , to achieve adaptive adjustment of high-frequency information. Due to the high-frequency features with HLFD feedback Different from the number of channels of the internal features of MEEM, to ensure the stability of fusion, channel matching is first performed through 1×1 convolution:

[0119]

[0120] MEEM uses a gating network to calculate the first adaptive fusion weight G MEEM , to adjust the contribution degree of high-frequency information to the initial features of MEEM:

[0121]

[0122] Among them, σ(·) is the Sigmoid activation function, ensuring that the gating coefficient G MEEM is within the range of [0,1].

[0123] The calculated gating coefficient G MEEM is used as a modulation factor to perform adaptive weighted fusion on high-frequency features, obtaining the first weighted fusion feature:

[0124]

[0125] (2) Convolutional feature transformation: Perform 1×1 convolution operation on the pooled features to further compress and transform the feature representation:

[0126]

[0127] Convolutional output feature

[0128] (3) Edge enhancement operation (Edge Enhancer): The purpose of edge enhancement is to enhance the details and edge information in the features. This operation combines the high-frequency part of the image (e.g., details such as edges and textures) with the low-frequency part of the image (e.g., background, global structure), thereby enhancing the edges and details in the image. In the specific implementation, the EdgeEnhancer module uses the input features and extracts edge information through convolution and pooling operations, and further strengthens these edge features, making the edges of the image clearer.

[0129]

[0130] Among them, the EdgeEnhancer module consists of a 3×3 convolutional layer and a non-linear activation function, which is used to enhance the edge attributes of the input features.

[0131] (4) Progressive feature aggregation: The output features of each stage are added to the output features of the previous stage, thereby achieving progressive feature enhancement:

[0132]

[0133] This progressive feature aggregation method ensures that the enhanced features at each stage can directly affect the final output features, enabling each enhancement to cumulatively improve the feature expression ability.

[0134] After the feature enhancement through all L stages, the output features of each stage are concatenated in the channel dimension to obtain the final multi-scale edge enhancement feature representation:

[0135]

[0136] The finally output multi-scale edge enhancement feature This feature integrates multi-scale information from each edge enhancement stage, providing rich and detail-enhanced features for the decoder module.

[0137] In the multi-scale feature enhancement module of this embodiment, features are enhanced progressively through multiple cascaded enhancement stages. Each stage includes convolution, pooling, and edge enhancement operations, ensuring the effective extraction and strengthening of feature information at different scales. In each stage, using the features from high-frequency decomposition, the EdgeEnhancer module further strengthens the edges and details, enabling the enhanced features to retain more rich detail information, thereby improving the quality of the finally generated image. Finally, by concatenating the enhanced features of each stage in the channel dimension, the effective fusion of multi-scale features is achieved, such that the output features contain both detailed local information and multi-scale global feature expressions, providing a more discriminative feature representation. These features will be evaluated for quality by the discriminator and guide the encoder to continuously optimize the feature extraction process.

[0138] 3. High-frequency and low-frequency feature decomposition module

[0139] The output results of the primary feature extraction module and the multi-scale edge enhancement feature output results are input into the HLFD module for high-frequency and low-frequency decomposition processing. The HLFD module decomposes the feature map into a high-frequency part and a low-frequency part through average pooling and differential operations, generally expressed as:

[0140]

[0141] where the high-frequency part F high_freq retains the edge and detail information in the image, while the low-frequency part F low_freq represents the overall structure information of the image. This decomposition can independently process the detail and overall information, thereby more precisely expressing different feature levels of the input image.

[0142] As a specific implementation, the main purpose of the high-low frequency feature decomposition module is to decompose the multi-scale edge enhancement features into low-frequency and high-frequency components, process and enhance each component separately, and finally fuse them to obtain more expressive features. This module decouples the features, enabling the low-frequency part to retain the overall information and the high-frequency part to capture details and edge features, such as Figure 6 as shown.

[0143] First, high-low frequency feature decomposition is performed. For extracting low-frequency features, the preliminary multi-scale features are downsampled through an average pooling operation to extract the low-frequency components. The formula for extracting low-frequency features is:

[0144]

[0145] where AvgPool(·) represents the average pooling operation, the pooling kernel size is k = 2, and the stride is s = 2, which is used to reduce the size of the feature map. The low-frequency feature retains the overall structure and global information of the input features.

[0146] For extracting high-frequency features, they are obtained by subtracting the upsampled low-frequency features from the preliminary multi-scale features to retain detail and edge information. The calculation formula for high-frequency features is:

[0147]

[0148] where Upsample(·) represents the upsampling operation, the stride is s = 2, and bilinear interpolation is used to restore the low-frequency features to the size of the original features, that is The high-frequency features obtained in this way mainly retain the details, edges, and fast-changing feature information of the input image.

[0149] Then, feature enhancement is performed. For the low-frequency feature F low further enhancement processing is carried out to extract and strengthen the global structure information. Low-frequency enhancement uses a convolutional module UNetConvBlock, which contains two convolutional layers in series. The operation of the first convolution is:

[0150]

[0151] where the convolutional kernel size of the convolutional layer is 3×3, the stride is 1, and the padding is 1. The output feature Then, batch normalization and activation are performed on the convolutional result:

[0152]

[0153] Then, the enhanced features are further processed using a second 3×3 convolutional layer:

[0154]

[0155] Output low-frequency enhanced features After enhancement through two convolutional layers, the overall structure and texture features of the low-frequency information are retained.

[0156] For high-frequency feature F high It is enhanced through a DenseBlock to strengthen the detail and edge information. The DenseBlock consists of multiple cascaded convolutional layers, and the input of each layer contains the outputs of all previous layers. Specifically, the input-output relationship of the l-th convolutional layer is:

[0157]

[0158] where Concat(·) represents the concatenation operation in the channel dimension, The convolutional kernel size of each layer is 3×3, the stride is 1, and the padding is 1. Through the cascaded connection method, the DenseBlock can make full use of the feature information of all previous convolutional layers, enhancing the diversity of features and the ability to describe details. Through the enhancement operations of multiple layers in the DenseBlock, high-frequency enhanced features are obtained

[0159] In this embodiment, based on the bidirectional interaction mechanism, when the high-low frequency feature decomposition module (HLFD) receives the edge features provided externally it will first perform channel adaptation and spatial size alignment on the edge features to ensure effective fusion with the current high-frequency features. Specifically, if the number of channels of is inconsistent with the number of high-frequency channels processed by this module, first perform channel mapping through a 1×1 convolution;

[0160]

[0161] If the spatial resolution of the edge features does not match the input features, bilinear interpolation is used to adjust them to the same size:

[0162]

[0163] After completing the channel and spatial alignment, the HLFD module concatenates the high-frequency enhanced features and the edge features and inputs them into the Gate Network. The Gate Network generates an adaptive gating coefficient through convolution and activation functions to weight and adjust the high-frequency enhanced features. Specifically, the high-frequency enhanced features and the edge features are concatenated and input into the Gate Network to generate a gating coefficient G∈[0,1]; subsequently, the high-frequency enhanced features are multiplied by this coefficient, so as to retain the global high-frequency information while strengthening or suppressing the local details in combination with the edge features.

[0164]

[0165] Using the calculated second adaptive fusion weight G, the high-frequency features and enhanced features are weighted and fused to achieve adaptive adjustment. The fused second weighted fusion feature F enhanced is calculated by the following formula:

[0166]

[0167] where ⊙ represents the element-wise pointwise multiplication operation, and the gating coefficient G is used to dynamically adjust the intensity of feature fusion to ensure that the output features retain both the clarity of details and stability in the global structure.

[0168] Thus, by fusing the external edge features with the high-frequency features, the HLFD module can further adaptively correct the detailed information of the high-frequency part on the basis of high-low frequency decomposition, retaining both the overall structure of the image and improving the accuracy of key details and contours using edge features, providing higher-quality feature inputs for subsequent high-frequency enhancement and feature fusion.

[0169] Finally, the low-frequency enhanced features and high-frequency enhanced features are fused. The enhanced low-frequency enhanced features are restored to the size of the original features through an upsampling operation:

[0170]

[0171] to obtain the output features The upsampled low-frequency enhanced features and high-frequency enhanced features are concatenated in channels to obtain a fused feature representation:

[0172]

[0173] where The fused features contain enhanced global information and edge detail information, providing more expressive features for the processing of subsequent modules.

[0174] In this embodiment, the bidirectional interaction encoder establishes a bidirectional interaction mechanism between the multi-scale feature enhancement module (MEEM module) and the high-low frequency decomposition module (HLFD module) to further enhance the features output by the primary feature extraction module. Specifically: First, it enters the MEEM module to obtain the edge information of the first step. Then, the edge information and the output of the initial feature extraction module are input into the HLFD module together. Then, the input features are decomposed into high and low frequencies. The high-frequency part is enhanced through a dense connection block and then passed to the MEEM module again. The MEEM module uses this high-frequency information for edge enhancement operations and feeds back the enhanced edge features to the HLFD module to guide its subsequent feature decomposition and fusion process. This bidirectional interaction forms a closed-loop optimization process: The MEEM module provides the initial edge information and uses a gated weight network to guide the HLFD module to better utilize the high-frequency detail information, which is then used as the input for the next round of the MEEM module. The MEEM module then generates more accurate edge features based on this information, and these edge features can help the HLFD module achieve better feature decomposition in the next iteration.

[0175] Through this bidirectional interaction mechanism, the two modules can promote each other and optimize the feature representation in a progressive manner: The HLFD module focuses on extracting high-quality high-frequency details while maintaining the global structural information, while the MEEM module focuses on enhancing the edge information in these high-frequency details. This synergistic effect ensures that the features are fully enhanced in both detail performance and structural integrity.

[0176] This bidirectional interaction mechanism of the codec integration ensures that the features can maintain effective enhancement of details and structure throughout the entire process from extraction to reconstruction, thus producing a higher-quality fusion result.

[0177] III. Dual-branch Decoder

[0178] In the decoder, the bidirectional interaction mechanism is also continued and strengthened. Specifically:

[0179] Each branch of the decoder contains a high-low frequency processing unit, which inherits the same structure as the HLFD module in the encoder. During the processing, the high-frequency part of the high-low frequency processing unit will interact with the edge-enhanced features output by the MEEM module. This interaction is manifested as: When the high-frequency features are enhanced through a dense connection block, they are modulated using the edge-enhanced information from the MEEM module at the same time, that is:

[0180] This design enables the decoder to: maintain sensitivity to details during the reconstruction process, ensuring that the high-frequency information and edge features extracted by the encoder are not lost during the reconstruction process, and through the dual-branch structure and independent parameters for each branch, the two branches can focus on reconstructing features of different scales and types respectively, thus achieving more comprehensive feature recovery. With continuous interaction with the MEEM module, the edge details are continuously optimized and refined during the upsampling and reconstruction processes.

[0181] As Figure 7 shown, the dual-branch decoder module of this embodiment includes two decoding branches with the same structure but independent parameters, which are respectively used for infrared feature reconstruction and visible light feature reconstruction. Through this design, the two branches can respectively maintain and enhance the feature expressions of their respective modalities, and finally obtain a high-quality fused image that retains both infrared thermal radiation information and visible light details through the weighted fusion method.

[0182] Through independent parameter learning and optimization processes, the dual-branch decoder enables one branch to focus on learning and reconstructing the thermal radiation features of the infrared image, and the other branch to focus on learning and reconstructing the texture detail features of the visible light image. During the training process, by applying targeted reconstruction losses to each branch respectively (such as comparing the output of branch1 with the infrared image and the output of branch2 with the visible light image), the two branches are guided to learn the feature expressions of the corresponding modalities respectively.

[0183] (I) Dual-branch decoder structure

[0184] Each branch consists of the following components:

[0185] 1. Initial feature conversion layer

[0186] The initial feature conversion layer of each branch consists of a convolutional layer, a batch normalization layer, and an activation function. This layer is used to reduce the input feature map from 256 channels to 128 channels to provide a more suitable feature expression for subsequent feature processing. Our encoder input contains And F edge.feat . The first skip connection feature comes from the output feature of the first stage of the encoder (features[0]), that is, F fused ; the second skip connection feature comes from the output feature of the second stage of the encoder (features[1]), that is, F enhanced , F edge.feat is the edge information obtained by the first MEEM module processing. First, the input feature multi-scale edge enhancement special module output feature is input into the initial feature conversion layer of the decoder module to obtain the initial conversion feature:

[0187]

[0188] Among them, is the initial conversion feature of each branch.

[0189] 2. High and low frequency processing unit

[0190] Initial conversion feature The initial conversion feature passes through the high and low frequency processing unit to further decompose and enhance the feature. In this unit, the feature is processed by the high and low frequency decomposition module respectively.

[0191] Low frequency feature processing: Obtain the low frequency conversion feature through average pooling operation:

[0192]

[0193] For the low frequency conversion feature Perform further enhancement processing, and use UNetConvBlock for convolution processing to maintain global structural information:

[0194]

[0195] For extracting the high frequency conversion feature, it is obtained by subtracting the upsampled low frequency conversion feature from the initial conversion feature to retain the detail and edge information. The calculation formula of the high frequency conversion feature is:

[0196]

[0197] For the high frequency conversion feature Use the DenseBlock for enhancement.

[0198] During the enhancement process, the high frequency feature Enters the HLFD module for high and low frequency decomposition, and combines with the edge enhancement feature F edge.feat from the MEEM module for auxiliary enhancement. The HLFD module further extracts the details of the high frequency feature through the DenseBlock, and finally obtains the enhanced high frequency feature:

[0199]

[0200] Among them, F edge.feat is the edge information calculated by MEEM, which selectively enhances the high frequency feature through the gating mechanism, so that the final high frequency enhanced feature has clearer edge and texture information.

[0201] 3. Feature fusion

[0202] The enhanced low frequency conversion feature and high frequency conversion feature are integrated through the feature fusion layer, and the fusion operation is concatenation in the channel dimension:

[0203]

[0204] Among them,

[0205] 4. Intermediate Feature Fusion Layer

[0206] After feature fusion, a 3×3 convolutional layer is used to integrate the fused features with the skip connection features of the corresponding layer in the encoder. The fusion process can be described as:

[0207]

[0208] Among them, is the corresponding stage feature extracted from the encoder, and Concat(·) represents concatenation in the channel dimension. The skip connection features include: the first skip connection feature comes from the output feature of the first stage of the encoder (features[0]), that is The second skip connection feature comes from the output feature of the second stage of the encoder (features[1]), that is When i = 1, it refers to the 64-channel feature output in the first stage (features[0]); when i = 2, it refers to the 128-channel feature output in the second stage (features[1]).

[0209] Subsequently, the fused features are further processed through convolution and batch normalization:

[0210]

[0211] Upsampling module: The spatial size of the feature map is restored through the upsampling module, and the size of the feature map is doubled using bilinear interpolation:

[0212]

[0213] The size of the upsampled feature is

[0214] 5. Fusion Result

[0215] At the end of each branch, a 1×1 convolutional layer is used to reduce the number of feature channels to 1, thereby generating a single-channel fusion result:

[0216]

[0217] The final output is a single-channel feature map, which is used for the output of the fusion branch.

[0218] (2) Working Process of the Dual-Branch Decoder

[0219] 1. First Feature Reconstruction Stage

[0220] Two branches simultaneously receive the feature maps from the encoder as inputs.

[0221] The feature dimension is reduced through the initial feature conversion layer.

[0222] The high - and low - frequency processing unit is used to enhance and reconstruct the features. During the process, it interacts with the edge enhancement information of the MEEM module to further strengthen the high - frequency details.

[0223] Channel concatenation is performed with the first skip - connection feature of the encoder to retain global and detailed information. Among them, the first skip - connection feature is the 64 - channel feature (features[0]) output in the first stage to retain the global structure information.

[0224] 2. Second Feature Reconstruction Stage

[0225] The output of the first stage is upsampled to restore part of the spatial resolution.

[0226] The second high - and low - frequency decomposition and processing are performed to enhance the detailed and structural information.

[0227] Channel concatenation is performed with the second skip - connection feature. Among them, the second skip - connection feature is the 128 - channel feature (features[1]) output in the second stage to retain the detailed texture information.

[0228] Finally, the stage output features are generated through a series of convolutional layers.

[0229] 3. Final Fusion Stage

[0230] Weighted fusion is performed on the final outputs from the two branches:

[0231]

[0232] The fused features are normalized through the Tanh activation function to obtain the final fused features:

[0233] I fusion_output = Tanh(F final 0;

[0234] In this embodiment, through the dual - branch decoder structure, combined with the collaborative effects of the high - and low - frequency decomposition module (HLFD module) and the multi - scale feature enhancement module (MEEM module), the quality of the fused image is significantly improved during the image reconstruction process.

[0235] Although the dual-branch decoder adopts the same network structure, through independent parameter optimization and feature processing paths, it realizes the reconstruction of features at different levels. Specifically, each branch processes through the HLFD module twice to achieve the gradual optimization of features at different scales. During the feature processing, edge-enhanced features are continuously utilized to enhance details, and at the same time, through phased feature channel compression (from 256 channels to 128 channels, and then to 64 channels), the fine reconstruction of features is gradually realized.

[0236] In terms of feature transfer between the encoder and decoder, this embodiment makes full use of the skip connection features of two different scales from the encoder. During the feature fusion process, complete feature information is retained through concatenation operations in the channel dimension, and bilinear interpolation is used for feature upsampling to ensure a smooth transition in the size of the feature map. This design effectively avoids information loss during feature reconstruction and provides the necessary feature support for high-quality reconstruction.

[0237] In terms of collaborative enhancement between modules, the HLFD module provides effective feature feedback by returning high-frequency features, while the edge features generated by the MEEM module run through the entire decoding process to continuously guide the reconstruction of details. Finally, through feature addition and Tanh activation operations, the effective fusion of information from the two branches is achieved. This simple and effective fusion strategy makes full use of the complementary features learned by the two branches respectively, ensuring the quality of the fusion result.

[0238] Through the organic combination of the above mechanisms, this embodiment not only ensures the structural integrity of the reconstructed image but also realizes the precise restoration of details, thus obtaining a high-quality fusion result. This design realizes multi-level extraction and optimization of features while keeping the model structure simple.

[0239] IV. Discriminator

[0240] The fused image discriminator module is used to discriminate the authenticity of the generated fused image. The discriminator includes a dual-head self-attention module, a layer-by-layer convolutional module, and a fully connected and discriminant output module, as Figure 8 shown, and the discrimination process is as Figure 9 shown.

[0241] Among them, the dual-head self-attention module enables the discriminator to dynamically focus on key regions of features, enhancing the ability to capture high-frequency details and global structures; the layer-by-layer convolutional module processes the enhanced features using multiple convolutional blocks, gradually downsampling, and abstracting high-level feature representations to enhance the discriminator's discrimination ability for input images; the fully connected and discriminant output module maps the extracted features to a scalar output through a fully connected layer, which is used to represent the discriminator's confidence in the "authenticity" of the input image.

[0242] (1) Dual-Head Self-Attention Module (DHSA):

[0243] By introducing the Dual-Head Self-Attention Module (DHSA), the discriminator can enhance its feature representation ability and better capture detailed information and global context features.

[0244] After the input layer of the discriminator, the Dual-Head Self-Attention Module (DHSA) is used to enhance the input features. The role of this module is to dynamically weight the input features through the self-attention mechanism to capture the dependencies between different feature dimensions, thereby improving the feature representation ability.

[0245] Input features: The discriminator receives the final fused features The input features are first enhanced through the DHSA module:

[0246] A = DHSA(I fusion_input )

[0247] Implementation of the attention mechanism: The Dual-Head Self-Attention Module first linearly transforms the input features through the convolutional layer qkv to generate queries, keys, and values:

[0248] qkv = Conv 1×1 (I fusion_imput )

[0249] Next, the qkv is further processed through depthwise separable convolutional operations and then split into five parts q1, k1, q2, k2, v, where q1, k1 are used for high-frequency feature calculation, q2, k2 are used for low-frequency feature calculation, and v is used as a shared value vector. Through feature sorting and matching operations, effective interaction between high- and low-frequency feature dimensions is ensured.

[0250] Attention calculation: After normalizing the queries and keys respectively, the attention score matrix is calculated and normalized through the softmax function:

[0251]

[0252] where d k represents the dimension of the key, and · represents matrix multiplication. Then, the attention score matrix is used to weight and sum the values v to obtain the enhanced features:

[0253] out1 = attn1 × v;

[0254] The same steps are also used for q2 and k2 to calculate attn2 and the corresponding enhanced output out2.

[0255] Output feature: After interacting out1 and out2, project them back to the original feature dimension through the convolutional layer project_out, and fuse the output with the input through the residual connection:

[0256] A′ = I fusion_input + project_out(out1 ⊙ out2);

[0257] where ⊙ represents element-wise multiplication, which is used to fuse two enhanced features. Finally, the output attention feature is used as the input of the subsequent convolutional block.

[0258] The discriminator in this embodiment adds a dual-head self-attention mechanism in the input feature processing stage, and enhances the feature representation through the interaction calculation of queries, keys, and values, so as to be able to capture the complex associations between features more effectively.

[0259] (2) Layer-by-layer convolutional module:

[0260] Based on the enhanced feature A′ by the DHSA module, the discriminator uses multiple convolutional layers with the same structure to gradually extract higher-level features, reduce the spatial dimension of the image, and increase the abstraction of the features.

[0261] Use the first convolutional layer with a kernel size of 3×3, a stride of 2, and a padding of 1 to downsample the input features:

[0262] B (1) = LeakyReLU(BatchNorm(Conv 3×3 (A′)), α = 0.2);

[0263] where α = 0.2 represents the negative slope of Leaky ReLU.

[0264] Use the second convolutional layer with the same structure to further downsample and abstract the features:

[0265] B (2) = LeakyReLU(BatchNorm(Conv 3×3 (B (1) ), α = 0.2);

[0266] Continue to process the features to extract the final discriminant features:

[0267] B (3) = LeakyReLU(BatchNorm(Conv 3×3 (B (2) ), α = 0.2);

[0268] After feature enhancement, the discriminator in this embodiment uses a layer-by-layer convolution module to downsample and abstract features from the input, enabling it to evaluate the authenticity of the input image at multiple scales. Each layer of convolution combines batch normalization and Leaky ReLU activation to ensure the effective flow and enhancement of features.

[0269] (III) Fully Connected and Discriminative Output Module:

[0270] Flatten the feature map output by the convolutional feature extraction layer into a vector:

[0271] C = B (3) ·view(B (3) ·size(0), -1);

[0272] Process the flattened features through two fully connected layers to generate the final discriminative result:

[0273] D (1) = LeakyReLU(Linear(C, 1024), α = 0.2);

[0274] D output = Linear(D (1) , 1);

[0275] where represents the discriminative result of the input image.

[0276] The fully connected and discriminative output module enhances the ability to evaluate the authenticity of generated images by fully combining global context and local detail features.

[0277] In this embodiment, by introducing a dual-head self-attention mechanism (DHSA) into the discriminator, the feature enhancement ability of the discriminator for fused images can be significantly improved. The attention mechanism enables the discriminator to dynamically focus on key regions during feature extraction, effectively capturing detail information and global context, thereby improving the judgment of the authenticity of generated images. The combination of convolutional feature abstraction and fully connected discriminative output ensures that the discriminator can handle both local features and enhance the quality of generated images through global features, ultimately achieving efficient and accurate discrimination of fused images.

[0278] V. Training Process

[0279] This embodiment uses the MS-COCO and LLVIP datasets for training. Among them, MS-COCO provides 12,025 natural images for training the decoder, and the LLVIP dataset provides 16,250 pairs of infrared and visible light image pairs for training the encoder.

[0280] All input images are uniformly adjusted to a size of 128×128, and the training is carried out with a configuration of batch size 12, number of training epochs 5, and learning rate 1e-4 (dynamically adjusted using the StepLR strategy with a decay factor of 0.92).

[0281] (I) Overview of the training process

[0282] The training process alternates between the adversarial training stage and the reconstruction training stage, ensuring the continuous improvement of the global consistency and detail accuracy of the fused images. The overall training process is as Figure 10 shown.

[0283] 1. Adversarial training stage:

[0284] As Figure 11 shown, first, the discriminator is trained. The discriminator receives data from real images and images generated by the generator, and by maximizing the output for real images and minimizing the misjudgment of generated images, it learns to distinguish real features from generated features, thereby calculating the discriminator loss

[0285] After the discriminator training is completed, the generator starts to be trained. Through the joint optimization of the generator adversarial loss identity preservation loss and edge feature loss the quality of the generated images is continuously improved. The main goal of the generator in this stage is to generate images with better fusion effects, making it difficult for the discriminator to distinguish.

[0286] 2. Reconstruction training stage:

[0287] As Figure 12 shown, in the reconstruction training stage, the natural image dataset is used to further optimize the decoder. The main focus is on improving the reconstruction quality of the generated images, especially the restoration of details and the overall consistency of features.

[0288] In this stage, the gradient reconstruction loss and the mean squared error loss (MSE loss) are used for optimization to more finely improve the reconstruction effect of the images, so that the generated fused images not only have a natural feel visually but also can maintain a high level of detail quality

[0289] (II) Setting of the loss function

[0290] During the entire training process, the loss functions of the generator and the discriminator are set as follows:

[0291] 1. Generator loss function

[0292] The generator loss function includes adversarial loss, identity preservation loss, and edge feature loss, with the goal of ensuring that the generated images have high quality both in terms of integrity and details.

[0293]

[0294] Among them, represents the adversarial loss, represents the identity preservation loss, represents the edge feature loss, represents the interaction loss. λ adv , λ idt , λ edge , λ inter represent the weight coefficients of each loss term, which can be set according to the actual situation. Preferably, in this embodiment, they are set to to adjust the influence degree of each loss in the overall loss.

[0295] 2. Adversarial Loss

[0296] The adversarial loss is used to make the images generated by the generator "deceive" the discriminator as much as possible, so that the generated images are visually as close as possible to the real images. It is defined as:

[0297]

[0298] Among them, I fusion_output is the fused image generated by the generator, and D is the discriminator. The goal is to make the discriminator unable to distinguish between the generated image and the real image, thereby improving the quality of the generated image.

[0299] 3. Identity Preservation Loss

[0300] The identity preservation loss is used to ensure that the generated images can retain the important features in the input images, especially when performing multi-modal fusion, to ensure that the main structural information of the images is maintained. The L1 loss is used to calculate the difference between the generated image and the input image:

[0301]

[0302] Among them, I fusion_output is the image generated by the generator, and I input is the input image. The goal of the identity preservation loss is to ensure that the generated image is consistent with the input image in terms of global structure and main features.

[0303] 4. Edge Feature Loss

[0304] The edge feature loss is used to enhance the edge details in the image, ensuring that the generated image is consistent with the real image in terms of high-frequency features. The Laplacian filter is used to extract the edge features of the generated image and the target image, and the L1 loss is calculated:

[0305]

[0306] where, represents the Laplacian filter operation, aiming to retain the details of the image edges. I target is the target image. I fusion_output represents the fused image output by the generator. I target represents the target image, that is, the real image expected to be generated. represents ensuring that the edges of the generated image are consistent with the real image, thereby enhancing the high-frequency details.

[0307] 5. Cross Loss

[0308] The interaction loss is used to evaluate and optimize the feature interaction between the internal modules of the generator, especially the two-way interaction process between the high-low frequency feature decomposition module (HLFD module) and the multi-scale feature enhancement module (MEEM module). Through the optimization of the interaction loss, the overall feature expression and detail performance of the generated image reach the optimal effect.

[0309] The specific definition and calculation of the interaction loss:

[0310] 6. High-Low Frequency Interaction Loss

[0311] In the encoder part of the generator, the HLFD module and the MEEM module form a closed-loop optimization through the mutual transfer and enhancement of high-low frequency features. To ensure the effectiveness of the high-frequency features in this interaction process, the feature interaction loss is introduced to evaluate the fidelity and effectiveness of the high-frequency features after being enhanced by the MEEM module:

[0312]

[0313] where, F enhanced_high represents the high-frequency features enhanced by the MEEM module, while F orig_high represents the high-frequency features originally output from the HLFD module.

[0314] 7. Edge Interaction Loss

[0315] When processing high-frequency features, by calculating the edge feature loss, the effectiveness of the edge feature adaptation and gating mechanism in the interaction process is evaluated. The edge interaction loss is used to evaluate the quality change of the edge features during the transfer between different modules:

[0316]

[0317] Among them, A edge represents the feature gating weight, which is used to adjust the weight of the loss according to the importance of the edge features. F edge_feat represents the enhanced edge features, and F orig_edge represents the original edge features.

[0318] 8. Combination of interaction losses:

[0319] The final interaction loss is composed of the high-frequency and low-frequency interaction losses and the edge interaction loss weighted:

[0320]

[0321] Among them, λ high and λ edge are the weight coefficients of the high-frequency and low-frequency interactions and the edge interaction respectively. By adjusting these coefficients, the intensity of feature interaction can be balanced during training to ensure the best quality of the generated images.

[0322] 9. Discriminator loss function

[0323] The goal of the discriminator is to distinguish real images from generated images, and improve the discriminative ability of the discriminator by maximizing the judgment accuracy of real images and minimizing the misjudgment of generated images. It is defined as follows:

[0324]

[0325] I target represents the samples from real images. I fusion_output represents the fused images generated by the generator. p data represents the distribution of real images. P c represents the distribution of generated images. The goal of the discriminator is to maximize the correct judgment probability of real images and minimize the wrong judgment probability of generated images, thereby promoting the improvement of the generator.

[0326] 10. Loss function in the reconstruction stage

[0327] In the reconstruction training stage, the decoder is mainly optimized, and a natural image dataset is used to improve the reconstruction quality. The following loss function is adopted:

[0328] Gradient reconstruction loss

[0329] The gradient reconstruction loss is used to improve the smoothness and detail transition effect of the generated images. By calculating the gradient difference between the generated images and the target images, it ensures better retention of details:

[0330]

[0331] Mean Squared Error Loss (MSE)

[0332] The MSE loss is used to measure the average squared difference between the generated image and the target image, ensuring the minimization of the overall reconstruction error, thereby generating a more realistic fused image:

[0333]

[0334] (III) Optimization Strategy

[0335] During the entire training process, the Adam optimizer is used to update the model parameters. The initial learning rate is set to 1×10 -4 , and the learning rate is dynamically adjusted through the StepLR learning rate adjustment strategy. The specific steps are as follows:

[0336] Adam optimizer: By setting the initial learning rate and momentum parameters (β1, β2), efficient update of the model parameters is achieved. Ensure the stability and convergence speed of the training process.

[0337] StepLR learning rate adjustment: The learning rate is reduced every fixed number of steps to better adapt to the changes in the model during the training process and avoid model oscillation caused by too large a learning rate in the later stage.

[0338] In this specific embodiment, through the alternating optimization of the adversarial training stage and the reconstruction training stage, the generator can gradually improve its ability to extract and reconstruct the features of the input image. The adversarial stage ensures that the fused image has better overall feature consistency, and the reconstruction stage further optimizes the detail performance, making the generated fused image have high visual quality and realism.

[0339] Adding edge feature loss to the generator loss function ensures the effective retention of edge and detail information in the generated image. Through the dual effects of edge feature loss and gradient reconstruction loss, the edge details of the generated image are significantly enhanced, making it visually more realistic and delicate.

[0340] By using infrared-visible light image pairs in the adversarial training stage and natural image datasets in the reconstruction training stage, the model not only retains multi-modal features but also effectively improves the image reconstruction ability and generalization performance, ensuring the consistency and detail performance of the generated image in various scenarios.

[0341] To verify the effect of the present invention, as Figure 13 shows the specific effect comparison of the method of the present invention in the infrared and visible light image fusion task, including: (a) the original infrared image, (b) the original visible light image, and (c) the fused result image.

[0342] It can be intuitively seen from the figure that the method of the present invention successfully takes into account the advantages of infrared and visible light images during the fusion process, making the fused image have both the target prominence of the infrared image and retain the rich details and natural colors of the visible light image. Compared with single-modal images, the fusion results are superior in terms of brightness uniformity, contrast optimization, and target clarity, ensuring the overall perception and visual comfort of the scene. Especially in the fused image (c), the infrared targets are fully retained, making them still have good visibility at night or in low-light environments, while the details in the visible light image, such as ground textures, building structures, trees, etc., are also restored to a high degree, avoiding the common problems of blurring or over-smoothing in traditional methods. In addition, the overall color and brightness transition of the fused image are natural, avoiding artifacts, over-enhancement, or information loss that may occur in traditional methods, making the final result more in line with the human eye perception habit in terms of visual quality, and providing more reliable data support for subsequent intelligent analysis, target detection, and scene understanding.

[0343] Compared with existing methods, the two-stage adversarial training strategy and edge perception mechanism adopted by the present invention effectively suppress the common problems of detail loss and edge blurring in traditional methods while ensuring the clear visibility of infrared targets. Specifically, this method performs multi-scale decomposition of infrared and visible light information through high-low frequency feature decoupling (HLFD), and further optimizes the fusion of high-frequency information with the help of multi-scale feature enhancement (MEEM) to ensure the clarity of target edges. In addition, combined with a dual-head self-attention discriminator (DHSA), this method can accurately extract and strengthen important details, making the fusion results superior to traditional methods in terms of visual consistency and information retention.

[0344] In summary, Figure 13 it fully verifies the effectiveness and superiority of the method of the present invention, and proves its practical application value in the field of infrared-visible light image fusion.

[0345] Embodiment 2

[0346] This embodiment provides an image fusion system based on two-stage adversarial training and edge perception, including:

[0347] A merging module, configured to select an infrared-visible light sample image pair dataset, and perform channel merging on the infrared image and the visible light image in the dataset to obtain an initial merged image;

[0348] A fusion model construction module, configured to construct a fusion feature generation adversarial model, and the model includes a bidirectional interaction encoder, a dual-branch decoder, and a discriminator;

[0349] A generator for inputting an initial merged image into a bidirectional interactive encoder, enhancing features based on a bidirectional interaction mechanism using a multi-scale feature enhancement module and a high-low frequency feature decomposition module to obtain multi-scale fusion features; inputting the multi-scale fusion features into a dual-branch decoder for fusion, reconstructing infrared and visible light image features respectively to obtain fusion features;

[0350] A discriminator for obtaining the fusion features and training the model for authenticity judgment based on a two-stage training strategy; when the loss function converges, the training of the model is completed;

[0351] A fusion module for inputting the merged image of the infrared image and the visible light image to be fused into the trained fusion feature generation adversarial model to obtain the fused image.

[0352] Embodiment III

[0353] This embodiment provides a computer-readable storage medium with a computer program stored thereon, and when the program is executed by a processor, it implements the steps in a method for image fusion based on two-stage adversarial training and edge perception as described in Embodiment I above.

[0354] Embodiment IV

[0355] This embodiment provides a computer device including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a method for image fusion based on two-stage adversarial training and edge perception as described in Embodiment I above.

[0356] The steps or modules involved in Embodiments II to IV above correspond to those in Embodiment I, and the specific implementation details can be referred to the relevant description part of Embodiment I. The term "computer-readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention. The above are only preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An image fusion method based on two-stage adversarial training and edge perception, characterized in that: include: Select an infrared-visible light sample image pair data set, and merge the channels of the infrared image and the visible light image in the data set to obtain an initial merged image; Constructing a fusion feature generation adversarial model, the model comprising a bidirectional interactive encoder, a dual-branch decoder and a discriminator; The initial merged image is input into the bidirectional interactive encoder. Based on the bidirectional interactive mechanism, the multi-scale feature enhancement module and the high- and low-frequency feature decomposition module are used to enhance the features to obtain multi-scale fusion features. The multi-scale fusion features are input into the dual-branch decoder for fusion to reconstruct the infrared and visible light image features respectively to obtain fusion features. The fused features are input into the discriminator, and the model is trained based on the two-stage training strategy for authenticity judgment; when the loss function converges, the model training is completed; The merged image of the infrared image and the visible light image to be fused is input into the trained fusion feature generation adversarial model to obtain the fused image.

2. The image fusion method based on two-stage adversarial training and edge perception as claimed in claim 1, characterized in that: The step of selecting an infrared-visible light sample image pair data set and performing channel merging on the infrared image and the visible light image in the data set to obtain an initial merged image specifically includes: Obtaining infrared sample images and visible light sample images collected at the same time; The visible light sample image is gray-scaled, the sizes of the gray-scale visible light sample image and the infrared sample image are unified, and after normalization, the two images are merged in the channel dimension to obtain an initial merged image.

3. The image fusion method based on two-stage adversarial training and edge perception as claimed in claim 1, characterized in that: The initial merged image is input into a bidirectional interactive encoder, and based on the bidirectional interactive mechanism, a multi-scale feature enhancement module and a high- and low-frequency feature decomposition module are used to perform feature enhancement to obtain a multi-scale fusion feature; specifically, the following steps are included: The bidirectional interactive encoder includes a primary feature extraction module, a multi-scale feature enhancement module and a high- and low-frequency feature decomposition module; The initial merged image is input into the primary feature extraction module based on the convolutional layer to obtain preliminary multi-scale features; The preliminary multi-scale features are input into the multi-scale feature enhancement module and the high- and low-frequency feature decomposition module respectively, the preliminary convolution features are obtained in the multi-scale feature enhancement module based on the initial feature conversion, the high-frequency enhancement features and the low-frequency enhancement features are obtained in the high- and low-frequency feature decomposition module based on the high- and low-frequency decomposition, and the high-frequency enhancement features are input into the multi-scale feature enhancement module; Based on the two-way interaction mechanism, in the multi-scale feature enhancement module, if the high-frequency enhancement feature fed back by the high- and low-frequency feature decomposition module is received, the first adaptive fusion weight is calculated based on the gating network, the preliminary convolution feature and the high-frequency enhancement feature are fused to obtain the first weighted fusion feature, and the first weighted fusion feature is average pooled; otherwise, the preliminary convolution feature is average pooled; The average pooled features are subjected to convolution transformation, edge enhancement and feature aggregation to obtain multi-scale edge enhancement features; In the high- and low-frequency feature decomposition module, if the multi-scale edge enhancement feature fed back by the multi-scale feature enhancement module is received, the second adaptive fusion weight is calculated based on the gating network, and the high-frequency enhancement feature and the multi-scale edge enhancement feature are fused to obtain the second weighted fusion feature; the second weighted fusion feature and the low-frequency enhancement feature are fused to obtain the multi-scale fusion feature.

4. The image fusion method based on two-stage adversarial training and edge perception as claimed in claim 3, characterized in that: The obtaining of high-frequency enhancement features and low-frequency enhancement features based on high- and low-frequency decomposition in the high- and low-frequency feature decomposition module specifically includes: Downsample the preliminary multi-scale features through average pooling to extract low-frequency features; The high-frequency features are obtained by subtracting the upsampled low-frequency components from the preliminary multi-scale features; The low-frequency features are convolved to obtain low-frequency enhanced features; the high-frequency features are processed through densely connected blocks to obtain high-frequency enhanced features.

5. The image fusion method based on two-stage adversarial training and edge perception as claimed in claim 3, characterized in that: The edge enhancement specifically includes: extracting edge information of features after convolution conversion through convolution and pooling operations, and using the EdgeEnhancer module to enhance the edge features, which is expressed as: in, is the feature after edge enhancement, is the feature after convolution transformation.

6. The image fusion method based on two-stage adversarial training and edge perception as claimed in claim 1, characterized in that: The multi-scale fusion features are input into a dual-branch decoder for fusion, and infrared and visible light image features are reconstructed respectively to obtain fusion features; Specifically include: The dual-branch decoder comprises two decoding branches with the same structure and different parameters; the dual-branch decoder comprises a first feature reconstruction stage and a second feature reconstruction stage; The first feature reconstruction stage includes: two branches simultaneously receiving feature maps from the encoder as input; Reduce feature dimensionality through primary feature transformation layers; The high and low frequency processing units are used to enhance and reconstruct the features. In the process, the edge enhancement information of the multi-scale edge enhancement module is interacted to enhance the high frequency details. Fuse with the output features of the first feature enhancement stage of the encoder to obtain preliminary fused features; The second feature reconstruction stage includes: upsampling the preliminary fusion features to restore part of the spatial resolution; Perform a second high and low frequency decomposition and processing to enhance details and structural information; Fuse with the output features of the second feature enhancement stage of the encoder to obtain enhanced fusion features; The enhanced fusion features of the two branches are weighted fused to obtain the final fusion features.

7. The image fusion method based on two-stage adversarial training and edge perception as claimed in claim 1, characterized in that: The discriminator includes a dual-head self-attention module, a layer-by-layer convolution module, and a full connection and discrimination output module; The fused features are input into the dual-headed self-attention module to obtain the attention fusion features; the attention fusion features are input into the layer-by-layer convolution module for convolution operation to obtain multi-scale fusion features; the multi-scale fusion features are input into the fully connected and discriminant output module, and are used for authenticity judgment by combining the global context and local detail features.

8. An image fusion system based on two-stage adversarial training and edge perception, characterized in that: include: A merging module is used to select an infrared-visible light sample image pair data set, and perform channel merging on the infrared image and the visible light image in the data set to obtain an initial merged image; A fusion model building module, used to build a fusion feature generation adversarial model, the model includes a bidirectional interactive encoder, a dual-branch decoder and a discriminator; A generator is used to input the initial merged image into a bidirectional interactive encoder, and based on the bidirectional interactive mechanism, use a multi-scale feature enhancement module and a high- and low-frequency feature decomposition module to perform feature enhancement to obtain multi-scale fusion features; the multi-scale fusion features are input into a dual-branch decoder for fusion, and infrared and visible light image features are reconstructed respectively to obtain fusion features; The discriminator is used to obtain fusion features and train the model based on the two-stage training strategy for authenticity judgment; when the loss function converges, the model training is completed; The fusion module is used to input the merged image of the infrared image and the visible light image to be fused into the trained fusion feature generation adversarial model to obtain a fused image.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the image fusion method based on two-stage adversarial training and edge perception as described in any one of claims 1 to 7 are implemented.

10. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the image fusion method based on two-stage adversarial training and edge perception as described in any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Infrared and visible light image fusion method based on multi-stage generative adversarial network

    CN116935177A

  • Infrared and visible light image fusion method based on Swin Transform and GAN

    CN117333410A

  • SAR and visible light image lightweight fusion method based on double-branch GAN network

    CN117495690A

  • Infrared and visible light image fusion method for inspection robot

    CN119559471A

  • Lightweight and efficient object segmentation and counting method based on generative adversarial network (GAN)

    GB2618876A

Cited By

  • Unified image fusion method and system based on adaptive distribution difference perception

    CN120410893A

  • Picture reconstruction method, system and device and storage medium

    CN120807284A

  • Image quality improvement method and device for low earth orbit satellite, and storage medium

    CN120852183A

  • Image quality improvement method and device for low earth orbit satellite and storage medium

    CN120852183B

  • Infrared light and visible light image fusion method, system and device based on double-branch cross-domain feature fusion and medium

    CN121032817A