Image fusion method and system based on two-stage adversarial training and edge perception

By employing a two-stage adversarial training and edge-aware image fusion method, utilizing a bidirectional interactive encoder and a dual-branch decoder, combined with multi-scale feature enhancement and high- and low-frequency feature decomposition modules, the quality and edge detail issues in existing infrared and visible light image fusion are resolved, achieving high-quality image fusion results.

CN120163719BActive Publication Date: 2026-02-06TAISHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510321415.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2026-02-06
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

Existing infrared and visible light image fusion methods suffer from problems such as unnatural fusion results, loss of details, computational complexity, unclear edge contours, insufficient contrast, blurred edge details, and insufficient multi-scale perception capabilities due to the simple discriminator structure.

Method used

An image fusion method based on two-stage adversarial training and edge awareness is adopted. By constructing a fusion feature generation adversarial model, including a bidirectional interactive encoder, a dual-branch decoder and a discriminator, feature enhancement is performed using a multi-scale feature enhancement module and a high- and low-frequency feature decomposition module. Multiple loss functions are used for training to optimize the discriminator structure.

Benefits of technology

It significantly improves the quality of fused images, ensuring that the overall structure and detail representation reach a high level. The generator has strong generalization ability, and the generated images perform well in terms of edge features, detail quality and overall visual effect, achieving effective fusion of multimodal information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163719B_ABST
    Figure CN120163719B_ABST
Patent Text Reader

Abstract

The application provides an image fusion method and system based on two-stage adversarial training and edge perception, and belongs to the field of image processing. Method: selecting an infrared-visible light sample image pair dataset, merging the image pairs to obtain an initial merged image; constructing a fusion feature generation adversarial model including a bidirectional interactive encoder, a double-branch decoder and a discriminator; inputting the initial merged image into the bidirectional interactive encoder, enhancing the features through a multi-scale feature enhancement and high-low frequency feature decomposition module, inputting the multi-scale fusion features into the double-branch decoder to fuse and reconstruct infrared and visible light image features, and forming fusion features. Then, the fusion features are inputted into the discriminator, the model is trained according to a two-stage training strategy until the loss function converges, and the trained model is used to fuse the to-be-fused merged image. Through the design of the two-stage training strategy and the introduction of the edge feature enhancement mechanism, fine processing in the frequency domain is realized, and the feature extraction and fusion effect and the quality of the fused image are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and in particular to an image fusion method and system based on two-stage adversarial training and edge perception. BACKGROUND

[0002] With the development of image processing technology, infrared and visible light image fusion plays a crucial role in many important fields such as night vision monitoring, security systems, military reconnaissance, etc. Visible light images can present rich texture details and color information, while infrared images can capture target thermal radiation information in low light environments. The fusion of the two can obtain more comprehensive scene information, helping to improve the accuracy of target detection and scene understanding.

[0003] Current infrared and visible light image fusion methods have various defects: traditional methods based on multi-scale decomposition (such as wavelet transform, pyramid transform, etc.) have unnatural fusion results and loss of details; methods based on sparse representation are computationally complex and difficult to preserve the integrity of the source image; methods based on pulse coupled neural networks have insufficient contrast and blurred edge details in the fused image.

[0004] With the rise of deep learning technology, image fusion methods based on generative adversarial networks (GAN) have become a research hotspot, but they also expose many problems: the single training stage cannot balance feature extraction and image reconstruction quality, resulting in poor fusion results in terms of details and structural integrity; the lack of edge feature processing mechanism makes the edge profile of the fused image unclear; the processing of different frequency components is rough, and high and low frequency information cannot be effectively separated and enhanced. At the same time, the existing discriminator structure is simple, relying only on simple convolutional networks to distinguish features, lacking multi-scale perception ability, limiting the effect of generative adversarial learning; the single reconstruction path used in the decoding stage cannot fully utilize different levels of feature information, limiting the quality of the reconstructed image. SUMMARY

[0005] To solve the above problems, the present application proposes an image fusion method and system based on two-stage adversarial training and edge perception, which realizes fine processing in the frequency domain and optimization of the discriminator structure by designing a two-stage training strategy and introducing an edge feature enhancement mechanism, improving feature extraction and fusion effect, and effectively improving the quality of the fused image.

[0006] To achieve the above purpose, the present application adopts the following technical solutions:

[0007] In a first aspect, the present application provides an image fusion method based on two-stage adversarial training and edge perception, comprising:

[0008] An infrared-visible light sample image pair dataset is selected, and channel merging is performed on the infrared images and the visible light images in the dataset to obtain an initial merged image;

[0009] A fusion feature generative adversarial model is constructed, and the model comprises a bidirectional interaction encoder, a double-branch decoder and a discriminator;

[0010] The initial merged image is input into the bidirectional interaction encoder, and feature enhancement is performed by using a multi-scale feature enhancement module and a high-low frequency feature decomposition module based on a bidirectional interaction mechanism to obtain multi-scale fusion features; the multi-scale fusion features are input into the double-branch decoder for fusion to respectively reconstruct infrared and visible light image features, and fusion features are obtained;

[0011] The fusion features are input into the discriminator, and the model is trained based on a two-stage training strategy for authenticity judgment; when the loss function converges, the training of the model is completed;

[0012] The merged image of the infrared image and the visible light image to be fused is input into the trained fusion feature generative adversarial model to obtain a fused image.

[0013] In a second aspect, the present application provides an image fusion system based on two-stage adversarial training and edge perception, comprising:

[0014] A merging module is configured to select an infrared-visible light sample image pair dataset, and perform channel merging on the infrared images and the visible light images in the dataset to obtain an initial merged image;

[0015] A fusion model construction module is configured to construct a fusion feature generative adversarial model, and the model comprises a bidirectional interaction encoder, a double-branch decoder and a discriminator;

[0016] A generator is configured to input the initial merged image into the bidirectional interaction encoder, perform feature enhancement by using a multi-scale feature enhancement module and a high-low frequency feature decomposition module based on a bidirectional interaction mechanism to obtain multi-scale fusion features; input the multi-scale fusion features into the double-branch decoder for fusion to respectively reconstruct infrared and visible light image features, and obtain fusion features;

[0017] A discriminator is configured to acquire fusion features, and train the model based on a two-stage training strategy for authenticity judgment; when the loss function converges, the training of the model is completed;

[0018] A fusion module is configured to input the merged image of the infrared image and the visible light image to be fused into the trained fusion feature generative adversarial model to obtain a fused image.

[0019] In a third aspect, the present application provides a computer readable storage medium having stored thereon a computer program which, when executed by a processor, implements the steps of the image fusion method based on two-stage adversarial training and edge perception according to the first aspect.

[0020] In a fourth aspect, the present application provides a computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the steps of the image fusion method based on two-stage adversarial training and edge perception according to the first aspect when executing the program.

[0021] Compared with the prior art, the present application has the following beneficial effects:

[0022] (1) The present application realizes a high-quality infrared and visible light image fusion method through an improved generative adversarial network (GAN) structure, which exhibits significant technical advantages and effects. A two-stage alternating training mechanism, i.e., an adversarial training stage and a reconstruction training stage, is adopted, and the model gradually optimizes the feature extraction and image reconstruction capabilities to ensure that the fused image reaches a high-quality level in overall structure and detail performance. The generator integrates an encoder, a decoder, a high-low frequency feature decomposition module (HLFD module), and a multi-scale feature enhancement module (MEEM module), which work collaboratively to effectively enhance the feature representation capability, making the fused image have better detail clarity and overall consistency.

[0023] (2) The present application generates images with high quality in both pixel level and edge details and global vision through the joint optimization of multiple loss functions such as adversarial loss, identity preservation loss, edge feature loss, gradient reconstruction loss, and mean square error loss. The model is trained using a diversified dataset, including infrared-visible light image pairs and natural images, which makes the generator have strong generalization ability and performs well in multiple scenarios. In addition, the use of the Adam optimizer combined with the StepLR learning rate dynamic adjustment strategy effectively improves the stability of the training process and the convergence speed of the model. Ultimately, the technical effect of the present application lies in significantly improving the edge feature, detail quality, and overall visual effect of the generated image, ensuring effective fusion of multi-modal information, having strong generalization ability and stability, and providing an efficient and practical solution for the field of multi-modal image fusion.

[0024] The advantages of the additional aspects of the present application will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0025] The accompanying drawings, which form a part of this specification, are included to provide a further understanding of the application and are incorporated in and constitute a part of this specification. The embodiments of the application, and their

[0026] Figure 1 A main flowchart of an image fusion method based on two-stage adversarial training and edge perception provided for an embodiment of the application;

[0027] Figure 2 A whole flowchart of an image fusion method based on two-stage adversarial training and edge perception provided for an embodiment of the application;

[0028] Figure 3 A schematic diagram of a bidirectional interactive encoder provided for an embodiment of the application;

[0029] Figure 4 A schematic diagram of a primary feature extraction module provided for an embodiment of the application;

[0030] Figure 5 A schematic diagram of a multi-scale feature enhancement module provided for an embodiment of the application;

[0031] Figure 6 A schematic diagram of a high-low frequency feature decomposition module provided for an embodiment of the application;

[0032] Figure 7 A schematic diagram of a double-branch decoder provided for an embodiment of the application;

[0033] Figure 8 A schematic diagram of a discriminator provided for an embodiment of the application;

[0034] Figure 9 A flowchart of a discriminator provided for an embodiment of the application;

[0035] Figure 10 A whole structural diagram of a training process provided for an embodiment of the application;

[0036] Figure 11 A schematic diagram of an encoder training module provided for an embodiment of the application;

[0037] Figure 12 A schematic diagram of a decoder training module provided for an embodiment of the application;

[0038] Figure 13 A schematic diagram of an effect display provided for an embodiment of the application. DETAILED DESCRIPTION

[0039] The application will be further described by examples with reference to the drawings attached.

[0040] Embodiment one

[0041] AsFigure 1 The embodiment shown discloses an image fusion method based on two-stage adversarial training and edge perception, comprising the following steps:

[0042] S1: Selecting an infrared-visible light sample image pair dataset, and merging the infrared images and visible light images in the dataset to obtain an initial merged image;

[0043] S2: Constructing a fusion feature generation adversarial model, wherein the generator in the model includes a bidirectional interactive encoder and a double-branch decoder;

[0044] S3: Inputting the initial merged image into the bidirectional interactive encoder for feature enhancement to obtain multi-scale enhanced features; inputting the multi-scale enhanced features into the double-branch decoder for fusion to obtain fusion features;

[0045] S4: Inputting the fusion features into the discriminator for authenticity judgment, based on the judgment result, using a loss function to alternately train the generator and the discriminator, and when the loss function converges, the training of the model is completed;

[0046] S5: Inputting the merged image of the infrared image and the visible light image to be fused into the trained fusion feature generation adversarial model to obtain a fused image.

[0047] Next, combined with Figure 2 , a detailed description of the image fusion method based on two-stage adversarial training and edge perception disclosed in the embodiment is given.

[0048] I. Data input and preprocessing

[0049] The data input and preprocessing module of the present application aims to preprocess the original input infrared and visible light images to ensure the effectiveness of the subsequent feature extraction and fusion steps. The specific steps of data input and preprocessing are as follows:

[0050] (I) Input image data

[0051] Respectively acquire infrared images and visible light images. The infrared images have unique thermal information and can effectively show the contours and thermal radiation of objects in insufficient visible light conditions. The visible light images are used to provide detailed texture and color information of the target object. Both images come from different spectral acquisition with the same field of view, and may have different spatial resolutions when input.

[0052] (II) Data preprocessing steps

[0053] 1. Grayscale processing

[0054] First, the visible light image I VISGrayscale processing is performed to convert the three-channel RGB image into a single-channel grayscale image. The pixel value of the grayscale image is calculated by the following formula:

[0055]

[0056] wherein, represents the red, green, and blue channel values of the visible light image at position (x, y), represents the pixel value of the grayscale image. The processed grayscale image is denoted as

[0057] 2. Size Adjustment

[0058] Since the size of the input image can be different, in order to unify the size and adapt to the subsequent feature extraction operation, the infrared image and the grayscale visible light image need to be adjusted in size.

[0059] The infrared image I IR and the grayscale visible light image are uniformly adjusted to HxW size, which is set to H=128, W=128 in this embodiment.

[0060]

[0061] The size adjustment is realized by bilinear interpolation to ensure that the details of the image are preserved as much as possible during scaling.

[0062] 3. Standardization Processing

[0063] The infrared image and the visible light image after grayscale processing and size adjustment need to be further standardized to improve the stability of the input data and the convergence speed of the training. The pixel value of the image is standardized to the interval [-1, 1], and the standardization formula is:

[0064]

[0065] wherein, I I ′ R ,I′ VIS ∈[-1,1] H×W×1 . This makes the mean value of the input image close to 0, avoiding the instability of training caused by the large range of different input data.

[0066] 4. Data Alignment and Merging

[0067] To ensure that the infrared and visible light images have consistent spatial information, the standardized infrared image and the visible light image are merged in the channel dimension to obtain the initial merged image I input :

[0068] Iinput = Concat(I I R ,I′ VIS );

[0069] wherein, The merged input data contains two channels: one for infrared information and one for visible light information, and B is the batch number; the data will be input to the encoder for feature extraction.

[0070] II. Bidirectional interaction encoder

[0071] As shown in Figure 3 , the main purpose of the bidirectional interaction encoder is to extract high-quality multi-scale features from the input merged data (infrared and visible light images), enhance feature details and texture information, and provide support for the decoder to generate high-quality fused images.

[0072] To achieve this goal, the bidirectional interaction encoder includes a primary feature extraction module, a high-low frequency feature decomposition module (HLFD module), and a multi-scale feature enhancement module (MEEM module). Through the close cooperation between these modules, the encoder can gradually extract and enhance image features.

[0073] (I) Overall architecture of the encoder

[0074] The bidirectional interaction encoder described in this embodiment adopts an innovative bidirectional interaction design, specifically the interaction between the HLFD module and the MEEM module, aiming to extract and enhance multi-scale features from the input infrared and visible light images. Through the cooperation and bidirectional interaction between the modules, the encoder forms a closed-loop optimization path, thereby improving the expression ability of the features and the quality of the fused images.

[0075] 1. Primary feature extraction module

[0076] First, the input features are subjected to preliminary convolution processing through the primary feature extraction module to obtain preliminary multi-scale feature representations.

[0077] The primary feature extraction module is used to extract preliminary features from the input initial merged image, and different scale and level feature representations are extracted through step-by-step down-sampling and feature enhancement.

[0078] As a specific implementation, the primary feature extraction module includes three first convolutional blocks, second convolutional blocks, and third convolutional blocks that are the same in structure and connected in sequence, each of which includes a convolutional layer with a kernel size of 3x3, a stride of 2, and a padding of 1, a batch normalization layer, and a nonlinear activation layer, as shown in Figure 4 .

[0079] First, the initial merged image​ The first convolutional block takes the input and performs initial feature extraction and down-sampling through convolutional layers.

[0080]

[0081] where the output of the convolution operation is The output channel number is 128, and the feature map size is reduced to one-half of the original. Next, a batch normalization (Batch Normalization) operation is applied to the convolution output to standardize the feature values, reduce the internal covariate shift during training, and improve the stability of the model. The output of batch normalization is:

[0082]

[0083] A Leaky ReLU activation function is used to perform non-linear mapping on the normalized features. The negative slope of the activation function is 0.2, which is used to preserve some negative value feature information to avoid sparsification. The activated features are represented as:

[0084]

[0085] The output feature map size is It contains preliminary spatial and texture features.

[0086] The output of the first convolutional block is taken as input and further feature extraction and down-sampling are performed through the convolutional layers of the second convolutional block:

[0087]

[0088] The output feature has a channel number of 64 and a feature map size of one-fourth of the input. A batch normalization layer is used to standardize the output features:

[0089]

[0090] A Leaky ReLU activation function is used to perform non-linear mapping on the normalized features:

[0091]

[0092] The output feature map size is

[0093] The output of the second convolutional block is taken as input, and further down-sampling and deep feature extraction are performed through the convolutional layers of the third convolutional block:

[0094]

[0095] Output feature The number of channels is reduced to 32, and the spatial size of the feature map is further reduced to one-eighth of the original. Batch normalization is applied to the output feature:

[0096]

[0097] The features are nonlinearly mapped using a Leaky ReLU activation function:

[0098]

[0099] The output feature map size is Preliminary multi-scale features.

[0100] The primary feature extraction module of this embodiment achieves multi-level abstraction and enhancement of input features through stepwise downsampling, nonlinear mapping, and feature channel changes. In each convolutional block, a convolution operation with a step size of 2 is used to downsample the feature map, effectively reducing the feature size and capturing more context information. After convolution, a Leaky ReLU activation function (negative slope of 0.2) is introduced for nonlinear enhancement, preserving negative feature information to avoid sparsification and ensuring that more feature information can be passed to the next layer. In terms of feature channels, the number of channels gradually increases and then gradually decreases with each layer of feature extraction. This design fully captures local details in the early stages and gradually compresses features to extract high-level semantic information in the later stages, effectively balancing the expression of feature details and overall structure.

[0101] 2. Multi-scale feature enhancement module (MEEM module)

[0102] The preliminary multi-scale features are input into the MEEM module for multi-scale edge enhancement.

[0103] As shown in Figure 5 , the MEEM module performs convolution operations on the feature map in parallel using different kernel sizes (e.g., 3x3, 5x5, 7x7), extracting edge information of different scales. Then, through a fusion mechanism, these edge features of different scales are fused to obtain multi-scale edge enhanced features , which is generally represented as:

[0104]

[0105] The goal of the Multi-Scale Edge Enhancement Module (MEEM) is to enhance the fused features in multiple stages and scales to further improve the edge and detail information of the features. This module uses a progressive multi-stage feature enhancement structure to capture the multi-scale properties of the input features and enhance the edge information.

[0106] As a specific embodiment, the enhancement process of the module is composed of multiple stages of progressive feature enhancement units, each of which processes and promotes the input features through convolution operations, pooling operations, and edge enhancement operations in multiple scales. The specific steps are as follows:

[0107] The input features of the multi-scale feature enhancement module come from the features extracted by the primary module, denoted as

[0108] First, initial feature conversion is performed. The input features are compressed in channels by a feature conversion unit with a structure of a 1x1 convolution layer to reduce the dimension and computational complexity of the features. The convolution operation formula is:

[0109]

[0110] wherein, C2 is the number of compressed channels. This represents the features after preliminary convolution, which are passed as input to the subsequent multi-scale processing stages.

[0111] Then, the converted features are subjected to cascaded enhancement. The cascaded enhancement unit contains L cascaded feature enhancement stages, and the specific steps of each stage are as follows:

[0112] (1) Feature pooling: average pooling operation is performed on the input converted features to extract low-frequency components in the features to capture more contextual information. The feature pooling of the 1st stage is denoted as:

[0113]

[0114] Output feature

[0115] If there is a high-frequency feature of HLFD feedback then

[0116]

[0117] The output feature is still and the subsequent process does not change.

[0118] Specifically, this embodiment introduces high-frequency feature adaptation and adaptive adjustment. Before multi-scale feature enhancement, based on the bidirectional interaction mechanism, if there is a high-frequency enhancement feature of HLFD feedback then the weight coefficient G is calculated through the gate network (Gate Network) MEEM to achieve adaptive adjustment of high-frequency information. Since the high-frequency feature of HLFD feedback To ensure the stability of fusion, a 1x1 convolution is performed to match the number of channels of the internal features of MEEM:

[0119]

[0120] MEEM adopts a gating network to calculate the first adaptive fusion weight G MEEM to adjust the contribution of high-frequency information to the initial features of MEEM:

[0121]

[0122] where σ(·) is the Sigmoid activation function, ensuring that the gating coefficient G MEEM is in the range [0, 1].

[0123] The calculated gating coefficient G MEEM is used as an adjustment factor to adaptively weight and fuse high-frequency features to obtain the first weighted fusion feature:

[0124]

[0125] (2) Convolution feature conversion: 1x1 convolution operation is performed on the pooled features to further compress and convert the feature representation:

[0126]

[0127] Convolution output feature

[0128] (3) Edge Enhancement Operation (Edge Enhancer): The purpose of edge enhancement is to enhance the details and edge information in the features. This operation calculates the high-frequency part of the image (such as edge, texture, etc. Detail information) and combines it with the low-frequency part of the image (such as background, global structure), thereby enhancing the edges and details in the image. In specific implementation, the EdgeEnhancer module uses the input features and extracts edge information through convolution and pooling operations, and further strengthens these edge features, making the edges of the image clearer.

[0129]

[0130] where the EdgeEnhancer module consists of a 3x3 convolution layer and a nonlinear activation function, used to enhance the edge properties of the input features.

[0131] (4) Progressive feature aggregation: The output features of each stage are added to the output features of the previous stage, thereby realizing progressive feature enhancement:

[0132]

[0133] This progressive feature aggregation manner ensures that the enhanced features of each stage directly affect the final output features, so that each enhancement can accumulate to improve the feature expression ability.

[0134] After the feature enhancement of all L stages, the output features of each stage are spliced in the channel dimension to obtain the final multi-scale edge enhanced feature representation:

[0135]

[0136] The final output multi-scale edge enhanced feature This feature integrates multi-scale information from each edge enhancement stage, providing rich and detail-enhanced features for the decoder module.

[0137] The multi-scale feature enhancement module of the embodiment progressively enhances features through multiple cascaded enhancement stages, each of which includes convolution, pooling, and edge enhancement operations, ensuring effective extraction and strengthening of feature information at different scales. In each stage, the features from the high-frequency decomposition are used to further enhance edges and details through the EdgeEnhancer module, so that the enhanced features can retain more rich detail information, thereby improving the quality of the finally generated image. Finally, by splicing the enhanced features of each stage in the channel dimension, effective fusion of multi-scale features is achieved, so that the output features contain both rich local information and multi-scale global feature expression, providing more discriminative feature representation. These features will be evaluated for quality by the discriminator and guide the encoder to continuously optimize the feature extraction process.

[0138] 3, High and low frequency feature decomposition module

[0139] The output results of the primary feature extraction module and the multi-scale edge enhanced feature output results are input into the HLFD module for high and low frequency decomposition processing. The HLFD module decomposes the feature map into high frequency and low frequency parts through average pooling and difference operation, which is generally represented as:

[0140]

[0141] Among them, the high frequency part F high_freq retains the edge and detail information of the image, while the low frequency part F low_freq represents the overall structure information of the image. This decomposition can process details and overall information independently, thereby more finely expressing different feature levels of the input image.

[0142] As a specific implementation, the main purpose of the high-low frequency feature decomposition module is to decompose the multi-scale edge-enhanced features into low-frequency and high-frequency components, and process and enhance each component respectively, and finally fuse them to obtain more expressive features. This module decouples the features, so that the low-frequency part retains the overall information, and the high-frequency part captures the detail and edge features, as shown in Figure 6

[0143] First, high-low frequency feature decomposition is performed. For extracting low-frequency features, the preliminary multi-scale features are down-sampled by an average pooling operation to extract low-frequency components. The extraction formula of low-frequency features is:

[0144]

[0145] where AvgPool(·) represents the average pooling operation, the pooling kernel size is k = 2, and the step size is s = 2, which is used to reduce the size of the feature map. The low-frequency features retain the overall structure and global information of the input features.

[0146] For extracting high-frequency features, the details and edge information are retained by subtracting the up-sampled low-frequency features from the preliminary multi-scale features. The calculation formula of high-frequency features is:

[0147]

[0148] where Upsample(·) represents the up-sampling operation, the step size is s = 2, and the bilinear interpolation is used to restore the low-frequency features to the size of the original features, i.e. The high-frequency features obtained in this way mainly retain the details, edges and rapidly changing feature information of the input image.

[0149] Then, feature enhancement is performed. For low-frequency features F low further enhancement processing is performed to extract and strengthen the global structure information. The low-frequency enhancement uses a convolution module UNetConvBlock, which contains two convolution layers in series. The operation of the first convolution is:

[0150]

[0151] where the convolution layer has a kernel size of 3x3, a step size of 1, and a padding of 1. The output features Then, batch normalization and activation are performed on the convolution results:

[0152]

[0153] Then, a second 3x3 convolution layer is used to continue processing the enhanced features:

[0154]

[0155] Output low-frequency enhancement feature After enhancement through two convolutional layers, the overall structure and texture features of low-frequency information are retained.

[0156] For high-frequency feature F high Enhanced by a dense connection block (DenseBlock) to strengthen the details and edge information. The dense connection block is composed of multiple cascaded convolutional layers, and the input of each layer contains the output of all previous layers. Specifically, the input and output relationship of the lth layer convolution is:

[0157]

[0158] where Concat(·) represents the concatenation operation in the channel dimension, The convolution kernel size of each layer is 3x3, the step is 1, and the padding is 1. Through the cascaded connection mode, the dense connection block can fully utilize the feature information of all previous convolutional layers, enhancing the diversity and detail description ability of the features. Through the enhancement operation of multiple layers in the dense connection block, the high-frequency enhancement feature F is obtained.

[0159] In this embodiment, based on the bidirectional interaction mechanism, when the high-low frequency feature decomposition module (HLFD) receives the edge feature provided externally , it will first perform channel adaptation and spatial size alignment on the edge feature to ensure that it can be effectively fused with the current high-frequency feature. Specifically, if the number of channels is inconsistent with the number of high-frequency channels processed by the module, a 1x1 convolution is first performed for channel mapping;

[0160]

[0161] If the spatial resolution of the edge feature does not match the input feature, it is adjusted to the same size using bilinear interpolation:

[0162]

[0163] After completing the channel and spatial alignment, the HLFD module splices the high-frequency enhancement feature and the edge feature, and inputs them into the gate network (Gate Network). The gate network produces adaptive gating coefficients through convolution and activation functions to weight and adjust the high-frequency enhancement feature. Specifically, the high-frequency enhancement feature and the edge feature are spliced and input into the gate network to produce a gating coefficient G∈[0,1]; then, the high-frequency enhancement feature is multiplied by the coefficient, so that the global high-frequency information is retained while the local details are strengthened or suppressed in combination with the edge feature.

[0164]

[0165] The calculated second adaptive fusion weight G is used to weight and fuse the high-frequency feature and the enhanced feature to achieve adaptive adjustment. The second weighted fusion feature F enhanced is calculated by the following formula:

[0166]

[0167] wherein represents the element-level point-by-point multiplication operation, and the gating coefficient G is used to dynamically adjust the strength of feature fusion to ensure that the output feature retains the clarity of details and maintains stability in the global structure.

[0168] Thus, by fusing the external edge feature with the high-frequency feature, the HLFD module can further adaptively correct the detail information of the high-frequency part on the basis of high-low frequency decomposition, retain the overall structure of the image, and improve the accuracy of key details and contours using the edge feature, providing higher quality feature input for subsequent high-frequency enhancement and feature fusion.

[0169] Finally, the low-frequency enhanced feature and the high-frequency enhanced feature are fused. The enhanced low-frequency enhanced feature is restored to the size of the original feature through upsampling operation:

[0170]

[0171] to obtain the output feature The upsampled low-frequency enhanced feature and the high-frequency enhanced feature are channel spliced to obtain the fused feature representation:

[0172]

[0173] wherein The fused feature contains enhanced global information and edge detail information, providing more expressive features for subsequent module processing.

[0174] In this embodiment, the bidirectional interaction encoder establishes a bidirectional interaction mechanism between the multi-scale feature enhancement module (MEEM module) and the high-low frequency decomposition module (HLFD module) to further enhance the features output by the primary feature extraction module. Specifically: first enter the MEEM module to obtain the first step of edge information, and then input the edge information and the initial feature extraction module output to the HLFD module together, and then perform high-low frequency decomposition on the input features, and the high frequency part is transmitted to the MEEM module again after being enhanced by the dense connection block. The MEEM module uses these high frequency information to perform edge enhancement operation, and feeds back the enhanced edge features to the HLFD module for guiding the subsequent feature decomposition and fusion process. This bidirectional interaction forms a closed-loop optimization process: the MEEM module provides the initial edge information, uses a gating weight network to guide the HLFD module to better utilize high frequency detail information, and then serves as the input of the next round of MEEM module, and the MEEM module generates more accurate edge features based on these information, which can help the HLFD module to achieve better feature decomposition in the next round of iteration.

[0175] Through this bidirectional interaction mechanism, the two modules can promote each other and progressively optimize the feature expression: the HLFD module focuses on maintaining the global structure information while extracting high-quality high-frequency details, and the MEEM module focuses on enhancing the edge information in these high-frequency details. This synergistic effect ensures that the features are fully enhanced in both detail performance and structural integrity.

[0176] This codec-integrated bidirectional interaction mechanism ensures that the features are effectively enhanced in both details and structure throughout the entire process from extraction to reconstruction, resulting in higher quality fusion results.

[0177] III. Dual-branch decoder

[0178] In the decoder, the bidirectional interaction mechanism is also continued and strengthened. Specifically:

[0179] Each branch of the decoder contains a high-low frequency processing unit, which inherits the same structure as the HLFD module in the encoder. In the processing process, the high frequency part of the high-low frequency processing unit will interact with the edge enhanced features output by the MEEM module. This interaction is manifested as: when the high frequency features are enhanced by the dense connection block, they are modulated by the edge enhancement information from the MEEM module at the same time, that is:

[0180] This design enables the decoder to maintain sensitivity to details in the reconstruction process, ensuring that the high-frequency information and edge features extracted by the encoder are not lost during reconstruction, and through the dual-branch structure and independent parameters, the two branches can focus on different scales and types of feature reconstruction, achieving more comprehensive feature recovery. With continuous interaction with the MEEM module, edge details are continuously optimized and refined during upsampling and reconstruction.

[0181] As Figure 7 shown, the dual-branch decoder module of the present embodiment includes two decoding branches that are structurally identical but have independent parameters, respectively used for infrared feature reconstruction and visible light feature reconstruction. Through this design, the two branches can maintain and enhance the feature expression of their respective modalities, and ultimately obtain a high-quality fused image that retains both infrared thermal radiation information and visible light details through weighted fusion.

[0182] The dual-branch decoder learns and optimizes through independent parameters, so that one branch focuses on learning and reconstructing the thermal radiation features of the infrared image, and the other branch focuses on learning and reconstructing the texture detail features of the visible light image. During the training process, by applying a targeted reconstruction loss to each branch (such as comparing the output of branch1 with the infrared image and the output of branch2 with the visible light image), the two branches are guided to learn the feature expression of the corresponding modality.

[0183] (I) Dual-branch decoder structure

[0184] Each branch is composed of the following components:

[0185] 1. Initial feature conversion layer

[0186] The initial feature conversion layer of each branch is composed of a convolution layer, a batch normalization layer, and an activation function. This layer is used to reduce the input feature map from 256 channels to 128 channels to provide more suitable feature expression for subsequent feature processing. Our encoder input contains In addition F edge.feat The first skip connection feature comes from the first stage output feature of the encoder (features[0]), i.e. F fused ; the second skip connection feature comes from the second stage output feature of the encoder (features[1]), i.e. F enhanced , and F edge.feat is the edge information obtained by the MEEM module for the first time. First, input the multi-scale edge enhancement module output feature to the initial feature conversion layer of the decoder module to obtain the initial conversion feature:

[0187]

[0188] where, is the initial converted feature for each branch.

[0189] 2. High-low frequency processing unit

[0190] Initial converted feature The initial converted feature is further decomposed and enhanced by the high-low frequency processing unit. In this unit, the feature is processed by the high-low frequency decomposition module respectively.

[0191] Low frequency feature processing: the low frequency converted feature is obtained by the average pooling operation:

[0192]

[0193] Further enhancement processing is performed on the low frequency converted feature using UNetConvBlock for convolution processing to maintain global structure information:

[0194]

[0195] For extracting high frequency converted feature, it is obtained by subtracting the up-sampled low frequency converted feature from the initial converted feature to preserve the detail and edge information. The calculation formula of the high frequency converted feature is:

[0196]

[0197] Enhancement is performed on the high frequency converted feature using DenseBlock.

[0198] In the enhancement process, the high frequency feature enters the HLFD module for high-low frequency decomposition, and combines the edge enhancement feature F edge.feat from the MEEM module for auxiliary enhancement. The HLFD module further extracts the details of the high frequency feature through DenseBlock, and finally obtains the enhanced high frequency feature:

[0199]

[0200] where, F edge.feat is the edge information calculated by MEEM, which selectively enhances the high frequency feature through the gating mechanism, so that the final high frequency enhanced feature has clearer edge and texture information.

[0201] 3. Feature fusion

[0202] The enhanced low frequency converted feature and high frequency converted feature are integrated through the feature fusion layer, and the fusion operation is the concatenation in the channel dimension:

[0203]

[0204] where,

[0205] 4. Intermediate feature fusion layer

[0206] After feature fusion, the fused feature is integrated with the skip connection feature of the corresponding layer in the encoder through a 3x3 convolution layer. The fusion process can be described as:

[0207]

[0208] where, is the corresponding stage feature extracted from the encoder, and Concat(·) represents concatenation in the channel dimension. The skip connection feature includes: the first skip connection feature comes from the first stage output feature (features[0]), that is, the second skip connection feature comes from the second stage output feature (features[1]), that is, When i=1, it refers to the 64-channel feature output by the first stage (features[0]); when i=2, it refers to the 128-channel feature output by the second stage (features[1]).

[0209] The fused feature is then further processed by convolution and batch normalization:

[0210]

[0211] Upsampling module: The spatial size of the feature map is restored through the upsampling module, and the size of the feature map is doubled by using the bilinear interpolation method:

[0212]

[0213] The size of the feature after upsampling is

[0214] 5. Fusion result

[0215] At the end of each branch, a 1x1 convolution layer is used to reduce the number of feature channels to 1, thereby generating a single-channel fusion result:

[0216]

[0217] The final output is a single-channel feature map, which is the output of the fusion branch.

[0218] (II) Workflow of the dual-branch decoder

[0219] 1. First feature reconstruction stage

[0220] Two branches receive feature maps from the encoder as input at the same time.

[0221] The feature dimension is reduced by the initial feature conversion layer.

[0222] The feature is enhanced and reconstructed by the high-low frequency processing unit, and interacts with the edge enhancement information of the MEEM module in the process, further strengthening the high-frequency details.

[0223] The channel splicing is performed with the first jump connection feature of the encoder to retain global and detail information. The first jump connection feature is the 64-channel feature (features[0]) output by the first stage to retain global structure information.

[0224] 2. Second feature reconstruction stage

[0225] The output of the first stage is upsampled to restore part of the spatial resolution.

[0226] Second high-low frequency decomposition and processing are performed to enhance detail and structure information.

[0227] The channel splicing is performed with the second jump connection feature. The second jump connection feature is the 128-channel feature (features[1]) output by the second stage to retain detail texture information.

[0228] The final stage output feature is generated through a series of convolution layers.

[0229] 3. Final fusion stage

[0230] The final outputs from the two branches are weighted and fused:

[0231]

[0232] The fused feature is normalized by the Tanh activation function to obtain the final fusion feature:

[0233] I fusion_output = Tanh(F final 0;

[0234] The embodiment significantly improves the quality of the fused image in the image reconstruction process through the synergistic effect of the double-branch decoder structure, the high-low frequency decomposition module (HLFD module), and the multi-scale feature enhancement module (MEEM module).

[0235] Although the two-branch decoder adopts the same network structure, it realizes the reconstruction of different levels of features through independent parameter optimization and feature processing paths. Specifically, each branch realizes the step-by-step optimization of features at different scales through the processing of the HLFD module twice, continuously utilizes the edge enhancement features for detail enhancement in the feature processing process, and gradually realizes the fine reconstruction of the features through the phased feature channel compression (from 256 channels to 128 channels, and then to 64 channels).

[0236] In terms of feature transmission between codecs, the embodiment makes full use of the two different scale jump connection features from the encoder. In the feature fusion process, the complete feature information is preserved through the channel dimension splicing operation, and the bilinear interpolation is used for feature upsampling to ensure the smooth transition of the feature map size. This design effectively avoids the loss of information of the features in the reconstruction process, and provides necessary feature support for high-quality reconstruction.

[0237] In terms of synergistic enhancement between modules, the HLFD module provides effective feature feedback by returning high-frequency features, and the edge features generated by the MEEM module run through the entire decoding process, continuously guiding the reconstruction of details. Finally, through feature addition and Tanh activation operation, the effective fusion of the information of the two branches is realized. This simple and effective fusion strategy makes full use of the complementary features learned by the two branches respectively, and ensures the quality of the fusion result.

[0238] Through the organic combination of the above mechanisms, the embodiment not only guarantees the structural integrity of the reconstructed image, but also realizes the accurate recovery of details, so as to obtain a high-quality fusion result. This design not only maintains the simplicity of the model structure, but also realizes the multi-level extraction and optimization of features.

[0239] IV. Discriminator

[0240] The fusion image discriminator module is used to discriminate the authenticity of the generated fusion image. The discriminator includes a double-head self-attention module, a layer-by-layer convolution module, and a full connection and discrimination output module, as shown in Figure 8 The discrimination process is as shown in Figure 9 .

[0241] Among them, the double-head self-attention module enables the discriminator to dynamically focus on the key areas of the features, enhancing the ability to capture high-frequency details and global structures; the layer-by-layer convolution module uses multiple convolution blocks to process the enhanced features, gradually down-samples, and abstracts high-level feature representations, to enhance the discrimination ability of the discriminator on the input image; the full connection and discrimination output module maps the extracted features through a full connection layer to a scalar output, which is used to represent the confidence of the discriminator on the "authenticity" of the input image.

[0242] (1) Double-Headed Self-Attention Module (DHSA):

[0243] By introducing the Double-Headed Self-Attention Module (DHSA), the discriminator can enhance the feature expression ability and better capture detailed information and global context features.

[0244] After the input layer of the discriminator, the input features are enhanced using the Double-Headed Self-Attention Module (DHSA). The role of this module is to dynamically weight the input features through self-attention mechanism to capture the dependency between different feature dimensions, thereby improving the expression ability of the features.

[0245] Input features: the discriminator receives the final fused features The input features are first enhanced by the DHSA module:

[0246] A = DHSA(I fusion_input );

[0247] Implementation of attention mechanism: the Double-Headed Self-Attention Module first performs linear transformation on the input features through the convolution layer qkv to generate queries, keys and values:

[0248] qkv = Conv 1×1 (I fusion_imput );

[0249] Next, the qkv is further processed through the depth separable convolution operation, and then it is split into five parts q1, k1, q2, k2, v, where q1, k1 are used for high-frequency feature calculation, q2, k2 are used for low-frequency feature calculation, and v is used as a shared value vector. Through feature sorting and matching operations, effective interaction between high and low frequency feature dimensions is ensured.

[0250] Attention calculation: after normalizing the queries and keys respectively, the attention score matrix is calculated, and then normalized through the softmax function:

[0251]

[0252] where d k represents the dimension of the key, and · represents matrix multiplication. Then the value v is weighted and summed using the attention score matrix to obtain the enhanced features:

[0253] out 1 = attn1 × v;

[0254] The same steps are also used for q2 and k2 to calculate attn2 and the corresponding enhanced output out2.

[0255] Output feature: After the interaction of out1 and out2, map back to the original feature dimension through the convolution layer project_out, and fuse the output with the input through the residual connection:

[0256] A' = I fusion_input + project_out(out1 ⊙ out2);

[0257] Where ⊙ represents element-level multiplication, which is used to fuse two enhanced features. The final output attention feature As the input of the subsequent convolution block.

[0258] The discriminator of the embodiment adds a double-head self-attention mechanism in the input feature processing stage, which can effectively capture the complex correlation between features by query, key and value interaction calculation.

[0259] (ii) Layer-by-layer convolution module:

[0260] Based on the enhanced feature A' after the DHSA module, the discriminator uses multiple convolution layers with the same structure to gradually extract higher-level features, reduce the spatial dimension of the image, and increase the abstractness of the features.

[0261] The first convolution layer with a convolution kernel size of 3x3, a step size of 2, and a padding of 1 is used to downsample the input feature:

[0262] B (1) = LeakyReLU(BatchNorm(Conv 3×3 (A'), α = 0.2);

[0263] Where α = 0.2 represents the negative slope of Leaky ReLU.

[0264] The second convolution layer with the same structure is used to further downsample and abstract the features:

[0265] B (2) = LeakyReLU(BatchNorm(Conv 3×3 (B (1) ), α = 0.2);

[0266] Continue to process the features to extract the final discriminative features:

[0267] B (3) = LeakyReLU(BatchNorm(Conv 3×3 (B (2) ), α = 0.2);

[0268] After feature enhancement, the discriminator of the embodiment uses a layer-by-layer convolution module to downsample and abstract the input, enabling it to evaluate the authenticity of the input image at multiple scales. Each layer of convolution combines batch normalization and Leaky ReLU activation to ensure effective feature flow and enhancement.

[0269] (iii) Fully connected and discriminative output module:

[0270] The feature map output by the convolution feature extraction layer is flattened into a vector:

[0271] C = B (3) view(B (3) size(0),-1);

[0272] The flattened features are processed through two fully connected layers to generate the final discriminative result:

[0273] D (1) = LeakyReLU(Linear(C,1024),a=0.2);

[0274] D output = Linear(D (1) ,1);

[0275] wherein, represents the discriminative result of the input image.

[0276] The fully connected and discriminative output module improves the evaluation of the authenticity of the generated image by fully combining global context and local detail features.

[0277] In the embodiment, the introduction of the double-head self-attention mechanism (DHSA) in the discriminator significantly improves the feature enhancement capability of the discriminator for the fused image. The attention mechanism enables the discriminator to dynamically focus on key areas during feature extraction, effectively capturing detailed information and global context, thereby improving the judgment of the authenticity of the generated image. The combination of convolution feature abstraction and fully connected discriminative output ensures that the discriminator can process local features and enhance the quality of the generated image through global features, ultimately achieving efficient and accurate discrimination of the fused image.

[0278] V. Training process

[0279] The embodiment uses the MS-COCO and LLVIP datasets for training, where the MS-COCO provides 12025 natural images for training the decoder, and the LLVIP dataset provides 16250 pairs of infrared and visible light image pairs for training the encoder.

[0280] All input images are uniformly adjusted to 128x128 size, and the training is performed with a batch size of 12, 5 rounds of training, and a learning rate of 1e-4 (using the StepLR strategy to dynamically adjust, with a decay factor of 0.92).

[0281] (I) Overview of the training process

[0282] The training process alternates between the adversarial training phase and the reconstruction training phase, ensuring the continuous improvement of the global consistency and detail accuracy of the fused image. The overall training process is shown in Figure 10

[0283] 1. Adversarial training phase:

[0284] As shown in Figure 11 , the discriminator is first trained. The discriminator receives data from real images and images generated by the generator, and learns to distinguish real features from generated features by maximizing the output for real images and minimizing the misjudgment for generated images, thereby calculating the discriminator loss

[0285] After the discriminator is trained, the generator is trained. Through the joint optimization of the generator adversarial loss identity preservation loss and edge feature loss , the quality of the generated image is continuously improved. The main goal of the generator in this stage is to generate images with better fusion effect, making it difficult for the discriminator to distinguish.

[0286] 2. Reconstruction training phase:

[0287] As shown in Figure 12 , in the reconstruction training phase, the decoder is further optimized using natural image datasets. The main goal is to improve the reconstruction quality of the generated image, especially the restoration of details and the overall consistency of features.

[0288] In this stage, gradient reconstruction loss and mean square error loss (MSE loss) are used for optimization to more finely improve the reconstruction effect of the image, so that the generated fused image not only has a natural visual effect, but also maintains a high level of detail quality

[0289] (ii) Loss function setting

[0290] In the entire training process, the loss function settings for the generator and the discriminator are as follows:

[0291] 1. Generator loss function

[0292] ​The generator loss function includes an adversarial loss, an identity preservation loss and an edge feature loss, and the goal is to ensure that the generated image has high quality in terms of overall and details.

[0293]

[0294] wherein, represents the adversarial loss, represents the identity preservation loss, represents the edge feature loss, represents the interaction loss. λ adv , λ idt , λ edge , λ inter represents the weight coefficient of each loss term, which can be set according to the actual situation, and is preferably set to for adjusting the influence degree of each loss in the overall loss.

[0295] 2. Adversarial loss

[0296] The adversarial loss is used to make the generated image by the generator as "deceptive" as possible to the discriminator, so that the generated image is as close as possible to the real image in vision. It is defined as:

[0297]

[0298] wherein, I fusion_output is the generated image by the generator, and D is the discriminator, and the goal is to make the discriminator unable to distinguish the generated image from the real image, thereby improving the quality of the generated image.

[0299] 3. Identity preservation loss

[0300] The identity preservation loss is used to ensure that the generated image can retain important features in the input image, especially when multi-modal fusion is performed, to ensure that the main structural information of the image is preserved. The L1 loss is used to calculate the difference between the generated image and the input image:

[0301]

[0302] wherein, I fusion_output is the generated image by the generator, and I input is the input image, and the goal of the identity preservation loss is to ensure that the generated image is consistent with the input image in terms of global structure and main features.

[0303] 4. Edge feature loss

[0304] Edge feature loss is used to enhance the edge details in the generated image, ensuring that the generated image is consistent with the real image in high-frequency features. Laplacian filter is used to extract the edge features of the generated image and the target image, and L1 loss is calculated:

[0305]

[0306] where, represents the Laplacian filter operation, aiming to preserve the details of the image edges. target is the target image. fusion_output represents the fusion image output by the generator. target represents the target image, i.e. the real image expected to be generated. represents the edge consistency between the generated image and the real image, thereby enhancing the high-frequency details.

[0307] 5. Cross loss

[0308] Cross loss is used to evaluate and optimize the feature interaction between the internal modules of the generator, especially the bidirectional interaction process between the high-low frequency feature decomposition module (HLFD module) and the multi-scale feature enhancement module (MEEM module). Through the optimization of the cross loss, the overall feature expression and detail performance of the generated image are optimized.

[0309] Specific definition and calculation of cross loss:

[0310] 6. High-low frequency interaction loss

[0311] In the encoder part of the generator, the HLFD module and the MEEM module form a closed loop optimization through the mutual transmission and enhancement of high-low frequency features. To ensure the effectiveness of the high-frequency features in this interaction process, a feature interaction loss is introduced to evaluate the fidelity and effectiveness of the high-frequency features after being enhanced by the MEEM module:

[0312]

[0313] where, enhanced_high represents the high-frequency features enhanced by the MEEM module, while orig_high represents the original high-frequency features output from the HLFD module.

[0314] 7. Edge interaction loss

[0315] When dealing with high-frequency features, edge feature loss is calculated to evaluate the effectiveness of the edge feature adaptation and gating mechanism in the interaction process. Edge interaction loss is used to evaluate the quality change of edge features during transmission between different modules:

[0316]

[0317] where A edge represents the characteristic gate weight, used to adjust the weight of the loss according to the importance of the edge features, F edge_feat represents the enhanced edge features, F orig_edge represents the original edge features.

[0318] 8. Combination of interaction loss:

[0319] The final interaction loss is a weighted combination of high-frequency interaction loss and edge interaction loss:

[0320]

[0321] where λ high and λ edge are the weight coefficients of high-frequency interaction and edge interaction, respectively. By adjusting these coefficients, the strength of feature interaction can be balanced during training, ensuring the quality of generated images to reach the best state.

[0322] 9. Discriminator loss function

[0323] The goal of the discriminator is to distinguish between real images and generated images, by maximizing the accuracy of judging real images and minimizing the misjudgment of generated images to improve the discrimination ability of the discriminator. The definition is as follows:

[0324]

[0325] I target represents the sample from the real image. I fusion_output represents the fusion image generated by the generator. p data represents the distribution of real images. P c represents the distribution of generated images. The goal of the discriminator is to maximize the correct judgment probability of real images and minimize the false judgment probability of generated images, thereby promoting the improvement of the generator.

[0326] 10. Loss function in reconstruction phase

[0327] In the reconstruction training phase, the main optimization is performed on the decoder, using natural image datasets to improve the reconstruction quality, and the following loss function is adopted:

[0328] Gradient reconstruction loss

[0329] Gradient reconstruction loss is used to improve the smoothness and detail transition effect of generated images, by calculating the gradient difference between generated images and target images, to better preserve details:

[0330]

[0331] Mean Squared Error Loss (MSE)

[0332] The MSE loss is used to measure the average squared difference between the generated image and the target image, ensuring that the overall reconstruction error is minimized, resulting in a more realistic fused image:

[0333]

[0334] (III) Optimization Strategy

[0335] During the entire training process, the Adam optimizer is used to update the model parameters. The initial learning rate is set to 1×10 -4 , and the learning rate is dynamically adjusted by the StepLR learning rate adjustment strategy. The specific steps are as follows:

[0336] Adam optimizer: By setting the initial learning rate and momentum parameters (β1, β2), efficient updating of model parameters is achieved. The stability and convergence speed of the training process are guaranteed.

[0337] StepLR learning rate adjustment: Reduce the learning rate every fixed number of steps to better adapt to the changes in the model during training, and avoid model oscillation due to excessive learning rate in the later stage.

[0338] This embodiment generates a generator that gradually improves its ability to extract and reconstruct features from input images through alternating optimization of the adversarial training phase and the reconstruction training phase. The adversarial phase ensures that the fused image has good overall feature consistency, and the reconstruction phase further optimizes the details, resulting in a fused image with high visual quality and realism.

[0339] The edge feature loss is added to the generator loss function to ensure the effective preservation of edge and detail information in the generated image. Through the dual action of edge feature loss and gradient reconstruction loss, the edge details of the generated image are significantly enhanced, making it more realistic and delicate in vision.

[0340] By using infrared-visible light image pairs in the adversarial training phase and natural image datasets in the reconstruction training phase, the model not only preserves multi-modal features but also effectively improves image reconstruction ability and generalization performance, ensuring the consistency and detail performance of the generated image in various scenarios.

[0341] To verify the effect of the present invention, as Figure 13 The specific effect comparison of the method of the present invention in the infrared and visible light image fusion task is shown, including: (a) original infrared image, (b) original visible light image, and (c) fused result image.

[0342] As can be seen intuitively from the figure, the method of the present application successfully takes into account the advantages of infrared and visible light images during fusion, so that the fused image not only has the target prominence of the infrared image, but also retains the rich details and natural colors of the visible light image. Compared with single-mode images, the fusion result is better in brightness balance, contrast optimization and target clarity, ensuring the overall perceptibility and visual comfort of the scene. In particular, in the fused image (c), the infrared target is fully retained, so that it still has good visibility at night or in low light environments, while the details in the visible light image, such as ground texture, building structure, trees, etc., are also restored to a high degree, avoiding the common problems of blurring or excessive smoothing in traditional methods. In addition, the overall color and brightness transition of the fused image is natural, avoiding the artifacts, excessive enhancement or information loss that may occur in traditional methods, so that the final result is more in line with the visual quality of human eye perception habits, providing more reliable data support for subsequent intelligent analysis, target detection and scene understanding.

[0343] Compared with existing methods, the two-stage adversarial training strategy and edge perception mechanism adopted by the present application effectively suppress the common problems of detail loss and edge blur in traditional methods while ensuring the clarity and visibility of the infrared target. Specifically, the present method decouples high and low frequency features (HLFD) to perform multi-scale decomposition of infrared and visible light information, and further optimizes the fusion of high frequency information with the help of multi-scale feature enhancement (MEEM) to ensure the clarity of target edges. In addition, combined with the double-head self-attention discriminator (DHSA), the present method can accurately extract and enhance important details, so that the fusion result is better than traditional methods in terms of visual consistency and information retention.

[0344] In summary, Figure 13 The effectiveness and superiority of the method of the present application are fully verified, and the practical application value of the method in the field of infrared-visible light image fusion is proved.

[0345] Embodiment Two

[0346] The embodiment provides an image fusion system based on two-stage adversarial training and edge perception, comprising:

[0347] The merging module is configured to select an infrared-visible light sample image pair dataset, and merge the channels of the infrared image and the visible light image in the dataset to obtain an initial merged image.

[0348] The fusion model construction module is configured to construct a fusion feature generation adversarial model, wherein the model comprises a bidirectional interactive encoder, a double-branch decoder and a discriminator.

[0349] The generator is configured to input the initial merged image into a bidirectional interaction encoder, perform feature enhancement based on a bidirectional interaction mechanism, and obtain multi-scale fusion features by using a multi-scale feature enhancement module and a high-low frequency feature decomposition module.

[0350] The discriminator is configured to obtain the fusion features, perform authenticity judgment training on the model based on a two-stage training strategy, and complete training of the model when a loss function converges.

[0351] The fusion module is configured to input the merged image of the infrared image and the visible light image to be fused into the trained fusion feature generation adversarial model to obtain a fused image.

[0352] Embodiment three

[0353] The embodiment provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement steps in the image fusion method based on two-stage adversarial training and edge perception.

[0354] Embodiment four

[0355] The embodiment provides a computer device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements steps in the image fusion method based on two-stage adversarial training and edge perception when executing the program.

[0356] The steps or modules involved in the above embodiments two to four correspond to embodiment one, and the specific embodiments can refer to the related description part of embodiment one. The term "computer readable storage medium" should be understood as including a single medium or multiple media of one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying instruction sets for execution by a processor and causing the processor to perform any method in the present application. The above only describes the preferred embodiments of the present application and is not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. An image fusion method based on two-stage adversarial training and edge awareness, characterized in that, include: Infrared-visible light sample images were selected for the dataset, and the infrared and visible light images in the dataset were channel-merged to obtain the initial merged image; A fusion feature generation adversarial model is constructed, which includes a bidirectional interactive encoder, a dual-branch decoder, and a discriminator. The initial merged image is input into a bidirectional interactive encoder. Based on the bidirectional interactive mechanism, feature enhancement is performed using a multi-scale feature enhancement module and a high- and low-frequency feature decomposition module to obtain multi-scale fused features. The multi-scale fused features are then input into a dual-branch decoder for fusion, and infrared and visible light image features are reconstructed separately to obtain fused features. The fused features are input into the discriminator, and the model is trained based on a two-stage training strategy to determine authenticity; when the loss function converges, the training of the model is complete. The combined image of the infrared and visible light images to be fused is input into the trained fusion feature generative adversarial model to obtain the fused image.

2. The image fusion method based on two-stage adversarial training and edge awareness as described in claim 1, characterized in that, The process involves selecting infrared-visible light sample images for the dataset and performing channel merging on the infrared and visible light images in the dataset to obtain an initial merged image; specifically, this includes: Acquire infrared and visible light sample images collected at the same time; The visible light sample image is processed to grayscale, and the sizes of the grayscale visible light sample image and the infrared sample image are unified. After standardization, the two images are merged in the channel dimension to obtain the initial merged image.

3. The image fusion method based on two-stage adversarial training and edge awareness as described in claim 1, characterized in that, The initial merged image is input into a bidirectional interactive encoder. Based on the bidirectional interactive mechanism, feature enhancement is performed using a multi-scale feature enhancement module and a high- and low-frequency feature decomposition module to obtain multi-scale fused features; specifically, this includes: The bidirectional interactive encoder includes a primary feature extraction module, a multi-scale feature enhancement module, and a high- and low-frequency feature decomposition module. The initial merged image is input into a primary feature extraction module built on convolutional layers to obtain preliminary multi-scale features; The initial multi-scale features are input into the multi-scale feature enhancement module and the high- and low-frequency feature decomposition module, respectively. In the multi-scale feature enhancement module, the initial convolutional features are obtained based on the initial feature transformation. In the high- and low-frequency feature decomposition module, the high-frequency enhancement features and low-frequency enhancement features are obtained based on the high- and low-frequency decomposition. The high-frequency enhancement features are then input into the multi-scale feature enhancement module. Based on the bidirectional interaction mechanism, in the multi-scale feature enhancement module, if a high-frequency enhanced feature is received from the high- and low-frequency feature decomposition module, the first adaptive fusion weight is calculated based on the gated network, and the preliminary convolutional feature and the high-frequency enhanced feature are fused to obtain the first weighted fusion feature, and the first weighted fusion feature is averaged; otherwise, the preliminary convolutional feature is averaged. After average pooling, the features are transformed by convolution, edge enhancement, and feature aggregation to obtain multi-scale edge enhancement features; In the high- and low-frequency feature decomposition module, if the multi-scale edge enhancement features fed back by the multi-scale feature enhancement module are received, the second adaptive fusion weight is calculated based on the gated network, and the high-frequency enhancement features and the multi-scale edge enhancement features are fused to obtain the second weighted fusion features; the second weighted fusion features and the low-frequency enhancement features are fused to obtain the multi-scale fusion features.

4. The image fusion method based on two-stage adversarial training and edge awareness as described in claim 3, characterized in that, The process of obtaining high-frequency enhanced features and low-frequency enhanced features based on high-frequency decomposition in the high- and low-frequency feature decomposition module specifically includes: The initial multi-scale features are downsampled using average pooling to extract low-frequency features; High-frequency features are obtained by subtracting the upsampled low-frequency components from the preliminary multi-scale features; Low-frequency features are convolved to obtain low-frequency enhanced features; high-frequency features are processed through dense connecting blocks to obtain high-frequency enhanced features.

5. The image fusion method based on two-stage adversarial training and edge awareness as described in claim 3, characterized in that, The edge enhancement specifically includes: extracting edge information from the convolutionally transformed features through convolution and pooling operations, and then enhancing the edge features using the EdgeEnhancer module, as shown below: in, Features after edge enhancement These are the features after convolution transformation.

6. The image fusion method based on two-stage adversarial training and edge awareness as described in claim 1, characterized in that, The multi-scale fusion features are input into a dual-branch decoder for fusion, and infrared and visible light image features are reconstructed separately to obtain fused features; Specifically, it includes: The dual-branch decoder includes two decoding branches with the same structure but different parameters; the dual-branch decoder includes a first feature reconstruction stage and a second feature reconstruction stage. The first feature reconstruction stage includes: two branches simultaneously receiving feature maps from the encoder as input; Reduce feature dimensionality through a primary feature transformation layer; The high- and low-frequency processing units are used to enhance and reconstruct the features, and during the process, they interact with the edge enhancement information of the multi-scale edge enhancement module to enhance high-frequency details. The features are fused with the output features from the first feature enhancement stage of the encoder to obtain preliminary fused features; The second feature reconstruction stage includes: upsampling the preliminary fused features to restore some spatial resolution; A second high- and low-frequency decomposition and processing is performed to enhance details and structural information; The enhanced fused features are obtained by fusing them with the output features of the second feature enhancement stage of the encoder. The enhanced fusion features of the two branches are weighted and fused to obtain the final fusion feature.

7. The image fusion method based on two-stage adversarial training and edge awareness as described in claim 1, characterized in that, The discriminator includes a dual-head self-attention module, a layer-by-layer convolution module, and a fully connected and discriminative output module; The fused features are input into a dual-head self-attention module to obtain attention fused features; the attention fused features are input into a layer-by-layer convolution module for convolution operation to obtain multi-scale fused features; the multi-scale fused features are input into a fully connected and discriminative output module, which combines global context and local detail features for authenticity judgment.

8. An image fusion system based on two-stage adversarial training and edge awareness, characterized in that, include: The merging module is used to select infrared-visible light sample images to the dataset and perform channel merging on the infrared and visible light images in the dataset to obtain an initial merged image; A fusion model construction module is used to construct a fusion feature generation adversarial model, which includes a bidirectional interactive encoder, a dual-branch decoder, and a discriminator. The generator takes the initial merged image as input to the bidirectional interactive encoder. Based on the bidirectional interactive mechanism, it uses a multi-scale feature enhancement module and a high- and low-frequency feature decomposition module to enhance features and obtain multi-scale fused features. The multi-scale fused features are then input into the dual-branch decoder for fusion, and infrared and visible light image features are reconstructed separately to obtain fused features. A discriminator is used to acquire fused features and train the model to judge authenticity based on a two-stage training strategy; when the loss function converges, the training of the model is completed. The fusion module is used to input the merged image of the infrared image and the visible light image to be fused into the trained fusion feature generative adversarial model to obtain the fused image.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the image fusion method based on two-stage adversarial training and edge awareness as described in any one of claims 1-7.

10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the image fusion method based on two-stage adversarial training and edge awareness as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Infrared and visible light image fusion method based on multi-stage generative adversarial network

    CN116935177A

  • Infrared and visible light image fusion method based on Swin Transform and GAN

    CN117333410A