Multi-modal image fusion method based on SwinTransform
By adopting a multimodal image fusion method based on SwinTransformer in image fusion, using the dual decoder Laplace pyramid structure and adaptive frequency fusion module, the problem of image fusion distortion in the prior art is solved, and the effect of image fusion and the generalization ability of the model are improved.
Patent Information
- Application Number
- CN202510015768.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-06
AI Technical Summary
The existing image fusion method ignores the potential possibility of information complementarity in multi-task fusion, resulting in images being distorted in certain frequency ranges, large computing overhead, weak local structural information perception ability, and insufficient feature fusion, resulting in redundancy in data information.
Using the multimodal image fusion method based on SwinTransformer, a dual-decoder Laplace pyramid structure and adaptive frequency fusion module are designed. Multi-scale features are extracted through the Laplace pyramid encoder and decoder, and feature fusion is performed using the differential information fusion module and the public information fusion module.
The effect of image fusion is improved, and the multi-scale different frequency information of the image is fully utilized, the performance and model complexity are balanced, and the generalization ability and feature extraction ability of the model are enhanced.
Smart Images

Figure CN119963957A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image fusion, and in particular to a multimodal image fusion method based on SwinTransformer. Background Art
[0002] Image fusion plays a vital role in enhancing the quality and accuracy of semantic segmentation, which is particularly critical for autonomous driving systems. By fusing information from multiple imaging sensors or modalities (such as infrared and visible light images), image fusion enriches the data and thus improves the perception capability of autonomous vehicles. However, existing image fusion methods mainly focus on solving the fusion problem of a single task, ignoring the potential possibility of information complementarity in multi-task fusion. Therefore, how to perform multimodal image fusion and explore information complementarity in multi-task fusion remains a challenge.
[0003] The Chinese authorization announcement number is "CN117173525B", and the name is "A general multimodal image fusion method and device". This method first uses convolution to extract shallow information, and then extracts multi-scale features through an encoder constructed by i+1 sequentially connected Transformer modules and the first i Transformer modules are connected to a wavelet downsampling module. The decoder is composed of i sequentially connected upsampling modules, and each of the upsampling modules is connected to a Transformer module to generate a high-quality fused image. This method uses wavelet downsampling to obtain multi-scale image features, but this operation will cause the image to be distorted in certain frequency ranges, especially the high-frequency part, and will cause the loss of detail information in the low-frequency area; in order to improve the network's extraction of global information, the Transformer module is used, which will lead to a larger computational overhead, higher requirements for computing resources, and weak perception of local structural information; in addition, simply splicing two images directly and sending them to the model will lead to the loss of spatial relationship, insufficient feature fusion, and redundancy of data information.
[0004] In summary, how to design a new multimodal image fusion method that can fully utilize the multi-scale and different frequency information of the image while improving the generalization ability of the network, balance the performance and complexity of the model, and utilize the differences and commonalities between the source images to achieve better fusion effects is an urgent problem to be solved in this field. Summary of the invention
[0005] The technical solution of the present invention to solve the above technical problem is to provide a multimodal image fusion method based on SwinTransformer, comprising the following steps:
[0006] Step 1, prepare training data: select multi-focus dataset, multi-exposure dataset, infrared image dataset and visible light image dataset, divide the dataset into training set, validation set and test set, and pre-process the original images and their corresponding labels in the multi-focus dataset and multi-exposure dataset;
[0007] Step 2, constructing a network model: including an image receiving module for receiving image data, a shallow feature extraction module for extracting coarse shallow information of the image, a global feature extraction module for extracting fine information with global information, an adaptive frequency fusion module for fusing difference information and common information of the image, and an image reconstruction module for reconstructing the fused image; the global feature extraction module includes two Laplacian pyramid encoders with the same structure, and the image reconstruction module includes a decoder with an inverse Laplacian pyramid structure;
[0008] Step 3, first-stage network training: Use labeled data from the multi-exposure dataset and the multi-focus dataset for multi-task supervised pre-training to enhance the model’s ability to extract complementary features from multi-exposure and multi-focus tasks;
[0009] Step 4, second-stage network training: Use unlabeled data from the infrared and visible light image datasets for unsupervised training to enhance the generalization ability of the model while saving data costs; use labeled data from the multi-exposure dataset and multi-focus dataset for supervised training to enhance the model's ability to retain source image details;
[0010] Step 5, determine the evaluation index: select the loss function to evaluate the quality of image fusion, and determine the evaluation index to evaluate the performance of the network.
[0011] Furthermore, the multimodal image fusion method based on SwinTransformer also includes the following steps: Step 6, solidifying the network model: after completing the network model adjustment, fix the network parameters and determine the final multimodal image fusion model; if image fusion tasks are required later, the image to be fused can be directly input into the network model to obtain the fusion result.
[0012] Furthermore, in the step 2, after receiving the two types of image data, the image input module sends them to the shallow feature extraction module to perform a 3×3 convolution operation on the input image to obtain local information and shallow features of the image; the shallow information is sent to the global feature extraction module; the global feature extraction module includes two Laplace pyramid encoders with the same structure, which continuously downsample and Laplace differencing the image, decompose feature maps of different frequencies, and then send them to the adaptive frequency fusion module; the adaptive frequency fusion module includes an image embedding module, a difference information fusion module, two public information fusion modules and a refining module, fuses images of the same frequency, and after fusion, sends them to the image reconstruction module for reconstruction, and the image reconstruction module includes a decoder with an inverse Laplace pyramid structure, which continuously upsamples and residually connects the image and then outputs the fused image.
[0013] Furthermore, in step 2, the encoder and the decoder form a Laplacian pyramid structure; the encoder obtains the lowest-scale feature Then upsample and calculate the residual with the features of the same scale to generate high-frequency features The structure of the decoder is completely symmetrical with that of the encoder;
[0014] The adaptive frequency fusion module includes an image embedding module, a difference information fusion module, two common information fusion modules and a refinement module. The image embedding module divides the image into patches, introduces position encoding after linear projection, and outputs a sequence containing patch embedding vectors as the input of the subsequent model. The embedding process can be expressed as:
[0015]
[0016] Introduce the cross attention architecture and build a difference information fusion module, using the frequency features decomposed by the encoder as and As input, and output difference information features; and The division into n local feature fragments is as follows:
[0017]
[0018] in and And s = h × w;
[0019] Using a linear layer to transform feature snippets into query Q, key K and value V, the linear projection can be expressed as:
[0020] Q i =LinearQ(Q i ),K i =LinearK(K i),V i =LinearV(V i );
[0021] Where i = 1, ..., n, Linear(·) represents a linear projection operation shared between different fragments;
[0022] The dot product attention layer is used to calculate the similarity matrix of the query Q and the key K, and then multiplied by the value V to infer the relevance between Q and V; it is expressed as:
[0023]
[0024] Among them, d k It is a scaling factor that can alleviate the convergence of the softmax function to the gradient minimum region when the dot product increases;
[0025] The difference information between Q and V is obtained by removing the common information; it is expressed as:
[0026] DV = Linear (V - CMV);
[0027] Inject the difference information into Q; expressed as:
[0028]
[0029] Among them, LN(·) represents the normalization layer, MLP(·) represents the multi-layer perceptron, and F dm is the output of the difference information fusion module;
[0030] Two identical information fusion modules are introduced, and the output fragments of the difference information fusion module are used to provide Q, while The fragment provides K and V, so F dm and The common information between them can be expressed as:
[0031]
[0032] Then F dm and Public information between CM joins F dm , expressed as:
[0033]
[0034] Among them, F cm represents the output of the first public information fusion module;
[0035] F cm and The common information between them is injected into the fusion feature to further enrich the fusion feature;
[0036] The fused features are input into the refinement module for further feature fusion. The refinement module consists of convolutional layer 1, mixing block, and convolutional layer 2, and then outputs the fusion result after residual connection and Solftmax.
[0037] Compared with the prior art, this application has the following beneficial effects:
[0038] 1. The present invention designs a novel image fusion framework of a dual-decoder Laplacian pyramid structure in the image processing module, which makes full use of the gradient image information of the input image and the detail information of the original image to improve the fusion effect, and utilizes the potential information complementarity between different fusion tasks, so that the model can have the ability to process multimodal image fusion, thereby improving the wide applicability of the model;
[0039] 2. The present invention designs an adaptive frequency fusion module, designs a difference information fusion module and a common information fusion module in the image fusion module, and utilizes the cross-attention mechanism to obtain the difference features and common features of the two images. On the premise of maintaining their respective characteristics, the feature extraction capability of the network is better improved, thereby effectively improving the effect of image fusion.
[0040] 3. The present invention proposes a two-stage training method. The supervised training in the first stage can help the model learn some basic characteristics of image fusion, provide a good foundation for the training in the second stage, and reduce the model instability in the initial stage; using unsupervised training in the second stage can effectively solve the problem that it is difficult to obtain high-quality labels for infrared and visible light image data. In addition, the data distribution of infrared and visible light images and multi-focus and multi-exposure are different. Unsupervised training can help the model adapt to new modalities and characteristics and improve the generalization ability of the model; in the second stage of unsupervised training, a part of the supervised multi-focus and multi-exposure data is integrated as a positive supervision signal to stabilize the unsupervised training process and avoid the problem of overfitting. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying creative work.
[0042] Figure 1 This is a flowchart of the steps of the multimodal image fusion method based on SwinTransformer of the present invention;
[0043] Figure 2It is a schematic diagram of the overall structure of the network model described in the present invention;
[0044] Figure 3 This is a schematic diagram of the structure of the adaptive frequency fusion module of the present invention;
[0045] Figure 4 This is a schematic diagram of the structure of the difference information fusion module of the present invention;
[0046] Figure 5 This is a schematic diagram of the structure of the public information fusion module of the present invention;
[0047] Figure 6 This is a schematic diagram of the composition of the refining module of the present invention;
[0048] Figure 7 It is a schematic diagram of the composition of the mixing block described in the present invention. DETAILED DESCRIPTION
[0049] The present invention proposes a multimodal image fusion method based on SwinTransformer, aiming to.
[0050] The multimodal image fusion method based on SwinTransformer proposed in the present invention will be described in a specific embodiment below:
[0051] Embodiment 1:
[0052] A multimodal image fusion method based on SwinTransformer, such as Figure 1 As shown, the following steps are included:
[0053] Step 1, prepare training data: select multi-focus dataset, multi-exposure dataset, infrared and visible light image dataset, divide the dataset into training set, validation set and test set, and pre-process the original images and their corresponding labels in the multi-focus and multi-exposure training sets;
[0054] Step 2, constructing a network model: including an image receiving module for receiving image data, a shallow feature extraction module for extracting coarse shallow information of the image, a global feature extraction module for extracting fine information with global information, an adaptive frequency fusion module for fusing difference information and common information of the image, and an image reconstruction module for reconstructing the fused image; the global feature extraction module includes two Laplacian pyramid encoders with the same structure, and the image reconstruction module includes a decoder with an inverse Laplacian pyramid structure;
[0055] Specifically, after receiving two kinds of image data, the image input module sends them to the shallow feature extraction module to perform a 3×3 convolution operation on the input image to obtain local information and shallow features of the image; the shallow information is sent to the global feature extraction module; the global feature extraction module includes two Laplace pyramid encoders with the same structure, which continuously downsample and Laplace differencing the image, decompose feature maps of different frequencies, and then send them to the adaptive frequency fusion module; the adaptive frequency fusion module includes an image embedding module, a difference information fusion module, two public information fusion modules and a refinement module, which fuses images of the same frequency and sends them to the image reconstruction module for reconstruction after fusion. The image reconstruction module includes a decoder with an inverse Laplace pyramid structure, which continuously upsamples and residually connects the image and then outputs the fused image.
[0056] Step 3, first-stage network training: Use labeled data from the multi-exposure dataset and the multi-focus dataset for multi-task supervised pre-training to enhance the model’s ability to extract complementary features from multi-exposure and multi-focus tasks;
[0057] Specifically, we use labeled data of multi-exposure and multi-focus tasks for multi-task supervised pre-training to enhance the model's ability to extract complementary features from multi-exposure and multi-focus tasks. In addition, labeled data helps the model converge better at this stage. We select a composite loss function for image fusion and image reconstruction. For image fusion loss, we select the structural similarity index metric (SSIM) loss. We use L 1 The norm is used as the reconstruction loss, and the total loss is calculated as a weighted combination of the above losses. Start network training and minimize the loss function value of the fusion result until the number of training times reaches the initial set threshold or the value of the loss function reaches the preset range. The network model is considered to have been trained and the network model parameters are saved.
[0058] Step 4, second-stage network training: Use unlabeled data from the infrared and visible light image datasets for unsupervised training to enhance the generalization ability of the model while saving data costs; use labeled data from the multi-exposure dataset and multi-focus dataset for supervised training to enhance the model's ability to retain source image details;
[0059] Specifically, semi-supervised fine-tuning is performed using labeled and unlabeled data. Visible light and infrared image data of unlabeled data are used to train infrared and visible light image fusion, and then the multi-task supervised training in the first stage is still used to enhance unsupervised learning. Unsupervised training improves the generalization ability of the model in tasks with scarce labeled data such as infrared and visible light fusion. At the same time, supervised training ensures the model's ability to maintain detail information, thereby helping the learning of unlabeled data. The total loss function in the second stage is a weighted sum of a supervised loss function and an unsupervised loss function, where the supervised loss function directly uses the loss function of the first stage, and the unsupervised loss function also includes the image fusion loss function and the image reconstruction loss function. The fusion loss includes the intensity loss function, the texture loss function and the result loss function. The fusion loss is obtained by weighting these three losses, and the reconstruction loss function uses L 1 Norm.
[0060] Step 5, determine the evaluation index: select the loss function to evaluate the quality of image fusion, and determine the evaluation index to evaluate the performance of the network.
[0061] Step 6, solidify the network model: After completing the adjustment of the network model, fix the network parameters and determine the final multimodal image fusion model; if image fusion tasks are required later, the image to be fused can be directly input into the network model to obtain the fusion result.
[0062] Furthermore, in step 2, the encoder and the decoder form a Laplacian pyramid structure; the encoder obtains the lowest-scale feature Then upsample and calculate the residual with the features of the same scale to generate high-frequency features The structure of the decoder is completely symmetrical with that of the encoder;
[0063] The adaptive frequency fusion module includes an image embedding module, a difference information fusion module, two common information fusion modules and a refinement module. The image embedding module divides the image into patches, introduces position encoding after linear projection, and outputs a sequence containing patch embedding vectors as the input of the subsequent model. The embedding process can be expressed as:
[0064]
[0065] Introduce the cross attention architecture and build a difference information fusion module, using the frequency features decomposed by the encoder as and As input, and output difference information features; and The division into n local feature fragments is as follows:
[0066]
[0067] in and And s = h × w;
[0068] Using a linear layer to transform feature snippets into query Q, key K and value V, the linear projection can be expressed as:
[0069] Q i =LinearQ(Q i ),K i =LinearK(K i ),V i =LinearV(V i );
[0070] Where i = 1, ..., n, Linear(·) represents a linear projection operation shared between different fragments;
[0071] The dot product attention layer is used to calculate the similarity matrix of the query Q and the key K, and then multiplied by the value V to infer the relevance between Q and V; it is expressed as:
[0072]
[0073] Among them, d k It is a scaling factor that can alleviate the convergence of the softmax function to the gradient minimum region when the dot product increases;
[0074] The difference information between Q and V is obtained by removing the common information; it is expressed as:
[0075] DV = Linear (V - CMV);
[0076] Inject the difference information into Q; expressed as:
[0077]
[0078] Among them, LN(·) represents the normalization layer, MLP(·) represents the multi-layer perceptron, and F dm is the output of the difference information fusion module;
[0079] Two identical information fusion modules are introduced, and the output fragments of the difference information fusion module are used to provide Q, while The fragment provides K and V, so F dm and The common information between them can be expressed as:
[0080]
[0081] Then F dm and Public information between CM joins Fdm , expressed as:
[0082]
[0083] Among them, F cm represents the output of the first public information fusion module;
[0084] F cm and The public information between them is injected into the fusion feature to further enrich the fusion feature;
[0085] The fused features are input into the refinement module for further feature fusion. The refinement module consists of convolutional layer 1, mixing block, and convolutional layer 2, and then outputs the fusion result after residual connection and Solftmax.
[0086] Embodiment 2:
[0087] A multimodal image fusion method based on SwinTransformer specifically includes the following steps:
[0088] Step 1, prepare training data: select the multi-focus datasets as RealMFF and MFI-WHU. The MFI-WHU dataset contains 190 samples, 96 of which are randomly selected for training. The RealMFF dataset contains 710 real samples. Since the sample distributions in the two datasets are very different, 96 samples are randomly selected from the RealMFF dataset for training. This ensures that the number and distribution of samples in the two datasets are basically the same, and the model trained with these samples will not overfit to a specific dataset. The multi-exposure dataset is the SICE dataset. The SICE dataset contains 229 samples, and 192 samples are randomly selected for training. In order to save computing resources, the image resolution of the first stage training is limited to 256×256. These datasets are used for supervised training. The datasets used for the visible light and infrared image fusion task are RoadScene and TNO datasets. The RoadScene dataset contains nearly 1,000 pairs of infrared and visible light image data pairs, and the TNO dataset contains about 2,400 pairs of infrared and visible light image data pairs. Similarly, in order to save computing resources, the image resolution of the second stage training is limited to 64×64. These data are used for unsupervised training.
[0089] Step 2: Build a network model: Figure 2As shown in the figure, the network model mainly includes two encoders, an adaptive frequency fusion module and a shared decoder; the two encoders take two source images as input respectively and decompose the source images into low-frequency features and high-frequency feature components; the adaptive frequency fusion module is used to effectively further fuse the extracted features; a shared decoder is used for image reconstruction and image fusion, and the structure of the decoder is completely symmetrical with the encoder.
[0090] The encoder consists of convolutional layer 1, hybrid block 1, hybrid block 2, hybrid block 3 and hybrid block 4. The convolutional layer is used to extract shallow features, and the hybrid block repeatedly extracts features from the source image and downsamples it to obtain frequency features at different levels. The structures of the two encoders are exactly the same. In contrast, the decoder extracts features layer by layer from the features of the encoder and upsamples them to reconstruct the image. The encoder and decoder form a Laplacian pyramid structure.
[0091] The encoder obtains the lowest-scale features Then upsample and calculate the residual with the features of the same scale to generate high-frequency features The structure of the decoder is completely symmetrical with that of the encoder;
[0092] Adaptive frequency fusion module such as Figure 3 As shown in Figure 1, the adaptive frequency fusion module consists of an image embedding module, a difference information fusion module, two common information fusion modules and a refinement module. The image embedding module divides the image into patches, introduces position encoding after linear projection, and outputs a sequence containing patch embedding vectors as the input of the subsequent model. The embedding process can be expressed as:
[0093]
[0094] In order to effectively obtain the difference features between images, a cross-attention architecture is introduced to construct a difference information fusion module, such as Figure 4 As shown in the figure, the frequency characteristics decomposed by the encoder are as follows and is input and outputs the difference information features. Specifically, and The division into n local feature fragments is as follows:
[0095] in and And s=h×w.
[0096] Next, a linear layer is used to transform these feature fragments into query Q, key K and value V. The linear projection can be expressed as:
[0097] Q i=LinearQ(Q i ),K i =LinearK(K i ),V i =LinearV(V i )
[0098] Where i=1,...,n, Linear(·) represents a linear projection operation shared among different fragments.
[0099] In order to extract the common information of the features of the two images and consider the long-range feature relationship, the dot product attention layer is used to calculate the similarity matrix of the query Q and the key K, and then multiplied by the value V to infer the correlation between Q and V. This process can be expressed as:
[0100]
[0101] Among them, d k It is a scaling factor that alleviates the problem of the softmax function converging to the gradient minimum region when the dot product increases.
[0102] Then, the difference information between Q and V is obtained by removing the common information. This process can be expressed as:
[0103] DV = Linear (V - CMV);
[0104] In order to Obtain complementary information and inject the difference information into Q, which is specifically expressed as:
[0105]
[0106] Among them, LN(·) represents the normalization layer, MLP(·) represents the multi-layer perceptron, and F dm It is the output of the difference information fusion module.
[0107] If we only extract the difference information of feature maps, it is easy to lose the background details, so we introduce two modules with the same information fusion. Figure 5 Specifically, the output fragment of the difference information fusion module is used to provide Q, and The fragment provides K and V, so F dm and The common information between them can be expressed as:
[0108]
[0109] Then F dm and Public information between CM joins F dm , this process can be expressed as:
[0110]
[0111] Among them, F cm Represents the output of the first common information fusion module.
[0112] After that, you need to add F cm and The common information between them is injected into the fusion feature to further enrich the fusion feature. The process is the same as the above process.
[0113] Finally, the fused features are input into the refinement module for further feature fusion. The refinement module is as follows: Figure 6 As shown in the figure, the convolution layer 1, the mixing block, and the convolution layer 2 are connected through residual connection and then Solftmax to output the fusion result.
[0114] Among them, the composition structure of all hybrid blocks is the same, such as Figure 7 As shown in Figure 1, the hybrid block consists of a convolutional layer, a SwinTransformer block, and a residual connection of the convolutional layer, where the convolution kernel size of all convolutional layers is 3×3, the depth of the SwinTransformer module is 2, the window size is 2, and the number of heads of the multi-head attention is 6.
[0115] Step 3: Perform the first stage of training. The first stage is high exposure and multi-focus fusion tasks. For these two tasks, the loss functions include image fusion loss and image reconstruction loss function. The structural similarity index measure (SSIM) is used as the image fusion loss function. L 1 The norm is used as the reconstruction loss function, and the loss function of the super exposure task can be summarized as:
[0116]
[0117] in is the fusion function, is the reconstructed source image, λ 1 is the trade-off parameter, SSIM(·) is the SSIM function, P·P 1 YesL 1 Norm. The loss function of the multi-focus task is as follows:
[0118]
[0119] Step 4, conduct the second stage of training. After the first stage of training is completed, fix the parameters of the encoder and decoder, save the image fusion module of the first stage to handle high exposure and multi-focus tasks, and train the second set of image fusion modules. In this stage, use the unlabeled infrared and visible light image pair datasets to perform unsupervised training on the model. However, it is not enough to rely solely on the unlabeled data of the infrared and visible light image pair datasets for unsupervised training. It is also necessary to combine the supervised learning of the first stage and use weight coefficients to balance the weights of multiple tasks in this training stage.
[0120] The total loss function of the second stage training is a weighted sum of a supervised loss function and an unsupervised loss function:
[0121] L S2 =L super +βL unsuper ;
[0122] Where β is a positive trade-off parameter. The loss function used in the first stage is directly used as the loss function of supervised training. The loss function of unsupervised learning also includes image fusion loss function and image reconstruction loss function. The unsupervised loss function is as follows:
[0123] L unsuper =L fuse +β 1 L recon ;
[0124] where β 1 is a weight coefficient. The fusion loss mainly consists of three parts: intensity loss function, texture loss function and structure loss function:
[0125]
[0126] Among them I 1u and I 2u is an unlabeled source image is the corresponding fusion result, β 2 , β 3 is a weight coefficient. The strength loss can be calculated as:
[0127]
[0128] Where MAX(·) is the element-wise maximum operation. The texture loss in the gradient domain can be calculated as:
[0129]
[0130] where |·| is the absolute value function and ▽ represents the gradient operator. 1norm so that the model can decompose the unlabeled source image into multi-frequency features:
[0131]
[0132] in and is the reconstructed unlabeled source image.
[0133] The hyperparameters in the loss function are set to λ=0, λ 1 =0.5,β=0.1,β 1 =1.25,β 2 =1,β 3 =0.5, the initial learning rate of the first stage is set to 1×10^-4 and halved after 200 cycles of training, and the initial learning rate of the second stage is set to 3×10^-5.
[0134] Step 5, select appropriate evaluation indicators: peak signal-to-noise ratio, structural similarity index, information entropy, mutual information, and root mean square error. The peak signal-to-noise ratio is used to calculate the difference between the fused image and the original reference image. The higher the peak signal-to-noise ratio, the better the image quality; the structural similarity index considers the similarity of brightness, contrast, and structural information, and reflects the visual quality of the image. The closer to 1, the more similar it is; information entropy is used to measure the complexity of the image. The larger the information entropy, the more information the image contains; mutual information calculates the correlation between the fused image and the source image, reflecting how much useful information of the original image is contained in the fused image. The higher the mutual information, the better the fusion effect; the root mean square error measures the average error between the original image and the fused image. The lower the better. The definitions of peak signal-to-noise ratio, structural similarity index, information entropy, mutual information, and root mean square error are as follows:
[0135]
[0136] Among them, μ x , μ y Represent the mean of images x and y respectively, and Represents the variance of images x and y, σ xy is the covariance of images x and y, C 1 and C 2 is a constant; p(x i ) is the value of each pixel x in the image i where n is the total number of different pixel values in the image; p(I(i,j),K(i,j)) is the joint probability distribution, p(I(i,j)) and p(K(i,j)) are separate probability distributions; N and M are the sizes of the image, I(i,j) and K(i,j) are the pixel values of the original image and the fused image, respectively.
[0137] Table 1 Performance comparison of different models
[0138]
[0139] Table I shows the objective evaluation of 20 pairs of images in the MSRS dataset using four different fusion methods. This method performs best in EN, SSIM and RMSE indicators, and performs second best in MI and PSNR indicators. Combining subjective and objective analysis, it can be concluded that this method outperforms other methods in both visual perception and objective indicators, and achieves the best fusion performance.
[0140] Step 6, solidify the network model: After completing the network model adjustment, fix the network parameters and determine the final multimodal image fusion model; if image fusion tasks are required later, the image to be fused can be directly input into the network model to obtain the fusion result.
[0141] In the above embodiments, implementations such as convolution, SwinTransformer, activation function, normalization, normalized exponential function, matrix multiplication operation, and corresponding element multiplication are algorithms well known to those skilled in the art, and the specific processes and methods can be found in corresponding textbooks or technical literature.
[0142] In addition, the present invention may be implemented as a system, method or computer program product. Therefore, the present disclosure may be specifically implemented in the following forms, namely: it may be complete hardware, it may be complete software (including firmware, resident software, microcode, etc.), or it may be a combination of hardware and software.
[0143] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A multimodal image fusion method based on SwinTransformer, characterized in that: The following steps are involved: Step 1, prepare training data: select multi-focus dataset, multi-exposure dataset, infrared image dataset and visible light image dataset, divide the dataset into training set, validation set and test set, and pre-process the original images and their corresponding labels in the multi-focus dataset and multi-exposure dataset; Step 2, constructing a network model: the network model includes an image receiving module for receiving image data, a shallow feature extraction module for extracting coarse shallow information of the image, a global feature extraction module for extracting fine information with global information, an adaptive frequency fusion module for fusing difference information and common information of the image, and an image reconstruction module for reconstructing the fused image; the global feature extraction module includes two Laplacian pyramid encoders with the same structure, and the image reconstruction module includes a decoder with an inverse Laplacian pyramid structure; Step 3, first-stage network training: Use labeled data from the multi-exposure dataset and the multi-focus dataset for multi-task supervised pre-training to enhance the model’s ability to extract complementary features from multi-exposure and multi-focus tasks; Step 4, second-stage network training: Use unlabeled data from the infrared and visible light image datasets for unsupervised training to enhance the generalization ability of the model while saving data costs; use labeled data from the multi-exposure dataset and multi-focus dataset for supervised training to enhance the model's ability to retain source image details; Step 5, determine the evaluation index: select the loss function to evaluate the quality of image fusion, and determine the evaluation index to evaluate the performance of the network.
2. The multimodal image fusion method based on SwinTransformer according to claim 1, characterized in that: The following steps are also included: Step 6, solidifying the network model: after completing the adjustment of the network model, fix the network parameters and determine the final multimodal image fusion model; If image fusion tasks are required later, the image to be fused is directly input into the network model to obtain the fusion result.
3. The multimodal image fusion method based on SwinTransformer according to claim 1, characterized in that: In step 2, the image input module receives the two types of image data and sends them to the shallow feature extraction module to perform a 3×3 convolution operation on the input image to obtain local information and shallow features of the image; the shallow information is sent to the global feature extraction module; The global feature extraction module includes two Laplacian pyramid encoders with the same structure, which continuously downsample and Laplacian differencing the image, decompose the feature maps of different frequencies, and then send them to the adaptive frequency fusion module; The adaptive frequency fusion module consists of an image embedding module, a difference information fusion module, two common information fusion modules and a refinement module. It fuses images of the same frequency and sends them to the image reconstruction module for reconstruction. The image reconstruction module includes a decoder with an inverse Laplacian pyramid structure, which continuously upsamples and residually connects the image and then outputs the fused image.
4. The multimodal image fusion method based on SwinTransformer according to claim 3, characterized in that: In step 2, the encoder and decoder form a Laplacian pyramid structure; the encoder obtains the lowest-scale feature Then upsample and calculate the residual with the features of the same scale to generate high-frequency features The structure of the decoder is completely symmetrical with that of the encoder; The adaptive frequency fusion module includes an image embedding module, a difference information fusion module, two common information fusion modules and a refinement module. The image embedding module divides the image into patches, introduces position encoding after linear projection, and outputs a sequence containing patch embedding vectors as the input of the subsequent model. The embedding process can be expressed as: where F token is a sequence of patch embedding vectors, and PE(·) represents the process of segmenting the image into patches and introducing position encoding after linear projection. Introduce the cross attention architecture and build a difference information fusion module, using the frequency features decomposed by the encoder as and As input, and output difference information features; F1 token and The division into n local feature fragments is as follows: Among them, F1 token and And s = h × w; Using a linear layer to transform feature snippets into query Q, key K and value V, the linear projection can be expressed as: Q i =LinearQ(Q i ),K i =LinearK(K i ),V i =LinearV(V i ); Where i = 1, ..., n, Linear(·) represents a linear projection operation shared between different fragments; The dot product attention layer is used to calculate the similarity matrix of the query Q and the key K, and then multiplied by the value V to infer the relevance between Q and V; it is expressed as: Among them, d k It is a scaling factor that can alleviate the convergence of the softmax function to the gradient minimum region when the dot product increases; The difference information between Q and V is obtained by removing the common information; it is expressed as: DV = Linear (V - CMV); Inject the difference information into Q; expressed as: Among them, LN(·) represents the normalization layer, MLP(·) represents the multi-layer perceptron, and F dm is the output of the difference information fusion module; Two identical information fusion modules are introduced, and the output fragments of the difference information fusion module are used to provide Q, while The fragment provides K and V, so F dm and The common information between them can be expressed as: Then F dm and F1 token Public information between CM joins F dm , expressed as: Among them, F cm represents the output of the first public information fusion module; F cm and F1 token The common information between them is injected into the fusion feature to further enrich the fusion feature; The fused features are input into the refinement module for further feature fusion. The refinement module consists of convolutional layer 1, mixing block, and convolutional layer 2, and then outputs the fusion result after residual connection and Solftmax.
Citation Information
Patent Citations
A universal multimodal image fusion method and device
CN117173525B
Light-weight multi-scale infrared image super-resolution reconstruction method
CN114092330A
Image fusion method based on RFN-Nest
CN114742739A
Infrared and visible light image fusion method combining Transform and CNN double encoders
CN117314808A
Global interaction hyperspectral multispectral cross-modal fusion method with spectral fidelity
CN117911830A
Cited By
Image fusion system and method based on multi-branch feature fusion and attention mechanism
CN120635658A
Multi-modal image fusion method based on dynamic pseudo supervision and semantic guidance
CN121788981A
A multi-modal image fusion method based on dynamic pseudo-supervision and semantic guidance
CN121788981B