Lightweight multi-modal image fusion method based on knowledge distillation technology

By applying knowledge distillation technology in image fusion, the ability of large language models is transferred to the lightweight student network, solving the problem of large computing overhead of existing methods, achieving high-quality image fusion and significant resource savings.

CN119992273APending Publication Date: 2025-05-13HEFEI INST OF TECH INNOVATION ENG CHINESE ACAD OF SCI
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510199586.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing text-guided image fusion method has a large computing overhead and a huge model scale, which cannot effectively reduce resource requirements, resulting in disproportionate performance improvement.

Method used

The lightweight multimodal image fusion method based on knowledge distillation technology is adopted to transfer the semantic understanding ability of large language models to the lightweight student network through the teacher-student network architecture, reducing computing overhead.

Benefits of technology

It realizes that high-quality image fusion effect is maintained without relying on large-scale models, significantly reducing calculation overhead, reducing the amount of model parameters to 10% of the teacher network, and maintaining performance above 90%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992273A_ABST
    Figure CN119992273A_ABST
Patent Text Reader

Abstract

The invention relates to a lightweight multi-modal image fusion method based on a knowledge distillation technology. Compared with the prior art, the lightweight multi-modal image fusion method based on the knowledge distillation technology overcomes the defects that an image fusion method based on text guidance is large in calculation overhead and large in model scale. The method comprises the following steps: obtaining and preprocessing a source image; constructing a lightweight image fusion model; training a lightweight image fusion model; obtaining a to-be-fused image; and obtaining a multi-model image fusion result. According to the method, through a teacher-student network architecture and a customized prior distillation process, the semantic understanding ability of a large language model is successfully transferred to a lightweight student network, and under the condition that text guidance in a reasoning stage is not needed, relatively high fusion quality is still kept, and meanwhile, the calculation overhead is remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image fusion processing, and in particular to a lightweight multimodal image fusion method based on knowledge distillation technology. Background Art

[0002] Image fusion plays an important role in the field of digital image processing. Taking visible light-infrared image fusion as an example, visible light images contain rich color detail information, which is easy for humans to understand and interpret; while infrared images contain radiation-based feature information, which has unique advantages in target detection and low-light conditions. By fusing the complementary information of these two modalities, a fused image with high-quality visual effects and enhanced detection performance can be generated.

[0003] However, in the actual imaging process, environmental and equipment limitations often lead to a degradation in the quality of the source image: visible light images may have problems such as low resolution and blur; infrared images are easily affected by multiple noises such as thermal noise, electronic noise, and environmental noise.

[0004] There are several main solutions to these problems: Traditional fusion methods, such as FusionGAN, U2Fusion, and SwinFusion, directly fuse the source images, but cannot effectively distinguish between noise and useful information, resulting in poor performance when processing low-quality input images; some methods rely on manual preprocessing to enhance the quality of source images, but lack flexibility. Recently emerged text-guided methods, such as TextFusion and Text-IF, use multimodal large language models (MLLMs) to generate text descriptions for source images (such as "low light"), enabling the model to adaptively distinguish between image content and degradation, thereby improving image quality. These methods enhance the model's understanding of image content by integrating semantic prior knowledge, effectively promoting the fusion process.

[0005] However, text-guided methods also have obvious flaws. Integrating large language models into the fusion process will bring huge resource overhead. In most cases, users only need to fuse images for downstream tasks (such as detection or segmentation) without a lot of interaction; the computational requirements of large language models are often ten to thousands of times that of the fusion module itself. Even the relatively lightweight visual-language model CLIP has about 5 times the number of parameters as the fusion module. This computational cost is disproportionate to the scale of the task. Due to the limitations of the quality and quantity of the dataset, the performance improvement brought by integrating large models is marginal relative to the significant decrease in efficiency.

[0006] Therefore, there is an urgent need for a technology that can achieve degradation-aware fusion and enhanced semantic understanding without relying on large-scale models. Summary of the invention

[0007] The purpose of the present invention is to solve the defects of high computational overhead and large model size of the text-guided image fusion method in the prior art, and to provide a lightweight multimodal image fusion method based on knowledge distillation technology to solve the above problems.

[0008] In order to achieve the above object, the technical solution of the present invention is as follows:

[0009] A lightweight multimodal image fusion method based on knowledge distillation technology includes the following steps:

[0010] Acquisition and preprocessing of source images: Acquire the multimodal source images to be fused and crop them according to the set size to generate a training set. The multimodal source images include visible light image pairs, infrared image pairs, MRI image pairs or CT image pairs;

[0011] Constructing a lightweight image fusion model: A lightweight image fusion model is constructed based on the teacher-student network architecture, in which the teacher network integrates the prior knowledge of the large language model, and the student network learns the fusion ability of the teacher network through knowledge distillation;

[0012] Training of lightweight image fusion model: Use the training set to train the lightweight image fusion model;

[0013] Acquire the image to be fused: acquire the multimodal image pair to be fused, and perform multimodal image pair registration and standardization processing;

[0014] Obtain multi-model image fusion results: Input the processed image pairs to be fused into the trained lightweight image fusion model to obtain multi-model image fusion results.

[0015] The construction of the lightweight image fusion model includes the following steps:

[0016] A lightweight image fusion model is constructed based on a teacher-student network architecture. The lightweight image fusion model includes a teacher network and a student network.

[0017] The teacher network is set to include a text-guided module, an encoder module, and a decoder module. The text-guided module uses a large language model and a CLIP model to generate semantic guidance information for the source image. The encoder module processes input images of different modalities through a two-stream Transformer, including a transposed self-attention module and a spatial self-attention module. The decoder module gradually reconstructs the fused image through text-guided feature modulation.

[0018] The student network is set to include a feature extraction module, a space-channel cross-fusion module and a reconstruction module. The feature extraction module uses a lightweight convolutional network to extract the features of multimodal images. The space-channel cross-fusion module cross-fuses the features in the spatial dimension and the channel dimension. The reconstruction module reconstructs the fused features into the final fused image.

[0019] Setting up the text guide module, encoder module and decoder module;

[0020] Set up feature extraction module, space-channel cross fusion module and reconstruction module.

[0021] The training of the lightweight image fusion model includes the following steps:

[0022] The loss function combination of the lightweight image fusion model is set as follows:

[0023] Basic loss:

[0024]

[0025] in, The structural similarity is measured by the SSIM index. preserves the maximum intensity information between source images, Use the Sobel operator to ensure gradient consistency. Maintain color fidelity in YCbCr space, weight λ t Dynamically adjust based on text guidance;

[0026] The detailed expressions of each loss component are as follows:

[0027]

[0028] Where H and W are the height and width of the image, I f is the fused image, and denote visible light and infrared guidance images, respectively;

[0029]

[0030] Among them, SSIM(·) calculates the structural similarity index, δ ir (t) is the infrared-guided text dependency weight;

[0031]

[0032] in, represents the gradient operator implemented in the horizontal and vertical directions using the Sobel filter;

[0033]

[0034] Among them, F CbCr Represents the conversion function from RGB to YCbCr;

[0035] These losses work together to ensure high-quality image fusion: retains the maximum intensity information from both source images, Preserve structural details through the SSIM metric, Ensure edge consistency through gradient preservation, Keeping the natural color appearance, the text-dependent weights α(t) allow the contribution of each loss component to be dynamically adjusted according to the specific fusion requirements described in the textual cues;

[0036] Distillation losses:

[0037] Feature-level adaptation loss: Align the multi-scale features of the teacher network and the student network through the adapter,

[0038]

[0039] in and Represents the i-th layer feature map of the student and teacher networks, Adapter i is the corresponding layer adapter, ∥·∥ 2 represents the L2 norm;

[0040] Output-level supervision loss: maintain the consistency of the final output,

[0041]

[0042] The total loss function is:

[0043]

[0044] where λ 1 -λ 3 ∈[0.1,1.0] is the weight, the default is 1:1:1;

[0045] Configure the optimizer:

[0046] The Adam optimizer is used, with an initial learning rate of 1e-4, a weight decay coefficient of 0.01, and a momentum parameter β = (0.9, 0.999).

[0047] Use the cosine annealing scheduler to dynamically adjust the learning rate. The formula is:

[0048]

[0049] Among them, η max =1e-4 is the maximum learning rate, ηmin =1e-6 is the minimum learning rate, T max is the maximum number of training cycles;

[0050] Apply gradient clipping to prevent gradient explosion, and set the threshold to 2.0;

[0051] Teacher network training includes the following steps:

[0052] Initialize the teacher network parameters and freeze the large language model parameters.

[0053] Set the loss function combination:

[0054]

[0055] in is the maximum strength loss, is the YCbCr color space fidelity loss,

[0056] Iterate over each image in the dataset,

[0057] Data enhancement processing,

[0058] The input image is randomly cropped to 128x128,

[0059] Apply color dithering enhancement, adjust brightness by ±0.1, and contrast by ±0.2.

[0060] Flip horizontally / vertically with a probability of 0.5,

[0061] Multimodal feature fusion:

[0062] The large language model generates a text description T, which is then passed through the CLIP text encoder to obtain F. text ,

[0063] Extract image features F through XRestormer encoder A F B , and F is used uniformly below. i ,

[0064] Up and down sampling through convolution G = Conv (F i ) is used for modulation decoding process,

[0065] Calculate the 4-layer attention fusion feature using the operator sparse spatial attention to obtain the text-guided feature map F guide =SpatialAttention(F i ,F text ),

[0066] Use the decoder to get the fused image I fused , used to calculate

[0067] Parameter update strategy:

[0068] Gradient clipping is used with a threshold of 2.0.

[0069] Use Adam optimizer, β = (0.9, 0.999),

[0070] Weight decay coefficient 0.01;

[0071] The student network training includes the following steps: loading the pre-trained teacher network parameters, freezing the teacher network parameters,

[0072] Knowledge distillation,

[0073] Feature-level adaptation training,

[0074] Enable only loss, learning rate 1e-4,

[0075] Feature adaptation strategy:

[0076] Layer 4: 384→128 channels, 1×1 convolution for dimensionality reduction,

[0077] Layer 3: 192 → 64 channels, 3×3 convolutional space-channel conversion,

[0078] Layer 2: 96→32 channels, 5×5 convolution to enhance local features,

[0079] Layer 1: 48 → 16 channels, 7×7 convolutions to capture context,

[0080] Joint training phase:

[0081] Enable Combined loss:

[0082]

[0083] Use the same update strategy as the teacher network;

[0084] Training process:

[0085] Input visible light image I vis and infrared image I ir ,

[0086] The teacher network generates fused images through text guidance I teacher ,

[0087] The student network processes the image independently to obtain I student ,

[0088] Compute feature-level and output-level distillation losses,

[0089] Back-propagation updates the student network parameters.

[0090] The setting of the text guiding module, the encoder module and the decoder module comprises the following steps:

[0091] Set up the text guidance module: Generate image description text T through a large language model and use the CLIP text encoder to convert text features F text ;

[0092] To set up the encoder module:

[0093] Use the transposed self-attention module to process,

[0094]

[0095] Among them, Q T , K, V are the products of input X and Wq, Wk, Wv respectively;

[0096] Use the spatial self-attention module to process, divide it into 8 windows, and add their respective position encodings;

[0097]

[0098] O = Conv(A·V),

[0099] Here, PosEmb uses RelPosEmb and sets the number of learnable parameter matrices W and H.

[0100] The XY coordinates of the internal position of Q are multiplied, and the two directions are respectively cosine-encoded or flattened and then cosine-encoded. Conv represents convolution operation, and the convolution kernel is 3x3. The SoftMax operation is to exponentially smooth all the results at the channel layer.

[0101] Feature downsampling: The connection between layers is completed by feature downsampling; the downsampling expression is as follows:

[0102]

[0103] The separate form used for convolution is implemented here;

[0104] Repeat the transposed self-attention module processing and the spatial self-attention module processing several times, doubling d each time;

[0105] Set up the decoder module:

[0106] Perform feature upsampling;

[0107] F up =Conv(F fused );

[0108] Perform detail enhancement:

[0109] F enhanced =F up +Conv(SSA(F up ))+Conv(TSA(F up ))

[0110] Here, SSA and TSA represent the formulas for extracting A and Y above. This module is implemented by repeating the operation twice, halving d each time.

[0111] The setting of the feature extraction module, the space-channel cross fusion module and the reconstruction module comprises the following steps:

[0112] Set up feature extraction module: extract features of different images separately:

[0113] F vis =Conv v (I vis )

[0114] F ir =Conv i (I ir );

[0115] Set up the space-channel cross-fusion module:

[0116] The top three layers use spatial dimension fusion:

[0117]

[0118] F fused =SSA(M s )+TSA(M s ),

[0119] Here F vis 、F ir From the feature extraction above, each layer has one copy, ⊙ is the element product, σ, The F of Clip is connected by two fully connected layers. text get,

[0120] In the sub-model, M s =Conv([F vis ,F ir ]), no T is required, and the d of each layer is reduced by 3 times. Because there are 4 layers, the corresponding layer square parameters of different layers are reduced by 2, so the final model is reduced to 0.1 times, and σ, information, that is, no large model and CLIP are required, and the number of parameters is reduced to 0.01 times;

[0121] The lowest layer uses cross-attention feature fusion:

[0122] F fused =CrossAttention(F vis ,F ir )

[0123] Here CrossAttention exchanges A and Y of the two, and then implements it according to the encoder's formula;

[0124] Setting up the reconstruction module involves the following steps:

[0125] Perform residual connection enhancement:

[0126] F enhanced =F up +Conv(SSA(F up ))+Conv(TSA(F up ))

[0127] Here F up It is the highest level;

[0128] Full connection reconstruction: I fused =Linear(F enhanced )

[0129] Linear is a fully connected layer.

[0130] Beneficial Effects

[0131] Compared with the prior art, the lightweight multimodal image fusion method of the present invention based on knowledge distillation technology successfully transfers the semantic understanding ability of the large language model to the lightweight student network through the teacher-student network architecture and the customized prior distillation process. The present invention maintains a high fusion quality without the need for text guidance in the reasoning stage, while significantly reducing the computational overhead.

[0132] The present invention reduces the number of model parameters to 10% of the teacher network (excluding the CLIP model, which itself requires 5 times the number of parameters of the fusion module), maintains more than 90% of the performance, and completely eliminates the dependence on large language models in the inference stage.

[0133] The spatial-channel cross-fusion module (SCFM) proposed in the present invention can make full use of text prior information in both spatial and channel dimensions: in the spatial dimension, long-range dependencies are captured through the overlapping window attention mechanism, and in the channel dimension, accurate feature fusion is achieved through feature adaptive modulation. The synergy of the two dimensions improves the model's ability to process objects of different scales.

[0134] The prior distillation loss function designed in the present invention dually constrains the student network from the feature level and the output level. The distillation at the feature level ensures that the student network learns the feature processing method of the teacher network. The supervision at the output level ensures the quality of the final fusion result. The distillation of multi-scale features improves the generalization ability of the model.

[0135] The method of the present invention is applicable to a variety of image fusion scenarios, such as visible light-infrared image fusion, medical image fusion, etc. The processing process is fully automated, no human intervention is required, the operation efficiency is high, and it is suitable for deployment on resource-constrained devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0136] Figure 1 is a method sequence diagram of the present invention;

[0137] Figure 2 It is a structural schematic diagram of the lightweight image fusion model involved in the present invention;

[0138] Figure 3 The figure is a comparison diagram of the effects of the present invention and the traditional method. DETAILED DESCRIPTION

[0139] In order to have a further understanding and recognition of the structural features and the effects achieved by the present invention, a preferred embodiment and accompanying drawings are used for detailed description as follows:

[0140] like Figure 1 As shown, the lightweight multimodal image fusion method based on knowledge distillation technology described in the present invention includes the following steps:

[0141] The first step is to obtain and preprocess the source images: obtain the multimodal source images to be fused, crop them according to the set size, and generate training sets and test sets. The multimodal source images include visible light image pairs, infrared image pairs, MRI image pairs, or CT image pairs.

[0142] The second step is to build a lightweight image fusion model: Figure 2 As shown in the figure, a lightweight image fusion model is constructed based on the teacher-student network architecture, in which the teacher network integrates the prior knowledge of the large language model, and the student network learns the fusion ability of the teacher network through knowledge distillation.

[0143] (1) A lightweight image fusion model is constructed based on a teacher-student network architecture. The lightweight image fusion model includes a teacher network and a student network.

[0144] (2) The teacher network is set to include a text-guided module, an encoder module and a decoder module. The text-guided module uses a large language model and a CLIP model to generate semantic guidance information for the source image. The encoder module processes input images of different modalities through a two-stream Transformer, including a transposed self-attention module and a spatial self-attention module. The decoder module gradually reconstructs the fused image through text-guided feature modulation.

[0145] (3) The student network is set to include a feature extraction module, a space-channel cross-fusion module and a reconstruction module. The feature extraction module uses a lightweight convolutional network to extract the features of the multimodal image. The space-channel cross-fusion module cross-fuses the features in the spatial dimension and the channel dimension. The reconstruction module reconstructs the fused features into the final fused image.

[0146] (4) Set the text guidance module, encoder module and decoder module.

[0147] Setting up the text guide module, encoder module, and decoder module includes the following steps:

[0148] A1) Setting up the text-guided module: Generate image description text T through a large language model, and use the CLIP text encoder to convert text features F text ;

[0149] A2) Set up the encoder module:

[0150] A21) is processed using the transposed self-attention module,

[0151]

[0152] Among them, Q T , K, V are the products of input X and Wq, Wk, Wv respectively;

[0153] A22) is processed using the spatial self-attention module, divided into 8 windows, and each window is encoded with its own position;

[0154]

[0155] O = Conv(A·V),

[0156] Here, PosEmb uses RelPosEmb and sets the number of learnable parameter matrices W and H.

[0157] The XY coordinates of the internal position of Q are multiplied, and the two directions are respectively added with cosine coding or flattened and then added with cosine coding. Conv represents convolution operation, and the convolution kernel uses 3x3. The SoftMax operation is to exponentially smooth all the results at the channel layer.

[0158] A23) Feature downsampling: The connection between layers is completed by feature downsampling; the downsampling expression is as follows:

[0159]

[0160] The separate form used for convolution is implemented here;

[0161] A24) Repeat the transposed self-attention module processing and the spatial self-attention module processing several times, doubling d each time;

[0162] A3) Set up the decoder module:

[0163] A31) performs feature upsampling;

[0164] F up =Conv(F fused );

[0165] A32) performs detail enhancement processing:

[0166] F enhanced =F up +Conv(SSA(F up ))+Conv(TSA(F up ))

[0167] Here, SSA and TSA represent the formulas for extracting A and Y above. This module is implemented by repeating the operation twice, halving d each time.

[0168] (5) Set the feature extraction module, space-channel cross fusion module and reconstruction module.

[0169] Setting the feature extraction module, the space-channel cross fusion module and the reconstruction module includes the following steps:

[0170] B1) Setting up feature extraction module: Extract features of different images separately:

[0171] F vis =Conv v (I vis )

[0172] F ir =Conv i (I ir );

[0173] B2) Setting up the space-channel cross-fusion module:

[0174] B21) The upper three layers use spatial dimension fusion:

[0175]

[0176] Ffused =SSA(M s )+TSA(M s ),

[0177] Here F vis 、F ir From the feature extraction above, each layer has one copy, ⊙ is the element product, σ, The F of Clip is connected by two fully connected layers. text get,

[0178] In the sub-model, M s =Conv([F vis ,F ir ]), no T is required, and the d of each layer is reduced by 3 times. Because there are 4 layers, the corresponding layer square parameters of different layers are reduced by 2, so the final model is reduced to 0.1 times, and σ, information, that is, no large model and CLIP are required, and the number of parameters is reduced to 0.01 times;

[0179] B22) The lowest layer uses cross-attention feature fusion:

[0180] F fused =CrossAttention(F vis ,F ir )

[0181] Here CrossAttention exchanges A and Y of the two, and then implements it according to the encoder's formula;

[0182] B3) Setting the reconstruction module includes the following steps:

[0183] Perform residual connection enhancement:

[0184] F enhanced =F up +Conv(SSA(F up ))+Conv(TSA(F up ))

[0185] Here F up It is the highest level;

[0186] Full connection reconstruction: I fused =Linear(F enhanced )

[0187] Linear is a fully connected layer.

[0188] The third step is training the lightweight image fusion model: using the training set to train the lightweight image fusion model.

[0189] (1) The loss function combination of the lightweight image fusion model is set as follows:

[0190] Basic loss:

[0191]

[0192] in, The structural similarity is measured by the SSIM index. preserves the maximum intensity information between source images, Use the Sobel operator to ensure gradient consistency. Maintain color fidelity in YCbCr space, weight λ t Dynamically adjust based on text guidance;

[0193] The detailed expressions of each loss component are as follows:

[0194]

[0195] Where H and W are the height and width of the image, I f is the fused image, and denote visible light and infrared guidance images, respectively;

[0196]

[0197] Among them, SSIM(·) calculates the structural similarity index, δ ir (t) is the infrared-guided text dependency weight;

[0198]

[0199] in, represents the gradient operator implemented in the horizontal and vertical directions using the Sobel filter;

[0200]

[0201] Among them, F CbCr Represents the conversion function from RGB to YCbCr;

[0202] These losses work together to ensure high-quality image fusion: retains the maximum intensity information from both source images, Preserve structural details through the SSIM metric, Ensure edge consistency through gradient preservation, Keeping the natural color appearance, the text-dependent weights α(t) allow the contribution of each loss component to be dynamically adjusted according to the specific fusion requirements described in the textual cues;

[0203] Distillation losses:

[0204] Feature-level adaptation loss: Aligning multi-scale features of teacher-student networks via adapters,

[0205]

[0206] in and Represents the i-th layer feature map of the student and teacher networks, Adapter i is the corresponding layer adapter, ∥·∥ 2 represents the L2 norm;

[0207] Output-level supervision loss: maintain the consistency of the final output,

[0208]

[0209] The total loss function is:

[0210]

[0211] where λ 1 -λ 3 ∈[0.1,1.0] is the weight, the default is 1:1:1.

[0212] (2) Configure the optimizer:

[0213] The Adam optimizer is used, with an initial learning rate of 1e-4, a weight decay coefficient of 0.01, and a momentum parameter β = (0.9, 0.999).

[0214] Use the cosine annealing scheduler to dynamically adjust the learning rate. The formula is:

[0215]

[0216] Among them, η max =1e-4 is the maximum learning rate, η min =1e-6 is the minimum learning rate, T max is the maximum number of training cycles;

[0217] Gradient clipping is applied to prevent gradient explosion, with a threshold of 2.0.

[0218] (3) Teacher model training includes the following steps:

[0219] Initialize the teacher network parameters and freeze the large language model parameters.

[0220] Set the loss function combination:

[0221]

[0222] in is the maximum strength loss, is the YCbCr color space fidelity loss,

[0223] Iterate over each image in the dataset,

[0224] Data enhancement processing,

[0225] The input image is randomly cropped to 128x128,

[0226] Apply color dithering enhancement, adjust brightness by ±0.1, and contrast by ±0.2.

[0227] Flip horizontally / vertically with a probability of 0.5,

[0228] Multimodal feature fusion:

[0229] The large language model generates a text description T, which is then passed through the CLIP text encoder to obtain F. text ,

[0230] Extract image features F through XRestormer encoder A F B , the following uniformly uses F i ,

[0231] Up and down sampling through convolution G = Conv (F i ) is used for modulation decoding process,

[0232] Calculate the 4-layer attention fusion feature using the operator sparse spatial attention to obtain the text-guided feature map F guide =SpatialAttention(F i ,F text ),

[0233] Use the decoder to get the fused image I fused , used to calculate

[0234] Parameter update strategy:

[0235] Gradient clipping (threshold 2.0) is used.

[0236] Using the Adam optimizer (β = (0.9, 0.999)),

[0237] The weight decay coefficient is 0.01.

[0238] (4) Student model training includes the following steps: loading pre-trained teacher network parameters, freezing teacher network parameters,

[0239] Knowledge distillation,

[0240] Feature-level adaptation training,

[0241] Enable only loss, learning rate 1e-4,

[0242] Feature adaptation strategy:

[0243] Layer 4: 384→128 channels, 1×1 convolution for dimensionality reduction,

[0244] Layer 3: 192 → 64 channels, 3×3 convolutional space-channel conversion,

[0245] Layer 2: 96→32 channels, 5×5 convolution to enhance local features,

[0246] Layer 1: 48 → 16 channels, 7×7 convolutions to capture context,

[0247] Joint training phase:

[0248] Enable Combined loss:

[0249]

[0250]

[0251] Use the same update strategy as the teacher model;

[0252] Training process:

[0253] Input visible light image I vis and infrared image I ir ,

[0254] The teacher network generates fused images through text guidance I teacher ,

[0255] The student network processes the image independently to obtain I student ,

[0256] Compute feature-level and output-level distillation losses,

[0257] Back-propagation updates the student network parameters.

[0258] The fourth step is to obtain the image to be fused: obtain the multimodal image pair to be fused, and perform multimodal image pair registration and standardization processing.

[0259] The fifth step is to obtain the multi-model image fusion result: input the processed image pair to be fused into the trained lightweight image fusion model to obtain the multi-model image fusion result.

[0260] In order to verify the effectiveness of the image fusion method based on knowledge distillation proposed in this paper, a large number of experiments were conducted on multiple data sets. They mainly include the following aspects:

[0261] 1. Infrared-visible light image fusion experiment

[0262] Experiments were conducted on three datasets: MSRS, M3FD, and RoadScene, and compared with the current mainstream image fusion methods. Table 1 shows the quantitative comparison results. The evaluation indicators include:

[0263] Information entropy (EN): used to measure the amount of information contained in an image.

[0264] Visual Information Fidelity (VIF): evaluates the perceptual quality of an image,

[0265] Gradient-based fusion quality (Q AB / F ):Evaluate the degree of preservation of edge information.

[0266] Table 1: Quantitative comparison with existing methods on infrared-visible image fusion tasks

[0267]

[0268]

[0269] The bold indicates the best result. The brackets are the non-training parameters used in inference (unit: million)

[0270] As can be seen from Table 1, the method proposed in the present invention has achieved the best or suboptimal results in multiple evaluation indicators. In particular, the small network after distillation not only maintains the same performance as the teacher network, but even exceeds the teacher network in some indicators, while the number of parameters is only one tenth of that of the teacher network. This shows that the knowledge distillation method proposed in the present invention can effectively transfer the performance of the large model to the small model.

[0271] 2. Medical image fusion experiment

[0272] In order to verify the generalization ability of the proposed method, we conducted experiments on medical image fusion tasks. The experiments included the fusion of medical images of three different modalities: PET-MRI, CT-MRI, and SPECT-MRI. Table 2 shows the quantitative comparison results. The evaluation indicators include:

[0273] Structural Similarity (SSIM): evaluates the degree of structural preservation of images.

[0274] Visual Information Fidelity (VIF): evaluates the perceptual quality of an image,

[0275] Gradient-based fusion quality (QAB / F ):Evaluate the degree of preservation of edge information.

[0276] Table 2: Quantitative comparison with existing methods on medical image fusion tasks

[0277]

[0278]

[0279] The bold indicates the best result, and the number of non-training parameters used during inference is in brackets (unit: million).

[0280] In the SPECT-MRI fusion task, the proposed method retains anatomical details well while maintaining functional information. In CT-MRI fusion, the proposed method achieves the highest SSIM score, indicating better structure preservation. The results of PET-MRI further confirm the effectiveness of the proposed method in processing multimodal medical images with different characteristics. In particular, the small network after distillation significantly reduces the computational requirements while maintaining these advantages, making it more suitable for deployment.

[0281] like Figure 3 As shown, Figure 3 This is a comparison chart about the fusion of visible light and infrared. It can be seen that our method (Troposed teacher, Troposed Distilled) effectively extracts edge contour information, such as the contour of the light in the upper left corner. At the same time, the method of the present invention retains better color information, and will not lose color information like U2Fusion and CDDFuse, nor will it have serious color cast when the text guides the biased structure like text-if. In addition, the distillation sub-model of the present invention successfully gets rid of the limitations of model size and large models and CLIP, and successfully restores the effect similar to the teacher model, but there is a problem of over-exerting in the color of details. The teacher model has a good structure and normal color, but the sub-model tends to over-focus on contour information, resulting in purple contours in some places. However, the sub-model still achieves performance comparable to that of the teacher model. The excessive focus on the color cast of the structure even allows the student model to surpass the main model in the running scores of structural classes such as vif and ssim in some cases.

[0282] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions only describe the principles of the present invention. The present invention may be subject to various changes and improvements without departing from the spirit and scope of the present invention. These changes and improvements fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the attached claims and their equivalents.

Claims

1. A lightweight multimodal image fusion method based on knowledge distillation technology, characterized in that: The following steps are involved: 11) Acquisition and preprocessing of source images: Acquire multimodal source images to be fused and crop them according to a set size to generate a training set. The multimodal source images include visible light image pairs, infrared image pairs, MRI image pairs or CT image pairs; 12) Constructing a lightweight image fusion model: A lightweight image fusion model is constructed based on the teacher-student network architecture, in which the teacher network integrates the prior knowledge of the large language model, and the student network learns the fusion ability of the teacher network through knowledge distillation; 13) Training of lightweight image fusion model: using the training set to train the lightweight image fusion model; 14) Obtaining images to be fused: obtaining multimodal image pairs to be fused, and performing multimodal image pair registration and standardization processing; 15) Obtaining the multi-model image fusion result: Inputting the processed image pair to be fused into the trained lightweight image fusion model to obtain the multi-model image fusion result.

2. According to claim 1, a lightweight multimodal image fusion method based on knowledge distillation technology is characterized in that: The construction of the lightweight image fusion model includes the following steps: 21) A lightweight image fusion model is constructed based on a teacher-student network architecture, where the lightweight image fusion model includes a teacher network and a student network; 22) The teacher network is set to include a text-guided module, an encoder module, and a decoder module. The text-guided module uses a large language model and a CLIP model to generate semantic guidance information for the source image. The encoder module processes input images of different modalities through a two-stream Transformer, including a transposed self-attention module and a spatial self-attention module. The decoder module gradually reconstructs the fused image through text-guided feature modulation; 23) The student network is set to include a feature extraction module, a space-channel cross fusion module and a reconstruction module. The feature extraction module uses a lightweight convolutional network to extract features of multimodal images. The space-channel cross fusion module cross-fuses features in the spatial dimension and the channel dimension. The reconstruction module reconstructs the fused features into a final fused image. 24) Setting the text guide module, encoder module and decoder module; 25) Set the feature extraction module, space-channel cross fusion module and reconstruction module.

3. According to claim 1, a lightweight multimodal image fusion method based on knowledge distillation technology is characterized in that: The training of the lightweight image fusion model includes the following steps: 31) The loss function combination of the lightweight image fusion model is set as follows: Basic loss: in, The structural similarity is measured by the SSIM index. preserves the maximum intensity information between source images, Use the Sobel operator to ensure gradient consistency. Maintain color fidelity in YCbCr space, weight λ t Dynamically adjust based on text guidance; The detailed expressions of each loss component are as follows: Where H and W are the height and width of the image, I f is the fused image, and denote visible light and infrared guidance images, respectively; Among them, SSIM(·) calculates the structural similarity index, δ ir (t) is the IR-guided text dependency weight; in, represents the gradient operator implemented in the horizontal and vertical directions using the Sobel filter; Among them, F CbCr Represents the conversion function from RGB to YCbCr; These losses work together to ensure high-quality image fusion: retains the maximum intensity information from both source images, Preserve structural details through the SSIM metric, Ensure edge consistency through gradient preservation, Keeping the natural color appearance, the text-dependent weights α(t) allow the contribution of each loss component to be dynamically adjusted according to the specific fusion requirements described in the textual cues; Distillation losses: Feature-level adaptation loss: Align the multi-scale features of the teacher network and the student network through the adapter, in and Represents the i-th layer feature map of the student and teacher networks, Adapter i is the corresponding layer adapter, ∥·∥2 represents the L2 norm; Output-level supervision loss: maintain the consistency of the final output, The total loss function is: Where λ1-λ3∈[0.1,1.0] is the weight, the default is 1:1:1; 32) Configure the optimizer: The Adam optimizer is used, with an initial learning rate of 1e-4, a weight decay coefficient of 0.01, and a momentum parameter β = (0.9, 0.999). Use the cosine annealing scheduler to dynamically adjust the learning rate. The formula is: Among them, η max =1e-4 is the maximum learning rate, η min =1e-6 is the minimum learning rate, T max is the maximum number of training cycles; Apply gradient clipping to prevent gradient explosion, and set the threshold to 2.0; 33) Teacher network training includes the following steps: Initialize the teacher network parameters and freeze the large language model parameters. Set the loss function combination: in is the maximum strength loss, is the YCbCr color space fidelity loss, Iterate over each image in the dataset, Data enhancement processing, The input image is randomly cropped to 128x128, Apply color dithering enhancement, adjust brightness by ±0.1, and contrast by ±0.

2. Flip horizontally / vertically with a probability of 0.5, Multimodal feature fusion: The large language model generates a text description T, which is then passed through the CLIP text encoder to obtain F. text , Extract image features F through XRestormer encoder A F B , and F is used uniformly below. i , Up and down sampling through convolution G = Conv (F i ) is used for modulation decoding process, Calculate the 4-layer attention fusion feature using the operator sparse spatial attention to obtain the text-guided feature map F guide =SpatialAttention(F i ,F text ), Use the decoder to get the fused image I fused , used to calculate Parameter update strategy: Gradient clipping is used with a threshold of 2.

0. Use Adam optimizer, β = (0.9, 0.999), Weight decay coefficient 0.01; 34) The student network training includes the following steps: loading the pre-trained teacher network parameters, freezing the teacher network parameters, Knowledge distillation, Feature-level adaptation training, Enable only loss, learning rate 1e-4, Feature adaptation strategy: Layer 4: 384→128 channels, 1×1 convolution for dimensionality reduction, Layer 3: 192 → 64 channels, 3×3 convolutional space-channel conversion, Layer 2: 96→32 channels, 5×5 convolution to enhance local features, Layer 1: 48 → 16 channels, 7×7 convolutions to capture context, Joint training phase: Enable Combined loss: Use the same update strategy as the teacher network; Training process: Input visible light image I vis and infrared image I ir , The teacher network generates fused images through text guidance I teacher , The student network processes the image independently to obtain I student , Compute feature-level and output-level distillation losses, Back-propagation updates the student network parameters.

4. According to claim 2, a lightweight multimodal image fusion method based on knowledge distillation technology is characterized in that: The setting of the text guiding module, the encoder module and the decoder module comprises the following steps: 41) Set up the text guidance module: Generate image description text T through the large language model, and use the CLIP text encoder to convert text features F text ; 42) Set the encoder module: 421) is processed using the transposed self-attention module, Among them, Q T , K, V are the products of input X and Wq, Wk, Wv respectively; 422) The spatial self-attention module is used to process the images, which are divided into 8 windows and then encoded with their respective positions; O = Conv(A·V), Here, PosEmb uses RelPosEmb and sets the number of learnable parameter matrices W and H. The XY coordinates of the internal position of Q are multiplied, and the two directions are respectively cosine-encoded or flattened and then cosine-encoded. Conv represents convolution operation, and the convolution kernel is 3x3. The SoftMax operation is to exponentially smooth all the results at the channel layer. 423) Feature downsampling: The connection between layers is completed through feature downsampling; the downsampling expression is as follows: The separate form used for convolution is implemented here; 424) Repeat the transposed self-attention module processing and the spatial self-attention module processing several times, doubling d each time; 43) Set up the decoder module: 431) perform feature upsampling; F up =Conv(F fused ); 432) for detail enhancement: F enhanced =F up +Conv(SSA(F up ))+Conv(TSA(F up )) Here, SSA and TSA represent the formulas for extracting A and Y above. This module is implemented by repeating the operation twice, halving d each time.

5. According to claim 2, a lightweight multimodal image fusion method based on knowledge distillation technology is characterized in that: The setting of the feature extraction module, the space-channel cross fusion module and the reconstruction module comprises the following steps: 51) Set up feature extraction module: perform feature extraction for different images: F vis =Conv v (I vis ) F ir =Conv i (I ir ); 52) Set up the space-channel cross fusion module: 521) The upper three layers use spatial dimension fusion: F fused =SSA(M s )+TSA(M s ), Here F vis 、F ir From the feature extraction above, each layer has one copy, ⊙ is the element product, σ, The F of Clip is connected by two fully connected layers. text get, In the sub-model, M s =Conv([F vis ,F ir ]), no T is required, and the d of each layer is reduced by 3 times. Because there are 4 layers, the corresponding layer square parameters of different layers are reduced by 2, so the final model is reduced to 0.1 times, and σ, information, that is, no large model and CLIP are required, and the number of parameters is reduced to 0.01 times; 522) The lowest layer uses cross-attention feature fusion: F fused =CrossAttention(F vis ,F ir ) Here CrossAttention exchanges A and Y of the two, and then implements it according to the encoder's formula; 53) Setting the reconstruction module includes the following steps: Perform residual connection enhancement: F enhanced =F up +Conv(SSA(F up ))+Conv(TSA(F up )) Here F up It is the highest level; Full connection reconstruction: I fused =Linear(F enhsnced ) Linear is a fully connected layer.

Citation Information

Cited By

  • Lightweight target detection method based on YOLO

    CN120182585A

  • Infrared and visible light image fusion method and system based on text-guided semantic perception

    CN120765479A

  • A Text-Guided Semantic Awareness-Based Method and System for Infrared and Visible Image Fusion

    CN120765479B