Image fusion method based on detail synergy module and continuous learning of knowledge replay
Patent Information
- Application Number
- CN202410793920.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-19
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-06-19
AI Technical Summary
U2fusion利用EWC正则化的方式保留对之前任务更重要的参数以保留处理之前任务的能力,但这种策略的问题是模型容量是固定的,惩罚损失会使模型的可塑性随着旧任务的增加而降低
[0015]本发明针对现有的图像融合任务在处理连续任务时出现的灾难性遗忘的问题,设置便于知识重放的记忆池,并且基于梯度的筛选机制选取样本,提高记忆池样本多样性;同时利用知识蒸馏的方式使旧知识能够更好得迁移到新模型中,使得图像融合任务能够对抗灾难性遗忘。
Smart Images

Figure CN118823530B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image fusion, specifically relating to an image fusion method based on continuous learning using detail co-operation modules and knowledge replay. Background Technology
[0002] Currently, due to the combined effects of weather, environment, and camera equipment, it is difficult to obtain high-quality images during image acquisition and transmission. Therefore, images often contain significant noise, failing to accurately depict objects within the image and hindering subsequent applications. To address these issues, various image processing techniques have been introduced to improve image quality, such as noise reduction, illumination filtering, image enhancement, and image fusion. Among these, image fusion processes multiple images according to specific rules, understanding the complementary information between them to maximize image quality. Compared to other image processing techniques, image fusion possesses undeniable advantages and has remained a research hotspot in the field, attracting considerable attention from scholars and leading to extensive research on related algorithms.
[0003] In real life, due to limitations in imaging principles or optical devices, images captured using a single type of sensor or a single shooting setup can only capture a portion of the information. For digital photographic images, images taken by digital cameras can only clearly present scenes within a predefined depth of field. The discrete blurring effect outside the predefined depth of field produces unclear semantics, and the limitation of low dynamic range makes the images visually inferior compared to natural scenes. Infrared sensors capture thermal radiation information emitted by objects, effectively highlighting salient targets such as pedestrians and vehicles, but lack detailed descriptions of the scene. Visible light images typically contain rich textural details, but are susceptible to extreme environments and occlusion, resulting in the loss of targets within the scene. Images produced by these single sensors all suffer from incomplete content, severely impacting downstream computer vision tasks.
[0004] Simultaneously, with the continuous development of information technology, various types of data are experiencing explosive growth. Traditional machine learning algorithms can only achieve good performance when the distribution of test data is similar to that of training data. They cannot continuously and adaptively learn in dynamic environments; however, this adaptive learning capability is a characteristic possessed by any intelligent system. Deep neural networks have shown the best learning ability in many applications; however, when using this method to incrementally update data, they face catastrophic interference or forgetting problems, causing the model to forget how to solve old tasks after learning new ones. To address this issue, many methods have been developed to solve the problem of catastrophic forgetting, including weight regularization and function regularization. Many image fusion algorithms have also been introduced, but the following problems still exist.
[0005] Current general image fusion methods, in order to be applicable to different fusion tasks, have downplayed the focus on the characteristics of the task itself and sought more generalizable feature modeling methods. As a result, general fusion frameworks often produce worse performance than frameworks designed for individual tasks.
[0006] Catastrophic forgetting: Most existing general image fusion frameworks are designed for fusion tasks but do not consider the continuous processing of a series of image fusion tasks. When faced with images from different fusion tasks being continuously input, the network exhibits catastrophic forgetting. U2fusion uses EWC regularization to retain parameters that are more important to previous tasks in order to preserve the ability to process previous tasks. However, the problem with this strategy is that the model capacity is fixed, and the penalty loss will reduce the model's plasticity as the number of old tasks increases. Summary of the Invention
[0007] To address the aforementioned technical problems, this invention provides an image fusion method based on continuous learning using a detail collaboration module and knowledge replay, comprising:
[0008] Step 1: Prepare the image dataset and preprocess the image dataset;
[0009] The image dataset includes a dataset of infrared-visible images, a dataset of multi-exposure images, and a dataset of multi-focus images;
[0010] Step 2: Construct a unified image fusion network for processing the image dataset fusion task;
[0011] The unified image fusion network consists of a shared feature encoder, a Transformer-based basic feature encoder, a CNN-based detail feature encoder, a basic feature fusion layer, a detail feature fusion layer, and a Restormer-based feature decoder.
[0012] Step 3: Construct a knowledge distillation framework based on cross-task detailed collaboration modules and knowledge replay, and process the data output by the unified image fusion network according to the constructed knowledge distillation framework;
[0013] The knowledge distillation framework consists of knowledge distillation for continuous learning of image fusion, a cross-task detail collaboration module, and a gradient-based sample selection mechanism.
[0014] The beneficial effects of this invention are:
[0015] This invention addresses the problem of catastrophic forgetting in existing image fusion tasks when handling consecutive tasks. It sets up a memory pool that facilitates knowledge replay and selects samples based on a gradient-based filtering mechanism to improve the diversity of the memory pool samples. At the same time, it uses knowledge distillation to enable old knowledge to be better transferred to the new model, thus enabling the image fusion task to resist catastrophic forgetting. Attached Figure Description
[0016] Figure 1 An overall flowchart provided for embodiments of the present invention;
[0017] Figure 2 A schematic diagram of the unified image fusion network structure provided in an embodiment of the present invention;
[0018] Figure 3 A knowledge distillation framework diagram provided for embodiments of the present invention;
[0019] Figure 4 The overall schematic diagram provided for embodiments of the present invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] like Figures 1-4 As shown, this embodiment of the invention provides an image fusion method based on continuous learning using a detail co-operation module and knowledge replay, comprising:
[0022] Step 1: Prepare the image dataset and preprocess the image dataset;
[0023] The image dataset includes a dataset of infrared-visible images, a dataset of multi-exposure images, and a dataset of multi-focus images;
[0024] Step 2: Construct a unified image fusion network for processing the image dataset fusion task;
[0025] The unified image fusion network consists of a shared feature encoder, a Transformer-based basic feature encoder, a CNN-based detail feature encoder, a basic feature fusion layer, a detail feature fusion layer, and a Restormer-based feature decoder.
[0026] Step 3: Construct a knowledge distillation framework based on cross-task detailed collaboration modules and knowledge replay, and process the data output by the unified image fusion network according to the constructed knowledge distillation framework;
[0027] The knowledge distillation framework consists of knowledge distillation for continuous learning of image fusion, a cross-task detail collaboration module, and a gradient-based sample selection mechanism.
[0028] like Figure 2 As shown: Preprocessing of the image dataset includes:
[0029] The visible light image portion, multi-exposure image dataset, and multi-focus image dataset in the infrared-visible light image dataset are converted from RGB color space to YCbCr color space, and the Y channel containing image brightness and contrast information is taken as the dataset for training. Then, the extracted Y channel images are cropped to a size of 128×128.
[0030] Specifically, the data is expanded by cutting the data to a size of 128×128 and then randomly flipping and rotating it.
[0031] like Figure 2 As shown: The shared feature encoder takes the two images to be fused as input, first performs convolution operation with 3×3 convolution kernels to expand the number of channels of the two images respectively, and then obtains preliminary image features through 4 Transformer blocks.
[0032] Specifically, the Transformer block calculates channel attention, which reduces computation while maximizing the use of global information from multiple channels to obtain preliminary image features.
[0033] like Figure 2 As shown: The Transformer-based basic feature encoder takes the output of the shared feature encoder as input. First, the preliminary image features are passed through a biased layer normalization module to normalize each feature. Then, multi-channel attention is calculated on the normalized features. The input features are linearly transformed through multiple different channel attention heads to obtain different queries, keys, and values. Attention weights are calculated for each attention head, and then normalized through the softmax function to obtain the attention weights on each feature dimension. The values are weighted and summed using the attention weights to obtain the output of the multi-channel attention. This allows the basic features to better focus on the global dependencies between features.
[0034] Specifically: Normalization is used to address the shift of internal covariates, thus avoiding instability and slow convergence during training.
[0035] like Figure 2 As shown: The CNN-based detail encoder: This module takes the output of the shared feature encoder as input. Its main structure is a reversible neural network. First, the initial image features are divided equally along the channel dimension to obtain two new features Z1 and Z2. The two new features are passed through three reversible neural network nodes to obtain detail features. Then, the scaling factor ρ, the bias factor φ, and the bias factor η are generated using reversible residual blocks. The bias factor φ is added to Z2, and Z1 is multiplied by the scaling factor ρ and then added to the bias factor η. Finally, the updated Z1 and Z2 are input to the next reversible neural network node. The forward and backward calculations of the reversible neural network are inversely related. The initial image features are accurately recovered through the backward calculation.
[0036] Specifically: Since the forward and backward computations of the network are inverses, even features mapped by the network can be accurately restored to the original features through backward computation, thus ensuring the integrity of detailed features.
[0037] like Figure 2 As shown: The basic feature fusion layer and the detail feature fusion layer add the two sets of basic features and detail features extracted from the two images to be fused, respectively, to achieve feature stitching.
[0038] like Figure 2 As shown: The Restormer-based feature decoder inputs the fused basic features and detail features into the feature decoder. The feature decoder first uses a tensor concatenation function to concatenate the detail features and basic features along the channel dimension to obtain preliminary fused features. Then, the preliminary fused features are aggregated through a 1×1 convolution kernel to obtain a feature map with more concentrated information.
[0039] Specifically: The aggregated features are input into a Restormer block of length 3. First, normalization is applied to the aggregated features, which may have distributional biases. The normalized result is then used as input to the attention mechanism. Within the attention mechanism, query, key, and value features are obtained through three convolutional operations. An attention score is calculated by assessing the correlation between the query and key. This attention score is then multiplied by the value, and the results are summed to obtain the self-attention output of the aggregated features. The final output is the fused result.
[0040] like Figure 1 As shown: A knowledge distillation framework based on cross-task detailed collaboration modules and knowledge replay is constructed. The data output by the unified image fusion network is processed according to the constructed knowledge distillation framework, including:
[0041] Step 3.1: Train the unified image fusion network using knowledge distillation through continuous learning of image fusion. Input image fusion datasets from different tasks sequentially into the unified image fusion network, and let the two input images to be fused be... Where t represents the task number, and F() represents the image fusion network. After the training of the dataset for a fusion task is completed, the model weights at this time are saved, and the parameters are frozen, setting the network parameters to not calculate gradients. Since the frozen model parameters were trained on the dataset of the previous fusion task and have the ability to process the previous task, they are used as the teacher model to guide the training of the next task. The fusion result of the teacher model is used to constrain the fusion result of the current training model. The knowledge of the existing teacher model is iteratively transferred to the current training model. L1 loss is selected as the distillation loss, and the distillation loss of the fusion result is expressed as:
[0042]
[0043] Specifically, F t-1 (.) represents the teacher model obtained after training on the (t-1)th task, F t (.) represents the model for the current task, where X1t and X2t are input to F. t After (.) modeling, first calculate the fusion loss of the data for task t in the current model. Then, check if the data in the current memory pool (including data from task 1 to task t-1) is empty. If not empty, call the data X1t-1 and X1t-2 from the current memory pool and input them into the teacher model and the current model respectively to obtain the corresponding fusion results:
[0044]
[0045] Step 3.2: The cross-task detail collaboration module directly transfers knowledge from the teacher model using L1 loss. The differences between these three tasks—multi-focus image fusion focusing on the clear regions of two images, infrared-visible image fusion focusing on salient targets and detailed textures, and multi-exposure image fusion focusing on global brightness information—are addressed by inputting a set of detail features f1 obtained from a CNN-based detail feature encoder into the two images to be fused. detail f2 detailAs input, the two detail features are first layer-normalized to remove distribution bias. Then, multi-channel attention is calculated on the normalized features. The input features are linearly transformed through multiple different channel attention heads to obtain different queries Q1 and Q2, keys K1 and K2, and values V1 and V2. The queries of the two detail features are swapped, i.e., Q1 and Q2 are swapped. The dot product of the swapped queries and keys is calculated to obtain the attention score map. Then, it is normalized by the softmax function to obtain the attention weight on each feature dimension. The values are weighted and summed using the attention weights to obtain the cross features of the two images to be fused in the detail features in the current task, i.e., the mutual attention.
[0046] Specifically: While the cross-task detail collaboration module uses L1 loss to directly transfer knowledge from the old model, it is efficient. However, the focus on images differs between different fusion tasks, especially on detailed features. Simply constraining the fusion results ignores the differences between tasks, making it impossible to effectively utilize the replayed knowledge.
[0047] The definition of obtaining cross features is... To preserve the attention paid to image detail regions by the old task, knowledge distillation is performed on the data in the memory pool using cross-features obtained from the teacher model and the current training model as constraints. L1 loss is also used, and the final distillation loss for the cross-detail features is:
[0048]
[0049] Among them, L detailKD The distillation loss represents the cross-detail feature. and These represent the first and second detail features in a set of detail features. Let represent the cross-feature function, and t represent the task number.
[0050] Step 3.3: A gradient-based sample selection mechanism is used to select samples from a fixed-size memory pool. This mechanism compares the gradient generated during network training for the current task with the gradients generated by data in the existing memory pool when passing through the network. A larger gradient angle indicates a greater difference between the current data and the data in the memory pool, making the current data more suitable as candidate data for storage. Specifically, when training for the first task ends, 600 pairs of samples (equal to the size of the memory pool) are randomly sampled. At the start of subsequent tasks, for each batch of data, the same number of data are randomly sampled from the memory pool during training. Simultaneously, the gradients of the current task's data and the data in the memory pool with respect to the current task are calculated, and an array of gradient cosine similarity values is maintained. After the current task ends, data with larger gradient differences are used to fill the memory pool. The selection formula is as follows:
[0051]
[0052] Where, r i g represents the similarity score between the current sample gradient and the gradients of samples in the memory pool. i G and G represent the gradients of the current sample and the sample set stored in the memory pool, respectively.
[0053] Specifically, a gradient-based sample selection mechanism is used to select effective samples from a fixed-size memory pool. Effective samples are those that have a greater impact on the update of network parameters during network training in previous tasks, and that are more representative of the features of the current task dataset.
[0054] Principle: Prepare datasets of infrared-visible light images, multi-exposure images, and multi-focus images, and perform data preprocessing on these datasets for fusion tasks.
[0055] Training a single task: Construct a unified image fusion network for handling image fusion tasks. The input images to be fused are processed by a shared feature encoder, a Transformer-based basic feature encoder, a CNN-based detail feature encoder, a basic feature fusion layer, a detail feature fusion layer, and a Restormer-based feature decoder to finally obtain the fused result, and random samples are selected and put into the memory pool.
[0056] Training on multiple tasks, a knowledge distillation framework based on cross-task detailed collaborative modules and knowledge replay is constructed. A training mode combining knowledge distillation and knowledge replay is adopted. While the current task is input into the network for training, data from previous tasks are processed using both the old and current models. The distillation loss of cross-features and fusion results is calculated, preserving the ability to handle detailed features for different tasks. The training data from previous tasks is stored in a fixed-size memory pool through a gradient-based filtering mechanism. Finally, the current task model is frozen as the teacher model for the next task, thus combating catastrophic forgetting.
[0057] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An image fusion method based on continuous learning using detail collaboration modules and knowledge replay, characterized in that, include: Step 1: Prepare the image dataset and preprocess the image dataset; The image dataset includes a dataset of infrared-visible images, a dataset of multi-exposure images, and a dataset of multi-focus images; Step 2: Construct a unified image fusion network for processing the image dataset fusion task; The unified image fusion network consists of a shared feature encoder, a Transformer-based basic feature encoder, a CNN-based detail feature encoder, a basic feature fusion layer, a detail feature fusion layer, and a Restormer-based feature decoder. The shared feature encoder takes the two images to be fused as input, first performs convolution operation with 3×3 convolution kernels to expand the number of channels of the two images respectively, and then obtains preliminary image features through 4 Transformer blocks; The Transformer-based basic feature encoder takes the output of the shared feature encoder as input. First, the preliminary image features are passed through a biased layer normalization module to normalize each feature. Then, multi-channel attention is calculated on the normalized features. The input features are linearly transformed through multiple different channel attention heads to obtain different queries, keys, and values. Attention weights are calculated for each attention head, and then normalized through the softmax function to obtain the attention weights on each feature dimension. The values are weighted and summed using the attention weights to obtain the output of the multi-channel attention, which is the basic feature. The CNN-based detail encoder takes the output of the shared feature encoder as input. The structure of the CNN-based detail encoder is a reversible neural network. First, the initial input image features are divided equally along the channel dimension to obtain two new features Z1 and Z2. The two new features are passed through three reversible neural network nodes to obtain detail features. Then, the scaling factor ρ, the bias factor φ, and the bias factor η are generated using reversible residual blocks. The bias factor φ is added to Z2, and Z1 is multiplied by the scaling factor ρ and then added to the bias factor η. Finally, the updated Z1 and Z2 are input to the next reversible neural network node. The forward and backward calculations of the reversible neural network are inversely related. The initial image features are accurately recovered through the backward calculation. The basic feature fusion layer and the detail feature fusion layer add the two sets of basic features and detail features extracted from the two images to be fused, respectively, to achieve feature stitching; The Restormer-based feature decoder inputs the fused basic and detail features into the feature decoder. The feature decoder first uses a tensor concatenation function to concatenate the detail features and basic features along the channel dimension to obtain preliminary fused features. Then, the preliminary fused features are aggregated by passing a 1×1 convolution kernel to obtain a feature map with more concentrated information. Step 3: Construct a knowledge distillation framework based on cross-task detailed collaboration modules and knowledge replay, and process the data output by the unified image fusion network according to the constructed knowledge distillation framework; The knowledge distillation framework consists of knowledge distillation for continuous learning of image fusion, a cross-task detail collaboration module, and a gradient-based sample selection mechanism.
2. The image fusion method based on continuous learning using a detail collaboration module and knowledge replay as described in claim 1, characterized in that, Preprocessing of the image dataset includes: The visible light image portion, multi-exposure image dataset, and multi-focus image dataset in the infrared-visible light image dataset are converted from RGB color space to YCbCr color space, and the Y channel containing image brightness and contrast information is taken as the dataset for training. Then, the extracted Y channel images are cropped to a size of 128×128.
3. The image fusion method based on continuous learning using a detail collaboration module and knowledge replay as described in claim 1, characterized in that, A knowledge distillation framework based on cross-task detailed collaboration modules and knowledge replay is constructed. Data output from a unified image fusion network is processed according to this knowledge distillation framework, including: Step 3.1: Train the unified image fusion network using knowledge distillation through continuous learning of image fusion. Input image fusion datasets from different tasks sequentially into the unified image fusion network, and let the two input images to be fused be... , , , where t represents the task number, This represents an image fusion network. After training on the dataset for a fusion task is completed, the model weights at this point are saved, and the parameters are frozen. The network parameters are set to not calculate gradients. Since the frozen model parameters were trained on the dataset of the previous fusion task, they have the ability to process the previous task. Therefore, they are used as the teacher model to guide the training of the next task. The fusion result of the teacher model is used to constrain the fusion result of the current training model. The knowledge of the existing teacher model is iteratively transferred to the current training model. L1 loss is selected as the distillation loss. Step 3.2: The cross-task detail collaboration module directly transfers knowledge from the teacher model using L1 loss. The differences between these three tasks—multi-focus image fusion focusing on the clear regions of two images, infrared-visible image fusion focusing on salient targets and detailed textures, and multi-exposure image fusion focusing on global brightness information—are addressed by inputting a set of detail features f1 obtained from a CNN-based detail feature encoder into the two images to be fused. detail f2 detail As input, the two detail features are first layer-normalized to remove distribution bias. Then, multi-channel attention is calculated on the normalized features. The input features are linearly transformed through multiple different channel attention heads to obtain different queries Q1 and Q2, keys K1 and K2, and values V1 and V2. The queries of the two detail features are swapped, i.e., Q1 and Q2 are swapped. The dot product of the swapped queries and keys is calculated to obtain the attention score map. Then, the softmax function is used for normalization to obtain the attention weights on each feature dimension. The attention weights are used to perform a weighted summation of the values to obtain the cross features of the two images to be fused on the detail features in the current task, i.e., the mutual attention. The definition of obtaining cross features is... To preserve the attention paid to image detail regions by the old task, the data in the memory pool is constrained by the cross features obtained from the teacher model and the current training model to perform knowledge distillation. The L1 loss is also used to obtain the distillation loss of the cross detail features. Step 3.3: Gradient-based sample selection mechanism selects samples from a fixed-size memory pool. By using a gradient-based sample selection mechanism, the gradient generated by the current task during network training is compared with the gradient generated by the data in the existing memory pool when passing through the current network. A larger gradient angle means that the current data has a greater difference from the data in the memory pool, and is more suitable as candidate data to be stored in the memory pool. At the beginning of subsequent tasks, for each batch of data, the same number of data are randomly sampled from the memory pool during training. At the same time, the gradient magnitudes of the current task data and the task data in the memory pool with respect to the current task are calculated, and a set of gradient cosine similarity arrays are maintained. After the current task ends, the data with greater gradient differences are used as data to fill the memory pool.
4. The image fusion method based on continuous learning using a detail collaboration module and knowledge replay as described in claim 3, characterized in that, The distillation loss of the cross-detail features is: ; in, The distillation loss represents the cross-detail feature. and These represent the first and second detail features in a set of detail features. Let represent the cross-feature function, and t represent the task number.
5. The image fusion method based on continuous learning using a detail collaboration module and knowledge replay as described in claim 3, characterized in that, The formula for the gradient-based sample selection mechanism is as follows: ; in, This represents the similarity score between the current sample gradient and the gradients of samples in the memory pool. and These represent the gradients of the current sample and the sample set stored in the memory pool, respectively.
Citation Information
Patent Citations
Infrared and visible light fusion method based on knowledge distillation
CN116739957A
Infrared and visible light image fusion system based on distillation-fusion-semantic joint driving
CN117274759A