An adaptive multi-modal image fusion method based on knowledge embedding
Patent Information
- Application Number
- CN202410745910.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-11
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-06-11
AI Technical Summary
[0005]针对现有技术的不足,本发明提供了一种基于知识嵌入的自适应多模态图像融合方法,具备生成清晰可靠的融合图像等优点,解决了上述技术问题
[0059]本发明通过采用Mean Teacher的自监督机制,有效提升了图像融合网络在恶劣天气下的鲁棒性,还引入了多层次协同自适应重建网络,通过多分枝、多尺度的设计,实现了对不同特征的差异化处理,从而既保留了图像丰富的纹理信息,又维持了图像的语义一致性,实验结果表明,本发明的方法在视觉质量和定量评价方面优于当前技术水平,为极端天气条件下的图像融合任务提供了更多有效信息,并且有助于促进下游视觉任务的发展。
Smart Images

Figure CN118762257B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, specifically to an adaptive multimodal image fusion method based on knowledge embedding. Background Technology
[0002] With the rapid development of big data and artificial intelligence technologies, single-modal images are no longer sufficient to meet the needs of more comprehensive image understanding. Image information from different modalities, such as visible light images, infrared images, and radar images, often contains complementary information. Effectively integrating these modalities can significantly address the degradation caused by environmental conditions and noise in single-modal images, enabling more comprehensive and accurate visual understanding. Therefore, multimodal image fusion technology has emerged as an important direction in the fields of computer vision and image processing.
[0003] When applying vision-based tasks such as autonomous driving and object tracking to real-world scenarios, they face challenges from complex and chaotic environments and severe weather conditions. Conventional image fusion methods in this area represent a critical research gap. Although existing multimodal image fusion methods perform well under normal imaging conditions, they are inevitably affected by weather factors in adverse weather conditions, leading to asymmetric distortion in the perceptual flow during feature fusion and thus impacting the fusion results. This is because existing training datasets are biased towards sunny weather conditions, and the network structure of the models is not specifically designed to handle adverse scenarios with asymmetric information flow. Summary of the Invention
[0004] (a) Technical problems to be solved
[0005] To address the shortcomings of existing technologies, this invention provides an adaptive multimodal image fusion method based on knowledge embedding, which has advantages such as generating clear and reliable fused images, and solves the aforementioned technical problems.
[0006] (II) Technical Solution
[0007] To achieve the above objectives, the present invention provides the following technical solution: an adaptive multimodal image fusion method based on knowledge embedding, comprising the following steps:
[0008] S1. Set up the training dataset, which includes the following steps:
[0009] S1.1, from M 3 Infrared images randomly selected from the FD and MSRS datasets. i ∈R H×W×1 and visible light image V i ∈R H×W×3 ;
[0010] Wherein, the subscript i represents the image number, the superscript H×W×1 represents the length, width and number of channels of the infrared image as H, W and 1 respectively, and the superscript H×W×3 represents the length, width and number of channels of the visible light image as H, W and 3 respectively;
[0011] S1.2. Based on the atmospheric scattering model, all selected visible light images are subjected to fogging processing to obtain fogged visible light images.
[0012] S1.3, transfer the infrared image I i ∈R H×W×1 Visible light image V i ∈R H×W×3 and fogged visible light images To form the training dataset;
[0013] S2. Construct a multi-level collaborative adaptive reconstruction network, including a multi-scale feature extractor, an information collaboration layer, and an image reconstruction decoder. The multi-scale feature extractor is used to extract shallow detail features and deep semantic features of the image. The information collaboration layer adopts a differentiated strategy to integrate information for inputs at different levels, allowing complementary information to be fused and redundant information to be filtered. The image reconstruction decoder is used to decode the output features of the information collaboration layer and reconstruct the fused features.
[0014] The multi-scale feature extractor includes a three-layer Restormer Block and a high-frequency perception enhancement module. The three-layer Restormer Block is stacked to extract features. The high-frequency perception enhancement module includes two branches, one of which contains a two-layer convolutional residual block and the other integrates the Sobel operator.
[0015] The process of shallow detail feature extraction by the information collaboration layer is as follows:
[0016] S2a.1. Element-wise difference of the infrared and visible light features extracted by the multi-scale feature extractor;
[0017] S2a.2, Using the difference portion D as the excitation signal to generate channel weight W D And perform two compensations on infrared and visible light features;
[0018] S2a.3, Obtain the compensated detailed features F t ;
[0019] The information collaboration layer performs the following steps for deep semantic feature fusion:
[0020] S2b.1. The element-wise sum of the input infrared features and the infrared and visible light features is used as the guiding feature and the feature to be screened, respectively, and then subjected to a 1x1 convolution.
[0021] S2b.2. Use 3x3 depthwise convolution to encode the spatial context features of the guiding features and the features to be selected, to obtain the query vector Q, key vector K and numerical vector V for attention calculation, respectively.
[0022] S2b.3. Perform a dot product operation between the query vector Q generated by the guided features and the transpose of the key vector K generated by the features to be filtered to generate a C*C attention map.
[0023] S2b.4. Multiply the attention map with the original features to obtain the compensated fused features;
[0024] S2b.5. Perform saliency calculation on the compensated fusion features to obtain the enhanced features F. s ;
[0025] S3. Construct a self-supervised image restoration and multimodal image fusion joint learning network, consisting of a multi-layered collaborative adaptive reconstruction network, forming a set of twin guidance and learning networks. The guidance and learning networks are trained on different data respectively. and Data s To learn from network datasets, Data t To guide the network dataset, N represents the image pair I. i , and I i V i The number of images, where the subscript i represents the i-th image;
[0026] S3.1 Training the multi-scale feature extractor and image reconstruction decoder: Using the feature encoder E of the learning network S and the guiding network T. S and E T Feature extraction is performed, and an image reconstruction decoder D is constructed using a learning network S and a guiding network T. S and D T Image reconstruction is performed using the following expression:
[0027]
[0028] in, and These represent the infrared and visible light images reconstructed by the learning network and the image reconstruction decoder, respectively. and These represent the infrared and visible light images reconstructed by the image reconstruction decoder after being guided by the network. Data s To learn from network datasets, Data t To guide the network dataset, N represents the image pair I.i , and I i V i The quantity, D S and D T Let S and T represent the image reconstruction decoders of the learning network S and the guiding network T, respectively.
[0029] S3.2 Training of the Information Collaboration Layer: Paired infrared and degraded visible light images After passing through the multi-scale feature extractor E s The information collaboration layer and the image reconstruction decoder obtain the fused image F; similarly, the guidance network T receives... Obtaining pseudo tags The specific expression is as follows:
[0030]
[0031]
[0032] in, and Representing the learning network and the bootstrap network respectively, Data s To learn from network datasets, Data t To guide the network dataset, F represents the fused image obtained by the information collaboration layer and the image reconstruction decoder. Indicates a pseudo-tag;
[0033] S3.3 Setting the parameters θ of the bootstrap network during training. t By learning network parameters θ s The exponential moving average is used for updating, and the specific expression is as follows:
[0034] θ t =ηθ t +(1-η)θ s
[0035] Where, θ t θ represents the parameters of the guiding network. s The parameter η represents the learning network parameter, η∈(0,1), where η represents the momentum term. Higher η gives more weight to new information, while lower η focuses more on historical parameters.
[0036] S4. Set the loss function during image reconstruction. The specific steps are as follows:
[0037] S4.1, Set in pairs and The images are fed into the learning network S and the guiding network T, respectively, and then processed by the feature encoder and the image reconstruction decoder to reconstruct clear infrared and visible light images. and The loss function of the process is expressed as follows:
[0038]
[0039]
[0040] in, L represents the consistency loss in the first stage, α and β represent the adjustment parameters, and L represents the consistency loss in the first stage. re_ir and L re_vi Let represent the reconstruction losses for infrared and visible light images, respectively, and let I and V represent the infrared and visible light images used in the training process, respectively. and These represent the reconstructed infrared and visible light images, respectively. L1(*,*) represents the L2 norm, γ represents the weights of the loss function, and L1(*,*) represents the L1 norm. SSIM (*,*)=1-SSIM(*,*), and L SSIM (*,*) represents the structural similarity loss. CC(*,*) represents the consistency loss in the first stage, and CC(*,*) represents the correlation coefficient, Φ s ,Φ t The sub-subjects represent the features output by the multi-scale feature extractors of the learning network and the guided network, respectively, and SSIM(*,*) represents the structural similarity index. This indicates that a clear infrared image has been reconstructed. This indicates that a clear visible light image has been reconstructed.
[0041] S4.2. Using paired infrared and visible light images as input, after passing through the multi-scale feature extractor trained in S3.1, their deep semantic features and shallow detail features are fused through an information synergy layer. Finally, the loss function in the image reconstruction decoder reconstructs the fused image is expressed as follows:
[0042]
[0043]
[0044] in, This represents the total loss function for the second stage, which is composed of the fusion loss L. fusion Consistency loss in the second phase It consists of two parts, with fusion loss L fusion Including intensity loss and gradient loss, L int L represents the pixel intensity loss between the fused image and the original image. grad This represents the gradient loss between the fused image and the original image, where ||*||1 represents the L1 norm. This represents the Sobel gradient operator with respect to *, max(*,*) indicates taking the maximum value, and H and W represent the height and width of the image, respectively. F represents the second-phase consistency loss. These represent the fused image and pseudo-label obtained in step S3.2, respectively. This represents the calculation of the L1 loss between the fused image and the pseudo-label.
[0045] As a preferred embodiment of the present invention, the calculation expression for the feature extraction process of the multi-scale feature extractor is as follows:
[0046]
[0047] in, and These represent the shallow detail features obtained from feature extraction of infrared and visible light images, respectively. They are obtained by concatenating the features extracted from the first two Restormer Block layers along the channel dimension. and These represent the deep semantic features extracted from infrared and visible light images, respectively. Concate(*) represents the feature map concatenation operation. RTB 1 RTB 2 RTB 3 These represent Restormer Blocks for layers 1-3, e I and e V These represent the infrared and visible light edge features extracted by the high-frequency perception enhancement module HPEM, respectively. HPEM(*,*) indicates that edge extraction is performed on the internal image data.
[0048] As a preferred embodiment of the present invention, the calculation expression for the shallow detail feature extraction process of the information collaboration layer is as follows:
[0049] D = F v -F i
[0050]
[0051] Where σ represents the sigmoid activation function, GAP represents the global average pooling operation, ReLU represents the ReLU activation function, D represents the difference component, and F i and F v W represents infrared and visible light characteristics, respectively. D Indicates channel weight, and These represent the infrared and visible light image features after initial compensation, respectively. g and w lThese represent the weights calculated using global attention and local attention, respectively. Conv represents the convolution operation, and ⊙ represents the dot product operation. For matrix multiplication, w i The weights of the infrared features, w v Weights of visible light features.
[0052] As a preferred embodiment of the present invention, the expression for deep semantic feature fusion in the information collaboration layer is as follows:
[0053]
[0054] Among them, F Guide and F filtered These represent the guiding features and the features to be selected, respectively. Reshape(*) represents the reshape function, Conv represents the convolution operation, Dcos(*) represents the depthwise convolution operation, Attention(*) represents the attention calculation, and Q... guide The query vector representing the semantic attention computation is generated from the guiding features, K. filtered and V filtered F represents the key vector and numerical vector used for attention calculation, generated from the features to be selected. f For fusion features, F s This indicates a fusion feature that has undergone saliency enhancement; Salient-computation(*) is the saliency enhancement operation.
[0055] As a preferred embodiment of the present invention, the expression for the significance calculation is as follows:
[0056]
[0057] Where Contrast represents saliency calculation, K represents the number of pixels, and x j This represents the value of the j-th pixel in the fused features. This represents the mean of the feature map.
[0058] Compared with existing technologies, this invention provides an adaptive multimodal image fusion method based on knowledge embedding, which has the following beneficial effects:
[0059] This invention effectively improves the robustness of image fusion networks under severe weather conditions by employing a Mean Teacher self-supervised mechanism. It also introduces a multi-level collaborative adaptive reconstruction network, which achieves differentiated processing of different features through multi-branch and multi-scale design. This preserves the rich texture information of the image while maintaining its semantic consistency. Experimental results show that the method of this invention outperforms the current technology in terms of visual quality and quantitative evaluation, provides more effective information for image fusion tasks under extreme weather conditions, and helps promote the development of downstream visual tasks. Attached Figure Description
[0060] Figure 1 This is the overall flowchart of the present invention;
[0061] Figure 2 This is a structural diagram of the multi-level feature extraction, information collaboration layer, and image reconstruction decoder of the present invention;
[0062] Figure 3 This is a structural diagram of the difference-aware detail feature fusion module (DADFM) in the information collaboration layer of this invention;
[0063] Figure 4 This is a structural diagram of the cross-modal semantic consistency maintenance module CSCPM in the information collaboration layer of this invention;
[0064] Figure 5 This invention is in M 3 Visualization of experiments on the FD dataset. Detailed Implementation
[0065] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0066] Please see Figure 1-5 An adaptive multimodal image fusion method based on knowledge embedding includes the following steps:
[0067] S1. Set up the training dataset, which includes the following steps:
[0068] S1.1, from M 3 Infrared images randomly selected from the FD and MSRS datasets. i ∈R H×W×1 and visible light image V i ∈R H×W×3 ;
[0069] Wherein, the subscript i represents the image number, the superscript H×W×1 represents the length, width and number of channels of the infrared image as H, W and 1 respectively, and the superscript H×W×3 represents the length, width and number of channels of the visible light image as H, W and 3 respectively;
[0070] S1.2. Based on the atmospheric scattering model, all selected visible light images are subjected to fogging processing to obtain fogged visible light images.
[0071] S1.3, transfer the infrared image I i ∈R H×W×1 Visible light image V i ∈R H×W×3 and fogged visible light images To form the training dataset;
[0072] S2. Construct a multi-level collaborative adaptive reconstruction network, including a multi-scale feature extractor, an information collaboration layer, and an image reconstruction decoder. The multi-scale feature extractor is used to extract shallow detail features and deep semantic features of the image. The information collaboration layer adopts a differentiated strategy to integrate information from different levels of input, allowing complementary information to be fused and redundant information to be filtered. The image reconstruction decoder is used to decode the output features of the information collaboration layer and reconstruct the fused features.
[0073] The multi-scale feature extractor includes a three-layer Restormer Block and a high-frequency perception enhancement module (HPEM). The three-layer Restormer Block is stacked to extract features. The high-frequency perception enhancement module includes two branches. One branch contains a two-layer convolutional residual block, and the other branch integrates the Sobel operator to effectively embed local and global information into the network, ensuring that the network can capture rich details and salient regions of the scene.
[0074] The computational expression for the feature extraction process of the multi-scale feature extractor is as follows:
[0075]
[0076] in, and These represent the shallow detail features obtained from feature extraction of infrared and visible light images, respectively. They are obtained by concatenating the features extracted from the first two Restormer Block layers along the channel dimension. and These represent the deep semantic features extracted from infrared and visible light images, respectively. Concate(*) represents the feature map concatenation operation. RTB 1 RTB 2 RTB 3These represent Restormer Blocks for layers 1-3, e I and e V These represent the infrared and visible light edge features extracted by the high-frequency perception enhancement module HPEM, respectively. HPEM(*,*) indicates that edge extraction operation is performed on the internal image data.
[0077] The process of shallow detail feature extraction by the difference-aware detail feature fusion module DADFM in the information collaboration layer is as follows:
[0078] S2a.1. Element-wise difference of the infrared and visible light features output by the multi-scale feature extractor;
[0079] S2a.2, Using the difference portion D as the excitation signal to generate channel weight W D And perform two compensations on infrared and visible light features;
[0080] S2a.3, Obtain the compensated detailed features F t ;
[0081] The calculation expression for the shallow detail feature extraction process of the information collaboration layer through the difference-aware detail feature fusion module DADFM is as follows:
[0082]
[0083] D = F v -F i
[0084] Where σ represents the sigmoid activation function, GAP represents the global average pooling operation, ReLU represents the ReLU activation function, D represents the difference component, and F i and F v W represents infrared and visible light characteristics, respectively. D Indicates channel weight, and These represent the infrared and visible light image features after initial compensation, respectively. g and w l These represent the weights calculated using global attention and local attention, respectively. Conv represents the convolution operation, and ⊙ represents the dot product operation. For matrix multiplication, w i The weights of the infrared features, w v Weights of visible light features.
[0085] The information collaboration layer performs the following steps for deep semantic feature fusion through the cross-modal semantic consistency maintenance module (CSCPM):
[0086] S2b.1. The element-wise sum of the input infrared features and the infrared and visible light features is used as the guiding feature and the feature to be screened, respectively, and then subjected to a 1x1 convolution.
[0087] S2b.2. A 3x3 depthwise convolution is used to encode the spatial context features of the guiding features and the features to be selected. The generated matrix is transformed into HW*C, which yields the query vector Q, key vector K and numerical vector V for attention calculation, respectively.
[0088] S2b.3. By performing a dot product operation between the query vector Q generated by the guided features and the transpose of the key vector K generated by the features to be filtered, a C*C attention map is generated. This attention map reflects the similarity between the two features to maintain the consistency of deep semantics.
[0089] S2b.4. Multiply the attention map with the original features to be screened to obtain the compensated fused features;
[0090] S2b.5. The saliency of the compensated fused features is calculated to obtain the enhanced features Fs; the expression for the information collaboration layer's processing of deep features is as follows:
[0091]
[0092] Among them, F Guide and F filtered These represent the guiding features and the features to be selected, respectively. Reshape(*) represents the reshape function, Conv represents the convolution operation, Dcos(*) represents the depthwise convolution operation, Attention(*) represents the attention calculation, and Q... guide The query vector representing the semantic attention computation is generated from the guiding features, K. filtered and V filtered F represents the key vector and numerical vector used for attention calculation, generated from the features to be selected. f For fusion features, F s This indicates a fusion feature that has undergone saliency enhancement; Salient-computation(*) is the saliency enhancement operation.
[0093] S3. Construct a self-supervised image restoration and multimodal image fusion joint learning network (SIRIFN), consisting of a multi-layered collaborative adaptive reconstruction network (MFCRNet), forming a set of twin guiding and learning networks. The guiding and learning networks are trained on different data respectively. and Data s To learn from network datasets, Data t To guide the network dataset, N represents the image pair I.i , and I i V i The number of images, where the subscript i represents the i-th image;
[0094] S3.1 Training the multi-scale feature extractor and image reconstruction decoder: Using the feature encoder E of the learning network S and the guiding network T. S and E T Feature extraction is performed, and an image reconstruction decoder D is constructed using a learning network S and a guiding network T. S and D T Image reconstruction is performed using the following expression:
[0095]
[0096] in, and These represent the infrared and visible light images reconstructed by the learning network and the image reconstruction decoder, respectively. and These represent the infrared and visible light images reconstructed by the image reconstruction decoder after being guided by the network. Data s To learn from network datasets, Data t To guide the network dataset, N represents the image pair I. i , and I i V i The quantity, D S and D T Let S and T represent the image reconstruction decoders of the learning network S and the guiding network T, respectively.
[0097] S3.2 Training of the Information Collaboration Layer: Paired infrared and degraded visible light images After passing through the multi-scale feature extractor E s The information collaboration layer and the image reconstruction decoder obtain the fused image F; similarly, the guidance network T receives... Obtaining pseudo tags The specific expression is as follows:
[0098]
[0099]
[0100] in, and Representing the learning network and the bootstrap network respectively, Data s To learn from network datasets, Data t To guide the network dataset, F represents the fused image obtained by the information collaboration layer and the image reconstruction decoder. Indicates a pseudo-tag;
[0101] S3.3 Setting the weights θ of the bootstrap network during training. t By learning the network parameter θ s The exponential moving average is used for updating, and the specific expression is as follows:
[0102] θ t =ηθ t +(1-η)θ s
[0103] Where, θ t θ represents the weights of the guiding network. s Let η represent the learned network parameters, η∈(0,1), where η represents the momentum term;
[0104] S4. Set the loss function during image reconstruction. The specific steps are as follows:
[0105] S4.1, Set in pairs and The images are fed into the learning network S and the guiding network T, respectively, and then processed by the feature encoder and the image reconstruction decoder to reconstruct clear infrared and visible light images. and The loss function of the process is expressed as follows:
[0106]
[0107]
[0108] in, L represents the consistency loss in the first stage, α and β represent the adjustment parameters, and L represents the consistency loss in the first stage. re_ir and L re_vi Let represent the reconstruction losses for infrared and visible light images, respectively, and let I and V represent the infrared and visible light images used in the training process, respectively. and These represent the reconstructed infrared and visible light images, respectively. L1(*,*) represents the L2 norm, γ represents the weights of the loss function, and L1(*,*) represents the L1 norm. SSIM (*,*)=1-SSIM(*,*), and L SSIM (*,*) represents the structural similarity loss. CC(*,*) represents the consistency loss in the first stage, and CC(*,*) represents the correlation coefficient, Φ s ,Φ t The sub-subjects represent the features output by the multi-scale feature extractors of the learning network and the guided network, respectively, and SSIM(*,*) represents the structural similarity index. This indicates that a clear infrared image has been reconstructed. This indicates that a clear visible light image has been reconstructed.
[0109] S4.2. Using paired infrared and visible light images as input, after passing through the multi-scale feature extractor trained in S3.1, their deep semantic features and shallow detail features are fused through an information synergy layer. Finally, the loss function in the image reconstruction decoder reconstructs the fused image is expressed as follows:
[0110]
[0111]
[0112]
[0113]
[0114] in, This represents the total loss function for the second stage, which is composed of the fusion loss L. fusion Consistency loss in the second phase It consists of two parts, with fusion loss L fusion Including intensity loss and gradient loss, L int L represents the pixel intensity loss between the fused image and the original image. grad This represents the gradient loss between the fused image and the original image, where ||*||1 represents the L1 norm. This represents the Sobel gradient operator with respect to *, max(*,*) indicates taking the maximum value, and H and W represent the height and width of the image, respectively. F represents the second-phase consistency loss. These represent the fused image and pseudo-label obtained in step S3.2, respectively. This represents the calculation of the L1 loss between the fused image and the pseudo-label.
[0115] Experimental data and preprocessing, this invention in M 3 Qualitative and quantitative experiments were conducted on four public datasets: FD, MSRS, RoadScene, and TNO, to comprehensively evaluate the effectiveness of the proposed method.
[0116] For evaluation metrics and comparison methods, this invention uses entropy (EN), standard deviation (SD), spatial frequency (SF), mutual information (MI), visual information fidelity (VIF), edge information transfer index (Qabf), peak signal-to-noise ratio (PSNR), and average gradient (AG) as evaluation metrics. Higher metrics generally indicate better quality of the fused image. Furthermore, this invention is compared with state-of-the-art methods, including FusionGAN, GANMcC, NestFuse, RFN-Nest, U2Fusion, SeAFusiont, LRRNet, and CDDFuse.
[0117] Comparative experiment:
[0118] In M 3 In-depth generalization analysis was performed on the FD test set (300 pairs), MSRS test set (361 pairs), RoadScene (50 pairs), and TNO (25 pairs). Experimental results are shown in Tables 1-4.
[0119] Table 1: M 3 Quantitative results on the FD dataset.
[0120]
[0121]
[0122] Table 2: Quantitative results on the MSRS dataset.
[0123]
[0124] Table 3: Quantitative results on the TNO dataset.
[0125]
[0126]
[0127] Table 4: Quantitative results on the RoadScene dataset.
[0128]
[0129]
[0130] The quantitative comparison results are shown in Tables 1-4. This invention utilizes M... 3 The FD and MSRS datasets were used to evaluate the metrics of different algorithms using their pre-defined test sets, while the TNO and RoadScene datasets were used with CDDFuse configuration as test images for quantitative comparison.
[0131] The method of this invention performs well on almost all metrics across four datasets. The optimal information entropy (EN) indicates a significant advantage in image information complexity, containing rich information. Simultaneously, the method ranks first in standard deviation (SD), spatial frequency (SF), and average gradient (AG), consistent with the qualitative analysis of the visualization results, demonstrating its sensitivity to image details and excellent preservation of image sharpness and contrast. Furthermore, thanks to the novel design of the network structure and the two-stage training strategy, superior mutual information (MI) values and high visual information fidelity (VIF) further ensure high correlation between the fused image and the original image, and outstanding performance in visual perception relevance and error sensitivity.
[0132] See the visualization results. Figure 5 Under normal daytime imaging conditions, visible light images accurately reflect the environmental information of the scene being captured. However, infrared imaging's ability to resolve target objects such as pedestrians and vehicles provides complementary information to visible light images. Nevertheless, under poor imaging conditions, interference from weather, lighting, and other factors can degrade the quality of visible light images and affect the performance of fusion algorithms. Figure 5 As shown, fusion methods such as FusionGAN, GANMcC, and U2Fuse can basically preserve the thermal target information of infrared images. However, the color information in the fused images is severely lacking, making it impossible to accurately present scene information such as streets and buildings.
[0133] In smoky environments, most methods are able to preserve salient objects in infrared images, but the performance of these fusion algorithms is also affected to varying degrees by interference information. For example, in smoky scenes, salient objects (such as pedestrians) in LRRNet and CDDFuse are inevitably weakened. NestFuse, U2Fuse, and SeAFusion, on the other hand, preserve thermal objects, but perform poorly in revealing the texture details of buildings behind the smoke.
[0134] In low-light, dark scenes, visible light images provide limited environmental information and objects are difficult to identify, while infrared images are insensitive to illumination. Therefore, the applicability of fusion algorithms in dark scenes is particularly important. Most methods preserve valuable information from infrared images but struggle to extract useful information from visible light images. Other methods, except for the one presented in this paper, are affected by overexposure from streetlights and insufficient light, resulting in images appearing shrouded in fog. Only the method presented in this paper achieves sufficient sharpness and contrast in image quality.
[0135] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An adaptive multimodal image fusion method based on knowledge embedding, characterized in that: Includes the following steps: S1. Set up the training dataset, which includes the following steps: S1.1, from M 3 Infrared images randomly selected from the FD and MSRS datasets. i ∈R H×W×1 and visible light image V i ∈R H ×W×3 ; Wherein, the subscript i represents the image number, the superscript H×W×1 represents the length, width and number of channels of the infrared image as H, W and 1 respectively, and the superscript H×W×3 represents the length, width and number of channels of the visible light image as H, W and 3 respectively; S1.
2. Based on the atmospheric scattering model, all selected visible light images are subjected to fogging processing to obtain fogged visible light images. S1.3, transfer infrared image I i ∈R H×W×1 Visible light image V i ∈R H×W×3 and fogged visible light images The training dataset is composed of... S2. Construct a multi-level collaborative adaptive reconstruction network, including a multi-scale feature extractor, an information collaboration layer, and an image reconstruction decoder. The multi-scale feature extractor is used to extract shallow detail features and deep semantic features of the image. The information collaboration layer uses a differentiated strategy to integrate information from different levels of input, which is used to fuse complementary information and filter out redundant information. The image reconstruction decoder is used to decode the output features of the information collaboration layer and reconstruct the fused features. The multi-scale feature extractor includes a three-layer Restormer Block and a high-frequency perception enhancement module. The three-layer Restormer Block is stacked to extract features. The high-frequency perception enhancement module includes two branches, one of which contains a two-layer convolutional residual block and the other integrates the Sobel operator. The process of fusing shallow detailed features in the information collaboration layer is as follows: S2a.
1. Element-wise difference of the infrared and visible light features extracted by the multi-scale feature extractor; S2a.2, Using the difference portion D as the excitation signal to generate channel weight W D And perform two compensations on infrared and visible light features; S2a.3, Obtain the fused detailed features F t ; The information collaboration layer performs the following steps for deep semantic feature fusion: S2b.
1. The element-wise sum of the input infrared features and the infrared and visible light features is used as the guiding feature and the feature to be screened, respectively, and then subjected to a 1x1 convolution. S2b.
2. Use 3x3 depthwise convolution to encode the spatial context features of the guiding features and the features to be selected, to obtain the query vector Q, key vector K and numerical vector V for attention calculation, respectively. S2b.
3. Perform a dot product operation between the query vector Q generated by the guided features and the transpose of the key vector K generated by the features to be filtered to generate a C*C attention map. S2b.
4. Multiply the attention map with the original features to be screened to obtain the compensated fused features; S2b.
5. Perform saliency calculation on the compensated fusion features to obtain the enhanced features F. s ; S3. Construct a self-supervised image restoration and multimodal image fusion joint learning network, consisting of a multi-layered collaborative adaptive reconstruction network, forming a set of twin guidance and learning networks. The guidance and learning networks are trained on different data respectively. and Data s To learn from network datasets, Data t To guide the network dataset, N represents the number of image pairs. and I i V i The number of images, where the subscript i represents the i-th image; S3.1 Training the multi-scale feature extractor and image reconstruction decoder: Using the feature encoder E of the learning network S and the guiding network T. S and E T Feature extraction is performed, and an image reconstruction decoder D is constructed using a learning network S and a guiding network T. S and D T Image reconstruction is performed using the following expression: in, and These represent the infrared and visible light images reconstructed by the learning network and the image reconstruction decoder, respectively. and These represent the infrared and visible light images reconstructed by the image reconstruction decoder after being guided by the network. Data s For learning from network datasets, Data t To guide the network dataset, N represents the number of image pairs. and I i V i The quantity, D S and D T Let S and T represent the image reconstruction decoders of the learning network S and the guiding network T, respectively. S3.2 Training of the Information Collaboration Layer: Paired infrared and degraded visible light images After passing through the multi-scale feature extractor E s The information collaboration layer and the image reconstruction decoder obtain the fused image F; similarly, the guidance network T receives... Obtaining pseudo tags The specific expression is as follows: in, and Representing the learning network and the bootstrap network respectively, Data s To learn from network datasets, Data t To guide the network dataset, F represents the fused image obtained by the information collaboration layer and the image reconstruction decoder. Indicates a pseudo-tag; S3.3 Setting the parameters θ of the bootstrap network during training. t By learning the network parameter θ s The exponential moving average is used for updating, and the specific expression is as follows: i t =eth t +(1-η)θ s Where, θ t θ represents the parameters of the guiding network. s The parameter η represents the learning network parameter, η∈(0,1), where η represents the momentum term. Higher η gives more weight to new information, while lower η focuses more on historical parameters. S4. Set the loss function during image reconstruction. The specific steps are as follows: S4.1, Set in pairs and The images are fed into the learning network S and the guiding network T, respectively, and then processed by the feature encoder and the image reconstruction decoder to reconstruct clear infrared and visible light images. and The loss function for this process is expressed as follows: in, L represents the consistency loss in the first stage, α and β represent the adjustment parameters, and L represents the consistency loss in the first stage. re_ir and L re_vi Let represent the reconstruction losses for infrared and visible light images, respectively, and let I and V represent the infrared and visible light images used in the training process, respectively. and These represent the reconstructed infrared and visible light images, respectively. L1(*,*) represents the L2 norm, γ represents the weights of the loss function, and L1(*,*) represents the L1 norm. SSIM (*,*)=1-SSIM(*,*), and L SSIN (*,*) represents the structural similarity loss. CC(*,*) represents the consistency loss in the first stage, and CC(*,*) represents the correlation coefficient, Φ s ,Φ t The sub-subjects represent the features output by the multi-scale feature extractors of the learning network and the guided network, respectively, and SSIM(*,*) represents the structural similarity index. This indicates that a clear infrared image has been reconstructed. This indicates that a clear visible light image has been reconstructed; S4.
2. Using paired infrared and visible light images as input, after passing through the multi-scale feature extractor trained in S3.1, their deep semantic features and shallow detail features are fused through an information synergy layer. Finally, the loss function in the image reconstruction decoder reconstructs the fused image is expressed as follows: in, This represents the total loss function for the second stage, which is composed of the fusion loss L. fusion Consistency loss in the second phase It consists of two parts, with fusion loss L fusion Including intensity loss and gradient loss, L int L represents the pixel intensity loss between the fused image and the original image. grad This represents the gradient loss between the fused image and the original image, where ||*||1 represents the L1 norm. This represents the Sobel gradient operator with respect to *, max(*,*) indicates taking the maximum value, and H and W represent the height and width of the image, respectively. F represents the second-phase consistency loss. These represent the fused image and pseudo-label obtained in step S3.2, respectively. This represents the calculation of the L1 loss between the fused image and the pseudo-label.
2. The adaptive multimodal image fusion method based on knowledge embedding according to claim 1, characterized in that: The calculation expression for the feature extraction process of the multi-scale feature extractor is as follows: in, and These represent the shallow detail features obtained from feature extraction of infrared and visible light images, respectively. They are obtained by concatenating the features extracted from the first two Restormer Block layers along the channel dimension. and These represent the deep semantic features extracted from infrared and visible light images, respectively. Concate(*) represents the feature map concatenation operation. RTB 1 RTB 2 RTB 3 These represent Restormer Blocks for layers 1-3, e I and e V These represent the infrared and visible light edge features extracted by the high-frequency perception enhancement module HPEM, respectively. HPEM(*,*) indicates that edge extraction is performed on the internal image data.
3. The adaptive multimodal image fusion method based on knowledge embedding according to claim 1, characterized in that: The calculation expression for the shallow detail feature extraction process of the information collaboration layer is as follows: D=F v -F i Where σ represents the sigmoid activation function, GAP represents the global average pooling operation, ReLU represents the ReLU activation function, D represents the difference component, and F i and F v W represents infrared and visible light characteristics, respectively. D Indicates channel weight, and These represent the infrared and visible light image features after initial compensation, respectively. g and w l These represent the weights calculated using global attention and local attention, respectively. Conv represents the convolution operation, and ⊙ represents the dot product operation. For matrix multiplication, w i The weights of the infrared features, w v Weights of visible light features.
4. The adaptive multimodal image fusion method based on knowledge embedding according to claim 1, characterized in that: The expression for deep semantic feature fusion in the information collaboration layer is as follows: Among them, F Guide and F filtered These represent the guiding features and the features to be selected, respectively. Reshape(*) represents the reshape function, Conv represents the convolution operation, Dcos(*) represents the depthwise convolution operation, Attention(*) represents the attention calculation, and Q... guide The query vector representing semantic attention computation is generated from guiding features, K filtered and V filtered F represents the key vector and numerical vector used for attention calculation, generated from the features to be selected. f For fusion features, F s This indicates a fusion feature that has undergone saliency enhancement; Salient_computation(*) is the saliency enhancement operation.
5. The adaptive multimodal image fusion method based on knowledge embedding according to claim 4, characterized in that: The expression for calculating significance is as follows: Where Contrast represents saliency calculation, K represents the number of pixels, and x j This represents the value of the j-th pixel in the fused features. This represents the mean of the feature map.