A Self-Supervised Visible and Infrared Image Fusion Method Based on Knowledge Distillation

Through a self-supervised method based on knowledge distillation, the teacher and student network is designed to fully integrate the characteristics of visible light and infrared images, and the problems of poor fusion effect and insufficient network performance in the existing technology are solved, and better perceived effect and real-time performance are achieved.

CN117372271BActive Publication Date: 2025-06-24SHANGHAI INST OF MICROSYSTEM & INFORMATION TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311207197.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-19
Publication Date
2025-06-24
Estimated Expiration
2043-09-19

AI Technical Summary

Technical Problem

When the prior art fuses visible light and infrared images, it is difficult to fully retain the useful features of the two types of images, and the human eye perception and image information maintenance effects are difficult to balance. In the fields of real-time video surveillance, the network burden is relatively large and the performance is insufficient.

Method used

Using a self-supervised image fusion method based on knowledge distillation, by designing the teacher network and the student network, the teacher network extracts visible and infrared image features and reconstructs the original image. The student network extracts features and generates fusion images under the guidance of the teacher network, and optimizes the fusion effect using perceived significance weights and multiple loss functions.

Benefits of technology

It achieves a more fully integrated useful information of visible light and infrared images, balances the human eye perception and image information retention effect, reduces network parameters, and improves network performance and real-time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117372271B_ABST
    Figure CN117372271B_ABST
Patent Text Reader

Abstract

The present invention relates to a self-supervised visible light and infrared image fusion method based on knowledge distillation, comprising: receiving a visible light image and an infrared image to be fused; splicing the visible light image and the infrared image to obtain a spliced image; inputting the spliced image into an image fusion model to obtain a fused image; wherein, the image fusion model includes a teacher network part and a student network part; the teacher network part includes: a teacher encoder for extracting features of the spliced image to obtain first feature information; a teacher first decoder and a teacher second decoder respectively for reconstructing the visible light image and the infrared image based on the first feature information; the student network part includes: a student encoder for extracting features of the spliced image under the guidance of the teacher encoder to obtain second feature information; a student decoder for generating a fused image of the visible light image and the infrared image based on the second feature information. The present invention can fully fuse the useful information of the two types of images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image fusion, and particularly to a self-supervised visible light and infrared image fusion method based on knowledge distillation. Background Art

[0002] In tasks such as video surveillance and pedestrian detection, there are often problems of insufficient night lighting intensity and obstacle occlusion, and even extreme weather such as sand, wind, rain, and snow in border surveillance scenarios. Due to the limitations of the imaging principles of single-type sensors, visible light images obtained by visible light sensors have rich textures, but are prone to losing target information under conditions of lack of light, object occlusion, and extreme weather. Infrared sensors can obtain temperature difference images by capturing temperature differences, which can avoid the influence of light and occlusion, and can often present prominent targets, but cannot obtain more abundant foreground and background texture details. Therefore, simultaneously using visible light and infrared sensors to combine visible light images and infrared images for fusion to obtain a fusion image with more sufficient information under various lighting conditions is of great significance for human eye perception and further advanced vision tasks such as target detection and semantic segmentation.

[0003] There is no standard fusion image as a guide for the image fusion task, which belongs to an unsupervised learning task. Currently, image fusion algorithms based on deep learning usually adopt a network structure of feature extraction, feature fusion, and feature reconstruction. By designing a loss function, it is guided that the fusion image obtained by the network is consistent with the original visible light and infrared images in terms of semantics, structure, and other features, and at the same time, as many texture details represented mathematically as gradients are retained as possible. Moreover, considering that the human eye is more accustomed to observing visible light images, in order to meet the needs of human eye observation, it is expected that the information-rich fusion image is more inclined to be consistent with the feature expression of the visible light image to obtain better visualization results. In existing algorithms, the following limitations mainly exist:

[0004] (1) The self-feature extraction of visible light and infrared images of two modalities is insufficient, resulting in difficulty in retaining useful features of each modality after further fusion;

[0005] (2) It is difficult to balance the effects of human eye perception and image information retention;

[0006] (3) In fields such as video surveillance, there is a high demand for real-time performance. How to minimize the network burden and improve network performance has always been an issue that needs attention. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to provide a self-supervised visible light and infrared image fusion method based on knowledge distillation, which can fully fuse the useful information of two types of images and better balance the effects of human eye perception and image information retention.

[0008] The technical solution adopted by the present invention to solve its technical problems is: to provide a self-supervised visible light and infrared image fusion method based on knowledge distillation, including the following steps:

[0009] Receive the visible light image and the infrared image to be fused;

[0010] Stitch the visible light image and the infrared image to obtain a stitched image;

[0011] Input the stitched image into an image fusion model to obtain a fused image;

[0012] Among them, the image fusion model includes a teacher network part and a student network part;

[0013] The teacher network part includes:

[0014] A teacher encoder, which is used to extract the features of the stitched image to obtain first feature information;

[0015] A teacher first decoder, which is used to reconstruct the visible light image based on the first feature information;

[0016] A teacher second decoder, which is used to reconstruct the infrared image based on the first feature information;

[0017] The student network part includes:

[0018] A student encoder, which is used to extract the features of the stitched image under the guidance of the teacher encoder to obtain second feature information;

[0019] A student decoder, which is used to generate a fused image of the visible light image and the infrared image based on the second feature information.

[0020] The teacher encoder includes a 1×1 convolutional layer and 4 ACmix modules arranged in sequence, where both the 1×1 convolutional layer and the 4 ACmix modules use LeakyReLU as the activation function; the ACmix module includes 1 3×3 convolutional kernel and 4 7×7 attention kernels.

[0021] The teacher first decoder and the teacher second decoder have the same structure, both including 5 reflect_conv blocks arranged in sequence. The 5 reflect_conv blocks are connected in a centrosymmetric manner, and the first 4 reflect_conv blocks use LeakyReLU as the activation function, and the last reflect_conv block uses the Tanh activation function; the reflect_conv block includes a 3×3 convolutional operation layer with a stride of 1.

[0022] The student encoder includes a 1×1 convolutional layer and 4 reflect_conv blocks arranged in sequence. Among them, the 1×1 convolutional layer and the 4 reflect_conv blocks both use LeakyReLU as the activation function.

[0023] The structure of the student decoder is the same as that of the teacher's first decoder.

[0024] When training the teacher network, self-supervised training is adopted. The overall objective function L of the teacher network T is: where L ref_vis represents the reconstruction loss between the reconstructed visible light image and the original visible light image, and L ref_inf represents the reconstruction loss between the reconstructed infrared image and the original infrared image; β is the weight coefficient; L1 is the pixel-level image reconstruction loss, expressed as L1 = ||T(I), I||1, and L grad is the gradient loss, expressed as: L vgg is the feature perception difference loss, expressed as: L vgg = ||VGG(T(I)), VGG(I)||1, and L ssim is the structural similarity loss, expressed as: L ssim = 1 - SSIM(T(I), I), where T(I) is the reconstructed visible light image or the reconstructed infrared image, I is the original visible light image or the original infrared image, || ||1 represents the L1 norm calculation; represents calculating the gradient using the sobel operator, VGG() represents extracting features using the trained VGG19 network, and SSIM() calculates the similarity in brightness, contrast, and structure between images; t1, t2, t3, t4 are weighting coefficients.

[0025] When training the student network, the parameters of the trained teacher network are loaded, and the teacher network does not participate in the gradient backpropagation and parameter update; the overall objective function L of the student network S is expressed as: L S = L fuse + αL teach where L fuse is the fusion loss, expressed as: L fuse = s1L′1 + s2L′ ssim + s3L′ grad , L′1 is the pixel-level image generation loss, expressed as: L′1 = ω||I fuse , I vis ||1 + (1 - ω)||I fuse , I inf ||1, L′ ssimIt is the structural similarity loss, denoted as: L′ ssim = ω(1 - SSIM(I fuse , I vis )) + (1 - ω)(1 - SSIM(I fuse , I inf ))), L′ grad is the gradient loss, denoted as: s1, s2, s3 are weighting coefficients, I fuse is the generated fused image, I vis is the original visible light image, I inf is the original infrared image, || ||1 represents the L1 norm calculation; represents calculating the gradient using the sobel operator, SSIM() calculates the similarity in brightness, contrast, and structure between images, ω is the perceptual saliency weight, denoted as PS vis and PS inf respectively represent the perceptual saliency metrics of the original visible light image and the original infrared image; L teach is the distillation loss, denoted as F T and F S respectively represent the output feature maps of the last layer of the teacher encoder and the output feature maps of the last layer of the student encoder, Γ represents the hyperparameter temperature in knowledge distillation, c represents the channel, C represents the total number of channels, i represents the position in the channel, W·H represents the width and height, φ() represents the calculation using the softmax function; α is the balance coefficient.

[0026] The splicing of the visible light image and the infrared image to obtain a spliced image is specifically as follows:

[0027] Convert the visible light image into a ycbcr image and extract the y-channel part to obtain a y-channel image;

[0028] Grayscale the infrared image to obtain an infrared grayscale image;

[0029] Splice the y-channel image and the infrared grayscale image to obtain a spliced image.

[0030] Beneficial effects

[0031] Due to the above technical solution, compared with the prior art, the present invention has the following advantages and positive effects: The present invention designs a teacher network and a student network through the idea of knowledge distillation. By training the teacher network to extract visible light and infrared image features, and respectively reconstructing the visible light and infrared images, the trained teacher network is used to guide the student network to extract features and generate a fused image. The process of the teacher network reconstructing the original image can use the original image input to the network as a reference to guide the training. Through sufficient training under the supervision of the reference image, useful features of the two modalities can be well learned. The features extracted by the powerful teacher network guide the feature extraction process of the student network, so that the student network can achieve good performance with a much lighter and smaller network while greatly reducing the network parameters. In addition, by designing a loss function and using perceptual saliency measurement as an important weight parameter, useful information of the two types of images can be more fully fused, and the human eye perception and image information preservation effect can be better balanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is a flowchart of the self-supervised visible light and infrared image fusion method based on knowledge distillation according to an embodiment of the present invention;

[0033] Figure 2 is a schematic structural diagram of the fusion model in an embodiment of the present invention;

[0034] Figure 3 is a comparison diagram of experimental results of "elecbike" fusion;

[0035] Figure 4 is a comparison diagram of experimental results of "labMan" fusion. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] The following further elaborates the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

[0037] An embodiment of the present invention relates to a self-supervised visible light and infrared image fusion method based on knowledge distillation, as Figure 1 shown, including the following steps:

[0038] Step 1, receiving a visible light image and an infrared image to be fused.

[0039] Step 2, splicing the visible light image and the infrared image to obtain a spliced image. This step specifically includes:

[0040] Convert the visible light image into a YCbCr image, and extract the Y channel part to obtain the Y channel image y_vis_image;

[0041] Grayscale the infrared image to obtain the infrared grayscale image inf_image;

[0042] Stitch the Y channel image y_vis_image and the infrared grayscale image inf_image to obtain a stitched image.

[0043] Step 3: Input the stitched image into an image fusion model to obtain a fused image fuse_image. Finally, in tasks that require color restoration, stitch the fused image fuse_image with the Cb and Cr channels of the visible light image and convert it back to the RGB format to obtain the final color fused image.

[0044] The image fusion model in this embodiment utilizes the knowledge distillation idea. As Figure 2 shown, it includes a teacher network and a student network. Train the teacher network to extract the features of the visible light image and the infrared image, and respectively reconstruct the visible light image and the infrared image. Use the trained teacher network to guide the student network to extract features and generate a fused image. The process of the teacher network reconstructing the original image can use the original image input to the network as a reference for training. Through sufficient training under the supervision of the reference image, useful features of the two modalities can be well learned. The features extracted by the powerful teacher network guide the feature extraction process of the student network, so that the student network can achieve good performance with a much lighter and smaller network while greatly reducing the network parameters.

[0045] In this embodiment, the teacher network and the student network have similar structures, both including an encoding part and a decoding part. After the two original images are stitched, they are sent to the encoding part to extract features, and then the image is reconstructed through the decoding part. The teacher network has a teacher encoder T_encoder and two teacher decoders, which are respectively denoted as the first teacher decoder T_vis_decoder and the second teacher decoder T_inf_decoder, and are respectively used to reconstruct the visible light image and the infrared image; the student network has a student encoder S_encoder and a student decoder S_decoder, and the student decoder S_decoder is used to generate a fused image.

[0046] The first layer of the T_encoder uses a 1*1 convolution to change the channel size, and then extracts global and local features through 4 ACmix modules. Each layer in the T_encoder uses LeakyReLU as the activation function. The ACmix module was proposed by Tsinghua University in 2021, integrating the advantages of convolutional and transformer networks. The ACmix module used in this embodiment contains a 3*3 convolutional kernel, a 7*7 attention kernel, 4 attention heads, and a stride of 1. The network structures of T_vis_decoder and T_inf_decoder are exactly the same, both containing 5 reflect_conv blocks. Each reflect_conv block contains a 3*3 convolution operation with a stride of 1, using a centrosymmetric connection method. The first four layers use the LeakyReLU activation function, and the last layer uses the Tanh activation function. Correspondingly, the first layer of the S_encoder uses a 1*1 convolution to change the channel size, and then extracts features through 4 simple reflect_conv blocks. Each layer in the S_encoder uses LeakyReLU as the activation function. S_decoder is exactly the same as T_vis_decoder and T_inf_decoder. Compared with the teacher network, the network parameters of the student network are greatly reduced.

[0047] In this embodiment, a two-stage training process is used to train the teacher network and the student network.

[0048] In the first stage, the teacher network is trained in a self-supervised manner. The visible light image is converted to the ycbcr format and the y channel is extracted to obtain y_vis_image. The y_vis_image and the single-channel infrared image inf_image are concatenated in channels and used as the input of the teacher network. First, features are extracted through the T_encoder, and the last layer of feature maps is used as the input of T_vis_decoder and T_inf_decoder. T_vis_decoder reconstructs the visible light image, and T_inf_decoder reconstructs the infrared image. Based on the two types of original images as references, the overall objective function for training the teacher network can be defined as:

[0049] L T =L ref_vis +βL ref_inf

[0050] L ref_vis / L ref_inf =t1L1+t2L grad +t3L vgg +t4L ssim

[0051] Among them, L ref_visDenote the reconstruction loss between the reconstructed visible light image and the original visible light image as \(L\). ref_inf Denote the reconstruction loss between the reconstructed infrared image and the original infrared image. \(\beta\) is the weight coefficient of the two reconstruction losses. In this embodiment, it is desired to reconstruct both types of images simultaneously, and they are of equal importance. Therefore, \(\beta\) is set to 1. \(L\). ref_vis and \(L\). ref_inf both include the pixel-level image reconstruction loss \(L_1\), expressed as: \(L_1 = ||T(I), I||_1\), which is used to reduce pixel-level differences; the gradient loss \(L\). grad , expressed as It calculates the gradient using the sobel operator and calculates the distance between the gradient maps using the \(L_1\) norm; the feature perception difference loss \(L\). vgg , expressed as: \(L\). vgg \(= ||VGG(T(I)), VGG(I)||_1\). It extracts features from the original image and the reconstructed image respectively using the pre-trained VGG19 network and measures the difference using the \(L_1\) norm; the structural similarity loss \(L\). ssim , expressed as \(L\). ssim \(= 1 - SSIM(T(I), I)\). It is based on the commonly used image quality evaluation index ssim and is used to constrain the similarity in brightness, contrast, and structure between the reconstructed image and the original image. Among them, \(T(I)\) is the reconstructed visible light image or the reconstructed infrared image, \(I\) is the original visible light image or the original infrared image, \(|| ||_1\) represents the calculation using the \(L_1\) norm. Denote the calculation of the gradient using the sobel operator, \(VGG()\) represents the extraction of features using the trained VGG19 network, and \(SSIM()\) calculates the similarity in brightness, contrast, and structure between images; \(t_1, t_2, t_3, t_4\) are the weighting coefficients.

[0052] In the second stage, train the student network. Load the parameters of the trained teacher network. The teacher network does not participate in the gradient backpropagation and parameter update. Similar to the teacher network, concatenate the channels of \(y_{vis\_image}\) and the single-channel infrared image \(inf\_image\) as the input of the student network and the frozen teacher network. The output feature map of \(T\_encoder\) of the teacher network guides \(S\_encoder\) to extract features. \(S\_decoder\) takes the last-layer feature map of \(S\_encoder\) as the input and generates the fused image. Calculate the distillation loss between the last-layer output feature maps of \(T\_encoder\) and \(S\_encoder\) so that the student network can learn the knowledge of the teacher and improve the performance of the student network. The student network obtains guidance from the teacher network and simultaneously uses the two original input images as references. Among them, calculate the perceptual saliency of the paired visible light and infrared images as the weight to provide a tendency to be closer in image features, so as to more fully fuse the useful information of the two types of images.

[0053] The overall objective function for student network training can be defined as:

[0054] L S = L fuse + αL teach

[0055] where α is the balance coefficient between the fusion loss L fuse of the generated fused image and the distillation loss L teach . The expression of L fuse is as follows:

[0056] L fuse = s1L′1 + s2L′ ssim + s3L′ grad

[0057] L′1 = ω||I fuse , I vis ||1 + (1 - ω)||I fuse , I inf ||1

[0058] L′ ssim = ω(1 - SSIM(I fuse , I vis )) + (1 - ω)(1 - SSIM(I fuse , I inf ))

[0059]

[0060] where I vis is the original visible light image, I inf is the original infrared image, I fuse is the generated fused image, s1, s2, s3 are the weighted coefficients of the three losses, L′1 is the pixel-level image generation loss, L′ ssim is the structural similarity loss, L′ grad is the gradient loss. ω is the perceptual saliency weight, defined based on the perceptual saliency metric PS in the GFCE paper. PS is calculated by multiplying the window-averaged power spectrum entropy and the gradient magnitude and normalizing, considering both the frequency domain features and edge information of the image. The larger the value, the richer the information in the image, and the smaller the value, the more distortion, noise, or missing details in the image. ω is defined as follows:

[0061]

[0062] where PS vis and PS inf represent the perceptual saliency metrics of the original visible light image and the original infrared image, respectively.

[0063] Pixel-level image generation loss L′1 and structural similarity loss L′ ssim , use ω to weight the calculation results of the generated fused image with the visible light and infrared images, so that the fused image is more inclined to be consistent with the image with richer information. The gradient loss always expects to be consistent with the one with larger gradient in the two original images to ensure that the fused image has more sufficient gradient information and retains high-frequency texture details to avoid image blurring.

[0064] Distillation loss L teach Adopt the idea of channel distillation and use KL divergence to compare the distribution differences between the feature maps of the teacher network and the student network. Compared with the distillation method in the spatial dimension, channel dimension distillation avoids over-strong constraints, makes the distillation process pay more attention to the most significant regions of each channel, ensures the consistency of the role of each channel, and requires less computational cost during the training process. Distillation loss L teach Is expressed as:

[0065]

[0066]

[0067]

[0068] Among them, F T and F S respectively represent the output feature maps of the last layer of the teacher encoder and the output feature maps of the last layer of the student encoder, c represents the channel, C represents the total number of channels, i represents the position in the channel, W·H represents the width and height, φ() represents the calculation of the softmax function, and Γ represents the hyperparameter temperature in knowledge distillation, which is used to adjust the degree of attention to negative samples during the training process. The larger Γ is, the more dispersed the probability is, that is, the larger the concerned space is.

[0069] To verify the effectiveness of this embodiment, a model was built and experimented based on the pytorch platform, trained on the MSRS dataset, and tested on the publicly available VIFB test dataset.

[0070] Experimental data: The training set MSRS contains 1444 pairs of aligned high-quality infrared and visible image pairs. The publicly available VIFB (Visible and Infrared Fusion Benchmark) is used as the test set, which contains 21 carefully selected image pairs, covering typical scenarios such as exposure, strong light, insufficient light, and occlusion.

[0071] Training and testing details: During the training process, the image pairs input into the teacher and student networks are cropped to a resolution of 96*96, and the input data dimension is 2. The output dimensions of each layer of the teacher network T_encoder are 32, 64, 128, 256, 512 respectively, and the output dimensions of each layer of T_decoder are 256, 128, 64, 32, 1 respectively. When training the teacher model, the Adam optimizer is used to train for 200 epochs, the number of samples in each training batch is 2, the base learning rate is 1e-4, and the learning rate is halved in the last 100 epochs. When training the student model, the Adam optimizer is also used to train for 200 epochs, the number of samples in each training batch is 4, the base learning rate is 1e-4, and the learning rate is halved in the last 100 epochs. The weights t1, t2, t3, t4 for training the teacher network are set to 20, 20, 2, 0.01. The weights s1, s2, s3 for training the teacher network are set to 2, 2, 40, α is set to 0.05, and Γ is set to 3. All experiments are completed on 2 NVIDIA 2080ti graphics cards.

[0072] Model evaluation metrics: Seven general evaluation metrics are selected. Mutual Information (Mutinf), Average Gradient (Avg_gradient), Edge Intensity, Quality Assessment by Boundary Focusing (Qabf), Standard Deviation (SD), Spatial Frequency, and Quality with Color and Brightness (Qcb). These metrics are used to measure different aspects of image features and quality, and can comprehensively evaluate factors such as image sharpness, detail richness, edge information, and color fidelity.

[0073] Experimental results: Six commonly used traditional methods (ADF, FPDE, MSVD, GFF, GTF, IFEVIP) and five latest deep learning methods (PIAfusion, SeAFusion, SwinFusion, U2Fusion, YDTR) for visible-light infrared image fusion are selected for comparison. Taking two images named "elecbike" and "labMan" as examples, the fusion results of each algorithm are as follows. The first row is the original visible-light image and infrared image, and the second to fourth rows are the fusion results of ADF, FPDE, MSVD, GFF, GTF, IFEVIP, PIAfusion, SeAFusion, SwinFusion, U2Fusion, YDTR, and the fusion result of this embodiment in sequence.

[0074] In Figure 3 this case, the high-contrast targets provided by the infrared image are well integrated in the strong exposure area of the vehicle lamp, and the texture details of the branches in the dark area below are also retained; in Figure 4 this case, the clearest and most complete human face is obtained by this embodiment. The details in the ceiling area are kept intact, and there is no large-area texture loss. Moreover, the temperature difference changes provided by the infrared image are well combined in the human body part.

[0075] The calculation results of 7 indicators for 21 pairs of test images on the VIFB test dataset are shown in Table 1 below (optimal indicators: bold + underlined, sub-optimal indicators: underlined).

[0076] Table 1 Calculation results of image indicators for 21 pairs of test images on the VIFB test dataset

[0077]

[0078] The method of this embodiment achieves the optimal results for 2 out of 7 indicators. The Mutinf and Spatial_frequency indicators have obvious advantages, indicating that the fused image retains more information of the original image and has rich texture and details; 5 indicators obtain sub-optimal results, indicating that the fused image has clear edge boundaries and good color fidelity, and the overall quality of the image remains good.

[0079] In summary, this embodiment can preferably maintain the target contrast in the strong exposure area, retain more texture details in the dark area, and can obtain clear edges, richer information and good visual effects. At the same time, since the training set and the test set are not from the same dataset, through qualitative and quantitative result analysis, it is proved that this embodiment has good performance and robustness.

Claims

1. A self-supervised visible and infrared image fusion method based on knowledge distillation, characterized in that It includes the following steps: Receive a visible light image and an infrared image to be fused; Stitch the visible light image and the infrared image to obtain a stitched image; Input the stitched image into an image fusion model to obtain a fused image; Wherein, the image fusion model includes a teacher network part and a student network part; The teacher network part includes: A teacher encoder for extracting the features of the stitched image to obtain first feature information; A first teacher decoder for reconstructing the visible light image based on the first feature information; A second teacher decoder for reconstructing the infrared image based on the first feature information; The student network part includes: A student encoder for extracting the features of the stitched image under the guidance of the teacher encoder to obtain second feature information; A student decoder for generating a fused image of the visible light image and the infrared image based on the second feature information; When training the teacher network, a self-supervised manner is adopted for training, and the overall objective function \(L\) of the teacher network T is as follows: Among them, L ref_vis represents the reconstruction loss between the reconstructed visible light image and the original visible light image, and L ref_inf represents the reconstruction loss between the reconstructed infrared image and the original infrared image; β is the weight coefficient; L1 is the pixel-level image reconstruction loss, expressed as L1 = ||T(I), I||1, and L grad is the gradient loss, expressed as: L grad = ||▽(T(I)), ▽(I)||1, and L vgg is the feature perception difference loss, expressed as: L vgg = ||VGG(T(I)), VGG(I)||1, and L ssim is the structural similarity loss, expressed as: L ssim = 1 - SSIM(T(I), I), where T(I) is the reconstructed visible light image or the reconstructed infrared image, I is the original visible light image or the original infrared image, || ||1 represents the L1 norm calculation; ▽() represents calculating the gradient using the sobel operator, VGG() represents extracting features using the trained VGG19 network, SSIM() calculates the similarity in brightness, contrast, and structure between images; t1, t2, t3, t4 are the weighting coefficients; When training the student network, load the parameters of the trained teacher network, and the teacher network does not participate in gradient backpropagation and parameter update; the overall objective function L of the student network S is expressed as: L S = L fuse + αL teach , where L fuse is the fusion loss, expressed as: L fuse = s1L′1 + s2L′ ssim + s3L′ grad , L′1 is the pixel-level image generation loss, expressed as: L′1 = ω||I fuse , I vis ||1 + (1 - ω)||I fuse , I inf ||1, L′ ssim is the structural similarity loss, expressed as: L′ ssim = ω(1 - SSIM(I fuse , I vis )) + (1 - ω)(1 - SSIM(I fuse , I inf )), L′ grad is the gradient loss, expressed as: L′ grad = ||▽(I fuse ), max(▽(I vis ), ▽(I inf ))||1, s1, s2, s3 are weighting coefficients, I fuse is the generated fused image, I vis is the original visible light image, I inf is the original infrared image, ||||1 represents the L1 norm calculation; ▽() represents calculating the gradient using the sobel operator, SSIM() calculates the similarity in brightness, contrast, and structure between images, and ω is the perceptual saliency weight, expressed as PS vis and PS inf respectively represent the perceptual saliency metrics of the original visible light image and the original infrared image; L teach is the distillation loss, expressed as: F T and F S respectively represent the output feature maps of the last layer of the teacher encoder and the output feature maps of the last layer of the student encoder. Γ represents the hyperparameter temperature in knowledge distillation, c represents the channel, C represents the total number of channels, i represents the position in the channel, W·H represents the width and height, and φ() represents the calculation of the softmax function; α is the balance coefficient.

2. The self-supervised visible and infrared image fusion method based on knowledge distillation according to claim 1, wherein The teacher encoder includes a 1×1 convolutional layer and 4 ACmix modules arranged in sequence, wherein both the 1×1 convolutional layer and the 4 ACmix modules use LeakyReLU as the activation function; the ACmix module includes 1 3×3 convolutional kernel and 4 7×7 attention kernels.

3. The self-supervised visible and infrared image fusion method based on knowledge distillation according to claim 1, wherein, The first teacher decoder and the second teacher decoder have the same structure, and both include 5 reflect_conv blocks arranged in sequence. The 5 reflect_conv blocks are connected in a centrosymmetric manner, and the first 4 reflect_conv blocks use LeakyReLU as the activation function, and the last reflect_conv block uses the Tanh activation function; the reflect_conv block includes a 3×3 convolutional operation layer with a stride of 1.

4. The self-supervised visible and infrared image fusion method based on knowledge distillation according to claim 1, wherein The student encoder includes a 1×1 convolutional layer and 4 reflect_conv blocks arranged in sequence, wherein both the 1×1 convolutional layer and the 4 reflect_conv blocks use LeakyReLU as the activation function.

5. The self-supervised visible and infrared image fusion method based on knowledge distillation according to claim 1 or 3, characterized in that, The structure of the student decoder is the same as that of the first teacher decoder.

6. The self-supervised visible and infrared image fusion method based on knowledge distillation according to claim 1, wherein The specific method of stitching the visible light image and the infrared image to obtain a stitched image is as follows: Convert the visible light image into a ycbcr image and extract the y-channel part to obtain a y-channel image; Perform grayscale processing on the infrared image to obtain an infrared grayscale image; Stitch the y-channel image and the infrared grayscale image to obtain a stitched image.

Citation Information

Patent Citations

  • Visible light and infrared fusion target detection method based on dynamic weight positioning distillation

    CN115661597A

  • Lightweight monocular near-infrared silent human face living body discrimination method

    CN116092201A