An infrared and visible light image fusion method based on multi-mode features
By constructing a multi-mode feature encoder-decoder network and embedded Transformer fusion strategy, combining entropy, gradient and significance information, a multi-mode adaptive loss function is designed, which solves the problem of lack of global information in the fusion of infrared and visible light images in the prior art, and achieves a high-quality image fusion effect.
Patent Information
- Application Number
- CN202210244332.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-14
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-03-14
AI Technical Summary
The existing infrared and visible image fusion methods lack global information during the fusion process, resulting in low quality of the fusion result, and deep learning models ignore global information during feature extraction, resulting in unclear image edges.
Using the infrared and visible image fusion method based on multi-mode features, a multi-mode feature encoder-decoder network is constructed, combining entropy, gradient and significance information, a multi-mode adaptive loss function is designed, and a fusion weight learning model embedded in the Transformer fusion strategy is used to optimize the model to enhance information transmission.
It realizes adaptive fusion of multi-modal features during the fusion process, enhances information transmission, and avoids the problem of weakening of information of different features. The resulting fusion image has better visual effects and robustness.
Smart Images

Figure CN114639002B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of image fusion, and specifically relates to an infrared and visible light image fusion method based on multi-modal features. Background Art
[0002] Image fusion refers to combining images obtained from different types of sensors to generate a robust or information-rich image for subsequent processing or decision-making. Complex applications require comprehensive information about a specific scenario to enhance the overall understanding of the scenario. Single-mode sensors can only perceive single-scenario information of the target and cannot perform multi-granularity perception of the target. Therefore, fusion technology plays an increasingly important role in modern applications and computer vision.
[0003] Due to the limitations of physical sensors, the scene information captured by infrared and visible light images is very different. Visible light images capture reflected light. Such images usually have high spatial resolution, rich color, texture details, and high-contrast features, which are suitable for human visual perception but are easily affected by light. For example, in scenes with insufficient light such as bad weather or at night, the image quality drops significantly. Infrared images capture thermal radiation. Infrared images that describe the thermal radiation of objects can resist interference such as bad weather and insufficient light, but usually have low spatial resolution and lack information such as image texture and color. The fusion of infrared and visible light images refers to combining infrared and visible light images of the same scene and using the complementarity of these two types of images to generate a highly robust and information-rich fused image. The infrared and visible light image fusion technology has been widely applied in fields such as target detection, image enhancement, video surveillance, and remote sensing.
[0004] The infrared and visible light image fusion methods are mainly divided into traditional methods and deep learning methods. Traditional image fusion methods mainly use multi-scale transform (MST), sparse representation (SR), saliency-based, hybrid models, and other methods. These methods have achieved good fusion performance, but problems such as the need for manual feature crafting and high computational complexity still exist. In deep learning-based methods, models such as FusionGAN, Attention FGAN, and Nestfuse improve the disadvantages of traditional methods, but also have certain limitations. First, deep learning networks usually directly extract feature maps from the previous convolutional layer, ignoring global information, resulting in low-quality fusion results. Second, in the encoder-decoder model with a fusion strategy, a simple fusion strategy may make the image edges unclear. Finally, the design of the loss function also affects the network fusion effect. An inappropriate loss function design will not only slow down the convergence speed but also cause problems such as artifacts and blurred boundaries in the fusion results. Summary of the Invention
[0005] The object of the present invention is to provide an infrared and visible light image fusion method based on multi-modal features, which solves the problem of lack of global information in the fusion process, adaptively fuses multi-modal features, combines saliency information, and finally achieves good fusion.
[0006] To achieve the above task, the present invention adopts the following technical solutions:
[0007] An infrared and visible light image fusion method based on multi-modal features, characterized by comprising the following steps:
[0008] Step 1, construct a feature extraction and image reconstruction network, based on a multi-scale convolutional network, and optimize to generate a multi-modal feature encoder-decoder network under the guidance of a loss function;
[0009] Step 2, extract infrared and visible light multi-modal features through the encoder-decoder network, measure the multi-modal features using entropy, gradient, and saliency, and design a multi-modal adaptive loss.
[0010] Step 3, construct a fusion weight learning model embedding a Transformer fusion strategy, and assign weights to the fusion model;
[0011] Step 4, obtain the saliency map of the infrared image as a label, and add the saliency label as the region selection for optimizing the fusion network;
[0012] Step 5, cascade the fusion weight learning model embedding the Transformer fusion strategy with the encoder-decoder to construct an infrared and visible light image fusion network, and train the infrared and visible light image fusion network using the saliency label and multi-modal loss.
[0013] According to the present invention, the structure of the encoder in Step 1 is as follows: it includes 1 1×1 convolutional layer and 4 encoding convolutional modules ECB10, ECB20, ECB30, and ECB40, and each encoding convolutional module includes 2 3×3 convolutional layers and a max pooling layer.
[0014] The structure of the decoder in Step 1 is as follows: it includes 1 1×1 convolutional layer and 6 decoding convolutional modules DCB30, DCB20, DCB21, DCB10, DCB11, and DCB12, and each decoding convolutional module includes two 3×3 convolutional layers.
[0015] Specifically, the specific connection method of the decoder in step 1 is as follows: In the first and second scales, horizontal dense skip connections are adopted. In the channel connection method, the final fused feature of the second scale is skip-connected to the input of DBC21, the final fused feature of the first scale is skip-connected to the inputs of DCB11 and DCB12, and the output of DCB10 is skip-connected to the input of DCB12. Through the horizontal dense skip connections, the depth features of all intermediate layers are used for feature reconstruction, improving the reconstruction ability of multi-scale depth features; in the decoding sub-network, vertical dense connections are established in all scales. In the upsampling method, the final fused feature of the fourth scale is connected to the input of DCB30, the final fused feature of the third scale is connected to the input of DCB20, the final fused feature of the second scale is connected to the input of DCB10, the output of DCB30 is connected to the input of DCB21, the output of DCB20 is connected to the input of DCB11, and the output of DCB21 is connected to the input of DCB12. Through the vertical dense upsampling connections, the features of all scales are used for feature reconstruction, further improving the reconstruction ability of multi-scale depth features.
[0016] Furthermore, the loss function L of the encoder-decoder network ED , which is the pixel consistency and structural similarity between the input image and the output image, as shown in formula (1):
[0017] L ED = L p + βL ssin (1)
[0018] where L p is the pixel consistency loss, and L ssin is the structural similarity loss;
[0019] The pixel consistency loss L p is as shown in formula (2):
[0020]
[0021] The structural similarity loss L ssin is as shown in formula (3):
[0022] L ssim = 1 - ssim(O, I) (3)
[0023] where O is the network output image and I is the input image.
[0024] Furthermore, the step of measuring the multi-modal features using entropy, gradient, and saliency in step 2 includes the following steps:
[0025] Step 2.1: Calculate the entropy of the features output by the encoder, compare the entropy values of the features at each scale, and classify the feature with the highest entropy, which contains the most content and details, as the content feature;
[0026] Step 2.2: Use the Sobel gradient operator to calculate the gradient of the input image of the encoder, downsample the gradient, subtract it from each feature, and calculate the mean. The feature with the smallest mean contains more structural features such as contours and edges, and is classified as the edge structural feature;
[0027] Step 2.3: Use the saliency extraction algorithm to calculate the saliency map of the input image of the encoder, downsample the saliency map, subtract it from each feature, and calculate the mean. The feature with the smallest mean has a certain distinction between the foreground object and the background, and is classified as the patch feature;
[0028] Furthermore, the multi-modal adaptive loss function in Step 2 includes content loss, correlation loss, and class saliency loss, as shown in Formula (4):
[0029] L fea =L con +λL corr +ρL sil-l (4)
[0030] where L con is the content loss, L corr is the correlation loss, L sil-l is the class saliency loss, and λ and ρ are hyperparameters used to balance the weights of the three losses;
[0031] The content loss L con enhances the fusion of features, as shown in Formula (5):
[0032]
[0033] where w ir and w vi are adaptive weights, w vi =1 - w ir ;
[0034] The correlation loss L corr enhances the fusion of edge structural features, as shown in Formula (6):
[0035]
[0036] where cov(·) is the covariance function and σ is the standard deviation function.
[0037] The class saliency loss Lsal-l Enhance the fusion of plaque features, as shown in formula (7):
[0038]
[0039] Among them, Φ ir is the infrared feature, Φ vi is the visible light feature, Φ f is the feature after fusing the infrared and visible light features through the fusion network, M ir and M vi are the masks for removing noise in the features, as shown in formulas (8) and (9):
[0040]
[0041]
[0042] Among them, θ is a constant.
[0043] Furthermore, the fusion network structure described in step 3 is as follows: It includes 4 Transformer modules, and each Transformer module consists of 2 1×1 convolutional layers and 1 FocalTransformer module; the first convolutional layer adjusts the feature channels, the FocalTransformer module combines local and global information to fuse the features, and the second convolutional layer increases the nonlinear characteristics of the network.
[0044] Preferably, adding the saliency label described in step 4 includes the following steps:
[0045] Step 4.1, use the LC saliency extraction algorithm to detect the input infrared image to obtain the saliency map M sal ;
[0046] Step 4.2, perform normalization processing on the saliency map to obtain
[0047] Step 4.3, design the saliency loss for the normalized saliency map as shown in formula (10):
[0048]
[0049] Among them, F is the fused image, I ir and I vi are the infrared and visible light images respectively.
[0050] Furthermore, the cascading of the encoder, fusion module, and decoder to form a complete image fusion network described in step 5 is represented as follows:
[0051] Extract the multi-modal features Φ of the infrared image and the visible light image through the trained encoder E ir and Φ vi , concatenate the infrared feature Φ ir and the visible light feature Φ vi on the channel and then input them into the feature fusion network F. The fused feature Φ f generated by the feature fusion network F is decoded by the trained decoder D to generate the fused image I f . The entire fusion process can be formalized as formula (11):
[0052] I f = D(F(E(I ir ), E(I vi ))) (11)
[0053] where I ir and I vi represent the infrared image and the visible light image respectively; E(·) represents the encoder function, F(·) represents the feature fusion network function, and D(·) represents the decoder function.
[0054] Furthermore, the fusion model training process in step 5 is as follows: taking the infrared image, the visible light image and the saliency map of the infrared image as inputs, the total loss function is shown in formula (12):
[0055] L = L ssim + αL fea + βL sal (12)
[0056] where L ssim is the structural similarity loss, and its calculation formula is L ssim = 1 - ssim(I f , I vi ), L fea is the multi-modal adaptive loss, L sal is the saliency loss, and α and β are hyperparameters used to balance the weights of the three losses.
[0057] The infrared and visible light image fusion method based on multi-modal features of the present invention, compared with the prior art, brings the following technological innovations:
[0058] (1) A new infrared and visible light image fusion network is proposed, which uses the FocalTransformer module to fuse and extract image features, taking into account both local and global information, and has better fusion performance;
[0059] (2) For the multi-modal features extracted by the encoder, a multi-modal adaptive loss is designed to optimize and learn the model, which strengthens the information transmission in the fusion process and effectively avoids the problem of weakening different feature information in existing fusion methods;
[0060] (3) Significance information is added in the fusion to optimize the model to adaptively enhance the weights of thermal targets in infrared images and texture details in visible light images, and finally a fusion result with good visual effects is obtained. Description of the Drawings
[0061] Figure 1 It is the structural diagram of the multi-modal feature encoder-decoder network;
[0062] Figure 2 It is the structural diagram of the network during the training of the encoder-decoder;
[0063] Figure 3 It is the structure of the convolutional block ECB in the encoder;
[0064] Figure 4 It is the structure of the convolutional block DCB in the decoder;
[0065] Figure 5 It is the result diagram of the first embodiment. Among them, figure (a) is the infrared image to be fused in the first embodiment; figure (b) is the visible light image to be fused in the first embodiment; figure (c) is the fused image based on the Laplacian pyramid (LP); figure (d) is the fused image based on the discrete wavelet transform (DWT); figure (e) is the fused image based on the curvelet transform (CVT); figure (f) is the fused image of FusionGAN; figure (g) is the fused image of DenseFuse; figure (h) is the fused image of the method of the present invention.
[0066] Figure 6 It is the result diagram of the second embodiment. Among them, figure (a) is the infrared image to be fused in the second embodiment; figure (b) is the visible light image to be fused in the second embodiment; figure (c) is the fused image based on the Laplacian pyramid (LP); figure (d) is the fused image based on the discrete wavelet transform (DWT); figure (e) is the fused image based on the curvelet transform (CVT); figure (f) is the fused image of FusionGAN; figure (g) is the fused image of DenseFuse; figure (h) is the fused image of the method of the present invention.
[0067] The present invention will be further described in detail below with reference to the drawings and embodiments. Detailed Embodiment
[0068] This embodiment provides an infrared and visible light image fusion method based on multi-modal features, including the following steps:
[0069] Step 1, construct a feature extraction and image reconstruction network. Based on a multi-scale convolutional network, guided by the loss function, optimize to generate a multi-modal feature encoder-decoder network. The encoder-decoder network structure is as Figure 2 shown. The encoder contains 1 1×1 convolutional layer and 4 encoding convolutional modules ECB10, ECB20, ECB30, and ECB40. Each encoding convolutional module contains 2 3×3 convolutional layers and a max pooling layer. The ECB structure is as Figure 3 shown. The decoder contains 1 1×1 convolutional layer and 6 decoding convolutional modules DCB30, DCB20, DCB21, DCB10, DCB11, and DCB12. Each decoding convolutional module contains two 3×3 convolutional layers. The DCB structure is as Figure 4 shown. The specific connection method of the decoder is as follows: In the first and second scales, use horizontal dense skip connections. Adopt the channel connection method to skip-connect the final fused feature of the second scale to the input of DBC21, skip-connect the final fused feature of the first scale to the inputs of DCB11 and DCB12, and skip-connect the output of DCB10 to the input of DCB12. Through horizontal dense skip connections, the deep features of all intermediate layers are used for feature reconstruction, improving the reconstruction ability of multi-scale deep features; in the decoding sub-network, establish vertical dense connections in all scales. Adopt the upsampling method to connect the final fused feature of the fourth scale to the input of DCB30, the final fused feature of the third scale to the input of DCB20, the final fused feature of the second scale to the input of DCB10, connect the output of DCB30 to the input of DCB21, the output of DCB20 to the input of DCB11, and the output of DCB21 to the input of DCB12. Through vertical dense upsampling connections, all scale features are used for feature reconstruction, further improving the reconstruction ability of multi-scale deep features.
[0070] The loss function L of the encoder-decoder network ED , which is the pixel consistency and structural similarity between the input image and the output image, as shown in formula (1):
[0071] L ED = L p + βL ssin (1)
[0072] where L p is the pixel consistency loss, and L ssin is the structural similarity loss;
[0073] The pixel consistency loss L p is shown in formula (2):
[0074]
[0075] Structural similarity loss L ssin As shown in formula (3):
[0076] L ssim = 1 - ssim(O, I) (3)
[0077] Where O is the network output image and I is the input image.
[0078] During the encoder-decoder training stage, the four features Φ 1 , Φ 2 , Φ 3 and Φ 4 output by the encoder are directly input into the decoder. The network is trained using the MS-COCO dataset. Eighty thousand images are selected, converted to grayscale, and then resized to 256×256 as the network input. The loss function is L ED . After training, the network parameters are frozen.
[0079] Step 2: Extract infrared and visible light multi-modal features through the encoder-decoder network, measure the multi-modal features using entropy, gradient, and saliency, and design a multi-modal adaptive loss. The steps for measuring the multi-modal features are as follows:
[0080] Step 2.1: Calculate the entropy of the features output by the encoder, compare the entropy values of the features at each scale. The feature with the highest entropy contains the most content and details and is classified as the content feature.
[0081] Step 2.2: Calculate the gradient of the encoder input image using the Sobel gradient operator, downsample the gradient, subtract it from each feature, and calculate the mean. The feature with the smallest mean contains more structural features such as contours and edges and is classified as the edge structural feature.
[0082] Step 2.3: Calculate the saliency map of the encoder input image using the saliency extraction algorithm, downsample the saliency map, subtract it from each feature, and calculate the mean. The feature with the smallest mean can distinguish the foreground object from the background to a certain extent and is classified as the patch feature.
[0083] The multi-modal adaptive loss function includes content loss, correlation loss, and class saliency loss, as shown in formula (4):
[0084] L fea = L con + λL corr + ρL sil-l (4)
[0085] Where Lcon is the content loss, L corr is the relevance loss, L sil-l is the class saliency loss, λ and ρ are hyperparameters used to balance the weights of the three losses;
[0086] Content loss L con Enhance the fusion of features, as shown in Equation (5):
[0087]
[0088] where, w ir and w vi are adaptive weights, w vi = 1 - w ir ;
[0089] Relevance loss L corr Enhance the fusion of edge structural features, as shown in Equation (6):
[0090]
[0091] where, cov(·) is the covariance function and σ is the standard deviation function.
[0092] Class saliency loss L sal-l Enhance the fusion of patch features, as shown in Equation (7):
[0093]
[0094] where, Φ ir is the infrared feature, Φ vi is the visible light feature, Φ f is the feature after fusing the infrared and visible light features through a fusion network, M ir and M vi are Masks for removing noise in the features, as shown in Equations (8) and (9):
[0095]
[0096]
[0097] where, θ is a constant.
[0098] Step 3: Construct a fusion weight learning model that embeds the Transformer fusion strategy and assign values to the weights of the fusion model.
[0099] The fusion network structure is as follows: It includes 4 Transformer modules, and each Transformer module consists of 2 1×1 convolutional layers and 1 Focal Transformer module; the first convolutional layer adjusts the feature channels, the Focal Transformer module combines local and global information to fuse the features, and the second convolutional layer increases the non-linearity of the network.
[0100] Step 4, Obtain the saliency map of the infrared image as the label. Adding the saliency label as the region selection for fusing network optimization includes the following steps:
[0101] Step 4.1, Use the LC saliency extraction algorithm to detect the input infrared image to obtain the saliency map M sal ;
[0102] Step 4.2, Perform normalization processing on the saliency map to obtain
[0103] Step 4.3, Design the saliency loss for the normalized saliency map as shown in formula (10):
[0104]
[0105] where F is the fused image, I ir and I vi are the infrared and visible light images respectively.
[0106] Step 5, Cascade the fusion weight learning model with the embedded Transformer fusion strategy and the encoder-decoder to construct an infrared and visible light image fusion network, and use the saliency label and multi-modal loss to train the infrared and visible light image fusion network. Extract the multi-modal features Φ ir and Φ vi of the infrared image and the visible light image through the trained encoder E, concatenate the infrared feature Φ ir and the visible light feature Φ vi on the channel and input them into the feature fusion network F. The fused feature Φ f generated by the feature fusion network F is decoded by the trained decoder D to generate the fused image I f , and the entire fusion process can be formalized as formula (11):
[0107] I f = D(F(E(I ir ), E(I vi ))) (11)
[0108] where I irand I vi represent an infrared image and a visible light image respectively; E(·) represents an encoder function, F(·) represents a feature fusion network function, and D(·) represents a decoder function.
[0109] Furthermore, the fusion model training process in step 5 is as follows: taking the infrared image, the visible light image, and the saliency map of the infrared image as inputs, the total loss function is as shown in formula (12):
[0110] L = L ssim + αL fea + βL sal (12)
[0111] where L ssim is the structural similarity loss, and its calculation formula is L ssim = 1 - ssim(I f , I vi ), L fea is the multi-modal adaptive loss, L sal is the saliency loss, and α and β are hyperparameters used to balance the weights of the three losses.
[0112] In the fusion network training stage, the FLIR dataset is used to train the network. 12,000 images are selected, converted to grayscale, and then resized to 256×256 as the network input, and the loss function is L total .
[0113] A specific embodiment is as follows:
[0114] Two groups of embodiments selected from the TNO dataset. The first group of data contains objects such as human bodies, grass, houses, and doors and windows, and the intensity differences of the same objects in the infrared and visible light are relatively large. The second group of data contains objects such as vehicles, houses, and clouds, and the texture details in the infrared image are richer.
[0115] The first group of embodiments
[0116] As Figure 5 shown, the input images of the first group of embodiments contain objects such as human bodies, grass, houses, and doors and windows. The infrared image has a higher significant contrast, and the visible light image has richer texture details. By Figure 5 in figure (c) - Figure 5From the analysis of Figure (h) in , it can be seen that the background brightness of the LP method is unnatural. Serious artifacts appear in both the DWT and CVT methods, and the fusion effect is poor. The human target in the FusionGAN method is blurred, and the texture details of the grass are also lost. The DenseFuse method has a good fusion effect, but it is overall on the dim side. Compared with other methods, the infrared and visible light image fusion method based on multi-mode features in this embodiment can highlight the target information of the infrared image while retaining more texture and detail information of the visible light image, and without introducing artifacts. This is because the fusion method based on multi-mode features of the present invention has corresponding adaptive fusion strategies for various features, enabling more information of the infrared and visible light images to be retained in the fused image.
[0117] The second group of embodiments
[0118] As Figure 6 shown, the second group of input images includes objects such as vehicles, houses, clouds, etc., and the texture details in the infrared image are richer. Through the analysis of Figure 6 in (c)- Figure 6 in (h) in , it can be seen that all fusion algorithms can achieve good overall fusion of the infrared image and the visible light image. The image obtained by the LP method has low brightness and is overall rather dim. Artifacts appear in the fused images obtained by the DWT and CVT methods. The background of FusionGAN is more blurred compared to the source images, losing most of the texture information of the background. The fusion effect of DenseFuse is relatively good, but the brightness is higher than that of the source images. The method of the present invention can well fuse the feature information of the infrared image and the visible light image. Visible light textures such as car windows and infrared textures such as clouds and floor tiles in the fused image are completely retained, and the visual effect is good.
[0119] To further verify the feasibility and effectiveness of the present invention, 21 pairs of infrared and visible light images were selected for fusion testing in this paper, and quantitatively evaluated and compared with five other methods. Quantitative evaluation objectively evaluates the fusion performance through some statistical indicators. In this embodiment, 6 evaluation indicators widely used in the field of image fusion were selected, such as Standard Deviation (SD), Entropy (EN), Spatial Frequency (SF), Mutual Information (MI), Structural Similarity (SSIM), and Root Mean Squared Error (RMSE). SD reflects the contrast of the fused image, and a large SD value indicates good contrast. EN measures the amount of information in the fused image. The larger the EN value, the more information the fused image contains. SF measures the overall detail richness of the fused image. The larger the SF, the richer the texture contained in the fused image. MI measures the amount of information from the source images contained in the fused image. The larger the MI, the more information from the source images the fused image contains. SSIM represents the structural correlation between the fused image and the source images. The larger the SSIM, the more similar the fused image is to the source images and the smaller the distortion. RMSE measures the error between the fused image and the source images. The smaller the RMSE index, the better the fusion performance, which means the fused image is closer to the source images and the error is minimized during the fusion process.
[0120] Table 1 shows the objective evaluation indicators of the experimental results of the selected 21 pairs of infrared and visible light images under different fusion methods. The bold and underlined data respectively represent the optimal and sub-optimal values of the evaluation indicators. From the data in Table 1, it can be seen that the infrared and visible light image fusion method based on multi-mode features given in this embodiment obtains the best scores in terms of SD, EN, MI, and RMSE indicators, and is only second to the DWT and DenseFuse methods in terms of SF and SSIM respectively. These results indicate that the method of the present invention transfers the most information from the source images to the fused image during the fusion process and can better preserve the edges. The fused image has the highest contrast, contains the most information, retains more global structural and edge features of the source images, and also has a better visual effect.
[0121] Table 1: Objective Evaluation Indicators of Infrared and Visible Light Image Fusion Results
[0122]
Claims
1. An infrared and visible light image fusion method based on multi-mode features, characterized in that, It includes the following steps: Step 1: Construct a feature extraction and image reconstruction network. Based on a multi-scale convolutional network, guided by a loss function, optimize and generate a multi-modal feature encoder-decoder network; Step 2: Extract multi-modal features of infrared and visible light through the encoder-decoder network, measure the multi-modal features using entropy, gradient, and saliency, and design a multi-modal adaptive loss function; the multi-modal adaptive loss function includes content loss, correlation loss, and class saliency loss, as shown in formula (4): L fea = L con +λL corr +ρL sil-l (4) Among them, L con is the content loss, L corr is the relevance loss, L sil-l is the class saliency loss, λ and ρ are hyperparameters used to balance the weights of the three losses; Content loss L con Enhance the fusion of features, as shown in formula (5): (5) Among them, w ir and w vi are adaptive weights, , w vi = 1 - w ir ; Relevance loss L corr Enhance the fusion of edge structural features, as shown in formula (6): (6) Among them, is the covariance function, σ is the standard deviation function; Class significance loss L sal-l Enhance the fusion of plaque features, as shown in formula (7): (7) Among them, is the infrared feature, is the visible light feature, is the feature after fusing the infrared and visible light features through a fusion network, M ir and M vi are Masks for removing noise in the feature, as shown in Formulas (8) and (9): (8) (9) Among them, θ is a constant; Step 3: Construct a fusion weight learning model integrating a Transformer fusion strategy and assign weights to the fusion model; the structure of the fusion weight learning model is as follows: it includes 4 Transformer modules, and each Transformer module consists of 2 1×1 convolutional layers and 1 Focal Transformer module; the first convolutional layer adjusts the feature channels, the Focal Transformer module combines local and global information to fuse the features, and the second convolutional layer increases the non-linearity of the network; Step 4: Obtain the saliency map of the infrared image as the label, and add the saliency label as the region selection for optimizing the fusion network; Step 5: Cascade the fusion weight learning model integrating the Transformer fusion strategy with the encoder-decoder to construct an infrared and visible light image fusion network, and train the infrared and visible light image fusion network using the saliency label and multi-modal loss.
2. The method according to claim 1, characterized in that, The structure of the encoder in Step 1 includes: 1 1×1 convolutional layer and 4 encoding convolutional modules ECB10, ECB20, ECB30, and ECB40, and each encoding convolutional module includes 2 3×3 convolutional layers and 1 max-pooling layer; The structure of the decoder in Step 1 includes: 1 1×1 convolutional layer and 6 decoding convolutional modules DCB30, DCB20, DCB21, DCB10, DCB11, and DCB12, and each decoding convolutional module includes two 3×3 convolutional layers.
3. The method according to claim 1, characterized in that, The specific connection method of the decoder network in Step 1 is as follows: In the first and second scales, horizontal dense skip connections are adopted. In the channel connection method, the final fusion feature of the second scale is skip-connected to the input of DBC21, the final fusion feature of the first scale is skip-connected to the inputs of DCB11 and DCB12, and the output of DCB10 is skip-connected to the input of DCB12; Through horizontal dense jump connections, the depth features of all intermediate layers are used for feature reconstruction, improving the reconstruction ability of multi-scale depth features; in the decoding sub-network, vertical dense connections are established at all scales. In an upsampling manner, the final fusion feature of the fourth scale is connected to the input of DCB30, the final fusion feature of the third scale is connected to the input of DCB20, the final fusion feature of the second scale is connected to the input of DCB10, the output of DCB30 is connected to the input of DCB21, the output of DCB20 is connected to the input of DCB11, and the output of DCB21 is connected to the input of DCB12. Through vertical dense upsampling connections, features of all scales are used for feature reconstruction, further improving the reconstruction ability of multi-scale depth features.
4. The method according to claim 1, characterized in that: The loss function of the encoder-decoder network , which is the pixel consistency and structural similarity between the input image and the output image, as shown in Equation (1): L ED = L p + βL ssin (1) Among them L p is the pixel consistency loss, L ssin is the structural similarity loss; Pixel consistency loss L p As shown in formula (2): (2) Structural similarity loss L ssin As shown in formula (3): L ssim =1-ssim ( O, I )(3) Among them, O is the network output image, I is the input image.
5. The method according to claim 1, characterized in that, The use of entropy, gradient, and saliency to measure the multi-modal features described in step 2 includes the following steps: Step 2.1, calculate the entropy of the features output by the encoder, compare the entropy values of the features at each scale. The feature with the highest entropy contains the most content and details, and it is classified as the content feature. Step 2.2, use the Sobel gradient operator to calculate the gradient of the input image of the encoder, downsample the gradient and then calculate the difference with each feature, and find the mean value. The feature with the smallest mean value contains more structural features such as contours and edges, and it is classified as the edge structural feature. Step 2.3, use the saliency extraction algorithm to calculate the saliency map of the input image of the encoder, downsample the saliency map and then calculate the difference with each feature, and find the mean value. The feature with the smallest mean value has a certain distinction between the foreground object and the background, and it is classified as the patch feature.
6. The method according to claim 1, wherein The addition of the saliency label described in step 4 includes the following steps: Step 4.1: Use the LC saliency extraction algorithm to detect the input infrared image and obtain the saliency map M sal ; Step 4.2, perform normalization processing on the saliency map to obtain ; Step 4.3, for the normalized saliency map design a saliency loss as shown in formula (10): (10) Among them, F is the fused image, I ir and I vi are the infrared and visible light images respectively.
7. The method according to claim 1, wherein The cascading of the encoder, fusion module, and decoder to form a complete image fusion network described in step 5 is shown as follows: Extract multi-modal features of infrared images and visible light images through the trained encoder E and , concatenate the infrared features and visible light features on the channels and then input them into the feature fusion network F. The fusion features generated by the feature fusion network F are decoded by the trained decoder D to generate the fused image I f . The entire fusion process is formalized as Equation (11): (11) Among them, I ir and I vi represent an infrared image and a visible light image respectively; represents an encoder function, represents a feature fusion network function, represents a decoder function.
8. The method according to claim 1, wherein The fusion model training process in step 5 is: taking the infrared image, visible light image, and the saliency map of the infrared image as inputs, and the total loss function is shown in formula (12): L = L ssim + αL fea + βL sal (12) Among them, L ssim is the structural similarity loss, and the calculation formula is L ssim = 1 - ssim ( I f , I vi ), , L fea is a multi-mode adaptive loss, L sal is a saliency loss, α and β are hyperparameters used to balance the weights of the three losses.
Citation Information
Cited By
Lightweight infrared and visible light image fusion method, system, equipment and medium
CN116363034A