Infrared and visible light fusion method and system based on edge guidance and improved Transform, and storage medium
By introducing edge attention modules and embedded edge fusion modules in infrared and visible light images, the problem of existing methods ignoring edge and structural information is solved, and a clearer fusion image is achieved, suitable for application scenarios where clear boundaries and target outlines are required.
Patent Information
- Application Number
- CN202510158842.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-06
AI Technical Summary
The existing deep learning image fusion method ignores the presentation of edge and structural information in the fusion of infrared and visible light images, resulting in unclear boundaries and target contours of the fused images in some application scenarios.
A method based on edge guidance and improvement of Transformer was designed to enhance and fuse the edge features of the image through the edge attention module and the fusion module embedded in the edge, thereby generating a clearer fusion image.
This method significantly improves the details and information integrity of the fused image by enhancing the edge structure and context information expression of the image, especially in application scenarios where clear boundaries and target outlines are required.
Smart Images

Figure CN120107077A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and image processing technology, and in particular to an infrared and visible light fusion method, storage medium and system based on edge guidance and improved Transformer. Background Art
[0002] Image fusion is the process of integrating information from multiple image sensors into a single fused image, which aims to improve image quality and maintain the integrity of important features. In the field of image fusion, the fusion of infrared and visible light images is one of the most widely studied key areas. Infrared images perform well in dark environments and bad weather and can clearly display heat source targets, but have low resolution, high noise and lack of texture details. Visible light images have high resolution and rich texture details under good lighting conditions, but the effect is poor in the dark or bad weather.
[0003] Therefore, a deep learning-based method is proposed to study the fusion of infrared and visible light images. The infrared and visible light image fusion technology integrates the complementary information of the two spectra to generate a fused image that has both the thermal radiation information of the infrared image and the texture details of the visible light image, thereby achieving a more complete presentation of the image information.
[0004] However, in actual research, it is found that the fusion of infrared and visible light images often has different requirements for different application scenarios. For example, in scenarios such as security monitoring, military reconnaissance and autonomous driving, in addition to visual quality, the fused image also requires clear boundaries and target contours to achieve detection and recognition of significant targets. The current deep learning image fusion method mainly gives priority to information in the feature extraction process such as multi-scale features or the deepest semantic information, ignoring the presentation of some specific information such as edges and structures in the image.
[0005] In order to solve the above problems, the present invention proposes an attention-based edge guidance and improved Transformer method, and designs an edge attention module and an edge-embedded fusion module to realize a solution for the fusion of infrared and visible light images. Summary of the invention
[0006] The purpose of the present invention is to provide an infrared and visible light fusion method, storage medium and system based on edge guidance and improved Transformer to solve the problem of ignoring the presentation of some specific information such as edge and structure in the image.
[0007] To achieve the above objectives, the present invention provides the following technical solutions: an infrared and visible light fusion method based on edge guidance and improved Transformer,
[0008] The infrared and visible light fusion method based on edge guidance and improved Transformer is as follows:
[0009] S1, respectively obtaining edge images of the original infrared image and the original visible light image, taking the maximum value in the original infrared image and the original visible light image for fusion, and obtaining an edge fusion image; the edge fusion image is represented by Ef;
[0010] S2, the original infrared image and the original visible light image are subjected to initial convolution of the encoder to obtain shallow features; the edge fusion image Ef is input to the edge feature processing module 1 for channel expansion to match the shallow feature size to obtain edge fusion feature 1; the edge fusion feature 1 is represented by EF1;
[0011] The formula of the edge fusion feature 1 is: EF1 = Conv (Ef)
[0012] Among them, Conv is a convolution operation with a scale of 1;
[0013] S3, through the pre-built edge attention module, the edge fusion feature 1EF1 and the shallow feature are processed to obtain the edge aggregation shallow feature; the shallow feature is processed by LF i (i=ir,vi) indicates;
[0014] The edge attention module is:
[0015] Among them, Conv represents a convolution operation with a scale of 1, expanding the number of channels, and D(·) represents a downsampling operation. represents element-wise addition, Represents element-wise multiplication;
[0016] S4, processing the edge aggregation shallow features by pre-building an improved Transformer block to obtain a deep fusion feature containing global context information of the edge aggregation shallow features;
[0017] S5, inputting the edge fusion image into the edge feature processing module 2 to obtain the edge fusion feature 2, and inputting the deep fusion feature obtained after the processing in step S4 and the edge fusion feature 2 into the edge embedding fusion module through the edge embedding fusion module to perform feature enhancement and fusion to obtain the fusion feature;
[0018] The formula of the edge feature processing module 2 is: EF2 = Conv (Conv (Conv (Ef)))
[0019] Among them, Conv represents a convolutional layer with a scale of 3, and Ef represents an edge fusion image;
[0020] The edge embedded fusion module is expressed as:
[0021] SA(F1,F2)=Conv(Concat(mean(F1,F2),max(F1,F2)))
[0022] EIDF i =SA(SA(DF i , EF2), DF i )
[0023] Among them, SA(·) is the formula of the spatial attention mechanism, mean(·) and max(·) represent global average pooling and maximum pooling respectively, Concat represents the concatenation operation, and Conv(·) represents the convolution operation with a convolution kernel size of 3; EF2, DF i They represent edge fusion feature 2 and deep fusion feature respectively;
[0024] S6, decoding the fusion feature obtained in step S5 through a pre-built feature reconstruction module, and outputting a fusion image;
[0025] S7, optimizing the fused image by inputting the fused image obtained in step S6, the original infrared image, and the original visible light image through a pre-constructed joint loss function combining content loss and structural loss;
[0026] S8. Set the number of iterations for the processing of step S1 to step S7, and output the image after the iterative processing as an infrared and visible light fusion image.
[0027] Preferably, the step S2 specifically includes:
[0028] Using a scale feature extraction module to process the infrared edge image and the visible light edge image corresponding to the infrared edge image respectively;
[0029] The feature extraction module uses maximum fusion and 1 convolution block, and the number of channels of the convolution block is set to 64; the original resolution is maintained; the feature extraction module formula is as follows:
[0030] Ef=max(E ir ,E vi )
[0031] EF1=Conv(Ef)
[0032] Among them, E ir , E vi are respectively an infrared edge image and a visible light edge image corresponding to the infrared edge image, Ef is a fused edge image, and Conv is a convolution operation with a scale of 1.
[0033] Preferably, the step S3 specifically comprises:
[0034] S31, the edge fusion feature 1EF1 and the shallow feature LF i (i=ir,vi) performs residual connection to obtain input features and input them into the subsequent processing module;
[0035] First, the edge processing module 1 is used to process the edge fusion image Ef to obtain the edge fusion feature 1EF1, and then the LF i (i=ir,vi) is residually connected with EF1 to obtain The corresponding function expression is:
[0036] EF1=Conv(Ef)
[0037]
[0038] Among them, Conv represents a convolution operation with a scale of 1, expanding the number of channels, and D(·) represents a downsampling operation. represents element-wise addition, Represents element-wise multiplication;
[0039] S32, the subsequent processing module includes channel expansion convolution at both ends and global average pooling in the middle to obtain weight coefficients related to edge information;
[0040] S33, activating the weight coefficient obtained in step S32 and adding it to the input feature, outputting the feature representation, and obtaining the edge aggregation shallow feature EISF i (i=ir,vi);
[0041] The channel attention module includes global average pooling and convolution operations, and its expression is:
[0042]
[0043] in, represents one-dimensional convolution, GAP(·) represents the global average pooling operation, and Sigmoid(·) is the activation function.
[0044] Preferably, the step S4 specifically comprises:
[0045] Based on the Transformer block, the feature processing method is optimized, and a convolution-based local feature extraction channel is introduced. The function expression of the optimized Transformer block is:
[0046] DF i =T(EISF i ), i=ir,vi
[0047] Among them, T(·) represents the Transformer operation, DF i Represents the deep features after the Transformer block.
[0048] Preferably, the step S7 specifically includes:
[0049] The content loss function and the structure loss function are combined to obtain a joint loss function for training. The expression of the joint loss function is:
[0050] L J =L c +αL s
[0051] Among them, L J , L C and L S They represent the joint loss function, the content loss function and the structural loss function respectively; α represents a hyperparameter used to control L C and L S The ratio between them; L C is the combined pixel intensity loss L P , gradient loss L D The corresponding function expression is as follows:
[0052] L C =L P +βL D
[0053] Among them, β represents a hyperparameter used to control L P and L D The ratio between them; L P and L D They are used to measure the difference between the generated image and the real image at the pixel level and to maintain the details and texture information in the image. The function expressions are as follows:
[0054]
[0055] Where H represents the height of the image, W represents the width of the image, and max(·) represents the operation of calculating the maximum value of an element;
[0056] L S It is used to increase the structural information in the fused image. The function expression is as follows:
[0057] L S =MSE_V+MSE_I
[0058] Among them, MSE_V and MSE_I represent the mean square error similarity loss of visible light image and infrared image, and their function expressions are as follows:
[0059] MSE_I=5*loss_ssim(I ir ,I fused )+MSELoss(I ir ,I fused )
[0060] MSE_V=5*loss_ssim(I vi ,I fused )+MSELoss(I vi ,I fused )
[0061] Among them, I ir and I vi Represents the original infrared and visible light images, I fused To output the fused image, loss_ssim(·) represents the structural similarity loss, and MSELoss(·) represents the auxiliary loss. The function expression is as follows:
[0062]
[0063] A storage medium for an infrared and visible light fusion method based on edge guidance and improved Transformer, the storage medium having a system of the infrared and visible light fusion method based on edge guidance and improved Transformer, the storage medium having a processor and a memory communicatively connected to the processor; the number of the processor and the number of the memory are both not less than 1;
[0064] The memory stores instructions that can be executed by the processor, and the instructions are used for the processor to execute the edge guidance and infrared and visible light fusion method of the improved Transformer.
[0065] A system for edge-guided and improved Transformer infrared and visible light fusion method, wherein the system is used for edge-guided and improved Transformer infrared and visible light fusion.
[0066] Compared with the prior art, the present invention has the following beneficial effects:
[0067] ① The present invention designs an edge attention module of a channel attention mechanism, which can connect edge information in the channel dimension and redistribute weights, thereby strengthening the edge information in shallow features.
[0068] ② The present invention designs an edge embedding fusion module based on the spatial attention mechanism, which combines the edge feature map in the spatial dimension to achieve adaptive fusion of edge information, aiming to capture more high-frequency feature information, which can more fully express the features and enhance the details and information integrity of the fused image.
[0069] ③ Combine the improved Transformer block and edge extraction network, and use its powerful global context information capture ability to enhance the expression of edge structure and context information of the fused image.
[0070] ④Compared with existing algorithms, this method performs well in multiple evaluation indicators, can effectively capture salient targets, and present clear texture details. The ablation experiment results also further verify the effectiveness of the proposed module. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 It is a schematic diagram of the overall structure of the network framework in an embodiment of the present invention;
[0072] Figure 2 Schematic diagram of an edge attention module in an embodiment of the present invention;
[0073] Figure 3 It is a schematic diagram of an edge embedding and fusion module in an embodiment of the present invention. DETAILED DESCRIPTION
[0074] The technical solutions in the embodiments of the present invention are described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention. The experimental methods described in the following embodiments are all conventional methods unless otherwise specified.
[0075] Based on the problems mentioned in the background technology, in the infrared and visible light image fusion solution, this paper proposes an edge attention module and designs a new attention combination mechanism to strengthen the key information in the shallow feature map. On this basis, an edge embedding fusion module is also proposed to integrate the complementary information and details of the edge feature map and the deep feature, so as to more fully express the features.
[0076] The present invention is based on the network of autoencoders and combines the improved Transformer block to strengthen the semantic and contextual information of the fused image and further improve the understanding of the fused image at the semantic level. It aims to preserve the key feature information and global dependencies in the image.
[0077] The technical solution of the present invention is:
[0078] Step 1: The infrared image and the visible light image are first subjected to feature extraction by the convolution block to generate a local feature map. At the same time, the input infrared edge image and the visible light edge image are respectively obtained by the edge extraction module RCF, and then fused to obtain an edge fusion feature map.
[0079] Step 2: The extracted local feature map and edge fusion feature map enter the designed edge attention module. The convolution layer and channel attention module in the edge attention module can effectively learn the edge feature distribution;
[0080] Step 3: The fused features are passed through the improved Transformer block to obtain global context information and obtain deep features;
[0081] Step 4: After that, the edge fusion features are processed and entered into the designed embedded edge fusion module with the deep features to enhance the edge information strength to generate fusion features;
[0082] Step 5: The fused features are passed through a feature reconstruction module to generate a fused image;
[0083] Step 6: Iterate according to steps 1 to 5 until stopping, and output the final fused image.
[0084] In the present invention, the iterative training of the fusion network is not described one by one. The existing data set can be combined to introduce a neural network for effective training, and the trained network can be output for use.
[0085] The specific contents of the above steps are as follows:
[0086] The overall network structure is as follows Figure 1 As shown in the figure, it mainly consists of three parts: edge extraction network (Richer Convolutional Features, RCF), edge attention module (Edge Attention Module, EAM), and edge embedded fusion module (Edge-embedded Fusion Module, EFM).
[0087] The key to the method is to introduce edge information to enhance the structural features in the fused image and further highlight the salient targets.
[0088] In the present invention, feature enhancement consists of an edge attention module, an improved Transformer block, and an embedded edge fusion module, which work together to generate the final fused image. The edge attention module is used to enhance the edge information of shallow features, the improved Transformer block is used to capture long-distance dependencies, and the embedded edge fusion module is responsible for further optimizing and integrating features.
[0089] The shallow feature extraction module in step 1 mainly uses a convolution-based approach to extract features from the source image. Figure 1 As shown, shallow feature extraction consists of one convolution, and the size of the feature map is also shown in the figure.
[0090] Compared with traditional feature extraction, the present invention designs an edge attention module to enhance the ability to extract key features in images. The module is based on the attention mechanism and adds channel attention after splicing channel information.
[0091] like Figure 2 As shown in Figure 1, the edge attention module is mainly composed of residual connection and channel attention mechanism. Among them, Conv3×3 represents the convolution with a convolution kernel size of 3, and Conv1d represents the convolution with a convolution kernel size of 1.
[0092] First, the edge fusion image Ef is processed again using the edge processing module 2, and the edge fusion feature EF1 is obtained after passing through the convolution block with a scale of 1. i (i=ir,vi) is residually connected with EF1 and input into the channel attention module to obtain EISF i (i=ir,vi), the corresponding function expression is:
[0093] EF 1 =Conv(Ef)
[0094]
[0095] Among them, Conv represents a convolution operation with a scale of 1, which expands the number of channels. D(·) represents a downsampling operation
[0096]
[0097] in, represents one-dimensional convolution, GAP(·) represents the global average pooling operation, and Sigmoid(·) is the activation function.
[0098] In order to effectively combine details and global information, the present invention designs an improved Transformer block.
[0099] In the present invention, the improved Transformer optimizes the feature processing method based on the traditional Transformer and introduces a local feature extraction channel based on convolution to achieve more comprehensive feature extraction. This design effectively overcomes the redundancy problem existing in the calculation of the self-attention mechanism of the traditional Transformer, so that the model can capture the global information in the image more easily while maintaining performance. The Transformer block is used to capture the global context information of the fused image. Its specific operations are:
[0100] DF i =T(EISF i ), i=ir,vi
[0101] Among them, T(·) represents the Transformer operation, DF i Represents the deep features after the Transformer block.
[0102] The convolution and downsampling operations in the feature extraction process will cause the loss of image texture details to a certain extent. Therefore, an edge embedding fusion module is designed to introduce edge priors to jointly guide the decoder to focus on rich details.
[0103] First, the edge fusion image Ef is processed again using the edge processing module 2. After a convolution with a scale of 1, the edge fusion feature 2EF2 is obtained. The edge fusion feature 2EF2 is combined with the depth feature DF after the Transformer block. i (i=ir,vi) Input embedding edge fusion feature to obtain edge aggregation deep feature EIDF i (i=ir,vi), its expression is as follows:
[0104] SA(F1,F2)=Conv(Concat(mean(F1,F2),max(F1,F2)))
[0105] EIDF i =SA(SA(DF i ,EF2),DF i )
[0106] Where SA(·) is the formula of the spatial attention mechanism, mean(·) and max(·) represent global average pooling and maximum pooling respectively, Concat represents the concatenation operation, and Conv(·) represents the convolution operation with a convolution kernel size of 3; EF 2 ,DF i They represent edge fusion features, depth fusion features of infrared images, or depth fusion features of visible light images respectively.
[0107] Finally, the feature reconstruction module is responsible for generating the fused image.
[0108] In order to improve the quality and semantic richness of the fused image, the present invention adopts a joint loss function consisting of a content loss function and a structure loss function to train the network. The expression of the joint loss function is:
[0109] L J =L c +αL s
[0110] Among them, L J , L C and L S They represent the joint loss function, content loss function and structural loss function respectively, and α represents a hyperparameter used to control L C and L S The ratio between them; L C is the combined pixel intensity loss L P , gradient loss L D The corresponding function expression is as follows:
[0111] L C =L P +βL D
[0112] Among them, β represents a hyperparameter used to control L P and L D The ratio between them; L P and L D They are used to measure the difference between the generated image and the real image at the pixel level and to maintain the details and texture information in the image. The function expressions are as follows:
[0113]
[0114] Where H represents the height of the image, W represents the width of the image, and max(·) represents the operation of calculating the maximum value of an element;
[0115] L S It is used to increase the structural information in the fused image. The function expression is as follows:
[0116] L S =MSE_V+MSE_I
[0117] Among them, MSE_V and MSE_I represent the mean square error similarity loss of visible light image and infrared image, and their function expressions are as follows:
[0118] MSE_I=5*loss_ssim(I ir ,I fused )+MSELoss(I ir ,Ifused )
[0119] MSE_V=5*loss_ssim(I vi ,I fused )+MSELoss(I vi ,I fused )
[0120] Among them, I ir and I vi Represents the original infrared and visible light images, I fused is the output fused image. , loss_ssim(·) represents the structural similarity loss, MSELoss(·) represents the auxiliary loss, and its function expression is as follows:
[0121] loss_ssim(I ir ,I fused )=1-SSIM(I ir ,I fused )
[0122]
[0123] The present invention also includes a storage medium for the infrared and visible light fusion method based on edge guidance and improved Transformer, the storage medium has a system of the infrared and visible light fusion method based on edge guidance and improved Transformer, the storage medium has a processor and a memory communicatively connected to the processor; the number of the processor and the memory is not less than 1;
[0124] The memory stores instructions that can be executed by the processor, and the instructions are used for the processor to execute the edge guidance and infrared and visible light fusion method of the improved Transformer.
[0125] The present invention also includes a system for edge-guided and improved Transformer-based infrared and visible light fusion method, wherein the system is used for edge-guided and improved Transformer-based infrared and visible light fusion.
[0126] In order to better explain the scheme and practical effect of the present invention, the present invention illustrates the implementation of the present invention through specific implementation examples. The training and testing of the present invention are both carried out on public data sets to ensure the reliability and fairness of the experimental results. Specifically, the experiment used two data sets: the MSRS data set and the LLVIP data set. The MSRS data set is specially designed for spectral road scenes and contains 1444 pairs of infrared and visible light images, covering various types of target objects such as cars, pedestrians, bicycles, etc. The data set has been pre-divided into a training set and a test set, where the training set contains 1083 pairs of images and the test set contains 361 pairs of images. In order to verify the generalization ability of the model, the experiment was tested on the LLVIP data set. In the LLVIP data set, a total of 44 pairs of images are used for testing. This data set covers multiple scene types in low-light environments, which can fully reflect the performance of the model in processing various complex scene images.
[0127] The training process uses the Adam optimizer to train the fusion framework, the initial learning rate is set to 0.001, the batch size is 4, and 50 full iterations are performed to update the framework parameters. The experiment is based on the Pytorch framework, and the computing environment is configured with an operating environment of Ubuntu 22.04, a GTX 4090 GPU, Intel (R) Xeon (R) Silver 4310, and 30GRAM. The parameters of the fusion graph on the MSRS dataset are shown in Table 1.
[0128] Table 1 Mean values of various indicators of different methods on the MSRS dataset
[0129]
[0130]
[0131] 361 pairs of images on the MSRS test set were quantitatively analyzed. Table 1 shows the average values of various evaluation indicators, and the best results are highlighted in bold. According to the data in Table 1, the method performs best in EN, SD, VIF and MS_SSIM indicators. This shows that the algorithm has good performance in preserving image details, contrast and information content. EN can reflect the information content and similarity of the fused image, while SD and MS_SSIM can reflect the clarity and contrast of the fused image. VIF is a comprehensive evaluation indicator. The data shows that the method proposed in the present invention can generate a fused image with highly similar information content to the original image.
[0132] In order to further illustrate the fusion performance of this method, this embodiment conducts a generalization experiment on 44 pairs of images in the LLVIP dataset. The parameters of the fusion graph on the LLVIP dataset are shown in Table 2.
[0133] Table 2 Mean values of various indicators of different methods on the LLVIP dataset
[0134]
[0135] As shown in Table 2, this implementation achieved the best performance in the EN, VIF and MS_SSIM indicators, indicating that its fused image is superior to other methods in terms of information content, contrast and structural similarity. The reason why the indicators are poor on the LLVIP dataset was analyzed. It was found that the 44 pairs of test images of the LLVIP dataset were taken in a low-light environment. The distinction between the salient targets and the background in the image was not obvious, and the edge information was weak. After analysis, it was found that the extracted edge information was limited, so the Qabf indicator for measuring the total edge information and the SF indicator for measuring the contrast were poor. The advantages of the present invention in the performance indicators EN, VIF and MS_SSIM are more obvious, indicating that this method has a strong advantage in overall fusion quality. In general, the present invention performs well in fusion performance, verifying its good fusion ability and generalization ability.
[0136] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for fusion of infrared and visible light based on edge guidance and improved Transformer, characterized by: The infrared and visible light fusion method based on edge guidance and improved Transformer is as follows: S1, respectively obtaining edge images of the original infrared image and the original visible light image, taking the maximum value in the original infrared image and the original visible light image for fusion, and obtaining an edge fusion image; the edge fusion image is represented by Ef; S2, the original infrared image and the original visible light image are subjected to initial convolution of the encoder to obtain shallow features; The edge fusion image Ef is input to the edge feature processing module 1 for channel expansion to match the shallow feature size, thereby obtaining an edge fusion feature 1; the edge fusion feature 1 is represented by EF1; The formula of the edge fusion feature 1 is: EF1 = Conv (Ef) Among them, Conv is a convolution operation with a scale of 1; S3, through the pre-built edge attention module, the edge fusion feature 1EF1 and the shallow feature are processed to obtain the edge aggregation shallow feature; the shallow feature is processed by LF i (i=ir,vi) indicates; The edge attention module is: Among them, Conv represents a convolution operation with a scale of 1, expanding the number of channels, and D(·) represents a downsampling operation. represents element-wise addition, Represents element-wise multiplication; S4, processing the edge aggregation shallow features by pre-building an improved Transformer block to obtain a deep fusion feature containing global context information of the edge aggregation shallow features; S5, inputting the edge fusion image into the edge feature processing module 2 to obtain the edge fusion feature 2, and inputting the deep fusion feature obtained after the processing in step S4 and the edge fusion feature 2 into the edge embedding fusion module through the edge embedding fusion module to perform feature enhancement and fusion to obtain the fusion feature; The formula of the edge feature processing module 2 is: EF2 = Conv (Conv (Conv (Ef))) Among them, Conv represents a convolutional layer with a scale of 3, and Ef represents an edge fusion image; The edge embedded fusion module is expressed as: SA(F1,F2)=Conv(Concat(mean(F1,F2),max(F1,F2))) EIDF i =SA(SA(DF i ,EF2),DF i ) Among them, SA(·) is the formula of the spatial attention mechanism, mean(·) and max(·) represent global average pooling and maximum pooling respectively, Concat represents the splicing operation, and Conv(·) represents the convolution operation with a convolution kernel size of 3; EF2, DF i They represent edge fusion feature 2 and deep fusion feature respectively; S6, decoding the fusion feature obtained in step S5 through a pre-built feature reconstruction module, and outputting a fusion image; S7, optimizing the fused image by inputting the fused image obtained in step S6, the original infrared image, and the original visible light image through a pre-constructed joint loss function combining content loss and structural loss; S8. Set the number of iterations for the processing of step S1 to step S7, and output the image after the iterative processing as an infrared and visible light fusion image.
2. According to claim 1, the infrared and visible light fusion method based on edge guidance and improved Transformer is characterized by: The step S2 specifically comprises: Using a scale feature extraction module to process the infrared edge image and the visible light edge image corresponding to the infrared edge image respectively; The feature extraction module uses maximum fusion and 1 convolution block, and the number of channels of the convolution block is set to 64; the original resolution is maintained; the feature extraction module formula is as follows: Ef=max(E ir ,AND vi ) EF1=Conv(Ef) Among them, E ir , E vi are respectively an infrared edge image and a visible light edge image corresponding to the infrared edge image, Ef is a fused edge image, and Conv is a convolution operation with a scale of 1.
3. According to claim 1, the infrared and visible light fusion method based on edge guidance and improved Transformer is characterized by: The step S3 specifically comprises: S31, the edge fusion feature 1EF1 and the shallow feature LF i (i=ir,vi) performs residual connection to obtain input features and input them into the subsequent processing module; First, the edge processing module 1 is used to process the edge fusion image Ef to obtain the edge fusion feature 1EF1, and then the LF i (i=ir,vi) is residually connected with EF1 to obtain The corresponding function expression is: EF1=Conv(Ef) Among them, Conv represents a convolution operation with a scale of 1, expanding the number of channels, and D(·) represents a downsampling operation. represents element-wise addition, Represents element-wise multiplication; S32, the subsequent processing module includes channel expansion convolution at both ends and global average pooling in the middle to obtain weight coefficients related to edge information; S33, activating the weight coefficient obtained in step S32 and adding it to the input feature, outputting the feature representation, and obtaining the edge aggregation shallow feature EISF i (i=ir,vi); The channel attention module includes global average pooling and convolution operations, and its expression is: in, represents one-dimensional convolution, GAP(·) represents the global average pooling operation, and Sigmoid(·) is the activation function.
4. According to claim 1, the infrared and visible light fusion method based on edge guidance and improved Transformer is characterized by: The step S4 specifically comprises: Based on the Transformer block, the feature processing method is optimized, and a convolution-based local feature extraction channel is introduced. The function expression of the optimized Transformer block is: DF i =T(EISF i ),i=ir,vi Among them, T(·) represents the Transformer operation, DF i Represents the deep features after the Transformer block.
5. According to claim 1, the infrared and visible light fusion method based on edge guidance and improved Transformer is characterized by: The step S7 specifically comprises: The content loss function and the structure loss function are combined to obtain a joint loss function for training. The expression of the joint loss function is: L J =L c +αL s Among them, L J , L C and L S They represent the joint loss function, the content loss function and the structural loss function respectively; α represents a hyperparameter used to control L C and L S The ratio between them; L C is the combined pixel intensity loss L P , gradient loss L D The corresponding function expression is as follows: L C =L P +βL D Among them, β represents a hyperparameter used to control L P and L D The ratio between them; L P and L D They are used to measure the difference between the generated image and the real image at the pixel level and to maintain the details and texture information in the image. The function expressions are as follows: Where H represents the height of the image, W represents the width of the image, and max(·) represents the operation of calculating the maximum value of an element; L S It is used to increase the structural information in the fused image. The function expression is as follows: L S =MSE_V+MSE_I Among them, MSE_V and MSE_I represent the mean square error similarity loss of visible light image and infrared image, and their function expressions are as follows: MSE_I=5*loss_ssim(I ir ,I fused )+MSELoss(I ir ,I fused ) MSE_V=5*loss_ssim(I vi ,I fused )+MSELoss(I vi ,I fused ) Among them, I ir and I vi Represents the original infrared and visible light images, I fused To output the fused image, loss_ssim(·) represents the structural similarity loss, and MSELoss(·) represents the auxiliary loss. The function expression is as follows: loss_ssim(I ir ,I fused )=1-SSIM(I ir ,I fused ) 6. A storage medium for the infrared and visible light fusion method based on edge guidance and improved Transformer as claimed in any one of claims 1 to 5, characterized in that: The storage medium has the system of the edge guidance and improved Transformer infrared and visible light fusion method, the storage medium has a processor and a memory connected to the processor in communication; the number of the processor and the memory is not less than 1; The memory stores instructions that can be executed by the processor, and the instructions are used for the processor to execute the edge guidance and infrared and visible light fusion method of the improved Transformer.
7. A system for infrared and visible light fusion method based on edge guidance and improved Transformer, characterized by: The system is used for edge guidance and improved Transformer's infrared and visible light fusion.