Infrared and visible light image fusion method based on multi-scale channel cross Transform, electronic equipment and readable storage medium
Multi-scale features of infrared and visible light images are extracted through multi-scale channel cross-transformer and dual-branch multi-scale encoder and effectively fused, solving the problem of poor connection with the decoder after fusion in the prior art, and achieving high-quality infrared and visible light images fusion effect.
Patent Information
- Application Number
- CN202510110308.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-13
AI Technical Summary
The existing infrared image and visible image fusion method have the problem of effective connection with the decoder after the feature fusion, resulting in poor fusion performance or damage to the effect.
Multi-scale channel crossover Transformer is used to extract multi-scale features of infrared and visible light images through a dual-branch multi-scale encoder, and fuse them through a feature fusion module. Using multi-scale channel cross-transformer and layer-by-layer decoding + upsampling, effective aggregation of fusion layer features and decoder features is achieved.
It significantly improves the fusion effect between infrared images and visible light images, and the generated fusion image targets are clear and the texture details are clear, which enhances the understanding of the scene, facilitates accurate identification of the target, and is suitable for high-level visual systems to work around the clock.
Smart Images

Figure CN120147794A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to an infrared and visible light image fusion method, an electronic device and a readable storage medium based on a multi-scale channel cross Transformer. Background Art
[0002] Image fusion refers to the process of integrating complementary or redundant image information obtained from two or more image sensors to obtain a fused image with high clarity and rich information, thereby providing support for subsequent image target positioning, recognition, detection, etc.; an infrared sensor forms an image by using the infrared radiation emitted by an object and can reflect hidden targets under poor lighting conditions, but the infrared image cannot show detailed information and has low clarity; a visible light sensor forms an image by using the reflected light of an object, and the visible light image has rich detailed information, but a clear image cannot be obtained under poor lighting conditions; fusing the infrared image and the visible light image can ensure that the obtained fused image has clear and prominent targets, clear texture details, can also enhance the understanding of the scene, facilitate accurate target recognition, and is conducive to the all-weather operation of a high-level vision system.
[0003] Existing infrared image and visible light image fusion methods can be divided into four main categories: models based on Generative Adversarial Networks (GANs), models based on AutoEncoders (AEs), unified models, and algorithm unfolding models; among them, the AE model is usually the most widely used network architecture in the field of infrared and visible light fusion, and a variety of AE-based image fusion technologies have been developed to fuse infrared images and visible light images, which can generate a comprehensive image that contains both rich radiation information and details and textures; existing AE-based feature fusion network models usually include four parts, namely feature encoding, feature fusion, skip connection, and feature decoding. After capturing low-level and high-level features using a dual-branch multi-scale encoder, they are sent to a feature fusion module for multi-scale feature fusion, and then the decoder uses the skip connection structure to construct the final result.
[0004] Traditional skip connections can help propagate the spatial information lost during the pooling operation and help restore the full spatial resolution after the encoding-decoding process to achieve a perfect infrared-visible light image fusion effect; however, after feature fusion, which fusion layer features are connected to the decoder; and how to effectively aggregate with the decoder features instead of simply splicing; but existing research has found that simple skip connections are not always beneficial for fusion in the task of infrared-visible light image fusion and may even damage the fusion performance, depending on the semantic information gap at different levels of the encoder and whether the semantic information at the same level can be well compatible; based on the above analysis, it is necessary to develop a more advanced image fusion method to overcome the limitations of existing AE technologies. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides an infrared and visible light image fusion method, an electronic device, and a readable storage medium based on a multi-scale channel cross Transformer.
[0006] Based on the above object, the present invention is realized through the following technical solutions:
[0007] In the first aspect of the present invention, an infrared and visible light image fusion method based on a multi-scale channel cross Transformer is provided, including the following steps:
[0008] S1. Extraction of multi-scale features: Use a dual-branch multi-scale encoder to extract five groups of infrared and visible light image features with different scales from infrared images and visible light images respectively.
[0009] S2. Fusion of multi-scale features: Input the five groups of features with different scales extracted from infrared images and visible light images into the feature fusion module respectively.
[0010] S3. Pass through the sub-module CCT (Channel-wise Crossfusion Transformer) of the multi-scale channel cross Transformer to generate four groups of output features O 1 、O 2 、O 3 、O 4 And restore; then gradually input the output O i And its corresponding decoder feature D i As the input of the sub-module CCA (Channel-wise Cross Attention) module of the cross Transformer, and obtain the output And the decoder feature D of this level i Perform channel splicing and decoding + upsampling to finally obtain the decoder feature of the previous level; use the multi-scale channel cross Transformer and layer-by-layer decoding + upsampling to obtain the fused image.
[0011] S4. Use the input visible light image, infrared image, and the fused image obtained by fusion to construct an intensity-gradient-similarity loss function, and use the loss function to guide the training of the feature fusion network model. When the loss function converges, a trained feature fusion network model can be obtained, and the feature fusion network model is used for the fusion of infrared images and visible light images.
[0012] According to the above infrared and visible light image fusion method based on a multi-scale channel cross Transformer, preferably, in step S1, the steps of using a dual-branch multi-scale encoder to extract multi-scale features of infrared images and visible light images are:
[0013] S11. Extraction of shallow features: The infrared image and the visible light image are respectively passed through a 3×3 convolutional layer to pre-extract the shallow features of the infrared image and the visible light image, and a convolutional module is used to perform preliminary processing to transform each pre-extracted shallow feature from to
[0014] S12. Extraction of multi-scale features: After the extraction of shallow features, a series of encoding modules and their corresponding downsampling operations are used to extract deep features of different scales; the dual-branch multi-scale encoder includes two parallel independent encoder branches, one of which is used to process the features of the infrared image and the other is used to process the features of the visible light image.
[0015] According to the above infrared and visible light image fusion method based on multi-scale channel cross Transformer, preferably, in step S12, each of the independent encoder branches includes five encoding modules with the same structure. The encoding module is used to perform feature encoding processing on the data of its respective modality to extract multi-scale features from the source image; starting from the high-resolution input, the dual-branch multi-scale encoder gradually reduces the spatial dimension through downsampling, that is, reducing the spatial resolution to one-half, one-fourth, and one-eighth of the original size, and the last layer is still one-eighth; at the same time, increasing the channel capacity, and finally obtaining a feature map of size ; ensuring that the feature information at different scales can be gradually encoded and extracting multi-level feature information therefrom; using the markers and to distinguish the multi-scale information of the infrared image and the visible light image; where i ∈ {1, 2, 3, 4, 5}.
[0016] According to the above infrared and visible light image fusion method based on multi-scale channel cross Transformer, preferably, in step S12, the operation steps of the encoding module are as follows: The encoding module adopts a dual-branch structure, and the Restormer module is used for each level of feature extraction. The Restormer module includes an attention mechanism and a feed-forward mechanism in a series connection relationship; first, through the attention mechanism, given the input aggregates the cross-channel context in the pixel direction through a 1×1 convolution, and then applies a 3×3 depth convolution to encode the spatial context in the channel direction to generate query (Q), key (K), and value (V) projections, that is, and convolution, depth convolution; reshape the query, key, and value projections so that their dot product interactions generate a size of The transposed attention map, rather than the large-scale conventional attention map of size .
[0017] The attention mechanism is as follows:
[0018]
[0019] where X and are the input and output feature maps respectively; the matrices and are both reshaped from the query (Q), key (K), and value (V) of the original size .
[0020] After passing through the feed-forward mechanism, given an input
[0021]
[0022] where ⊙ represents element-wise multiplication; φ represents the GELU non-linearity, and LN represents layer normalization.
[0023] According to the above infrared and visible light image fusion method based on multi-scale channel cross Transformer, preferably, in step S2, the multi-scale feature fusion includes semi-coupled extraction and convolutional fusion, specifically:
[0024]
[0025] where and represent the shared convolutional kernel and the private convolutional kernel corresponding to the vi branch and the ir branch respectively; represents concatenation in the channel dimension; * represents convolution.
[0026] After the above feature extraction, the semi-coupled features of the infrared image and the visible light image are fused using a convolutional operator to obtain the fused feature where i ∈ {1, 2, 3, 4, 5}, and the superscript f represents fusion.
[0027] According to the above infrared and visible light image fusion method based on multi-scale channel cross Transformer, preferably, in step S3, the multi-scale channel cross Transformer includes a sub-module CCT for multi-scale encoder feature fusion, a sub-module CCA for decoder feature D i and a sub-module for sub-module CCT feature fusion; the steps for the multi-scale channel cross Transformer to perform improved skip connections include:
[0028] S31. The fused feature After feature cross - fusion by sub - module CCT, multi - head attention O is generated i 。
[0029] S32. Restore the generated O i to Then, pass through sub - module CCA. Take the i - th level output and the i - th level decoder feature map as the input of sub - module CCA; Compress the space, which is performed by the global average pooling (GAP) layer, to generate a vector Among them, the global pooling operation for the k - th channel is Use this operation to embed global spatial information, and then generate an attention mask for recalibrating or exciting O i ;
[0030]
[0031] Among them, and are the weight matrices of the linear layer; δ(·) and σ(·) are both activation functions.
[0032] S33. Output is concatenated with the up - sampled feature D of the i - th level decoder i in the channel dimension and then decoded + up - sampled to obtain D i-1 ;
[0033]
[0034] The decoder structure is symmetric to the encoder structure. Using multi - scale channel - cross Transformer and layer - by - layer decoding + up - sampling, finally, the final output is completed through a 3×3 convolution operation with a stride of 1, which can effectively reconstruct a high - resolution fused image.
[0035] According to the above - mentioned infrared and visible light image fusion method based on multi - scale channel - cross Transformer, preferably, in step S31, the feature cross - fusion step of the sub - module CCT includes:
[0036] S311. Multi - scale feature embedding;
[0037] Given four fused features Among them, i ∈ {1, 2, 3, 4}. First, by using patches with a patch size of to serially process respectively to obtain Among them, i ∈ {1, 2, 3, 4}; During this process, keep the original channel size Then, connect the four - layer tokens together as where \(i = 1, 2, 3, 4\); \(d\) is the sequence length;
[0038] S312. Obtaining multi - head attention;
[0039] The tokens are fed into the multi - head channel cross - attention module, encoding channels and dependencies through a multi - layer perceptron (MLP) with a residual structure;
[0040]
[0041] where C i \((i = 1, 2, 3, 4)\) is the number of channels.
[0042] Obtain After projection, attention is obtained;
[0043]
[0044] where \(\psi(\cdot)\) and \(\sigma(\cdot)\) represent instance normalization and softmax functions respectively.
[0045]
[0046] where \(N\) represents the number of heads of attention.
[0047] The multi - head attention is;
[0048] O i = MCA i + MLP(Q i + MCA i ) (9).
[0049] According to the above - mentioned infrared and visible light image fusion method based on multi - scale channel cross - Transformer, preferably, in step S4, the loss function adopts an intensity - gradient - similarity joint loss, and the loss function is specifically:
[0050] S41. The intensity loss is:
[0051]
[0052] where \(M(\cdot)\) is the element - wise maximum operation; the intensity loss is used to help the feature fusion network model obtain appropriate intensity information.
[0053] S42. The gradient loss is:
[0054]
[0055] where denotes the Sobel gradient operator; ||·|| 1 , denotes the l 1 norm; |·| represents the absolute value operation; max(·) represents the element maximum selection; the gradient loss is used to guide the feature fusion network model to retain the maximum amount of gradient information.
[0056] S43. The similarity loss is as follows:
[0057]
[0058] wherein, ssim(·) represents the structural similarity operation; set a 1 = a 2 = 0.5; is used to constrain I f and I ir , I vi ; the structural similarity index SSIM reflects the defects of infrared images and visible light images from three aspects: brightness, contrast, and structure.
[0059] S44. The overall loss function of the feature fusion network model is a combination of all sub-item losses, and each sub-item is assigned a specific weight;
[0060] L total = λ 1 L ssim + λ 2 L text + λ 3 L int (13);
[0061] wherein, λ 1 , λ 2 and λ 3 are hyperparameters used to control the balance of each sub-item loss.
[0062] The second aspect of the present invention provides an electronic device, including at least one control processor and a memory for communicating with the at least one control processor; the memory stores instructions executable by the at least one control processor, and when the instructions are executed by the at least one control processor, the at least one control processor can execute any step in the infrared and visible light image fusion method based on the multi-scale channel cross Transformer as described in the first aspect.
[0063] The third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computer processor, any step in the infrared and visible light image fusion method based on the multi-scale channel cross Transformer as described in the first aspect is implemented.
[0064] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0065] 1. The present invention solves the problem of how to effectively connect with the decoder after feature fusion in the existing image fusion method through the multi-scale channel cross Transformer, and can significantly improve the fusion effect of infrared images and visible light images; the dual-branch multi-scale attention mechanism can extract multi-scale global and local features from the input source images, and features of various scales are processed by the feature fusion network model; the multi-scale channel cross Transformer can reasonably connect the features of the fusion layer to the decoder and can effectively aggregate with the decoder features, that is, bridge the semantic information gap at different scales of the encoder and make the semantic information at the same level compatible; the decoder structure is symmetric with the encoder structure, and the multi-scale channel cross Transformer and layer-by-layer decoding + upsampling are used to effectively reconstruct the high-resolution fusion image; it can ensure that the obtained fusion image has clear and prominent targets, clear texture details, can also enhance the understanding of the scene, facilitate accurate target recognition, and is conducive to the all-weather operation of the high-level vision system.
[0066] 2. By comparing this method with eight advanced methods, namely CSF, DenseFuse, FusionGAN, GANMcC, IFCNN, PMGI, U2Fusion, and UMF-CMGR, in the daytime scene, infrared thermal radiation data can serve as a supplement to visible light images. Therefore, an advanced fusion algorithm should enhance key targets without causing spectral contamination while maintaining the texture details of the Vi image. FusionGAN and GANMcC do not retain the texture information of the Vi image. Although CSF, DenseFuse, IFCNN, PMGI, U2Fusion, and UMF-CMGR integrate the texture of the Vi image with infrared information, the background area is affected by spectral noise pollution to varying degrees. This method performs excellently in retaining texture and highlighting target information. In the nighttime scene, the detail information in the Vi image is limited. In contrast, the infrared image can not only capture prominent targets but also capture detailed textures, thus enhancing the texture information of the Vi image. CSF, DenseFuse, IFCNN, U2Fusion, and UMF-CMGR do not effectively retain prominent target information. At the same time, FusionGAN and GANMcC lose background texture information, and PMGI suffers from overexposed backgrounds resulting in texture distortion. This method can effectively retain valuable information.
[0067] 3. The present invention can solve the problem of how to effectively connect with the decoder after feature fusion in existing image fusion methods to achieve high-quality fusion of infrared images and visible light images. This method first uses a dual-branch multi-scale attention mechanism to extract multi-scale global and local features from the input source images, and features of various scales are processed by the feature fusion network model. Subsequently, the multi-scale channel cross Transformer can reasonably connect the features of the fusion layer to the decoder and effectively aggregate them with the decoder features, that is, bridge the semantic information gap at different scales of the encoder and make the semantic information at the same level compatible. Finally, the decoder structure is symmetric with the encoder structure. Using the multi-scale channel cross Transformer and layer-by-layer decoding + upsampling, a high-resolution fusion image is effectively reconstructed, which can ensure that the obtained fusion image has clear and prominent targets, clear texture details, can also enhance the understanding of the scene, facilitate accurate target recognition, and is conducive to the all-weather operation of the high-level vision system. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 It is a schematic flowchart of the present invention;
[0069] Figure 2 It is a schematic diagram of the overall architecture of the present invention;
[0070] Figure 3 It is a schematic diagram for comparing the fusion effect in the daytime scene;
[0071] Figure 4 Schematic diagram for comparing night scenes of fusion effects. Detailed implementation manners
[0072] The present invention will be further described in detail below through specific embodiments, but the scope of the present invention is not limited.
[0073] Embodiment 1
[0074] An infrared and visible light image fusion method based on a multi-scale channel cross Transformer, the process of which is as Figure 1 shown, Figure 2 Schematic diagram of the overall architecture of the infrared image and visible light image fusion method, including a dual-branch multi-scale encoder, a feature fusion module, a multi-scale channel cross Transformer, and a decoder, including the following steps:
[0075] S1. Extraction of multi-scale features; using a dual-branch multi-scale encoder to extract five groups of infrared and visible light image features with different scales from the infrared image and the visible light image respectively.
[0076] The steps of using a dual-branch multi-scale encoder to extract multi-scale features of the infrared image and the visible light image are:
[0077] S11. Extraction of shallow features; passing the infrared image and the visible light image through a 3×3 convolutional layer respectively to pre-extract the shallow features of the infrared image and the visible light image, and using a convolutional module to perform preliminary processing to transform each pre-extracted shallow feature from to
[0078] S12. Extraction of multi-scale features; after the extraction of shallow features is completed, different scales of deep features are extracted through a series of encoding modules and their corresponding downsampling operations; the dual-branch multi-scale encoder includes two parallel independent encoder branches, one of which is used to process the features of the infrared image, and the other independent encoder branch is used to process the features of the visible light image.
[0079] Each of the independent encoder branches includes five encoding modules with the same structure. The encoding module is used to perform feature encoding processing on the data of its respective modality to extract multi-scale features from the source image; starting from the high-resolution input, the dual-branch multi-scale encoder gradually reduces the spatial dimension through downsampling, that is, reducing the spatial resolution to one-half, one-fourth, and one-eighth of the original size, and the last layer is still one-eighth; at the same time, increasing the channel capacity, and finally obtaining a feature map with a size of ; ensuring that the feature information at different scales is gradually encoded, and fully extracting multi-level feature information from it; using the label Fi ir and F i vi Distinguish the multi-scale information of infrared images and visible light images; where, i ∈ {1, 2, 3, 4, 5}.
[0080] The operation steps of the encoding module are as follows: The encoding module adopts a dual-branch structure, and each level of feature extraction uses the Restormer module. The Restormer module includes an attention mechanism and a feed-forward mechanism in a series relationship before and after; first, through the attention mechanism, given the input Aggregate cross-channel context in the pixel direction through 1×1 convolution, and then apply 3×3 depth convolution to encode spatial context in the channel direction to generate query (Q), key (K), value (V) projections, that is and convolution, Depth convolution; reshape the query, key, and value projections so that their dot product interaction generates a transposed attention map of size instead of a large-scale conventional attention map of size
[0081] The attention mechanism is:
[0082]
[0083] Among them, X and are the input and output feature maps respectively; the matrices and are both reshaped from query (Q), key (K), value (V) of the original size
[0084] After passing through the feed-forward mechanism, given an input
[0085]
[0086] Among them, ⊙ represents element-wise multiplication; φ represents GELU non-linearity, and LN represents layer normalization.
[0087] S2. Fusion of multi-scale features; Input the five groups of different-scale features extracted from infrared images and visible light images into the feature fusion module respectively.
[0088] The multi-scale feature fusion includes semi-coupled extraction and convolutional fusion. Specifically, first perform feature extraction through the shared convolutional kernel and private convolutional kernel of the vi branch and the ir branch:
[0089]
[0090] Among them, and represent the shared convolutional kernel and the private convolutional kernel corresponding to the vi branch and the ir branch; represents concatenation in the channel dimension; * represents convolution.
[0091] After the above feature extraction, the semi-coupled features of the infrared image and the visible light image are fused using a convolutional operator to obtain fused features where i ∈ {1, 2, 3, 4, 5}, and the superscript f represents fusion.
[0092] S3. Through the sub-module CCT (Channel-wise Crossfusion Transformer) of the multi-scale channel cross Transformer, four groups of output features O with different scales are generated 1 、O 2 、O 3 、O 4 and restored; then the output O i is gradually taken and the corresponding decoder feature D i is used as the input of the sub-module CCA (Channel-wise Cross Attention) module of the cross Transformer to obtain the output and the current-level decoder feature D i are subjected to channel splicing and decoding + upsampling to finally obtain the upper-level decoder feature; using the multi-scale channel cross Transformer and layer-by-layer decoding + upsampling, a fused image is obtained.
[0093] The multi-scale channel cross Transformer includes a sub-module CCT for multi-scale encoder feature fusion, a sub-module CCA for decoder feature D i and sub-module CCT feature fusion; the fused features are connected to the decoder through the CCT feature cross-fusion CCA channel attention mechanism, replacing the traditional simple skip connection; the steps for the multi-scale channel cross Transformer to perform improved skip connection include:
[0094] Multi-scale encoder feature fusion: Feature embedding is performed on four fused features (where i ∈ {1, 2, 3, 4}) to obtain (where i ∈ {1, 2, 3, 4}), and they are connected together as where Through matrix operation, the three elements of the attention mechanism at each scale are obtained { K, V} where (i = 1, 2, 3, 4) calculates the multi-head attention at each scale, and the obtained output O 1 、O2 , O 3 , O 4 ; This enables the features of the fusion layer to be reasonably connected to the decoder.
[0095] Decoder feature D i and the fusion of the output of the CCT module: After restoring O 1 , O 2 , O 3 , O 4 , first perform spatial compression by the global average pooling (GAP) layer, generating a vector where the global pooling operation for the k-th channel is Use this operation to embed global spatial information, and then generate an attention mask for recalibrating or exciting O i , obtaining the output and the decoder feature D of this level i Perform channel concatenation and decoding + upsampling, and finally obtain the decoder feature D of the previous level i-1 . The decoder structure is symmetric to the encoder structure. Utilize multi-scale channel cross Transformer and layer-by-layer decoding + upsampling, and finally complete the final output through a 3×3 convolution operation (stride of 1), effectively reconstructing a high-resolution fused image.
[0096] S4. Use the input visible light image, infrared image, and the fused image obtained by fusion to construct an intensity-gradient-similarity loss function, and use the loss function to guide the training of the feature fusion network model. When the loss function converges, a trained feature fusion network model can be obtained, and the feature fusion network model is used for the fusion of infrared images and visible light images.
[0097] To improve the quality of the fused image, a weighted combination of intensity loss, gradient loss, and similarity loss is used as the loss function for training. The loss function adopts an intensity-gradient-similarity joint loss, and the loss function is specifically:
[0098] S41. Intensity loss: By comparing the fused image with the input infrared image and visible light image, evaluate the pixel-level content difference of the images; the intensity loss is:
[0099]
[0100] where M(·) is an element-wise maximum operation; the intensity loss is used to help the feature fusion network model obtain appropriate intensity information.
[0101] S42. Gradient Loss: By calculating the gradient information of the image, it is ensured that the fused image retains clear edge and structural information. Specifically, the Sobel operator is used to calculate the gradient of the image, and the optimal gradient information is retained; the gradient loss is:
[0102]
[0103] where represents the Sobel gradient operator; ||·|| 1 , represents the l 1 norm; |·| represents the absolute value operation; max(·) represents the element-wise maximum selection; the gradient loss is used to guide the feature fusion network model to retain the maximum amount of gradient information.
[0104] S43. Similarity Loss: The Structural Similarity Index (SSIM) reflects the differences between the source image and the fused image in terms of brightness, contrast, and structure. It is ensured that the output fused image has an optimal distribution in terms of brightness, contrast, and structure; the similarity loss is:
[0105]
[0106] where ssim(·) represents the structural similarity operation; set a 1 = a 2 = 0.5; is used to constrain I f and I ir , I vi between the structural similarity; the Structural Similarity Index SSIM reflects the defects of infrared images and visible light images in terms of brightness, contrast, and structure.
[0107] S44. The overall loss function is a weighted combination of intensity-gradient-similarity losses, and the feature fusion network model is trained with this until the loss function converges to obtain a trained feature fusion network model; the overall loss function of the feature fusion network model is a composite of all sub-item losses, and each sub-item is assigned a specific weight;
[0108]
[0109] where λ 1 , λ 2 and λ 3 are hyperparameters used to control the balance of each sub-item loss.
[0110] For example Figure 3 , 4As shown in the figure, by comparing this method with eight advanced methods, namely CSF, DenseFuse, FusionGAN, GANMcC, IFCNN, PMGI, U2Fusion, and UMF-CMGR, in the daytime scene, infrared thermal radiation data can be used as a supplement to visible light images. Therefore, an advanced fusion algorithm should enhance key targets without causing spectral contamination while maintaining the texture details of the Vi image; as Figure 3 shown, FusionGAN and GANMcC do not retain the texture information of the Vi image. Although CSF, DenseFuse, IFCNN, PMGI, U2Fusion, and UMF-CMGR integrate the texture of the Vi image with infrared information, the background area is affected by spectral noise pollution to varying degrees. This method performs well in retaining texture and highlighting target information. In the nighttime scene, the detail information in the Vi image is limited. In contrast, the infrared image can not only capture prominent targets but also capture detailed textures, thus enhancing the texture information of the Vi image; as Figure 4 shown, CSF, DenseFuse, IFCNN, U2Fusion, and UMF-CMGR do not effectively retain prominent target information. At the same time, FusionGAN and GANMcC lose background texture information, and PMGI suffers from overexposed backgrounds resulting in texture distortion. This method can effectively retain valuable information.
[0111] Embodiment 2
[0112] An electronic device includes a memory and a processor. A computer program is stored on the memory. When the processor executes the computer program, any step in the method of infrared and visible light image fusion based on multi-scale channel cross Transformer described in Embodiment 1 is implemented.
[0113] Furthermore, the process of the method of infrared and visible light image fusion based on multi-scale channel cross Transformer described in Embodiment 1 can be implemented as a computer software program. For example, this embodiment includes a computer program product, which includes a computer program carried on a computer-readable medium. The computer program contains program code for executing the method. In such an embodiment, the computer program can be downloaded and installed from the network and / or installed from a removable medium. When the computer program is executed by the processor, the above functions defined in the method of this application are executed.
[0114] Embodiment 3
[0115] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, any step in implementing an infrared and visible light image fusion method based on a multi-scale channel cross Transformer as described in Embodiment 1 is implemented.
[0116] The computer-readable medium described in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in this application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0117] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Python and C++, and also include conventional procedural programming languages or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0118] The computer-readable storage medium of this embodiment can be accelerated by hardware such as a GPU, and the parallel computing advantage of the GPU is utilized to accelerate any step in implementing an infrared and visible light image fusion method based on a multi-scale channel cross Transformer as described in Embodiment 1.
[0119] In summary, the present invention effectively overcomes the deficiencies in the prior art and has high industrial utilization value. The role of the above embodiments is to illustrate the substantial content of the present invention, but does not limit the protection scope of the present invention. Those of ordinary skill in the art should understand that the technical solution of the present invention can be modified or equivalently replaced without departing from the essence and protection scope of the technical solution of the present invention.
[0120] The above embodiments are specific implementation manners of the present invention, but the implementation manners of the present invention are not limited by the above embodiments. Any combination, change, modification, substitution, or simplification that does not exceed the design concept of the present invention falls within the protection scope of the present invention.
Claims
1. A method for fusion of infrared and visible light images based on multi-scale channel cross Transformer, characterized in that: The following steps are involved: S1, extraction of multi-scale features; A dual-branch multi-scale encoder is used to extract five sets of infrared and visible light image features of different scales from infrared images and visible light images respectively; S2, fusion of multi-scale features; The five groups of features of different scales extracted from the infrared image and the visible light image are input into the feature fusion module respectively; S3, after the submodule CCT of the multi-scale channel cross transformer, generates four sets of output features O1, O2, O3, O4 of different scales and restores them; then the output O i The corresponding decoder feature D i As the input of the CCA module of the cross Transformer submodule, the output is and the decoder characteristics D i Perform channel splicing and decoding + upsampling to finally obtain the features of the previous level decoder; use multi-scale channel cross Transformer and layer-by-layer decoding + upsampling to obtain the fused image; S4. Use the input visible light image, infrared image and fused image to construct an intensity-gradient-similarity loss function, and use the loss function to guide the training of the feature fusion network model to obtain a trained feature fusion network model. The feature fusion network model is used for the fusion of infrared images and visible light images.
2. The infrared and visible light image fusion method based on multi-scale channel cross Transformer according to claim 1 is characterized in that: In step S1, the steps of extracting multi-scale features of infrared images and visible light images using a dual-branch multi-scale encoder are as follows: S11, extraction of shallow features; the infrared image and the visible light image are respectively passed through a 3×3 convolution layer to pre-extract shallow features of the infrared image and the visible light image, and the convolution module is used to perform preliminary processing to extract each pre-extracted shallow feature from Transformed to S12, extraction of multi-scale features; After the shallow feature extraction is completed, deep features of different scales are extracted through a series of encoding modules and their corresponding downsampling operations; the dual-branch multi-scale encoder includes two parallel independent encoder branches, one of which is used to process the features of infrared images, and the other is used to process the features of visible light images.
3. The infrared and visible light image fusion method based on multi-scale channel cross Transformer according to claim 2 is characterized in that: In step S12, the independent encoder branches each include five encoding modules with the same structure, which are used to perform feature encoding processing on the data of each modality and extract multi-scale features from the source image; starting from the high-resolution input, the dual-branch multi-scale encoder gradually reduces the spatial dimension by downsampling, that is, reducing the spatial resolution to half, one quarter, one eighth of the original size, and the last layer is still one eighth; at the same time, the channel capacity is increased, and the final size is feature map; ensure that the feature information at different scales can be gradually encoded, and fully extract multi-level feature information from it; use markers and Distinguish the multi-scale information of infrared images and visible light images; where i∈{1,2,3,4,5}.
4. The infrared and visible light image fusion method based on multi-scale channel cross Transformer according to claim 3 is characterized in that: In step S12, the operation steps of the encoding module are as follows: the encoding module adopts a dual-branch structure, and each level of feature extraction uses a Restormer module, which includes an attention mechanism and a feedforward mechanism in a front-to-back series relationship; first, the attention mechanism is used to give an input The cross-channel context is aggregated in the pixel direction through 1×1 convolution, and then a 3×3 depthwise convolution is applied to encode the spatial context in the channel direction to produce query (Q), key (K), value (V) projections, i.e. and is the 1×1 convolution, and s the 3×3 depthwise convolution; reshapes the query, key, and value projections so that their dot product interaction produces a size of The transposed attention map of The attention mechanism is: Among them, X and are the input and output feature maps respectively; the matrix and All are from the original size query(Q), key(K), value(V); after the feedforward mechanism, given an input Among them, ⊙ represents element-wise multiplication; φ represents GELU nonlinearity, and LN represents layer normalization.
5. The infrared and visible light image fusion method based on multi-scale channel cross Transformer according to claim 1 is characterized in that: In step S2, multi-scale feature fusion includes semi-coupled extraction and convolution fusion, specifically: in, and Represents the shared convolution kernel and private convolution kernel corresponding to the vi branch and ir branch; Represents cascade on the channel dimension; * represents convolution; After the above feature extraction, the convolution operator is used to fuse the semi-coupled features of the infrared image and the visible light image to obtain the fused feature Among them, i∈{1,2,3,4,5}, and the superscript f represents fusion.
6. The infrared and visible light image fusion method based on multi-scale channel cross Transformer according to claim 1 is characterized in that: In step S3, the multi-scale channel cross Transformer includes a submodule CCT for multi-scale encoder feature fusion, a submodule CCT for decoder feature D i Submodule CCA fused with submodule CCT features; The steps of improving skip connections in multi-scale channel cross Transformer include: S31. Fusion features After the submodule CCT performs feature cross fusion, a multi-head attention O is generated. i ; S32, the generated O i Restore to Then, through the submodule CCA, the i-th level output and the i-th level decoder feature map As the input of the submodule CCA; the spatial compression is performed by the global average pooling (GAP) layer, producing a vector Among them, the global pooling operation of the kth channel is This operation is used to embed the global spatial information and then generate an attention mask for recalibration or stimulation of O i ; in, and is the linear layer weight matrix; δ(·),σ(·) are both activation functions; S33, Output and the upsampled features D of the i-th level decoder i Cascade decoding + upsampling in the channel dimension to obtain D i-1 ; The decoder structure is symmetrical to the encoder structure. It uses multi-scale channel cross-Transformer and layer-by-layer decoding + upsampling, and finally completes the final output through a 3×3 convolution operation with a stride of 1.
7. The infrared and visible light image fusion method based on multi-scale channel cross Transformer according to claim 6 is characterized in that: In step S31, the feature cross-fusion step of the submodule CCT includes: S311, multi-scale feature embedding; Given four fusion features Among them, i∈{1,2,3,4}, first by using the patch size P, The patches are respectively Serialize to get Among them, i∈{1,2,3,4}; in this process, the original channel size is maintained Then, the four layers The tokens are concatenated as Where i = 1, 2, 3, 4; d is the sequence length; S312, Acquisition of multi-head attention; The tokens are fed into a multi-head channel cross attention module, encoding channels and dependencies through a multi-layer perceptron with a residual structure; in, is the number of channels; get After projection, get attention; Among them, ψ(·) and σ(·) represent instance normalization and softmax function respectively; Among them, N represents the number of attention heads; The multi-head attention is; O i =MCA i +MLP(Q i +MCA i ) (9)。 8. The infrared and visible light image fusion method based on multi-scale channel cross Transformer according to claim 1 is characterized in that: In step S4, the loss function adopts the strength-gradient-similarity joint loss, and the loss function is specifically: S41, strength loss is: where M(·) is the element-wise maximum operation; the strength loss Used to help the feature fusion network model obtain appropriate intensity information; S42, gradient loss is: in, represents the Sobel gradient operator; ||·||1, represents the l1 norm; |·| represents the absolute value operation; max(·) represents the maximum element selection; gradient loss Used to guide the feature fusion network model to retain the maximum amount of gradient information; S43, similarity loss is: Where ssim(·) represents the structural similarity operation; set a1=a2=0.5; For constraint I f and I ir , I vi The structural similarity between them; the structural similarity index SSIM reflects the defects of infrared images and visible light images from three aspects: brightness, contrast and structure; S44. The overall loss function of the feature fusion network model is the composite of all sub-item losses, and each sub-item is assigned a specific weight; Among them, λ1, λ2, and λ3 are hyperparameters used to control the balance of each sub-item loss.
9. An electronic device, characterized in that: It includes at least one control processor and a memory for communicating with the at least one control processor; the memory stores instructions that can be executed by the at least one control processor, and the instructions are executed by the at least one control processor so that the at least one control processor can implement any step in the infrared and visible light image fusion method based on multi-scale channel cross Transformer as described in any one of claims 1-8 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer processor, it implements any step in the infrared and visible light image fusion method based on a multi-scale channel cross Transformer as described in any one of claims 1 to 8.
Citation Information
Cited By
Boiler scale detection method and device based on multi-modal image
CN120385686A
Multi-level attention visible light guided infrared image super-resolution method and system
CN120495087A
Dark visible light and infrared image fusion method, device, equipment and medium
CN120598803A
Guiding registration method, system and device for RGB and infrared image fusion and storage medium
CN121095299A