Infrared and visible light image fusion method based on cross-domain Transform

Through the encoding-fusion-decoding structure of the cross-domain Transformer, combined with axial attention and frequency domain information processing, the limitations of texture preservation and detail expression in the fusion of infrared and visible light images are solved, high-quality fused images are generated, and the fusion effect in complex scenes is improved.

CN120708011APending Publication Date: 2025-09-26FUZHOU UNIV
View PDF 0 Cites 9 Cited by

Patent Information

Application Number
CN202510834805.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing infrared and visible light image fusion methods have limitations in texture preservation, structure enhancement and detail expression, especially in complex scenes, the fusion effect is not ideal, and traditional CNN models find it difficult to capture long-distance dependencies and global context information.

Method used

A cross-domain Transformer-based method is adopted. Through the encoding-fusion-decoding structure, an axial attention mechanism and frequency domain information processing are introduced. By combining spatial and frequency domain features, a cross-attention mechanism and spatial and frequency domain collaborative modulation blocks are designed to enhance the structure and detail expression of the image.

Benefits of technology

It generates fused images with clear details and rich information, significantly improving the interpretability and reliability of downstream tasks and improving the fusion effect in complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708011A_ABST
    Figure CN120708011A_ABST
Patent Text Reader

Abstract

The invention relates to an infrared and visible light image fusion method based on a cross-domain Transform, and belongs to the field of computer image processing. The method comprises the following steps: respectively carrying out preprocessing operation on an infrared image and a visible light image to obtain a training data set; an end-to-end image generator network is designed, an encoder module is used for extracting deep semantic features of an infrared image and a visible light image, a fusion module introduces an axial attention mechanism to enhance the global modeling capability of the features, and feature fusion is carried out in combination with information of a spatial domain and a frequency domain; the fused features are gradually recovered to an image space through a decoder module, and a fused image is generated; constructing a fusion loss function module, and guiding the network to focus a significant feature difference between the source image and the fusion image based on a comparative learning idea; and finally, inputting the infrared and visible light image Y channel into the network model, generating a fusion image, completing a training process, and realizing unified optimization of fusion performance and visual quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer image processing, and specifically relates to a method for fusing infrared and visible light images based on a cross-domain Transformer. Background Art

[0002] With the continuous development of artificial intelligence and sensor technology, infrared and visible light image fusion technology has demonstrated significant value in many practical application scenarios and has been widely used in night vision monitoring, target recognition, environmental perception, medical imaging, smart transportation, and remote sensing monitoring. Infrared images can reflect the thermal radiation characteristics of objects and maintain excellent target perception capabilities under complex lighting or low-light conditions. Visible light images, on the other hand, contain rich details such as texture and color, which helps enhance the image's visual expressiveness. Therefore, the effective fusion of infrared and visible light images aims to take into account the complementary advantages of both and produce high-quality images with both structural clarity and information richness. This is of great significance for subsequent target detection, scene understanding, and intelligent analysis.

[0003] Early image fusion methods primarily relied on traditional image processing techniques, such as multi-scale transforms (wavelet and pyramid decomposition), image filtering, feature selection, and reconstruction strategies. While these methods can fuse key information from infrared and visible light images to a certain extent, they often have limitations in texture preservation, structure enhancement, and detail expression due to their lack of ability to model deep semantic information in the images. This results in unsatisfactory fusion results, especially in complex scenes.

[0004] In recent years, the rapid development of deep learning has injected new vitality into the field of image fusion. Image fusion methods based on convolutional neural networks (CNNs) leverage their powerful feature extraction and nonlinear modeling capabilities to automatically learn semantically significant information from the original images, significantly improving the quality of the fused image. Mainstream methods typically construct an end-to-end encoding-fusion-decoding structure, extracting deep features from infrared and visible light images through parallel or shared networks, and designing attention mechanisms, residual connections, or guidance strategies during the fusion stage to enhance the complementarity between images. However, the receptive field of traditional CNN models is limited, and capturing long-range dependencies and global contextual information remains challenging. In addition, the differences in the representation of infrared and visible light images in the spatial and frequency domains are often ignored, resulting in certain defects in the fusion results in detail reconstruction and edge preservation.

[0005] To this end, the present invention proposes a method for fusion of infrared and visible light images that integrates an axial attention mechanism with frequency domain information processing. This method is based on a feed-forward encoding-fusion-decoding structure. First, the encoder module extracts the multi-scale semantic features of infrared and visible light images. Subsequently, the axial attention mechanism is introduced into the fusion module to enhance the global modeling capability, and the spatial domain structure and frequency domain detail expression of the image are jointly considered in the fusion process. In order to further improve the fusion quality, the fusion module also designs a cross-attention mechanism and spatial and frequency domain collaborative modulation blocks to more finely model the complementarity and correlation between infrared and visible light images. Finally, the decoder module restores the image space to generate a fused image with clear details and rich information, which significantly improves the interpretability and reliability of downstream tasks. Summary of the Invention

[0006] The purpose of the present invention is to provide an infrared and visible light image fusion method based on a cross-domain Transformer. This method proposes an infrared and visible light image fusion method that integrates the axial attention mechanism and frequency domain information processing, and jointly considers the spatial domain structure and frequency domain detail expression of the image in the fusion process.

[0007] To achieve the above objectives, the technical solution of the present invention is: a method for fusion of infrared and visible light images based on a cross-domain Transformer, comprising the following steps:

[0008] Step A: Preprocess the infrared image and the visible light image to obtain a training data set;

[0009] Step B: Design an end-to-end image generator network using an encoder-fusion-decoder architecture. First, the encoder module extracts deep features from infrared and visible light images. During the fusion phase, an axial attention mechanism is introduced, combining spatial and frequency domain information to enhance image structure and detail expression to generate fused features. The fused features are then restored by the decoder module to generate a high-quality fused image.

[0010] Step C: Design a loss function based on contrastive learning. By constructing positive and negative sample pairs, the network is guided to learn the significant feature differences between the source image and the fused image, guiding the parameter optimization of the network designed in step B to obtain a trained infrared and visible fusion model.

[0011] Step D: Input the infrared and visible light images to be measured into the model trained in step C to generate a fused image.

[0012] Furthermore, step A includes the following steps:

[0013] Step A1: Pair the infrared images and visible light images used for training; then convert all visible light images to the YCbCr color space and extract the Y channel; and normalize the Y channels of all infrared images and visible light images to the range [0, 1].

[0014] Step A2: Scale the normalized infrared image and visible light image to a size of H×W, where H is the image height and W is the width. The processed paired infrared image and visible light image are used as a training data set.

[0015] Furthermore, the specific implementation steps of step B are as follows:

[0016] Step B1: Construct an encoder module to extract the input infrared image I r and Y channel I in the visible light image v Multi-scale feature information;

[0017] Step B2: Design a fusion module that combines spatial and frequency domain features to achieve efficient cross-modal information fusion through axial attention mechanism, frequency domain transformation processing, cross attention mechanism, and spatial and frequency domain collaborative modulation strategy;

[0018] Step B3: Construct a decoder module and use the dilated spatial pyramid pooling ASPP and multi-level convolution operations to reconstruct the Y channel I of the fused image. f ;

[0019] Step B4: Integrate the encoder module, fusion module and decoder module to build a complete end-to-end image generator network to generate the fused Y channel, and then convert the Y channel I of the fused image into f Together with the Cb channel and Cr channel of the visible light image, it is used as the final fused image.

[0020] Furthermore, the specific implementation steps of step B1 are as follows:

[0021] Step B11, construct an encoder module, which consists of a convolution layer and a downsampling operation; first, the Y channel of the infrared image or visible light image is used as the input image of the encoder module, the input image is edge-padded to compensate for the boundary, and then a first-layer convolution operation with a 3×3 convolution kernel and a stride of 1 is performed, and then the output result is instance normalized and an activation function is applied; then edge padding is performed again, and a second-layer convolution operation with a 3×3 convolution kernel and a stride of 1, instance normalization, and activation are performed; the output size remains unchanged at H×W, and the number of channels is C1=64. 1 , the specific formula is as follows:

[0022] ConvB3=LeakyReLU(InstanceNorm(Conv3(x)))

[0023] F 1 =ConvB3(ReflectionPad(ConvB3(ReflectionPad(I))))

[0024] Where I represents the Y channel of the input visible light image or infrared image, that is, I∈{I r , I v}, ReflectionPad(·) indicates edge padding achieved by mirroring, LeakyReLU(·) is the activation function, Conv3(·) indicates 3x3 convolution, InstanceNorm(·) indicates instance normalization, ConvB3(·) indicates a 3x3 convolution block consisting of 3x3 convolution, instance normalization, and LeakyReLU, and x indicates the input of the convolution block;

[0025] Step B12: Feature map F 1 First, spatial downsampling is achieved through a 3×3 convolutional layer with a stride of 2, and the image size is reduced to Then instance normalization and LeakyReLU activation are performed; after edge padding, the 3×3 convolution operation with preserved size, instance normalization and LeakyReLU activation are performed, and the output resolution is Feature map F with channel number C2=128 2 , the specific formula is as follows:

[0026] F 2 =ConvB3(ReflectionPad(Down scale=2 (F 1 )))

[0027] Down scale=2 (·) represents downsampling by reducing the width and height of the feature map by half, consisting of a 3×3 convolution with a stride of 2, instance normalization, and LeakyReLU activation;

[0028] Step B13: The feature map F 2 The same process as in step B12 yields a resolution of Feature map F with channel number C3=128 3 , the specific formula is as follows:

[0029] F 3 =ConvB3(ReflectionPad(Down scale=2 (F 2 )))

[0030] Step B14: transform the feature map F 3 The same process as in step B11 yields a resolution of Feature map F with C4=128 channels 4 , the specific formula is as follows:

[0031] F 4 =ConvB3(ReflectionPad(ConvB3(ReflectionPad(F 3 ))))

[0032] The outputs of the encoder module include Represents a real number.

[0033] Furthermore, the specific implementation steps of step B2 are as follows:

[0034] Step B21, design a fusion module, including four fusion blocks, each fusion block includes a spatial domain fusion block, a frequency domain fusion block, a cross attention mechanism and a spatial domain and frequency domain collaborative modulation block; the inputs of the four fusion blocks are the feature maps of four different scales output by the Y channels of the infrared image and the visible light image after step B1, that is, each fusion block processes the feature map of the same scale; the spatial domain fusion block contains two branches, in the first branch, the feature maps corresponding to the Y channels of the infrared image and the visible light image are respectively subjected to 1x1 convolution blocks, longitudinal-axial attention blocks, and transverse-axial attention blocks in sequence. force, 1x1 convolution block, residual connection, and then the obtained features are spliced ​​and passed through a 1x1 convolution block as the output feature of the first branch; in the second branch, the feature maps corresponding to the Y channel of the infrared image and the visible light image are first passed through a 1x1 convolution block respectively, and the obtained features are spliced, and then passed through a 1x1 convolution block, depth convolution, GELU activation function, matrix element-by-element multiplication, 1x1 convolution block and residual connection as the output feature of the second branch; then the output features of the two branches are added as the output feature of the spatial domain fusion block; specifically, and is the output feature of the Y channel of the input infrared image and visible image after step B1, where i∈{1, 2, 3, 4}, representing four different scales, The specific formula of the first branch is as follows:

[0035] ConvB1=LeakyReLU(InstanceNorm(Conv1(x)))

[0036]

[0037] Where ConvB1(·) represents a 1x1 convolution block consisting of 1x1 convolution, instance normalization, and LeakyReLU. is a matrix addition operation, Concat(·,·) is a concatenation operation along the channel dimension, Conv1(·) represents a 1x1 convolution, PA h (·),PA w (·) longitudinal-axial attention calculation and transverse-axial attention calculation, respectively, is the intermediate process feature of the first branch, A i is the output feature of the first branch,

[0038] The specific formula for the second branch and the sum of the output features of the two branches is as follows:

[0039]

[0040] F i “,F i "'=Split(DWConv(Conv1(F i ')))

[0041]

[0042] S i =D i +A i

[0043] in is a matrix element-by-element multiplication operation, DWConv(·) is a depthwise convolution, Split(·) means splitting the features into two groups of features in the channel dimension, with the number of channels in each group halved, and using 1x1 convolution before the depthwise convolution to increase the number of channels to twice the input channels, GELU(·) is the activation function, F i ', F i “,F i "' is the intermediate process characteristic of the second branch, D i is the output feature of the second branch, S i is the output feature of the spatial domain fusion block,

[0044] Step B22, design a frequency domain fusion block, including Fourier transform, 3x3 convolution block and inverse Fourier transform; first, the Y channels of the infrared image and the visible light image are subjected to the feature maps of four different scales output in step B1, and Fourier transform is used in the frequency domain to obtain the phase and amplitude of the infrared image and the phase and amplitude of the Y channel of the visible light image, and then the phase and amplitude of the Y channel of the infrared image and the visible light image are fused. The fused phase and amplitude are converted into spatial domain features after inverse Fourier transform, and then output after passing through two 3x3 convolution blocks; specifically, and is the output feature of the Y channel of the input infrared image and visible image after step B1, where i∈{1, 2, 3, 4}. The specific formula of the frequency domain fusion block is as follows:

[0045]

[0046] in are the amplitude and phase of the infrared image, is the amplitude and phase of the Y channel of the visible light image, FFT(·) is the Fourier transform operation, IFFT(·) is the inverse Fourier transform operation, ConvB3(·) represents a 3x3 convolution block, represents the output features of the frequency domain fusion block,

[0047] Step B23, design a cross attention mechanism; use the cross attention mechanism at four scales respectively, and the output feature S of the spatial domain fusion block at the same scale i Output features of the frequency domain fusion block Perform cross attention calculation to obtain the cross attention result m. ​​The specific formula of the cross attention mechanism is as follows:

[0048]

[0049] where CA(·,·,·) represents the cross attention mechanism and d is the embedding dimension.(·) T represents the transposition operation, Softmax(·) is the activation function;

[0050] Step B24, spatial domain and frequency domain collaborative modulation block; the attention result obtained in step B23 is input into the spatial domain and frequency domain collaborative modulation block CM for further decoupling processing; CM consists of a set of simple and efficient convolution and activation functions, specifically including a 1×1 convolution layer for capturing local context information, followed by a ReLU activation function to enhance the nonlinear expression ability, and then a 1×1 convolution layer for outputting the final modulation factor; after the spatial domain and frequency domain collaborative modulation block, the modulation factor μ,v is obtained; finally, the output feature S of the spatial domain fusion block obtained in step B21 is iModulate and obtain the fusion features of the final output of the fusion module The specific formula of the spatial domain and frequency domain coordinated modulation block is as follows:

[0051] μ,v=CM(m)

[0052]

[0053] in, It is the output feature of the spatial domain and frequency domain collaborative modulation block, and is also the fusion feature finally output by the fusion module.

[0054] Furthermore, the specific implementation steps of step B3 are as follows:

[0055] First, the fusion features of the third and fourth layers output by the fusion module and The encoder module output features of the third and fourth layers with the encoder module output Splicing is performed in the channel dimension to obtain the splicing feature F b , and then processed by the void space pyramid pooling ASPP to obtain the void space pyramid pooling feature Then Entering the layer-by-layer decoding stage, the resolution of the original image is gradually restored, and finally the fused Y channel is obtained; the layer-by-layer decoding stage includes two processes of layer-by-layer upsampling and feature fusion; first, the feature It is upsampled twice and outputs the fusion features of the second layer with the fusion module The second layer features output by the encoder module The concatenation is performed on the channel dimension and then processed through a 3x3 convolution block to obtain the intermediate feature F m ; Next, the intermediate feature F m Perform double upsampling again and output the fusion features of the first layer with the fusion module The first layer of features output by the encoder module The intermediate feature F is obtained by concatenating the layers in the channel dimension and processing them through a 3x3 convolution block. t ; Finally, the intermediate feature F t Through a 3x3 convolution block, a 3×3 convolution layer, and then the Sigmoid function to output the final fusion Y channel I f , the specific formula is as follows:

[0056]

[0057] I f =Sigmoid(Conv3(ConvB3((F t )))

[0058] in, Up scale=2 (·) represents the upsampling operation, which uses bilinear interpolation to double the width and height of the feature map; ASPP(·) is the atrous spatial pyramid pooling operation; Among them, C4=128, C5=512, C6=384, C7=256, and Sigmoid is the activation function operation.

[0059] Furthermore, the specific implementation steps of step C are as follows:

[0060] The total loss function is defined to consist of four parts: strength loss Gradient Loss Structural similarity loss and contrast loss The specific formula is as follows

[0061]

[0062] Among them, λ i represents the weight parameter, i∈{1,2,3,4};

[0063] Strength loss The specific formula is:

[0064]

[0065] Among them, I v , I r are the Y channel of the visible light image and the infrared image input in step B1, H represents the height, W represents the width, max{·,·} represents the maximum value operation, ∑ represents the sum operation of all pixel positions, and |·,·| represents the mean absolute error, i.e., L1 loss;

[0066] Gradient Loss The specific formula is:

[0067]

[0068] Where G(·) represents the spatial gradient of the calculated image;

[0069] Structural similarity loss The specific formula is:

[0070]

[0071] The structural similarity index SSIM is used to measure the similarity between the fused Y channel and the Y channels of the infrared image and the visible light image. SSIM(·,·) represents the structural similarity index calculation operation;

[0072] Contrastive loss Based on a contrastive learning mechanism with a saliency binary mask, weighted cosine similarity is calculated for "salient regions" and "non-salient regions" respectively. The salient regions are represented by a mask M, which is obtained by performing salient target detection on infrared images. The detected salient targets are regarded as salient regions, and other regions are regarded as non-salient regions.

[0073] For the salient region, namely the mask M:

[0074] Positive contrast loss for infrared images:

[0075] Negative sample contrast loss of the Y channel of the visible light image:

[0076] Where weighted_sim(·,·,·) is the weighted average of the cosine similarity calculated in the corresponding area;

[0077] For non-salient regions (1-M):

[0078] Positive sample contrast loss of the Y channel of the visible light image:

[0079] Negative contrast loss for infrared images:

[0080] In summary, the contrast loss The specific formula is as follows:

[0081]

[0082] Furthermore, the specific implementation steps of step D are as follows:

[0083] Step D1: randomly divide the infrared image and visible light image fusion training dataset constructed in step A into several batches, each batch containing N pairs of aligned infrared image and visible light image inputs;

[0084] Step D2: For each batch of image pairs, input the infrared image and the visible light image into the image fusion network described in step B to obtain a fused image;

[0085] Step D3: Based on the total loss function defined in step C, the gradient of the network parameters is calculated using the backpropagation method, and the full network parameters are updated in combination with the Adam optimizer.

[0086] Step D4: Repeat steps D1 to D3 in batches until the loss function converges or the preset number of training rounds is met, and save the image fusion network parameters obtained from the final training as the trained infrared and visible light image fusion model.

[0087] The present invention also provides an infrared and visible light image fusion system based on a cross-domain Transformer, comprising a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, it can implement any of the method steps described above.

[0088] The present invention also provides a computer-readable storage medium on which computer program instructions that can be executed by a processor are stored. When the processor executes the computer program instructions, any of the method steps described above can be implemented.

[0089] Compared with the existing technology, the present invention has the following beneficial effects: the method of the present invention proposes an infrared and visible light image fusion method that integrates the axial attention mechanism and frequency domain information processing, and jointly considers the spatial domain structure and frequency domain detail expression of the image in the fusion process. BRIEF DESCRIPTION OF THE DRAWINGS

[0090] Figure 1 It is a flow chart for realizing the method of the present invention.

[0091] Figure 2 This is a structural diagram of the infrared and visible light image fusion network based on the cross-domain Transformer in an embodiment of the present invention.

[0092] Figure 3 It is a structural diagram of the fusion module in an embodiment of the present invention.

[0093] Figure 4 4 is a structural diagram of the spatial domain fusion block in the fusion module in an embodiment of the present invention.

[0094] Figure 5 4 is a structural diagram of the frequency domain fusion block in the fusion module in an embodiment of the present invention. DETAILED DESCRIPTION

[0095] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.

[0096] The present invention provides a method for fusion of infrared and visible light images based on a cross-domain Transformer, comprising the following steps:

[0097] Step A: Preprocess the infrared image and the visible light image to obtain a training data set;

[0098] Step B: Design an end-to-end image generator network using an encoder-fusion-decoder architecture. First, the encoder module extracts deep features from infrared and visible light images. During the fusion phase, an axial attention mechanism is introduced, combining spatial and frequency domain information to enhance image structure and detail expression to generate fused features. The fused features are then restored by the decoder module to generate a high-quality fused image.

[0099] Step C: Design a loss function based on contrastive learning. By constructing positive and negative sample pairs, the network is guided to learn the significant feature differences between the source image and the fused image, guiding the parameter optimization of the network designed in step B to obtain a trained infrared and visible fusion model.

[0100] Step D: Input the infrared and visible light images to be measured into the model trained in step C to generate a fused image.

[0101] The following is a specific implementation process of the present invention.

[0102] The present invention provides a method for fusion of infrared and visible light images based on cross-domain Transformer, such as Figure 1 As shown, the following steps are included:

[0103] Step A: Preprocess the infrared image and visible light image separately, including channel extraction, normalization, and size unification, to obtain a training data set;

[0104] Step B: Design an end-to-end image generator network (i.e. Figure 2 The infrared and visible light image fusion network based on a cross-domain Transformer (shown in Figure 2) uses an encoder-fusion-decoder structure. The encoder first extracts deep features from the infrared and visible light images. During the fusion phase, an axial attention mechanism is introduced, combining spatial and frequency domain information to enhance image structure and detail. The fused features are then restored by the decoder to produce a high-quality fused image.

[0105] Step C: Design a loss function based on contrastive learning. By constructing positive and negative sample pairs, the network is guided to learn the significant feature differences between the source image and the fused image, guiding the parameter optimization of the network designed in step B to obtain a trained infrared and visible fusion model.

[0106] Step D: Input the infrared and visible light images to be measured into the model trained in step C to generate a fused image;

[0107] Furthermore, step A includes the following steps:

[0108] Step A1: Pair the infrared and visible light images used for training. Then, convert all visible light images to the YCbCr color space and extract the Y channel. Normalize the Y channels of all infrared and visible light images to the range [0, 1].

[0109] Step A2: Scale the normalized infrared image and visible light image (YCbCr color space) to H×W size, where H is the image height and W is the width. The paired infrared image and visible light image after the above processing are used as the training data set. Further, step B includes the following steps:

[0110] Step B1: Construct an encoder module to extract the input infrared image I r and Y channel I in the visible light image v Multi-scale feature information;

[0111] Step B2, design fusion module (such as Figure 3-5 As shown in the figure), combining spatial and frequency domain features, the efficient fusion of cross-modal information is achieved through axial attention mechanism, frequency domain transformation processing, cross attention mechanism and spatial and frequency domain collaborative modulation strategy;

[0112] Step B3: Construct a decoder module and use Atrous Spatial Pyramid Pooling (ASPP) and multi-level convolution operations to reconstruct the Y channel I of the fused image. f ;

[0113] Step B4: Integrate the encoder module, fusion module and decoder module to construct a complete image fusion network to generate a fused Y channel, and then use the fused Y channel and the Cb and Cr channels of the visible light image as the final fused image.

[0114] Furthermore, step B1 includes the following steps:

[0115] Step B11, construct the encoder module, which mainly consists of convolution layers and downsampling operations. First, the Y channel of the infrared image or visible light image is used as the input image of the encoder module, the input image is edge-filled to compensate for the boundary, and then the first layer of convolution operation with a 3×3 convolution kernel and a stride of 1 is performed. Then, the output result is instance normalized and an activation function is applied; then edge filling, the second layer of convolution operation (also using a 3×3 convolution kernel and a stride of 1 configuration), instance normalization, and activation are performed again. The output size remains unchanged H×W, and the number of channels is C1=64, which is a feature map F. 1 , the specific formula is as follows:

[0116] ConvB3=LeakyReLU(InstanceNorm(Conv3(x)))

[0117] F 1 =ConvB3(ReflectionPad(ConvB3(ReflectionPad(I))))

[0118] Where I represents the Y channel of the input visible light image or infrared image, that is, I∈{I r , I v}, ReflectionPad(·) indicates edge padding achieved by mirroring, LeakyReLU(·) is the activation function, Conv3(·) indicates 3x3 convolution, InstanceNorm(·) indicates instance normalization, ConvB3(·) indicates a 3x3 convolution block consisting of 3x3 convolution, instance normalization, and LeakyReLU, and x indicates the input of the convolution block;

[0119] Step B12: Feature map F 1 First, spatial downsampling is achieved through a 3×3 convolutional layer with a stride of 2 Then instance normalization and LeakyReLU activation are performed; after edge padding, a 3×3 convolution operation (stride 1) with preserved size is performed, followed by instance normalization and LeakyReLU activation. The output resolution is Feature map F with channel number C2=128 2 , the specific formula is as follows:

[0120] F 2 =ConvB3(ReflectionPad(Down scale=2 (F 1 )))

[0121] Down scale=2 (·) represents downsampling by reducing the width and height of the feature map by half, consisting of a 3×3 convolution with a stride of 2, instance normalization, and LeakyReLU activation;

[0122] Step B13: The feature map F 2 The same process as in step B12 yields a resolution of Feature map F with channel number C3=128 3 , the specific formula is as follows:

[0123] F 3 =ConvB3(ReflectionPad(Down scale=2 (F 2 )))

[0124] Step B14: transform the feature map F 3 The same process as in step B11 yields a resolution of Feature map F with C4=128 channels 4 , the specific formula is as follows:

[0125] F 4 =ConvB3(ReflectionPad(ConvB3(ReflectionPad(F 3 ))))

[0126] The outputs of the encoder module include represents a real number;

[0127] Furthermore, step B2 includes the following steps:

[0128] Step B21: Design a fusion module. This module includes four fusion blocks. Each fusion block mainly includes a spatial domain fusion block, a frequency domain fusion block, a cross-attention mechanism, and a spatial and frequency domain co-modulation block. The inputs of the four fusion blocks in this module are the feature maps of four different scales output by the Y channels of the infrared image and the visible light image after step B1, that is, each fusion block processes feature maps of the same scale. The spatial domain fusion block contains two branches. In the first branch, the feature maps corresponding to the Y channels of the infrared image and the visible light image are respectively subjected to a 1x1 convolution block, longitudinal-axial attention, transverse-axial attention, a 1x1 convolution block, and a residual connection. The resulting features are then concatenated and passed through a 1x1 convolution block as the output features of the first branch, thereby enhancing interactive capabilities at a finer granularity while obtaining global dependency information. In the second branch, the feature maps corresponding to the Y channel of the infrared image and the visible light image are first passed through a 1x1 convolution block respectively. The obtained features are then concatenated and passed through a 1x1 convolution block to restore the number of channels, depthwise convolution, GELU activation function, matrix element-wise multiplication, 1x1 convolution block and residual connection to serve as the output features of the second branch. The output features of the two branches are then added together as the output features of the spatial domain fusion block. Specifically, and is the output feature of the Y channel of the input infrared image and visible image after step B1, where i∈{1, 2, 3, 4}, representing four different scales, The specific formula of the first branch is as follows:

[0129] ConvB1=LeakyReLU(InstanceNorm(Conv1(x)))

[0130]

[0131] Where ConvB1(·) represents a 1x1 convolution block consisting of 1x1 convolution, instance normalization, and LeakyReLU. is a matrix addition operation, Concat(·,·) is a concatenation operation along the channel dimension, Conv1(·) represents a 1x1 convolution, PA h (·), PA w (·) longitudinal-axial attention calculation and transverse-axial attention calculation, respectively, is the intermediate process feature of the first branch, A i is the output feature of the first branch, The specific formula for the second branch and the sum of the output features of the two branches is as follows:

[0132]

[0133] F i “,F i "'=Split(DWConv(Conv1(F i ')))

[0134]

[0135] S i =D i +A i

[0136] in is a matrix element-by-element multiplication operation, DWConv(·) is a depthwise convolution, Split(·) means splitting the features into two groups of features in the channel dimension, with the number of channels in each group halved, and using 1x1 convolution before the depthwise convolution to increase the number of channels to twice the input channels, GELU(·) is the activation function, F i ', F i “,F i "' is the intermediate process characteristic of the second branch, D i is the output feature of the second branch, S i is the output feature of the spatial domain fusion block,

[0137] Step B22, design the frequency domain fusion block, which mainly includes Fourier transform, 3x3 convolution block and inverse Fourier transform. First, the Y channels of the infrared image and the visible light image are subjected to the four different scale feature maps output by step B1, and Fourier transform is used in the frequency domain to obtain the phase and amplitude of the infrared image and the phase and amplitude of the Y channel of the visible light image. Then, the phase and amplitude of the Y channel of the infrared image and the visible light image are fused. The fused phase and amplitude are converted into spatial domain features after inverse Fourier transform, and then output after passing through two 3x3 convolution blocks. Specifically, and is the output feature of the Y channel of the input infrared image and visible image after step B1, where i∈{1, 2, 3, 4}. The specific formula of the frequency domain fusion block is as follows:

[0138]

[0139] in are the amplitude and phase of the infrared image, is the amplitude and phase of the Y channel of the visible light image, FFT(·) is the Fourier transform operation, IFFT(·) is the inverse Fourier transform operation, ConvB3(·) represents a 3x3 convolution block, represents the output features of the frequency domain fusion block,

[0140] Step B23, design a cross attention mechanism. The cross attention mechanism is used at four scales respectively, and the output feature S of the spatial domain fusion block at the same scale is i Output features of the frequency domain fusion block Perform cross-attention calculations to obtain the cross-attention result m. ​​The core function of the cross-attention mechanism is to guide spatial features to actively focus on the important information contained in the fusion of the frequency domain fusion block, especially the potential features such as brightness contrast, structural edges, and texture details in amplitude and phase, thereby improving the spatial features' ability to perceive high-frequency and structural signals. The specific formula of the cross-attention mechanism is as follows:

[0141]

[0142] where CA(·,·,·) represents the cross attention mechanism and d is the embedding dimension.(·) T represents the transposition operation, and Softmax(·) is the activation function.

[0143] Step B24, spatial domain and frequency domain collaborative modulation block. The attention result obtained in step B23 is input into the spatial domain and frequency domain collaborative modulation block (CM) for further decoupling processing. CM consists of a set of simple and efficient convolution and activation functions, specifically including a 1×1 convolution layer for capturing local context information, followed by a ReLU activation function to enhance the nonlinear expression ability, and then a 1×1 convolution layer to output the final modulation factor. After the spatial domain and frequency domain collaborative modulation block, the modulation factor μ,v is obtained. Finally, the output feature S of the spatial domain fusion block obtained in step B21 is i Modulate and obtain the fusion features of the final output of the fusion module The specific formula of the spatial domain and frequency domain coordinated modulation block is as follows:

[0144] μ,v=CM(m)

[0145]

[0146] in, It is the output feature of the spatial domain and frequency domain collaborative modulation block, and is also the fusion feature finally output by the fusion module.

[0147] Furthermore, step B3 includes the following steps:

[0148] First, the fusion features of the third and fourth layers output by the fusion module and The encoder module output features of the third and fourth layers with the encoder module output Splicing is performed in the channel dimension to obtain the splicing feature F b , and then processed by Atrous Spatial Pyramid Pooling (ASPP) to obtain the Atrous Spatial Pyramid Pooling feature Then Enter the layer-by-layer decoding stage, gradually restore the resolution of the original image, and finally obtain the fused Y channel. The layer-by-layer decoding stage includes two layers of upsampling and feature fusion processes. First, the feature It is upsampled twice and outputs the fusion features of the second layer with the fusion module The second layer features output by the encoder module The concatenation is performed on the channel dimension and then processed through a 3x3 convolution block to obtain the intermediate feature F m Then, the intermediate feature F m Perform double upsampling again and output the fusion features of the first layer with the fusion module The first layer of features output by the encoder module The intermediate feature F is obtained by concatenating the layers in the channel dimension and processing them through a 3x3 convolution block. tFinally, the intermediate feature F t Through a 3x3 convolution block, a 3×3 convolution layer, and then the Sigmoid function to output the final fusion Y channel F out , the specific formula is as follows:

[0149]

[0150] F out =Sigmoid(Conv3(ConvB3((F t )))

[0151] in, Up scale=2 (·) represents an upsampling operation, which uses bilinear interpolation to double the width and height of the feature map. ASPP(·) is a dilated spatial pyramid pooling operation. Among them, C4=128, C5=512, C6=384, C7=256, and Sigmoid is the activation function operation.

[0152] Furthermore, step C includes the following steps:

[0153] The total loss function is defined to consist of four parts: strength loss Gradient Loss Structural similarity loss and contrast loss The specific formula is as follows

[0154]

[0155] Among them, λ i represents the weight parameter, i∈{1,2,3,4};

[0156] The specific formula for strength loss is:

[0157]

[0158] Among them I f Indicates the fused Y channel obtained in step B3, I v , I r are the Y channel of the visible light image and the infrared image input in step B1, H represents the height, W represents the width, max{·,·} represents the maximum value operation, Σ represents the sum operation of all pixel positions, and |·,·| represents the mean absolute error, i.e., L1 loss;

[0159] Gradient Loss The specific formula is:

[0160]

[0161] Where G(·) represents the spatial gradient of the calculated image;

[0162] Structural similarity loss The specific formula is:

[0163]

[0164] The structural similarity index SSIM is used to measure the similarity between the fused Y channel and the Y channels of the infrared image and the visible light image. SSIM(·,·) represents the structural similarity index calculation operation;

[0165] is the contrast loss, which is based on a contrastive learning mechanism with a saliency binary mask. The weighted cosine similarity is calculated for the "salient region" and "non-salient region" respectively. The salient region is represented by a mask M, which is obtained by performing salient target detection on the infrared image. The detected salient targets are regarded as salient regions, and other regions are regarded as non-salient regions. For the salient region (mask M):

[0166] Positive contrast loss for infrared images:

[0167] Negative sample contrast loss of the Y channel of the visible light image:

[0168] Where weighted_sim(·,·,·) is the weighted average of the cosine similarity calculated in the corresponding area. For non-salient areas (1-M):

[0169] Positive sample contrast loss of the Y channel of the visible light image:

[0170] Negative contrast loss for infrared images:

[0171] In summary, the contrast loss The specific formula is as follows:

[0172]

[0173] Furthermore, step D comprises the following steps:

[0174] Step D1: randomly divide the infrared image and visible light image fusion training dataset constructed in step A into several batches, each batch containing N pairs of aligned infrared image and visible light image inputs.

[0175] Step D2: For each batch of image pairs, the infrared image and the visible light image are respectively input into the image fusion network described in step B to obtain a fused image.

[0176] Step D3: Based on the total loss function defined in step C, the gradient of the network parameters is calculated using the backpropagation method, and the Adam optimizer is used to update the parameters of the entire network.

[0177] Step D4: Repeat steps D1 to D3 in batches until the loss function converges or the preset number of training rounds is met, and save the image fusion network parameters obtained from the final training as the trained infrared and visible light image fusion model.

[0178] The present invention also provides an infrared and visible light image fusion system based on a cross-domain Transformer, comprising a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, it can implement any of the method steps described above.

[0179] The present invention also provides a computer-readable storage medium on which computer program instructions that can be executed by a processor are stored. When the processor executes the computer program instructions, any of the method steps described above can be implemented.

[0180] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions and effects do not exceed the scope of the technical solution of the present invention, shall fall within the scope of protection of the present invention.

Claims

1. A cross-domain Transformer-based infrared and visible light image fusion method, characterized in that: The following steps are involved: Step A: Preprocess the infrared image and the visible light image to obtain a training data set; Step B: Design an end-to-end image generator network using an encoder-fusion-decoder architecture. First, the encoder module extracts deep features from infrared and visible light images. During the fusion phase, an axial attention mechanism is introduced, combining spatial and frequency domain information to enhance image structure and detail expression to generate fused features. The fused features are then restored by the decoder module to generate a high-quality fused image. Step C: Design a loss function based on contrastive learning. By constructing positive and negative sample pairs, the network is guided to learn the significant feature differences between the source image and the fused image, guiding the parameter optimization of the network designed in step B to obtain a trained infrared and visible fusion model. Step D: Input the infrared and visible light images to be measured into the model trained in step C to generate a fused image.

2. The infrared and visible light image fusion method based on cross-domain Transformer according to claim 1 is characterized in that: Step A includes the following steps: Step A1: Pair the infrared images and visible light images used for training; then convert all visible light images to the YCbCr color space and extract the Y channel; and normalize the Y channels of all infrared images and visible light images to the range [0, 1]. Step A2: Scale the normalized infrared image and visible light image to a size of H×W, where H is the image height and W is the width. The processed paired infrared image and visible light image are used as a training data set.

3. The infrared and visible light image fusion method based on cross-domain Transformer according to claim 1 is characterized in that: The specific implementation steps of step B are as follows: Step B1: Construct an encoder module to extract the input infrared image I r and Y channel I in the visible light image v Multi-scale feature information; Step B2: Design a fusion module that combines spatial and frequency domain features to achieve efficient cross-modal information fusion through axial attention mechanism, frequency domain transformation processing, cross attention mechanism, and spatial and frequency domain collaborative modulation strategy; Step B3: Construct a decoder module and use the dilated spatial pyramid pooling ASPP and multi-level convolution operations to reconstruct the Y channel I of the fused image. f ; Step B4: Integrate the encoder module, fusion module and decoder module to build a complete end-to-end image generator network to generate the fused Y channel, and then convert the Y channel I of the fused image into f Together with the Cb channel and Cr channel of the visible light image, it is used as the final fused image.

4. The infrared and visible light image fusion method based on cross-domain Transformer according to claim 3 is characterized in that: The specific implementation steps of step B1 are as follows: Step B11, construct an encoder module, which consists of a convolution layer and a downsampling operation; first, the Y channel of the infrared image or visible light image is used as the input image of the encoder module, the input image is edge-padded to compensate for the boundary, and then a first-layer convolution operation with a 3×3 convolution kernel and a stride of 1 is performed, and then the output result is instance normalized and an activation function is applied; then edge padding is performed again, and a second-layer convolution operation with a 3×3 convolution kernel and a stride of 1, instance normalization, and activation are performed; the output size remains unchanged at H×W, and the number of channels is C1=64. 1 , the specific formula is as follows: ConvB3=LeakyReLU(InstanceNorm(Conv3(x))) F 1 =ConvB3(ReflectionPad(ConvB3(ReflectionPad(I)))) Where I represents the Y channel of the input visible light image or infrared image, that is, I∈{I r , I v }, ReflectionPad(·) indicates edge padding achieved by mirroring, LeakyReLU(·) is the activation function, Conv3(·) indicates 3x3 convolution, InstanceNorm(·) indicates instance normalization, ConvB3(·) indicates a 3x3 convolution block consisting of 3x3 convolution, instance normalization, and LeakyReLU, and x indicates the input of the convolution block; Step B12: Feature map F 1 First, spatial downsampling is achieved through a 3×3 convolutional layer with a stride of 2, and the image size is reduced to Then instance normalization and LeakyReLU activation processing are performed; After edge padding, the 3×3 convolution operation, instance normalization and LeakyReLU activation are performed to maintain the size, and the output resolution is Feature map F with channel number C2=128 2 , the specific formula is as follows: F 2 =ConvB3(ReflectionPad(Down scale=2 (F 1 ))) Down scale=2 (·) represents downsampling by reducing the width and height of the feature map by half, consisting of a 3×3 convolution with a stride of 2, instance normalization, and LeakyReLU activation; Step B13: The feature map F 2 The same process as in step B12 yields a resolution of Feature map F with channel number C3=128 3 , the specific formula is as follows: F 3 =ConvB3(ReflectionPad(Down scale=2 (F 2 ))) Step B14: transform the feature map F 3 The same process as in step B11 yields a resolution of Feature map F with C4=128 channels 4 , the specific formula is as follows: F 4 =ConvB3(ReflectionPad(ConvB3(ReflectionPad(F 3 )))) The outputs of the encoder module include Represents a real number.

5. The infrared and visible light image fusion method based on cross-domain Transformer according to claim 4 is characterized in that: The specific implementation steps of step B2 are as follows: Step B21: Design a fusion module, including four fusion blocks. Each fusion block includes a spatial domain fusion block, a frequency domain fusion block, a cross-attention mechanism, and a spatial and frequency domain co-modulation block. The inputs of the four fusion blocks are the feature maps of four different scales output by the Y channels of the infrared image and the visible light image after step B1, that is, each fusion block processes the feature map of the same scale; the spatial domain fusion block contains two branches. In the first branch, the feature maps corresponding to the Y channels of the infrared image and the visible light image are respectively subjected to 1x1 convolution blocks, longitudinal-axial attention, transverse-axial attention, 1x1 convolution blocks, and residual connections in sequence. The obtained features are then concatenated and then passed through a 1x1 convolution block as the output features of the first branch; in the second branch, the feature maps corresponding to the Y channels of the infrared image and the visible light image are first respectively passed through a 1x1 convolution block, and the obtained features are concatenated, and then passed through a 1x1 convolution block, depthwise convolution, GELU activation function, matrix element-by-element multiplication, 1x1 convolution blocks and residual connections as the output features of the second branch; Then the output features of the two branches are added as the output features of the spatial domain fusion block; specifically, and is the output feature of the Y channel of the input infrared image and visible image after step B1, where i∈1,2,3,4}, representing four different scales, The specific formula of the first branch is as follows: ConvB1=LeakyReLU(InstanceNorm(Conv1(x))) Where ConvB1(·) represents a 1x1 convolution block consisting of 1x1 convolution, instance normalization, and LeakyReLU. is a matrix addition operation, Concat(·,·) is a concatenation operation along the channel dimension, Conv1(·) represents a 1x1 convolution, PA h (·),PA w (·) longitudinal-axial attention calculation and transverse-axial attention calculation, respectively, is the intermediate process feature of the first branch, A i is the output feature of the first branch, The specific formula for the second branch and the sum of the output features of the two branches is as follows: F i ″,F i ″′=Split(DWConv(Conv1(F i ′))) S i =D i +A i in is a matrix element-by-element multiplication operation, DWConv(·) is a depthwise convolution, Split(·) means splitting the features into two groups of features in the channel dimension, with the number of channels in each group halved, and using 1x1 convolution before the depthwise convolution to increase the number of channels to twice the input channels, GELU(·) is the activation function, F i ', F i ",F i "' is the intermediate process feature of the second branch, D i is the output feature of the second branch, S i is the output feature of the spatial domain fusion block, Step B22, design a frequency domain fusion block, including Fourier transform, 3x3 convolution block and inverse Fourier transform; first, the Y channels of the infrared image and the visible light image are subjected to the feature maps of four different scales output in step B1, and Fourier transform is used in the frequency domain to obtain the phase and amplitude of the infrared image and the phase and amplitude of the Y channel of the visible light image, and then the phase and amplitude of the Y channel of the infrared image and the visible light image are fused. The fused phase and amplitude are converted into spatial domain features after inverse Fourier transform, and then output after passing through two 3x3 convolution blocks; specifically, and is the output feature of the Y channel of the input infrared image and visible image after step B1, where i∈{1, 2, 3, 4}. The specific formula of the frequency domain fusion block is as follows: in are the amplitude and phase of the infrared image, is the amplitude and phase of the Y channel of the visible light image, FFT(·) is the Fourier transform operation, IFFT(·) is the inverse Fourier transform operation, ConvB3(·) represents a 3x3 convolution block, represents the output features of the frequency domain fusion block, Step B23, design a cross attention mechanism; use the cross attention mechanism at four scales respectively, and the output feature S of the spatial domain fusion block at the same scale i Output features of the frequency domain fusion block Perform cross attention calculation to obtain the cross attention result m. ​​The specific formula of the cross attention mechanism is as follows: where CA(·,·,·) represents the cross attention mechanism and d is the embedding dimension.(·) T represents the transposition operation, Softmax(·) is the activation function; Step B24, spatial and frequency domain co-modulation block; the attention results obtained in step B23 are input into the spatial and frequency domain co-modulation block CM for further decoupling processing; CM consists of a set of simple and efficient convolution and activation functions, specifically including a 1×1 convolution layer for capturing local context information, followed by a ReLU activation function to enhance nonlinear expression capabilities, and then a 1×1 convolution layer to output the final modulation factor; After the spatial domain and frequency domain collaborative modulation blocks, the modulation factor μ,v is obtained; finally, the output feature S of the spatial domain fusion block obtained in step B21 is i Modulate and obtain the fusion features of the final output of the fusion module i∈{1, 2, 3, 4}; the specific formula of the spatial domain and frequency domain coordinated modulation block is as follows: μ,v=CM(m) in, It is the output feature of the spatial domain and frequency domain collaborative modulation block, and is also the fusion feature finally output by the fusion module.

6. The infrared and visible light image fusion method based on cross-domain Transformer according to claim 5 is characterized in that: The specific implementation steps of step B3 are as follows: First, the fusion features of the third and fourth layers output by the fusion module and The encoder module output features of the third and fourth layers with the encoder module output Splicing is performed in the channel dimension to obtain the splicing feature F b , and then processed by the void space pyramid pooling ASPP to obtain the void space pyramid pooling feature Then Entering the layer-by-layer decoding stage, the resolution of the original image is gradually restored, and finally the fused Y channel is obtained; the layer-by-layer decoding stage includes two processes of layer-by-layer upsampling and feature fusion; first, the feature It is upsampled twice and outputs the fusion features of the second layer with the fusion module The second layer features output by the encoder module The concatenation is performed on the channel dimension and then processed through a 3x3 convolution block to obtain the intermediate feature F m ; Next, the intermediate feature F m Perform double upsampling again and output the fusion features of the first layer with the fusion module The first layer of features output by the encoder module The intermediate feature F is obtained by concatenating the layers in the channel dimension and processing them through a 3x3 convolution block. t ; Finally, the intermediate feature F t Through a 3x3 convolution block, a 3×3 convolution layer, and then the Sigmoid function to output the final fusion Y channel I f , the specific formula is as follows: I f =Sigmoid(Conv3(ConvB3((F t ))) in, Up scale=2 (·) represents the upsampling operation, which uses bilinear interpolation to double the width and height of the feature map; ASPP(·) is the atrous spatial pyramid pooling operation; Among them, C4=128, C5=512, C6=384, C7=256, and Sigmoid is the activation function operation.

7. The infrared and visible light image fusion method based on cross-domain Transformer according to claim 3 is characterized in that: The specific implementation steps of step C are as follows: The total loss function is defined to consist of four parts: strength loss Gradient Loss Structural similarity loss and contrast loss The specific formula is as follows Among them, λ i represents the weight parameter, i∈{1,2,3,4}; Strength loss The specific formula is: Among them, I v , I r are the Y channel of the visible light image and the infrared image input in step B1, H represents the height, W represents the width, max{·,·} represents the maximum value operation, ∑ represents the sum operation of all pixel positions, and |·,·| represents the mean absolute error, i.e., L1 loss; Gradient Loss The specific formula is: Where G(·) represents the spatial gradient of the calculated image; Structural similarity loss The specific formula is: The structural similarity index SSIM is used to measure the similarity between the fused Y channel and the Y channels of the infrared image and the visible light image. SSIM(·,·) represents the structural similarity index calculation operation; Contrastive loss Based on a contrastive learning mechanism with a saliency binary mask, weighted cosine similarity is calculated for "salient regions" and "non-salient regions" respectively. The salient regions are represented by a mask M, which is obtained by performing salient target detection on infrared images. The detected salient targets are regarded as salient regions, and other regions are regarded as non-salient regions. For the salient region, namely the mask M: Positive contrast loss for infrared images: Negative sample contrast loss of the Y channel of the visible light image: Where weighted_sim(·,·,·) is the weighted average of the cosine similarity calculated in the corresponding area; For non-salient regions (1-M): Positive sample contrast loss of the Y channel of the visible light image: Negative contrast loss for infrared images: In summary, the contrast loss The specific formula is as follows:

8. The infrared and visible light image fusion method based on cross-domain Transformer according to claim 1 is characterized in that: The specific implementation steps of step D are as follows: Step D1: randomly divide the infrared image and visible light image fusion training dataset constructed in step A into several batches, each batch containing N pairs of aligned infrared image and visible light image inputs; Step D2: For each batch of image pairs, input the infrared image and the visible light image into the image fusion network described in step B to obtain a fused image; Step D3: Based on the total loss function defined in step C, the gradient of the network parameters is calculated using the backpropagation method, and the full network parameters are updated in combination with the Adam optimizer. Step D4: Repeat steps D1 to D3 in batches until the loss function converges or the preset number of training rounds is met, and save the image fusion network parameters obtained from the final training as the trained infrared and visible light image fusion model.

9. A cross-domain Transformer-based infrared and visible light image fusion system, characterized by: The method comprises a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the method steps according to any one of claims 1 to 8 can be implemented.

10. A computer-readable storage medium storing computer program instructions that can be executed by a processor, wherein when the processor executes the computer program instructions, the method steps according to any one of claims 1 to 8 can be implemented.

Citation Information

Cited By

  • Infrared-visible light image fusion method based on pseudo twin network

    CN121121382A

  • An Infrared-Visible Image Fusion Method Based on Pseudo-Twin Networks

    CN121121382B

  • Infrared and visible light image fusion method and device based on lightweight model

    CN121329792A

  • Improvement-based YOLOV8 steel wire rope damage detection method

    CN121353746A

  • Remote sensing image target detection method, system and program product based on dual-drive multi-mode fusion network

    CN121353934A