Infrared and visible light image fusion method based on cross focusing linear attention
By adopting an image fusion method based on cross-focusing linear attention, this paper solves the problem of modal feature interaction in existing image fusion techniques, realizes image fusion technology, generates semantically rich fused images, improves image fusion technology, generates semantically rich image fusion methods, solves the problem of insufficient modal feature interaction in existing image fusion methods, and realizes the sufficiency of feature interaction and the preservation of high-frequency texture information in the image fusion process.
Patent Information
- Application Number
- CN202511079249.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-02
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-02
AI Technical Summary
Existing infrared and visible light image fusion methods fail to fully utilize the long-distance dependence and cross-modal feature interaction between the two modes, resulting in the loss of detailed information in the fused image.
An image fusion method based on cross-focusing linear attention is adopted, which includes a four-level encoder and decoder. It is trained by combining a dual-branch feature extraction module, an adaptive feature correction module, a cross-focusing linear attention fusion module, and a frequency-aware feature aggregation module with a joint loss function to achieve feature correction and training, optimize model parameters, and realize image fusion.
By employing a cross-focused linear attention approach, we achieved sufficient feature interaction and preservation of high-frequency texture information during image fusion, generating semantically rich fused images and improving image usability and visual quality.
Smart Images

Figure CN120976036A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, specifically to a method for fusing infrared and visible light images based on cross-focusing linear attention. Background Technology
[0002] Image fusion is a technique that integrates complementary information from multiple source images into a single output image. Its goal is to generate an image that is richer in information and has higher perceptual value, possessing stronger expressive power than any single source image. Infrared and visible light image fusion, as an important branch of image fusion, primarily focuses on fusing thermal infrared images with ordinary visible light images. Infrared cameras, based on the principle of thermal radiation, can highlight targets with thermal characteristics even at night, in obstructed, or harsh environments, while visible light cameras can provide high-resolution texture and color details under good lighting conditions. By fusing infrared and visible light images, a composite image can be obtained that simultaneously possesses high-contrast thermal target information and rich detail. This fused image significantly enhances situational awareness and is widely used in applications such as night vision, search and rescue, perimeter protection, and intelligent transportation. For example, infrared and visible light image fusion systems, by integrating complementary information from the infrared thermal channel and the visible light channel, can reveal the thermal radiation characteristics of concealed or camouflaged targets and provide semantic context of the scene, thereby improving the accuracy of target detection and recognition. However, infrared images typically have low spatial resolution, insufficient texture information, and are susceptible to noise; while visible light images perform poorly under low light or strong glare conditions. Therefore, a successful infrared and visible light image fusion method needs to effectively preserve the characteristics of thermal targets while taking into account the texture details and contrast of visible light, and properly handle noise differences and data alignment issues between different modalities.
[0003] Existing infrared and visible light image fusion methods often neglect the interrelationships and long-distance dependency modeling between the two modalities, resulting in insufficient feature interaction between the two modalities and causing the fused image to lose detailed information. Summary of the Invention
[0004] The purpose of this invention is to provide an infrared and visible light image fusion method based on cross-focusing linear attention, which solves the problems of insufficient long-distance dependency modeling, insufficient cross-modal feature interaction, and loss of high-frequency texture information in existing infrared and visible light image fusion methods. This invention can make full use of the information between the two modalities to achieve sufficient feature interaction, prevent the loss of modal information, and make full use of high-frequency texture detail information, thereby obtaining a semantically rich fused image. At the same time, it has an important promoting effect on downstream tasks, such as image segmentation.
[0005] The technical solution adopted in this invention is: an infrared and visible light image fusion method based on cross-focusing linear attention, comprising:
[0006] Step 1: Construct an image fusion network, including a four-level encoder and decoder. The encoder includes a dual-branch feature extraction module, an adaptive feature correction module, and a cross-focusing linear attention fusion module. The decoder includes a frequency-aware feature aggregation module and an upsampling module.
[0007] Step 2: Preprocess the acquired infrared and visible light image datasets;
[0008] Step 3: Train the constructed image fusion network using the designed joint loss function, optimize the model parameters, and retain the optimal model parameters to obtain the trained image fusion network;
[0009] Step 4: Input the preprocessed infrared and visible light images into the trained image fusion network to obtain the image fusion result.
[0010] Step 4 specifically involves:
[0011] The preprocessed infrared and visible light images are respectively input into the dual-branch feature extraction module for feature extraction, and infrared and visible light image features are obtained.
[0012] The obtained infrared and visible light features are used as input to the adaptive feature correction module for feature correction.
[0013] The feature-corrected infrared and visible light features are processed through a cross-focusing linear attention feature fusion module to obtain fused features;
[0014] The fused features obtained from the four-level encoder are input into the frequency-aware feature aggregation module for feature aggregation. Finally, the aggregated features are upsampled to restore the feature size and obtain the final fused image.
[0015] Specifically, the dual-branch feature extraction module includes multiple extraction blocks; each extraction block includes overlapping block embedding, Flatten Transformer, and downsampling;
[0016] The overlapping block embedding is used to map the input image to initial transformed features through an overlapping convolution operation:
[0017]
[0018] Where X0 is the initial transformation feature; OverlapEmbeding(*) is the overlapping block operation on the input image; input is the input image; H and W are the height and width of the input image, respectively; C0 is the number of feature channels of X0; R is the real number space;
[0019] The Flatten Transformer is used to perform Flatten Transformer operations on features at each scale, including aggregated linear attention and a hybrid feedforward network:
[0020]
[0021] Among them, X i C represents the output of the i-th Flatten Transformer. i This represents the number of channels of the feature output of the i-th FlattenTransformer;
[0022] Flatten Transformer consists of a stack of multiple submodules:
[0023]
[0024] in, This represents the output of the Flatten Transformer. The input to the Flatten Transformer is represented by MFA(·), which stands for Multi-Head Focus Linear-Attention. MFA is responsible for performing linear self-attention calculation on the normalized features and residual branch features to capture global long-range dependencies. LN(·) represents Layer Norm, F... SA The self-attention output and input residual are aggregated as an intermediate representation, and after passing through FTB, they are passed through a downsampling module to halve the feature size;
[0025] The downsampling is used to perform overlapping downsampling of mixed features:
[0026] X i Att =OverlapPatchMerging(X i ), i = 1, 2, 3;
[0027] Among them, X i Att The attention features are for the downsampled output; OverlapPatchMerging(*) is the downsampling process;
[0028] Furthermore, the focused linear attention is as follows:
[0029] Q = Sim(Q,K) V = φ P (Q)φ P (K) T V+DWC(V);
[0030] φ p (x)=f p (ReLU(x)),
[0031] Where O is the output linear attention feature; Q, K, and V are the Query, Key, and Value used for attention calculation; φ p (·) is a mapping function; x **P represents element-wise exponentiation, where P is a hyperparameter controlling the degree of feature focusing; Sim(Q,K) represents similarity calculation after focusing function mapping; DWC represents depthwise convolution; (K) T This indicates that K is transposed; x represents the function's independent variables, here referring to Q and (K). T ; RELU(x) represents the ReLU activation function; f p (x) represents the power normalization function; ||·|| represents the L2 norm operation.
[0032] The adaptive feature correction module specifically includes channel-wise correction and spatial-wise correction, with the further channel-wise correction specifically expressed as follows:
[0033]
[0034] in, Let F represent the channel correction weights for infrared and visible light images, respectively. MLP(*) represents the multilayer perceptron, σ(*) represents the sigmoid activation function, and F... split (*) indicates that the weights are divided into infrared and visible light weights, and Y represents the stitched infrared and visible light features used to generate channel correction weights. This indicates the visible light characteristics after Detail-Preserving Pooling. This represents the visible light characteristics after spectral pooling. This indicates the infrared features after detail-preserving pooling. The symbol represents the infrared features after spectral pooling, and || represents the splicing operation.
[0035]
[0036] in, and Used for subsequent correction of visible light and infrared image channel correction components, IR in and VIS in These are the infrared and visible light inputs for the feature correction module, respectively.
[0037] Furthermore, the infrared and visible light features are first concatenated, and then processed through convolution and activation functions to obtain an intermediate fused feature map F, which is used to subsequently obtain a spatial correction weight map, thus obtaining weight maps for visible light and infrared light in the spatial dimension. and The specific formula is as follows:
[0038]
[0039] F = Conv(RELU(Conv(IR) in ||VIS in )
[0040] Similar to channel correction, spatial correction is represented as
[0041]
[0042] in, and Used for subsequent correction of the visible light and infrared image channel correction components;
[0043] The entire feature correction module is specifically represented as follows, λ C and λ S The relative weights used for automatically adjusting the channel domain and spatial domain corrections are all set to 0.5;
[0044]
[0045] Among them, VIS out and IR out This represents the visible light and infrared features after adaptive feature correction.
[0046] The cross-focusing linear attention feature fusion module is specifically represented as follows:
[0047] G Vis =φ p (Q Ir )φ p (K Vis ) T
[0048] G Ir =φ p (Q Vis )φ p (K Ir ) T
[0049] Among them, G Vis and G Ir Let φ represent the cross-attention matrices for infrared and visible light, respectively. p (*) denotes the mapping function, Q Ir and QVis The tables represent the infrared and visible light queries used for attention calculation, K. Ir and K Vis These represent the infrared and visible light keys used for attention calculation, respectively, and (*)T represents the transpose operation;
[0050] Based on this, the cross-aggregation linear attention result U Vis and U Ir Represented as:
[0051] U Vis =LN(G Vis V Vis +DWC(V Vis )+X Vis )
[0052] U Ir =LN(G Ir V Ir +DWC(V Ir )+X Ir )
[0053] Where LN(*) represents LayerNorm, DWC(*) represents depthwise separable convolution, and V Vis and V Ir Value X represents the visible light and infrared values used for attention calculation. Vis and X Ir These represent the visible light and infrared features input to the cross-focusing linear attention feature fusion module, respectively.
[0054] Two output U Vis and U Ir F is obtained by splicing along the channel dimension cat After global interaction, a local fusion is performed again. The feature fusion network based on cross-focusing linear attention is represented as follows:
[0055] F cat =Concat(U Vis U Ir )
[0056] F out =LN(F cat +Conv(DWC(Conv(F cat ))))
[0057] Among them, F out This represents the output of the cross-focusing linear attention feature fusion module.
[0058] Specifically, the frequency-aware feature aggregation module is represented as follows:
[0059] f_2 = FreqFusion(f4, f3)
[0060] f_1 = FreqFusion(f_2, f2)
[0061] Fused = FreqFusion(f_1, f1)
[0062] Where f4, f3, f2, f1 represent the output of the four-level encoder, f_2, f_1 represent the intermediate aggregation results, Fused represents the final aggregation result, and FreqFusion(*) represents the frequency-aware feature fusion operation.
[0063] Specifically, FreqFusion(*) is divided into two phases of fusion, denoted as initial fusion and final fusion, where the initial fusion is... Init This is represented as X after upsampling. l+1 and the high-pass filtered result generated by the Adaptive High-Pass Filter Generator (AHPF) Element-by-element addition:
[0064]
[0065] Specifically, the enhanced high-frequency characteristics of the adaptive high-pass filter (AHPF) are represented as follows:
[0066] V l =Conv 3×3 (X l )
[0067] W l =E-Softmax(V l )
[0068]
[0069] Among them, X l V represents the output of the l-th stage encoder. l For the compressed features, Softmax(*) represents the Softmax activation function, E represents the unit kernel, and W... l Represents the high-pass filter weights
[0070] The obtained initial fusion features Fusion Init Used as a high-pass filter for AHPF generation in the second stage for high-frequency texture enhancement, the final fused representation is as follows:
[0071] Z l =Conv 3×3 (Fusion Init )
[0072]
[0073]
[0074] Where Up(*) represents the upsampling operation, Fusion Final This represents the final fusion result of FreqFusion.
[0075] Step 3 specifically includes:
[0076] The constructed image fusion network is trained using the designed joint loss function to optimize its parameters. The specific loss function is as follows:
[0077] L total =λ1L ssim +λ2L grad +λ3L int
[0078] Where λ1, λ2, and λ3 are used to control the balance of the three sub-losses in the joint loss, where L ssim L represents the structural similarity loss. grad L represents the gradient loss. int Indicates pixel intensity loss;
[0079] Among them, structural similarity loss L ssim Used to maintain the fused image I f and source image (I v ,I i The structural similarity between them is specifically represented as follows:
[0080] L ssim =α(1-ssim(I) f ,I i ))+β(1-ssim(I f ,I v ))
[0081] Where α and β represent balance factors, and I f Represents the fused image, (I v ,I i () represent visible light source images and infrared source images, respectively;
[0082] Meanwhile, the gradient loss is specifically expressed as
[0083]
[0084] in, The Sobel gradient operator is represented, max(·) represents element-wise maximum selection, ||·|| represents L1 norm operation, H and W represent the length and width of the fused image, and HW represents their product.
[0085] Pixel intensity loss was added, specifically expressed as
[0086]
[0087] Here, mean(·) represents the element-wise averaging operation.
[0088] The beneficial effects of this invention are:
[0089] This network efficiently integrates complementary information from heterogeneous modalities. Specifically, the encoder comprises four cascaded stages, each consisting of a Flatten Transformer Block (FTB), a Feature Correction Module (FRM), and a feature fusion module based on cross-focusing linear attention. The FTB extracts global and local contextual features, while the FRM eliminates inconsistencies in channel distribution and spatial alignment through channel and spatial correction mechanisms, significantly improving feature consistency. Furthermore, a cross-focusing linear attention mechanism is introduced into the feature fusion module to promote deep interaction between cross-modal features, thereby enhancing feature representation capabilities. Subsequently, the decoder dynamically enhances edge details and high-frequency textures through a frequency-aware feature aggregation module, gradually fusing multi-scale semantic information and effectively preserving significant hot targets and visible details. This effectively addresses the issues of insufficient utilization of long-range contextual information, inadequate utilization of the correlation between two modal features, and insufficient feature interaction, significantly improving image usability and visual quality. It exhibits excellent robustness and generalization ability, highlighting its potential in advanced vision applications. Attached Figure Description
[0090] Figure 1 A partial flowchart illustrating the infrared and visible light image fusion method based on cross-aggregation linear attention provided in an embodiment of the present invention;
[0091] Figure 2 This is a schematic diagram of the fusion network structure in an embodiment of the present invention;
[0092] Figure 3 This is a schematic diagram of the adaptive feature correction module provided in an embodiment of the present invention;
[0093] Figure 4 This is a schematic diagram of the structure of the frequency sensing feature aggregation module provided in an embodiment of the present invention;
[0094] Figure 5 This is a schematic diagram comparing the fused images generated by the method of the present invention with those generated by other methods. Detailed Implementation
[0095] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0096] Example 1: As Figure 1 As shown, an infrared and visible light image fusion network based on cross-aggregation linear attention includes the following:
[0097] Step 1: Construct an image fusion network, such as Figure 2 As shown, it includes a four-level encoder and decoder. The encoder includes a dual-branch feature extraction module, an adaptive feature correction module, and a cross-focusing linear attention fusion module. The decoder includes a frequency-aware feature aggregation module and an upsampling module. The upsampling module is used to restore the features to their original size.
[0098] Step 2: Preprocess the acquired infrared and visible light image datasets;
[0099] Step 3: Train the constructed image fusion network using the designed joint loss function, optimize the model parameters, and retain the optimal model parameters as the trained image fusion network for final image fusion.
[0100] Step 4: Input the infrared and visible light images into the trained image fusion network to obtain the image fusion result.
[0101] Furthermore, Step 3 specifically includes:
[0102] To achieve end-to-end unsupervised training and thus optimize the fusion model, a joint loss function was designed, the specific loss function of which is as follows:
[0103] L total =λ1L ssim +λ2L grad +λ3L int
[0104] Where λ1, λ2, and λ3 are used to control the balance of the three sub-losses in the joint loss, and are set as: 3, 50, 7, where L ssim L represents the structural similarity loss. grad L represents the gradient loss. int This indicates pixel intensity loss.
[0105] Among them, structural similarity loss L ssim Used to maintain the fused image I f and source image (Iv ,I i The structural similarity between them can be specifically represented as:
[0106] L ssim =α(1-ssim(I) f ,I i ))+β(1-ssim(I f ,I v ))
[0107] Where α and β represent balancing factors used to balance the influence of source images on the fusion result, both set to 0.5, indicating that infrared and visible light images contribute equally to structural maintenance. f Represents the fused image, (I v ,I i () represent visible light source images and infrared source images, respectively.
[0108] At the same time, gradient loss is used to preserve as much texture detail as possible in the fused image. Specifically, gradient loss can be expressed as...
[0109]
[0110] in, This represents the Sobel gradient operator, and max(·) indicates element-wise maximum selection. ||·|| represents the L1 norm operation. H and W represent the length and width of the input image, and HW represents their product.
[0111] Furthermore, to ensure the network retains more valuable pixel intensity information, a pixel intensity loss is incorporated, which can be specifically expressed as...
[0112]
[0113] Here, mean(·) represents the element-wise averaging operation.
[0114] Furthermore, Step 4 specifically includes:
[0115] The preprocessed infrared and visible light images are respectively input into the dual-branch feature extraction module for feature extraction, and infrared and visible light image features are obtained.
[0116] The obtained infrared and visible light features are used as input to the adaptive feature correction module for feature correction.
[0117] The feature-corrected infrared and visible light features are processed through a cross-focusing linear attention feature fusion module to obtain fused features;
[0118] The fused features obtained from the four-level encoder are input into the frequency-aware feature aggregation module for feature aggregation. Finally, the aggregated features are upsampled to restore the feature size and obtain the final fused image.
[0119] Furthermore, each of the dual-branch extraction modules includes multiple extraction blocks; each extraction block includes overlapping block embedding, Flatten Transformer, and downsampling;
[0120] The overlapping block embedding is used to map the input image to initial transformed features through an overlapping convolution operation:
[0121]
[0122] Where X0 is the initial transformation feature; OverlapEmbeding(*) is the overlapping block operation on the input image; input is the input image; H and W are the height and width of the input image, respectively; C0 is the number of feature channels of X0; R is the real number space;
[0123] The Flatten Transformer is used to perform Flatten Transformer operations on features at each scale, including aggregated linear attention and a hybrid feedforward network:
[0124]
[0125] Among them, X i C represents the output of the i-th Flatten Transformer. i This represents the number of channels of the feature output of the i-th FlattenTransformer.
[0126] Flatten Transformer consists of a stack of multiple submodules:
[0127]
[0128] in, This represents the output of the Flatten Transformer. This represents the input to the Flatten Transformer. MFA(·) denotes Multi-Head Focus Linear-Attention, which is responsible for performing linear self-attention calculation on the normalized features and residual branch features to capture global long-range dependencies. LN(·) denotes Layer Norm,F SA The self-attention output and input residuals are aggregated as an intermediate representation. After passing through FTB, a downsampling module is used to halve the feature size.
[0129] The downsampling is used to perform overlapping downsampling of mixed features:
[0130] X i Att =OverlapPatchMerging(X i ), i = 1, 2, 3;
[0131] Among them, X i Att The attention features are the downsampled output; OverlapPatchMerging(*) is the downsampling process.
[0132] Furthermore, the focused linear attention is as follows:
[0133] O=Sim(Q,K)V=φ P (Q)φ P (K) T V+DWC(V);
[0134] φ p (x)=f p (ReLU(x)),
[0135] Where O is the output linear attention feature; Q, K, and V are the Query, Key, and Value used for attention calculation; φ p (·) is a mapping function; x **P represents element-wise exponentiation, where P is a hyperparameter controlling the degree of feature focusing; Sim(Q,K) represents similarity calculation after focusing function mapping; DWC represents depthwise convolution; (K) T This indicates that K is transposed; x represents the function's independent variables, here referring to Q and (K). T ; RELU(x) represents the ReLU activation function; f p (x) represents the power normalization function; ||·|| represents the L2 norm operation.
[0136] Considering that information between different modalities is often complementary, and in order to better utilize the complementary nature of information between two modalities, this invention designs an adaptive feature correction module, such as... Figure 3As shown, feature correction is performed using both channel and spatial methods to achieve better modal information interaction. For channel-by-channel correction, infrared and visible light features are embedded into two attention vectors along the spatial axis, respectively. Unlike previous channel attention methods, detail-preserving pooling (DPP) and spectral pooling operations are performed on infrared and visible light features along the channel dimension, respectively. This operation adaptively balances max pooling and average pooling, preserving important high-frequency "detail" information while achieving a certain degree of smoothness to reduce noise interference. Let G be the result of the two pooling methods. DDP With G Spect Then, they are concatenated along the channel dimension: this preserves key high-frequency details while also smoothing the surface to retain more texture information. Subsequently, channel correction weights are generated using multilayer perceptron mapping (MLP) and a sigmoid activation function, and then split into two sets of channel weights for visible light and infrared light. and Specifically, it can be expressed as:
[0137]
[0138] in, Let F represent the channel correction weights for infrared and visible light images, respectively. MLP(*) represents the multilayer perceptron, σ(*) represents the sigmoid activation function, and F... split (*) indicates that the weights are divided into infrared and visible light weights, and Y represents the stitched infrared and visible light features used to generate channel correction weights. This indicates the visible light characteristics after Detail-Preserving Pooling (DPP). This represents the visible light characteristics after spectral pooling. This indicates the infrared features after detail-preserving pooling. represents the infrared features after spectral pooling, and || represents the splicing operation;
[0139]
[0140] in, and Used for subsequent correction of visible light and infrared image channel correction components, IR in and VIS in These are the infrared and visible light inputs for the feature correction module, respectively.
[0141] Furthermore, the infrared and visible light features are first concatenated, and then processed through convolution and activation functions to obtain an intermediate fused feature map F, which is used to subsequently obtain a spatial correction weight map, thus obtaining weight maps for visible light and infrared light in the spatial dimension. and The specific formula is as follows:
[0142]
[0143] F = Conv(RELU(Conv(IR) in ||VIS in )
[0144] Similar to channel correction, spatial correction can be expressed as
[0145]
[0146] in, and Used for subsequent correction of the visible light and infrared image channel correction components.
[0147] The entire feature correction module can be specifically represented as follows, λ C and λ S The relative weights used for automatically adjusting the channel domain and spatial domain corrections are all set to 0.5 here;
[0148]
[0149] Among them, VIS out and IR out This represents the visible light and infrared features after adaptive feature correction.
[0150] To enhance information interaction and fusion between the two modalities, this invention designs a cross-focusing linear attention module to fuse infrared and visible light features, enabling the network to explicitly model the interrelationships between the features of the two modalities. Most traditional cross-attention mechanisms use traditional self-attention mechanisms to model the two modalities; however, the computational complexity of traditional self-attention mechanisms increases quadratically with sequence length, leading to high computational costs when using attention mechanisms with global receptive fields. Therefore, this invention designs a cross-focusing linear attention feature fusion module to fuse the features of the two modalities, specifically as follows:
[0151] G Vis =φ p (Q Ir )φ p (K Vis ) T
[0152] GIr =φ p (Q Vis )φ p (K Ir ) T
[0153] Among them, G Vis and G Ir Let φ represent the cross-attention matrices for infrared and visible light, respectively. p (*) denotes the mapping function, Q Ir and Q Vis K represents the infrared and visible light queries used for attention calculation, respectively. Ir and K Vis These represent the infrared and visible light keys used for attention calculations, respectively. (*) T This indicates the transpose operation.
[0154] Based on this, the cross-aggregation linear attention result U Vis and U Ir It can be represented as:
[0155] U Vis =LN(G Vis V Vis +DWC(V Vis )+X Vis )
[0156] U Ir =LN(G Ir V Ir +DWC(V Ir )+X Ir )
[0157] Where LN(*) represents LayerNorm, DWC(*) represents depthwise separable convolution, and V Vis and V Ir Value X represents the visible light and infrared values used for attention calculation. Vis and X Ir These represent the visible light and infrared features input to the cross-focusing linear attention feature fusion module, respectively.
[0158] Two output U Vis and U Ir F is obtained by splicing along the channel dimension cat After global interaction, a local fusion is performed again, which not only fully mixes the concatenated high-dimensional features but also preserves details and spatial coherence through separable convolution. The feature fusion network based on cross-focusing linear attention can be represented as:
[0159] F cat =Concat(U Vis UIr )
[0160] F out =LN(F cat +Conv(DWC9Conv(F cat ))))
[0161] Among them, F out This represents the output of the cross-focusing linear attention feature fusion module.
[0162] This mechanism allows for the effective combination of thermal targets in infrared images and detailed textures in visible light images in the fusion result, enhancing key information while avoiding information omissions or conflicts that may result from simple fusion.
[0163] Existing multi-scale feature fusion methods mostly employ simple stitching, which leads to insufficient utilization of high-frequency and low-frequency information from various stages, resulting in information loss. Therefore, this invention designs a frequency-aware feature aggregation module to aggregate features from various stages of a four-level encoder. The frequency-aware feature aggregation module consists of three layers of frequency-aware feature fusion modules, such as... Figure 4 As shown, the specific operation of the frequency-aware feature aggregation module can be represented as follows:
[0164] f_2 = FreqFusion(f4, f3)
[0165] f_1 = FreqFusion(f_2, f2)
[0166] Fused = FreqFusion(f_1, f1)
[0167] Where f4, f3, f2, f1 represent the output of the four-level encoder, f_2, f_1 represent the intermediate aggregation results, Fused represents the final aggregation result, and FreqFusion(*) represents the frequency-aware feature fusion operation.
[0168] Specifically, FreqFusion(*) can be roughly divided into two stages of fusion, denoted as initial fusion and final fusion, where the initial fusion is... Init This can be roughly expressed as upsampling X l+1 And the high-pass filtered result generated by the Adaptive High-Pass Filter Generator (AHPF) Element-by-element addition:
[0169]
[0170] Specifically, the enhanced high-frequency characteristics of the adaptive high-pass filter (AHPF) can be expressed as follows:
[0171] V l=Conv 3×3 (X l )
[0172] W l =E-Softamx(V l )
[0173]
[0174] Among them, X l V represents the output of the l-th stage encoder. l For the compressed features, Softmax(*) represents the Softmax activation function, E represents the unit kernel, and W... l Represents the high-pass filter weights
[0175] The obtained initial fusion features Fusion Init The high-pass filter used as the second-stage AHPF generation high-pass filter for high-frequency texture enhancement, and the final fusion can be expressed as:
[0176] Z l =Conv 3×3 (Fusion Init )
[0177]
[0178]
[0179] Where Up(*) represents the upsampling operation, Fusion Final This indicates the final fusion result of FreqFusion.
[0180] Two high-pass filters are used to repeatedly enhance the high-frequency information of different layers, avoiding the loss of details due to simple splicing.
[0181] To further verify the feasibility of this method, a qualitative and quantitative comparative analysis was conducted between the method described in Example 1 and nine state-of-the-art methods: DenseFuse, DIDFuse, FusionGAN, U2Fusion, SwinFuse, SwinFusion, DATFuse, CDDFuse, and MBHFuse. MBHFuse was reproduced using its source code. All comparative experiments were implemented according to the requirements of the original paper. The model of this scheme was trained using the TNO dataset, and experiments were conducted on the MSRS and LLVIP datasets, two mainstream infrared and visible light image fusion datasets. During the training phase, preprocessing was performed first, with training samples cropped to 128×128 pixel blocks at a stride of 200 to obtain a sufficient number of images for model training. The training batch size was 120 epochs with a batch size of 8. The initial learning rate was set to 0.0001, employing a multi-learning rate decay strategy and using the Adam optimizer to adaptively adjust the learning rate.
[0182] like Figure 5 As shown in the experimental results, the proposed scheme can generate fused images that effectively preserve both low-frequency global features and high-frequency local features. This is due to the adaptive feature correction module fully mining the complementary information of the source images, while the cross-aggregation linear attention feature fusion module dynamically extracts representative feature representations. This enables a deeper mining of latent features when generating fused images, providing a more comprehensive description of the imaging scene and generating fused images that are rich in detail, globally consistent, high in contrast, and comprehensively integrated in information.
[0183] Quantitative comparative analysis experiments were conducted, using commonly used fusion metrics in image fusion tasks for evaluation, including entropy (EN), standard deviation (SD), spatial frequency (SF), mutual information (MI), sum of differential correlations (SCD), visual information fidelity (VIF), edge-based metrics (Qab / f), and average gradient (AG). These metrics are positive indicators; higher values indicate better fused image results, as shown in Tables 1 and 2. Table 1 presents the MSRS quantitative comparison results, and Table 2 presents the LLVIP quantitative comparison results.
[0184]
[0185] Table 1
[0186]
[0187] Table 2
[0188] The results show that our method has significant advantages in most metrics, with SD, MI and VIF being significantly better than the comparison methods. This indicates that our method can depict the details of the target scene more richly and retain more comprehensive image information, especially in terms of scene details and edge features.
[0189] This invention fully utilizes the interactive information between two modalities to obtain semantically rich fused images. The network includes a feature extraction network, an attention-based adaptive feature correction module, and a frequency-aware feature aggregation module. Another innovation is the design of a cross-attention fusion module based on Focus Linear attention, which can adaptively fuse infrared thermal radiation information and visible light image texture information, thus better realizing the information interaction between the two features.
[0190] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A method for fusion of infrared and visible light images based on cross-focusing linear attention, characterized in that, include: Step 1: Construct an image fusion network, including a four-level encoder and decoder. The encoder includes a dual-branch feature extraction module. The system includes an adaptive feature correction module, a cross-focusing linear attention fusion module, and a decoder comprising a frequency-aware feature aggregation module and an upsampling module. Step 2: Preprocess the acquired infrared and visible light image datasets; Step 3: Train the constructed image fusion network using the designed joint loss function, optimize the model parameters, and retain the optimal model parameters to obtain the trained image fusion network; Step 4: Input the preprocessed infrared and visible light images into the trained image fusion network to obtain the image fusion result.
2. The infrared and visible light image fusion method based on cross-focusing linear attention as described in claim 1, characterized in that, Step 4 specifically involves: The preprocessed infrared and visible light images are respectively input into the dual-branch feature extraction module for feature extraction, and infrared and visible light image features are obtained. The obtained infrared and visible light features are used as input to the adaptive feature correction module for feature correction. The feature-corrected infrared and visible light features are processed through a cross-focusing linear attention feature fusion module to obtain fused features; The fused features obtained from the four-level encoder are input into the frequency-aware feature aggregation module for feature aggregation. Finally, the aggregated features are upsampled to restore the feature size and obtain the final fused image.
3. The infrared and visible light image fusion method based on cross-focusing linear attention according to claim 2, characterized in that, The dual-branch feature extraction module includes multiple extraction blocks; each extraction block includes overlapping block embedding, Flatten Transformer, and downsampling. The overlapping block embedding is used to map the input image to initial transformed features through an overlapping convolution operation: Where X0 is the initial transformation feature; OverlapEmbeding(*) is the overlapping block operation on the input image; input is the input image; H and W are the height and width of the input image, respectively; C0 is the number of feature channels of X0; R is the real number space; The Flatten Transformer is used to perform Flatten Transformer operations on features at each scale, including aggregated linear attention and a hybrid feedforward network: Among them, X i C represents the output of the i-th Flatten Transformer. i This represents the number of channels of the feature output of the i-th Flatten Transformer; Flatten Transformer consists of a stack of multiple submodules: in, This represents the output of the Flatten Transformer. The input to the Flatten Transformer is represented by MFA(·), which stands for Multi-Head Focus Linear-Attention. MFA is responsible for performing linear self-attention calculations on the normalized features and residual branch features to capture global long-range dependencies. LN(·) represents LayerNorm,F SA The self-attention output and input residual are aggregated as an intermediate representation, and after passing through FTB, they are passed through a downsampling module to halve the feature size; The downsampling is used to perform overlapping downsampling of mixed features: X i Att =OverlapPatchMerging(X i ),i=1,2,3; Among them, X i Att The attention features are for the downsampled output; OVerlapPatchMerging(*) is the downsampling process; Furthermore, the focused linear attention is as follows: O=Sim(Q,K)V=φ P (Q)φ P (K) T V+DWC(V); Where O is the output linear attention feature; Q, K, and V are the Query, Key, and Value used for attention calculation; φ p (·) is a mapping function; x **P represents element-wise exponentiation, where P is a hyperparameter controlling the degree of feature focusing; Sim(Q,K) represents similarity calculation after focusing function mapping; DWC represents depthwise convolution; (K) T This indicates that K is transposed; x represents the function's independent variables, here referring to Q and (K). T ; RELU(x) represents the ReLU activation function; f p (x) represents the power normalization function; ||·|| represents the L2 norm operation.
4. The infrared and visible light image fusion method based on cross-focusing linear attention according to claim 2, characterized in that, The adaptive feature correction module specifically includes channel-wise correction and spatial-wise correction, with the further channel-wise correction specifically expressed as follows: in, Let F represent the channel correction weights for infrared and visible light images, respectively. MLP(*) represents the multilayer perceptron, σ(*) represents the sigmoid activation function, and F... split (*) indicates that the weights are divided into infrared and visible light weights, and Y represents the stitched infrared and visible light features used to generate channel correction weights. This indicates the visible light characteristics after Detail-Preserving Pooling. This represents the visible light characteristics after spectral pooling. This indicates the infrared features after detail-preserving pooling. The symbol represents the infrared features after spectral pooling, and || represents the splicing operation. in, and Used for subsequent correction of visible light and infrared image channel correction components, IR in and VIS in These are the infrared and visible light inputs for the feature correction module, respectively. Furthermore, the infrared and visible light features are first concatenated, and then processed through convolution and activation functions to obtain an intermediate fused feature map F, which is used to subsequently obtain a spatial correction weight map, thus obtaining weight maps for visible light and infrared light in the spatial dimension. and The specific formula is as follows: F=Conv(RELU(Conv(IR in ||VIS in ) Similar to channel correction, spatial correction is represented as in, and Used for subsequent correction of the visible light and infrared image channel correction components; The entire feature correction module is specifically represented as follows, λ C and λ S Used to automatically adjust the relative weights of channel domain and spatial domain corrections. Among them, VIS out and IR out This represents the visible light and infrared features after adaptive feature correction.
5. The infrared and visible light image fusion method based on cross-focusing linear attention as described in claim 2, characterized in that, The cross-focusing linear attention feature fusion module is specifically represented as follows: G Vis =φ p (Q Ir )φ p (K Vis ) T G Ir =φ p (Q Vis )φ p (K Ir ) T Among them, G Vis and G Ir Let φ represent the cross-attention matrices for infrared and visible light, respectively. p (*) denotes the mapping function, Q Ir and Q Vis The tables represent the infrared and visible light queries used for attention calculation, K. Ir and K Vis These represent the infrared and visible light keys used for attention calculations, respectively. (*) T Indicates the transpose operation; Based on this, the cross-aggregation linear attention result U Vis and U Ir Represented as: U Vis =LN(G Vis V Vis +DWC(V Vis )+X Vis ) U Ir =LN(G Ir V Ir +DWC(V Ir )+X Ir ) Where LN(*) represents LayerNorm, DWC(*) represents depthwise separable convolution, and V Vis and V Ir Value X represents the visible light and infrared values used for attention calculation. Vis and X Ir These represent the visible light and infrared features input to the cross-focusing linear attention feature fusion module, respectively. Two output U Vis and U Ir F is obtained by splicing along the channel dimension cat After global interaction, a local fusion is performed again. The feature fusion network based on cross-focusing linear attention is represented as follows: F cat =Concat(U Vis ,U Ir ) F out =LN(F cat +Conv(DWC(Conv(F cat )))) Among them, F out This represents the output of the cross-focusing linear attention feature fusion module.
6. The infrared and visible light image fusion method based on cross-focusing linear attention as described in claim 2, characterized in that, The frequency-aware feature aggregation module is specifically represented as follows: f_2 = FreqFusion(f4, f3) f_1 = FreqFusion(f_2, f2) Fused = FreqFusion(f_1, f1) Where f4, f3, f2, f1 represent the output of the four-level encoder, f_2, f_1 represent the intermediate aggregation results, Fused represents the final aggregation result, and FreqFysion(*) represents the frequency-aware feature fusion operation. Specifically, FreqFusion(*) is divided into two phases of fusion, denoted as initial fusion and final fusion, where the initial fusion is... Init This is represented as X after upsampling. l+1 and the high-pass filtered result generated by the Adaptive High-Pass Filter Generator (AHPF) Element-by-element addition: Specifically, the enhanced high-frequency characteristics of the adaptive high-pass filter (AHPF) are represented as follows: V l =Conv 3×3 (X l ) W l =E-Softmax(V l ) Among them, X l V represents the output of the l-th stage encoder. l For the compressed features, Softmax(*) represents the Softmax activation function, E represents the unit kernel, and W... l Represents the high-pass filter weights The obtained initial fusion features Fusion Init Used as a high-pass filter for AHPF generation in the second stage for high-frequency texture enhancement, the final fused representation is as follows: Z l =Conv 3×3 (Fusion Init ) Where Up(*) represents the upsampling operation, Fusion Final This indicates the final fusion result of FreqFusion.
7. The infrared and visible light image fusion method based on cross-focusing linear attention according to claim 1, characterized in that, Step 3 specifically includes: The constructed image fusion network is trained using the designed joint loss function to optimize its parameters. The specific loss function is as follows: L total =λ1L ssim +λ2L grad +λ3L int Where λ1, λ2, and λ3 are used to control the balance of the three sub-losses in the joint loss, where L ssin L represents the structural similarity loss. grad L represents the gradient loss. int Indicates pixel intensity loss; Among them, structural similarity loss L ssim Used to maintain the fused image I f and source image (I v I i The structural similarity between them is specifically represented as follows: L ssim =α(1-ssim(I f ,I i ))+β(1-ssim(I f ,I v )) Where α and β represent balance factors, and I f Represents the fused image, (I v I i () represent visible light source images and infrared source images, respectively; Meanwhile, the gradient loss is specifically expressed as in, The Sobel gradient operator is represented, max(·) represents element-wise maximum selection, ||·|| represents L1 norm operation, H and W represent the length and width of the fused image, and HW represents their product. Pixel intensity loss was added, specifically expressed as Here, mean(·) represents the element-wise averaging operation.
Citation Information
Patent Citations
Image semantic segmentation method and system based on frequency perception feature fusion
CN116704180A
Semi-supervised medical image segmentation method based on data enhancement and focusing linear attention
CN118781137A
Hyperspectral image super-resolution reconstruction method based on double-branch network
CN119027314A
Cited By
Farm disease identification method and device based on multi-modal data fusion
CN121685464A
An infrared and visible light image fusion method and system based on linear attention
CN122390991A
An infrared and visible light image fusion method and system based on linear attention
CN122390991B