Infrared and visible image fusion method based on text semantic guidance and frequency domain compensation

CN122529995APending Publication Date: 2026-08-07SOUTHWEAT UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTHWEAT UNIV OF SCI & TECH
Filing Date
2026-07-10
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0006]针对现有技术中的上述不足,本发明提供的基于文本语义引导与频域补偿的红外可见光图像融合方法解决了多模态信息容易相互干扰、跨模态交互不足以及复杂环境下细节信息易丢失的问题

Benefits of technology

(1)本发明提出了基于文本语义引导与频域补偿的红外可见光图像融合方法,流程上采用编码–融合–解码的层级结构,通过双分支特征编码与空间上下文增强、跨模态交互融合与文本引导调制、频域细节增强与跨层传播三大核心机制,有效实现了红外目标热显著性与可见光结构纹理细节的稳定互补融合,解决了多模态信息容易相互干扰、跨模态交互不足以及复杂环境下细节信息易丢失的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122529995A_ABST
    Figure CN122529995A_ABST
Patent Text Reader

Abstract

The application discloses an infrared and visible light image fusion method based on text semantic guidance and frequency domain compensation, and relates to the technical field of infrared and visible light image fusion technology, and the method comprises the following steps: inputting an infrared image and a visible light image into a double-branch feature coding module, strengthening feature expression capability by adopting a spatial context enhancement strategy, and generating multi-scale features of the infrared image and the visible light image; inputting the multi-scale features of the infrared image and the visible light image into a cross-modal interaction fusion module, establishing a correlation between infrared and visible light features through an attention mechanism, and obtaining a fusion feature map; and inputting the fusion feature map into a frequency domain detail enhancement and cross-layer propagation decoding module, adaptively supplementing and deeply restoring high-frequency information, and generating a final fusion image. The method effectively realizes stable complementary fusion of infrared target thermal saliency and visible light structure texture details, and solves the problems of easy mutual interference of multi-modal information, insufficient cross-modal interaction and easy loss of detail information in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of infrared and visible light image fusion technology, specifically involving an infrared and visible light image fusion method based on text semantic guidance and frequency domain compensation. Background Technology

[0002] Image fusion is a key technology in the field of image processing. By effectively utilizing complementary information from multiple sensors, the limitations of single-modal sensors in information acquisition can be largely overcome. In this research area, Infrared-Visible Image Fusion (IVIF) has received widespread attention. IVIF fuses important thermal radiation information from infrared (IR) images with rich texture details from visible (VI) images, thereby providing a richer and more complete scene representation. By integrating information from different modalities and generating fusion results with good visual effects, IVIF has been widely applied in various fields such as security monitoring, military reconnaissance, scene understanding, and assisted driving.

[0003] In recent years, IVIF has gradually become a research hotspot, driving the rapid development of related methods. From the perspective of the network structure used, existing methods can be broadly classified into: convolutional neural network-based methods, autoencoder-based methods, generative adversarial network-based methods, Transformer-based methods, and diffusion model-based methods. Furthermore, from the perspective of functional objectives, these methods can be further divided into: visual quality-oriented methods, joint registration and fusion methods, semantic-driven methods, and degradation-robust methods.

[0004] While IVIF has made some progress, it still has significant limitations under complex environmental conditions. Degradation factors such as fog, rain, and low illumination significantly reduce the contrast of visible light images and weaken texture details, while infrared images, although able to stably reflect the thermal radiation information of targets, lack clear structural representation. When the two modalities are fused in the network, without a reasonable information organization mechanism, insufficient information utilization or mutual interference often occurs. For example, the salience of the target in the infrared image is weakened by the background texture during the fusion process, or details in the visible light image are gradually lost during feature transformation.

[0005] Infrared and visible light images inherently differ in imaging mechanisms and information distribution. If cross-modal interaction relies solely on simple feature stitching or weighted fusion, effective semantic associations between different modalities become difficult to form, significantly limiting the ability to express information in key regions. Furthermore, in adverse weather conditions, different scene conditions alter the importance distribution of image information, but most fusion methods still employ fixed feature fusion strategies, lacking the ability to adaptively adjust to environmental changes. This also makes it easy for high-frequency structural information to gradually decay during multi-layer feature extraction and fusion, resulting in insufficient edge and texture detail in the final fused image. Summary of the Invention

[0006] To address the aforementioned shortcomings in existing technologies, the infrared and visible light image fusion method based on text semantic guidance and frequency domain compensation provided by this invention solves the problems of easy mutual interference of multimodal information, insufficient cross-modal interaction, and easy loss of detailed information in complex environments.

[0007] To achieve the aforementioned objectives, the present invention employs the following technical solution: an infrared-visible light image fusion method based on text semantic guidance and frequency domain compensation, comprising the following steps: S1. Input the infrared image and the visible light image into the dual-branch feature encoding module. In the encoding stage, the spatial context enhancement strategy is used to strengthen the feature expression capability and generate multi-scale features of the infrared image and the visible light image. S2. Input the multi-scale features of infrared and visible light images into the cross-modal interactive fusion module, establish the correlation between infrared and visible light features through the attention mechanism, and obtain the fused feature map; S3. Input the fused feature map into the frequency domain detail enhancement and cross-layer propagation decoding module. In the decoding stage, adaptive supplementation and depth restoration of high-frequency information are performed to generate the final fused image.

[0008] Furthermore: In S1, the dual-branch feature encoding module includes encoders for infrared images and visible light images. The two encoders have the same structure, each including a first Transformer Based Block, a second Transformer Based Block, a first Transformer EFFN Block, and a second Transformer EFFN Block connected in sequence. The first Transformer EFFN Block and the second Transformer EFFN Block both contain enhanced frequency domain feedforward networks. Spatial channel attention modules are set after the second Transformer Based Block, the first Transformer EFFN Block and the second Transformer EFFN Block to improve the spatial structure modeling ability and contextual semantic awareness ability of feature representation.

[0009] Furthermore, the specific workflow of the enhanced frequency domain feedforward network is as follows: A1. Input characteristics of enhanced frequency domain feedforward networks Perform layer normalization, and through Convolution and Depthwise separable convolutions generate the first intermediate features separately. Second intermediate features ; In the formula, , Represents the space of real numbers. Indicates batch size, Indicates the number of feature channels. For layer normalization, Indicates the height of the feature map. Indicates the width of the feature map. Indicates the kernel size as The first convolution weight matrix, Indicates the kernel size as The second convolution weight matrix is ​​used for channel mapping. Indicates the kernel size as The first depthwise separable convolution, Indicates the kernel size as The second depthwise separable convolution is used to introduce local spatial context information; A2. To further utilize shallow structural information to construct spatially modulated signals, the frequency domain feedforward network is enhanced to extract spatial context information from auxiliary features and generate spatial modulation features. ; In the formula, This indicates auxiliary features originating from shallow or auxiliary branches. This represents the average pooling operation, used to extract global statistics. This represents a mapping function consisting of convolutional layers, normalized layers, and nonlinear activation functions. This indicates an upsampling operation used to recover the spatial size of features, which modulates the spatial features. With the first intermediate feature The data is concatenated along the channel dimension, and gating weights are generated. : In the formula, This indicates a channel splicing operation. The first convolutional mapping weights, The activation function for the Gaussian error linear unit; A3. Obtain the local frequency domain spatial enhancement features of the output of the enhanced frequency domain feedforward network through element-wise multiplication. ; In the formula, This is an element-wise multiplication operation.

[0010] Furthermore: The spatial channel attention module includes interconnected high- and wide-dimensional gating attention modules and channel attention gating modules. The specific workflow of the spatial channel attention module is as follows: B1. Input features of the spatial channel attention module The input is then fed into a height-width-dimensional gated attention module, where global average pooling is performed along both the height and width directions to obtain contextual descriptions in two directions, including the feature descriptions compressed along the height direction. Feature description obtained by compression along the width direction ; In the formula, , This represents the average pooling operation along the height direction, with an output size of... , This represents an average pooling operation along the width direction, with an output size of... ; The context description is divided into several sub-feature groups along the channel dimension, and one-dimensional deep convolution with different receptive fields is applied to each group to generate sub-features after deep convolution, including the first sub-feature representation. Second sub-feature representation ; In the formula, express No. Features of individual channel groups express No. Features of individual channel groups , The total number of channel sub-feature groups; Indicates the kernel size as One-dimensional depthwise convolution operation, ; After concatenating all sub-features along the channel dimension, spatial attention weights, including horizontal spatial attention weights, are generated through grouping normalization and a sigmoid activation function. Spatial attention weight in the vertical direction ; In the formula, for No. Features of individual channel groups for No. Features of individual channel groups for No. Features of individual channel groups for No. Features of individual channel groups This indicates a channel splicing operation. This indicates a grouping normalization operation. This represents the Sigmoid activation function; Computational Space Augmentation Features , as the output feature of the high- and wide-dimensional gated attention module; B2. Input the spatially enhanced features into the channel attention gating module, and compress the features through spatial pooling to obtain a low-resolution representation. : In the formula, express Space pooling operation, through Convolution generates query, key, and value vectors: In the formula, Generate a general attention query vector for the current fused features. Generic attention key vector generated for the current fused features. A generic attention value vector generated for the current fused features. The weights are the mapping weights for the second convolution. The weights are mapped to the third convolution. For the fourth convolution mapping weights, calculate channel self-attention. The specific expression is: In the formula, Represents the normalization function. The channel weights are generated using global pooling and the sigmoid activation function, which acts as a scaling factor for the channel dimension. These weights are then modulated to obtain the spatially enhanced feature input and output features of the channel attention gating module. This is used as the output feature of the spatial channel attention module; In the formula, This indicates a global average pooling operation.

[0011] Furthermore, in S2, the cross-modal interaction fusion module includes a difference-aware feature selector module, a text cross-scale fusion module, an enhanced feature fusion module, and a spatial structure enhancement module connected in sequence.

[0012] Furthermore, the workflow of the difference-aware feature selector module is as follows: C1. Perform channel-by-channel mapping on the multi-scale features of the infrared and visible light images to obtain the initial visible light features after channel mapping. and infrared initial features : In the formula, For visible light modes Convolution mapping function, For infrared mode The convolution mapping function is then used to calculate the joint response features of the two modalities. : C2. Generating modulation features through local aggregation operations : In the formula, This represents depthwise separable convolution. express Convolution operation, For layer normalization, The Sigmoid activation function is used, followed by the generation of dynamic weights for both modalities. : In the formula, Represents the dynamic weighting coefficients of the visible light modes. Represents the dynamic weighting coefficients of the infrared modes. Represents a linear mapping. This is the normalization function; C3. Perform a Fourier transform to generate a frequency domain representation of the visible light characteristics. Frequency domain representation of infrared features : In the formula, This represents a Fourier transform, through conditional weighting, to obtain a dynamic filter, including dynamic frequency filters for visible light modes. Dynamic frequency filter for infrared modes : In the formula, For the first One basic filter, Indicates the number of filters. The dynamic weighting coefficients of the visible light modes The generated filter combination coefficients; The dynamic weighting coefficients of the infrared modes The generated filter combination coefficients are multiplied point-by-point in the frequency domain to obtain the enhancement features of the two modes, which are then used as the output of the difference-aware feature selector module. These enhancement features include visible light enhancement features. and infrared enhancement features : In the formula, For inverse Fourier transform, This is element-wise multiplication.

[0013] Furthermore: The text cross-scale fusion module calculates semantically enhanced visual features based on the enhanced features of the two input modalities, including semantically enhanced visible light modal features and infrared modal features. The methods for calculating semantically enhanced visible light modal features and infrared modal features are the same. Specifically, the method for calculating semantically enhanced visible light modal features is as follows: D1. Map the enhanced features of the two modalities to construct the query, key, and value in the attention mechanism, and calculate the cross-modal attention weights from visible light to infrared. : In the formula, The query vector generated for visible light enhancement features. The key vector generated for infrared enhancement features, For attention feature dimension, The transpose symbol is used; infrared features are weighted and aggregated using attention weights to obtain cross-modal interactive features. : In the formula, Value vectors generated for infrared enhancement features; D2. Construct a cue set using several learnable cue vectors, and dynamically generate semantic cue representations based on text semantic features. : In the formula, For learnable cue vectors, This represents the set of Top-k hints with the highest semantic similarity. Text features For the prompt vector The weight, , For text feature dimensions; By fusing cross-modal attention features with semantic cue information, semantically enhanced visible light modal features are obtained. : .

[0014] Furthermore, the workflow of the enhanced feature fusion module is as follows: E1. Construct mode-shared features based on the semantically enhanced visible light mode features and infrared mode features. Modal difference characteristics : In the formula, These are the semantically enhanced infrared modal features. Represents the convolution operation; generates a conditional vector through global average pooling. : In the formula, Indicates global average pooling. This represents a multilayer perceptron; E2. Modulate the frequency features using a dynamic filter to generate enhanced shared features. and enhanced differential features : In the formula, For the frequency domain representation of shared features, , The frequency domain representation of the difference features. , Represents the condition vector The generated shared feature filter, Represents the condition vector Generate the difference feature filter; then calculate the fusion feature. : In the formula, This represents a learnable parameter used to balance the contribution of differential features to the fusion result.

[0015] Furthermore: The specific workflow of the space structure enhancement module is as follows: F1. For the fusion features of the input spatial structure enhancement module, expand them into a sequence form and perform spatial enhancement to obtain the local attention output features. : In the formula, Generate a general attention query vector for the current fused features. Generic attention key vector generated for the current fused features. A generic attention value vector generated for the current fused features. For local spatial attention weights, It is the transpose symbol; F2. To capture long-range dependencies, a global semantic prototype is further introduced. : In the formula, For the first 1 semantic prototype vector, and construct global cross-attention output features. : In the formula, Generate a general attention query vector for the current fused features. The key vector is obtained by mapping the global semantic prototype set. A value vector obtained by mapping the global semantic prototype set; F3. Calculate the spatial enhancement features after fusion. The feature map is then restored to a two-dimensional feature map, resulting in a spatially enhanced fused feature map. : In the formula, This is a tensor rearrangement operation.

[0016] Furthermore: In S3, the frequency domain detail enhancement and cross-layer propagation decoding module includes a first decoding layer to a fourth decoding layer connected in sequence, namely the third Transformer EFFN Block, the fourth Transformer EFFN Block, the third Transformer Based Block, and the fourth Transformer Based Block. The output of each decoding layer is connected to the dynamic high-frequency feature extractor of the current layer, and the dynamic high-frequency feature extractors of the first layer to the fourth layer are connected in sequence. The specific workflow of the frequency domain detail enhancement and cross-layer propagation decoding module is as follows: G1. After performing decoding operations at each decoding layer, the decoded features of the current layer are input into the dynamic high-frequency feature extractor of the current layer to generate high-frequency residuals passed from the current layer to the previous layer. These residuals are then accumulated with the high-frequency details of the previous layer, and this accumulation is performed layer by layer to obtain the accumulated high-frequency residual. High-frequency details; Among them, the generation from the first Layer pass to the first The high-frequency residuals of the layer are obtained by summing the results of the first layer. The process of detailing high-frequency information in a layer is as follows: Let the first layer be... Layer decoding features are , No. Layer dynamic high-frequency feature extractor from The high-frequency details extracted are ,pass Convolution generates the first Layer gate weight graph : In the formula, For feature-level indexing, For the first Layer decoding features, This is a 1×1 convolution operation. Represents the Sigmoid activation function; through gated weight graph fusion, the th... Enhanced features after layer fusion : In the formula, This represents element-wise multiplication; To achieve the transmission of details from deep to shallow layers, a cross-layer propagation path is constructed: [The path is then used to] transmit the first layer's details to the shallow layer's details. Enhanced features after layer fusion pass Channel projection and bilinear upsampling generate from the first... Layer pass to the first High-frequency residuals of the layer ; In the formula, For bilinear upsampling, Used with the The high-frequency residuals of the layers are accumulated, and the accumulated values ​​are the first layer. High-frequency details The specific expression is: In the formula, For the first Layer gate weight map For the first High-frequency details output by the layer dynamic high-frequency feature extractor; G2, for the first Layer decoding features And after accumulation, the first High-frequency details Line splicing, through Convolution generates spatial attention maps : In the formula, Values , express Convolution; Calculate the final fused image : The dynamic high-frequency feature extractor calculates high-frequency details based on the input decoded features as follows: H1. Perform a two-dimensional Fourier transform on the input decoding features and then shift the frequency to obtain the frequency domain features after frequency shifting. : In the formula, This is a frequency domain rearrangement operation. The decoding features are the input. Then, the amplitude spectrum is calculated and averaged across channels to obtain the amplitude spectrum after cross-channel averaging. : H2. Calculate the total frequency domain energy of the current feature map: In the formula, These are the row coordinates of the two-dimensional frequency domain feature map. The column coordinates of the two-dimensional frequency domain feature map; given the energy parameters of this layer. Minimum radius of the low-frequency region of the dynamic search center To satisfy: In the formula, The x-coordinate of the frequency domain center is... The ordinate of the frequency domain center is used; a binary mask is then constructed. : After applying a binary mask to the frequency domain features, the high-frequency details are output through inverse frequency shift and inverse Fourier transform. : In the formula, This is the inverse Fourier transform (iFFT). It is a frequency domain inverse shift function. For taking the mold.

[0017] The beneficial effects of this invention are as follows: (1) This invention proposes an infrared and visible light image fusion method based on text semantic guidance and frequency domain compensation. The process adopts a hierarchical structure of encoding-fusion-decoding. Through three core mechanisms, namely dual-branch feature encoding and spatial context enhancement, cross-modal interactive fusion and text-guided modulation, and frequency domain detail enhancement and cross-layer propagation, it effectively realizes the stable complementary fusion of infrared target thermal saliency and visible light structural texture details, and solves the problems of easy mutual interference of multimodal information, insufficient cross-modal interaction, and easy loss of detail information in complex environments.

[0018] 1) Two-branch feature encoding and spatial context enhancement mechanism: A dual-branch feature coding structure for infrared and visible light is established based on the dual-branch feature coding module. Multi-scale feature representations of different modalities are extracted by two independent encoders to obtain spatially enhanced features. While maintaining the integrity of modal difference information, a spatial context enhancement strategy is introduced to perform context modeling and background suppression on the spatially enhanced features, thereby strengthening the feature expression capability and providing a more stable and richer feature foundation for subsequent cross-modal fusion.

[0019] 2) Cross-modal interaction fusion and text-guided modulation mechanism: A cross-modal interaction fusion module is constructed in a high-level semantic feature space. An attention mechanism is used to establish the correlation between infrared and visible light features, enabling effective interaction of complementary information. At the same time, textual semantic prompts are introduced. Semantic embeddings are extracted through a pre-trained vision-language model, and the fusion features are dynamically modulated, enabling the network to adaptively adjust the fusion strategy according to different degradation environments.

[0020] 3) Frequency domain detail enhancement and cross-layer propagation mechanism: In the decoding stage, a frequency domain high-frequency information enhancement method is introduced. After each layer of decoding operation is completed, the high-frequency residual is immediately extracted from the current decoding feature, and the high-frequency information is propagated layer by layer in combination with the cross-layer feature propagation strategy to achieve adaptive supplementation and depth restoration of high-frequency details. This effectively restores the high-frequency texture information that is easily lost during the fusion process and improves the structural integrity and visual clarity of the fused image.

[0021] (2) Numerous qualitative and quantitative experiments show that the proposed method achieves significant superiority in key no-reference indicators such as CLIP-IQA, MUSIQ, and TReS under mixed degradation scenarios such as fog, rain, low illumination, and fog and low light, demonstrating excellent perception quality, structural integrity and information retention capabilities, and fully verifying its robustness and generalization performance in complex environments.

[0022] (3) Compared with existing methods, this framework effectively solves the problems of cross-modal information interference, detail loss, and insufficient degradation adaptability, providing an effective solution for multimodal image fusion under adverse weather conditions. In the future, the model's lightweight and real-time performance can be further optimized, and its expansion to more sensor modalities can be explored to promote the application of this technology in practical scenarios such as safety monitoring and assisted driving. Attached Figure Description

[0023] Figure 1 This is a flowchart of the infrared and visible light image fusion method based on text semantic guidance and frequency domain compensation according to the present invention.

[0024] Figure 2 The infrared and visible light image fusion network architecture constructed for this invention.

[0025] Figure 3 A schematic diagram of a frequency domain feedforward network.

[0026] Figure 4 This is a schematic diagram of the spatial channel attention module.

[0027] Figure 5 This is a schematic diagram of the difference-aware feature selector module.

[0028] Figure 6 This is a schematic diagram of the text cross-scale fusion module.

[0029] Figure 7 A schematic diagram of the feature fusion module.

[0030] Figure 8 This is a schematic diagram of the space structure enhancement module.

[0031] Figure 9 This is a schematic diagram of a dynamic high-frequency feature extractor. Detailed Implementation

[0032] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0033] like Figure 1 As shown, in one embodiment of the present invention, the infrared-visible light image fusion method based on text semantic guidance and frequency domain compensation includes the following steps: S1. Input the infrared image and the visible light image into the dual-branch feature encoding module. In the encoding stage, the spatial context enhancement strategy is used to strengthen the feature expression capability and generate multi-scale features of the infrared image and the visible light image. S2. Input the multi-scale features of infrared and visible light images into the cross-modal interactive fusion module, establish the correlation between infrared and visible light features through the attention mechanism, and obtain the fused feature map; S3. Input the fused feature map into the frequency domain detail enhancement and cross-layer propagation decoding module. In the decoding stage, adaptive supplementation and depth restoration of high-frequency information are performed to generate the final fused image.

[0034] When fusing infrared and visible light images under adverse weather conditions, traditional methods often struggle to simultaneously preserve effective information from both modalities. Fog, rain, or low-light environments lead to decreased contrast and blurred texture details in visible light images. While infrared images can stably represent target thermal information, they lack structural details and background hierarchy. When the two modalities are directly fused, problems such as insufficient information utilization or mutual interference can easily arise. On the one hand, the thermal features of the target area may be weakened by background texture; on the other hand, the detailed information in the visible light is over-smoothed during the fusion process. Furthermore, complex weather conditions can introduce noise and contrast imbalances, making it difficult for traditional fusion strategies to maintain stable performance in different scenarios. This results in a struggle to achieve a balance between target saliency, structural integrity, and detail preservation in the fusion result.

[0035] To address the aforementioned issues, this paper adopts a multi-module collaborative optimization design approach to construct an infrared-visible light image fusion network architecture, thereby achieving effective integration of multimodal information. Its overall structure mainly consists of a dual-branch feature encoding module, a cross-modal interactive fusion module, and a frequency domain detail enhancement and cross-layer propagation decoding module. The dual-branch feature encoding module extracts multi-scale features from infrared and visible light images using two independent encoders to maintain the integrity of information across different modalities. The cross-modal interactive fusion module establishes the correlation between the two modalities in the high-level semantic feature space, promoting effective interaction of complementary information through cross-modal attention mechanisms and feature fusion strategies. The text cross-scale fusion module uses external semantic cues to adaptively modulate features, enabling the network to dynamically adjust the fusion strategy according to different degradation scenarios. The frequency domain detail enhancement and cross-layer propagation decoding module extracts and propagates high-frequency information layer by layer during the decoding stage to supplement structural details easily lost during the fusion process, thereby improving the overall target representation ability and texture integrity of the fused image. The specific workflow of this architecture is as follows: Figure 2 As shown.

[0036] In terms of overall structure, the proposed network adopts a hierarchical framework of Encoder-Fusion-Decoder to achieve the gradual integration of multimodal information. Its core idea can be summarized as: "First, extract modal differentiation features, then achieve deep fusion through semantic interaction, and finally restore and enhance high-frequency detail information layer by layer." The network architecture of this invention can effectively restore structural texture information while maintaining target saliency, thereby generating a fused image that balances semantic expressiveness and detail integrity. Furthermore, the network achieves effective fusion of infrared and visible light information through a step-by-step information processing flow of "feature enhancement in the encoding stage—cross-modal interaction in the semantic space—frequency domain detail restoration in the decoding stage."

[0037] In the encoding stage of infrared-visible image fusion, the core task is to fully mine and retain the complementary feature information of the two modalities. Simultaneously, considering the inherent differences in the imaging mechanisms of infrared and visible light, the feature extraction process should enhance spatial structure modeling capabilities and suppress noise interference from adverse weather conditions. This requires the encoding structure to not only achieve independent multi-scale extraction of bimodal features but also provide stable and discriminative feature representations for subsequent cross-modal interactive fusion. Based on this, this paper designs an encoding strategy combining bi-branch feature encoding and spatial context enhancement to construct a bi-branch feature encoding module.

[0038] In common coding structures, while deep features possess stronger semantic expressive power, their spatial structural information often weakens gradually during layer-by-layer abstraction. Furthermore, in Transformer-based network structures, feed-forward networks (FFNs) primarily perform feature transformations through linear mappings along the channel dimension, resulting in relatively limited ability to model spatial structural relationships. This problem is particularly pronounced in infrared-visible light fusion tasks: if deep features lack sufficient spatial structural constraints, it can easily lead to blurred target boundaries or loss of detailed information. On the other hand, under complex environmental conditions, significant background interference weakens the stability of feature representation. Therefore, it is necessary to introduce effective contextual modeling mechanisms during the encoding process to enhance the network's ability to perceive key structural information and suppress irrelevant background information.

[0039] To address the aforementioned issues, this paper employs a dual-branch feature extraction structure in the encoding stage, performing multi-scale feature encoding on infrared and visible light images respectively. Furthermore, it introduces an enhanced frequency-domain feedforward (EFFN) network and a spatial channel attention (SCA) module into the deep feature space to enhance the spatial structure modeling capability and contextual semantic awareness capability of feature representation.

[0040] In S1, the dual-branch feature encoding module includes encoders for infrared images and visible light images. The two encoders have the same structure, each including a first Transformer Based Block, a second Transformer Based Block, a first Transformer EFFN Block, and a second Transformer EFFN Block connected in sequence. The first Transformer EFFN Block and the second Transformer EFFN Block are equipped with enhanced frequency domain feedforward networks. A spatial channel attention module is set after the second Transformer Based Block, the first Transformer EFFN Block and the second Transformer EFFN Block.

[0041] Different levels of features have significant differences in spatial structure and semantic expression. Shallow features (such as Level-1 and Level-2) usually have high spatial resolution and can better preserve edge and texture details, but their semantic expression ability is relatively limited. Conversely, deep features (such as Level-3 and Level-4) have stronger semantic expression ability after multiple downsampling and feature aggregation, but their spatial structure information is often gradually weakened in the process of layer-by-layer abstraction, which can easily lead to blurred target boundaries or degradation of detailed structure.

[0042] In infrared-visible fusion tasks, target regions in infrared images typically rely on high-level semantic information for identification, while fine-grained textures in visible light images require spatial structural information for supplementation. If deep semantic features lack effective spatial constraints, structural ambiguity or unclear target outlines can easily occur during the fusion process. Therefore, compared with shallow features, deep features require additional spatial structural enhancement and contextual modeling mechanisms.

[0043] This paper introduces an enhanced frequency domain feedforward network (EFFN) in the deeper stages of the encoder (Level-3 and Level-4), such as... Figure 3 As shown, the EFFN replaces the feedforward network FFN in the traditional Transformer structure. The reason why this structure is mainly applied to deep features rather than directly to shallow features is that EFFN enhances the structural expression ability of semantic features through spatial modulation mechanism. Shallow features already contain relatively rich local structural information, and introducing complex spatial modulation too early may destroy the original texture distribution.

[0044] Furthermore, shallow features possess high spatial resolution, and directly introducing complex spatial enhancement structures at low levels would significantly increase computational overhead. Therefore, applying them to deeper features (Level-3 / Level-4) with stronger semantics and lower resolution can achieve more significant structural enhancement effects while maintaining efficiency. That is, this paper chooses to introduce EFFN in the deep feature stage of the encoder to supplement spatial structural information while maintaining high-level semantic expressiveness, thereby obtaining a more stable semantic feature representation. The specific workflow of the enhanced frequency domain feedforward network is as follows: A1. Input characteristics of enhanced frequency domain feedforward networks Perform layer normalization, and through Convolution and Depthwise separable convolutions generate the first intermediate features separately. Second intermediate features ; In the formula, , Represents the space of real numbers. Indicates the batch size. Indicates the number of feature channels. For layer normalization, Indicates the height of the feature map. Indicates the width of the feature map. Indicates the kernel size as The first convolution weight matrix, Indicates the kernel size as The second convolution weight matrix is ​​used for channel mapping. Indicates the kernel size as The first depthwise separable convolution, Indicates the kernel size as The second depthwise separable convolution is used to introduce local spatial context information; A2. To further utilize shallow structural information to construct spatially modulated signals, the frequency domain feedforward network is enhanced to extract spatial context information from auxiliary features and generate spatial modulation features. ; In the formula, This indicates auxiliary features originating from shallow or auxiliary branches. This represents the average pooling operation, used to extract global statistics. This represents a mapping function consisting of convolutional layers, normalized layers, and nonlinear activation functions. This indicates an upsampling operation used to recover the spatial size of features, which modulates the spatial features. With the first intermediate feature The data is concatenated along the channel dimension, and gating weights are generated. : In the formula, This indicates a channel splicing operation. The first convolutional mapping weights, The activation function for the Gaussian error linear unit; A3. Obtain the local frequency domain spatial enhancement features of the output of the enhanced frequency domain feedforward network through element-wise multiplication. ; In the formula, This is an element-wise multiplication operation. Through this gating mechanism, EFFN can adaptively adjust the feature responses of different regions based on spatial structure information, thereby enhancing the expressive power of target boundaries and fine-grained structures.

[0045] While EFFN can enhance local structural information in features through spatial modulation mechanisms, relying solely on local spatial enhancement is insufficient to fully model long-range contextual relationships under complex environmental conditions. For example, in severe weather scenarios, numerous irrelevant regions may interfere with feature representation. Therefore, after obtaining spatially enhanced features, a global context modeling mechanism is needed to improve the network's ability to focus on important structural regions and suppress background noise. After obtaining spatially enhanced features, the features are further modeled using a Spatial Channel Attention (SCA) module for contextualization and background suppression, such as... Figure 4 As shown, the spatial channel attention module consists of two parts: a high-width-to-high-dimensional gated attention (HWGA) module and a channel attention gate (CAG) module. The HWGA module is responsible for modeling spatial structural relationships, while the CAG module is used to capture semantic dependencies in the channel dimension.

[0046] The workflow of the spatial channel attention module is as follows: B1. Input features of the spatial channel attention module The input is then fed into a height-width-dimensional gated attention module, where global average pooling is performed along both the height and width directions to obtain contextual descriptions in two directions, including the feature descriptions compressed along the height direction. Feature description obtained by compression along the width direction ; In the formula, , This represents the average pooling operation along the height direction, with an output size of... , This represents an average pooling operation along the width direction, with an output size of... ; The context description is divided into several sub-feature groups (usually 4) along the channel dimension, and one-dimensional depthwise convolutions with different receptive fields are applied to each group to generate sub-features after depthwise convolution, including the first sub-feature representation. Second sub-feature representation ; In the formula, express No. Features of individual channel groups express No. Features of individual channel groups , The total number of channel sub-feature groups; Indicates the kernel size as One-dimensional depthwise convolution operation, ; After concatenating all sub-features along the channel dimension, spatial attention weights, including horizontal spatial attention weights, are generated through grouping normalization and a sigmoid activation function. Spatial attention weight in the vertical direction ; In the formula, for No. Features of individual channel groups for No. Features of individual channel groups for No. Features of individual channel groups for No. Features of individual channel groups This indicates a channel splicing operation. This indicates a grouping normalization operation. This represents the Sigmoid activation function; Computational Space Augmentation Features , as the output feature of the high- and wide-dimensional gated attention module; B2. Input the spatially enhanced features into the Channel Attention Gate (CAG) module to further model the semantic dependencies between features in the channel dimension, and compress the features through spatial pooling to obtain a low-resolution representation. : In the formula, express Space pooling operation, through Convolution generates a vector of queries, keys, and values: In the formula, Generate a general attention query vector for the current fused features. Generic attention key vector generated for the current fused features. A generic attention value vector generated for the current fused features. The weights are the mapping weights for the second convolution. The weights are mapped to the third convolution. For the fourth convolution mapping weights, calculate channel self-attention. The specific expression is: In the formula, Represents the normalization function. The channel weights are generated using global pooling and the sigmoid activation function, which acts as a scaling factor for the channel dimension. These weights are then modulated to obtain the spatially enhanced feature input and output features of the channel attention gating module. This is used as the output feature of the spatial channel attention module; In the formula, This indicates a global average pooling operation.

[0047] Through this spatial and channel collaborative modeling approach, the spatial channel attention module can effectively suppress background noise while preserving structural information, thereby obtaining a more stable feature representation.

[0048] By collaboratively introducing an enhanced frequency-domain feedforward network and a spatial channel attention module into a dual-branch coding structure, the network can simultaneously achieve spatial structure enhancement and contextual relationship modeling during the deep semantic feature stage. The enhanced frequency-domain feedforward network primarily compensates for the spatial structure information lost during the layer-by-layer abstraction of deep features, while the spatial channel attention module further models long-distance contextual dependencies through a spatial and channel collaborative attention mechanism, thereby suppressing interference from complex backgrounds. Combined with the dual-branch multi-scale feature coding structure, the coding stage ultimately yields a more stable and discriminative multi-scale modal feature representation, providing a reliable semantic foundation for subsequent cross-modal interaction and fusion.

[0049] Although the dual-branch coding structure can extract deep feature representations of visible light and infrared modes separately, significant differences remain between the two modes in terms of statistical distribution, frequency structure, and semantic response. Direct feature fusion not only fails to fully utilize the complementary information between the two modes but may also introduce redundant features due to mode conflicts. Especially under adverse weather conditions such as low light, rain, and fog, visible light images are often affected by illumination and noise, while infrared images, although possessing strong target response capabilities, lack sufficient texture detail. Therefore, simple stitching or weighted fusion is unlikely to yield stable fusion results.

[0050] To address this issue, this paper designs a cross-modal interaction fusion framework consisting of difference modeling, semantic guidance, frequency fusion, and spatial enhancement. The framework first enhances modal complementarity information using a difference-aware feature selector module, then guides modal interactions using textual semantics, achieves adaptive feature integration through a frequency domain fusion mechanism, and finally strengthens the structural expressive power of the fused features through a spatial structure enhancement module.

[0051] In S2, the cross-modal interactive fusion module includes a Diff-Aware Feature Selector (DAFS) module, a Text Cross Multi-Scale Fusion (TCMF) module, a Frequency Exhaustive Fusion Mechanism (EFF) module, and an Interactive2D Channel Attention (I2DCA) module connected in sequence.

[0052] Because visible light and infrared modes differ significantly in frequency structure, visible light images typically contain rich low-frequency texture and color information, while infrared images more readily highlight high-frequency structures and target responses. Therefore, before cross-modal interaction, it is necessary to explicitly model the differences between the two modalities. Thus, a difference-aware feature selector module is first introduced, such as... Figure 5 As shown.

[0053] The workflow of the difference-aware feature selector module is as follows: C1. Perform channel-by-channel mapping on the multi-scale features of the infrared and visible light images to obtain the initial visible light features after channel mapping. and infrared initial features : In the formula, For visible light modes Convolution mapping function, For infrared mode The convolution mapping function is then used to calculate the joint response features of the two modalities. : C2. Generating modulation features through local aggregation operations : In the formula, This represents depthwise separable convolution. express Convolution operation, For layer normalization, The Sigmoid activation function is used, followed by the generation of dynamic weights for both modalities. : In the formula, Represents the dynamic weighting coefficients of the visible light modes. Represents the dynamic weighting coefficients of the infrared modes. Represents a linear mapping. This is the normalization function; C3. To further enhance modal difference information, this paper performs dynamic filtering on the features in the frequency domain. A Fourier transform is performed to generate a frequency domain representation of the visible light features. Frequency domain representation of infrared features : In the formula, This represents a Fourier transform, through conditional weighting, to obtain a dynamic filter, including dynamic frequency filters for visible light modes. Dynamic frequency filter for infrared modes : In the formula, For the first One basic filter, Indicates the number of filters. The dynamic weighting coefficients of the visible light modes The generated filter combination coefficients; The dynamic weighting coefficients of the infrared modes The generated filter combination coefficients are multiplied point-by-point in the frequency domain to obtain the enhancement features of the two modes, which are then used as the output of the difference-aware feature selector module. These enhancement features include visible light enhancement features. and infrared enhancement features : In the formula, For inverse Fourier transform, This is an element-wise multiplication. Through this structure, the difference-aware feature selector module can highlight modal difference information at the frequency level, providing a more stable feature representation for subsequent cross-modal interactions.

[0054] Although the difference-aware feature selector module can enhance modal difference information, the process still mainly relies on the interaction between visual features. However, in complex environments, some target regions may not be visually significant, but semantic information has a strong discriminative ability.

[0055] Traditional cross-modal interaction typically relies solely on attentional relationships between visual features to establish modal connections, essentially remaining an information matching process within the visual domain. In complex and degenerate environments, this visual feature-dependent interaction method is susceptible to factors such as noise, low contrast, and texture loss, leading to unstable cross-modal attentional relationships and consequently affecting the effective extraction of complementary modal information.

[0056] Therefore, this paper further introduces external semantic information as an auxiliary guide for cross-modal interaction by constructing a Text Cross Multi-Scale Fusion (TCMF) module, such as... Figure 6 As shown, textual semantic embedding is used to modulate the attentional relationships between visual features, so that cross-modal interaction not only depends on visual similarity but is also constrained by semantic information, thereby obtaining more stable and discriminative modal associations in complex scenes.

[0057] The text cross-scale fusion module calculates semantically enhanced visual features based on the enhanced features of the two input modalities, including semantically enhanced visible light modal features and infrared modal features. The methods for calculating semantically enhanced visible light modal features and infrared modal features are the same. Specifically, the method for calculating semantically enhanced visible light modal features is as follows: D1. Map the enhanced features of the two modalities to construct the query, key, and value in the attention mechanism. Specifically, the visible light enhanced features... and infrared enhancement features Through mapping function An attention vector is generated, and then the cross-modal attention weights from visible light to infrared are calculated. : In the formula, The query vector generated for visible light enhancement features. The key vector generated for infrared enhancement features, For attention feature dimension, The transpose symbol is used; infrared features are weighted and aggregated using attention weights to obtain cross-modal interactive features. : In the formula, Value vectors generated for infrared enhancement features; D2. To further introduce semantic constraints, this paper constructs text cue words to enhance the ability of text semantics to guide visual features. Specifically, a cue set is constructed using several learnable cue vectors, and semantic cue representations are dynamically generated based on text semantic features. : In the formula, For learnable cue vectors, This represents the set of Top-k hints with the highest semantic similarity. Text features For the prompt vector The weight, , Extracted by a text encoder, For text feature dimensions; By fusing cross-modal attention features with semantic cue information, semantically enhanced visible light modal features are obtained. : Similarly, the semantically enhanced infrared modal features can be obtained using the methods described above. .

[0058] By introducing a semantic guidance mechanism, the cross-modal feature interaction process no longer relies solely on the matching relationship between visual features, but is also constrained by semantic information. This allows for more stable highlighting of target region features under complex environmental conditions, effective suppression of background interference, and improved reliability of cross-modal information fusion.

[0059] After completing cross-modal interaction, the contributions of the two modalities still differ in different spatial regions. Therefore, it is necessary to further construct an adaptive fusion mechanism to achieve more stable feature integration. This paper designs an enhanced feature fusion module (EFF) to simultaneously model modal shared and differential information in the frequency domain, such as... Figure 7 As shown.

[0060] The workflow of the enhanced feature fusion module is as follows: E1. Construct mode-shared features based on the semantically enhanced visible light mode features and infrared mode features. Modal difference characteristics : In the formula, Represents the convolution operation; generates a conditional vector through global average pooling. : In the formula, This indicates Global Average Pooling. This represents a multilayer perceptron; E2. Modulate the frequency features using a dynamic filter to generate enhanced shared features. and enhanced differential features : In the formula, For the frequency domain representation of shared features, , The frequency domain representation of the difference features. , Represents the condition vector The generated shared feature filter, Represents the condition vector Generate the difference feature filter; then calculate the fusion feature. : In the formula, This represents learnable parameters used to balance the contribution of dissimilar features to the fusion result. By simultaneously modeling shared frequency information and modal dissimilar information, the enhanced feature fusion module can supplement key detail information while maintaining structural consistency, thereby obtaining a more complete fused feature representation.

[0061] Although the above steps have completed modal fusion, the fused features may still have shortcomings in spatial dependency modeling. To further enhance the target structural information, a spatial structure enhancement module is introduced to spatially enhance the fused features, such as... Figure 8 As shown. The specific workflow of the space structure enhancement module is as follows: F1. For the fusion features of the input spatial structure enhancement module, expand them into a sequence form. In the formula For tensor rearrangement operations, spatial augmentation is performed to obtain local attention output features. : In the formula, Generate a general attention query vector for the current fused features. Generic attention key vector generated for the current fused features. A generic attention value vector generated for the current fused features. For local spatial attention weights, It is the transpose symbol; F2. To capture long-range dependencies, a global semantic prototype is further introduced. : In the formula, For the first A semantic prototype vector, in this embodiment, a global semantic prototype. It is initialized with learnable parameters and dynamically optimized during training through a clustering update mechanism based on feature similarity, enabling it to adaptively represent the global semantic structure in the input feature space. The construction method is as follows: First, a randomly generated initial prototype is initialized; that is, a prototype is randomly initialized directly when the network is first set up. A fixed-dimensional vector is used as the initial global semantic prototype. Each vector is the global semantic prototype. The initial value is then used; the prototype is then dynamically updated iteratively using clustering during the training phase. During training, for each batch of input image features, the prototype is first calculated to match the current image features. The similarity of each semantic prototype is used to assign features to the most similar prototypes; then, for each group of features, the cluster center is recalculated, and the new center replaces the original one. Continuously transform the global semantic prototype The corrections are made more accurate so that it can represent the global semantic patterns in the scene.

[0062] Construct global cross-attention output features : In the formula, Generate a general attention query vector for the current fused features. The key vector is obtained by mapping the global semantic prototype set. A value vector obtained by mapping the global semantic prototype set; F3. Calculate the spatial enhancement features after fusion. The feature map is then restored to a two-dimensional feature map, resulting in a spatially enhanced fused feature map. : In the formula, This is a tensor rearrangement operation. In this embodiment, the spatial structure enhancement module simultaneously models local and global spatial relationships. This module can further enhance the structural information of the target region and effectively suppress background noise.

[0063] A cross-modal interactive fusion framework, comprised of feature difference modeling, semantically guided interaction, enhanced feature fusion, and spatial structure enhancement, is constructed using a difference-aware feature selector module, a text cross-scale fusion module, an enhanced feature fusion module, and a spatial structure enhancement module. This framework first models modal differences in the frequency domain using the difference-aware feature selector module. Then, it introduces semantic information to guide visual interaction using the text cross-scale fusion module. Based on this, it achieves adaptive frequency fusion through the enhanced feature fusion module. Finally, it strengthens spatial structure dependencies through the spatial structure enhancement module. Through multi-stage progressive modeling, the network can more fully mine the complementary information between visible light and infrared modalities, thereby obtaining a more stable and discriminative fused feature map under adverse weather conditions.

[0064] During the decoding stage, although the network gradually recovers the spatial resolution through layer-by-layer upsampling and skip connections, the unavoidable low-pass filtering effect during the encoding-decoding process leads to the loss of a large amount of high-frequency structural details (such as target boundaries and texture abrupt changes). Especially under severe weather conditions, the superposition of visible light texture blurring and infrared thermal response easily results in problems such as edge smoothing and detail degradation in the fused image, making it difficult to balance target saliency and structural integrity.

[0065] To address this core bottleneck, this paper proposes a frequency domain detail enhancement and cross-layer propagation mechanism. Its greatest innovation lies in directly integrating the Dynamic High Freq Extractor (DHFE) into each layer of the main fusion decoder, rather than using independent frequency branches or performing frequency domain processing only at the input. Through a closed-loop design of "post-decoding high-frequency extraction—gated residual fusion—cross-layer progressive propagation—final spatial attention weighting," adaptive supplementation and depth restoration of high-frequency details are achieved. This design significantly differs from existing frequency domain fusion methods (such as independent frequency branches or fixed threshold filtering), enabling dynamic adaptation to different degradation levels in multimodal fusion scenarios, ensuring that the fusion result possesses both strong semantic expression and fine texture.

[0066] In S3, the frequency domain detail enhancement and cross-layer propagation decoding module includes a first decoding layer to a fourth decoding layer connected in sequence, namely the third Transformer EFFN Block, the fourth Transformer EFFN Block, the third Transformer Based Block, and the fourth Transformer Based Block. The output of each decoding layer is connected to the current layer's Dynamic High Freq Extractor (DHFE), and the first layer DHFE to the fourth layer DHFE are connected in sequence. The specific workflow of the frequency domain detail enhancement and cross-layer propagation decoding module is as follows: G1. After decoding at each decoding layer, the current layer's decoded features are input into the current layer's dynamic high-frequency feature extractor to generate high-frequency residuals that are passed from the current layer to the previous layer. These residuals are then accumulated with the high-frequency details from the previous layer. Through a gating mechanism and cross-layer propagation path, layer-by-layer accumulation compensation is achieved, resulting in the accumulated i-th feature. High-frequency details; Among them, the generation from the first Layer pass to the first The high-frequency residuals of the layer are obtained by summing the results of the first layer. The process of detailing high-frequency information in a layer is as follows: Let the first layer be... Layer decoding features are , No. Layer dynamic high-frequency feature extractor from The high-frequency details extracted are ,pass Convolution generates the first Layer gate weight graph : In the formula, For feature-level indexing, For the first Layer decoding features (i.e., features after upsampling + skip connection + decoder block are completed). This is a 1×1 convolution operation (generating a gating system). This represents the Sigmoid activation function. The value range is [0,1], used to control the intensity of high-frequency detail injection. The result is obtained through gating weight graph fusion. Enhanced features after layer fusion : In the formula, This represents element-wise multiplication (Hadamard product). The gating design enables the network to dynamically adjust the injection ratio of high-frequency details based on the semantic strength of the current feature, avoiding excessive amplification of noise. To achieve detail transfer from deep to shallow layers, this invention constructs a cross-layer propagation path: [The path is then described in the original text, which is incomplete and lacks context. A more accurate translation would require the full text.] Enhanced features after layer fusion pass Channel projection and bilinear upsampling generate from the first... Layer pass to the first High-frequency residuals of the layer ; In the formula, Bilinear upsampling (scale_factor=2) Used with the The high-frequency residuals of the layers are accumulated, and the accumulated values ​​are the first layer. High-frequency details The specific expression is: In the formula, For the first Layer gate weight map For the first High-frequency details output by the layer dynamic high-frequency feature extractor; G2, for the first Layer decoding features And after accumulation, the first High-frequency details Line splicing, through Convolution generates spatial attention maps : In the formula, Values , express Convolution; Calculate the final fused image : This spatial attention weighting mechanism further enhances the high-frequency response of the target area while suppressing residual background interference.

[0067] Through the above-mentioned post-decoding frequency domain enhancement and cross-layer propagation framework, the network achieves deep synergy between high-level semantic information and multi-scale high-frequency details, significantly improving the edge sharpness and texture fidelity of the fused image, which constitutes the core innovation of this paper that distinguishes it from traditional fusion methods.

[0068] like Figure 9 As shown, the dynamic high-frequency feature extractor is responsible for accurately extracting high-frequency details from the frequency domain. Its implementation is entirely based on the adaptive energy threshold mechanism designed in this paper. The specific method by which the dynamic high-frequency feature extractor calculates high-frequency details based on the input decoded features is as follows: H1. Perform a two-dimensional Fourier transform on the input decoding features and then shift the frequency to obtain the frequency domain features after frequency shifting. : In the formula, This is a frequency domain rearrangement operation. Decoding features for input (arbitrary layers) ), Then, the amplitude spectrum is calculated and averaged across channels to obtain the amplitude spectrum after cross-channel averaging. : H2. Calculate the total frequency domain energy of the current feature map: In the formula, These are the row coordinates of the two-dimensional frequency domain feature map. The column coordinates of the two-dimensional frequency domain feature map; given the energy parameters of this layer. (i.e., the target proportion of low-frequency energy to be preserved), the energy parameters of the current layer from layer 4 to layer 1 are set to 0.7, 0.2, 0.3 and 0.1 respectively, with deeper layers biased towards suppressing the background and shallower layers biased towards preserving details, and the minimum radius of the low-frequency region in the dynamic search center is determined. To satisfy: In the formula, The coordinates of the frequency domain center are (H / 2, W / 2). The x-coordinate of the frequency domain center is... Using the ordinate of the frequency domain center, a binary mask is constructed. : After applying a binary mask to the frequency domain features, the high-frequency details are output through inverse frequency shift and inverse Fourier transform. : In the formula, This is the inverse Fourier transform (iFFT). It is a frequency domain inverse shift function. For taking the model. When the feature size is too small ( or To avoid artifacts, return zero tensor directly.

[0069] The dynamic high-frequency feature extractor design eliminates the need for a preset fixed cutoff frequency. Instead, it adaptively determines the filtering range based on the actual spectral energy distribution of each feature map, exhibiting stronger robustness and scene adaptability compared to existing fixed threshold or static high-pass methods. Combining cross-layer propagation and gating mechanisms, the dynamic high-frequency feature extractor achieves progressive enhancement from coarse-scale background suppression to fine-scale detail restoration, providing rich structural information for the final fused image.

[0070] Through frequency domain detail enhancement and cross-layer propagation mechanisms, the network effectively recovers the fine texture of the visible light mode while maintaining the saliency of the target thermal information, forming a complete fusion process that is highly complementary to the dual-branch coding and cross-modal interaction modules.

[0071] To verify the effectiveness of the present invention, the following experimental data are provided in this embodiment: The fusion results of this invention were compared with seven other fusion methods: DDFM (Dual-Domain Feature Fusion Method), DRMF (Deep Robust Multimodal Fusion Method), EMMA (Enhanced Multimodal Attention Method), LRRNet (Low-Rank Representation Network), SegMiF (Semantic Segmentation Guided Fusion Method), Text-IF (Text-Guided Image Fusion Method), and Text-DiFuse (Text-Guided Diffusion Fusion Method). In order to comprehensively evaluate the effectiveness of the proposed fusion method, this embodiment conducted an in-depth quantitative analysis.

[0072] The quantitative analysis results strongly corroborated the qualitative observations. The results are summarized in Tables 1 and 2. The method achieved leading or highly competitive performance on most key indicators, including CLIP-IQA, MUSIQ, TReS (Take-Reference Image Quality Assessment), standard deviation (SD), and information entropy (EN).

[0073] Table 1 Comparison Experiments with Individual Severe Weather Table 2 Comparison Experiment of Mixed Severe Weather In single severe weather scenarios, various contrast methods exhibit performance fluctuations across different metrics. In rainy conditions, some methods maintain good structural information, but still fall short in overall perceptual quality or semantic consistency. In foggy and low-light scenarios, due to decreased contrast and blurred details, some methods struggle to simultaneously capture texture detail and target information. In contrast, the proposed method maintains relatively stable performance across all three scenarios, demonstrating significant advantages in perceptual quality metrics while also maintaining high levels of structural information and image contrast. Therefore, the proposed fusion framework can more effectively mine complementary information between infrared and visible light images, thereby achieving a more balanced fusion effect under different degradation conditions.

[0074] In mixed severe weather scenarios, the superposition of multiple degradation factors further increases the difficulty of the fusion task. As can be observed from Table 2, most of the comparative methods show a significant decline in performance under combined degradation environments, especially exhibiting large fluctuations in semantic consistency and overall perceptual quality. However, the proposed method still maintains leading performance in both mixed degradation scenarios, indicating that it has stronger robustness and adaptability in complex environments. This advantage is mainly due to the language-visual degradation cue mechanism introduced into the network, which enables the model to dynamically adjust the feature fusion strategy according to different degradation types, thereby more effectively preserving target information in infrared images and structural details in visible light images.

[0075] The proposed method not only achieves stable advantages on multiple evaluation metrics, but also demonstrates good consistency across different scenarios. This indicates that the proposed fusion framework can achieve more reliable information integration under complex environmental conditions, thereby generating fused images with higher visual quality and richer information expression capabilities.

[0076] In the description of this invention, the above are merely preferred embodiments and are not intended to limit the scope of protection of this invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. An infrared-visible image fusion method based on text semantic guidance and frequency domain compensation, characterized in that, Includes the following steps: S1. Input the infrared image and the visible light image into the dual-branch feature encoding module. In the encoding stage, the spatial context enhancement strategy is used to strengthen the feature expression capability and generate multi-scale features of the infrared image and the visible light image. S2. Input the multi-scale features of infrared and visible light images into the cross-modal interactive fusion module, establish the correlation between infrared and visible light features through the attention mechanism, and obtain the fused feature map; S3. Input the fused feature map into the frequency domain detail enhancement and cross-layer propagation decoding module. In the decoding stage, adaptive supplementation and depth restoration of high-frequency information are performed to generate the final fused image.

2. The infrared-visible image fusion method based on text semantic guidance and frequency domain compensation according to claim 1, characterized in that, In S1, the dual-branch feature encoding module includes encoders for infrared images and visible light images. The two encoders have the same structure, each including a first Transformer Based Block, a second Transformer Based Block, a first Transformer EFFN Block, and a second Transformer EFFN Block connected in sequence. The first Transformer EFFN Block and the second Transformer EFFN Block both contain enhanced frequency domain feedforward networks. Spatial channel attention modules are set after the second Transformer Based Block, the first Transformer EFFN Block and the second Transformer EFFN Block to improve the spatial structure modeling ability and contextual semantic awareness ability of feature representation.

3. The infrared-visible image fusion method based on text semantic guidance and frequency domain compensation according to claim 2, characterized in that, The specific workflow of the enhanced frequency domain feedforward network is as follows: A1. Input characteristics of enhanced frequency domain feedforward networks Perform layer normalization, and through Convolution and Depthwise separable convolutions generate the first intermediate features separately. Second intermediate features ; In the formula, , Represents the space of real numbers. Indicates batch size, Indicates the number of feature channels. For layer normalization, Indicates the height of the feature map. Indicates the width of the feature map. Indicates the kernel size as The first convolution weight matrix, Indicates the kernel size as The second convolution weight matrix is ​​used for channel mapping. Indicates the kernel size as The first depthwise separable convolution, Indicates the kernel size as The second depthwise separable convolution is used to introduce local spatial context information; A2. To further utilize shallow structural information to construct spatially modulated signals, the frequency domain feedforward network is enhanced to extract spatial context information from auxiliary features and generate spatial modulation features. ; In the formula, This indicates auxiliary features originating from shallow or auxiliary branches. This represents the average pooling operation, used to extract global statistics. This represents a mapping function consisting of convolutional layers, normalized layers, and nonlinear activation functions. This indicates an upsampling operation used to recover the spatial size of features, which modulates the spatial features. With the first intermediate feature The data is concatenated along the channel dimension, and gating weights are generated. : In the formula, This indicates a channel splicing operation. The first convolutional mapping weights, The activation function for the Gaussian error linear unit; A3. Obtain the local frequency domain spatial enhancement features of the output of the enhanced frequency domain feedforward network through element-wise multiplication. ; In the formula, This is an element-wise multiplication operation.

4. The infrared-visible image fusion method based on text semantic guidance and frequency domain compensation according to claim 3, characterized in that, The spatial channel attention module includes interconnected high- and wide-dimensional gating attention modules and channel attention gating modules. The specific workflow of the spatial channel attention module is as follows: B1. Input features of the spatial channel attention module The input is then fed into a height-width-dimensional gated attention module, where global average pooling is performed along both the height and width directions to obtain contextual descriptions in two directions, including the feature descriptions compressed along the height direction. Feature description obtained by compression along the width direction ; In the formula, , This represents the average pooling operation along the height direction, with an output size of... , This represents an average pooling operation along the width direction, with an output size of... ; The context description is divided into several sub-feature groups along the channel dimension, and one-dimensional deep convolution with different receptive fields is applied to each group to generate sub-features after deep convolution, including the first sub-feature representation. Second sub-feature representation ; In the formula, express No. Features of individual channel groups express No. Features of individual channel groups , The total number of channel sub-feature groups; Indicates the kernel size as One-dimensional depthwise convolution operation, ; After concatenating all sub-features along the channel dimension, spatial attention weights, including horizontal spatial attention weights, are generated through grouping normalization and a sigmoid activation function. Spatial attention weight in the vertical direction ; In the formula, for No. Features of individual channel groups for No. Features of individual channel groups for No. Features of individual channel groups for No. Features of individual channel groups This indicates a channel splicing operation. This indicates a grouping normalization operation. This represents the Sigmoid activation function; Computational Space Augmentation Features , as the output feature of the high- and wide-dimensional gated attention module; B2. Input the spatially enhanced features into the channel attention gating module, and compress the features through spatial pooling to obtain a low-resolution representation. : In the formula, express Space pooling operation, through Convolution generates query, key, and value vectors: In the formula, Generate a general attention query vector for the current fused features. Generic attention key vector generated for the current fused features. A generic attention value vector generated for the current fused features. The weights are the mapping weights for the second convolution. The weights are mapped to the third convolution. For the fourth convolution mapping weights, calculate channel self-attention. The specific expression is: In the formula, Represents the normalization function. The channel weights are generated using global pooling and the sigmoid activation function, which acts as a scaling factor for the channel dimension. These weights are then modulated to obtain the spatially enhanced feature input and output features of the channel attention gating module. This is used as the output feature of the spatial channel attention module; In the formula, This indicates a global average pooling operation.

5. The infrared-visible image fusion method based on text semantic guidance and frequency domain compensation according to claim 4, characterized in that, In S2, the cross-modal interaction fusion module includes a difference-aware feature selector module, a text cross-scale fusion module, an enhanced feature fusion module, and a spatial structure enhancement module, which are connected in sequence.

6. The infrared-visible image fusion method based on text semantic guidance and frequency domain compensation according to claim 5, characterized in that, The workflow of the difference-aware feature selector module is as follows: C1. Perform channel-by-channel mapping on the multi-scale features of the infrared and visible light images to obtain the initial visible light features after channel mapping. and infrared initial features : In the formula, For visible light modes Convolution mapping function, For infrared mode The convolution mapping function is then used to calculate the joint response features of the two modalities. : C2. Generating modulation features through local aggregation operations : In the formula, This represents depthwise separable convolution. express Convolution operation, For layer normalization, The Sigmoid activation function is used, followed by the generation of dynamic weights for both modalities. : In the formula, Represents the dynamic weighting coefficients of the visible light modes. Represents the dynamic weighting coefficients of the infrared modes. Represents a linear mapping. This is the normalization function; C3. Perform a Fourier transform to generate a frequency domain representation of the visible light characteristics. Frequency domain representation of infrared features : In the formula, This represents a Fourier transform, through conditional weighting, to obtain a dynamic filter, including dynamic frequency filters for visible light modes. Dynamic frequency filter for infrared modes : In the formula, For the first One basic filter, Indicates the number of filters. The dynamic weighting coefficients of the visible light modes The generated filter combination coefficients; The dynamic weighting coefficients of the infrared modes The generated filter combination coefficients are multiplied point-by-point in the frequency domain to obtain the enhancement features of the two modes, which are then used as the output of the difference-aware feature selector module. These enhancement features include visible light enhancement features. and infrared enhancement features : In the formula, For inverse Fourier transform, This is element-wise multiplication.

7. The infrared-visible image fusion method based on text semantic guidance and frequency domain compensation according to claim 6, characterized in that, The text cross-scale fusion module calculates semantically enhanced visual features based on the enhanced features of the two input modalities, including semantically enhanced visible light modal features and infrared modal features. The methods for calculating semantically enhanced visible light modal features and infrared modal features are the same. Specifically, the method for calculating semantically enhanced visible light modal features is as follows: D1. Map the enhanced features of the two modalities to construct the query, key, and value in the attention mechanism, and calculate the cross-modal attention weights from visible light to infrared. : In the formula, The query vector generated for visible light enhancement features. The key vector generated for infrared enhancement features, For attention feature dimension, The transpose symbol is used; infrared features are weighted and aggregated using attention weights to obtain cross-modal interactive features. : In the formula, Value vectors generated for infrared enhancement features; D2. Construct a cue set using several learnable cue vectors, and dynamically generate semantic cue representations based on text semantic features. : In the formula, For learnable cue vectors, This represents the set of Top-k hints with the highest semantic similarity. Text features For the prompt vector The weight, , For text feature dimensions; By fusing cross-modal attention features with semantic cue information, semantically enhanced visible light modal features are obtained. : 。 8. The infrared-visible image fusion method based on text semantic guidance and frequency domain compensation according to claim 7, characterized in that, The workflow of the enhanced feature fusion module is as follows: E1. Construct mode-shared features based on the semantically enhanced visible light mode features and infrared mode features. Modal difference characteristics : In the formula, These are the semantically enhanced infrared modal features. Represents the convolution operation; generates a conditional vector through global average pooling. : In the formula, Indicates global average pooling. This represents a multilayer perceptron; E2. Modulate the frequency features using a dynamic filter to generate enhanced shared features. and enhanced differential features : In the formula, For the frequency domain representation of shared features, , The frequency domain representation of the difference features. , Represents the condition vector The generated shared feature filter, Represents the condition vector Generate the difference feature filter; then calculate the fusion feature. : In the formula, This represents a learnable parameter used to balance the contribution of differential features to the fusion result.

9. The infrared-visible image fusion method based on text semantic guidance and frequency domain compensation according to claim 8, characterized in that, The specific workflow of the space structure enhancement module is as follows: F1. For the fusion features of the input spatial structure enhancement module, expand them into a sequence form and perform spatial enhancement to obtain the local attention output features. : In the formula, Generate a general attention query vector for the current fused features. Generic attention key vector generated for the current fused features. A generic attention value vector generated for the current fused features. For local spatial attention weights, It is the transpose symbol; F2. To capture long-range dependencies, a global semantic prototype is further introduced. : In the formula, For the first 1 semantic prototype vector, and construct global cross-attention output features. : In the formula, Generate a general attention query vector for the current fused features. The key vector is obtained by mapping the global semantic prototype set. A value vector obtained by mapping the global semantic prototype set; F3. Calculate the spatial enhancement features after fusion. The feature map is then restored to a two-dimensional feature map, resulting in a spatially enhanced fused feature map. : In the formula, This is a tensor rearrangement operation.

10. The infrared-visible image fusion method based on text semantic guidance and frequency domain compensation according to claim 9, characterized in that, In S3, the frequency domain detail enhancement and cross-layer propagation decoding module includes a first decoding layer to a fourth decoding layer connected in sequence, namely the third Transformer EFFN Block, the fourth Transformer EFFN Block, the third Transformer Based Block, and the fourth Transformer Based Block. The output of each decoding layer is connected to the dynamic high-frequency feature extractor of the current layer, and the dynamic high-frequency feature extractors of the first layer to the fourth layer are connected in sequence. The specific workflow of the frequency domain detail enhancement and cross-layer propagation decoding module is as follows: G1. After performing decoding operations at each decoding layer, the decoded features of the current layer are input into the dynamic high-frequency feature extractor of the current layer to generate high-frequency residuals passed from the current layer to the previous layer. These residuals are then accumulated with the high-frequency details of the previous layer, and this accumulation is performed layer by layer to obtain the accumulated high-frequency residual. High-frequency details; Among them, the generation from the first Layer pass to the first The high-frequency residuals of the layer are obtained by summing the results of the first layer. The process of detailing high-frequency information in a layer is as follows: Let the first layer be... Layer decoding features are , No. Layer dynamic high-frequency feature extractor from The high-frequency details extracted are ,pass Convolution generates the first Layer gate weight graph : In the formula, For feature-level indexing, For the first Layer decoding features, This is a 1×1 convolution operation. Represents the Sigmoid activation function; through gated weight graph fusion, the th... Enhanced features after layer fusion : In the formula, This represents element-wise multiplication; To achieve the transmission of details from deep to shallow layers, a cross-layer propagation path is constructed: [The path is then used to] transmit the first layer's details to the shallow layer's details. Enhanced features after layer fusion pass Channel projection and bilinear upsampling generate from the first... Layer pass to the first High-frequency residuals of the layer ; In the formula, For bilinear upsampling, Used with the The high-frequency residuals of the layers are accumulated, and the accumulated values ​​are the first layer. High-frequency details The specific expression is: In the formula, For the first Layer gate weight map For the first High-frequency details output by the layer dynamic high-frequency feature extractor; G2, for the first Layer decoding features And after accumulation, the first High-frequency details Line splicing, through Convolution generates spatial attention maps : In the formula, Values , express Convolution; Calculate the final fused image : The dynamic high-frequency feature extractor calculates high-frequency details based on the input decoded features as follows: H1. Perform a two-dimensional Fourier transform on the input decoding features and then shift the frequency to obtain the frequency domain features after frequency shifting. : In the formula, This is a frequency domain rearrangement operation. The decoding features are the input. Then, the amplitude spectrum is calculated and averaged across channels to obtain the amplitude spectrum after cross-channel averaging. : H2. Calculate the total frequency domain energy of the current feature map: In the formula, These are the row coordinates of the two-dimensional frequency domain feature map. The column coordinates of the two-dimensional frequency domain feature map; given the energy parameters of this layer. Minimum radius of the low-frequency region of the dynamic search center To satisfy: In the formula, The x-coordinate of the frequency domain center is... The ordinate of the frequency domain center is used; a binary mask is then constructed. : After applying a binary mask to the frequency domain features, the high-frequency details are output through inverse frequency shift and inverse Fourier transform. : In the formula, This is the inverse Fourier transform (iFFT). It is a frequency domain inverse shift function. For taking the mold.