Infrared and visible light image fusion method based on cross-level perception
By employing a dynamic feature balancing design with a three-branch encoder and a hierarchical perception fusion module, the real-time performance and robustness issues of infrared and visible light image fusion in autonomous driving were resolved. This enabled efficient adaptive fusion of multimodal features and improved the perception capabilities of the autonomous driving system.
Patent Information
- Application Number
- CN202510904557.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-10-31
AI Technical Summary
Existing deep learning methods are insufficient to meet the real-time, robust, and scene adaptability requirements of infrared and visible light image fusion in autonomous driving. The advantages of multimodal complementarity are difficult to fully exploit, and cross-level key features cannot be adaptively focused, resulting in an imbalance between the consistency of the global traffic scene and the fidelity of local key target details in the fused image.
A three-branch encoder is used to extract infrared intensity features, visible light gradient features, and cross-modal interaction features respectively. Redundant information is dynamically filtered through hardware-aware SSM blocks, and the contribution of feature map information is autonomously determined by learnable weights in the hierarchical perception fusion module. Combined with an adaptive gating unit to adjust the fusion weights, a dynamic balance between shallow details and deep semantic features is achieved.
It significantly improves the global consistency and local detail fidelity of fused images in complex scenes, enhancing the real-time performance, robustness, and cross-scene generalization ability of autonomous driving perception systems.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of image fusion technology, specifically to a method for fusion of infrared and visible light images based on cross-level perception. Background Technology
[0002] In the context of autonomous driving technology striving for safe and reliable operation in all weather and all scenarios, overcoming the bottleneck of perception in extreme environments is crucial. Dense fog, complete darkness, and glare from strong light severely restrict traditional perception systems that rely on a single visible light camera. Infrared and visible light image fusion technology, due to its complementary characteristics, has become a key direction for overcoming this bottleneck: Infrared imaging relies on thermal radiation characteristics and can clearly present heat sources in complete darkness, smoke-covered conditions, or when targets are camouflaged, but it has low spatial resolution and lacks texture; visible light imaging provides rich details and color information consistent with human vision, but it is easily affected by changes in lighting and severe weather. Current image fusion methods are mainly divided into traditional methods based on the spatial domain and methods based on deep learning. Traditional methods rely on manually designed fixed fusion rules, which are effective in specific scenarios but are difficult to adapt to complex dynamic environments and have limited control over feature interactions. In contrast, deep learning-based methods construct deep neural networks to automatically learn complex mappings from multimodal inputs to high-quality fusion outputs, achieving adaptive deep integration of cross-modal information. Mainstream approaches have formed three paradigms based on architectural innovation: convolutional neural network architecture focuses on local feature extraction and spatial correlation modeling; the Transformer architecture utilizes self-attention to establish global contextual dependencies, effectively modeling cross-modal semantic relationships; and generative adversarial networks drive the fusion results to approximate the joint distribution characteristics of multimodal data through adversarial training between the generator and discriminator. These methods, by designing cross-modal feature interaction modules, multi-scale fusion strategies, and task-driven optimization objectives, have achieved a paradigm shift from shallow pixel overlay to deep semantic collaboration, significantly improving the balance between target saliency preservation and detail restoration in complex scenes, and providing reliable cross-modal data for intelligent perception systems.
[0003] In existing technologies, driven by the urgent need for reliable perception in all weather and all scenarios for autonomous driving, infrared and visible light image fusion technology has become a key path to overcome the bottleneck of environmental perception. However, existing deep learning methods still struggle to meet the stringent real-time, robustness, and scene adaptability requirements of autonomous driving. Specifically, while CNN architecture can capture local road details, its limited receptive field leads to insufficient understanding of the global structure and semantic relationships in complex traffic scenes, affecting overall environmental situational awareness. Although Transformer architecture can model long-range dependencies, its inherent high computational complexity results in significant inference latency, making it difficult to support the real-time perception and decision-making needs of vehicles traveling at high speeds. At the level of cross-modal feature interaction, CNN does not sufficiently extract the significant differences between infrared thermal radiation features and visible light reflection features, limiting the deep utilization of complementary information of key targets; at the same time, Transformer's global attention mechanism easily introduces a large amount of redundant computation and may over-smooth or weaken fine features that are crucial to autonomous driving safety, such as pedestrian contours, lane line edges, and traffic sign textures. This makes it difficult for existing methods to efficiently balance the integrity of multi-scale feature representation, resulting in an imbalance between the consistency of the global traffic scene and the fidelity of details of key local targets in the fused image. More critically, the lack of a scene-aware dynamic adaptive mechanism makes it impossible to adjust the fusion strategy in real time according to the complex and ever-changing driving environment. In addition, the model performs poorly in suppressing real-world road noise, and its performance is highly dependent on the coverage of extreme conditions by limited training data. It is difficult to accurately fit nonlinear degradation factors such as dense fog, darkness, glare, and dynamic targets, and its cross-scene generalization ability is weak. These defects severely restrict the performance of fused images in terms of information integrity and visual quality, ultimately affecting the reliability and safety of core downstream tasks of autonomous driving, and becoming a major obstacle to the deployment of advanced autonomous driving.
[0004] To address the core issues of insufficiently leveraging the multimodal complementarity advantages and the inability to adaptively focus key cross-level features in infrared and visible light image fusion, this invention proposes an infrared and visible light image fusion method based on a hierarchical perception strategy. To achieve this goal, this invention proposes an infrared and visible light image fusion method based on cross-level perception, comprising six parts: input preprocessing, feature extraction by a three-branch encoder, a visual state space module, a feature enhancement module, a hierarchical perception fusion module, and reconstruction of the fused image. This method extracts infrared intensity features, visible light gradient features, and cross-modal interaction features through a three-branch encoder. In the encoding stage, it replaces traditional convolutional layers with hardware-aware SSM blocks and utilizes a selective scanning mechanism to dynamically filter redundant information, thereby enhancing cross-modal complementarity. Meanwhile, this invention features a specially designed hierarchical perception fusion module that autonomously determines the information contribution of feature maps at different levels through learnable weights, enabling a dynamic balance between shallow detail features and deep semantic features during the fusion process. The adaptive gating unit acts as a "buffer" for modal differences, automatically adjusting the fusion weights of infrared and visible light features through gating coefficients during feature interaction. This prevents strong modal features from excessively suppressing weak modal information while ensuring the synergistic enhancement of key target radiation intensity and texture details. This closed-loop design paradigm of "decoupling-interaction-rebalancing" systematically solves key problems such as multimodal feature representation conflicts and information fusion granularity mismatch. Summary of the Invention
[0005] Based on the aforementioned technical problems, such as the inability of existing deep learning methods to meet the stringent real-time, robustness, and scene adaptability requirements of autonomous driving, the difficulty in fully leveraging the complementary advantages of multimodal learning, and the inability to adaptively focus on key cross-level features, this invention provides a cross-level perception-based infrared and visible light image fusion method. This invention primarily utilizes a three-branch encoder to extract infrared intensity features, visible light gradient features, and cross-modal interaction features, and uses hardware-aware SSM blocks to replace traditional convolutional layers to dynamically filter redundant information. Simultaneously, it leverages learnable weights in the hierarchical perception fusion module to discriminate the contribution of feature map information, and combines this with an adaptive gating unit to adjust the fusion weights, achieving a dynamic balance between shallow details and deep semantic features. This systematically solves the problems of multimodal feature representation conflicts and information fusion granularity mismatch, significantly improving the global consistency and local detail fidelity of the fused image in complex scenes, and enhancing the real-time performance, robustness, and cross-scene generalization ability of the autonomous driving perception system.
[0006] The technical means employed in this invention are as follows: An infrared and visible light image fusion method based on cross-level perception includes the following steps: Infrared and visible light images are acquired and preprocessed to obtain basic features of the infrared and visible light images. These basic features are then passed through a branch encoder to obtain infrared features, visible light features, and cross-modal interaction features. The branch encoder includes an infrared intensity branch, a visible light gradient branch, and a cross-modal interaction branch. These features are then passed through a visual state space module to achieve dynamic feature modeling and spatial enhancement, resulting in output features of the visual state space module. These output features include visually enhanced infrared features, visible light features, and cross-modal interaction features. The output features of the visual state space module are then input to a feature enhancement module to obtain its output features, which also include enhanced infrared features, visible light features, and cross-modal interaction features. Finally, the output features of the feature enhancement module are input to a hierarchical perception fusion module to obtain fused image features. The spatial resolution of the fused image features is restored through skip connections and upsampling, and then decoded. The decoded features are then input to a block expansion module to generate the final fused image.
[0007] Furthermore, the preprocessing of the infrared and visible light images yields basic features of the infrared image and basic features of the visible light image, including: The infrared image and the visible light image are respectively subjected to convolution to increase their dimensionality, thereby obtaining the basic features of the infrared image and the basic features of the visible light image.
[0008] Furthermore, the step of obtaining infrared features, visible light features, and cross-modal interaction features by passing the infrared image basic features and visible light image basic features through a branch encoder includes: The infrared intensity branch performs layer normalization on the basic features of the infrared image to obtain normalized infrared image features, and the visible light gradient branch performs layer normalization on the basic features of the visible light image to obtain normalized visible light image features. Directional information is filtered from the normalized infrared and visible light image features. Spatiotemporal features are captured through bidirectional scanning in both horizontal and vertical directions, and dynamically fused using a gating mechanism to obtain filtered infrared and normalized visible light image features. The basic infrared image features and the filtered infrared image features are then added together and processed via GE. The LU function obtains preliminary infrared features. The basic features of the visible light image and the filtered visible light image features are added together, and the preliminary visible light features are obtained through the GELU function. Global average pooling is performed on the preliminary visible light features, and channel weights are learned through convolution and activated using the ReLU activation function. In the cross-modal interaction branch, the infrared features and visible light features are linearly transformed and layer normalized respectively to obtain query features, key features, and value features. The query features, key features, and value features are used to calculate the first attention weight. The first attention weight is multiplied by the value feature to obtain the cross-modal interaction features.
[0009] Furthermore, the method for calculating the filtered infrared image features is as follows:
[0010]
[0011]
[0012] in, Let h be the state vector of the infrared image at time step t in the horizontal direction, where t is the time step and h is the horizontal direction. This is the state transition matrix of the infrared image in the horizontal direction. Let be the state vector of the infrared image in the horizontal direction at time step t-1. The input matrix for the infrared image. For normalized infrared image features, The output features of the infrared image at time step t in the horizontal direction. This is the output transformation matrix of the infrared image in the horizontal direction. Let v be the state vector of the infrared image at time step t in the vertical direction, where v represents the vertical direction. This is the state transition matrix of the infrared image in the vertical direction. Let be the state vector of the infrared image in the vertical direction at time step t-1. This is the input matrix of the infrared image in the vertical direction. The output features of the infrared image at time step t in the vertical direction are... This is the output transformation matrix of the infrared image in the vertical direction. The feature components that are gated and activated for infrared images. For gated activation functions, For variables related to the gating mechanism of infrared images, For feature dimension related identifiers, Features of the filtered infrared image; The method for calculating the features of the filtered visible light image is as follows:
[0013]
[0014]
[0015] in, Let h be the state vector of the visible light image at time step t in the horizontal direction, where t is the time step and h is the horizontal direction. This is the state transition matrix of a visible light image in the horizontal direction. Let be the state vector of the visible light image in the horizontal direction at time step t-1. The input matrix for the visible light image, For normalized visible light image features, The output features of the visible light image at time step t in the horizontal direction. This is the horizontal output transformation matrix of the visible light image. Let v be the state vector of the visible light image at time step t in the vertical direction, where v is the vertical direction. This is the state transition matrix of a visible light image in the vertical direction. Let be the state vector of the visible light image in the vertical direction at time step t-1. This is the input matrix of the visible light image in the vertical direction. The output features of the visible light image at time step t in the vertical direction are... This is the output transformation matrix of the visible light image in the vertical direction. For the gated activation feature components of a visible light image, To control the activation function, For variables related to the gating mechanism of visible light images, For feature dimension related identifiers, These are the features of the filtered visible light image.
[0016] Furthermore, the step of passing infrared features, visible light features, and cross-modal interaction features through the visual state space module to obtain the output features of the visual state space module includes: The infrared features, visible light features, and cross-modal interaction features are sequentially processed through convolutional layers and normalization layers to obtain processed infrared features, visible light features, and cross-modal interaction features. These processed features are then further processed using horizontal and vertical state transition mechanisms to capture spatiotemporal features, resulting in transferred infrared features, visible light features, and cross-modal interaction features. These transferred features are then sequentially processed through depthwise separable convolutional layers and SiLU activation functions to obtain enhanced infrared features, visible light features, and cross-modal interaction features. Finally, these enhanced features are subjected to linear transformation to obtain the output features of the visual state space module.
[0017] Furthermore, the step of obtaining the output features of the visual state space module through linear transformation of the enhanced infrared features, visible light features, and cross-modal interaction features includes:
[0018]
[0019]
[0020] in, For visually enhanced infrared signatures, For the enhanced infrared signature, Visible light features for visual enhancement To enhance the visible light characteristics, For visually enhanced cross-modal interaction features, for To enhance cross-modal interaction features, It is a linear transformation.
[0021] Further, the step of inputting the output features of the visual state space module to the feature enhancement module to obtain the output features of the feature enhancement module includes: The output features of each level of the visual state space module are processed by local detail capture to extract local detail information from the feature map. The output features of the last level of the visual state space module are then global attention aligned. The aligned output features are then global average pooled, and the average value of the pooled output features in the spatial dimension is calculated to obtain a channel-dimensional vector. This vector is then input into a multilayer perceptron and passed through a sigmoid activation function to obtain a second attention weight. The second attention weight is then used to perform element-wise multiplication on the output features of each level of the visual state space module to obtain the output features of the feature enhancement module.
[0022] Further, the step of inputting the output features of the feature enhancement module into the hierarchical perception fusion module to obtain fused image features includes: The enhanced infrared and visible light features are normalized to obtain preprocessed infrared and visible light features. These preprocessed features are then enhanced using ES2D to obtain ES2D-enhanced infrared and visible light features. Infrared branch attention weights and visible light branch attention weights are calculated based on the ES2D-enhanced features. The enhanced infrared and visible light features are then weighted and fused using these weights to obtain weighted fused features. A learnable parameter is introduced to fuse the weighted fused features with the enhanced infrared features to obtain fused image features.
[0023] Furthermore, the formula for calculating the infrared branch attention weight is as follows:
[0024] in, For the infrared branch attention weights, Visible light features enhanced by ES2D For the learnable parameters of the infrared branch, The feature dimension scaling factor; The formula for calculating the visible light branch attention weight is as follows:
[0025] in, For the infrared branch attention weights, The infrared signature is enhanced by ES2D. For learnable parameters of the visible light branch, This is the feature dimension scaling factor.
[0026] Further, the process of restoring the spatial resolution of the fused image features through skip connections and upsampling, performing decoding, and inputting the decoded features into the block expansion module to generate the final fused image includes: By restoring the spatial resolution of the fused image features through skip connections and upsampling, a feature map with restored resolution is obtained:
[0027] in, To recover the feature map with resolution, This is the resolution feature map for the next time step. To fuse image features; The feature map of the restored resolution is processed by several layers of decoders to obtain the decoded features. The decoded features are then input into the block expansion module to generate the final fused image.
[0028] Compared with the prior art, the present invention has the following advantages: This invention proposes an infrared and visible light image fusion method based on a hierarchical perception strategy. It accurately preserves thermal radiation features by using a three-branch encoder and selective SSM to dynamically scan the direction and filter redundant background noise. The visible light branch combines gradient operators and channel attention to enhance high-frequency textures while suppressing irrelevant details. The cross-modal branch uses an attention mechanism to dynamically align the key regions of both images and resolve modal conflicts.
[0029] This invention addresses the issues of inaccurate cross-modal feature extraction and non-adaptive fusion by proposing a differentiated branch feature extraction strategy. This method designs dedicated processing branches for infrared, visible light, and cross-modal interactions: the infrared branch employs a dynamic filtering mechanism (SSM) to effectively suppress background thermal noise and focus on extracting the core thermal radiation features of the target; the visible light branch fuses gradient operators and channel attention mechanisms to significantly enhance the expressive power of key edges and texture details; and the cross-modal branch dynamically aligns and associates key feature information from infrared and visible light through an attention mechanism, effectively bridging modal differences. This strategy ensures accurate and differentiated extraction of key features for each modality from the outset.
[0030] This invention addresses the issues of imbalanced feature contributions across different levels and inefficient cross-level fusion by proposing a hierarchical perceptual fusion module. In the encoding stage, this module replaces traditional convolution with hardware-aware SSM blocks, efficiently modeling long-range dependencies through multi-directional scanning and gating mechanisms, selectively preserving important features. In the decoding stage, it integrates a lightweight UNet structure and skip connections to achieve efficient fusion of multi-scale features, avoiding detail loss. The core of the module intelligently balances and integrates features from different levels through learnable weights and adaptive gating units, ultimately reconstructing visible light texture details while significantly preserving the salience of infrared targets and substantially improving computational efficiency.
[0031] Based on the above reasons, this invention can be widely applied in fields such as image fusion. Attached Figure Description
[0032] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0033] Figure 1 This is a schematic diagram of the overall structure of the fusion network of the present invention.
[0034] Figure 2 This is a schematic diagram of the visual state space module of the present invention.
[0035] Figure 3 This is a schematic diagram of the feature enhancement module of the present invention.
[0036] Figure 4 This is a schematic diagram of the hierarchical perception fusion module of the present invention.
[0037] Figure 5 The examples show the fusion comparison results of different algorithms on images. Detailed Implementation
[0038] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0039] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0040] To address the feature conflict caused by modal differences and the lack of adaptive selection capability for cross-level features, this invention proposes an infrared and visible light image fusion method based on a hierarchical perception strategy. Figure 1 As shown, the infrared and visible light image fusion method based on a hierarchical perception strategy consists of six steps: input preprocessing, three-branch encoder feature extraction, visual state space module, feature enhancement module, hierarchical perception fusion module, and reconstructed fused image. First, input preprocessing converts the infrared and visible light images into a format suitable for network processing and initially extracts key information. The three-branch encoder feature extraction extracts infrared thermal radiation, visible light texture, and cross-modal interaction features through differential extraction.
[0041] The visual state space module enhances global context by modeling the spatiotemporal dependencies of cross-modal features. The feature enhancement module enhances cross-modal complementarity by aligning multi-scale features. The reconstructed fused image restores details and optimizes image quality by rebuilding the fused image features. Through the above process, this invention effectively restores visible light texture while maintaining the saliency of infrared targets, solving the problems of detail loss and modal conflict in traditional methods.
[0042] like Figure 1-4 As shown, this invention provides a method for fusing infrared and visible light images based on cross-level sensing, the specific steps of which are as follows: S1. Acquire and preprocess infrared and visible light images to obtain basic features of infrared and visible light images.
[0043] The input preprocessing operation converts infrared and visible light source images into a format suitable for network processing and initially extracts key information.
[0044] Step S1 is as follows: S11. The infrared image and the visible light image are respectively increased in dimensionality through convolution to obtain the basic features of the infrared image and the basic features of the visible light image.
[0045] S12. Extract the gradients of the infrared and visible light images using the Sobel operator.
[0046] S2. The infrared image basic features and the visible light image basic features are processed by a branch encoder to obtain infrared features, visible light features and cross-modal interaction features. The branch encoder includes an infrared intensity branch, a visible light gradient branch and a cross-modal interaction branch.
[0047] The three-branch encoder achieves efficient extraction and complementary fusion of multimodal features through differentiated design. Its function can be divided into three feature processing branches: infrared, visible light, and cross-modal. When the three work together, the infrared branch preserves target saliency, the visible light branch enhances background texture, and the cross-modal branch aligns the two through global context. At the same time, the residual U-block structure and multi-scale fusion strategy balance local details and global semantics, ultimately providing high-quality multimodal input for subsequent fusion.
[0048] The infrared intensity branch dynamically filters redundant directional information through residual connections and a selective state-space model (SSM), focusing on extracting target thermal radiation features from infrared images. For example, it uses horizontal / vertical gating mechanisms to enhance key thermal radiation distributions related to the target. The visible light gradient branch combines the Sobel operator to extract texture details and enhances the weight of high-frequency information through a channel attention mechanism. While multi-scale residual block groups progressively extract edge features, layer normalization (LN) stabilizes the training process. The cross-modal interaction branch uses an attention mechanism to dynamically associate infrared features with visible light features as keys and values. It selects texture details in visible light that are related to the infrared target through cross-attention weights, resolving feature conflicts caused by modal differences.
[0049] Specifically, step S2 includes: S21. The infrared intensity branch performs layer normalization on the basic features of the infrared image to obtain normalized infrared image features. The visible light gradient branch performs layer normalization on the basic features of the visible light image to obtain normalized visible light image features. Layer normalization is performed to eliminate scale differences in the input data and improve the stability of subsequent calculations. The visible light gradient branch incorporates the fundamental features of the visible light image. Layer normalization is performed to eliminate scale differences in the input data and improve the stability of subsequent calculations. The formulas for calculating normalized infrared image features and normalized visible light image features are as follows: ,
[0050] in, As a basic feature of infrared images, As the basic features of normalized infrared images, As a fundamental feature of visible light images, For the basic features of normalized visible light images, This is a layer normalization operation.
[0051] S22. Based on the gradients extracted from infrared and visible light images using the Sobel operator, the directional information of the normalized infrared and visible light image features is filtered. Spatiotemporal features are captured by bidirectional scanning in the horizontal and vertical directions. Dynamic fusion is performed through a gating mechanism to obtain the filtered infrared and normalized visible light image features.
[0052] The method for calculating the features of the filtered infrared image is as follows:
[0053]
[0054]
[0055] in, Let h be the state vector of the infrared image at time step t in the horizontal direction, where t is the time step and h is the horizontal direction. This is the state transition matrix of the infrared image in the horizontal direction. Let be the state vector of the infrared image in the horizontal direction at time step t-1. The input matrix for the infrared image. For normalized infrared image features, The output features of the infrared image at time step t in the horizontal direction. This is the output transformation matrix of the infrared image in the horizontal direction. Let v be the state vector of the infrared image at time step t in the vertical direction, where v represents the vertical direction. This is the state transition matrix of the infrared image in the vertical direction. Let be the state vector of the infrared image in the vertical direction at time step t-1. This is the input matrix of the infrared image in the vertical direction. The output features of the infrared image at time step t in the vertical direction are... This is the output transformation matrix of the infrared image in the vertical direction. The feature components that are gated and activated for infrared images. For gated activation functions, For variables related to the gating mechanism of infrared images, For feature dimension related identifiers, These are the features of the filtered infrared image.
[0056] The method for calculating the features of the filtered visible light image is as follows:
[0057]
[0058]
[0059] in, Let h be the state vector of the infrared image at time step t in the horizontal direction, where t is the time step and h is the horizontal direction. This is the state transition matrix of the infrared image in the horizontal direction. Let be the state vector of the infrared image in the horizontal direction at time step t-1. The input matrix for the infrared image. For normalized infrared image features, The output features of the infrared image at time step t in the horizontal direction. This is the output transformation matrix of the infrared image in the horizontal direction. Let v be the state vector of the infrared image at time step t in the vertical direction, where v represents the vertical direction. This is the state transition matrix of the infrared image in the vertical direction. Let be the state vector of the infrared image in the vertical direction at time step t-1. This is the input matrix of the infrared image in the vertical direction. The output features of the infrared image at time step t in the vertical direction are... This is the output transformation matrix of the infrared image in the vertical direction. The feature components that are gated and activated for infrared images. For gated activation functions, For variables related to the gating mechanism of infrared images, For feature dimension related identifiers, Features of the filtered infrared image S23. Add the basic features of the infrared image and the filtered infrared image features together, and obtain the preliminary infrared features through the GELU function. Add the basic features of the visible light image and the filtered visible light image features together, and obtain the preliminary visible light features through the GELU function.
[0060] ,
[0061] in, As a basic feature of infrared images, As a fundamental feature of visible light images, Features of the filtered infrared image These are the features of the filtered visible light image. Preliminary infrared signature, These are preliminary visible light characteristics.
[0062] S24. Perform global average pooling on the initial visible light features, learn the channel weights through 1×1 convolution, and activate using the ReLU activation function:
[0063] Where s is an intermediate variable. These are preliminary visible light characteristics.
[0064] S25. In the cross-modal interaction branch, infrared and visible light features are subjected to linear transformation and layer normalization respectively to obtain query features, key features, and value features. Feature extraction in the cross-modal interaction branch consists of two core steps: feature projection and layer normalization, and cross-modal attention calculation. First, the features of the infrared intensity branch and the features of the visible light gradient branch are subjected to linear transformation and layer normalization respectively. The calculation formulas for query features, key features, and value features are as follows:
[0065]
[0066] in, To query features, To query the weights corresponding to the features, Infrared characteristics, Key features, For value characteristics, The weights are the corresponding key features. It has visible light characteristics. The weights are the values corresponding to the features.
[0067] S26. Calculate the first attention weight using query features, key features, and value features.
[0068]
[0069] in, To query features, As the first attention weight, The feature dimension scaling factor. Key features.
[0070] S27. Multiply the first attention weight by the value feature to obtain the cross-modal interaction feature.
[0071]
[0072] in, As the first attention weight, For value characteristics, This refers to cross-modal interaction features.
[0073] S3. Infrared features, visible light features, and cross-modal interaction features are processed through the visual state space module to achieve dynamic modeling and spatial enhancement of features, resulting in the output features of the visual state space module (VSS). The output features of the visual state space module include visually enhanced infrared features, visible light features, and cross-modal interaction features.
[0074] To address the need for refined processing of multimodal features output by a three-branch encoder to achieve dynamic modeling and spatial enhancement, a visual state space module is proposed. This module refines the multimodal features output by the encoder through three core steps. First, local detail capture is used to extract local detail information from the feature map using 3×3 convolutions, which is then fused with the original features and normalized to enhance feature stability and local expressiveness. Next, the state space model is extended to dynamically model spatiotemporal features through a bidirectional state transition mechanism in both horizontal and vertical directions. A gating fusion strategy is then used to integrate the bidirectional features, generating a more globally relevant feature representation. Finally, the spatial interaction enhancement module employs depthwise separable convolutions and linear transformations to further strengthen the spatial interaction capabilities of the features. Residual connections are combined to preserve the original information, ultimately outputting high-quality, multi-level fused features. This process progressively optimizes features from local to global, achieving dynamic modeling and spatial enhancement of multimodal information. The visual state space operation steps are as follows.
[0075] like Figure 2 As shown, step S3 is as follows: S31. The infrared features, visible light features, and cross-modal interaction features are processed sequentially through a 3×3 convolutional layer and a normalization layer to obtain the processed infrared features, visible light features, and cross-modal interaction features.
[0076] S32. The processed infrared features, visible light features, and cross-modal interaction features are further captured by the state transition mechanism in the horizontal and vertical directions to obtain the transferred infrared features, visible light features, and cross-modal interaction features.
[0077] S33. The transferred infrared features, visible light features, and cross-modal interaction features are sequentially passed through a depth-separable convolutional layer and a SiLU activation function to obtain enhanced infrared features, visible light features, and cross-modal interaction features.
[0078] S34. The enhanced infrared features, visible light features, and cross-modal interaction features are linearly transformed to obtain the output features of the visual state space module. The formula for calculating the output features of the visual state space module is:
[0079]
[0080]
[0081] in, For visually enhanced infrared signatures, For the enhanced infrared signature, Visible light features for visual enhancement To enhance the visible light characteristics, For visually enhanced cross-modal interaction features, for To enhance cross-modal interaction features, It is a linear transformation.
[0082] S4. Input the output features of the visual state space module into the feature enhancement module (AFFM) to obtain the output features of the feature enhancement module. The output features of the feature enhancement module include feature-enhanced infrared features, visible light features, and cross-modal interaction features.
[0083] The feature enhancement module receives and processes the multi-level features output by the hierarchical perceptual fusion module. Located after the hierarchical perceptual fusion module, its core function is to further refine these fused features, including local detail enhancement, global channel alignment, and effective integration of multi-level features. This generates higher-quality, more discriminative feature representations for subsequent tasks such as detection or recognition. Local detail capture enhances the local expressive power of features, global attention alignment ensures reasonable weight allocation between channels, and feature fusion achieves effective integration of multi-level features through learnable weights. This process embodies a multi-level feature optimization strategy from local to global, and from detail to whole.
[0084] like Figure 3 As shown, S4 specifically refers to: S41. Local detail capture processing is applied to the output features of each level of the visual state space module to extract local detail information from the feature maps. The local detail capture module extracts local detail information from each feature map. This process aims to enhance the local expressive power of the features, enabling subsequent processing to better capture fine-grained information.
[0085]
[0086] in, For local details, This is a linear deformable convolution operation. Output features for each level.
[0087] S42. Perform global attention alignment on the output features of the last level of the visual state space module, perform global average pooling on the aligned output features, and calculate the average value of the pooled output features in the spatial dimension to obtain a channel-dimensional vector:
[0088] in, For the output features of each level, It is a vector with one channel dimension, where i and j are index variables of spatial coordinates, used to traverse the height and width directions of the feature map.
[0089] S43. Input the vector into a multilayer perceptron (MLP), and then pass it through a sigmoid activation function to obtain the second attention weights.
[0090]
[0091] Where σ represents the sigmoid activation function, and MLP is a fully connected network used to learn the relationships between channels. It is a vector with one channel dimension. This is the second attention weight.
[0092] S44. Use the second attention weight to perform element-wise multiplication on the output features of each level of the visual state space module to obtain the output features of the feature enhancement module.
[0093]
[0094] in, For the output features of each level, As the second attention weight, This refers to the output features of the feature enhancement module.
[0095] S5. Input the output features of the feature enhancement module into the hierarchical perception fusion module to obtain the fused image features.
[0096] The hierarchical perceptual fusion module receives and processes refined multi-level feature maps output from the feature enhancement module. Within this module, five steps—layer normalization and linear transformation, ES2D enhancement, dynamic weight generation, feature aggregation, and asymmetric mixing—gradually achieve refined processing and efficient fusion of multi-level features. This process not only enhances the spatial interaction capabilities of features but also improves the model's flexibility and robustness through dynamic weights and asymmetric mixing strategies, ultimately outputting high-quality multimodal fused features.
[0097] like Figure 4 As shown, step S5 is as follows: S51. Normalize the enhanced infrared and visible light features respectively to obtain the preprocessed infrared and visible light features.
[0098] S52. Perform ES2D enhancement on the preprocessed infrared and visible light features to obtain ES2D enhanced infrared and visible light features.
[0099] S53. Based on the enhanced infrared and visible light features from ES2D, calculate the attention weights for the infrared branch and the visible light branch. To achieve dynamic feature fusion, calculate the attention weights between the infrared intensity branch and the visible light gradient branch. Based on the infrared intensity branch features... Visible light gradient branching characteristics The attention weights are calculated using the dot product, a process that ensures the infrared intensity branch features can dynamically adjust their weights based on the visible light gradient branch. Similarly, the attention weights are calculated based on the infrared intensity branch features. and visible light gradient branching features The dot product is used to calculate the attention weights, which enables a bidirectional cross-modal attention mechanism, ensuring that the features of the two branches can guide and enhance each other.
[0100] The formula for calculating the attention weight of the infrared branch is:
[0101] in, For the infrared branch attention weights, Visible light features enhanced by ES2D For the learnable parameters of the infrared branch, This is the feature dimension scaling factor.
[0102] The formula for calculating the attention weights of the visible light branch is:
[0103] in, For the infrared branch attention weights, The infrared signature is enhanced by ES2D. For learnable parameters of the visible light branch, This is the feature dimension scaling factor.
[0104] S54. The infrared and visible light features enhanced by the infrared branch attention weights and visible light branch attention weights are weighted and fused to obtain the weighted fused features. This weighted fusion method can dynamically adjust the importance of the two branch features according to the needs of the current task, thereby achieving more flexible and efficient multimodal feature fusion. The formula for calculating the weighted fused features is:
[0105] in, For the infrared branch attention weights, For the infrared branch attention weights, For feature enhancement infrared features, For feature enhancement of visible light features, These are the features after weighted fusion.
[0106] S55. Further optimize the fusion result through an asymmetric mixing strategy. Learnable parameters are introduced to fuse the weighted fused features with the feature-enhanced infrared features, resulting in fused image features.
[0107]
[0108] in, The features after weighted fusion For feature enhancement infrared features, For learnable parameters, To fuse image features.
[0109] S6. The spatial resolution of the fused image features is restored through skip connections and upsampling, and then decoded. The decoded features are input into the Patch Expanding module to generate the final fused image.
[0110] In the decoder and output stages, the spatial resolution of features is restored through skip connections and upsampling. The final fused image is generated using PatchExpanding, and the fusion effect is optimized through a multi-objective loss function. This process ensures that the fused image retains both the thermal radiation information of the infrared image and the details and textures of the visible light image, achieving high-quality multimodal visual information fusion.
[0111] In the field of image fusion, Patch Expanding is a core operation that achieves feature map upsampling through learnable channel reorganization: first, the number of input channels is expanded by r2 times, and then the spatial dimension is reorganized by r×r grid, which increases the feature map resolution by r times and compresses the channels by r2 times, thereby efficiently restoring the resolution in the encoding and decoding architecture and realizing deep fusion of multimodal / multiscale features and pixel-level prediction and reconstruction.
[0112] S6 specifically includes: S61. By using skip connections and upsampling to restore the spatial resolution of the fused image features, multi-scale feature fusion and spatial information recovery are achieved, resulting in a feature map with restored resolution:
[0113] in, To recover the feature map with resolution, This is the resolution feature map for the next time step. To fuse image features.
[0114] S62. After the feature map of the restored resolution is processed by several layers of decoders, the decoded features are obtained. The decoded features are then input into the Patch Expanding module to generate the final fused image.
[0115] To further demonstrate and verify the fusion performance of this invention, six pairs of images were randomly selected from the TNO test set for comparison of fusion effects. The results are as follows: Figure 5 As shown in the figure, early deep learning algorithms performed poorly in fusion, exhibiting problems such as loss of infrared thermal radiation features, significant loss of image structure, and poor visual perception quality. This invention demonstrates superior performance in preserving texture details, improving global contrast, and balancing grayscale distribution in the fused infrared and visible light images. It can generate images with richer details and more accurate target representations, indicating that this invention offers better fusion results and visual experience compared to other algorithms.
[0116] To further verify the effectiveness of this invention, 40 pairs of images from the aforementioned TNO dataset were quantitatively analyzed using seven indicators, and the results are shown in Table 1. The optimal value for each indicator is marked in red, and the second-best value is marked in blue. The U2Fusion and PMGI fused images are affected by some image noise, resulting in a better SF index. The algorithm in this chapter achieves optimal values in four indicators: MI, SD, VIF, and SSIM. The MI index indicates that the fused image retains more information about the source image; SD indicates that the fused result has high contrast and fully integrates the prominent features of the infrared image; the optimal VIF and SSIM indicate that the algorithm in this chapter can minimize structural loss and deformation, and reduce distortion between the fused image and the source image.
[0117] Table 1. Mean values of quantitative evaluation metrics for 40 pairs of fused images in the TNO dataset.
[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for fusing infrared and visible light images based on cross-level perception, characterized in that, Includes the following steps: Infrared and visible light images are acquired and preprocessed to obtain the basic features of the infrared and visible light images. The infrared image basic features and the visible light image basic features are passed through a branch encoder to obtain infrared features, visible light features and cross-modal interaction features. The branch encoder includes an infrared intensity branch, a visible light gradient branch and a cross-modal interaction branch. The infrared features, visible light features, and cross-modal interaction features are passed through the visual state space module to achieve dynamic modeling and spatial enhancement of the features, resulting in the output features of the visual state space module. The output features of the visual state space module include visually enhanced infrared features, visible light features, and cross-modal interaction features. The output features of the visual state space module are input to the feature enhancement module to obtain the output features of the feature enhancement module. The output features of the feature enhancement module include feature-enhanced infrared features, visible light features, and cross-modal interaction features. The output features of the feature enhancement module are input into the hierarchical perception fusion module to obtain the fused image features; The spatial resolution of the fused image features is restored by skip connections and upsampling, and then decoded. The decoded features are input into the block expansion module to generate the final fused image.
2. The infrared and visible light image fusion method based on cross-level perception according to claim 1, characterized in that, The preprocessed infrared and visible light images yield basic features of the infrared and visible light images, including: The infrared image and the visible light image are respectively subjected to convolution to increase their dimensionality, thereby obtaining the basic features of the infrared image and the basic features of the visible light image.
3. The infrared and visible light image fusion method based on cross-level perception according to claim 1, characterized in that, The process involves using a branch encoder to obtain infrared features, visible light features, and... Cross-modal interaction features include: The infrared intensity branch performs layer normalization on the basic features of the infrared image to obtain normalized infrared image features, and the visible light gradient branch performs layer normalization on the basic features of the visible light image to obtain normalized visible light image features. The normalized infrared image features and normalized visible light image features are filtered for directional information. Spatiotemporal features are captured by bidirectional scanning in the horizontal and vertical directions. Dynamic fusion is performed through a gating mechanism to obtain filtered infrared image features and filtered normalized visible light image features. The basic features of the infrared image and the filtered infrared image features are added together, and the preliminary infrared features are obtained by using the GELU function. The basic features of the visible light image and the filtered visible light image features are added together, and the preliminary visible light features are obtained by using the GELU function. Global average pooling is performed on the initial visible light features, and channel weights are learned through convolution, followed by activation using the ReLU activation function. In the cross-modal interaction branch, infrared features and visible light features are subjected to linear transformation and layer normalization respectively to obtain query features, key features and value features; The first attention weight is calculated using query features, key features, and value features; Multiplying the first attention weight by the value feature yields the cross-modal interaction feature.
4. The infrared and visible light image fusion method based on cross-level perception according to claim 3, characterized in that, The method for calculating the features of the filtered infrared image is as follows: in, Let h be the state vector of the infrared image at time step t in the horizontal direction, where t is the time step and h is the horizontal direction. This is the state transition matrix of the infrared image in the horizontal direction. Let be the state vector of the infrared image in the horizontal direction at time step t-1. The input matrix for the infrared image. For normalized infrared image features, The output features of the infrared image at time step t in the horizontal direction. This is the output transformation matrix of the infrared image in the horizontal direction. Let v be the state vector of the infrared image at time step t in the vertical direction, where v represents the vertical direction. This is the state transition matrix of the infrared image in the vertical direction. Let be the state vector of the infrared image in the vertical direction at time step t-1. This is the input matrix of the infrared image in the vertical direction. The output features of the infrared image at time step t in the vertical direction are... This is the output transformation matrix of the infrared image in the vertical direction. The feature components that are gated and activated for infrared images. For gated activation functions, For variables related to the gating mechanism of infrared images, For feature dimension related identifiers, Features of the filtered infrared image; The method for calculating the features of the filtered visible light image is as follows: in, Let h be the state vector of the visible light image at time step t in the horizontal direction, where t is the time step and h is the horizontal direction. This is the state transition matrix of a visible light image in the horizontal direction. Let be the state vector of the visible light image in the horizontal direction at time step t-1. The input matrix for the visible light image, For normalized visible light image features, The output features of the visible light image at time step t in the horizontal direction. This is the horizontal output transformation matrix of the visible light image. Let v be the state vector of the visible light image at time step t in the vertical direction, where v is the vertical direction. This is the state transition matrix of a visible light image in the vertical direction. Let be the state vector of the visible light image in the vertical direction at time step t-1. This is the input matrix of the visible light image in the vertical direction. The output features of the visible light image at time step t in the vertical direction are... This is the output transformation matrix of the visible light image in the vertical direction. For the gated activation feature components of a visible light image, To control the activation function, For variables related to the gating mechanism of visible light images, For feature dimension related identifiers, These are the features of the filtered visible light image.
5. The infrared and visible light image fusion method based on cross-level perception according to claim 1, characterized in that, The infrared features, visible light features and Cross-modal interaction features are obtained through the visual state space module, resulting in the output features of the visual state space module, including: The infrared features, visible light features, and cross-modal interaction features are processed sequentially through convolutional layers and normalization layers to obtain the processed infrared features, visible light features, and cross-modal interaction features. The processed infrared features, visible light features, and cross-modal interaction features are further captured using horizontal and vertical state transition mechanisms to obtain the transferred infrared features, visible light features, and cross-modal interaction features. The transferred infrared features, visible light features, and cross-modal interaction features are sequentially passed through a depth-separable convolutional layer and a SiLU activation function to obtain enhanced infrared features, visible light features, and cross-modal interaction features; The enhanced infrared features, visible light features, and cross-modal interaction features are linearly transformed to obtain the output features of the visual state space module.
6. The infrared and visible light image fusion method based on cross-level perception according to claim 5, characterized in that, The enhanced infrared features, visible light features and Cross-modal interaction features are transformed linearly to obtain the output features of the visual state space module, including: in, For visually enhanced infrared signatures, For the enhanced infrared signature, Visible light features for visual enhancement To enhance the visible light characteristics, For visually enhanced cross-modal interaction features, for To enhance cross-modal interaction features, It is a linear transformation.
7. The infrared and visible light image fusion method based on cross-level perception according to claim 1, characterized in that, The step of inputting the output features of the visual state space module to the feature enhancement module to obtain the output features of the feature enhancement module includes: The output features of each level of the visual state space module are processed by local detail capture to extract local detail information from the feature map; Global attention alignment is performed on the output features of the last level of the visual state space module. The aligned output features are then subjected to global average pooling. The average value of the pooled output features in the spatial dimension is calculated to obtain a vector in the channel dimension. The vector is input into a multilayer perceptron and then passed through a sigmoid activation function to obtain the second attention weights; The second attention weight is used to perform element-wise multiplication on the output features of each level of the visual state space module to obtain the output features of the feature enhancement module.
8. The infrared and visible light image fusion method based on cross-level perception according to claim 1, characterized in that, The step of inputting the output features of the feature enhancement module into the hierarchical perception fusion module to obtain fused image features includes: The enhanced infrared and visible light features are normalized to obtain the preprocessed infrared and visible light features. The preprocessed infrared and visible light features are enhanced by ES2D to obtain the ES2D enhanced infrared and visible light features. Based on the enhanced infrared and visible light features of ES2D, calculate the infrared branch attention weights and the visible light branch attention weights. The infrared and visible light features enhanced by the infrared branch attention weights and visible light branch attention weights are weighted and fused to obtain the weighted fused features. Learnable parameters are introduced to fuse the weighted fused features with the enhanced infrared features to obtain fused image features.
9. The infrared and visible light image fusion method based on cross-level perception according to claim 8, characterized in that, The formula for calculating the infrared branch attention weight is as follows: in, For the infrared branch attention weights, Visible light features enhanced by ES2D For the learnable parameters of the infrared branch, The feature dimension scaling factor; The formula for calculating the visible light branch attention weight is as follows: in, For the infrared branch attention weights, The infrared signature is enhanced by ES2D. For learnable parameters of the visible light branch, This is the feature dimension scaling factor.
10. The infrared and visible light image fusion method based on cross-level perception according to claim 1, characterized in that, The process of restoring the spatial resolution of the fused image features through skip connections and upsampling, performing decoding, and inputting the decoded features into the block expansion module to generate the final fused image includes: By restoring the spatial resolution of the fused image features through skip connections and upsampling, a feature map with restored resolution is obtained: in, To recover the feature map with resolution, This is the resolution feature map for the next time step. To fuse image features; The feature map of the restored resolution is processed by several layers of decoders to obtain the decoded features. The decoded features are then input into the block expansion module to generate the final fused image.
Citation Information
Cited By
Method and system for detecting multiple types of defects of automobile damping rod based on machine vision
CN121453783A