Semantic-guided Mama infrared and visible light image fusion method and system

By constructing an image reconstruction and fusion network and using a loss function to train feature extraction and fusion of infrared and visible light images, the problem of low efficiency and insufficient fusion performance of existing methods in high-resolution image fusion is solved, and high-quality image fusion effect is achieved.

CN121481864APending Publication Date: 2026-02-06GUANGDONG OCEAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610004010.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-05
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing infrared and visible light image fusion methods suffer from low efficiency and insufficient fusion performance in feature extraction and fusion, especially when processing high-resolution images, where computational complexity is high and it is difficult to achieve full interaction and complementary fusion of depth features between infrared and visible light modes.

Method used

An image reconstruction network consisting of a reconstruction encoder and a reconstruction decoder is constructed. The encoder with high-quality feature extraction capability is trained by the first-stage loss function. Then, an image fusion network consisting of a fusion module and a semantically guided dual-branch decoder is constructed and trained using the second-stage loss function to generate a high-quality fused image.

Benefits of technology

It significantly improves the visual effect and objective evaluation index of image fusion, outperforming existing methods. While reducing computational complexity, it improves the fusion quality of infrared thermal target information and visible light texture details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121481864A_ABST
    Figure CN121481864A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to a semantic-guided Mama infrared and visible light image fusion method and system. The method comprises the following steps: constructing an image reconstruction network, and training the image reconstruction network based on a first-stage loss function to obtain a trained image reconstruction network; constructing an image fusion network, wherein the image fusion network comprises a trained reconstruction encoder, a fusion module and a semantic guidance double-branch decoder; keeping the parameters of the reconstruction encoder fixed, and training the image fusion network based on a second-stage loss function to obtain a trained image fusion network; fusing the to-be-fused infrared image and the to-be-fused visible light image by adopting the trained image fusion network to generate a fused image; according to the invention, the image fusion quality can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, specifically to a semantically guided Mamba infrared and visible light image fusion method and system. Background Technology

[0002] Image fusion technology aims to effectively combine image information from different modalities within the same scene to generate a fused image that comprehensively and accurately describes the scene. In infrared and visible light image fusion tasks, images from different modalities provide complementary information: infrared images, by capturing the thermal radiation information of objects, can work effectively in complex environments such as low light and inclement weather, highlighting the temperature difference between thermal targets and the scene; while visible light images can provide rich texture details and color information that conform to human visual perception. However, both have inherent limitations: infrared images typically lack texture details and are easily affected by ambient temperature fluctuations; visible light images experience a sharp decline in performance at night or under low light conditions. Therefore, effectively fusing the two can comprehensively utilize their complementary advantages, significantly improving the perception capability of complex environments in applications such as security monitoring, autonomous driving, and military reconnaissance.

[0003] Existing infrared and visible light image fusion methods can be broadly categorized into two types: traditional methods and conventional methods. Traditional methods (such as those based on multi-scale transformation and sparse representation) typically rely on manually designed fusion rules and prior knowledge. While simple and effective in controlled scenarios, they have poor generalization ability and struggle to meet the complex and ever-changing demands of real-world applications. With the development of deep learning, data-driven methods (such as those utilizing autoencoders, convolutional neural networks (CNNs), and generative adversarial networks (GANs)) have significantly improved the quality of fused images through their powerful feature learning capabilities. These methods can automatically learn deep features in images, thereby generating fusion results that are superior in both detail preservation and feature enhancement.

[0004] Currently, research in this field is shifting from simple low-level pixel fusion to semantically enhanced image fusion. By incorporating semantic information from high-level vision tasks (such as object detection and semantic segmentation), the fusion process can be guided more precisely, making the results more semantically plausible. However, existing semantically driven fusion methods still face two major challenges:

[0005] Firstly, at the feature extraction level, although Transformer-based models are used due to their powerful global context modeling capabilities, the inherent computational complexity of their self-attention mechanism increases quadratically with the size of the input image, which severely restricts their efficiency and application prospects when processing high-resolution infrared and visible light images.

[0006] Secondly, at the feature fusion level, how to design an effective mechanism to achieve full interaction and complementary fusion of deep features between infrared and visible light modes remains a challenge in current research, and there is still considerable room for improvement in the fusion efficiency of existing methods. Summary of the Invention

[0007] To address the aforementioned problems, this invention provides a semantically guided Mamba infrared and visible light image fusion method and system, which can improve the quality of image fusion.

[0008] To achieve the above objectives, the present invention provides the following technical solution:

[0009] On one hand, embodiments of the present invention provide a semantically guided Mamba infrared and visible light image fusion method, the method comprising the following steps:

[0010] An image reconstruction network is constructed and trained based on a first-stage loss function to obtain a trained image reconstruction network; the image reconstruction network includes a reconstruction encoder and a reconstruction decoder.

[0011] An image fusion network is constructed, which includes a trained reconstruction encoder, a fusion module, and a semantically guided dual-branch decoder.

[0012] Keeping the parameters of the reconstruction encoder fixed, the image fusion network is trained based on the second-stage loss function to obtain a trained image fusion network;

[0013] A trained image fusion network is used to fuse the infrared and visible light images to be fused, generating a fused image.

[0014] Optionally, the step of using a trained image fusion network to fuse the infrared image and the visible light image to be fused to generate a fused image includes:

[0015] Acquire the infrared and visible light images to be fused;

[0016] The infrared image and the visible light image are input into the reconstruction encoder, which outputs a multi-layer feature map; the feature map includes an infrared feature map and a visible light feature map.

[0017] The infrared feature maps and visible light feature maps of the multiple layers are input into the fusion module of the corresponding layer for fusion processing to obtain the fused feature maps of each layer;

[0018] The fused feature map is input into a semantically guided dual-branch decoder for decoding to obtain the fused image.

[0019] Optionally, the reconstruction encoder includes an infrared encoder and a visible light encoder, and both the infrared encoder and the visible light encoder include five high-efficiency Mamba feature extraction modules cascaded in sequence;

[0020] The infrared image and the visible light image are input into the reconstruction encoder, and a multi-layer feature map is output; the feature map includes an infrared feature map and a visible light feature map, including:

[0021] The infrared feature map is calculated using the following formula:

[0022] ;

[0023] ;

[0024] The visible light feature map is calculated using the following formula:

[0025] ;

[0026] ;

[0027] in, and These represent the infrared encoder and the visible light encoder at the [missing information - likely a specific timeframe or timeframe]. Feature map of the layer; and These represent the input infrared image and the visible light image, respectively. , , , These represent the image's height, width, and number of channels, respectively. and These represent the first and second parts of the infrared encoder and visible light encoder, respectively. A high-efficiency Mamba feature extraction module for layers; and These represent the wavelet downsampling modules in the infrared encoder and the visible light encoder, respectively. and These represent the feature maps of the infrared encoder and the visible light encoder at layer 1, respectively.

[0028] Optionally, the fusion module is a 5-layer spatial spectral attention mechanism;

[0029] The infrared and visible light feature maps from the multiple layers are input into the fusion module of the corresponding layer for fusion processing to obtain the fused feature maps of each layer, including:

[0030] The fused feature map is calculated using the following formula:

[0031] ;

[0032] in, Indicates the first Layer fusion feature map; Indicates the first Spatial spectral attention mechanism of layers.

[0033] Optionally, the semantically guided dual-branch decoder includes a semantic-aware branch, and the image fusion branch includes five cascaded high-efficiency Mamba feature extraction modules and a high-efficiency upsampling module; the step of inputting the fused feature map into the semantically guided dual-branch decoder for decoding to obtain the fused image includes:

[0034] The fused image is calculated using the following formula:

[0035] ;

[0036] ;

[0037] ;

[0038] in, Indicates a fused image. Indicates the first The fused feature map of the layer after upsampling; This indicates a high-efficiency upsampling module.

[0039] Optionally, the semantically guided dual-branch decoder further includes a semantically aware branch, which is used to output semantic segmentation results, edge segmentation results, and binary segmentation results;

[0040] The calculation formula for the semantic perception branch is:

[0041] ;

[0042] ;

[0043] ;

[0044] ;

[0045] ;

[0046] ;

[0047] ;

[0048] ;

[0049] ;

[0050] ;

[0051] in, Indicates the semantic segmentation result. This represents the result of binary segmentation. This indicates the edge segmentation result; Indicates the fusion encoder at the 1st Features of layer output, This indicates a high-efficiency upsampling module. Indicates upsampling Second-rate, Indicates feature splicing, This indicates element-wise multiplication; This represents the initial multi-scale fusion result of semantic features. For passing through three layers Module-enhanced semantic features This is the result of binary segmentation. This is the semantic segmentation result. This is the result of boundary decomposition.

[0052] Optionally, the formula for the loss function in the first stage is:

[0053] ;

[0054] ;

[0055] ;

[0056] in, This represents the loss function in the first stage. and These represent pixel loss and multi-scale structural similarity loss, respectively. and These represent the output image and the input image, respectively. and These represent the height and width of the image, respectively. This represents a multi-scale structural similarity index.

[0057] Optionally, the formula for the loss function in the second stage is:

[0058] ;

[0059] in, To mitigate image fusion loss and ensure the quality of the fused image; To assist in segmentation loss;

[0060] Fusion loss The formula is:

[0061] ;

[0062] in, The mask-guided fusion loss is formulated as follows:

[0063] ;

[0064] in, Image size, It is a binary mask. This represents element-wise multiplication. Represents the L1 norm;

[0065] The gradient loss is given by the following formula:

[0066] ;

[0067] in, For gradient operators, This indicates taking the absolute value. This indicates taking the maximum of the two values;

[0068] Auxiliary segmentation loss It contains three components:

[0069] ;

[0070] in, For 9-class semantic segmentation loss, This is the edge segmentation loss; The binary segmentation loss is defined as follows:

[0071] ;

[0072] ;

[0073] ;

[0074] in, For the difficult example sample set, and These are the cross-entropy and binary cross-entropy loss functions, respectively. , and For real labels, , and To predict probabilities.

[0075] On the other hand, embodiments of the present invention provide a semantically guided Mamba infrared and visible light image fusion system, comprising:

[0076] At least one processor;

[0077] At least one memory for storing at least one program;

[0078] When the at least one program is executed by the at least one processor, the at least one processor performs the method described above.

[0079] On the other hand, embodiments of the present invention provide a computer-readable storage medium storing a processor-executable program, which, when executed by a processor, is used to perform the above-described method.

[0080] The embodiments of the present invention have the following beneficial effects:

[0081] This invention constructs an image reconstruction network comprising a reconstruction encoder and a reconstruction decoder, with separate reconstruction networks for infrared and visible light images. First, a reconstruction encoder with high-quality feature extraction capabilities is trained based on a first-stage loss function. Then, an image fusion network comprising a reconstruction encoder, a fusion module, and a semantically guided dual-branch decoder is constructed and trained using a second-stage loss function. This enables the image fusion network to fully extract and fuse thermal target information from the infrared image and texture details from the visible light image. Finally, the trained image fusion network is used to fuse the infrared and visible light images to generate a high-quality fused image. Compared with existing fusion methods, this invention demonstrates superior fusion performance in terms of both visual effects and objective evaluation metrics. Attached Figure Description

[0082] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0083] Figure 1 This is a flowchart illustrating the semantically guided Mamba infrared and visible light image fusion method in an embodiment of the present invention.

[0084] Figure 2 This is a network structure diagram of the semantically guided Mamba infrared and visible light image fusion method in this embodiment of the invention;

[0085] Figure 3 This is a structural diagram of the efficient Mamba feature extraction module in an embodiment of the present invention;

[0086] Figure 4This is a structural diagram of the spatial spectral attention mechanism in an embodiment of the present invention;

[0087] Figure 5 This is a structural diagram of the implementation process of the first-stage training strategy in this embodiment of the invention;

[0088] Figure 6 This is a structural diagram of the implementation process of the second-stage training strategy in an embodiment of the present invention;

[0089] Figure 7 This is a schematic diagram of the semantically guided Mamba infrared and visible light image fusion system in an embodiment of the present invention. Detailed Implementation

[0090] The following will provide a clear and complete description of the concept, specific structure, and technical effects of the present invention in conjunction with embodiments and accompanying drawings, so as to fully understand the purpose, solution, and effects of the present invention. It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other.

[0091] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of this invention; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this invention as detailed in the appended claims.

[0092] It is understood that the terms “first,” “second,” etc., used in this invention may be used herein to describe various concepts, but unless specifically stated otherwise, these concepts are not limited by these terms. These terms are used only to distinguish one concept from another. For example, first information may also be referred to as second information without departing from the scope of embodiments of the invention, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to determination” as used herein may be interpreted as “when…” or “when…” or “in response to determination.”

[0093] The terms “at least one,” “multiple,” “each,” “any,” etc., used in this invention, “at least one” includes one, two, or more than two; “multiple” includes two or more than two; “each” refers to each of the corresponding multiple; and “any” refers to any one of the multiple.

[0094] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein is for the purpose of describing embodiments of the invention only and is not intended to limit the invention.

[0095] refer to Figure 1 and Figure 2 ,like Figure 1 The image shown is a semantically guided Mamba infrared and visible light image fusion method provided by an embodiment of the present invention. The method includes the following steps:

[0096] S100, Construct an image reconstruction network, train the image reconstruction network based on the first-stage loss function to obtain a trained image reconstruction network; the image reconstruction network includes a reconstruction encoder and a reconstruction decoder.

[0097] In this example, the main task of the first stage is to train the infrared encoder and the visible light encoder. This stage sets up independent image reconstruction networks for the infrared and visible light images respectively, with the reconstruction encoder and decoder forming a U-shaped structure. By constructing this stage as an image reconstruction task, the aim is to optimize the network parameters so that the reconstructed output image visually approximates the original input image, thereby effectively improving the encoder's ability to extract image features.

[0098] S200, Construct an image fusion network, which includes a reconstruction encoder, a fusion module, and a semantically guided dual-branch decoder from a trained image reconstruction network;

[0099] S300, keeping the parameters of the reconstruction encoder fixed, train the image fusion network based on the second-stage loss function to obtain the trained image fusion network.

[0100] In this example, the goal of the second stage is to construct an image fusion network comprising an infrared encoder, a visible light encoder, a fusion module, and a semantically guided dual-branch decoder. This network is trained using a loss function designed in the second stage, aiming to fully extract and fuse thermal target features from the infrared image with texture details from the visible light image. During this process, the parameters of the infrared and visible light encoders remain fixed; only the fusion module and the semantically guided dual-branch decoder are optimized, thereby achieving high-quality multimodal image fusion while preserving multi-scale feature representation capabilities.

[0101] The S400 uses a trained image fusion network to fuse the infrared and visible light images to be fused, generating a fused image.

[0102] This invention proposes a semantically guided Mamba infrared and visible light image fusion method and system. The method first constructs an image reconstruction network comprising an encoder and a reconstruction decoder, with separate reconstruction networks for infrared and visible light images. First, an encoder with high-quality feature extraction capabilities is trained based on a first-stage loss function. Then, an image fusion network comprising dual encoders for infrared and visible light, a fusion module, and a semantically guided dual-branch decoder is constructed and trained using a second-stage loss function, enabling the network to fully extract and fuse thermal target information from the infrared image and texture details from the visible light image. Finally, the trained image fusion network is used to fuse the input image, generating a high-quality fused image. Compared with existing fusion methods, this invention demonstrates superior fusion performance in terms of visual effects and objective evaluation metrics, exhibiting high application value.

[0103] In some embodiments, S400, the step of fusing the infrared image and the visible light image to be fused using a trained image fusion network to generate a fused image includes:

[0104] S410, acquire the input infrared image and visible light image;

[0105] S420: Input infrared and visible light images into the reconstruction encoder and output multi-layer feature maps; the feature maps include infrared feature maps and visible light feature maps.

[0106] Specifically, the infrared image used as a training sample is input into the infrared encoder, and the output is a multi-layer infrared feature map; the visible light image used as a training sample is input into the visible light encoder, and the output is a multi-layer visible light feature map.

[0107] S430, the infrared feature map and visible light feature map of the multilayer are input into the fusion module of the corresponding layer for fusion processing to obtain the fused feature map;

[0108] S440, the fused feature map is input into the semantically guided dual-branch decoder for decoding processing to obtain the fused image.

[0109] In some embodiments, both the infrared encoder and the visible light encoder consist of five cascaded, high-efficiency Mamba feature extraction modules.

[0110] like Figure 3As shown, the efficient Mamba feature extraction module consists of a first depthwise separable convolutional layer, a hidden state mixer SSD layer (HSM-SSD), a second depthwise separable convolutional layer, and a feedforward network (FFN). The first depthwise separable convolutional layer has a kernel size of 3, a stride of 1, and a padding of 1. The hidden state mixer SSD layer (HSM-SSD) is used to implement long-range dependency modeling based on the state space model. The second depthwise separable convolutional layer has a kernel size of 3, a stride of 1, and a padding of 1. The feedforward network (FFN) consists of two 1×1 convolutional layers. The outputs of each submodule are weighted and fused using learnable layer scaling factors to enhance feature representation capabilities.

[0111] Given a pair of infrared and visible light images, the formula for calculating the infrared feature map for an infrared encoder is:

[0112] ;

[0113] ;

[0114] The formula for calculating the visible light feature map is:

[0115] ;

[0116] ;

[0117] in, and These represent the infrared encoder and the visible light encoder at the [missing information - likely a specific timeframe or timeframe]. Feature map of the layer; and These represent the input infrared image and the visible light image, respectively. , , , These represent the number of channels, height, and width of the image, respectively. and These represent the first and second parts of the infrared encoder and visible light encoder, respectively. A high-efficiency Mamba feature extraction module for layers; and These represent the wavelet downsampling modules in the infrared encoder and the visible light encoder, respectively. and These represent the feature maps of the infrared encoder and the visible light encoder at layer 1, respectively.

[0118] Specifically, the image reconstruction encoder is used to extract multi-scale features from infrared and visible light images, and consists of five layers. The number of feature map channels in each layer is 1, 32, 64, 128, 256, and 512 respectively, with the spatial resolution decreasing layer by layer, representing the original image's... , , , and .

[0119] In some embodiments, the step of inputting the multi-layered stitched image into the fusion module for fusion processing to obtain a fused feature map includes:

[0120] Constructing fusion modules: The fusion modules are located at the jump connections of the network, with a total of 5 locations. They are designed to adjust and balance the contributions of multi-scale feature maps from the infrared and visible light encoders to improve the feature fusion effect.

[0121] The fusion module consists of five layers of spatial spectral attention mechanisms. For example... Figure 4 As shown, the spatial spectral attention mechanism consists of a channel attention layer, a spatial convolutional attention layer, a feature weighting layer, and a fusion convolutional layer. The channel attention layer extracts channel weights through global average pooling and global max pooling. The spatial convolutional attention layer extracts spatial weights through dual-path convolution and spatial feature compression. The feature weighting layer multiplies the input features by the channel and spatial weights respectively. The fusion convolutional layer has a kernel size of 3, a stride of 1, and a padding of 1, and is used to fuse the weighted dual-path features into a single output.

[0122] Given an infrared encoder and a visible light encoder in the first... Infrared and visible light feature maps of the layer ,in , , , These represent the height, width, number of channels, and layer number of the feature map, respectively.

[0123] The formula for the spatial spectral attention mechanism is as follows:

[0124] ;

[0125] in, Indicates the first Layer fusion feature map; Indicates the first Spatial spectral attention mechanism of layers.

[0126] In some embodiments, the decoder and image reconstruction encoder are structurally symmetrical, together forming a U-shaped architecture. For example... Figure 2As shown, the reconstruction decoder consists of five cascaded, high-efficiency Mamba feature extraction modules;

[0127] The image reconstruction process can be represented by the following formula:

[0128] ;

[0129] ;

[0130] ;

[0131] in, Indicates a fused image. Indicates the first The fused feature map of the layer after upsampling; Indicates a high-efficiency upsampling module; symbol This indicates a splicing operation.

[0132] refer to Figure 5 In some embodiments, the training objective of the first stage is to lay the foundation for the subsequent fusion task in the second stage, namely, to obtain an infrared and visible light encoder with strong feature representation capabilities. Given the fundamental differences in information representation between infrared and visible light images, this stage designs independent image reconstruction networks for each modality, forcing the encoder to learn and strengthen its ability to extract key features (such as thermal targets and texture details) of its respective modality. Therefore, the first stage training mainly employs the following loss function:

[0133] Pixel Loss: Pixel loss is used to measure the difference between the fused image and the input image at the pixel level, in order to ensure that the output image is as close as possible to the input image in numerical terms.

[0134] The formula for the loss function in the first stage is:

[0135] ;

[0136] in, This represents the loss function in the first stage. and These represent pixel loss and multi-scale structural similarity loss, respectively.

[0137] The formula for the pixel loss function is:

[0138] ;

[0139] in, and These represent the output image and the input image, respectively. and These represent the height and width of the image, respectively.

[0140] Multi-Scale Structural Similarity Loss (SSIM Loss): While pixel loss can effectively measure the pixel-level numerical differences between the fused image and the input image, it struggles to reflect the perceptual quality of the human eye's multi-scale structure. To address this limitation, we introduce multi-scale structural similarity loss. This loss simulates the multi-scale observation mechanism of the human eye, calculating the similarity of brightness, contrast, and structural information at multiple downsampling scales, and using their weighted product as the final evaluation metric. This provides a more comprehensive measure of the perceptual consistency between the output and input images across multiple resolutions. The formula for the multi-scale structural similarity loss function is:

[0141] ;

[0142] in, This represents a multi-scale structural similarity index.

[0143] like Figure 6 As shown, the second-stage training process aims to fully utilize the multi-scale features extracted by the encoder and integrate the information through the fusion module to generate a high-quality fused image.

[0144] Second-stage loss function design: In this example, the second-stage task is to train the fusion module and the semantically guided dual-branch decoder. During this stage, the parameters of the infrared encoder and the visible light encoder are frozen, and only the fusion module and the semantically guided dual-branch decoder are optimized to ensure that the fusion network can fully utilize multi-scale features to achieve high-quality fused image reconstruction.

[0145] In some embodiments, the formula for the second-stage loss function is:

[0146] ;

[0147] in, To mitigate image fusion loss and ensure the quality of the fused image; To assist in segmentation loss and improve the model's semantic understanding ability.

[0148] Fusion loss It consists of two parts:

[0149] ;

[0150] in, To guide the fusion loss using a mask, a binary mask is used. Selectively supervised image fusion With infrared images or visible light image The difference. The mask-guided fusion loss is defined as follows:

[0151] ;

[0152] in Image size, It is a binary mask. This represents element-wise multiplication. This represents the L1 norm.

[0153] The gradient loss ensures that the fused image retains important edge information from the source images, and is defined as follows:

[0154] ;

[0155] in, For gradient operators, This indicates taking the absolute value. This indicates taking the maximum of the two values.

[0156] Auxiliary segmentation loss It contains three components:

[0157] ;

[0158] in For 9-category semantic segmentation loss, an online hard example mining strategy is adopted; This is the edge segmentation loss; The binary segmentation loss is defined as follows:

[0159] ;

[0160] ;

[0161] ;

[0162] in For the difficult example sample set, and These are the cross-entropy and binary cross-entropy loss functions, respectively. , and For real labels, , and To predict probabilities.

[0163] To further verify the performance of this method, in TNO, M 3 Comparative experiments were conducted on the FD, RoadScene, and three datasets. Tables 1, 2, and 3 present the experimental results of the proposed fusion method on these three datasets and compare it with existing methods to objectively evaluate its performance and effectiveness.

[0164] Table 1: Comparison of quantitative evaluation metrics between the present invention and existing fusion methods on the TNO dataset.

[0165]

[0166] Table 2: Comparison of the present invention and existing fusion methods in M 3 Comparison of quantitative evaluation metrics on the FD dataset.

[0167]

[0168] Table 3: Comparison of quantitative evaluation metrics between the present invention and existing fusion methods on the RoadScene dataset.

[0169]

[0170] The relevant literature in Tables 1, 2, and 3 is as follows: [1]Xydeas CS, Petrovic V. Objective image fusion performance measure[J]. Electronics letters, 2000, 36(4): 308-309. [2]Qu G, Zhang D, Yan P. Information measure for performance of imagefusion[J]. Electronics letters, 2002, 38(7): 313-315. [3]Aslantas V, Bendes E. A new image quality metric for image fusion:The sum of the correlations of differences[J]. Aeu-international Journal ofelectronics and communications, 2015, 69(12): 1890-1896. [4]Wang Q, Shen Y, Jin J. Performance evaluation of image fusiontechniques[J]. Image fusion: algorithms and applications, 2008, 19: 469-492. [5]Chen Y, Blum R S. A new automated quality assessment algorithm forimage fusion[J]. Image and vision computing, 2009, 27(10): 1421-1432. [6]Eskicioglu A M, Fisher P S. Image quality measures and theirperformance[J]. IEEE Transactions on communications, 2002, 43(12): 2959-2965. [7]Sheikh H R, Bovik A C. Image information and visual quality[J].IEEE Transactions on image processing, 2006, 15(2): 430-444. [8]Cui G, Feng H, Xu Z, et al. Detail preserved fusion of visible andinfrared images using regional saliency extraction and multi-scale imagedecomposition[J]. Optics Communications, 2015, 341: 199-209. [9]Hossny M, Nahavandi S, Creighton D. Comments on ‘Informationmeasure for performance of image fusion’[J]. Electronics letters, 2008, 44(18): 1066-1067.

[10] Gao Z, Zhang C. Texture clear multi-modal image fusion with jointsparsity model[J]. Optik, 2017, 130: 255-265.

[11] Li H, Wu X J, Kittler J. MDLatLRR: A novel decomposition methodfor infrared and visible image fusion[J]. IEEE Transactions on ImageProcessing, 2020, 29: 4733-4746.

[12] Li H, Wu X J, Kittler J. Infrared and visible image fusion usinga deep learning framework[C] / / 2018 24th international conference on patternrecognition (ICPR). IEEE, 2018: 2705-2710.

[13] Li H, Wu X, Durrani T S. Infrared and visible image fusion withResNet and zero-phase component analysis[J]. Infrared Physics&Technology,2019, 102: 103039.

[14] Li H, Wu X J. DenseFuse: A fusion approach to infrared andvisible images[J]. IEEE Transactions on Image Processing, 2018, 28(5): 2614-2623.

[15] Ma J, Yu W, Liang P, et al. FusionGAN: A generative adversarialnetwork for infrared and visible image fusion[J]. Information fusion, 2019,48: 11-26.

[16] Ma J, Zhang H, Shao Z, et al. GANMcC: A generative adversarialnetwork with multiclassification constraints for infrared and visible imagefusion[J]. IEEE Transactions on Instrumentation and Measurement, 2020, 70: 1-14.

[17] Li H, Wu X J, Kittler J. RFN-Nest: An end-to-end residual fusionnetwork for infrared and visible images[J]. Information Fusion, 2021, 73: 72-86.

[18] Xu H, Wang X, Ma J. DRF: Disentangled representation for visibleand infrared image fusion[J]. IEEE Transactions on Instrumentation andMeasurement, 2021, 70: 1-13.

[19] Liu J, Fan X, Huang Z, et al. Target-aware dual adversariallearning and a multi-scenario multi-modality benchmark to fuse infrared andvisible for object detection[C] / / Proceedings of the IEEE / CVF conference oncomputer vision and pattern recognition. 2022: 5802-5811.

[20] Liang P, Jiang J, Liu X, et al. Fusion from decomposition: Aself-supervised decomposition approach for image fusion[C] / / Europeanconference on computer vision. Cham: Springer Nature Switzerland, 2022: 719-735.

[0171] Compared with the prior art, the present invention has the following beneficial effects:

[0172] 1. This invention introduces an efficient Mamba feature extraction module, which, compared with Transformer-based methods, can significantly reduce computational complexity while further improving the model's feature representation capabilities.

[0173] 2. By introducing a spatial spectral attention mechanism, the model can adaptively enhance the response of key features, thereby effectively improving the fusion quality of infrared thermal target information and visible light texture details.

[0174] 3. Through TNO, M 3 A comparison on the FD and RoadScene datasets (a total of 223 pairs of infrared and visible light images) shows that the algorithm proposed in this invention outperforms 11 existing fusion algorithms and achieves state-of-the-art performance in terms of visual quality and objective evaluation metrics.

[0175] See Figure 7 This invention also provides a semantically guided Mamba infrared and visible light image fusion system, comprising:

[0176] At least one processor;

[0177] At least one memory for storing at least one program;

[0178] When the at least one program is executed by the at least one processor, the at least one processor performs the method described above.

[0179] The content of the above method embodiments is applicable to this embodiment. The specific functions implemented in this embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments. Therefore, they will not be repeated here.

[0180] This invention also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0181] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0182] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0183] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0184] This invention also provides a computer program product, including a computer program or computer instructions, which are stored in a memory. A processor of a computer device reads the computer program or computer instructions from the memory and executes the computer program or computer instructions, causing the computer device to perform the above-described method.

[0185] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0186] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0187] It will be understood by those skilled in the art that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically include computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0188] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

Claims

1. A semantic guided Mamba infrared and visible image fusion method, characterized in that, The method comprises the following steps: constructing an image reconstruction network, training the image reconstruction network based on a first stage loss function to obtain a trained image reconstruction network; the image reconstruction network comprises a reconstruction encoder and a reconstruction decoder; An image fusion network is constructed, and the image fusion network comprises the trained reconstruction encoder, a fusion module, and a semantic guidance double-branch decoder; The parameters of the reconstruction encoder are kept fixed, and the image fusion network is trained based on a second stage loss function to obtain a trained image fusion network; The trained image fusion network is used to fuse an infrared image and a visible light image to be fused to generate a fused image.

2. The method of claim 1, wherein, The trained image fusion network is used to fuse an infrared image and a visible light image to be fused to generate a fused image, which comprises: An infrared image and a visible light image to be fused are obtained; The infrared image and the visible light image are input into the reconstruction encoder to output multi-layer feature maps; the feature maps comprise infrared feature maps and visible light feature maps; The multi-layer infrared feature maps and visible light feature maps are input into the corresponding layer fusion module for fusion processing to obtain fusion feature maps of each layer; The fusion feature maps are input into the semantic guidance double-branch decoder for decoding processing to obtain a fused image.

3. The method of claim 2, wherein, The reconstruction encoder comprises an infrared encoder and a visible light encoder, and the infrared encoder and the visible light encoder each comprise five high-efficiency Mamba feature extraction modules which are sequentially cascaded; The infrared image and the visible light image are input into the reconstruction encoder to output multi-layer feature maps; the feature maps comprise infrared feature maps and visible light feature maps, which comprises: The infrared feature maps are calculated by the following formula: ; ; The visible light feature maps are calculated by the following formula: ; ; wherein, and denote the feature maps of the 1st layer in the infrared encoder and the visible light encoder, respectively; denote the input infrared image and the visible light image, respectively, , , , denote the height, width and channel number of the image, respectively; denote the high efficient Mamba feature extraction module of the 1st layer in the infrared encoder and the visible light encoder, respectively; denote the wavelet down-sampling module in the infrared encoder and the visible light encoder, respectively, denote the feature maps of the 1st layer in the infrared encoder and the visible light encoder, respectively.​​​​​​ 4. The method of claim 3, wherein, The fusion module is a five-layer spatial spectrum attention mechanism; The multi-layer infrared feature maps and visible light feature maps are input into the corresponding layer fusion module for fusion processing to obtain fusion feature maps of each layer, which comprises: The fusion feature maps of each layer are calculated by the following formula: ; wherein, represents the fusion feature map of the i-th layer; represents the fusion feature map of the i-th layer; represents the fusion feature map of the i-th layer; represents the spatial-spectral attention mechanism of the i-th layer.

5. The method of claim 4, wherein, The semantic guidance double-branch decoder comprises a semantic perception branch, and the image fusion branch comprises five high-efficiency Mamba feature extraction modules and a high-efficiency up-sampling module which are sequentially cascaded; The fusion feature maps are input into the semantic guidance double-branch decoder for decoding processing to obtain a fused image, which comprises: The fused image is calculated by the following formula: ; ; ; wherein, denotes a fusion image, denotes a first layer of the fusion feature map after up-sampling processing; denotes a high-efficiency up-sampling module.

6. The method of claim 5, wherein, The semantic guidance double-branch decoder further comprises a semantic perception branch, and the semantic perception branch is used to output a semantic segmentation result, an edge segmentation result, and a binary segmentation result; The calculation formula of the semantic perception branch is: ; ; ; ; ; ; ; ; ; ; wherein, denotes a semantic segmentation result, denotes a binary segmentation result, denotes an edge segmentation result; denotes a feature output by a fusion encoder at the layer, denotes a high-efficiency upsampling module, denotes upsampling times, denotes feature splicing, denotes element-wise multiplication; is an initial multi-scale fusion result of semantic features, is a semantic feature after being reinforced by three modules, is a binary segmentation result, is a semantic segmentation result, is a boundary decomposition result.

7. The method of claim 1, wherein, The formula of the first stage loss function is: ; ; ; wherein, represents the first stage loss function, and respectively represent the pixel loss and the multi-scale structural similarity loss; and respectively represent the output image and the input image; and respectively represent the height and the width of the image; represents the multi-scale structural similarity index.

8. The method of claim 1, wherein, The formula of the second stage loss function is: ; wherein, is an image fusion loss, ensuring the quality of the fused image; is an auxiliary segmentation loss; Fusion loss The formula is: ; wherein, is a mask-guided fusion loss, whose formula is: ; wherein, is the image size, is a binary mask, denotes element-wise multiplication, denotes the LI norm; For the gradient loss, the formula is: ; wherein is the gradient operator, denotes taking the absolute value, denotes taking the maximum of both. auxiliary segmentation loss comprises three components: ; wherein, is the 9-class semantic segmentation loss, is the edge segmentation loss; is the binary segmentation loss, respectively defined as: ; ; ; wherein, is a set of difficult example samples, and are cross-entropy and binary cross-entropy loss functions, respectively, , and are true labels, , and are predicted probabilities.

9. A semantically guided Mamba infrared and visible image fusion system characterized by, It comprises: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method of any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. The computer program is executed by the processor to implement the method of any one of claims 1 to 8.