Visible light and infrared image fusion method based on scene-semantic dissociation
By adopting a multi-level encoding-fusion network and meta-feature embedding module based on scene-semantic dissociation in image fusion, the problems of insufficient cross-modal feature extraction and neglected scene information are solved, and richer feature fusion and optimized performance of downstream tasks are achieved.
Patent Information
- Application Number
- CN202510171582.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art fails to adequately process the feature relationship between different modes in image fusion, resulting in insufficient cross-modal feature extraction, and ignores the complementary effects of scene information and semantic context, affecting the performance of downstream tasks.
A kind of image fusion method based on scene-semantic dissociation is proposed. Through multi-level encoding-fusion network and scene-semantic dissociation and fusion strategies, scenes and semantic related features are extracted and fused, and semantic features are injected into scene features to generate fused features. At the same time, a meta feature embedding module is introduced to associate the feature information of downstream tasks with semantic-related features, and optimize the extraction of semantic-related features.
Rich scene information and context information are retained in the fusion feature, and the performance of downstream tasks is optimized, especially in the semantic segmentation task, the result quality is significantly improved.
Smart Images

Figure CN120107079A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of image fusion, and in particular to a visible light and infrared image fusion method based on scene-semantic dissociation. Background Art
[0002] Visible and infrared image fusion combines information from both image modalities to obtain richer data, thereby enhancing the performance of various applications. Visible images provide rich texture details but are susceptible to lighting conditions, while infrared images reveal thermal information, are not affected by lighting, but lack texture details. The goal of visible and infrared image fusion is to integrate information from both modalities to produce a more information-rich fused image. Given that downstream vision applications benefit from the more accurate scene and object representations obtained from the fused images, many efforts have been made to combine image fusion with these applications, such as semantic segmentation, object detection, salient object detection, and remote sensing image description.
[0003] In recent years, the powerful information representation ability of deep learning has introduced many new methods for image fusion, such as image fusion based on autoencoder (AE), image fusion based on convolutional neural network (CNN), image fusion based on generative adversarial network (GAN), and image fusion based on diffusion model. Among them, the most representative method is the fusion method based on the autoencoder architecture. These methods usually use convolutional neural network (CNN) to extract features from images, then design various fusion strategies to integrate visible light and infrared features, and finally reconstruct the image based on the integrated features. Although better fusion results can be obtained based on this method, there is still a major problem: this method tends to improve the fusion quality of images only at the pixel level, without fully processing the feature relationship between different modalities, resulting in insufficient cross-modal feature extraction. Therefore, more and more work is devoted to incorporating feature dissociation technology into image fusion methods based on deep learning.
[0004] Image fusion methods based on feature dissociation aim to dissociate source features into modality-shared and modality-specific components and apply different fusion strategies to these features. In order to enable deep learning-based image fusion to use feature dissociation technology, DIDFuse first proposed to dissociate multimodal images into background and detail feature maps containing low-frequency and high-frequency information in the spatial domain, effectively improving the problem of information loss in cross-modal fusion. However, DIDFuse mainly distinguishes images based on the frequency of information during the dissociation process, without optimizing for specific image content, which may introduce unnecessary noise or cause detail loss during the fusion process.
[0005] To bridge this gap, CUFD dissociates image features into common and unique parts based on feature mapping, and adopts different fusion rules for each part before reconstruction. DRF dissociates images based on the source of information in multimodal images, and dissociates images into scene features and sensor features by defining attribute categories of unique information, effectively alleviating the problem of improper extraction of unique information. Subsequently, to solve the problems of small receptive field range and high-frequency information loss in the forward propagation process of feature extraction in the above methods, CDDFuse proposed a dual-branch Transformer-CNN feature extractor. Although these methods have achieved good performance in terms of visual effects and quantitative indicators, the above methods mainly focus on the dissociation and reconstruction of pixel features of the image itself. Therefore, an important challenge is how to combine downstream tasks with the feature dissociation process to preserve the semantics and scene details of the source modality image while being driven by downstream tasks. To address these challenges, our work explores a more reasonable image fusion paradigm based on feature dissociation, and is guided by its performance in downstream applications.
[0006] Despite significant progress, existing methods such as DeFusion, MetaFusion, and CDDFuse still face significant challenges. For example, pixel-level fusion strategies usually only focus on improving image quality without fully capturing cross-modal feature relationships, resulting in poor representation performance in downstream tasks such as semantic segmentation and object detection. In addition, these methods often ignore the complementary role of scene information and semantic context, which are critical for fidelity and optimization of specific tasks. To address these problems, the present invention introduces downstream application guidance into image fusion based on feature dissociation and proposes a new image fusion framework based on scene-semantic dissociation. Summary of the invention
[0007] The present invention proposes a visible light and infrared image fusion method based on scene-semantic dissociation. The infrared and visible light fused image enhanced by this method can retain rich scene information in the fused features, enhance the contextual information in the fused features, and optimize the performance of downstream tasks.
[0008] The present invention adopts the following technical solutions.
[0009] A visible light and infrared image fusion method based on scene-semantic dissociation is disclosed. The method adopts a multi-level encoding-fusion network and realizes scene-semantic dissociation and fusion strategy through a fusion module. Specifically, scene and semantic related features are extracted and fused respectively, and the fused semantics are injected into the scene features to enrich the context information in the fused features while maintaining the fidelity of the fused image. Meta-feature embedding is further introduced, and the encoding-fusion network is connected with the downstream application network during the training process. The fusion effect is optimized by enhancing the semantic extraction capability and serving the semantic segmentation task.
[0010] The method is used to enhance an infrared and visible light fused image obtained by fusing an infrared image with a visible light image, by retaining more scene information in the fused features and enhancing context information in the fused features to optimize the performance of downstream tasks, and includes the following steps:
[0011] Step S1, designing a multi-level encoding-fusion network to extract coarse-grained features of visible light images and infrared images;
[0012] Step S2, using a scene-semantic dissociation and fusion strategy to dissociate and fuse the coarse-grained features into scene-related features and semantic-related features, and inject the semantic-related features into the scene-related features to generate fused features;
[0013] Step S3, generating a fused image using the fused features through an image reconstruction decoder;
[0014] Step S4, designing a two-stage training strategy. During the training process, the feature information of the downstream task is associated with the semantically related features through a meta-feature embedding module to optimize the extraction of the semantically related features.
[0015] In step S1, the multi-level encoding-fusion network includes multiple levels of coarse-grained feature encoders; the encoder includes two types of feature extraction blocks, namely: a coarse extraction block RCEB based on Restormer and a coarse extraction block MCEB based on MSCA.
[0016] The Restormer-based coarse extraction block uses a self-attention mechanism in the channel dimension to extract global features from the image, and promotes the forward propagation of effective features through a gating mechanism, making it perform well in image restoration and super-resolution tasks;
[0017] The MSCA-based coarse extraction block uses a large kernel attention mechanism to extract local contextual information in the spatial and channel dimensions and capture features at different scales to achieve optimal performance and maintain low computational complexity when used as a backbone network building block in semantic segmentation tasks.
[0018] The process of performing coarse feature extraction in the coarse extraction block is expressed as follows:
[0019]
[0020] in and represents the rough visible and infrared characteristics of the i-th level, ↓ n It means to downsample the input by 1 / n times. Represents a connection along the channel dimension.
[0021] The multi-level encoding-fusion network in step S1 further includes a plurality of levels of scene-semantic dissociation and fusion modules, and the scene-semantic dissociation and fusion strategy in step S2 uses the scene-semantic dissociation and fusion module to dissociate the coarse-grained features into scene-related features and semantic-related features;
[0022] Among the scene-related features, the scene branch uses a Transformer-based encoder to extract global scene features. Among the semantic-related features, the semantic branch uses a convolutional neural network (CNN)-based encoder to extract local semantic features. The features of the scene branch and the semantic branch are fused through the fusion layer, and the spatial adaptive normalization (SPADE) module is used to inject the semantic features into the scene features to generate the fused features.
[0023] The scene-semantic dissociation and fusion module includes:
[0024] The scene branch uses a Transformer-based encoder to extract global scene features;
[0025] The semantic branch uses an encoder based on a convolutional neural network (CNN) to extract local semantic features;
[0026] A fusion layer, which is used to fuse the features of the scene branch and the semantic branch;
[0027] A spatial adaptive normalization module, which is used to inject the semantic features into the scene features;
[0028] The process of dissociating the coarse-grained features into scene-related features and semantic-related features can be expressed as:
[0029]
[0030] in, It is a pair of Transformer-based encoders that form a scene branch. It is a pair of CNN-based encoders that form a semantic branch, and CA* is a channel attention layer;
[0031] The decomposed scene and semantic features are fused through their respective fusion layers, which can be expressed as follows:
[0032]
[0033] Among them, F sce and F sem They are scene and semantic fusion layers respectively.
[0034] The image reconstruction decoder in step S3 uses the hole residual dense block DRDB and the Restormer block to reconstruct the image; for a given set of fusion features, the image reconstruction decoder applies these features to reconstruct the fused image, specifically:
[0035] First, each fused feature is enhanced by dense blocks to improve feature representation, the formula is:
[0036] Φ E,i =DRDB(Φ F,) , Formula 7;
[0037] where Φ F,i is a given set of fused features;
[0038] After upsampling, multi-level features are aggregated using point-wise convolution, expressed as:
[0039]
[0040] Among them↑ n (·) means upsampling the input by a factor of n.
[0041] The Restormer block decodes the aggregated features Φ A ; Finally, the channel is downsampled through the convolution layer, and then the sigmoid activation function is applied to generate the fused image I fu :
[0042] Φ D =RB(Φ A ), formula 8;
[0043] I fu =sigmoid(Conv(Φ D )) Formula 9;
[0044] Where RB stands for Restormer Block.
[0045] The meta-feature embedding module in step S4 includes:
[0046] A meta-feature generator, which is used to extract meta-features from the semantically related features and the features of the downstream task; a feature transformation module, which is used to convert the semantically related features into intermediate features and associate the intermediate features with the meta-features by optimizing the guided loss;
[0047] In step S4, the training phase also includes the following steps:
[0048] Step A1: In the image reconstruction stage and the fusion learning stage, the encoder is guided to learn effective feature representation by a loss function, and the quality of the fused image is optimized by a fusion loss function;
[0049] Step A2: In the meta-feature embedding stage, the fusion model is associated with the downstream task model through the meta-feature embedding module, and the contextual information in the fusion feature is increased through the guided loss.
[0050] In step A1, the loss function used by the method in the image reconstruction stage is:
[0051]
[0052] Among them I vi and I ir is the input visible light image and infrared image, and is the image reconstructed by the image reconstruction decoder, Φ sce and Φ sem These are the features extracted for the scene branch and the semantic branch respectively;
[0053] Furthermore, the light intensity loss in formula 11 is defined as follows:
[0054]
[0055] Where I is the source image, is the image reconstructed by the network. Represents the square of the L2 norm of the element;
[0056] The structural similarity loss in Formula 11 is defined as follows:
[0057]
[0058] Where SSIM(·,·) represents the calculation of the structural similarity coefficient between two images;
[0059] The gradient loss in Formula 11 is defined as follows:
[0060]
[0061] in represents the Sobel operation, ||·|| 1 Represents the L1 norm of the element;
[0062] The dissociation loss in Equation 11 is defined as follows:
[0063]
[0064] in Represents the calculation of the correlation coefficient between two elements, ε is a very small constant used to stabilize the value;
[0065] The loss function used by the method in the image fusion stage is:
[0066] Among them I fu is the fused image output by the network, is the mask-based intensity loss function; the gradient loss in formula 7 is defined as follows:
[0067]
[0068] Among them, max(·,·) represents taking the maximum value of each element. Definition as in formula 15;
[0069] In step A2, the training process of the meta-feature embedding stage includes two stages: an external update stage and an internal update stage. In the external update stage, the parameters of the entire fusion network are updated; in the internal update stage, the parameters of the meta-feature embedding module are optimized;
[0070] The loss function used in the external update phase of the meta-feature embedding phase is as shown in Formula 7, and the loss function used in the internal update phase is:
[0071]
[0072] where Φ m is the meta-feature from the meta-feature embedding module, Φ t is the intermediate feature, Same as formula 16.
[0073] Furthermore, the guided loss in Formula 18 is defined as follows
[0074]
[0075] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0076] (1) Through the scene semantic decomposition and fusion strategy, the coarse features are decomposed into scene-related components and semantic-related components, and the two components are effectively fused, thereby retaining rich scene information in the fused features and enhancing the contextual information in the fused features. In addition, the decomposition and fusion of the semantic branches can be further fine-tuned by downstream applications to optimize their performance;
[0077] (2) The multi-level encoder-fusion network can effectively utilize the feature information of the downstream application network and combines the scene-semantic feature decomposition and fusion strategy. This network enables the features from the downstream application to effectively guide the semantic component of the fusion strategy and bridges the gap between the downstream application features and the extracted semantic features through the meta-feature embedding module;
[0078] (3) Qualitative and quantitative experimental results on image fusion tasks demonstrate that our method achieves leading performance and effectively promotes the performance of downstream applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:
[0080] Attached Figure 1 It is a schematic diagram of the overall workflow of the present invention;
[0081] Attached Figure 2 It is a schematic diagram of the overall framework of the present invention;
[0082] Attached Figure 3 It is a schematic diagram of the scene-semantic decomposition and fusion module architecture of the present invention;
[0083] Attached Figure 4 It is a schematic diagram of the architecture of the meta-feature embedding module of the present invention;
[0084] Attached Figure 5 It is a schematic diagram of the qualitative results of the present invention on the fusion of three different scene images on the MSRS dataset;
[0085] Attached Figure 6 It is a schematic diagram of the qualitative results of the present invention on the fusion of three different scene images on the FMB dataset;
[0086] Attached Figure 7 It is a schematic diagram of the qualitative results of the fusion of three different scene images on the RoadScene dataset of the present invention;
[0087] Attached Figure 8 It is a schematic diagram of the quantitative results of image fusion on the MSRS test set of the present invention;
[0088] Attached Fig. 9 It is a schematic diagram of the quantitative results of image fusion on the FMB test set of the present invention;
[0089] Attached Fig.10 It is a schematic diagram of the quantitative results of image fusion on the RoadScene test set of the present invention. DETAILED DESCRIPTION
[0090] As shown in the figure, a visible light and infrared image fusion method based on scene-semantic dissociation is described. The method adopts a multi-level encoding-fusion network and implements scene-semantic dissociation and fusion strategy through a fusion module. Specifically, scene and semantic related features are extracted and fused respectively, and the fused semantics are injected into the scene features to enrich the contextual information in the fused features while maintaining the fidelity of the fused image. Meta-feature embedding is further introduced, and the encoding-fusion network is connected with the downstream application network during the training process. The fusion effect is optimized by enhancing the semantic extraction capability and serving the semantic segmentation task.
[0091] The method is used to enhance an infrared and visible light fused image obtained by fusing an infrared image with a visible light image, by retaining more scene information in the fused features and enhancing context information in the fused features to optimize the performance of downstream tasks, and includes the following steps:
[0092] Step S1, designing a multi-level encoding-fusion network to extract coarse-grained features of visible light images and infrared images;
[0093] Step S2, using a scene-semantic dissociation and fusion strategy to dissociate and fuse the coarse-grained features into scene-related features and semantic-related features, and inject the semantic-related features into the scene-related features to generate fused features;
[0094] Step S3, generating a fused image using the fused features through an image reconstruction decoder;
[0095] Step S4, designing a two-stage training strategy. During the training process, the feature information of the downstream task is associated with the semantically related features through a meta-feature embedding module to optimize the extraction of the semantically related features.
[0096] In step S1, the multi-level encoding-fusion network includes multiple levels of coarse-grained feature encoders; the encoder includes two types of feature extraction blocks, namely: a coarse extraction block RCEB based on Restormer and a coarse extraction block MCEB based on MSCA.
[0097] The Restormer-based coarse extraction block uses a self-attention mechanism in the channel dimension to extract global features from the image, and promotes the forward propagation of effective features through a gating mechanism, making it perform well in image restoration and super-resolution tasks;
[0098] The MSCA-based coarse extraction block uses a large kernel attention mechanism to extract local contextual information in the spatial and channel dimensions and capture features at different scales to achieve optimal performance and maintain low computational complexity when used as a backbone network building block in semantic segmentation tasks.
[0099] The process of performing coarse feature extraction in the coarse extraction block is expressed as follows:
[0100]
[0101] in and represents the rough visible and infrared characteristics of the i-th level, ↓ n It means to downsample the input by 1 / n times. Represents a connection along the channel dimension.
[0102] The multi-level encoding-fusion network in step S1 further includes a plurality of levels of scene-semantic dissociation and fusion modules, and the scene-semantic dissociation and fusion strategy in step S2 uses the scene-semantic dissociation and fusion module to dissociate the coarse-grained features into scene-related features and semantic-related features;
[0103] Among the scene-related features, the scene branch uses a Transformer-based encoder to extract global scene features. Among the semantic-related features, the semantic branch uses a convolutional neural network (CNN)-based encoder to extract local semantic features. The features of the scene branch and the semantic branch are fused through the fusion layer, and the spatial adaptive normalization (SPADE) module is used to inject the semantic features into the scene features to generate the fused features.
[0104] The scene-semantic dissociation and fusion module includes:
[0105] The scene branch uses a Transformer-based encoder to extract global scene features;
[0106] The semantic branch uses an encoder based on a convolutional neural network (CNN) to extract local semantic features;
[0107] A fusion layer, which is used to fuse the features of the scene branch and the semantic branch;
[0108] A spatial adaptive normalization module, which is used to inject the semantic features into the scene features;
[0109] The process of dissociating the coarse-grained features into scene-related features and semantic-related features can be expressed as:
[0110]
[0111] in, It is a pair of Transformer-based encoders that form a scene branch. It is a pair of CNN-based encoders that form a semantic branch, and CA* is a channel attention layer;
[0112] The decomposed scene and semantic features are fused through their respective fusion layers, which can be expressed as follows:
[0113]
[0114] Among them, F sce and F sem They are scene and semantic fusion layers respectively.
[0115] The image reconstruction decoder in step S3 uses the hole residual dense block DRDB and the Restormer block to reconstruct the image; for a given set of fusion features, the image reconstruction decoder applies these features to reconstruct the fused image, specifically:
[0116] First, each fused feature is enhanced by dense blocks to improve feature representation, the formula is:
[0117] Φ E,i =DRDB(Φ F,i ), Formula 7;
[0118] where Φ F,i is a given set of fused features;
[0119] After upsampling, multi-level features are aggregated using point-wise convolution, expressed as:
[0120]
[0121] Among them↑ n (·) means upsampling the input by a factor of n.
[0122] The Restormer block decodes the aggregated features Φ A ; Finally, the channel is downsampled through the convolution layer, and then the sigmoid activation function is applied to generate the fused image I fu :
[0123] Φ D =RB(Φ A ), formula 8;
[0124] I fu =sigmoid(Conv(Φ D )) Formula 9;
[0125] Where RB stands for Restormer Block.
[0126] The meta-feature embedding module in step S4 includes:
[0127] A meta-feature generator, which is used to extract meta-features from the semantically related features and the features of the downstream task; a feature transformation module, which is used to convert the semantically related features into intermediate features and associate the intermediate features with the meta-features by optimizing the guided loss;
[0128] In step S4, the training phase also includes the following steps:
[0129] Step A1: In the image reconstruction stage and the fusion learning stage, the encoder is guided to learn effective feature representation by a loss function, and the quality of the fused image is optimized by a fusion loss function;
[0130] Step A2: In the meta-feature embedding stage, the fusion model is associated with the downstream task model through the meta-feature embedding module, and the contextual information in the fusion feature is increased through the guided loss.
[0131] In step A1, the loss function used by the method in the image reconstruction stage is:
[0132]
[0133] Among them I vi and I ir is the input visible light image and infrared image, and is the image reconstructed by the image reconstruction decoder, Φ sce and Φ sem These are the features extracted for the scene branch and the semantic branch respectively;
[0134] Furthermore, the light intensity loss in formula 11 is defined as follows:
[0135]
[0136] Where I is the source image, is the image reconstructed by the network. Represents the square of the L2 norm of the element;
[0137] The structural similarity loss in Formula 11 is defined as follows:
[0138]
[0139] Where SSIM(·,·) represents the calculation of the structural similarity coefficient between two images;
[0140] The gradient loss in Formula 11 is defined as follows:
[0141]
[0142] in represents the Sobel operation, ||·|| 1 Represents the L1 norm of the element;
[0143] The dissociation loss in Equation 11 is defined as follows:
[0144]
[0145] in Represents the calculation of the correlation coefficient between two elements, ε is a very small constant used to stabilize the value;
[0146] The loss function used by the method in the image fusion stage is:
[0147]
[0148] Among them I fu is the fused image output by the network, is the light intensity loss function based on the mask;
[0149] The gradient loss in Formula 7 is defined as follows:
[0150]
[0151] Among them, max(·,·) represents the maximum value of each element. Definition as in formula 15;
[0152] In step A2, the training process of the meta-feature embedding stage includes two stages: an external update stage and an internal update stage. In the external update stage, the parameters of the entire fusion network are updated; in the internal update stage, the parameters of the meta-feature embedding module are optimized;
[0153] The loss function used in the external update phase of the meta-feature embedding phase is as shown in Formula 7, and the loss function used in the internal update phase is:
[0154]
[0155] where Φ m is the meta-feature from the meta-feature embedding module, Φ t is the intermediate feature, Same as formula 16.
[0156] Furthermore, the guided loss in Formula 18 is defined as follows
[0157]
[0158] Example
[0159] In this case, the fusion of visible light images and infrared images aims to generate a fused image with comprehensive scene understanding and detailed contextual information.
[0160] Considering that existing methods often have difficulty in fully handling the relationship between different modalities and optimizing for downstream applications. This example proposes a new method for visible and infrared image fusion based on scene semantic decomposition, called SSDFusion. The specific method is: a multi-level encoder-fusion network is adopted, and the fusion module implements the proposed scene semantic decomposition and fusion strategy, extracts and fuses scene-related and semantic-related components respectively, and injects the fused semantics into the scene features, enriching the contextual information in the fused features while maintaining the fidelity of the fused image. In addition, the meta-feature embedding is further connected with the downstream application network during the training process, which enhances the ability of our method to extract semantics, optimize fusion effects, and serve tasks such as semantic segmentation. Extensive experiments show that SSDFusion achieves the state-of-the-art level in image fusion performance, while enhancing the results on semantic segmentation tasks. The method of this example bridges the gap between image fusion based on feature decomposition and advanced vision applications, providing a more effective paradigm for multimodal image fusion.
[0161] The method described in this example uses the MSRS dataset, FMB dataset, and RoadScene dataset as data for comparative experiments and generalization experiments. In the two training stages, the image reconstruction and fusion learning stages involve cropping the source image to 128×128 pixels and then randomly rotating and flipping it as input.
[0162] Step 1: Design a new multi-level encoding-fusion network to extract coarse-grained features of visible light images and infrared images;
[0163] Step 2, using a scene-semantic dissociation and fusion strategy, dissociating and fusion the coarse-grained features into scene-related features and semantic-related features, and injecting the semantic-related features into the scene-related features to generate fused features;
[0164] Step 3, generating a fused image using the fused features through an image reconstruction decoder;
[0165] Step 4, designing a new two-stage training strategy. During the training process, the feature information of the downstream task is associated with the semantically related features through a meta-feature embedding module to optimize the extraction of the semantically related features.
[0166] Furthermore, the multi-level encoding-fusion network in step 1 comprises a plurality of levels of coarse-grained feature encoders, which include two types of feature extraction blocks: a Restormer-based coarse extraction block (RCEB) and an MSCA-based coarse extraction block (MCEB).
[0167] Furthermore, the Restormer-based coarse extraction block adopts a self-attention mechanism in the channel dimension to extract global features from the image, and promotes the forward propagation of effective features through a gating mechanism, achieving outstanding performance in image restoration and super-resolution tasks.
[0168] Furthermore, the MSCA-based coarse extraction block uses a large kernel attention mechanism to extract local contextual information in the spatial and channel dimensions and capture features at different scales, which has been shown to have superior performance as a backbone network building block in semantic segmentation tasks while maintaining low computational complexity.
[0169] When this example is implemented, the design of the visible light and infrared image fusion method based on scene-semantic dissociation can be realized by software. In order to objectively measure the fusion performance of the method proposed in the present invention, the performance of each method is evaluated from both qualitative and quantitative aspects. Qualitative evaluation is a subjective evaluation method that depends on human visual perception. A good fusion result should include both the significant contrast of the infrared image and the rich texture of the visible light image. Quantitative evaluation objectively evaluates the fusion performance through some statistical indicators. This paper selects information entropy (EN), standard deviation (SD), visual information fidelity (VIF), feature mutual information (FMI) and two objective image fusion performance indicators Q AB / F 、N AB / F .
[0170] Qualitative comparison: Figure 5 , 6 , 7 show the qualitative comparison of the fusion results on the MSRS test set, FMB test set and RoadScene test set respectively. Figure 5It can be seen that in daytime scenes, the results of DeFusion, DePF, CDDFuse, CrossFuse and MRFS all have obvious detail blurring, especially the texture of distant buildings, which is degraded due to light noise interference. In night scenes, DeFusion and MetaFusion lose obvious details in low-brightness areas, while DePF and CDDFuse retain edge information weakly, resulting in blurred edges of vehicles and pedestrians. In glare scenes, DePF, MetaFusion, and PSFusion have poor suppression effects on high-brightness glare, resulting in serious loss of details in vehicle and pedestrian areas, while MRFS exacerbates the blurring effect caused by excessive brightness in glare areas. Our method can dynamically adapt to changes in scene brightness, while enhancing infrared information while retaining visible light details, and by adjusting local contrast, successfully reduce the impact of glare, making the fusion result more natural and more in line with human visual perception.
[0171] from Figure 6 It can be seen that in normal scenes, the results of DeFusion, DePF, CDDFuse, CrossFuse, and MRFS all have the problem of insufficient contrast, and the details of distant vehicles and road signs are blurred or lost, and the key information in the scene cannot be clearly presented. In the smoke scene, DeFusion, MetaFusion, and CDDFuse cannot effectively remove the interference of smoke, resulting in blurred pedestrian outlines and obvious degradation of the details of distant buildings. At the same time, the fusion results of MRFS and CrossFuse are too bright, further amplifying the problem of detail loss. In the fog scene, DePF, CDDFuse, and PSFusion have weak fog processing capabilities, resulting in blurred edges of vehicles and road signs, and distant objects are blocked by fog and are not clear enough. In contrast, our method dynamically adjusts the brightness and contrast in all three scenes, effectively removing smoke and fog interference while clearly retaining key details.
[0172] from Figure 7 It can be seen that in the overexposure scenario, the results of DeFusion, DePF, CDDFuse, CrossFuse and MRFS failed to effectively suppress the light interference in the high-brightness area, resulting in serious loss of details in the target area (such as cyclists and vehicles), and the background information was also degraded due to overexposure. Our method significantly alleviates the overexposure problem by dynamically adjusting the brightness and contrast, which not only retains the clear edge details of the target area, but also restores the natural contrast of the background, achieving the best fusion effect. Overall, the proposed method is very suitable for all-weather image fusion tasks while maintaining the contrast of the fused image.
[0173] Quantitative comparison: Figure 8 ,9 As shown in Figure 10, our method achieves the best or second best results in almost all metrics. Specifically, the EN metric measures the amount of information in the fused image. Our method achieves the best results on the RoadScene dataset with a higher average brightness, and the second best performance on the MSRS and FMB datasets with a lower average brightness. This is because our method does not perform low-light enhancement on the fused result in low-light scenes, so even if it contains details from both modalities, the overall result is lower than methods with low-light enhancement, such as MRFS. The SD metric measures the contrast of the fused image. Our method achieves the best results on MSRS, but the second best results on FMB and RoadScene. This is because our method aims to generate images with natural contrast. Therefore, when the scene is bright, such as during the day, our method does not include too many infrared features, but tends to ensure the fidelity of the fused result. In contrast, PSFusion and CDDFuse achieve the best results in the SD metric because their preference for infrared features leads to large contrast variations in the scene, making them optimal. The VIF metric measures the amount of information shared between the fused image and the source image based on the human visual system. We achieve the best results on this metric on all datasets, which shows that our proposed method effectively and comprehensively fuses features from the source scenes through the scene semantic decomposition and fusion strategy. AB / F The amount of information transferred from the source image to the fused image is measured from the perspective of feature mutual information and edge information, respectively. Our method achieves the best or second best results on all datasets in terms of these two metrics, which indicates that our method can transfer sufficient information to the fused image. AB / F The metric analyzes the distortion of the fused image from the perspective of ghosting. Our proposed method also achieves the best results on all datasets, which shows that our proposed method can introduce the least artifacts into the fused image, and also confirms the excellent performance of our method in scene fidelity. Overall, our method has made great progress in scene fidelity and detail preservation, thanks to our proposed fusion strategy, which not only fully fuses the characteristics of the source scene, but also enhances the contextual information in the fusion result.
[0174] The above description is only a preferred embodiment of the present invention and does not limit the technical scope of the present invention. Therefore, any slight modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A visible light and infrared image fusion method based on scene-semantic dissociation, characterized by: The method adopts a multi-level encoding-fusion network and implements a scene-semantic dissociation and fusion strategy through a fusion module. Specifically, scene and semantic related features are extracted and fused respectively, and the fused semantics are injected into the scene features to enrich the contextual information in the fused features while maintaining the fidelity of the fused image. Meta-feature embedding is further introduced, and the encoding-fusion network is connected with the downstream application network during the training process. The fusion effect is optimized by enhancing the semantic extraction capability and serving the semantic segmentation task.
2. The visible light and infrared image fusion method based on scene-semantic dissociation according to claim 1, characterized in that: The method is used to enhance an infrared and visible light fused image obtained by fusing an infrared image with a visible light image, by retaining more scene information in the fused features and enhancing context information in the fused features to optimize the performance of downstream tasks, and includes the following steps: Step S1, designing a multi-level encoding-fusion network to extract coarse-grained features of visible light images and infrared images; Step S2, using a scene-semantic dissociation and fusion strategy to dissociate and fuse the coarse-grained features into scene-related features and semantic-related features, and inject the semantic-related features into the scene-related features to generate fused features; Step S3, generating a fused image using the fused features through an image reconstruction decoder; Step S4, designing a two-stage training strategy. During the training process, the feature information of the downstream task is associated with the semantically related features through a meta-feature embedding module to optimize the extraction of the semantically related features.
3. The visible light and infrared image fusion method based on scene-semantic dissociation according to claim 2, characterized in that: In step S1, the multi-level encoding-fusion network includes multiple levels of coarse-grained feature encoders; the encoder includes two types of feature extraction blocks, namely: a coarse extraction block RCEB based on Restormer and a coarse extraction block MCEB based on MSCA.
4. The visible light and infrared image fusion method based on scene-semantic dissociation according to claim 3 is characterized by: The Restormer-based coarse extraction block uses a self-attention mechanism in the channel dimension to extract global features from the image, and promotes the forward propagation of effective features through a gating mechanism, making it perform well in image restoration and super-resolution tasks; The MSCA-based coarse extraction block uses a large kernel attention mechanism to extract local contextual information in the spatial and channel dimensions and capture features at different scales to achieve optimal performance and maintain low computational complexity when used as a backbone network building block in semantic segmentation tasks.
5. The visible light and infrared image fusion method based on scene-semantic dissociation according to claim 4, characterized in that: The process of performing coarse feature extraction in the coarse extraction block is expressed as follows: in and represents the rough visible and infrared characteristics of the i-th level, ↓ n It means to downsample the input by 1 / n times. Represents a connection along the channel dimension.
6. The visible light and infrared image fusion method based on scene-semantic dissociation according to claim 2, characterized in that: The multi-level encoding-fusion network in step S1 further includes a plurality of levels of scene-semantic dissociation and fusion modules, and the scene-semantic dissociation and fusion strategy in step S2 uses the scene-semantic dissociation and fusion module to dissociate the coarse-grained features into scene-related features and semantic-related features; Among the scene-related features, the scene branch uses a Transformer-based encoder to extract global scene features. Among the semantic-related features, the semantic branch uses a convolutional neural network (CNN)-based encoder to extract local semantic features. The features of the scene branch and the semantic branch are fused through the fusion layer, and the spatial adaptive normalization (SPADE) module is used to inject the semantic features into the scene features to generate the fused features.
7. The visible light and infrared image fusion method based on scene-semantic dissociation according to claim 6, characterized in that: The scene-semantic dissociation and fusion module includes: The scene branch uses a Transformer-based encoder to extract global scene features; The semantic branch uses an encoder based on a convolutional neural network (CNN) to extract local semantic features; A fusion layer, which is used to fuse the features of the scene branch and the semantic branch; A spatial adaptive normalization module, which is used to inject the semantic features into the scene features; The process of dissociating the coarse-grained features into scene-related features and semantic-related features can be expressed as: in, It is a pair of Transformer-based encoders that form a scene branch. It is a pair of CNN-based encoders that form a semantic branch. * is a channel attention layer; The decomposed scene and semantic features are fused through their respective fusion layers, which can be expressed as follows: Among them, F sce and F sem They are scene and semantic fusion layers respectively.
8. The visible light and infrared image fusion method based on scene-semantic dissociation according to claim 2, characterized in that: The image reconstruction decoder in step S3 uses the hole residual dense block DRDB and the Restormer block to reconstruct the image; for a given set of fusion features, the image reconstruction decoder applies these features to reconstruct the fused image, specifically: First, each fused feature is enhanced by dense blocks to improve feature representation, the formula is: Φ E,i =DRDB(Φ F,i ), Formula 7; where Φ F,i is a given set of fused features; After upsampling, multi-level features are aggregated using point-wise convolution, expressed as: Among them↑ n (·) means upsampling the input by a factor of n. The Restormer block decodes the aggregated features Φ A ; Finally, the channel is downsampled through the convolution layer, and then the sigmoid activation function is applied to generate the fused image I fu : Φ D =RB(Φ A ), formula 8; I fu =sigmoid(Conv(Φ D )) Formula 9; Where RB stands for Restormer Block.
9. The visible light and infrared image fusion method based on scene-semantic dissociation according to claim 8, characterized in that: The meta-feature embedding module in step S4 includes: A meta-feature generator, which is used to extract meta-features from the semantically related features and the features of the downstream task; a feature transformation module, which is used to convert the semantically related features into intermediate features and associate the intermediate features with the meta-features by optimizing the guided loss; In step S4, the training phase also includes the following steps: Step A1: In the image reconstruction stage and the fusion learning stage, the encoder is guided to learn effective feature representation by a loss function, and the quality of the fused image is optimized by a fusion loss function; Step A2: In the meta-feature embedding stage, the fusion model is associated with the downstream task model through the meta-feature embedding module, and the contextual information in the fusion feature is increased through the guided loss.
10. The visible light and infrared image fusion method based on scene-semantic dissociation according to claim 9, characterized in that: In step A1, the loss function used by the method in the image reconstruction stage is: Among them I vi and I ir is the input visible light image and infrared image, and is the image reconstructed by the image reconstruction decoder, Φ sce and Φ sem These are the features extracted for the scene branch and the semantic branch respectively; Furthermore, the light intensity loss in formula 11 is defined as follows: Where I is the source image, is the image reconstructed by the network. Represents the square of the L2 norm of the element; The structural similarity loss in Formula 11 is defined as follows: Where SSIM(·,·) represents the calculation of the structural similarity coefficient between two images; The gradient loss in Formula 11 is defined as follows: in represents the Sobel operation, ||·||1 represents the L1 norm of the element; The dissociation loss in Equation 11 is defined as follows: in Represents the calculation of the correlation coefficient between two elements, ε is a very small constant used to stabilize the value; The loss function used by the method in the image fusion stage is: Among them I fu is the fused image output by the network, is the light intensity loss function based on the mask; The gradient loss in Formula 7 is defined as follows: Among them, max(·,·) represents the maximum value of each element. Definition as in formula 15; In step A2, the training process of the meta-feature embedding stage includes two stages: an external update stage and an internal update stage. In the external update stage, the parameters of the entire fusion network are updated; in the internal update stage, the parameters of the meta-feature embedding module are optimized; The loss function used in the external update phase of the meta-feature embedding phase is as shown in Formula 7, and the loss function used in the internal update phase is: where Φ m is the meta-feature from the meta-feature embedding module, Φ t is the intermediate feature, Same as formula 16. Furthermore, the guided loss in Formula 18 is defined as follows
Citation Information
Cited By
Multi-granularity semantic-driven infrared and visible light image fusion method based on multi-vision language large model
CN121169709A