Cross-modal image difference fusion method and system based on dual-branch feature decomposition

Through the cross-modal image differential fusion method of dual-branch feature decomposition, channel weights are dynamically allocated, which solves the problem of information loss in cross-modal image fusion, achieves efficient target recognition and detail retention under extreme conditions, and improves the quality of the fused image.

CN120510479BActive Publication Date: 2025-09-23SOUTHWEAT UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510990275.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-09-23
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

Existing cross-modal image fusion methods ignore the characteristics of images in different modalities, resulting in significant information loss in the fusion results, especially in extreme conditions where target recognition becomes more difficult.

Method used

A cross-modal image differential fusion method based on dual-branch feature decomposition is adopted. Channel weights are dynamically allocated through the differential compensation module. Combined with global average pooling and sigmoid function, cross-modal image feature fusion is achieved, retaining the specific features of different modalities.

Benefits of technology

It significantly improves the discrimination and detail fidelity of the fused image, enhances the saliency of infrared targets and the retention of visible light texture details, improves the clarity and information integrity of the fused image, and is suitable for resource-constrained scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510479B_ABST
    Figure CN120510479B_ABST
Patent Text Reader

Abstract

The present application discloses a cross-modal image differential fusion method and system based on dual-branch feature decomposition, which relates to the field of image processing technology. The method comprises extracting low-frequency basic features and high-frequency detail features of a first modality image and a second modality image to be fused respectively; inputting the low-frequency basic features and the high-frequency detail features into a feature fusion module respectively to obtain corresponding basic fusion features and detail fusion features; the feature fusion module comprises a differential compensation module, which calculates differential features according to input features, generates channel weights according to the differential features after global average pooling using a sigmoid function, calculates feature compensation values, adds the feature compensation values ​​to the corresponding input features as compensation features, and reconstructs the compensated features into fusion features by the feature fusion module; splices the basic fusion coding features and the detail fusion coding features; and obtains a target fusion image after decoding the splicing result, thereby realizing dynamic weight allocation for cross-modal image fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a cross-modal image differential fusion method and system based on dual-branch feature decomposition. Background Art

[0002] Cross-modal image fusion methods combine images from different modalities, integrating the characteristics of the same target detection area at different angles to offset the limited feature representation of single-modal images. This makes cross-modal image fusion valuable in a variety of fields. For example, in security monitoring and autonomous driving, the fusion of infrared and visible light images effectively reduces the difficulty of target recognition under extreme conditions, maintaining stable performance at night, in inclement weather, and even under extreme lighting conditions. In the medical field, the fusion of thermal imaging physiological information and visible light anatomical structures improves the accuracy of lesion localization.

[0003] Currently, cross-modal image fusion methods primarily use a simple linear addition approach during the image feature fusion stage. While this approach is simple to operate, it ignores the image characteristics of different modalities through a simple equally weighted addition operation. This can easily lead to feature conflicts between the modal images during image fusion, suppressing key information from the different modal images, and causing significant information loss in the fusion result. For example, when fusing infrared and visible light images, thermal targets against bright backgrounds may be completely obliterated, while details in highly textured areas may be oversmoothed, resulting in defects such as thermal target blurring, loss of texture detail, and edge artifacts in the fused image. Summary of the Invention

[0004] The purpose of the present invention is to solve the technical problem that the existing technology ignores the characteristics of different modal images when fusing image features, resulting in significant information loss in the fusion result. A cross-modal image differential fusion method and system based on dual-branch feature decomposition is provided, and a differential compensation module is introduced into the image fusion module. The differential features between different modal images are globally averaged and pooled to generate corresponding channel weights, thereby realizing dynamic weight allocation of cross-modal image feature fusion, so as to achieve the purpose of effectively retaining the specific features of different modal images and improving the discriminant line and detail fidelity of the fused image.

[0005] According to a first aspect of the present invention, the present invention claims protection for a cross-modal image difference fusion method based on dual-branch feature decomposition, comprising:

[0006] Acquiring a first modality image and a second modality image to be fused;

[0007] Encoding the first modality image and the second modality image to extract corresponding low-frequency basic features and high-frequency detail features respectively;

[0008] The low-frequency basic features of the first modality image and the low-frequency basic features of the second modality image are input into the feature fusion module to obtain basic fusion features; the high-frequency detail features of the first modality image and the high-frequency detail features of the second modality image are input into the feature fusion module to obtain detail fusion features;

[0009] Among them, the feature fusion module includes a differential compensation module, which calculates differential features based on input features, uses a sigmoid function to generate channel weights based on the differential features after global average pooling, calculates feature compensation values ​​based on the differential features and corresponding channel weights, and adds the feature compensation values ​​to the corresponding input features as compensation features. The feature fusion module reconstructs the compensation features of all input features to obtain fused features;

[0010] Concatenate the basic fusion coding features and the detail fusion coding features;

[0011] After decoding the splicing result, the target fused image is obtained.

[0012] In one embodiment of the present application, encoding the first modality image and the second modality image further includes: extracting shallow features of the input image by a shared feature extractor, and extracting corresponding low-frequency basic features from the shallow features by a basic encoder, wherein the basic encoder includes a first extraction module, a global modeling module, a second extraction module, and a feedforward neural network;

[0013] Wherein, the first extraction module and the second extraction module are used to extract local structural information of the input features;

[0014] The global modeling module includes a flattening layer, a state modeling layer, and a residual fusion layer. The flattening layer flattens the two-dimensional feature map into a sequence vector. The state modeling layer adds the state vector to the dynamic offset and normalizes it into an attention weight. The attention weight and state feature are multiplied with the input feature of the state modeling layer. The transformation strength is dynamically adjusted by the scaling factor, and then the result is element-by-element multiplication with the projection matrix.

[0015] The feedforward neural network is used to perform nonlinear transformation on input features.

[0016] In one embodiment of the present application, the global modeling module further includes a normalization layer, which is used to normalize the output of the flattening layer using the LayerNorm method.

[0017] In one embodiment of the present application, the first extraction module, the global modeling module, the second extraction module and the feedforward neural network of the basic encoder further include a fusion layer, which performs a residual operation on the output and input of the corresponding modules.

[0018] In one embodiment of the present application, encoding the first modality image and the second modality image further includes extracting corresponding high-frequency detail features from shallow features using a detail encoder: the detail encoder divides the shallow features into three input sub-features along the channel dimension, extracts detail features from each input sub-feature, and then performs a splicing operation on all extracted results to obtain corresponding high-frequency detail features;

[0019] The step of extracting detail features from the first input sub-feature includes:

[0020] Extracting global information of the first input sub-feature through a bidirectional scanning Mamba module to obtain a first feature map;

[0021] Performing Haar wavelet transform on the first input sub-feature according to a preset low-pass filter and a high-pass filter to obtain a frequency domain feature;

[0022] Convolution is used to extract frequency domain features to extract local information;

[0023] Perform inverse wavelet transform on local information to obtain the transformed result;

[0024] The first feature map and the second feature map are summed to obtain the feature extraction result.

[0025] In one embodiment of the present application, extracting detail features from the second input sub-feature includes:

[0026] Divide the second input sub-feature into several second input grandchild features along the channel dimension;

[0027] Each second input sub-feature is subjected to a local convolution operation with different kernel sizes to obtain a feature extraction sub-result;

[0028] All feature extraction sub-results are spliced ​​according to channels to obtain the feature extraction result.

[0029] In one embodiment of the present application, an identity mapping method is used to extract detail features of the third input sub-feature.

[0030] In one embodiment of the present application, the images to be fused are an infrared image and a visible light image.

[0031] According to a second aspect of the present invention, the present invention claims protection for a cross-modal image difference fusion system based on dual-branch feature decomposition, comprising:

[0032] an acquisition unit, which acquires the first modality image and the second modality image to be fused;

[0033] an encoding unit, encoding the first modality image and the second modality image to extract corresponding low-frequency basic features and high-frequency detail features respectively;

[0034] A fusion unit, wherein the low-frequency basic features of the first modality image and the low-frequency basic features of the second modality image are input into a feature fusion module to obtain a basic fusion feature; the high-frequency detail features of the first modality image and the high-frequency detail features of the second modality image are input into the feature fusion module to obtain a detail fusion feature; and the basic fusion coding feature and the detail fusion coding feature are spliced ​​together;

[0035] Among them, the feature fusion module includes a differential compensation module, which calculates differential features based on input features, uses a sigmoid function to generate channel weights based on the differential features after global average pooling, calculates feature compensation values ​​based on the differential features and corresponding channel weights, and adds the feature compensation values ​​to the corresponding input features as compensation features. The feature fusion module reconstructs the compensation features of all input features to obtain fused features;

[0036] The decoding unit decodes the splicing result to obtain the target fused image.

[0037] In one embodiment of the present application, the encoding unit further includes a shared feature extractor and a basic encoder: the shared feature extractor extracts shallow features of the input image, and the basic encoder extracts corresponding low-frequency basic features from the shallow features: the basic encoder includes a first extraction module, a global modeling module, a second extraction module and a feedforward neural network;

[0038] Wherein, the first extraction module and the second extraction module are used to extract local structural information of the input features;

[0039] The global modeling module includes a flattening layer, a state modeling layer, and a residual fusion layer. The flattening layer flattens the two-dimensional feature map into a sequence vector. The state modeling layer adds the state vector to the dynamic offset and normalizes it into an attention weight. The attention weight and state feature are multiplied with the input feature of the state modeling layer. The transformation strength is dynamically adjusted by the scaling factor, and then the result is element-by-element multiplication with the projection matrix.

[0040] The feedforward neural network is used to perform nonlinear transformation on input features.

[0041] In one embodiment of the present application, the global modeling module further includes a normalization layer, which is used to normalize the output of the flattening layer using the LayerNorm method.

[0042] In one embodiment of the present application, the first extraction module, the global modeling module, the second extraction module and the feedforward neural network of the basic encoder further include a fusion layer, which performs a residual operation on the output and input of the corresponding modules.

[0043] In one embodiment of the present application, the encoding unit further includes a detail encoder, which extracts corresponding high-frequency detail features from the shallow features: the detail encoder divides the shallow features into three input sub-features along the channel dimension, extracts detail features from each input sub-feature, and then performs a splicing operation on all extracted results to obtain corresponding high-frequency detail features;

[0044] The step of extracting detail features from the first input sub-feature includes:

[0045] Extracting global information of the first input sub-feature through a bidirectional scanning Mamba module to obtain a first feature map;

[0046] Performing Haar wavelet transform on the first input sub-feature according to a preset low-pass filter and a high-pass filter to obtain a frequency domain feature;

[0047] Convolution is used to extract frequency domain features to extract local information;

[0048] Perform inverse wavelet transform on local information to obtain the transformed result;

[0049] The first feature map and the second feature map are summed to obtain the feature extraction result.

[0050] In one embodiment of the present application, extracting detail features from the second input sub-feature includes:

[0051] Divide the second input sub-feature into several second input grandchild features along the channel dimension;

[0052] Each second input sub-feature is subjected to a local convolution operation with different kernel sizes to obtain a feature extraction sub-result;

[0053] All feature extraction sub-results are spliced ​​according to channels to obtain the feature extraction result.

[0054] In one embodiment of the present application, an identity mapping method is used to extract detail features of the third input sub-feature.

[0055] In one embodiment of the present application, the images to be fused are an infrared image and a visible light image.

[0056] According to the third aspect of the present invention, the present invention seeks protection for a cross-modal image differential fusion device based on dual-branch feature decomposition, comprising a processor and a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the steps in the method described in the first aspect above are executed.

[0057] This application has the following beneficial effects:

[0058] 1. A differential compensation module is introduced into the image fusion module, utilizing a dynamic adaptive fusion mechanism to achieve cross-modal image feature fusion. The core of this module is to use global average pooling to globally calculate the response of each channel in the spatial dimension and extract global semantic information at the channel level. This approach not only reduces local noise interference but also effectively reduces the computational cost of summarizing the overall importance of the feature map, thereby achieving stable modeling of regions with modal differences. A sigmoid function is also used to compress the global statistical information obtained by global average pooling to the range [0, 1]. A gating mechanism is then used to weight the backbone features. The sigmoid function is not only simple and efficient, but also independently adjusts the response of each channel, reducing inter-channel competition. This makes significant differences more prominent while suppressing redundant background, maximizing the preservation of key structures and textures, and significantly improving the clarity, contrast, and information integrity of the fused image. For example, the contribution of infrared features is enhanced in salient infrared target regions, while detail preservation is strengthened in texture-rich visible regions, thereby achieving a natural integration of complementary information. This differential compensation mechanism is more robust, lightweight, and easy to train.

[0059] 2. The basic encoder introduces a groundbreaking global modeling module that integrates latent state mixing and state-space duality mechanisms. While retaining the ability to capture global features, this module significantly reduces computational complexity to a linear level, or O(N), significantly reducing computing resource consumption and improving training and inference speed. It also maintains or even enhances the ability to model long-range spatial dependencies, enhancing background consistency and subject integrity, and effectively improving the structural coherence and overall visual quality of the fused image. This module is particularly computationally inefficient when processing high-resolution images, making this method suitable for resource-constrained applications.

[0060] 3. Adding a normalization layer based on LayerNorm to the global modeling module can make the feature distribution more stable, provide a unified scale output for subsequent modeling, and ensure modeling stability.

[0061] 4. A residual fusion operation is introduced after each module in the basic encoder. Its purpose is to fully integrate global context information and local structural features, effectively improving the richness and accuracy of feature expression.

[0062] 5. In the feature extraction process of the first input sub-feature, a long-range wavelet transform enhanced modeling mechanism is introduced, combined with the state-space modeling method, to effectively enhance the responsiveness of high-frequency edge features while capturing global dependency information.

[0063] 6. During the feature extraction process of the second input sub-feature, multi-core depthwise separable convolution operations are used to extract local information with variable effective receptive fields, improve the model's ability to perceive spatial details under different receptive fields, and realize the joint expression of multi-scale structural information.

[0064] 7. During the feature extraction process of the third input sub-feature, the identity mapping method is used for direct transmission to avoid redundant calculations, reduce repeated feature expressions, effectively control the computational complexity, and improve the overall reasoning efficiency.

[0065] 8. When the three-branch structure works together, it takes into account both global structural modeling and local detail preservation, ensuring good computational efficiency while ensuring performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0067] Figure 1 Schematic diagram of the process of a cross-modal image difference fusion method based on dual-branch feature decomposition according to an embodiment of the present application;

[0068] Figure 2 This is a schematic diagram of the structure of the decoder and encoder involved in the embodiments of the present application;

[0069] Figure 3 This is a schematic diagram of the structure of the feature fusion module involved in the embodiment of the present application;

[0070] Figure 4 This is a schematic diagram of the structure of the basic encoder involved in the embodiment of the present application;

[0071] Figure 5 This is a schematic diagram of the structure of the state modeling layer involved in the embodiments of the present application;

[0072] Figure 6 A schematic diagram of the structure of a detailed encoder involved in an embodiment of the present application;

[0073] Figure 7 This is a schematic diagram of a fused image of a daytime environment involved in an embodiment of the present application;

[0074] Figure 8 This is a schematic diagram of a fused image of a night environment involved in an embodiment of the present application;

[0075] Figure 9 Schematic diagram of the structure of a cross-modal image difference fusion system based on dual-branch feature decomposition involved in an embodiment of the present application;

[0076] Figure 10 This is a schematic diagram of the structure of an electronic device involved in an embodiment of the present application. DETAILED DESCRIPTION

[0077] The present invention provides a cross-modal image difference fusion method and system based on dual-branch feature decomposition. To make the above-mentioned objects, features, and advantages of the present application more clearly understood, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. The components of the embodiments of the present application generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings with reference to the terms "one embodiment," "some embodiments," "implementation," "embodiment," "illustrative embodiment," "example," "specific example," or "some examples" is not intended to limit the scope of the claimed application, but merely indicates that the specific features, structures, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of the present application.

[0078] It should be noted that similar reference numerals and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures. At the same time, in the description of this application, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.

[0079] According to the first aspect of the present invention, the present invention claims a cross-modal image difference fusion method based on dual-branch feature decomposition, referring to the attached Figure 1 As shown, including:

[0080] S1: Acquire a first modality image and a second modality image to be fused.

[0081] It should be noted that the first modal image and the second modal image are used to represent two different image data sources with different imaging principles or information dimensions, such as infrared images, visible light images, radar images, microwave images, thermal imaging images, near-infrared images, multispectral images, hyperspectral images and depth images, etc.

[0082] In this embodiment, an infrared image is used as the first modality image and a visible light image is used as the second modality image for detailed description. The first modality image is denoted as I and the second modality image is denoted as V.

[0083] S2: Input the first modality image and the second modality image into the encoder respectively to extract the corresponding low-frequency basic features and high-frequency detail features respectively, and record the low-frequency basic features of the first modality image as , and its high-frequency detail features are recorded as , and the low-frequency basic features of the second modality image are recorded as , and its high-frequency detail features are recorded as .

[0084] It should be noted that the encoder is used to extract image features of the first modality image and the second modality image, and can be constructed based on a Transformer, a convolutional neural network model, etc.

[0085] In this embodiment, refer to the attached Figure 2 As shown, the encoder includes a shared feature extractor and a feature encoder, but it is not limited thereto. The shared feature extractor is constructed based on Restormer Block to extract shallow features of the first modality image and the second modality image respectively, which are recorded as and The feature encoder includes a basic feature encoder and a detail feature encoder. The basic feature encoder is constructed according to the Transformer model to extract the corresponding low-frequency basic features from the shallow features of the first modality image and the second modality image, that is, and The detail feature encoder is constructed according to a convolutional neural network model to extract the corresponding high-frequency detail features from the shallow features of the first modality image and the second modality image, that is, and .

[0086] S3: low-frequency basic features of the first modality image and low-frequency basic features of the second modality image Input feature fusion module to obtain basic fusion features, denoted as . The high frequency detail features of the first modality image and high-frequency detail features of the second modality image Input feature fusion module to obtain detail fusion features, denoted as .

[0087] In this embodiment, the feature fusion module is used to fuse the complementary and common information of the first modality image and the second modality image. The complementary and common information of the first modality image and the second modality image can be expressed as:

[0088] ;

[0089] ;

[0090] in, For common information, and is complementary information.

[0091] In this embodiment, the feature fusion module adopts a progressive fusion architecture based on a convolutional neural network framework, first extracting and fusing features progressively and then reconstructing features based on the fused features. The feature fusion module includes a differential compensation module, a splicing module, and a reconstruction module.

[0092] It should be noted that, refer to the attached Figure 3 As shown, the differential compensation module is used to dynamically weight the complementary features of the first modality image and the second modality image. The differential compensation module includes a differential layer, a global average pooling layer, and a sigmoid function normalization layer. Taking the feature fusion of the low-frequency basic features of the first modality image and the low-frequency basic features of the second modality image as an example, the input of the differential compensation module is the low-frequency basic features. and , the output of the differential compensation module is and :

[0093] ;

[0094] ;

[0095] Among them, GAP represents global average pooling, which is used for global average pooling to compress complementary features into a vector, and δ represents the sigmoid function, which normalizes the vector to [0, 1] to generate channel weights.

[0096] In this embodiment, the output of the feature fusion module It can be expressed as:

[0097] ;

[0098] Among them, R represents the reconstruction module, which is composed of 3*3 convolution operations; concat represents the splicing module, which is used for tensor splicing.

[0099] It should be noted that, similar to the feature fusion of the above low-frequency basic features, the detail fusion feature obtained after the feature fusion operation of the high-frequency detail features can be expressed as:

[0100] ;

[0101] in, Represents the high-frequency detail features of the first modality image High-frequency detail features of the second modality image The complementary characteristics in the differential compensation module are dynamically weighted results.

[0102] S4: Input the basic fusion feature into the basic feature encoder to obtain the basic fusion encoding feature, which is recorded as , the detail fusion feature is input into the detail feature encoder to obtain the detail fusion coding feature, which is recorded as .

[0103] It should be noted that the basic feature encoder and the detail feature encoder can be constructed based on Transformer, or based on a convolutional neural network model, or can adopt a structure similar to the basic encoder and the detail encoder, or other feasible structures. This embodiment does not further limit the specific structure of the basic feature encoder and the detail feature encoder.

[0104] S5: After performing a splicing operation on the basic fusion coding features and the detail fusion coding features, the splicing operation is input into a feature decoder to obtain a target fusion image.

[0105] It should be noted that the feature decoder is used to reconstruct the image fusion result of the first modality image and the second modality image by concatenating the tensor of the basic fusion coding feature and the detail fusion coding feature. The feature decoder is constructed based on a two-layer Restormer Block.

[0106] In a feasible embodiment, refer to the attached Figure 4 As shown, the feature encoder includes a basic encoder. The basic encoder is used to encode the shallow features output by the shared feature extractor to obtain the corresponding low-frequency basic features, that is, the shallow features of the first modality image. Input the basic encoder to obtain the corresponding low-frequency basic features , the shallow features of the second modality image The basic encoder is input to obtain the corresponding low-frequency basic features.

[0107] In this embodiment, the basic encoder is sequentially composed of a first extraction module, a global modeling module, a second extraction module and a feedforward neural network.

[0108] In this embodiment, both the first extraction module and the second extraction module are constructed based on deep convolution and are used to extract local structural information of the input features of the module. The global modeling module is used to perform global modeling of the input and extract long-range dependencies. The feedforward neural network introduces nonlinear transformations and high-order semantics to produce feature maps with stronger expressive and perceptual capabilities.

[0109] It should be noted that the first extraction module, the global modeling module, the second extraction module and the feedforward neural network of the basic encoder also include a fusion layer, which is used to perform residual operations on the output and input of the corresponding modules to fully integrate global context information and local structural features.

[0110] It should be noted that the shallow features of infrared images Extract low-frequency basic features from The input of the basic encoder is the shallow features of the infrared image , shallow features of infrared images First, after the first extraction module undergoes a depthwise convolution operation, a residual fusion operation is performed on the fusion layer, and the result of the residual fusion operation is recorded as , the calculation method can be expressed as:

[0111] ;

[0112] Here, α0 represents the gating coefficient of the fusion layer of the first extraction module, and DWConv1 represents the depth convolution operation of the first extraction module.

[0113] It should be noted that the global modeling module operates on the residual results The modeling results are recorded as , the calculation method can be expressed as:

[0114] ;

[0115] Wherein, SSD represents the global modeling module.

[0116] It should be noted that the modeling results of the global modeling module are input into the feedforward neural network after the second extraction module extracts the local structure information, and the low-frequency basic features of the infrared image are output. , the processing method can be expressed as:

[0117] ;

[0118] Among them, α2 and α3 represent the fusion ratio of the fusion layer of the feedforward neural network, DWConv2 represents the depth convolution operation of the second extraction module, and FFN represents a feedforward neural network.

[0119] In this embodiment, the global modeling module includes a flattening layer, a normalization layer, a linear transformation layer, a state modeling layer, and a fusion layer. The flattening layer is used to flatten the two-dimensional feature map into a sequence vector. The normalization layer uses LayerNorm to normalize the flattened result. The linear transformation layer is used to linearly reconstruct the normalized one-dimensional time series features. The specific method can be expressed as follows:

[0120] ;

[0121] Among them, Lin represents the linear transformation layer, xin represents the one-dimensional time series feature input to the linear transformation layer, that is, the output of the normalization layer, represents the transposed result of the linear transformation weight, and b represents the offset of the linear transformation.

[0122] It should be noted that the purpose of adding a normalization layer based on LayerNorm in the global modeling module is to make the feature distribution more stable, provide a unified scale output for the subsequent modeling, and ensure modeling stability.

[0123] In this embodiment, refer to the attached Figure 5 As shown, the state modeling layer adds the state vector and the dynamic offset and normalizes it into attention weights, multiplies it with the state features, and uses the multiplication result of the multiplication result with the sequence features after linear transformation as the latent variable, thereby realizing the extraction of long-distance dependencies and enhancing the global modeling capability. In addition, the state modeling layer also combines the latent variable with the scaling factor to dynamically adjust the transformation intensity, and performs element-by-element multiplication operation with the projection matrix to complete the context restoration process from state to space. The specific method of the state modeling layer can be expressed as:

[0124] ;

[0125] ;

[0126] ;

[0127] in, Represents the sequence features after linear transformation; h represents the latent variable; Y represents the output result of the state modeling layer; LayerNorm represents the normalization layer; Flattern represents the flattening layer; softmax represents the normalization operation using the softmax function; dt represents the dynamic offset; A represents the state vector; B represents the state attention weight; C represents the projection matrix, which is used to map the latent variable h to the output. The convolution block can be selected as the projection matrix to achieve the purpose of maintaining attention to the noteworthy parts of the original input value and reducing the degree of attention to the unnoticeable parts; D represents the scaling factor.

[0128] In this embodiment, the fusion layer is used to perform weighted fusion on the output of the state modeling layer and the output of the first extraction module to fully integrate global context information and local structural features, effectively improving the richness and accuracy of feature expression. The processing method of the residual fusion layer can be expressed as:

[0129] ;

[0130] Wherein, α1 represents the gating parameter of the fusion layer of the global modeling module.

[0131] It should be noted that the gate parameters α0-α3, state vector A, dynamic offset dt, state feature B, feature C, and scaling factor D of the fusion layer are all parameters obtained through training.

[0132] In a feasible embodiment, refer to the attached Figure 6 As shown, the feature encoder also includes a detail encoder. The detail encoder is used to encode the shallow features output by the shared feature extractor to obtain corresponding high-frequency detail features, that is, the shallow features of the first modality image. Input the detail encoder to obtain the corresponding high-frequency detail features , the shallow features of the second modality image The detail encoder is input to obtain corresponding high-frequency detail features.

[0133] In this embodiment, the detail encoder divides the input feature into three input sub-features along the channel dimension, and performs differential processing on each input sub-feature to achieve efficient extraction and multi-scale expression of detail features. The three input sub-features are respectively recorded as the first input sub-feature, the second input sub-feature, and the third input sub-feature. represents the first input sub-feature, represents the second input sub-feature, represents the third input sub-feature.

[0134] It should be noted that the value range of the global channel ratio ξ of the first input sub-feature is [0, 1], that is, ,in Represents a shape A tensor whose elements belong to the real number domain, h represents the feature height, w represents the feature width, c represents the number of channels, and R represents the real number domain. The value of the global channel ratio μ of the second input sub-feature is not greater than 1-ξ, that is, The third input sub-feature is expressed as .

[0135] In this embodiment, processing the first input sub-feature includes: performing bidirectional scanning on the first input sub-feature by a Mamba module. The processing of the Mamba module can be expressed as follows:

[0136] ;

[0137] ;

[0138] ;

[0139] Among them, SSM represents the state space model (Selective State Space Model, referred to as SSM) in the Mamba module, σ represents the nonlinear activation function, and the ReLU function can be selected. Conv represents the convolution operation, and Linear represents the linear conversion operation. represents the direct product operation, Indicates the features after modeling by the bidirectional scanning Mamba module. Represents the modulation vector, used for weight adjustment, Represents the first feature map.

[0140] In this embodiment, the first input sub-feature Haar wavelet transform is also performed through low-pass filter and high-pass filter to obtain frequency domain features , frequency domain characteristics Then the frequency domain features are extracted by local convolution information and then inverse wavelet transform (denoted as IWT) is performed to obtain the second feature map ,Right now The specific process can be expressed as:

[0141] ;

[0142] ;

[0143] Where WT represents Haar wavelet transform; IWT represents inverse wavelet transform; Represents frequency domain features; It is a low-pass filter; , , It is a set of high-pass filters.

[0144] In this embodiment, the feature map and feature maps The sum is calculated to obtain the output corresponding to the first input sub-feature.

[0145] It should be noted that the processing of the first input sub-feature is to further improve the ability to extract fine-grained information such as high-frequency edge details on the basis of global modeling, thereby enhancing the expression effect of local structure and texture features in the image.

[0146] In this embodiment, the second input sub-feature is divided into a plurality of second input grandchild features along the channel dimension, and the jth second input grandchild feature is expressed as Each of the second input features is subjected to local convolution operations with different kernel sizes, and the outputs of each convolution operation are channel-joined using convolution operations to obtain the fused features. . It can be specifically expressed as:

[0147] ;

[0148] ;

[0149] It should be noted that the second input sub-feature realizes feature interaction of multiple receptive fields by extracting local information with a variable effective receptive field.

[0150] In this embodiment, the third input sub-feature is subjected to identity mapping, and the mapping result is recorded as .

[0151] In this embodiment, the output of the detail encoding module It can be expressed as:

[0152] ;

[0153] Among them, concat represents the concatenation operation.

[0154] It should be noted that the purpose of the third input sub-feature is to reduce feature redundancy in high-dimensional space, so as to minimize unnecessary calculations and improve operation efficiency.

[0155] It should be noted that the values ​​of the above-mentioned global channel ratios ξ and μ are obtained by pre-setting. The larger the value of the global channel ratio ξ, the better the global modeling and wavelet high-frequency enhancement effects. The larger the value of the global channel ratio μ, the better the multi-scale convolution extraction of local receptive field features. When the values ​​of the global channel ratios ξ and μ are smaller, the training speed can be accelerated. In this embodiment, the value of the global channel ratio ξ is 0.7, and the value of the global channel ratio μ is 0.2.

[0156] In this embodiment, in order to verify the effectiveness of the cross-modal image fusion method proposed in the present invention, systematic training and testing were carried out on the public benchmark dataset MSRS, which is widely used in the field of infrared and visible light image fusion. A total of 1,083 pairs of data were included in the training dataset and 361 pairs were included in the test dataset.

[0157] In this implementation, seven existing image fusion methods, DIDFuse, U2Fusion, SDNet, RFNet, TarDAL, DeFusion, and ReCoNet, are selected for comparison. Figure 7 The image fusion results of each image fusion method in daytime environment are shown. Figure 8 The image fusion results of each image fusion method in a nighttime environment are presented. The figure clearly demonstrates the superior performance of the proposed image fusion method in integrating the dual-modal information of infrared and visible light images. It successfully highlights the thermal radiation signatures of key targets (such as pedestrians and vehicles) in infrared images. This is particularly true in dark areas with poor lighting conditions or complex backgrounds, effectively highlighting the target objects, significantly improving foreground-background distinction and target detectability. Furthermore, the method fully preserves the rich texture details and background structure information found in visible light images. Even in areas difficult to discern due to insufficient illumination (such as building outlines, vegetation textures, and road signs), the fused image exhibits sharp edges and rich contour details, significantly enhancing the understanding of the scene's geometric layout and the visual naturalness of the image. This enhanced saliency of thermal targets complements the precise preservation of background texture details, resulting in a more comprehensive and visually enhanced fusion result.

[0158] In this implementation, eight metrics are used as quantitative evaluation criteria for fused images: entropy (EN), standard deviation (SD), spatial frequency (SF), mutual information (MI), correlation sum of differences (SCD), visual information fidelity (VIF), QAB / F, and structural similarity index (SSIM). A higher evaluation result indicates a better fused image.

[0159] It should be noted that entropy (EN) is an information-theoretic metric that measures the information richness of the fused image. A higher entropy value indicates more information in the fused image, reflecting better fusion results; a lower entropy value indicates poorer fusion results. Standard deviation (SD) measures the grayscale and contrast distribution in the fused image. A larger standard deviation indicates better visual quality in the fused image, while a lower standard deviation indicates poorer visual quality. Spatial frequency (SF) comprehensively measures the level of local grayscale variation or detail richness in the fused image across both row and column directions, reflecting the image's spatial clarity and texture information. A higher spatial frequency indicates that the fused image contains more high-frequency details such as edges and textures. Mutual information (MI) measures the amount of information retained by the fused image F from the source images (infrared image IR and visible light image VIS). A higher mutual information indicates that the fused image more effectively integrates key information from the different source images. Sum of Difference Correlation (SCD) evaluates fusion performance by calculating the correlation between the difference maps of the fused image and the source images. A higher Sum of Difference Correlation indicates that the fused image better preserves the structural features of the source images. Visual information fidelity (VIF) simulates the characteristics of the human visual system (HVS) and evaluates the fidelity of the fused image in visual perception information relative to the source image. The higher the visual information fidelity, the closer the amount of visual information contained in the fused image is to the source image, and the better the perceptual quality. QAB / F is a fusion quality indicator based on local structural similarity. The closer the QAB / F value is to 1, the higher the fusion quality. The structural similarity index measure (SSIM) evaluates the similarity between the fused image and the ideal reference image (usually one of the source images is selected as the reference) in terms of brightness (luminance), contrast (contrast), and structure (structure). The higher the structural similarity index measure (maximum 1), the closer the fused image is to the reference image in terms of structural information. The experimental results are shown in Table 1:

[0160] Table 1 Comparative experimental results

[0161]

[0162] As shown in Table 1, the proposed image fusion method achieves leading or highly competitive performance across most key metrics: High EN values ​​confirm a significant increase in the information content of the fused image; outstanding SD and SF values ​​reflect its excellent overall contrast and rich spatial details (edges and textures); excellent MI values ​​demonstrate its ability to effectively preserve key information from the source image; leading SCD values ​​indicate a high correlation with the source image in terms of structural information; good VIF and QAB / F values ​​confirm the high quality of the fusion effect from the perspective of human visual perception models and local structural similarity; and a high SSIM value further consolidates its advantage in structural information fidelity. Comprehensive qualitative and quantitative evaluations show that this method not only clearly presents thermal radiation targets and preserves rich texture details in subjective visual perception, but also demonstrates excellent overall performance, strong robustness, and wide applicability under objective metrics. It can effectively cope with diverse lighting conditions and complex scenes containing various targets, providing a high-quality fused image foundation for subsequent visual perception and analysis tasks.

[0163] According to the second aspect of the present invention, the present invention claims protection for a cross-modal image differential fusion system based on dual-branch feature decomposition, with reference to the attached Figure 9 As shown, including:

[0164] an acquisition unit, which acquires the first modality image and the second modality image to be fused;

[0165] an encoding unit, encoding the first modality image and the second modality image to extract corresponding low-frequency basic features and high-frequency detail features respectively;

[0166] A fusion unit, wherein the low-frequency basic features of the first modality image and the low-frequency basic features of the second modality image are input into a feature fusion module to obtain a basic fusion feature; the high-frequency detail features of the first modality image and the high-frequency detail features of the second modality image are input into the feature fusion module to obtain a detail fusion feature; and the basic fusion coding feature and the detail fusion coding feature are spliced ​​together;

[0167] Among them, the feature fusion module includes a differential compensation module, which calculates differential features based on input features, uses a sigmoid function to generate channel weights based on the differential features after global average pooling, calculates feature compensation values ​​based on the differential features and corresponding channel weights, and adds the feature compensation values ​​to the corresponding input features as compensation features. The feature fusion module reconstructs the compensation features of all input features to obtain fused features;

[0168] The decoding unit decodes the splicing result to obtain the target fused image.

[0169] In a feasible embodiment, the encoding unit further includes a shared feature extractor and a basic encoder: the shared feature extractor extracts shallow features of the input image, and the basic encoder extracts corresponding low-frequency basic features from the shallow features: the basic encoder includes a first extraction module, a global modeling module, a second extraction module and a feedforward neural network;

[0170] Wherein, the first extraction module and the second extraction module are used to extract local structural information of the input features;

[0171] The global modeling module includes a flattening layer, a state modeling layer, and a residual fusion layer. The flattening layer flattens the two-dimensional feature map into a sequence vector. The state modeling layer adds the state vector to the dynamic offset and normalizes it into an attention weight. The attention weight and state feature are multiplied with the input feature of the state modeling layer. The transformation strength is dynamically adjusted by the scaling factor, and then the result is element-by-element multiplication with the projection matrix.

[0172] The feedforward neural network is used to perform nonlinear transformation on input features.

[0173] In a feasible implementation, the global modeling module further includes a normalization layer, which is used to normalize the output of the flattening layer using a LayerNorm method.

[0174] In a feasible implementation, the first extraction module, the global modeling module, the second extraction module and the feedforward neural network of the basic encoder further include a fusion layer, and the fusion layer performs a residual operation on the output and input of the corresponding modules.

[0175] In a feasible embodiment, the encoding unit further includes a detail encoder, which extracts corresponding high-frequency detail features from the shallow features: the detail encoder divides the shallow features into three input sub-features along the channel dimension, extracts detail features from each input sub-feature respectively, and then performs a splicing operation on all extracted results to obtain the corresponding high-frequency detail features;

[0176] The step of extracting detail features from the first input sub-feature includes:

[0177] Extracting global information of the first input sub-feature through a bidirectional scanning Mamba module to obtain a first feature map;

[0178] Performing Haar wavelet transform on the first input sub-feature according to a preset low-pass filter and a high-pass filter to obtain a frequency domain feature;

[0179] Convolution is used to extract frequency domain features to extract local information;

[0180] Perform inverse wavelet transform on local information to obtain the transformed result;

[0181] The first feature map and the second feature map are summed to obtain the feature extraction result.

[0182] In a feasible implementation, extracting detail features from the second input sub-feature includes:

[0183] Divide the second input sub-feature into several second input grandchild features along the channel dimension;

[0184] Each second input sub-feature is subjected to a local convolution operation with different kernel sizes to obtain a feature extraction sub-result;

[0185] All feature extraction sub-results are spliced ​​according to channels to obtain the feature extraction result.

[0186] In a feasible implementation, the detail features of the third input sub-feature are extracted by using identity mapping.

[0187] In a feasible implementation manner, the images to be fused are infrared images and visible light images.

[0188] Refer to the attached Figure 10 As shown, an embodiment of the present application provides an electronic device, including: a processor and a memory, the processor and the memory are interconnected and communicate with each other through a communication bus and / or other forms of connection mechanisms (not shown), the memory stores a computer program executable by the processor, and when the computing device is running, the processor executes the computer program to execute the system in any optional implementation mode of the above embodiment.

[0189] The embodiment of the present application provides a storage medium, and when the computer program is executed by a processor, the system of any optional implementation of the above embodiment is executed. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0190] In the embodiments provided in this application, it should be understood that the disclosed system can be implemented in other ways. The system embodiments described above are merely schematic. For example, the division of the modules is only a logical function division, and can be implemented in another way. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, indirect coupling or communication connection of the system or unit, which can be electrical, mechanical or other forms.

[0191] In addition, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0192] Furthermore, the functional modules in the various embodiments of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0193] Flowcharts are used herein to illustrate the steps of the methods of the embodiments of the present disclosure. It should be understood that the preceding or following steps do not necessarily need to be performed in exact order. Instead, the various steps may be evaluated in reverse order or simultaneously. Furthermore, other operations may be added to these processes.

[0194] Unless otherwise defined, all terms used herein have the same meaning as commonly understood by those skilled in the art to which the present disclosure belongs. It should also be understood that terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and should not be interpreted in an idealized or highly formal sense unless expressly defined as such herein.

[0195] The above is a detailed introduction to the cross-modal image differential fusion method and system based on dual-branch feature decomposition. This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only an embodiment of this application. It is only used to help understand the cross-modal image differential fusion method and system based on dual-branch feature decomposition of this application, and is not used to limit the scope of protection of this application. At the same time, for those skilled in the art, this application can have various changes and variations. Any modifications and equivalent substitutions made within the spirit and principles of this application should be included in the scope of protection of this application.

Claims

1. A cross-modal image differential fusion method based on dual-branch feature decomposition, characterized by: include: Acquiring a first modality image and a second modality image to be fused; Encoding the first modality image and the second modality image to extract corresponding low-frequency basic features and high-frequency detail features respectively; The low-frequency basic features of the first modality image and the low-frequency basic features of the second modality image are input into the feature fusion module to obtain basic fusion features; the high-frequency detail features of the first modality image and the high-frequency detail features of the second modality image are input into the feature fusion module to obtain detail fusion features; Among them, the feature fusion module includes a differential compensation module, which calculates differential features based on input features, uses a sigmoid function to generate channel weights based on the differential features after global average pooling, calculates feature compensation values ​​based on the differential features and corresponding channel weights, and adds the feature compensation values ​​to the corresponding input features as compensation features. The feature fusion module reconstructs the compensation features of all input features to obtain fused features; Concatenate the basic fusion coding features and the detail fusion coding features; After decoding the splicing result, the target fused image is obtained; In encoding the first modality image and the second modality image, the method further includes: extracting shallow features of the input image by a shared feature extractor, and extracting corresponding low-frequency basic features from the shallow features by a basic encoder, wherein the basic encoder includes a first extraction module, a global modeling module, a second extraction module, and a feedforward neural network; Wherein, the first extraction module and the second extraction module are used to extract local structural information of the input features; The global modeling module includes a flattening layer, a state modeling layer, and a residual fusion layer. The flattening layer flattens the two-dimensional feature map into a sequence vector. The state modeling layer adds the state vector to the dynamic offset and normalizes it into an attention weight. The attention weight and state feature are multiplied with the input feature of the state modeling layer. The transformation strength is dynamically adjusted by the scaling factor, and then the result is element-by-element multiplication with the projection matrix. Wherein, the feedforward neural network is used to perform nonlinear transformation on the input features; The encoding of the first modality image and the second modality image further includes extracting corresponding high-frequency detail features from the shallow features using a detail encoder: the detail encoder divides the shallow features into three input sub-features along the channel dimension, extracts detail features from each input sub-feature, and then performs a splicing operation on all extracted results to obtain corresponding high-frequency detail features; The step of extracting detail features from the first input sub-feature includes: Extracting global information of the first input sub-feature through a bidirectional scanning Mamba module to obtain a first feature map; Performing Haar wavelet transform on the first input sub-feature according to a preset low-pass filter and a high-pass filter to obtain a frequency domain feature; Convolution is used to extract frequency domain features to extract local information; Perform inverse wavelet transform on local information to obtain the transformed result; The first feature map and the second feature map are summed to obtain the feature extraction result.

2. The cross-modal image difference fusion method based on dual-branch feature decomposition according to claim 1 is characterized in that: The global modeling module also includes a normalization layer, which is used to normalize the output of the flattening layer using the LayerNorm method.

3. The cross-modal image difference fusion method based on dual-branch feature decomposition according to claim 2 is characterized in that: The first extraction module, the global modeling module, the second extraction module and the feedforward neural network of the basic encoder further include a fusion layer, which performs a residual operation on the output and input of the corresponding modules.

4. The cross-modal image difference fusion method based on dual-branch feature decomposition according to any one of claims 1 to 3, characterized in that: Extracting detail features from the second input sub-feature includes: Divide the second input sub-feature into several second input grandchild features along the channel dimension; Each second input sub-feature is subjected to a local convolution operation with different kernel sizes to obtain a feature extraction sub-result; All feature extraction sub-results are spliced ​​according to channels to obtain the feature extraction result.

5. The cross-modal image difference fusion method based on dual-branch feature decomposition according to claim 4 is characterized in that: The detail features of the third input sub-feature are extracted using the identity mapping method.

6. The cross-modal image difference fusion method based on dual-branch feature decomposition according to claim 1, characterized in that: The images to be fused are infrared images and visible light images.

7. A cross-modal image differential fusion system based on dual-branch feature decomposition, characterized by: include: an acquisition unit, which acquires the first modality image and the second modality image to be fused; an encoding unit, encoding the first modality image and the second modality image to extract corresponding low-frequency basic features and high-frequency detail features respectively; A fusion unit, wherein the low-frequency basic features of the first modality image and the low-frequency basic features of the second modality image are input into a feature fusion module to obtain a basic fusion feature; the high-frequency detail features of the first modality image and the high-frequency detail features of the second modality image are input into the feature fusion module to obtain a detail fusion feature; and the basic fusion coding feature and the detail fusion coding feature are spliced ​​together; Among them, the feature fusion module includes a differential compensation module, which calculates differential features based on input features, uses a sigmoid function to generate channel weights based on the differential features after global average pooling, calculates feature compensation values ​​based on the differential features and corresponding channel weights, and adds the feature compensation values ​​to the corresponding input features as compensation features. The feature fusion module reconstructs the compensation features of all input features to obtain fused features; A decoding unit decodes the splicing result to obtain the target fused image; The encoding unit also includes a shared feature extractor and a basic encoder: the shared feature extractor extracts shallow features of the input image, and the basic encoder extracts corresponding low-frequency basic features from the shallow features: the basic encoder includes a first extraction module, a global modeling module, a second extraction module and a feedforward neural network; Wherein, the first extraction module and the second extraction module are used to extract local structural information of the input features; The global modeling module includes a flattening layer, a state modeling layer, and a residual fusion layer. The flattening layer flattens the two-dimensional feature map into a sequence vector. The state modeling layer adds the state vector to the dynamic offset and normalizes it into an attention weight. The attention weight and state feature are multiplied with the input feature of the state modeling layer. The transformation strength is dynamically adjusted by the scaling factor, and then the result is element-by-element multiplication with the projection matrix. Wherein, the feedforward neural network is used to perform nonlinear transformation on the input features; The encoding unit also includes a detail encoder, which extracts corresponding high-frequency detail features from shallow features: the detail encoder divides the shallow features into three input sub-features along the channel dimension, extracts detail features from each input sub-feature, and then performs a splicing operation on all the extracted results to obtain the corresponding high-frequency detail features; The step of extracting detail features from the first input sub-feature includes: Extracting global information of the first input sub-feature through a bidirectional scanning Mamba module to obtain a first feature map; Performing Haar wavelet transform on the first input sub-feature according to a preset low-pass filter and a high-pass filter to obtain a frequency domain feature; Convolution is used to extract frequency domain features to extract local information; Perform inverse wavelet transform on local information to obtain the transformed result; The first feature map and the second feature map are summed to obtain the feature extraction result.

Citation Information

Patent Citations

  • Feature weighted face identification algorithm

    CN104866831A

  • Infrared and visible light image fusion method based on feature difference compensation and fusion

    CN116883303A