A Multimodal Image Fusion Method, System, Medium and Device
Through the domain adaptive feature decomposition and adaptive fusion methods, the feature inconsistency and information retention problems in multimodal image fusion are solved, and efficient multimodal image fusion is achieved, and image quality and information integrity are improved.
Patent Information
- Application Number
- CN202510599737.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The existing multimodal image fusion technology has many challenges in dealing with the differences in different modal features, balancing basic features and detailed features, batch normalization of design domain adaptability, and achieving efficient fusion, including modal feature inconsistency, balance of basic features and detailed features, insufficient adaptability of cross-domain feature statistical distribution, and insufficient information retention and enhancement.
The domain adaptive feature decomposition mechanism is adopted to clearly distinguish features into two categories: basic features and detailed features, and different domain-specific batch normalization and fusion strategies are applied to them. Combined with adaptive fusion methods, multi-stage training strategies and multiple attention mechanisms, it is processed through encoders, decoders and multi-stage training frameworks.
The quality and information integrity of multimodal image fusion are improved, the adaptability and robustness of the model to different modal images are enhanced, and the high-quality image fusion effect is achieved.
Smart Images

Figure CN120107084B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology. Specifically, it relates to a multi-modal image fusion method, system, medium and device. Background Art
[0002] The statements in this section merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] Multi-modal image fusion aims to integrate image information from different sensors or imaging devices to generate a fused image containing complementary information of each source image. In particular, the fusion of visible light images and infrared images has important value in many application scenarios.
[0004] With the rise of deep learning technology, image fusion methods based on deep neural networks have shown great potential. These methods usually adopt an autoencoder structure, including an encoder, a fusion layer, and a decoder, which can automatically learn the high-level feature representation of images and optimize the fusion process through end-to-end training. Existing deep learning methods include multi-focus image fusion methods based on convolutional neural networks, which fuse image features of different foci by learning weight maps; infrared and visible light image fusion models based on generative adversarial networks, which improve the visual quality of the fused image through adversarial training, etc.
[0005] Although the prior art has made certain progress, it still faces the following problems that need to be solved urgently:
[0006] (1) The problem of dealing with the inconsistency and difference of different modal features. Visible light images and infrared images come from different imaging principles and sensors, and there are significant differences in feature expression, statistical distribution, and information content. Visible light images are rich in texture, edges, and color information, while infrared images highlight thermal radiation targets and can penetrate certain obstacles. Traditional fusion methods often treat different modal features as homogeneous information, ignoring the differences in feature expression of different modal images, and are unable to effectively distinguish and adapt to these differences, resulting in feature conflicts or information loss during the fusion process.
[0007] (2) The balance problem between the consistency of the basic feature domain and the specificity of the detail feature domain. In multi-modal image fusion, the ideal situation is to retain the complementary advantages of different modalities while eliminating redundancy and conflicts. However, existing methods often adopt a unified strategy when dealing with basic features and detail features, lacking targeted differential processing. In fact, for basic features (such as main structures and contours), the consistency between domains should be pursued to ensure the integrity of the structure of the fused image; while for detail features (such as texture and hot spots), their domain specificity should be retained to enrich the fusion result.
[0008] (3) Cross - domain feature statistical distribution adaptation problem. Traditional batch normalization methods have limitations when dealing with multi - modal data and cannot effectively adapt to the statistical characteristics of different domains. During the fusion process, if the inter - domain statistical distribution differences are not considered, it is easy to lead to a decline in feature expression ability or bias towards a specific domain.
[0009] (4) Information retention and enhancement problems during the fusion process. During the feature fusion stage, how to effectively integrate information from different modalities without causing information loss is a key challenge. Simple weighted averaging or feature concatenation often cannot fully utilize the complementarity of multi - modal information.
[0010] (5) Training strategy and objective function design problems. The training of multi - modal image fusion models faces challenges such as scarce samples and complex objective function design. Simple reconstruction loss cannot fully guide the model to learn effective fusion strategies.
[0011] In summary, there are still many challenges in existing multi - modal image fusion technologies in dealing with the differences of different - modality features, balancing basic features and detailed features, designing domain - adaptive batch normalization, and achieving efficient fusion. Summary of the Invention
[0012] To solve the above problems, the present invention provides a multi - modal image fusion method, system, medium and device, and innovatively proposes a domain - adaptive feature decomposition mechanism. By using different feature extraction methods, features are clearly divided into two categories: basic features and detailed features, and different domain - specific batch normalization and fusion strategies are applied to them, enabling the model to more finely process different types of image information and enhancing the quality and information integrity of the fused image.
[0013] To achieve the above object, the present invention adopts the following technical solutions:
[0014] The first aspect of the present invention provides a multi - modal image fusion method, which includes:
[0015] Obtain visible - light images and infrared images;
[0016] For visible light images and infrared images, basic feature extraction and detailed feature extraction are respectively performed through an encoder; for the basic feature extraction, self-attention operation and channel-local composite attention operation are respectively performed on the image features first, and after the two attention outputs are fused through a gating mechanism, they are sequentially processed through a feed-forward network and domain-specific batch normalization to obtain basic features; for the detailed feature extraction, the image features are first divided into two branches, and after detailed node operation is applied to each branch, channel dimension splicing operation, channel-local composite attention operation, attention weight fusion operation and domain-specific batch normalization are sequentially performed to obtain detailed features; wherein, the domain-specific batch normalization adjusts the scaling parameter and the offset parameter by using a specificity factor, and the specificity factor used in the detailed feature extraction is greater than the specificity factor used in the basic feature extraction;
[0017] The basic features of the visible light image and the infrared image are adaptively fused to obtain fused basic features; the detailed features of the visible light image and the infrared image are adaptively fused to obtain fused detailed features;
[0018] Taking the visible light image as a reference image, combining the fused basic features and the fused detailed features, and through a decoder, a fused output image is obtained.
[0019] Further, the step of adaptively fusing the basic features of the visible light image and the infrared image includes: respectively applying a feature analysis network to the basic features of the visible light image and the basic features of the infrared image, calculating the channel attention weight, and performing smoothing processing to obtain the smoothed channel attention weight; calculating the spatial attention weight and the cross-domain interaction feature based on the basic features of the visible light image and the basic features of the infrared image; calculating the attention weight fusion feature based on the smoothed channel attention weights of the visible light image and the infrared image; after fusing the attention weight fusion feature, the spatial attention weight and the cross-domain interaction feature, feature recalibration and domain batch normalization processing are sequentially performed.
[0020] Further, the step of adaptively fusing the detailed features of the visible light image and the infrared image includes: respectively applying a domain analysis network to the detailed features of the visible light image and the detailed features of the infrared image, calculating the domain analysis score; calculating the intensity difference based on the detailed features of the visible light image and the detailed features of the infrared image; calculating the domain difference based on the domain analysis score of the visible light image and the domain analysis score of the infrared image; creating a selection mask based on the intensity difference and the domain difference; after calculating and obtaining the preliminary fused detailed features based on the selection mask, the detailed features of the visible light image and the detailed features of the infrared image, cross-domain interaction addition, feature recalibration and domain batch normalization processing are sequentially performed.
[0021] Further, the decoder includes: merging the fused basic features and the fused detailed features to generate unified features, enhancing the features by using a channel attention mechanism, processing the enhanced unified features through multiple transformer blocks, introducing a local attention mechanism to extract local features, processing the local features through an output layer, performing a residual connection with a reference image, and obtaining a fused output image through an activation function.
[0022] Further, the encoder, the adaptive fusion, and the decoder form a multi-modal image fusion model, and the multi-modal image fusion model adopts a multi-stage training strategy.
[0023] Further, the multi-stage training strategy includes a self-reconstruction training stage, and the self-reconstruction training stage adopts a reconstruction loss, a domain constraint loss, an SSIM loss, an MSE loss, and a feature stability loss.
[0024] Further, the multi-stage training strategy includes a fusion training stage, and the fusion training stage adopts a channel correlation loss and a fusion quality loss.
[0025] The second aspect of the present invention provides a multi-modal image fusion system, which includes:
[0026] An image acquisition module, which is configured to: acquire a visible light image and an infrared image;
[0027] An encoding module, which is configured to: respectively perform basic feature extraction and detailed feature extraction on the visible light image and the infrared image through an encoder; for the basic feature extraction, first perform self-attention operations and channel-local composite attention operations on the image features respectively, and fuse the two attention outputs through a gating mechanism, and then sequentially perform processing through a feed-forward network and domain-specific batch normalization to obtain basic features; for the detailed feature extraction, first divide the image features into two branches, and after applying detailed node operations to each branch, sequentially perform channel dimension splicing operations, channel-local composite attention operations, attention weight fusion operations, and domain-specific batch normalization to obtain detailed features; wherein, the domain-specific batch normalization adjusts the scaling parameter and the offset parameter by using a specificity factor, and the specificity factor used in the detailed feature extraction is greater than the specificity factor used in the basic feature extraction;
[0028] An adaptive fusion module, which is configured to: adaptively fuse the basic features of the visible light image and the infrared image to obtain fused basic features; adaptively fuse the detailed features of the visible light image and the infrared image to obtain fused detailed features;
[0029] A decoding module, which is configured to: use the visible light image as a reference image, combine the fused basic features and the fused detailed features, and obtain a fused output image through a decoder.
[0030] The third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps in a multi-modal image fusion method as described above are implemented.
[0031] The fourth aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the program, the steps in a multi-modal image fusion method as described above are implemented.
[0032] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0033] The present invention provides a multi-modal image fusion method, which breaks through the limitation of homogeneous processing of all features in traditional image fusion, and innovatively proposes a domain adaptation feature decomposition mechanism. This mechanism uses different feature extraction methods to clearly classify features into two categories: basic features and detail features, and applies different processing strategies to them; the basic features mainly include the structural and contour information of the image. By encouraging inter-domain consistency, the basic features from different modalities tend to be similar, providing a stable structural basis for fusion; the detail features include key detail information such as texture and hot spots. By maintaining domain specificity, it ensures that the fusion result can retain the unique advantages of each modality; this feature decomposition mechanism enables the model to process different types of image information more finely, enhancing the quality and information integrity of the fused image.
[0034] The present invention provides a multi-modal image fusion method, which designs an innovative domain-specific batch normalization method to solve the problem of inconsistent statistical distributions in multi-modal data processing. Different from traditional batch normalization, this method introduces multi-dimensional statistical features including mean, standard deviation, and local contrast when calculating statistics, enhancing the expression ability of image features. More importantly, through domain label encoding and feature modulation mechanisms, it dynamically adjusts the normalization parameters according to the input domain and feature type, achieving adaptive processing of data from different domains: for basic features, batch normalization tends to weaken domain specificity to enhance consistency; while for detail features, it retains or even enhances domain specificity to preserve unique information. This differential batch normalization strategy significantly improves the model's ability to process multi-modal data.
[0035] The present invention provides a multimodal image fusion method, which innovatively designs an adaptive fusion method. Different fusion strategies are adopted for the basic features and the detailed features respectively, replacing the traditional simple feature fusion method. The adaptive fusion method not only considers the importance analysis of the channel dimension, but also introduces a spatial attention mechanism and cross-domain feature interaction, realizing multi-dimensional feature enhancement and fusion. In the basic feature fusion, it tends to find the common parts of the features in different domains and generates a consistent structural representation through weighted fusion; in the detailed feature fusion, the most significant and discriminative detailed information in each domain is retained through feature intensity and domain specificity analysis. This differential fusion strategy based on feature types effectively improves the quality and information integrity of the fused image.
[0036] The present invention provides a multimodal image fusion method, which introduces a variety of attention mechanisms in the feature extraction and fusion process, including Channel Attention, Local Attention, CLAttention (channel-local composite attention), and the self-attention mechanism based on Transformer, forming a multi-level and multi-dimensional feature enhancement system. These attention mechanisms work together to enhance the model's ability to focus on key information at different levels: Channel Attention identifies important feature channels, Local Attention highlights key regions in space, and self-attention captures long-range dependencies. Through the integrated application of this multiple attention mechanism, the model can more accurately identify and retain the key information in different modal images, improving the fusion effect.
[0037] The present invention provides a multimodal image fusion method, which proposes a dual optimization strategy of domain constraint and feature stability. The multimodal image fusion model is guided to learn an ideal feature representation through a carefully designed loss function: the domain constraint loss encourages the basic features to maintain consistency between different domains by calculating the statistical differences of the features between domains, while allowing the detailed features to maintain domain specificity; the feature stability loss further enhances the similarity of the basic features and the difference of the detailed features through cosine similarity calculation. This dual optimization strategy not only improves the quality of the feature representation, but also enhances the stability and generalization ability of the multimodal image fusion model in cross-domain processing, laying a foundation for high-quality image fusion.
[0038] The present invention provides a multi-modal image fusion method, which designs an innovative multi-stage training framework, including a self-reconstruction training stage and a fusion training stage, to achieve a progressive improvement in the model's capabilities. In the self-reconstruction training stage, the multi-modal image fusion model learns to reconstruct images in different domains respectively, enhancing the expression ability of features in each domain and establishing domain constraints; in the fusion training stage, the multi-modal image fusion model learns how to effectively fuse features in different domains to generate high-quality fused images. This progressive learning framework enables the model to gradually master the ability to express single-domain features and fuse multi-domain features, improving the training efficiency and the performance of the multi-modal image fusion model; at the same time, by sharing the encoder and decoder parameters, knowledge transfer between different tasks is realized, further enhancing the generalization ability of the model.
[0039] The present invention provides a multi-modal image fusion method, which solves key problems in existing multi-modal image fusion technologies such as feature inconsistency processing, balance between basic and detailed features, cross-domain feature statistical adaptation, and information retention and enhancement, significantly improving the quality, robustness, and adaptability of multi-modal image fusion, and providing a new technical solution for the fields of computer vision and image processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute a limitation to the present invention.
[0041] Figure 1 It is an architecture diagram of a multi-modal image fusion method according to Embodiment 1 of the present invention;
[0042] Figure 2 It is an overall network structure diagram of Embodiment 1 of the present invention;
[0043] Figure 3 It is an encoder network structure diagram of Embodiment 1 of the present invention;
[0044] Figure 4 It is a domain-specific batch normalization network structure diagram of Embodiment 1 of the present invention;
[0045] Figure 5 It is an adaptive fusion network structure diagram of Embodiment 1 of the present invention;
[0046] Figure 6 It is a decoder network structure diagram of Embodiment 1 of the present invention;
[0047] Figure 7 It is a schematic diagram of an infrared image according to Embodiment 1 of the present invention;
[0048] Figure 8 It is a schematic diagram of a visible light image according to Embodiment 1 of the present invention;
[0049] Figure 9 Schematic diagram of the fused output image of the first embodiment of the present invention. Specific implementation manner
[0050] The present invention will be further described below in conjunction with the drawings and embodiments.
[0051] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0052] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be further described below in conjunction with the drawings and embodiments.
[0053] First Embodiment
[0054] The purpose of the first embodiment is to provide a multi-modal image fusion method.
[0055] A multi-modal image fusion method provided by this embodiment aims to improve the effect and robustness of multi-modal image fusion through innovative feature decomposition and fusion mechanisms. Traditional image fusion methods often treat features of different modalities as homogeneous information, making it difficult to effectively capture and utilize the differences and complementarities between modalities. To address this issue, this embodiment is based on the transfomer architecture and introduces a domain adaptation feature decomposition mechanism, an innovative batch normalization method, adaptive fusion, and a multi-stage training strategy, effectively solving the above technical problems and providing a new solution for the field of multi-modal image fusion.
[0056] A multi-modal image fusion method provided by this embodiment has a core idea of decomposing multi-modal image features into basic features and detail features. It pursues domain consistency for basic features and maintains domain specificity for detail features, achieving efficient fusion of different modal images through this differential processing strategy. Specifically, the encoder extracts basic features and detail features from the input image and distinguishes different input modalities through domain labels; after the features are processed by domain-specific batch normalization, they are reconstructed or fused in the decoder. Through the designed multi-stage training strategy, the model first learns the self-reconstruction ability within each domain, then optimizes the cross-domain feature fusion ability, and finally achieves high-quality multi-modal image fusion. This architecture design enables the model to retain the respective advantageous features when processing visible light and infrared images and generate high-quality fused images through an adaptive fusion strategy.
[0057] A multi-modal image fusion method provided by this embodiment, such as Figure 1As shown, it includes the following steps:
[0058] Step 1: Data input and preprocessing.
[0059] First, collect the original image data from the multi-modal sensors, including visible light images and infrared images , where each image satisfies , is the image height, is the image width, is the number of channels (for example, when it is an RGB image = 3).
[0060] Then, preprocess the collected original image data. The preprocessing process mainly includes normalization, size adjustment, noise removal, and enhancement operations.
[0061] In this embodiment, normalize each image, map the pixel values to a fixed range, for example, scale the pixel values to the interval of [0, 1] or [-1, 1]. The normalization formula is: , where, and respectively represent the minimum and maximum pixel values of the image.
[0062] In this embodiment, adjust the images to a unified size × through interpolation or cropping operations to ensure that all images have the same size in the subsequent processing, thus simplifying the requirements for batch processing and network input.
[0063] In this embodiment, to reduce noise and enhance contrast, methods such as filtering and histogram equalization are used to optimize the image quality, so that the input images can better adapt to the feature extraction of the deep network.
[0064] On this basis, in order to distinguish images of different modalities, corresponding domain labels are assigned to each image during the preprocessing; specifically, the domain label of the visible light image is assigned as = 1, while the domain label of the infrared image is assigned as = 0. This assignment of domain labels can be represented by the function , that is: if is a visible light image, then ; if is an infrared image, then .
[0065] After preprocessing, the obtained image data are respectively denoted as and , and the corresponding domain labels and , where Preprocess represents preprocessing.
[0066] These standardized, noise-suppressed, and enhanced input data provide a solid data foundation for subsequent feature extraction, decomposition, and fusion by the encoder, ensuring stable and consistent data flow in the end-to-end multi-modal image fusion process.
[0067] Step 2: Feature decomposition and domain adaptation mechanism.
[0068] One of the core innovations of this embodiment is the introduction of a domain-adaptive feature decomposition mechanism, which effectively solves the problem of feature inconsistency in multi-modal image fusion by decomposing the feature representation into basic features and detailed features and imposing different domain constraints.
[0069] Decomposition of basic features and detailed features. Feature decomposition is located inside the encoder, which decomposes the input feature map into two types of basic features and detailed features, and applies different processing strategies, such as Figure 3 shown, specifically including:
[0070] , here, Encoder represents the encoder, represents the basic features in the visible light domain extracted by the basic feature extractor, mainly including the structure and contour information of the image; represents the detailed features in the visible light domain extracted by the detailed feature extractor, including local information such as texture and edges; is the visible light input image, is the visible light domain label (value is 1).
[0071] Similarly, , represents the basic features in the infrared domain extracted by the basic feature extractor, represents the detailed features in the infrared domain extracted by the detailed feature extractor, is the infrared input image, is the infrared domain label (value is 0).
[0072] (1) The encoder first extracts the feature map from the input image: ; where is the feature map, represents the computational mapping function of the encoder, is the input image (visible light domain image or infrared domain image), are the network parameters. This step extracts the high-level feature representation of the image through multi-layer convolution and transformer blocks. Among them, for inputs in different domains, when and , the feature representation in the visible light domain is obtained; when And At this time, the feature representation in the infrared domain is obtained.
[0073] (2) Subsequently, the basic feature extractor and the detailed feature extractor respectively perform feature decomposition on the image features in the visible light domain and the infrared domain after feature mapping.
[0074] Among them, the basic feature extractor adopts the multi-head attention mechanism, and its calculation process:
[0075] First, apply the self-attention operation: , where attn_out is the feature enhanced by self-attention, attn represents the self-attention operation, norm1 is layer normalization, is the input image feature, and the self-attention operation can capture the long-range dependence relationship between different positions of the feature;
[0076] At the same time, apply channel-local composite attention: , where cl_attn_out is the feature enhanced by channel-local composite attention, cl_attn represents the channel-local composite attention operation, and this attention mechanism simultaneously considers the relationship between channels and local spatial information to enhance the attention to local important features;
[0077] The outputs of the two kinds of attention are fused through a gating mechanism: , , where gate is the calculated fusion gating parameter, represents the Sigmoid activation function, is the learnable parameter, and attn_combined is the fused feature. This gating fusion mechanism enables the model to adaptively combine the advantages of the two kinds of attention;
[0078] Subsequently, apply a feed-forward network to process and obtain the basic features: , where, mlp is a multi-layer perceptron, and norm 2 is layer normalization. This step enhances the expression ability of the features and introduces non-linear transformation;
[0079] Finally, apply domain-specific batch normalization to the basic features: , where, represents the domain-specific batch normalization operation, is the domain label, out_B represents the basic features after domain-specific batch normalization. This operation dynamically adjusts the normalization parameters according to the domain label to make the basic features enhance their expression ability while maintaining the consistency between domains.
[0080] Among them, the detailed feature extractor realizes feature decomposition through a multi-layer detailed node structure, and its calculation process is as follows:
[0081] First, the input image features are divided into two branches: , where and are the two branches of the features, is the number of channels. This channel splitting enables the model to process different types of feature information separately;
[0082] Apply the detailed node operation to each branch: , where layer is the detailed node operation, including non-linear transformation and interaction between branches:
[0083] First, merge and reallocate features through the shuffle convolution operation: , where represents the channel dimension concatenation.
[0084] Split the shuffled features into two branches: ;
[0085] Apply the first detailed transformation to the second branch: , where is the non-linear transformation layer, capturing the interaction information between branches;
[0086] Apply the composite transformation to the first branch: , where: Introduce non-linear modulation through the exponential mapping, Provide additional feature offsets, As the adaptive scaling factor, Provide residual enhancement.
[0087] This detailed node operation enhances the model's ability to capture and represent detailed features through complex non-linear transformations and interactions between branches, providing rich feature representations for multi-modal image fusion.
[0088] Feature branch fusion: , where represents the channel dimension concatenation operation, is the concatenated feature. The purpose of the concatenation operation is to retain all the feature information of the two branches, providing rich feature representations for subsequent attention enhancement.
[0089] Apply channel-local composite attention enhancement: , where cl _ out is the feature after attention enhancement, cl _ attnIt is a channel-local composite attention operation, which simultaneously considers the inter-channel relationship and local spatial information, enhancing the attention to key regions.
[0090] Attention weight fusion is carried out through learnable weight parameters: , , where is the attention weight, is the learnable parameter, is the activation function, is the feature after fusion through the gating mechanism. This weighted fusion mechanism enables the model to adaptively control the influence degree of attention;
[0091] Finally, domain-specific batch normalization is applied to the detailed features to obtain the detailed features after domain-specific batch normalization: , where represents domain-specific batch normalization, is the domain label. In this way, the model can dynamically adjust the feature representation according to the feature characteristics of different domains, retaining the specific information of each domain.
[0092] Among them, the steps of domain-specific batch normalization include:
[0093] (1) Introduction and role of domain labels.
[0094] The present invention introduces the domain label to guide the domain adaptation processing of features, defined as: , , , where represents the visible light domain, represents the infrared domain, represents the fusion domain. This numerical representation enables the model to clearly distinguish different input modalities.
[0095] The domain label is mapped to an embedding vector through an encoder: , where is the embedding representation of the domain label, is the domain label encoder, represents the domain label, which is 1, 0 or 0.5. This embedding representation contains rich information about domain characteristics.
[0096] The embedding vector is used to generate modulation parameters: , where is the modulation parameter, is the modulation parameter generator. These parameters are used for subsequent feature modulation.
[0097] The modulation parameters are divided into a scaling part and an offset part: , where is the scaling parameter, is the offset parameter, and these two parameters control the amplitude and reference value of the feature respectively.
[0098] (2)Domain-specific batch normalization.
[0099] (201)When implementing domain-specific batch normalization, it is necessary to calculate the multi-dimensional statistical characteristics of the input features to provide a comprehensive description of the features. The present invention designs the domain-specific batch normalization to calculate the multi-dimensional statistical characteristics of the input features: , where represents the input feature, including three statistics: mean, standard deviation, and local contrast, providing a multi-dimensional description of the feature.
[0100] Among them, the calculation method of the mean is: , represents the value at the i-th row and j-th column of the input feature, and are the height and width of the feature map respectively, and the mean reflects the global brightness level of the feature.
[0101] Among them, the calculation method of the standard deviation is: , and the standard deviation describes the global contrast and variation degree of the feature.
[0102] Among them, the calculation method of the local contrast is: , and the local contrast reflects the local details and texture information of the feature. These multi-dimensional statistical features interact with the feature through learnable correction coefficients, affecting the output of batch normalization. Specifically, the calculation method of statistical feature correction is:
[0103] (202)Perform calibration value calculation: , where represents the k-th statistic, represents the correction coefficient corresponding to the k-th statistic; perform feature correction application: , where is the output of standard batch normalization for the input feature , is the Sigmoid activation function, is the corrected feature (basic feature or detail feature).
[0104] This batch normalization mechanism guided by multi-dimensional statistical features enables the model to dynamically adjust the normalization effect according to the statistical characteristics of the features, enhancing the adaptability to features in different domains. For basic features and detail features, this mechanism shows different behaviors when applying domain-specific parameters, thus achieving domain consistency for basic features and maintaining domain specificity for detail features.
[0105] For the base features, domain-specific batch normalization first calculates the domain-specific factor: , where is the domain-specific factor, is the learnable parameter, is the Sigmoid function, and the base features use a lower domain-specific factor (about 0.1); then the scale and offset parameters are adjusted: , , this adjustment weakens the influence of the domain-specific parameters and enhances domain consistency; finally, the adjusted parameters are applied: , where is the base feature after domain-specific batch normalization, that is out _ B , represents the base feature, and the tanh function restricts the range of the scale parameter.
[0106] (204)For the detail features, first calculate the domain-specific factor: , the detail features use a higher domain-specific factor (about 0.8) to enhance domain specificity; then the scale and offset parameters are adjusted: , , this adjustment enhances the influence of the domain-specific parameters and retains domain specificity; the processing strategy is further adjusted according to the domain label:
[0107] When processing infrared domain features (i.e., obtained from infrared images): , , where is the infrared domain enhancement term, which enhances the characteristic patterns unique to infrared;
[0108] When processing visible light domain features (i.e., obtained from visible light images): , , where is the detail feature after domain-specific batch normalization, that is out_D , is the visible light domain enhancement term, which enhances the characteristic patterns unique to visible light;
[0109] When processing the fused domain features ( or ): , , is the fused domain feature after domain-specific batch normalization, and the fused domain adopts medium-strength adjustment to maintain a balanced feature representation.
[0110] Through the above design, the domain adaptation feature decomposition mechanism of the present invention effectively addresses the differences in feature expression of multimodal images, providing an ideal input for subsequent feature fusion and serving as the key technical foundation for achieving high-quality multimodal image fusion.
[0111] Step 3: Adaptive fusion.
[0112] In this embodiment, adaptive fusion is the core component of multimodal image fusion, used to intelligently fuse features from different domains. It adopts a differential strategy to process the base features and detail features, achieving efficient feature integration and enhancement.
[0113] Adaptive fusion receives the base features and detail features from the visible light domain and the infrared domain, applies different fusion strategies to them respectively, and then outputs the fused features for image reconstruction. Its overall structure can be expressed as:
[0114] , where is the fused base feature, and are the base features of the infrared domain and the visible light domain respectively, is the fusion domain label (value is 0.5);
[0115] , where is the fused detail feature, and are the detail features of the infrared domain and the visible light domain respectively.
[0116] The fused features are fed into the decoder to generate the final output: , where data_Fuse is the output image after fusion, feature_F is the intermediate feature representation, is provided to the decoder by the visible light reference image to guide the generation of the final fused image. The reason for choosing the visible light image as the reference is mainly based on three considerations: First, visible light images usually contain richer structural and texture details, which can provide more complete spatial structure information for the fused image; Second, the human visual system is more accustomed to the natural performance of visible light imaging, and using the visible light image as the reference can maintain the visual naturalness of the fusion result; Third, in many application scenarios (such as night driving assistance, security monitoring, etc.), infrared images are mainly used to supplement the missing thermal target information in visible light images, while the basic structure is still mainly based on visible light imaging. Through this design, the decoder can effectively integrate the thermal target information in the infrared image while retaining the structure of the visible light image, generating a fused image with the advantages of both modalities.
[0117] (1) Base feature fusion strategy, namely AdaptiveFusion.forward_base.
[0118] As Figure 5 shown, the basic feature fusion strategy mainly focuses on domain consistency, aiming to integrate the common structural information in different domains.
[0119] (101) First, calculate the channel attention weights: , where is the channel attention weight of the infrared domain basic features.
[0120] FeatureAnalysis is a feature analysis network that calculates the channel importance through adaptive average pooling and convolutional layers. Its structure contains multiple levels: . Specifically, the feature analysis network consists of the following parts: , the graph is compressed to a fixed size (1×1), only retaining the channel dimension information. This step captures the global statistical characteristics of each channel. , where is the feature after dimensionality reduction. is a 1×1 convolutional layer that reduces the number of channels to 1 / 4 of the original. ReLU is the activation function. This step realizes the dimensionality reduction of the feature channels, reduces the number of parameters, and extracts the key correlations between channels. After dimensionality reduction, through another 1×1 convolutional layer and Sigmoid activation function, the final channel attention weights are generated: , where is the final channel attention weight. is the Sigmoid activation function, and the output range is between 0 and 1, representing the importance score of each channel. This step generates the normalized channel attention weights for subsequent feature weighting. Through this "squeeze-and-excitation" mechanism, the feature analysis network can effectively evaluate the importance of different feature channels and provide adaptive weight guidance for feature fusion.
[0121] (102) Similarly, , V _ attn is the channel attention weight of the visible light domain basic features.
[0122] (103) To stabilize the fusion process, smooth the attention weights: , , this smoothing operation ensures that all feature channels have a basic contribution and prevents some channels from being completely ignored.
[0123] (104) Then, calculate the spatial attention weights: , , where max_feat is the element-wise maximum of the two domain features, and avg_feat is their average; , where spatial_attn is the spatial attention weight, and SpatialAttention is the spatial attention module that calculates the spatial importance of features through a convolutional layer.
[0124] (105) Add cross-domain interaction to enhance feature representation: , where cross_features is the cross-domain interaction feature, and CrossDomainInteraction is a 1×1 convolutional layer that captures the interaction relationship between the features of two domains.
[0125] (106) Finally, fuse features based on the attention weight: , , where weighted_feat is the attention-weighted fused feature, and weight_sum is the weight normalization factor.
[0126] (107) Introduce spatial attention and cross-domain interaction to further enhance the fusion effect: , where is the final fused basic feature.
[0127] This fusion strategy combines channel attention, spatial attention, and cross-domain interaction, and can effectively integrate the common information of different domains and retain the structural consistency.
[0128] (108) Perform feature recalibration and domain batch normalization on the fused basic feature in sequence: , , where Rescale is the feature recalibration operation, and DomainBN is the domain-specific batch normalization. These operations make the fused features more suitable for subsequent decoding processing.
[0129] (2) Detail feature fusion strategy, namely AdaptiveFusion.forward_detail.
[0130] As Figure 6 shown, different from the basic features, the detail feature fusion strategy focuses on retaining domain specificity, and the goal is to select the most significant detail information.
[0131] (201) First, calculate the domain analysis score: , where I_ domain_score is the domain analysis score of the detail features in the infrared domain.
[0132] The domain analysis network (DomainAnalysis) is specifically used to evaluate the domain-specific intensity of features and is a key component for processing detail feature fusion. Its structural design takes into account the unique properties of domain features. The specific composition of the domain analysis network is as follows: , where The feature is obtained through adaptive average pooling, which captures the global statistical information of the entire feature map; , where is the intermediate feature layer, Conv is a 1×1 convolutional layer, and LeakyReLU is an activation function with a small slope (usually 0.2). Different from the feature analysis network, LeakyReLU is used here instead of the standard ReLU, which can retain negative value information and is crucial for evaluating domain differences; , where is the domain analysis score, and Tanh is the hyperbolic tangent activation function with an output range between -1 and 1. Different from the feature analysis network that uses the Sigmoid function, the Tanh function is used here because domain analysis needs to distinguish features of different domains (positive and negative values), rather than just evaluating importance (0 to 1); Through this design, the domain analysis network can learn the feature differences between different domains and generate domain analysis with positive and negative distinctions.
[0133] (202) Similarly, , V_domain_score is the domain analysis score of the visible light domain detail features.
[0134] (203) Calculate the feature intensity difference and domain difference: , , where I_intensity and V_intensity are the intensities of the infrared domain and visible light domain detail features respectively; , , where intensity_diff is the intensity difference and domain_diff is the domain difference. These differences are used to guide feature selection. dim=1 specifies to calculate the mean in the channel dimension, and keepdim=True means to retain the calculated dimension shape (the dimension becomes 1 but is not removed), which ensures that the dimension structure of the calculated tensor matches the original feature and facilitates subsequent calculation of the feature intensity difference and mask creation.
[0135] (204) Then create a selection mask: , where mask is the selection mask, is the Sigmoid function. The area where the mask value is close to 1 selects the infrared features, and the area close to 0 selects the visible light features.
[0136] (205) Calculate the mask-based fused detail features: , where fused_feature is the preliminary fused detail feature, which adaptively selects the most significant information based on the feature intensity and domain specificity.
[0137] (206) Then add cross - domain interaction to enhance feature representation: , , where cross_features enhances the interaction between the two domain features and adds additional information to the fused features.
[0138] (207) Perform feature recalibration and domain batch normalization on the fused features in sequence: , , These operations make the fused features more suitable for subsequent decoding processing while maintaining domain adaptability.
[0139] Step 4, Decoding and reconstruction.
[0140] As Figure 6 shown, decoding and reconstruction is the last link of the processing flow, responsible for converting the fused features into the final output image. It receives the output of adaptive fusion and reconstructs a high - quality fused image through a series of transformation operations.
[0141] The input of the decoder (Restormer_Decoder) is the fused base feature and fused detail feature generated by adaptive fusion: , where, is the output image after fusion, that is, data_Fuse, is the intermediate feature representation, that is, feature_F, is the optional reference image, that is img_VI , and are the fused base feature and detail feature respectively, and Decoder represents the decoder. As Figure 7 , Figure 8 and Figure 9 shown, they are the infrared image, visible light image and output image after fusion respectively.
[0142] Step 401, The decoder first combines these two types of features to generate a unified feature: , where concat represents concatenating features along the channel dimension, and ReduceChannel is a 1×1 convolutional layer that halves the number of channels to reduce computational complexity.
[0143] Step 402, Feature enhancement and processing.
[0144] To enhance the feature representation ability, the decoder uses a channel attention mechanism to optimize the feature channel weights: , where ChannelAttention enhances the expression of key feature channels by learning the importance of different channels.
[0145] The enhanced unified feature is further processed through multiple transformer blocks: , where Encoder Level 2 consists of multiple Transformer blocks, each of which contains a self-attention mechanism and a feed-forward network, and can capture long-range dependencies of features.
[0146] Step 403, Local attention and output generation.
[0147] To further enhance the expression of important regions, the decoder introduces a local attention mechanism to extract local features: , where LocalAttention focuses on the spatial importance of the feature map, enabling the model to focus on regions with high information content.
[0148] Among them, the local attention mechanism The calculation process is as follows: , , where is the spatial attention weight, is the gating parameter, and the spatial importance of the feature map is calculated through a 7×7 convolution.
[0149] Step 404, Finally, the final fused image is generated through the output layer: .
[0150] Step 405, If there is a reference image, residual connection is adopted: ; Otherwise, directly output the processing result: ; Finally, the output range is restricted through the Sigmoid activation function: .
[0151] The decoder, together with the previous encoder, feature decomposition, and adaptive fusion, forms a complete processing chain, and realizes end-to-end optimization through parameter sharing and gradient backpropagation. During the multi-stage training process, the decoder learns both how to reconstruct single-domain images and how to generate high-quality fused images, so as to adapt to different processing requirements.
[0152] Through this design, the decoder can effectively convert the fused features into a fused image with high visual quality and rich information, providing a complete technical solution for multi-modal image fusion.
[0153] In this embodiment, the encoder, adaptive fusion, and decoder form a complete multi-modal image fusion model, and the multi-modal image fusion model adopts a multi-stage training strategy.
[0154] The multi-stage training strategy adopted in this embodiment includes a self-reconstruction training stage and a fusion training stage, enabling the multi-modal image fusion model to gradually master the ability to express single-domain features and fuse multi-domain features, improving the training efficiency and model performance.
[0155] (1)Self-reconstruction training stage.
[0156] Self-reconstruction training is the first stage of multi-stage training. The main purpose is to enhance the expression ability of the multi-modal image fusion model for the features of each domain and establish domain constraints.
[0157] In this stage, the multi-modal image fusion model first clears the optimizer gradient, and then extracts the basic features and detailed features from the visible light domain and infrared domain images: , , where and are the basic features of the visible light domain and infrared domain respectively, and are the corresponding detailed features.
[0158] Apply the enhanced domain constraint loss to guide the basic features to be consistent while keeping the detailed features different: , where is the total domain constraint loss, is the domain constraint loss of the basic features, is the domain constraint loss of the detailed features, and are the weight coefficients of the basic feature and detailed feature constraints respectively. The enhanced domain constraint loss enhancedDomainConstraints is one of the key innovations of the present invention. It realizes domain adaptability constraints by calculating the statistical differences of the basic features and detailed features respectively, specifically including:
[0159] (a)For the basic features, calculate the difference in the means of the two domain features: , , , where c , i and j represent the channel index, height index, and width index of the feature tensor respectively; represents the pixel value of the visible light domain basic feature at the c-th channel, i-th row, and j-th column; represents the pixel value of the infrared domain basic feature at the c-th channel, i-th row, and j-th column; and represent the channel means of the visible light domain and infrared domain basic features respectively. This mean difference loss encourages the basic features to maintain statistical consistency between different domains and helps to capture common structural information.
[0160] (b)For the detailed features, calculate the cosine similarity of the normalized feature directions: , , , , , By minimizing the absolute value of the cosine similarity, this loss encourages the directions of the detailed features to remain orthogonal, thus preserving the domain-specific information.
[0161] This design enables the model to achieve domain consistency in the basic features while maintaining domain specificity in the detailed features, which is very suitable for the multi-modal image fusion task.
[0162] (c) Reconstruct the original image using the extracted features and calculate the reconstruction loss: , , . Among them, represents the visible light image reconstructed by the decoder using the visible light image, the basic features in the visible light domain, and the detailed features; represents the infrared image reconstructed by the decoder using the infrared image, the basic features in the infrared domain, and the detailed features; represents the reconstruction loss, is the weight parameter of the reconstruction loss, and the default value is 1.
[0163] (d) In addition to the reconstruction loss, this embodiment also introduces a variety of loss functions for joint optimization, including the SSIM loss, the MSE loss, and the feature stability loss: , ; among them, and are both weight parameters; the feature stability loss is another innovation point, which evaluates the domain consistency and specificity of the features by calculating the cosine similarity of the feature channel means: , , , among them, and are both weight parameters. The design of the feature stability loss : For the basic feature : It is hoped that the basic features are similar between different domains (the loss is low when the similarity is high); for the detailed feature : It is hoped that the detailed features maintain differences between different domains (the loss is low when the similarity is low); through this loss design, the model can pursue domain consistency at the basic feature level and maintain domain differences at the detailed feature level.
[0164] (e) Calculate the total loss and perform backpropagation and parameter update: .
[0165] (2) Fusion training stage.
[0166] The fusion training is the second stage of the multi-stage training. The main purpose is to optimize the feature fusion ability of the multi-modal image fusion model. In this stage, the multi-modal image fusion model reextracts features, creates a fusion domain label (with a value of 0.5), and then adaptively fuses the features of different domains: , , .
[0167] A specific loss function is introduced in the fusion training stage. First is the channel correlation loss: , , , where cc represents the channel correlation loss calculation function. This innovative correlation loss design has two purposes: on the one hand, it penalizes the low correlation of the base features through a low denominator (small cc_B value), and on the other hand, it penalizes the high correlation of the detail features through a high numerator (large cc_D value); adding 1.01 to the denominator is for numerical stability. This design further strengthens the domain consistency of the base features and the domain specificity of the detail features, forming a complement to the domain constraint loss.
[0168] Another key loss is the fusion quality loss: , and the fusion quality loss function criteria_fusion is a comprehensive evaluation index. This multi-faceted evaluation ensures that the fused image not only retains the key information of the source image but also has good visual effects.
[0169] The total loss in the fusion training stage is: .
[0170] (3) Monitoring the training process and optimizing parameters.
[0171] To monitor the change of domain consistency during the training process, this embodiment introduces a domain distance calculation function. By calculating the difference in feature means and covariances, the inter-domain distance is quantified: , .
[0172] Among them, the domain distance calculation function computeDomainDistance quantifies the inter-domain distance by calculating the statistical differences between features of different domains, and it can be expressed as: , and are the inputs to the domain distance calculation function, and mean_dist and cov_dist are the outputs of the domain distance calculation function.
[0173] Calculation of the mean distance: First, the channel mean of the feature needs to be obtained: , where H and W are the height and width of the feature map respectively, Denote the feature map All channel values at the \(i\)-th row and \(j\)-th column; similarly, calculate the channel mean of the second feature: , where Denote the feature map All channel values at the \(i\)-th row and \(j\)-th column; then calculate the mean squared error between the means as the mean distance: .
[0174] To calculate the covariance distance, first obtain the channel covariance of the feature: . Where, reshape( B , C , -1) reshapes the tensor into the shape of ( B , C , -1), B is the batch size, C is the number of channels, and the var( ) function calculates the variance along the 2nd dimension (channel dimension). The purpose is to evaluate the degree of change of the feature in the channel dimension. A lower variance means that the feature changes less across different channels, and a higher variance means that the feature changes more across different channels.
[0175] During the training process, gradient clipping is adopted to prevent gradient explosion, a learning rate scheduler is used to gradually reduce the learning rate, and the multi-modal image fusion model parameters are saved regularly; the initial learning rate is set to 10 -4 , and the learning rate is reduced to 50% of the original every 20 training steps; this multi-stage training strategy enables the model to first master the single-domain feature expression ability and then learn an effective feature fusion strategy, thus achieving high-quality multi-modal image fusion. Through a carefully designed loss function system, the multi-modal image fusion model can learn an ideal feature representation and fusion rule, improving the quality and information integrity of the fused image.
[0176] To address the problem of dealing with the inconsistency and difference of different modal features, in this embodiment, by introducing a domain adaptation decomposition mechanism, the feature is divided into two parts: the basic feature and the detail feature, and different processing strategies are applied respectively, enabling the model to better adapt to the characteristics of different modal features.
[0177] To address the balance problem between the domain consistency of the basic feature and the domain specificity of the detail feature, in this embodiment, by designing different domain constraint strategies, the basic features tend to be consistent while the detail features remain different, thus ensuring both the structural stability and the retention of the unique information of each modality during the fusion process.
[0178] For the problem of cross - domain feature statistical distribution adaptation, in this embodiment, a domain - specific batch normalization mechanism is designed to dynamically adjust the normalization parameters according to the feature type and domain label, enabling the network to adaptively adjust its statistical characteristics according to different domains and feature types and enhancing its adaptability to multimodal data.
[0179] For the problem of information retention and enhancement in the fusion process, in this embodiment, adaptive fusion is designed. Different fusion strategies are adopted for the basic features and detailed features, and an attention mechanism and cross - domain interaction are introduced to enhance the expression of key information, suppress redundant information, and improve the fusion effect.
[0180] For the problem of training strategy and objective function design, in this embodiment, a multi - stage training method is proposed. First, the expression ability of each domain's features is enhanced through self - reconstruction training, and then the cross - domain feature fusion ability is optimized through fusion training. At the same time, a composite loss function including reconstruction loss, domain constraint loss, structural similarity loss, etc. is designed to comprehensively guide the optimization of the model.
[0181] Embodiment 2
[0182] The purpose of this Embodiment 2 is to provide a multimodal image fusion system, including:
[0183] An image acquisition module, which is configured to: acquire visible - light images and infrared images;
[0184] An encoding module, which is configured to: for the visible - light images and infrared images, respectively perform basic feature extraction and detailed feature extraction through an encoder. For the basic feature extraction, self - attention operation and channel - local composite attention operation are first performed on the image features respectively, and after fusing the two attention outputs through a gating mechanism, they are successively processed by a feed - forward network and domain - specific batch normalization to obtain basic features. For the detailed feature extraction, the image features are first divided into two branches, and after applying detailed node operations to each branch, they are successively subjected to channel - dimension splicing operation, channel - local composite attention operation, attention weight fusion operation, and domain - specific batch normalization to obtain detailed features. Among them, the domain - specific batch normalization uses a specificity factor to adjust the scaling parameter and the offset parameter, and the specificity factor used in the detailed feature extraction is greater than that used in the basic feature extraction;
[0185] An adaptive fusion module, which is configured to: adaptively fuse the basic features of the visible - light images and infrared images to obtain fused basic features; adaptively fuse the detailed features of the visible - light images and infrared images to obtain fused detailed features;
[0186] A decoding module, which is configured to: use the visible - light image as a reference image, and combine the fused basic features and fused detailed features to obtain a fused output image through a decoder.
[0187] It should be noted here that each module in this embodiment corresponds to each step in the first embodiment one by one, and their specific implementation processes are the same, so they will not be repeated here.
[0188] Embodiment Three
[0189] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps in a multi-modal image fusion method as described in the above-mentioned Embodiment One are implemented.
[0190] Embodiment Four
[0191] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the program, the steps in a multi-modal image fusion method as described in the above-mentioned Embodiment One are implemented.
[0192] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0193] Although the specific implementation manners of the present invention are described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solutions of the present invention, various modifications or deformations that can be made without creative efforts by those skilled in the art are still within the protection scope of the present invention.
Claims
1. A multimodal image fusion method, characterized in that Including: Obtain visible light images and infrared images; For the visible light images and infrared images, respectively perform basic feature extraction and detailed feature extraction through an encoder; for the basic feature extraction, first perform self-attention operations and channel-local composite attention operations on the image features respectively, and after fusing the outputs of the two attentions through a gating mechanism, sequentially process through a feed-forward network and domain-specific batch normalization to obtain basic features; for the detailed feature extraction, first divide the image features into two branches, and after applying detailed node operations to each branch, sequentially perform channel dimension splicing operations, channel-local composite attention operations, attention weight fusion operations and domain-specific batch normalization to obtain detailed features; wherein, the domain-specific batch normalization adjusts the scaling parameter and the offset parameter using a specific factor, and the specific factor used for detailed feature extraction is greater than the specific factor used for basic feature extraction; Adaptively fuse the basic features of the visible light images and infrared images to obtain fused basic features; adaptively fuse the detailed features of the visible light images and infrared images to obtain fused detailed features; Use the visible light image as a reference image, and combine the fused basic features and fused detailed features, and through a decoder, obtain the fused output image; The step of adaptively fusing the basic features of the visible light images and infrared images includes: applying a feature analysis network to the basic features of the visible light images and the basic features of the infrared images respectively, calculating channel attention weights, and performing smoothing processing to obtain the smoothed channel attention weights; calculating spatial attention weights and cross-domain interaction features based on the basic features of the visible light images and the basic features of the infrared images; calculating attention weight fusion features based on the smoothed channel attention weights of the visible light images and infrared images; after fusing the attention weight fusion features, spatial attention weights and cross-domain interaction features, sequentially perform feature recalibration and domain batch normalization processing; The step of adaptively fusing the detailed features of the visible light images and infrared images includes: applying a domain analysis network to the detailed features of the visible light images and the detailed features of the infrared images respectively, calculating domain analysis scores; calculating intensity differences based on the detailed features of the visible light images and the detailed features of the infrared images; calculating domain differences based on the domain analysis scores of the visible light images and the domain analysis scores of the infrared images; creating a selection mask based on the intensity differences and domain differences; after calculating and obtaining the preliminary fused detailed features based on the selection mask, the detailed features of the visible light images and the detailed features of the infrared images, sequentially perform cross-domain interaction addition, feature recalibration and domain batch normalization processing.
2. The multimodal image fusion method according to claim 1, wherein The decoder includes: combining the fused basic features and fused detailed features to generate unified features, and performing feature enhancement using a channel attention mechanism, processing the enhanced unified features through multiple transformer blocks, introducing a local attention mechanism to extract local features, after the local features are processed through an output layer, performing a residual connection with the reference image, and obtaining the fused output image through an activation function.
3. A multimodal image fusion method according to claim 1, characterized in that, The encoder, adaptive fusion, and decoder form a multi-modal image fusion model, and the multi-modal image fusion model adopts a multi-stage training strategy.
4. The multimodal image fusion method according to claim 3, characterized in that, The multi-stage training strategy includes a self-reconstruction training stage, and the self-reconstruction training stage adopts reconstruction loss, domain constraint loss, SSIM loss, MSE loss, and feature stability loss.
5. The multimodal image fusion method according to claim 3, characterized in that, The multi-stage training strategy includes a fusion training stage, and the fusion training stage adopts channel correlation loss and fusion quality loss.
6. A multimodal image fusion system adopting a multimodal image fusion method as described in claim 1, characterized in that, It includes: An image acquisition module configured to acquire visible light images and infrared images; An encoding module configured to perform basic feature extraction and detailed feature extraction on visible light images and infrared images respectively through an encoder; for the basic feature extraction, self-attention operation and channel-local composite attention operation are first performed on the image features respectively, and after the two attention outputs are fused through a gating mechanism, they are sequentially processed through a feed-forward network and domain-specific batch normalization to obtain basic features; for the detailed feature extraction, the image features are first divided into two branches, and after applying detailed node operations to each branch, channel dimension concatenation operation, channel-local composite attention operation, attention weight fusion operation, and domain-specific batch normalization are sequentially performed to obtain detailed features; wherein, the domain-specific batch normalization adjusts the scaling parameter and the offset parameter using a specificity factor, and the specificity factor used in the detailed feature extraction is greater than the specificity factor used in the basic feature extraction; An adaptive fusion module configured to adaptively fuse the basic features of visible light images and infrared images to obtain fused basic features; and adaptively fuse the detailed features of visible light images and infrared images to obtain fused detailed features; A decoding module configured to use the visible light image as a reference image, and combine the fused basic features and the fused detailed features to obtain a fused output image through a decoder.
7. A computer-readable storage medium having a computer program stored thereon, the program being executed by a processor, characterized in that, When the program is executed by a processor, it implements the steps in a multi-modal image fusion method as described in any one of claims 1-5.
8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in a multi-modal image fusion method as described in any one of claims 1-5.
Citation Information
Patent Citations
Multi-modal vehicle environment image fusion segmentation method, device, equipment and medium
CN119027666A
Target detection method and device based on multi-modal fusion, electronic equipment and medium
CN119625263A