Multi-modal image fusion method, system, medium and equipment

By adopting domain adaptive feature decomposition mechanism and adaptive fusion method in multimodal image fusion, combining multiple attention mechanisms and domain-specific batch normalization, problems such as feature differential processing, feature balance and information retention in multimodal image fusion are solved, and the fusion effect and quality are significantly improved.

CN120107084AActive Publication Date: 2025-06-06TAISHAN UNIV

Patent Information

Application Number
CN202510599737.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-06-06
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

The existing multimodal image fusion technology still has many challenges in dealing with the differences in different modal features, balancing basic features and detailed features, design domain adaptability batch normalization, and achieving efficient fusion.

Method used

The domain adaptive feature decomposition mechanism is adopted to clearly distinguish features into two categories: basic features and detailed features, and different domain-specific batch normalization and fusion strategies are imposed on them. Through the adaptive fusion method, multi-dimensional feature enhancement and fusion are achieved by combining channel attention, local attention and cross-domain feature interaction.

Benefits of technology

It significantly improves the quality, robustness and adaptability of multimodal image fusion, enhances the quality and information integrity of the fusion image, and solves key problems such as feature inconsistency, balance of basic and detailed features, statistical adaptation of cross-domain features, and information retention and enhancement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107084A_ABST
    Figure CN120107084A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, and discloses a multi-modal image fusion method and system, a medium and equipment, and the method comprises the steps: carrying out the basic feature extraction and detail feature extraction of a visible light image and an infrared image; according to the basic feature extraction, after self-attention operation and channel-local composite attention operation are carried out on image features, gating mechanism fusion, feedforward network processing and domain specific batch normalization are carried out in sequence; according to the detail feature extraction, image features are divided into two branches, and then detail node operation, channel dimension splicing operation, channel-local composite attention operation, attention weight fusion operation and domain specific batch normalization are carried out in sequence. Carrying out adaptive fusion on the basic features; and carrying out adaptive fusion on the detail features, and obtaining a fused output image through a decoder. And the quality and information integrity of the fused image are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a multimodal image fusion method, system, medium and device. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] Multimodal image fusion aims to integrate image information from different sensors or imaging devices to generate a fused image containing complementary information of each source image. In particular, the fusion of visible light images and infrared images is of great value in many application scenarios.

[0004] With the rise of deep learning technology, image fusion methods based on deep neural networks have shown great potential. These methods usually adopt an autoencoder structure, including three parts: encoder, fusion layer and decoder. They can automatically learn the high-level feature representation of the image and optimize the fusion process through end-to-end training. Existing deep learning methods include multi-focus image fusion methods based on convolutional neural networks, which fuse image features of different focuses by learning weight maps; infrared and visible light image fusion models based on generative adversarial networks, which improve the visual quality of fused images through adversarial training, etc.

[0005] Although the existing technology has made some progress, it still faces the following problems that need to be solved urgently: (1) Problems in processing inconsistencies and differences in features of different modalities. Visible light images and infrared images come from different imaging principles and sensors, and have significant differences in feature expression, statistical distribution, and information content. Visible light images are rich in texture, edge, and color information, while infrared images highlight thermal radiation targets and can penetrate certain obstructions. Traditional fusion methods often treat features of different modalities as homogeneous information, ignoring the differences in feature expression of images of different modalities. They are unable to effectively distinguish and adapt to these differences, resulting in feature conflicts or information loss during the fusion process.

[0006] (2) The balance between domain consistency of basic features and domain specificity of detailed features. In multimodal image fusion, the ideal situation is to retain the complementary advantages of different modalities while eliminating redundancy and conflict. However, existing methods often adopt a unified strategy when dealing with basic features and detailed features, lacking targeted differentiated processing. In fact, for basic features (such as main structures and contours), inter-domain consistency should be pursued to ensure the integrity of the fused image structure; while for detailed features (such as textures and hotspots), their domain specificity should be retained to enrich the fusion results.

[0007] (3) Cross-domain feature statistical distribution adaptation problem. Traditional batch normalization methods have limitations when processing multimodal data and cannot effectively adapt to the statistical characteristics of different domains. In the fusion process, if the statistical distribution differences between domains are not considered, it is easy to lead to a decrease in feature expression ability or bias towards a specific domain.

[0008] (4) Information preservation and enhancement during the fusion process. In the feature fusion stage, how to effectively integrate information from different modalities without causing information loss is a key challenge. Simple weighted averaging or feature concatenation often cannot fully utilize the complementarity of multimodal information.

[0009] (5) Training strategy and objective function design issues. The training of multimodal image fusion models faces challenges such as lack of samples and complex objective function design. Simple reconstruction loss cannot fully guide the model to learn effective fusion strategies.

[0010] In summary, the existing multimodal image fusion technology still faces many challenges in dealing with the differences in features of different modalities, balancing basic features and detail features, designing domain-adaptive batch normalization, and achieving efficient fusion. Summary of the invention

[0011] In order to solve the above problems, the present invention provides a multimodal image fusion method, system, medium and device, innovatively proposes a domain-adaptive feature decomposition mechanism, adopts different feature extraction methods, clearly distinguishes features into two categories: basic features and detail features, and applies different domain-specific batch normalization and fusion strategies to them, so that the model can process different types of image information more finely, thereby enhancing the quality and information integrity of the fused image.

[0012] In order to achieve the above object, the present invention adopts the following technical solution: A first aspect of the present invention provides a multimodal image fusion method, comprising: Acquire visible light images and infrared images; For visible light images and infrared images, basic feature extraction and detail feature extraction are performed respectively through the encoder; the basic feature extraction first performs self-attention operation and channel-local composite attention operation on the image features respectively, and after the two attention outputs are fused through the gating mechanism, they are processed by the feedforward network and domain-specific batch normalization in sequence to obtain the basic features; the detail feature extraction first divides the image features into two branches, and after applying the detail node operation to each branch, the channel dimension splicing operation, the channel-local composite attention operation, the attention weight fusion operation and the domain-specific batch normalization are performed in sequence to obtain the detail features; wherein the domain-specific batch normalization uses a specific factor to adjust the scaling parameter and the offset parameter, and the specific factor used for the detail feature extraction is greater than the specific factor used for the basic feature extraction; Adaptively fuse the basic features of the visible light image and the infrared image to obtain fused basic features; adaptively fuse the detail features of the visible light image and the infrared image to obtain fused detail features; The visible light image is used as the reference image, combined with the fusion basic features and the fusion detail features, and passed through the decoder to obtain the fused output image.

[0013] Furthermore, the step of adaptively fusing the basic features of the visible light image and the infrared image includes: applying a feature analysis network to the basic features of the visible light image and the basic features of the infrared image, calculating the channel attention weights, and performing smoothing to obtain the smoothed channel attention weights; calculating the spatial attention weights and cross-domain interaction features based on the basic features of the visible light image and the basic features of the infrared image; calculating the attention weight fusion features based on the smoothed channel attention weights of the visible light image and the infrared image; after fusing the attention weight fusion features, the spatial attention weights and the cross-domain interaction features, performing feature recalibration and domain batch normalization processing in sequence.

[0014] Furthermore, the step of adaptively fusing the detail features of the visible light image and the infrared image includes: applying a domain analysis network to the detail features of the visible light image and the detail features of the infrared image, respectively, to calculate the domain analysis score; calculating the intensity difference based on the detail features of the visible light image and the detail features of the infrared image; calculating the domain difference based on the domain analysis score of the visible light image and the domain analysis score of the infrared image; creating a selection mask based on the intensity difference and the domain difference; after calculating the preliminary fused detail features based on the selection mask, the detail features of the visible light image and the detail features of the infrared image, cross-domain interactive addition, feature recalibration and domain batch normalization are performed in sequence.

[0015] Furthermore, the decoder includes: merging the fused basic features and the fused detail features to generate a unified feature, and using a channel attention mechanism to enhance the features, after the enhanced unified features are processed by multiple transformer blocks, a local attention mechanism is introduced to extract local features, after the local features are processed by the output layer, a residual connection is performed with the reference image, and a fused output image is obtained through an activation function.

[0016] Furthermore, the encoder, adaptive fusion and decoder form a multimodal image fusion model, and the multimodal image fusion model adopts a multi-stage training strategy.

[0017] Furthermore, the multi-stage training strategy includes a self-reconstruction training stage, and the self-reconstruction training stage adopts reconstruction loss, domain constraint loss, SSIM loss, MSE loss and feature stability loss.

[0018] Furthermore, the multi-stage training strategy includes a fusion training stage, and the fusion training stage adopts channel correlation loss and fusion quality loss.

[0019] A second aspect of the present invention provides a multimodal image fusion system, comprising: An image acquisition module, which is configured to: acquire a visible light image and an infrared image; The encoding module is configured as follows: for a visible light image and an infrared image, basic feature extraction and detail feature extraction are performed respectively through an encoder; the basic feature extraction first performs self-attention operation and channel-local composite attention operation on the image features respectively, and after fusing the two attention outputs through a gating mechanism, the basic features are obtained by sequentially processing through a feedforward network and performing domain-specific batch normalization; the detail feature extraction first divides the image features into two branches, and after applying a detail node operation to each branch, sequentially performs a channel dimension splicing operation, a channel-local composite attention operation, an attention weight fusion operation and a domain-specific batch normalization to obtain detail features; wherein the domain-specific batch normalization uses a specific factor to adjust a scaling parameter and an offset parameter, and the specific factor used for the detail feature extraction is greater than the specific factor used for the basic feature extraction; The adaptive fusion module is configured to: adaptively fuse the basic features of the visible light image and the infrared image to obtain the fused basic features; adaptively fuse the detail features of the visible light image and the infrared image to obtain the fused detail features; The decoding module is configured to: use the visible light image as a reference image, combine the fusion basic features and the fusion detail features, and obtain a fused output image through a decoder.

[0020] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, the program being executed by a processor, and the program implementing the steps in a multimodal image fusion method as described above when executed by the processor.

[0021] A fourth aspect of the present invention provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the program, the steps in a multimodal image fusion method as described above are implemented.

[0022] Compared with the prior art, the present invention has the following beneficial effects: The present invention provides a multimodal image fusion method, which breaks through the limitation of traditional image fusion that all features are processed homogeneously, and innovatively proposes a domain-adaptive feature decomposition mechanism, which adopts different feature extraction methods to clearly distinguish features into two categories: basic features and detail features, and applies different processing strategies to them; the basic features mainly include the structure and contour information of the image, and by encouraging inter-domain consistency, the basic features from different modalities tend to be similar, providing a stable structural basis for fusion; the detail features include key detail information such as texture and hot spots, and by maintaining domain specificity, it is ensured that the fusion result can retain the unique advantages of each modality; this feature decomposition mechanism enables the model to process different types of image information more finely, thereby enhancing the quality and information integrity of the fused image.

[0023] The present invention provides a multimodal image fusion method, which designs an innovative domain-specific batch normalization method to solve the problem of inconsistent statistical distribution in multimodal data processing. Unlike traditional batch normalization, this method introduces multidimensional statistical features including mean, standard deviation and local contrast when calculating statistics, which enhances the ability to express image features. More importantly, it dynamically adjusts normalization parameters according to the input domain and feature type through domain label encoding and feature modulation mechanism, and realizes adaptive processing of data in different domains: for basic features, batch normalization tends to weaken domain specificity to enhance consistency; for detail features, domain specificity is retained or even enhanced to preserve unique information. This differentiated batch normalization strategy significantly improves the model's ability to process multimodal data.

[0024] The present invention provides a multimodal image fusion method, which innovatively designs an adaptive fusion method, adopts different fusion strategies for basic features and detail features, and replaces the traditional simple feature fusion method. The adaptive fusion method not only considers the importance analysis of the channel dimension, but also introduces the spatial attention mechanism and cross-domain feature interaction to achieve multi-dimensional feature enhancement and fusion. In basic feature fusion, = tends to find the common parts of features in different domains, and generates a consistent structural representation through weighted fusion; in detail feature fusion, the most significant and discriminative detail information in each domain is retained through feature strength and domain specificity analysis. This differentiated fusion strategy based on feature type effectively improves the quality and information integrity of the fused image.

[0025] The present invention provides a multimodal image fusion method, which introduces multiple attention mechanisms in the feature extraction and fusion process, including channel attention (ChannelAttention), local attention (LocalAttention), CLAttention (channel-local composite attention) and self-attention mechanism based on Transformer, forming a multi-level and multi-dimensional feature enhancement system. These attention mechanisms work together to enhance the model's ability to focus on key information at different levels: channel attention identifies important feature channels, local attention highlights key areas in space, and self-attention captures long-distance dependencies. Through the integrated application of this multiple attention mechanism, the model can more accurately identify and retain key information in different modal images and improve the fusion effect.

[0026] The present invention provides a multimodal image fusion method, which proposes a dual optimization strategy of domain constraint and feature stability, and guides the multimodal image fusion model to learn ideal feature representation through a carefully designed loss function: the domain constraint loss encourages the basic features to remain consistent between different domains by calculating the statistical differences of inter-domain features, while allowing the detailed features to maintain domain specificity; the feature stability loss is calculated through cosine similarity to further enhance the similarity of basic features and the difference of detailed features. This dual optimization strategy not only improves the quality of feature representation, but also enhances the stability and generalization ability of the multimodal image fusion model in cross-domain processing, laying the foundation for high-quality image fusion.

[0027] The present invention provides a multimodal image fusion method, which designs an innovative multi-stage training framework, including a self-reconstruction training stage and a fusion training stage, to achieve a progressive improvement in model capabilities. In the self-reconstruction training stage, the multimodal image fusion model learns to reconstruct images of different domains respectively, enhances the ability to express features of each domain and establishes domain constraints; in the fusion training stage, the multimodal image fusion model learns how to effectively fuse features of different domains to generate high-quality fused images. This progressive learning framework enables the model to gradually master the ability from single-domain feature expression to multi-domain feature fusion, improving training efficiency and multimodal image fusion model performance; at the same time, by sharing encoder and decoder parameters, knowledge transfer between different tasks is achieved, further enhancing the generalization ability of the model.

[0028] The present invention provides a multimodal image fusion method, which solves key problems in existing multimodal image fusion technology, such as feature inconsistency processing, basic and detail feature balance, cross-domain feature statistical adaptation, and information retention and enhancement, significantly improves the quality, robustness and adaptability of multimodal image fusion, and provides a new technical solution for the field of computer vision and image processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The accompanying drawings, which constitute a part of the specification of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but do not constitute limitations of the present invention.

[0030] Figure 1 This is an architecture diagram of a multimodal image fusion method according to the first embodiment of the present invention; Figure 2 This is an overall network structure diagram of the first embodiment of the present invention; Figure 3 This is a diagram of an encoder network structure according to Embodiment 1 of the present invention; Figure 4 This is a diagram of a domain-specific batch normalization network structure according to the first embodiment of the present invention; Figure 5 This is a diagram of an adaptive fusion network structure according to the first embodiment of the present invention; Figure 6 This is a decoder network structure diagram of the first embodiment of the present invention; Figure 7 is a schematic diagram of an infrared image according to the first embodiment of the present invention; Figure 8 is a schematic diagram of a visible light image according to the first embodiment of the present invention; Fig. 9 It is a schematic diagram of the fused output image of the first embodiment of the present invention. DETAILED DESCRIPTION

[0031] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0032] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0033] In the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other. The present invention is further described below with reference to the accompanying drawings and embodiments.

[0034] Embodiment 1 The purpose of this embodiment is to provide a multimodal image fusion method.

[0035] This embodiment provides a multimodal image fusion method, which aims to improve the effect and robustness of multimodal image fusion through innovative feature decomposition and fusion mechanisms. Traditional image fusion methods often treat the features of different modalities as homogeneous information processing, which makes it difficult to effectively capture and utilize the differences and complementarities between modalities. To address this problem, this embodiment is based on the transfomer architecture, and on this basis introduces a domain-adaptive feature decomposition mechanism, an innovative batch normalization method, adaptive fusion, and a multi-stage training strategy, which effectively solves the above technical problems and provides a new solution for the field of multimodal image fusion.

[0036] The core idea of ​​a multimodal image fusion method provided in this embodiment is to decompose the multimodal image features into two parts: basic features and detail features, pursue inter-domain consistency for the basic features, and maintain domain specificity for the detail features. This differentiated processing strategy is used to achieve efficient fusion of images of different modalities. Specifically, the encoder extracts basic features and detail features from the input image, and distinguishes different input modalities through domain labels; after the features are processed by domain-specific batch normalization, they are reconstructed or fused in the decoder. Through the designed multi-stage training strategy, the model first learns the self-reconstruction capability within each domain, and then optimizes the cross-domain feature fusion capability, and finally achieves high-quality multimodal image fusion. This architectural design enables the model to retain the respective advantageous features when processing visible light and infrared images, and generate high-quality fused images through an adaptive fusion strategy.

[0037] This embodiment provides a multimodal image fusion method, such as Figure 1 As shown, the following steps are included: Step 1: Data input and preprocessing.

[0038] First, raw image data is collected from multimodal sensors, including visible light images and infrared images , where each image satisfies , is the image height, is the image width, is the number of channels (for example, for RGB images =3).

[0039] Then, the collected original image data is preprocessed, and the preprocessing process mainly includes normalization, size adjustment, noise removal and enhancement operations.

[0040] In this embodiment, each image is normalized to map the pixel values ​​to a fixed range, for example, the pixel values ​​are scaled to the interval [0, 1] or [-1, 1]. The normalization formula is: ,in, and Represent the minimum and maximum pixel values ​​of the image respectively.

[0041] In this embodiment, the image is adjusted to a uniform size by interpolation or cropping. × , to ensure that each image has the same size in subsequent processing, thus simplifying batch processing and network input requirements.

[0042] In this embodiment, in order to reduce noise and enhance contrast, methods such as filtering and histogram equalization are used to optimize image quality so that the input image can better adapt to feature extraction of the deep network.

[0043] On this basis, in order to distinguish images of different modalities, a corresponding domain label is assigned to each image during the preprocessing process; specifically, the domain label assignment of the visible light image is =1, and the domain label of the infrared image is =0, the assignment of this domain label can be done with the function Indicates that: if is a visible light image, then ;like is an infrared image, then .

[0044] After preprocessing, the image data obtained are recorded as and , and the corresponding domain labels and , where Preprocess means preprocessing.

[0045] These standardized, noise-suppressed and enhanced input data provide a solid data foundation for the encoder's subsequent feature extraction, decomposition and fusion, ensuring stable and consistent data flow in the end-to-end multimodal image fusion process.

[0046] Step 2: Feature decomposition and domain adaptation mechanism.

[0047] One of the core innovations of this embodiment is the introduction of a domain-adaptive feature decomposition mechanism, which effectively solves the feature inconsistency problem in multimodal image fusion by decomposing feature representation into basic features and detail features and imposing different domain constraints.

[0048] Decomposition of basic features and detailed features. Feature decomposition is located inside the encoder, which decomposes the input feature map into two categories: basic features and detailed features, and applies different processing strategies to each, such as Figure 3 As shown, specifically including: , here, Encoderrepresents the encoder, Represents the basic features in the visible light domain extracted by the basic feature extractor, which mainly contain the structure and contour information of the image; Represents the detail features in the visible light domain extracted by the detail feature extractor, including local information such as texture and edge; is the visible light input image, is the visible light domain label (value is 1).

[0049] Similarly, , represents the basic features of the infrared domain extracted by the basic feature extractor, represents the detail features in the infrared domain extracted by the detail feature extractor, is the infrared input image, is the infrared domain tag (value is 0).

[0050] (1) The encoder first extracts a feature map from the input image: ;in, is the feature map, represents the encoder’s computational mapping function, is the input image (visible light domain image or infrared domain image), is the network parameter. This step extracts high-level feature representations of the image through multiple layers of convolution and transformer blocks. and When , the characteristic representation of the visible light domain is obtained; when and When , the feature representation in the infrared domain is obtained.

[0051] (2) Subsequently, the basic feature extractor and the detail feature extractor perform feature decomposition on the image features in the visible light domain and the infrared domain after feature mapping, respectively.

[0052] Among them, the basic feature extractor adopts a multi-head attention mechanism, and its calculation process is: First, apply the self-attention operation: , where attn_out is the feature enhanced by self-attention, attn represents the self-attention operation, norm1 is the layer normalization, is the input image feature, and the self-attention operation can capture the long-distance dependencies between different positions of the feature; At the same time, channel-local compound attention is applied: , where cl_attn_out is the feature enhanced by channel-local compound attention, and cl_attn represents the channel-local compound attention operation. This attention mechanism considers both the relationship between channels and local spatial information, and enhances the focus on local important features; The two attention outputs are fused through a gating mechanism: , , where gate is the calculated fusion gating parameter, represents the Sigmoid activation function, is a learnable parameter, attn_combined is the fused feature, and this gated fusion mechanism enables the model to adaptively combine the advantages of the two attentions; Then the feedforward network is applied to obtain the basic features: ,in, mlp It is a multi-layer perceptron, and norm 2 is layer normalization. This step enhances the expressiveness of features and introduces nonlinear transformations. Finally, domain-specific batch normalization is applied to the base features: ,in, represents the domain-specific batch normalization operation, is the domain label, out_B represents the base features after domain-specific batch normalization, which is done according to the domain label The normalization parameters are dynamically adjusted to enhance the expressive power of basic features while maintaining inter-domain consistency.

[0053] Among them, the detail feature extractor realizes feature decomposition through a multi-layer detail node structure, and its calculation process is: First, the input image features are divided into two branches: ,in, and There are two branches of features. is the number of channels. This channel division enables the model to process different types of feature information separately; Apply a detail node operation to each branch: , where layer is a detail node operation, including nonlinear transformations and interactions between branches: First, the features are merged and redistributed through a shuffled convolution operation: ,in, Represents channel dimension concatenation.

[0054] Split the shuffled features into two branches: ; Apply the first detail transformation to the second branch: ,in, It is a nonlinear transformation layer that captures the interaction information between branches; Apply a compound transformation to the first branch: ,in: Nonlinear modulation is introduced by exponential mapping, Provides additional feature offsets, As the adaptive scaling factor, Provides residual enhancement.

[0055] This detail node operation enhances the model's ability to capture and represent detail features through complex nonlinear transformations and interactions between branches, providing rich feature representation for multimodal image fusion.

[0056] Feature branch fusion: ,in, represents the channel dimension concatenation operation, It is the concatenated feature. The purpose of the concatenation operation is to retain all the feature information of the two branches and provide rich feature representation for subsequent attention enhancement.

[0057] Application channel - local compound attention enhancement: ,in, cl _ out is a feature that has been enhanced through attention. cl _ attn It is a channel-local composite attention operation, which considers both the inter-channel relationship and local spatial information, and enhances the attention to key areas.

[0058] Attention weight fusion through learnable weight parameters: , ,in, is the attention weight, is a learnable parameter, is the activation function, It is the feature fused through the gating mechanism. This weighted fusion mechanism enables the model to adaptively control the influence of attention. Finally, domain-specific batch normalization is applied to the detail features to obtain the domain-specific batch normalized detail features: ,in, represents domain-specific batch normalization, is the domain label. In this way, the model can dynamically adjust the feature representation according to the characteristic characteristics of different domains and retain the specific information of each domain.

[0059] The steps of domain-specific batch normalization include: (1) Introduction and function of domain labels.

[0060] The present invention introduces domain tags The domain adaptation process of the guidance feature is defined as: , , ,in, represents the visible light domain, represents the infrared domain, represents the fusion domain, and this numerical representation enables the model to clearly distinguish between different input modalities.

[0061] The domain labels are mapped into embedding vectors through the encoder: ,in, is the embedding representation of the domain label, is the domain label encoder, Represents the domain label, which is 1, 0, or 0.5. This embedding representation contains rich information about domain characteristics.

[0062] The embedding vector is used to generate the modulation parameters: ,in, is the modulation parameter, It is the modulation parameter generator, which is used for the subsequent feature modulation.

[0063] Split the modulation parameters into two parts: scale and offset: ,in, is the scaling parameter, is the offset parameter, and these two parameters control the amplitude and base value of the feature respectively.

[0064] (2) Domain-specific batch normalization.

[0065] (201) When implementing domain-specific batch normalization, it is necessary to calculate the multidimensional statistical characteristics of the input features to provide a comprehensive description of the features. The present invention designs domain-specific batch normalization to calculate the multidimensional statistical characteristics of the input features: ,in, Represents the input features, including three statistics: mean, standard deviation, and local contrast, providing a multi-dimensional description of the features.

[0066] The mean value is calculated as follows: , represents the value of the i-th row and j-th column of the input feature, and are the height and width of the feature map, respectively, and the mean reflects the global brightness level of the feature.

[0067] The standard deviation is calculated as follows: , the standard deviation describes the global contrast and variation of the features.

[0068] Among them, the local contrast is calculated as: , local contrast reflects the local details and texture information of the feature. These multidimensional statistical features interact with the features through learnable correction coefficients, affecting the output of batch normalization. Specifically, the calculation method of statistical feature correction is: (202) Calculate the calibration value: ,in, represents the kth statistic, Represents the correction coefficient corresponding to the k-th statistic; perform feature correction application: ,in, For the input features The output of standard batch normalization, is the Sigmoid activation function, is the corrected feature (basic feature or detail feature).

[0069] This batch normalization mechanism guided by multi-dimensional statistical features enables the model to dynamically adjust the normalization effect according to the statistical characteristics of the features, enhancing its adaptability to features from different domains. For basic features and detailed features, this mechanism exhibits differentiated behavior when applying domain-specific parameters, thereby achieving domain consistency of basic features and domain specificity of detailed features.

[0070] (203) For the base features, domain-specific batch normalization first calculates the domain-specific factor: ,in, is a domain-specific factor, is a learnable parameter, It is a Sigmoid function, and the base features use a low domain-specific factor (about 0.1); then adjust the scaling and offset parameters: , , this adjustment weakens the influence of domain-specific parameters and enhances domain consistency; finally, the adjusted parameters are applied: ,in, is the basic feature after domain-specific batch normalization, that is, out _ B , Represents the basic features, and the tanh function limits the range of the scaling parameter.

[0071] (204) For the detail features, the domain-specific factor is first calculated: , the detail features use a higher domain specificity factor (about 0.8) to enhance domain specificity; then adjust the scaling and offset parameters: , ,This adjustment enhances the influence of domain-specific parameters and retains domain specificity; the processing strategy is further adjusted according to the domain label: When dealing with infrared domain features (i.e., From the infrared image): , ,in, It is an infrared domain enhancement item, which enhances the characteristic patterns unique to infrared; When dealing with visible-domain features (i.e., When the image is obtained from visible light: , ,in, is the detailed feature after domain-specific batch normalization, that is, out_D , It is a visible light domain enhancement item, which enhances the characteristic patterns unique to visible light; When processing fusion domain features ( or )hour: , , It is the fusion domain feature after domain-specific batch normalization. The fusion domain adopts medium-intensity adjustment to maintain balanced feature expression.

[0072] The domain-adaptive feature decomposition mechanism of the present invention effectively handles the differences in feature expression of multimodal images through the above design, provides an ideal input for subsequent feature fusion, and is the key technical foundation for achieving high-quality multimodal image fusion.

[0073] Step 3: Adaptive fusion.

[0074] In this embodiment, adaptive fusion is the core component of multimodal image fusion, which is used to intelligently fuse features from different domains. It adopts a differentiated strategy to process basic features and detail features to achieve efficient feature integration and enhancement.

[0075] Adaptive fusion receives basic features and detail features from the visible light domain and infrared domain, applies different fusion strategies to them respectively, and then outputs the fused features for image reconstruction. Its overall structure can be expressed as: ,in, It is the fusion basic feature. and They are the basic characteristics of infrared and visible light domains, is the fusion domain label (value is 0.5); ,in, It is the integration of detail features. and They are the detailed features in the infrared domain and the visible light domain respectively.

[0076] The fused features are fed into the decoder to generate the final output: , where data_Fuse is the fused output image, feature_F is the intermediate feature representation, The visible light reference image is provided to the decoder to guide the generation of the final fused image. The reason for choosing the visible light image as a reference is mainly based on three considerations: first, visible light images usually contain richer structural and texture details, which can provide more complete spatial structure information for the fused image; second, the human visual system is more accustomed to the natural performance of visible light imaging, and using visible light images as a reference can maintain the visual naturalness of the fusion result; third, in many application scenarios (such as night driving assistance, security monitoring, etc.), infrared images are mainly used to supplement the missing thermal target information in visible light images, while the basic structure is still based on visible light imaging. Through this design, the decoder can effectively integrate the thermal target information in the infrared image while retaining the structure of the visible light image, and generate a fused image with the advantages of both modalities.

[0077] (1) Basic feature fusion strategy, namely AdaptiveFusion.forward_base.

[0078] like Figure 5 As shown in Figure 3, the basic feature fusion strategy mainly focuses on domain consistency, and its goal is to integrate common structural information in different domains.

[0079] (101) First, calculate the channel attention weight: ,in, is the channel attention weight of the infrared domain basic features.

[0080] FeatureAnalysis is a feature analysis network that calculates channel importance through adaptive average pooling and convolutional layers. Its structure contains multiple levels: Specifically, the feature analysis network consists of the following parts: , the graph is compressed to a fixed size (1×1), retaining only the channel dimension information. This step captures the global statistical characteristics of each channel. ,in, is the feature after dimensionality reduction, It is a 1×1 convolution layer, which reduces the number of channels to 1 / 4 of the original number. ReLU is the activation function. This step realizes the dimensionality reduction of feature channels, reduces the number of parameters and extracts the key correlation between channels. After dimensionality reduction, another 1×1 convolution layer and Sigmoid activation function are used to generate the final channel attention weight: ,in, is the final channel attention weight, It is a Sigmoid activation function, and the output range is between 0 and 1, which indicates the importance score of each channel. This step generates normalized channel attention weights for subsequent feature weighting. Through this "compression-excitation" mechanism, the feature analysis network can effectively evaluate the importance of different feature channels and provide adaptive weight guidance for feature fusion.

[0081] (102) Similarly, , V _ attn It is the channel attention weight of the basic features in the visible light domain.

[0082] (103) In order to stabilize the fusion process, the attention weights are smoothed: , ,This smoothing operation ensures that all feature channels have substantial contributions,preventing some channels from being completely ignored.

[0083] (104) Then, the spatial attention weight is calculated: , , where max_feat is the element-wise maximum of the two domain features and avg_feat is their average; , where spatial_attn is the spatial attention weight and SpatialAttention is the spatial attention module, which calculates the spatial importance of features through the convolutional layer.

[0084] (105) Add cross-domain interaction to enhance feature expression: , where cross_features is the cross-domain interaction feature and CrossDomainInteraction is a 1×1 convolutional layer that captures the interaction between the features of the two domains.

[0085] (106) Finally, the features are fused based on the attention weights: , , where weighted_feat is the attention weight fusion feature and weight_sum is the weight normalization factor.

[0086] (107) Introducing spatial attention and cross-domain interaction further enhances the fusion effect: ,in, It is the final fusion basic feature.

[0087] This fusion strategy combines channel attention, spatial attention and cross-domain interaction, which can effectively integrate common information from different domains and preserve structural consistency.

[0088] (108) The fused basic features are subjected to feature recalibration and domain batch normalization in turn: , , where Rescale is the feature recalibration operation and DomainBN is the domain-specific batch normalization. These operations make the fused features more suitable for subsequent decoding processing.

[0089] (2) Detail feature fusion strategy, namely AdaptiveFusion.forward_detail.

[0090] like Figure 6 As shown in Figure 2, unlike basic features, the detail feature fusion strategy focuses on preserving domain specificity and aims to select the most significant detail information.

[0091] (201) First, the domain analysis score is calculated: ,in, I_ domain_score is the domain analysis score of infrared domain detail features.

[0092] The domain analysis network (DomainAnalysis) is specifically used to evaluate the domain-specific strength of features. It is a key component for processing detailed feature fusion. Its structural design takes into account the unique properties of domain features. The specific composition of the domain analysis network is as follows: ,in, It is the feature obtained by adaptive average pooling, which captures the global statistical information of the entire feature map; ,in, It is the intermediate feature layer, Conv is a 1×1 convolutional layer, and LeakyReLU is an activation function with a small slope (the slope is usually 0.2). Unlike the feature analysis network, LeakyReLU is used here instead of standard ReLU to retain negative information, which is crucial for evaluating domain differences; ,in, is the domain analysis score. Tanh is the hyperbolic tangent activation function with an output range between -1 and 1. Different from the feature analysis network using the Sigmoid function, the Tanh function is used here because domain analysis needs to distinguish the features of different domains (positive and negative values), rather than just evaluating the importance (0 to 1). Through this design, the domain analysis network can learn the feature differences between different domains and generate domain analysis with positive and negative distinctions.

[0093] (202) Similarly, , V_domain_score is the domain analysis score of the visible light domain detail features.

[0094] (203) Calculate feature intensity differences and domain differences: , ,in, I_intensity and V_intensityare the intensities of detail features in the infrared and visible light domains, respectively; , , where intensity_diff is the intensity difference and domain_diff is the domain difference. These differences are used to guide feature selection. dim=1 specifies that the mean is calculated on the channel dimension. keepdim=True means that the calculated dimension shape is retained (the dimension becomes 1 but is not removed). This ensures that the calculated tensor dimension structure matches the original feature, facilitating subsequent feature intensity difference calculation and mask creation.

[0095] (204) Then create the selection mask: , where mask is the selection mask, It is a Sigmoid function. The area with mask value close to 1 selects infrared features, and the area close to 0 selects visible light features.

[0096] (205) Calculate the mask-based fusion detail features: , where fused_feature is the preliminary fused detail feature, which adaptively selects the most significant information based on feature strength and domain specificity.

[0097] (206) Then add cross-domain interactions to enhance feature expression: , Among them, cross_features enhances the interaction between the two domain features and adds additional information to the fusion features.

[0098] (207) The fused features are subjected to feature recalibration and domain batch normalization in turn: , ,These operations make the fused features more suitable for subsequent decoding processing while maintaining domain adaptability.

[0099] Step 4: Decode and reconstruct.

[0100] like Figure 6 As shown, decoding and reconstruction is the last link in the processing flow, which is responsible for converting the fused features into the final output image. It receives the output of adaptive fusion and reconstructs a high-quality fused image through a series of transformation operations.

[0101] The input of the decoder (Restormer_Decoder) is the fused basic features and fused detail features generated by adaptive fusion: ,in, is the output image after fusion, i.e. data_Fuse, is the intermediate feature representation, i.e. feature_F, is an optional reference image, i.e. img_VI , and They are the basic features and detail features after fusion, and Decoder represents the decoder. Figure 7 , Figure 8 and Fig. 9 As shown, they are infrared image, visible light image and fused output image respectively.

[0102] Step 401: The decoder first merges these two types of features to generate a unified feature: , where concat represents concatenating features along the channel dimension, and ReduceChannel is a 1×1 convolutional layer that reduces the number of channels by half to reduce computational complexity.

[0103] Step 402: Feature enhancement and processing.

[0104] To enhance the feature representation capability, the decoder uses a channel attention mechanism to optimize the feature channel weights: Among them, ChannelAttention enhances the expression of key feature channels by learning the importance of different channels.

[0105] The enhanced unified features are further processed through multiple transformer blocks: , where Encoder Level 2 consists of multiple Transformer blocks, each of which contains a self-attention mechanism and a feedforward network, which can capture the long-distance dependencies of features.

[0106] Step 403: local attention and output generation.

[0107] To further enhance the expression of important areas, the decoder introduces a local attention mechanism to extract local features: , where LocalAttention focuses on the spatial importance of feature maps, enabling the model to focus on areas with high information content.

[0108] Among them, the local attention mechanism The calculation process is: , ,in, is the spatial attention weight, is the gating parameter, and the spatial importance of the feature map is calculated through 7×7 convolution.

[0109] Step 404: Finally, the final fused image is generated through the output layer: .

[0110] Step 405: If a reference image exists, a residual connection is used: ; Otherwise, directly output the processing results: ;Finally, the output range is limited by the Sigmoid activation function: .

[0111] The decoder forms a complete processing chain with the previous encoder, feature decomposition and adaptive fusion, and achieves end-to-end optimization through parameter sharing and gradient back propagation. During the multi-stage training process, the decoder learns how to reconstruct single-domain images and how to generate high-quality fused images to adapt to different processing requirements.

[0112] Through this design, the decoder can effectively convert the fused features into a fused image with high visual quality and rich information, providing a complete technical solution for multimodal image fusion.

[0113] In this embodiment, the encoder, adaptive fusion and decoder form a complete multimodal image fusion model, and the multimodal image fusion model adopts a multi-stage training strategy.

[0114] The multi-stage training strategy adopted in this embodiment includes a self-reconstruction training stage and a fusion training stage, which enables the multimodal image fusion model to gradually master the ability from single-domain feature expression to multi-domain feature fusion, thereby improving training efficiency and model performance.

[0115] (1) Self-reconstruction training phase.

[0116] Self-reconstruction training is the first stage of multi-stage training. Its main purpose is to enhance the expression ability of the multimodal image fusion model for the features of each domain and to establish domain constraints.

[0117] At this stage, the multimodal image fusion model first clears the optimizer gradient and then extracts basic features and detail features from visible and infrared domain images: , ,in, and They are the basic characteristics of the visible light domain and the infrared domain, and is the corresponding detail feature.

[0118] Applying enhanced domain constraint loss guides basic features to be consistent while detail features remain different: ,in, is the total domain-constrained loss, is the domain-constrained loss of the base features, is the domain-constrained loss of detail features, and are the weight coefficients of the basic features and detail features constraints respectively. The enhanced domain constraint loss enhancedDomainConstraints is one of the key innovations of the present invention. The domain adaptability constraint is realized by calculating the statistical differences of the basic features and detail features respectively, which specifically includes: (a) For the basic features, calculate the difference between the means of the two domain features: , , ,in, c , i and j Represent the channel index, height index and width index of the feature tensor respectively; Represents the pixel value of the basic feature in the visible light domain in the cth channel, the ith row and the jth column; Represents the pixel value of the infrared domain basic feature in the cth channel, i-th row and j-th column; and They represent the channel means of the basic features in the visible light domain and the infrared domain respectively. This mean difference loss encourages the basic features to maintain statistical consistency between different domains and helps to capture common structural information.

[0119] (b) For detail features, calculate the cosine similarity of the normalized feature direction: , , , , , By minimizing the absolute value of the cosine similarity, this loss encourages the directions of detail features to remain orthogonal, thereby preserving domain-specific information.

[0120] This design enables the model to achieve domain consistency in basic features while maintaining domain specificity in detailed features, which is very suitable for multimodal image fusion tasks.

[0121] (c) Reconstruct the original image using the extracted features and calculate the reconstruction loss: , , .in, represents a visible light image reconstructed by a decoder using a visible light image, basic features in the visible light domain, and detail features; represents an infrared image reconstructed by a decoder using an infrared image, infrared domain basic features, and detail features; represents the reconstruction loss, For the reconstruction loss weight parameter, the default value is 1.

[0122] (d) In addition to the reconstruction loss, this embodiment also introduces multiple loss functions for joint optimization, including SSIM loss, MSE loss, and feature stability loss: , ;in, and are all weight parameters; feature stability loss is another innovation, which evaluates the domain consistency and specificity of features by calculating the cosine similarity of the feature channel means: , , ,in, and are all weight parameters. Feature stability loss Design: For basic features :We hope that the basic features are similar between different domains (the loss is low when the similarity is high); for the detailed features :We hope that the detailed features will remain different between different domains (low loss when the similarity is low); through this loss design, the model can pursue inter-domain consistency at the basic feature level and maintain inter-domain differences at the detailed feature level.

[0123] (e) Calculate the total loss and perform backpropagation and parameter update: .

[0124] (2) Fusion training phase.

[0125] Fusion training is the second stage of multi-stage training. Its main purpose is to optimize the feature fusion capability of the multimodal image fusion model. In this stage, the multimodal image fusion model re-extracts features, creates a fusion domain label (value is 0.5), and then adaptively fuses features from different domains: , , .

[0126] Specific loss functions are introduced in the fusion training phase, the first of which is the channel correlation loss: , , , where cc represents the channel correlation loss calculation function. This innovative correlation loss design has two purposes: on the one hand, by using a low denominator (small cc_B value) penalizes low relevance of basic features, while high molecular weight (large cc_D The denominator is added with 1.01 for numerical stability. This design further strengthens the domain consistency of basic features and the domain specificity of detailed features, which complements the domain constraint loss.

[0127] Another key loss is the fusion quality loss: ,The fusion quality loss function,criteria_fusion,is a comprehensive evaluation indicator.,This multi-faceted evaluation ensures that the fused image not only retains the key information of the,source image, but also has good visual effects.

[0128] The total loss during the fusion training phase is: .

[0129] (3) Training process monitoring and parameter optimization.

[0130] In order to monitor the changes in domain consistency during training, this embodiment introduces a domain distance calculation function to quantify the distance between domains by calculating the difference between feature means and covariances: , .

[0131] Among them, the domain distance calculation function computeDomainDistance quantifies the inter-domain distance by calculating the statistical differences between different domain features, which can be expressed as: , and It is the input of the domain distance calculation function, and mean_dist and cov_dist are the outputs of the domain distance calculation function.

[0132] Calculation of mean distance: First, we need to obtain the channel mean of the feature: ,in, H and W are the height and width of the feature map, respectively. Representation feature map All channel values ​​in row i and column j; similarly, calculate the channel mean of the second feature: ,in, Representation feature map All channel values ​​in row i and column j; then calculate the mean square error between the means as the mean distance: .

[0133] The calculation of covariance distance first obtains the channel covariance of the feature: Among them, reshape( B , C ,-1) is to reshape the tensor into ( B , C ,-1) shape, B is the batch size, C is the number of channels, and the var() function calculates the variance along the second dimension (channel dimension). The purpose is to evaluate the degree of variation of the feature in the channel dimension. A lower variance means that the feature varies less across channels, and a higher variance means that the feature varies more across channels.

[0134] Gradient clipping is used in the training process to prevent gradient explosion, and the learning rate scheduler is used to gradually reduce the learning rate, and the multimodal image fusion model parameters are saved regularly; the initial learning rate is set to 10 -4, reducing the learning rate to 50% of the original rate every 20 training steps; this multi-stage training strategy enables the model to first master the single-domain feature expression capabilities, and then learn effective feature fusion strategies, thereby achieving high-quality multimodal image fusion. Through a carefully designed loss function system, the multimodal image fusion model can learn ideal feature representation and fusion rules, improving the quality and information integrity of the fused image.

[0135] To address the inconsistency and difference issues of different modal features, this embodiment introduces a domain adaptive decomposition mechanism to divide the features into two parts: basic features and detailed features, and applies different processing strategies to each part, so that the model can better adapt to the characteristics of different modal features.

[0136] To address the balance between the consistency of the basic feature domain and the specificity of the detailed feature domain, this embodiment designs different domain constraint strategies to make the basic features tend to be consistent while the detailed features remain different, thereby ensuring the stability of the structure and retaining the unique information of each mode during the fusion process.

[0137] To address the problem of cross-domain feature statistical distribution adaptation, this embodiment designs a domain-specific batch normalization mechanism to dynamically adjust normalization parameters according to feature types and domain labels, so that the network can adaptively adjust statistical characteristics according to different domains and feature types, thereby enhancing its adaptability to multimodal data.

[0138] In order to solve the problem of information preservation and enhancement in the fusion process, this embodiment designs adaptive fusion, adopts different fusion strategies for basic features and detail features, and introduces attention mechanism and cross-domain interaction to enhance the expression of key information, suppress redundant information, and improve the fusion effect.

[0139] In response to the problems of training strategy and objective function design, this embodiment proposes a multi-stage training method. First, the expressive power of each domain feature is enhanced through self-reconstruction training, and then the cross-domain feature fusion capability is optimized through fusion training. At the same time, a composite loss function including reconstruction loss, domain constraint loss, structural similarity loss, etc. is designed to comprehensively guide the optimization of the model.

[0140] Embodiment 2 The purpose of the second embodiment is to provide a multimodal image fusion system, including: An image acquisition module, which is configured to: acquire a visible light image and an infrared image; The encoding module is configured as follows: for a visible light image and an infrared image, basic feature extraction and detail feature extraction are performed respectively through an encoder; the basic feature extraction first performs self-attention operation and channel-local composite attention operation on the image features respectively, and after fusing the two attention outputs through a gating mechanism, the basic features are obtained by sequentially processing through a feedforward network and performing domain-specific batch normalization; the detail feature extraction first divides the image features into two branches, and after applying a detail node operation to each branch, sequentially performs a channel dimension splicing operation, a channel-local composite attention operation, an attention weight fusion operation and a domain-specific batch normalization to obtain detail features; wherein the domain-specific batch normalization uses a specific factor to adjust a scaling parameter and an offset parameter, and the specific factor used for the detail feature extraction is greater than the specific factor used for the basic feature extraction; The adaptive fusion module is configured to: adaptively fuse the basic features of the visible light image and the infrared image to obtain the fused basic features; adaptively fuse the detail features of the visible light image and the infrared image to obtain the fused detail features; The decoding module is configured to: use the visible light image as a reference image, combine the fusion basic features and the fusion detail features, and obtain a fused output image through a decoder.

[0141] It should be noted here that each module in this embodiment corresponds to each step in Example 1 one by one, and the specific implementation process is the same, which will not be repeated here.

[0142] Embodiment 3 This embodiment provides a computer-readable storage medium on which a computer program is stored. The program is executed by a processor. When the program is executed by the processor, the steps in the multimodal image fusion method described in the above embodiment 1 are implemented.

[0143] Embodiment 4 This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the program, the steps in the multimodal image fusion method described in the above embodiment 1 are implemented.

[0144] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

[0145] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.

Claims

1. A multimodal image fusion method, characterized in that: include: Acquire visible light images and infrared images; For visible light images and infrared images, basic feature extraction and detail feature extraction are performed respectively through the encoder; the basic feature extraction first performs self-attention operation and channel-local composite attention operation on the image features respectively, and after the two attention outputs are fused through the gating mechanism, they are processed by the feedforward network and domain-specific batch normalization in sequence to obtain the basic features; the detail feature extraction first divides the image features into two branches, and after applying the detail node operation to each branch, the channel dimension splicing operation, the channel-local composite attention operation, the attention weight fusion operation and the domain-specific batch normalization are performed in sequence to obtain the detail features; wherein the domain-specific batch normalization uses a specific factor to adjust the scaling parameter and the offset parameter, and the specific factor used for the detail feature extraction is greater than the specific factor used for the basic feature extraction; Adaptively fuse the basic features of the visible light image and the infrared image to obtain fused basic features; adaptively fuse the detail features of the visible light image and the infrared image to obtain fused detail features; The visible light image is used as the reference image, combined with the fusion basic features and the fusion detail features, and passed through the decoder to obtain the fused output image.

2. A multimodal image fusion method as claimed in claim 1, characterized in that: The step of adaptively fusing the basic features of the visible light image and the infrared image includes: applying a feature analysis network to the basic features of the visible light image and the basic features of the infrared image, calculating the channel attention weights, and performing smoothing to obtain the smoothed channel attention weights; calculating the spatial attention weights and cross-domain interaction features based on the basic features of the visible light image and the basic features of the infrared image; calculating the attention weight fusion features based on the channel attention weights after smoothing of the visible light image and the infrared image; after fusing the attention weight fusion features, the spatial attention weights and the cross-domain interaction features, performing feature recalibration and domain batch normalization in sequence.

3. A multimodal image fusion method as claimed in claim 1, characterized in that: The step of adaptively fusing the detail features of the visible light image and the infrared image comprises: applying a domain analysis network to the detail features of the visible light image and the detail features of the infrared image, respectively, to calculate a domain analysis score; calculating an intensity difference based on the detail features of the visible light image and the detail features of the infrared image; calculating a domain difference based on the domain analysis score of the visible light image and the domain analysis score of the infrared image; creating a selection mask based on the intensity difference and the domain difference; and after calculating a preliminary fused detail feature based on the selection mask, the detail features of the visible light image and the detail features of the infrared image, sequentially performing cross-domain interactive addition, feature recalibration and domain batch normalization processing.

4. The multimodal image fusion method according to claim 1, characterized in that: The decoder includes: merging the fused basic features and the fused detail features to generate a unified feature, and using a channel attention mechanism to enhance the features. After the enhanced unified features are processed by multiple transformer blocks, a local attention mechanism is introduced to extract local features. After the local features are processed by the output layer, a residual connection is performed with the reference image, and a fused output image is obtained through an activation function.

5. The multimodal image fusion method according to claim 1, characterized in that: The encoder, adaptive fusion and decoder form a multimodal image fusion model, and the multimodal image fusion model adopts a multi-stage training strategy.

6. A multimodal image fusion method as claimed in claim 5, characterized in that: The multi-stage training strategy includes a self-reconstruction training stage, which uses reconstruction loss, domain constraint loss, SSIM loss, MSE loss and feature stability loss.

7. A multimodal image fusion method as claimed in claim 5, characterized in that: The multi-stage training strategy includes a fusion training stage, which adopts channel correlation loss and fusion quality loss.

8. A multimodal image fusion system, characterized in that: include: An image acquisition module, which is configured to: acquire a visible light image and an infrared image; The encoding module is configured as follows: for a visible light image and an infrared image, basic feature extraction and detail feature extraction are performed respectively through an encoder; the basic feature extraction first performs self-attention operation and channel-local composite attention operation on the image features respectively, and after fusing the two attention outputs through a gating mechanism, the basic features are obtained by sequentially processing through a feedforward network and performing domain-specific batch normalization; the detail feature extraction first divides the image features into two branches, and after applying a detail node operation to each branch, sequentially performs a channel dimension splicing operation, a channel-local composite attention operation, an attention weight fusion operation and a domain-specific batch normalization to obtain detail features; wherein the domain-specific batch normalization uses a specific factor to adjust a scaling parameter and an offset parameter, and the specific factor used for the detail feature extraction is greater than the specific factor used for the basic feature extraction; The adaptive fusion module is configured to: adaptively fuse the basic features of the visible light image and the infrared image to obtain the fused basic features; adaptively fuse the detail features of the visible light image and the infrared image to obtain the fused detail features; The decoding module is configured to: use the visible light image as a reference image, combine the fusion basic features and the fusion detail features, and obtain a fused output image through a decoder.

9. A computer-readable storage medium having a computer program stored thereon, the program being executed by a processor, characterized in that: When the program is executed by a processor, the steps in a multimodal image fusion method as described in any one of claims 1 to 7 are implemented.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the multimodal image fusion method as described in any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Visible light-infrared pedestrian re-identification method and system

    CN113887353A

  • Cross-domain face generation method based on multi-stage attention correlation learning

    CN116798102A

  • Infrared and visible light image fusion method and system based on relevant attention guidance

    CN116912649A

  • Face recognition method and system based on bimodal fusion

    CN117542097A

  • Regional security system and method fused with target detection

    CN118865564A

Cited By

  • Multi-modal image defogging method and device, electronic equipment and storage medium

    CN121032837A

  • A multi-modal image defogging method and device, electronic equipment and storage medium

    CN121032837B