Multimodal Image Fusion Method and Model Based on Global and Local Text Perception

By combining the global and local text perception modules of CLIP and BLIP visual language models, the problem of global and local detail processing in the fusion of infrared images and visible light images is solved, and efficient image fusion in complex scenes and inclement weather conditions is achieved, improving the adaptability and generalization performance of the model.

CN119809954BActive Publication Date: 2025-07-08FOSHAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510298599.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-08
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively process global and local detailed information in complex scenes in the fusion of infrared images and visible light images, and lacks unified weight processing capabilities under severe weather conditions, resulting in waste of computing resources and degradation of generalization performance.

Method used

The multimodal image fusion model based on global and local text perception is adopted, combined with CLIP and BLIP visual language models, the overall scene understanding is enhanced through the global text perception module, the local text perception module improves the local detail processing capability, and uses global and local text features for feature extraction and fusion.

Benefits of technology

The adaptability and generalization performance of the model in complex scenes and inclement weather conditions is improved, high-quality fusion of images is achieved, and the processing ability of global and local information is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119809954B_ABST
    Figure CN119809954B_ABST
Patent Text Reader

Abstract

This application belongs to the field of image processing technology, and discloses a multi-modal image fusion method and model based on global and local text perception. By combining two vision-language models, CLIP and BLIP, to process the global and local information of images respectively, effective fusion of images under complex scenarios and adverse weather conditions is achieved. The global text perception module enhances the model's understanding of the overall scene using CLIP features, while the local text perception module improves the ability to process local details using BLIP features. This dual text perception mechanism enables the model to more comprehensively utilize the advantages of vision-language models, avoiding problems such as relying solely on simple text prompts or overemphasizing local details. It can make full use of the advantages of vision-language models, while taking into account both global and local information processing, improving the adaptability and generalization performance of the model under complex scenarios and adverse weather conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology. Specifically, it relates to a multi-modal image fusion method and model based on global and local text perception. Background Art

[0002] The purpose of fusing infrared images and visible light images is to utilize information from different sensors or data sources to generate more expressive, accurate, and robust image representations. This technology is widely used in tasks such as object detection, semantic segmentation, and depth estimation.

[0003] In recent years, with the development of large language models, it has been possible to efficiently promote the establishment of visual models through the guidance of text semantic information. Text usually has a clear semantic structure and can reveal important information in the image with a clear description. Introducing text descriptions in the establishment of visual models helps the model associate with the object information and states in the scene and improves the model's information capture ability in image processing.

[0004] Currently, some studies have explored introducing visual language models into the field of image fusion to improve their performance in practical applications. However, there are still some unsolved problems. In the existing technology, some methods only use simple text as a degradation hint for embedding, which fails to fully utilize the advantages of visual language models. In complex scenes, the degradation situation involves not only global information but also numerous local details, such as the color change, motion state, and blur degree of objects. These details are usually not fully reflected in simplified text descriptions (such as captions). Therefore, only regarding the visual language model as a control condition for adjusting the weights of image features may lead to a waste of computing resources.

[0005] Some other existing technologies use multiple task labels in the source image to generate detailed scene description text embeddings. This method may cause the model to overemphasize intricate local details, thus reducing the model's ability to adapt to global patterns and affecting the generalization performance. During the training process, embedding rich text information can greatly enhance the influence of text on feature extraction. However, due to the complexity of natural images, the detailed description of individual images often cannot well generalize other images.

[0006] In addition, the current multi-modal image fusion models based on visual language models are difficult to adapt to bad weather. Bad weather will bring various attenuations, affect visible light images, and cause a significant decrease in the contrast of infrared images. The existing multi-modal image fusion models based on visual language models lack the ability to simultaneously handle multiple extreme degradation situations. It is still a challenge to achieve effective processing with unified weights under various bad weather conditions, especially in practical applications.

[0007] In view of the above problems, the existing technologies urgently need to be improved. Summary of the Invention

[0008] The purpose of this application is to provide a multi-modal image fusion method and model based on global and local text perception, which can make full use of the advantages of vision-language models, take into account both global and local information processing, and improve the adaptability and generalization performance of the model in complex scenarios and adverse weather conditions.

[0009] In a first aspect, this application provides a multi-modal image fusion model based on global and local text perception, including:

[0010] An input layer, a global text perception module, a CLIP image encoder, a CLIP text encoder, a BLIP text encoder, a feature extraction network, a decoder, and an output layer;

[0011] The input layer is used to input a source image and input the source image into the global text perception module, the CLIP image encoder, the CLIP text encoder, and the BLIP text encoder respectively; the source image includes a registered infrared source image and a visible light source image;

[0012] The CLIP image encoder is used to generate CLIP image features of the source image and input them into the global text perception module and the feature extraction network respectively; the CLIP text encoder is used to generate CLIP text features of the source image and input them into the global text perception module; the BLIP text encoder is used to generate BLIP text features of the source image and input them into the feature extraction network;

[0013] The global text perception module is used to fuse the source image and integrate the CLIP image features into the fusion result according to the CLIP text features to obtain a preliminary fusion result, so as to enrich the global information of the preliminary fusion result, and input the preliminary fusion result into the feature extraction network;

[0014] The feature extraction network is embedded with at least one local text perception module. The feature extraction network is used to extract features from the preliminary fusion result, and integrate the CLIP image features into the feature extraction result according to the BLIP text features during the feature extraction process, so as to enrich the local information of the feature extraction result, and input the feature extraction result into the decoder;

[0015] The decoder is used to decode the feature extraction result to generate a final fusion image, and output the final fusion image through the output layer.

[0016] By combining two vision - language models, CLIP and BLIP, to process the global and local information of images respectively, this model achieves effective fusion of images in complex scenarios and under adverse weather conditions. The global text perception module enhances the model's understanding of the overall scene using CLIP features, while the local text perception module improves the ability to process local details using BLIP features. This dual - text perception mechanism enables the model to make more comprehensive use of the advantages of vision - language models, avoiding problems such as relying solely on simple text prompts or over - emphasizing local details. It can fully utilize the advantages of vision - language models, taking into account both global and local information processing, and improving the adaptability and generalization performance of the model in complex scenarios and under adverse weather conditions.

[0017] Preferably, the global text perception module includes a first convolutional layer, a first splitting layer, a first max - pooling layer, a first residual module, a first cross - attention module, a second convolutional layer, a second splitting layer, a second max - pooling layer, a second residual module, and a second cross - attention module;

[0018] The first convolutional layer is used to input the visible light source image and output a visible light feature map to the first splitting layer. The first splitting layer is used to divide the visible light feature map into two parts. The second convolutional layer is used to input the infrared source image and output an infrared feature map to the second splitting layer. The second splitting layer is used to divide the infrared feature map into two parts. After the corresponding parts of the visible light feature map and the infrared feature map are subjected to max - pooling processing by the first max - pooling layer and the second max - pooling layer respectively, they are concatenated along the channel dimension and input into the first residual module. The corresponding other parts of the visible light feature map and the infrared feature map are concatenated along the channel dimension and then input into the second residual module;

[0019] The output feature of the first residual module is added to the CLIP image feature and then input into the first cross - attention module. The output feature of the second residual module is added to the CLIP image feature and then input into the second cross - attention module;

[0020] The first cross - attention module uses the CLIP text feature as the query vector and the image feature input into the first cross - attention module as the key vector and value vector to perform cross - attention calculation to obtain a first fusion feature. The second cross - attention module uses the CLIP text feature as the query vector and the image feature input into the second cross - attention module as the key vector and value vector to perform cross - attention calculation to obtain a second fusion feature. The first fusion feature and the second fusion feature are added together to obtain the preliminary fusion result.

[0021] This global text perception module realizes the effective fusion of the source image through the cooperation of multiple sub-modules, and integrates the CLIP text features and CLIP image features into the fusion result. This design makes full use of text information to guide the fusion of image features through multi-path parallel processing and text-guided attention mechanism, improving the quality and semantic relevance of the fusion result.

[0022] Preferably, both the first residual module and the second residual module include a plurality of first extraction layers and a second extraction layer. The plurality of first extraction layers are connected in series in sequence. The input end of the first first extraction layer is connected to the input end of the corresponding residual module. The output features of the last first extraction layer are added to the input of the corresponding residual module and then input into the second extraction layer. The output end of the second extraction layer is connected to the output end of the corresponding residual module.

[0023] The first extraction layer includes a third convolutional layer and a first ReLU activation function layer, and the second extraction layer includes a fourth convolutional layer and a second ReLU activation function layer.

[0024] This design can effectively improve the feature extraction ability while maintaining the trainability of the network. The cascading of multiple first extraction layers increases the network depth, and the residual connection helps the effective transmission of information. The second extraction layer, as the last layer, can make final adjustments to the features. This structural design improves the expression ability and performance of the model while maintaining computational efficiency.

[0025] Preferably, the feature extraction network includes multiple residual spatial state modules. Among them, one local text perception module is embedded after every N layers of the residual spatial state modules, where N is a positive integer. The CLIP image encoder inputs the CLIP image features into each local text perception module. The BLIP text encoder inputs the BLIP text features into each local text perception module.

[0026] Preferably, the residual spatial state module includes a first regularization layer, a visual state space module, a first scaling layer, a second regularization layer, a seventh convolutional layer, a channel attention module, and a second scaling layer.

[0027] The first regularization layer and the visual state space module are connected in sequence. The second regularization layer, the seventh convolutional layer, and the channel attention module are connected in sequence. The input ends of the first regularization layer and the first scaling layer are both connected to the input end of the residual spatial state module. The output of the visual state space module and the output of the first scaling layer are added and then input into the second regularization layer and the second scaling layer. The output of the channel attention module and the output of the second scaling layer are added and then output from the output end of the residual spatial state module.

[0028] Preferably, the visual state space module includes a first linear layer, a depthwise separable convolutional layer, a first Swish activation function layer, a 2D selective scanning module, a layer normalization module, a second linear layer, a second Swish activation function layer, and a third linear layer;

[0029] The first linear layer, the depthwise separable convolutional layer, the first Swish activation function layer, the 2D selective scanning module, and the layer normalization module are connected in sequence, and the second linear layer and the second Swish activation function layer are connected in sequence; the input ends of the first linear layer and the second linear layer are both connected to the input end of the visual state space module, the output of the layer normalization module and the output of the second Swish activation function layer are multiplied and then input into the third linear layer, and the output end of the third linear layer is connected to the output end of the visual state space module.

[0030] Preferably, the local text perception module includes several squeeze-and-excitation modules, an embedding module, a third cross-attention module, several dilated convolutional layers, and a fifth convolutional layer;

[0031] Several squeeze-and-excitation modules, the embedding module, and the third cross-attention module are connected in sequence, and the input end of the first squeeze-and-excitation module in the first layer is connected to the input end of the local text perception module;

[0032] The squeeze-and-excitation module is used to enhance important features and suppress redundant information of the image features input into the squeeze-and-excitation module through an adaptive channel attention mechanism;

[0033] The embedding module is used to map the CLIP image features into the semantic space of the image features input into the embedding module through a multi-layer perceptron and then embed the image features input into the embedding module;

[0034] The third cross-attention module uses the BLIP text features as query vectors and the image features input into the third cross-attention module as key vectors and value vectors to perform cross-attention calculations, and inputs the calculation results into each dilated convolutional layer;

[0035] The dilation rates of each dilated convolutional layer are different, and the outputs of all dilated convolutional layers are connected in channels and then input into the fifth convolutional layer; the output end of the fifth convolutional layer is connected to the output end of the local text perception module.

[0036] In a second aspect, the present application provides a multi-modal image fusion method based on global and local text perception, including the steps:

[0037] A1. Construct an initial multi-modal image fusion model based on global and local text perception; the initial multi-modal image fusion model based on global and local text perception is the multi-modal image fusion model based on global and local text perception described above;

[0038] A2. Train the initial multi-modal image fusion model based on global and local text perception to obtain a trained multi-modal image fusion model based on global and local text perception;

[0039] A3. Input the source images to be fused into the trained multi-modal image fusion model based on global and local text perception to obtain a fused image output by the trained multi-modal image fusion model based on global and local text perception.

[0040] Preferably, step A2 includes:

[0041] A201. Obtain a plurality of training samples; the training samples include sample source images and corresponding clean source images; the sample source images include sample infrared source images and sample visible light source images, the clean source images include clean infrared source images and clean visible light source images, the sample source images contain degradation information, the clean source images do not contain degradation information, and the degradation information includes weather interference information;

[0042] A202. Select the training samples, input the corresponding sample source images into the initial multi-modal image fusion model based on global and local text perception to obtain a fused image output by the initial multi-modal image fusion model based on global and local text perception, denoted as the first fused image;

[0043] A203. Input the first fused image into the CLIP image encoder to obtain CLIP image features corresponding to the first fused image, denoted as the first CLIP image features;

[0044] A204. Input the clean source images of the selected training samples into the CLIP text encoder to obtain CLIP text features corresponding to the clean source images, denoted as the first CLIP text features;

[0045] A205. Align the first CLIP text features and the first CLIP image features in the same feature space;

[0046] A206. Calculate a loss function according to the first fused image, the aligned first CLIP text features, and the clean source images of the selected training samples;

[0047] Optimize the model parameters of the initial multi-modal image fusion model based on global and local text perception according to the loss function;

[0048] A208. Repeat steps A202 - A207 until the loss function is less than a preset loss function threshold or the number of iterations reaches a preset number threshold.

[0049] Preferably, the loss function is:

[0050] ;

[0051] ;

[0052] ;

[0053] where is the loss function value, is the pixel-level loss, is the driving loss, is the first fused image, is the aligned first CLIP text feature, represents the norm, is the clean infrared source image of the selected training sample, is the clean visible light source image of the selected training sample, represents the 1-norm.

[0054] Beneficial effects: The multi-modal image fusion method and model based on global and local text perception provided by this application, by combining two vision-language models, CLIP and BLIP, to process the global and local information of images respectively, realizes the effective fusion of images under complex scenarios and adverse weather conditions. The global text perception module uses CLIP features to enhance the model's understanding of the overall scene, while the local text perception module uses BLIP features to improve the processing ability of local details. This dual text perception mechanism enables the model to more comprehensively utilize the advantages of vision-language models, avoiding the problems of relying only on simple text prompts or overemphasizing local details. It can make full use of the advantages of vision-language models, while taking into account both global and local information processing, and improving the adaptability and generalization performance of the model under complex scenarios and adverse weather conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 is a schematic structural diagram of the multi-modal image fusion model based on global and local text perception provided by the embodiment of this application.

[0056] Figure 2 is a schematic structural diagram of the global text perception module.

[0057] Figure 3 It is a schematic structural diagram of the first residual module and the second residual module.

[0058] Figure 4 It is a schematic structural diagram of the feature extraction network.

[0059] Figure 5 It is a schematic structural diagram of the residual space state module.

[0060] Figure 6 It is a schematic structural diagram of the visual state space module.

[0061] Figure 7 It is a schematic structural diagram of the local text perception module.

[0062] Figure 8 It is a schematic structural diagram of the decoder.

[0063] Figure 9 It is a flowchart of the multi-modal image fusion method based on global and local text perception provided by the embodiments of the present application.

[0064] Figure 10 It is a comparison chart of the first group of experimental results.

[0065] Figure 11 It is a comparison chart of the second group of experimental results.

[0066] Figure 12 It is a comparison chart of the third group of experimental results.

[0067] Figure 13 It is a comparison chart of the fourth group of experimental results.

[0068] Figure 14 It is a comparison chart of the fifth group of experimental results.

[0069] Figure 15 It is a comparison chart of the sixth group of experimental results.

[0070] Label description: 1. Input layer;

[0071] 2. Global text perception module; 201. First convolutional layer; 202. First separation layer; 203. First max pooling layer; 204. First residual module; 205. First cross-attention module; 206. Second convolutional layer; 207. Second separation layer; 208. Second max pooling layer; 209. Second residual module; 210. Second cross-attention module;

[0072] 3. CLIP image encoder; 4. CLIP text encoder; 5. BLIP text encoder; 6. Feature extraction network;

[0073] 7. Decoder; 701. Decoding module; 702. Wavelet convolutional layer; 703. Third ReLU activation function layer; 704. Sixth convolutional layer;

[0074] 8. Output layer;

[0075] 9. Local text perception module; 901. Squeeze-and-excitation module; 902. Embedding module; 903. Third cross-attention module; 904. Dilated convolutional layer; 905. Fifth convolutional layer; 906. SE attention module;

[0076] 10. First extraction layer; 1001. Third convolutional layer; 1002. First ReLU activation function layer;

[0077] 11. Second extraction layer; 1101. Fourth convolutional layer; 1102. Second ReLU activation function layer;

[0078] 12. Residual spatial state module; 1201. First regularization layer; 1202. Visual state space module; 1203. First scaling layer; 1204. Second regularization layer; 1205. Seventh convolutional layer; 1206. Channel attention module; 1207. Second scaling layer;

[0079] 13. First linear layer; 14. Depthwise separable convolutional layer; 15. First Swish activation function layer; 16. 2D selective scanning module; 17. Layer normalization module; 18. Second linear layer; 19. Second Swish activation function layer; 20. Third linear layer. Detailed implementation manner

[0080] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Usually, the components of the embodiments of the present application described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application to be protected, but only represents the selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.

[0081] It should be noted that: Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0082] In the field of multimodal image fusion, especially in the fusion of infrared and visible light images, using vision-language models to improve the fusion effect has become an important research direction. However, existing multimodal image fusion technologies based on vision-language models still face significant challenges in processing global and local information and adapting to adverse weather conditions. Specifically, these technologies often oversimplify or overcomplicate the use of text information, resulting in the inability to fully utilize the advantages of vision-language models. Among them, simple text embedding methods are difficult to capture detailed information in complex scenes, while overly detailed scene descriptions may cause the model to over-focus on local features, affecting its generalization ability. In addition, under adverse weather conditions, existing models are difficult to handle multiple extreme degradation situations simultaneously, which severely limits their performance in practical applications.

[0083] Therefore, referring to Figures 1 - 8 , this application provides a multimodal image fusion model based on global and local text perception, including:

[0084] Input layer 1, global text perception module 2, CLIP image encoder 3 (CLIP stands for Contrastive Language-Image Pre-Training, a contrastive learning-based language-image pre-training), CLIP text encoder 4, BLIP text encoder 5 (BLIP stands for Bootstrapping Language-Image Pre-training, a bootstrapping language-image pre-training), feature extraction network 6, decoder 7, and output layer 8;

[0085] The input layer 1 is used to input the source images and input the source images into the global text perception module 2, CLIP image encoder 3, CLIP text encoder 4, and BLIP text encoder 5 respectively; the source images include registered infrared source images and visible light source images;

[0086] The CLIP image encoder 3 is used to generate the CLIP image features of the source images (O1 in the figure represents the CLIP image features) and input them into the global text perception module 2 and the feature extraction network 6 respectively; the CLIP text encoder 4 is used to generate the CLIP text features of the source images (O2 in the figure represents the CLIP text features) and input them into the global text perception module 2; the BLIP text encoder 5 is used to generate the BLIP text features of the source images (O3 in the figure represents the BLIP text features) and input them into the feature extraction network 6;

[0087] The global text perception module 2 is used to fuse the source images and integrate the CLIP image features (i.e., O1) into the fusion result according to the CLIP text features (i.e., O2) to obtain a preliminary fusion result, so as to enrich the global information of the preliminary fusion result, and input the preliminary fusion result into the feature extraction network 6;

[0088] The feature extraction network 6 is embedded with at least one local text perception module 9. The feature extraction network 6 is used to extract features from the preliminary fusion result, and integrate the CLIP image feature (i.e., O1) into the feature extraction result according to the BLIP text feature (i.e., O3) during the feature extraction process, so as to enrich the local information of the feature extraction result, and input the feature extraction result into the decoder 7;

[0089] The decoder 7 is used to decode the feature extraction result to generate the final fused image, and output the final fused image through the output layer 8.

[0090] By combining two vision-language models, CLIP and BLIP, this model processes the global and local information of images respectively, and realizes the effective fusion of images under complex scenes and bad weather conditions. The global text perception module 2 uses the CLIP feature to enhance the model's understanding of the overall scene, while the local text perception module 9 uses the BLIP feature to improve the processing ability of local details. This dual text perception mechanism enables the model to make more comprehensive use of the advantages of vision-language models, avoiding the problems of relying only on simple text prompts or overemphasizing local details. It can make full use of the advantages of vision-language models, while taking into account global and local information processing, and improve the adaptability and generalization performance of the model under complex scenes and bad weather conditions.

[0091] Therefore, the input source image can be a source image containing degradation information, where the degradation information includes at least one of weather interference information such as rain, snow, fog, etc.

[0092] Among them, the CLIP image encoder 3, the CLIP text encoder 4, and the BLIP text encoder 5 can all adopt existing technologies, and will not be elaborated here.

[0093] Specifically, see Figure 2 , the global text perception module 2 includes a first convolutional layer 201, a first separation layer 202, a first max pooling layer 203, a first residual module 204, a first cross-attention module 205, a second convolutional layer 206, a second separation layer 207, a second max pooling layer 208, a second residual module 209, and a second cross-attention module 210;

[0094] The first convolutional layer 201 is used to input the visible light source image (i.e., Iv in the figure) and output a visible light feature map to the first separation layer 202. The first separation layer 202 is used to divide the visible light feature map into two parts (such as the upper half and the lower half, or the left half and the right half); the second convolutional layer 206 is used to input the infrared source image (i.e., Ir in the figure) and output an infrared feature map to the second separation layer 207. The second separation layer 207 is used to divide the infrared feature map into two parts (correspondingly, the upper half and the lower half, or the left half and the right half); corresponding parts of the visible light feature map and the infrared feature map are respectively subjected to max pooling processing by the first max pooling layer 203 and the second max pooling layer 208 (for example, the left half of the visible light feature map is subjected to max pooling processing by the first max pooling layer 203, and the left half of the infrared feature map is subjected to max pooling processing by the second max pooling layer 208), then channel connection is performed and the result is input into the first residual module 204; the corresponding other parts of the visible light feature map and the infrared feature map (such as the right half) are subjected to channel connection and then input into the second residual module 209;

[0095] The output feature of the first residual module 204 is added to the CLIP image feature (i.e., O1) and then input into the first cross-attention module 205. The output feature of the second residual module 209 is added to the CLIP image feature (i.e., O1) and then input into the second cross-attention module 210;

[0096] The first cross-attention module 205 uses the CLIP text feature (i.e., O2) as the query vector and the image feature input into the first cross-attention module 205 as the key vector and the value vector, and performs cross-attention calculation to obtain the first fusion feature; the second cross-attention module 210 uses the CLIP text feature (i.e., O2) as the query vector and the image feature input into the second cross-attention module 210 as the key vector and the value vector, and performs cross-attention calculation to obtain the second fusion feature; the first fusion feature and the second fusion feature are added to obtain the preliminary fusion result.

[0097] This global text perception module 2 realizes the effective fusion of the source images through the cooperation of multiple sub-modules, and integrates the CLIP text feature and the CLIP image feature into the fusion result. This design makes full use of the text information to guide the fusion of image features through multi-way parallel processing and text-guided attention mechanism, improving the quality and semantic relevance of the fusion result.

[0098] Among them, the first convolutional layer 201 and the second convolutional layer 206 respectively process the visible light source image and the infrared source image to extract preliminary features. These two convolutional layers can use 1×1 convolutional kernels to expand the feature channels of the source images.

[0099] Among them, the first separation layer 202 and the second separation layer 207 divide their respective feature maps into two parts to prepare for subsequent parallel processing. The separation operation can be achieved through channel splitting or spatial splitting. Channel splitting can evenly divide the feature map into two parts according to the number of channels, while spatial splitting can be performed according to the spatial position of the feature map, such as splitting it into upper and lower or left and right parts.

[0100] Among them, the first max pooling layer 203 and the second max pooling layer 208 downsample a part of the feature map to reduce the computational amount. The pooling operation can not only reduce the spatial dimension of the feature map but also improve the robustness of the model to small image transformations.

[0101] Among them, the channel connection operation fuses visible light and infrared features. This step can be achieved through simple channel concatenation or by using weighted summation to assign weights according to the importance of different modality features.

[0102] Among them, the first residual module 204 processes the pooled features, and the second residual module 209 processes the unpooled features, which helps to retain information at different scales, enhance the representativeness of features, and minimize feature loss. The design of the residual module can include multiple convolutional layers and activation functions, and the gradient vanishing problem is alleviated through shortcut connections, improving the effect of feature extraction and the robustness of features.

[0103] For example, the first residual module 204 and the second residual module 209 can adopt Figure 3 the shown residual module, which includes several first extraction layers 10 and a second extraction layer 11. Several first extraction layers 10 are connected in series in sequence. The input end of the first first extraction layer 10 is connected to the input end of the corresponding residual module (the first residual module 204 or the second residual module 209). The output features of the last first extraction layer 10 are added to the input of the corresponding residual module and then input into the second extraction layer 11. The output end of the second extraction layer 11 is connected to the output end of the corresponding residual module;

[0104] The first extraction layer 10 includes a third convolutional layer 1001 and a first ReLU activation function layer 1002, and the second extraction layer 11 includes a fourth convolutional layer 1101 and a second ReLU activation function layer 1102.

[0105] This design can effectively improve the feature extraction ability while maintaining the trainability of the network. The concatenation of multiple first extraction layers 10 increases the network depth, and the residual connection helps with the effective transmission of information. The second extraction layer 11, as the last layer, can perform final adjustments on the features. This structural design improves the model's expressive ability and performance while maintaining computational efficiency. Among them, the number of first extraction layers 10 and the kernel sizes of the third convolutional layer 1001 and the fourth convolutional layer 1101 can be set according to actual needs. For example, the number of first extraction layers 10 is 2, the kernel size of the third convolutional layer 1001 is 3×3, and the kernel size of the fourth convolutional layer 1101 is 1×1, but it is not limited to this.

[0106] Among them, the CLIP image features are added to the output of the residual module to ensure the consistency of multimodal information.

[0107] Among them, the first cross-attention module 205 and the second cross-attention module 210 use the CLIP text features as query vectors to integrate semantic information, achieve text-guided feature fusion, and improve the model's understanding ability. The cross-attention calculation can adopt the standard scaled dot-product attention mechanism, or a multi-head attention mechanism can be introduced to capture attention information in different subspaces.

[0108] The cross-attention calculation performed by the first cross-attention module 205 can be expressed as:

[0109] ;

[0110] Among them, is the cross-attention function, is the image feature input to the first cross-attention module 205, is corresponding to the key vector, is corresponding to the value vector, is the dimension of, is the softmax function, is the query vector corresponding to O2.

[0111] Similarly, the cross-attention calculation performed by the second cross-attention module 210 can be expressed as:

[0112] ;

[0113] Among them, is the image feature input to the second cross-attention module 210, is corresponding to the key vector, is corresponding to The value vector, is the dimension of.

[0114] Through this module design and feature fusion process, the global text perception module 2 can effectively solve the technical problem of how to integrate CLIP text features and CLIP image features into the fusion result. This method not only considers the visual features of the image, but also introduces text information as the global semantic guidance, thus achieving a more intelligent and semantic-aware image fusion.

[0115] Furthermore, see Figure 4 , the feature extraction network 6 includes multiple layers of residual spatial state modules 12, where a local text perception module 9 is embedded after every N layers of residual spatial state modules 12, and N is a positive integer; the CLIP image encoder 3 inputs the CLIP image features (i.e., O1) into each local text perception module 9; the BLIP text encoder 5 inputs the BLIP text features (i.e., O3) into each local text perception module 9.

[0116] The feature extraction network 6 adopts the structure of multiple layers of residual spatial state modules 12, and this structure has good feature extraction ability. By embedding local text perception modules 9 in this structure, the present application realizes multi-scale and multi-level text guidance for image features. Specifically, a local text perception module 9 is embedded after every N layers of residual spatial state modules 12, where N is a positive integer. This layout method has flexibility, and the value of N can be adjusted according to actual needs to balance the efficiency of feature extraction and the frequency of text perception. For example, when N = 2, a local text perception module 9 is embedded after every two layers of residual spatial state modules 12. This configuration can introduce text information frequently while maintaining a high feature extraction efficiency. When N = 3 or larger, the distribution of local text perception modules 9 will be sparser, which may be suitable for scenarios with limited computing resources or high real-time requirements.

[0117] The CLIP image encoder 3 and the BLIP text encoder 5 respectively provide different types of features for the local text perception module 9. CLIP image features usually contain higher-level semantic information, while BLIP text features may contain more fine-grained text descriptions. By inputting these two types of features into each local text perception module 9 simultaneously, the present application realizes the deep fusion of image and text information. This design enables the feature extraction network 6 to capture the semantic associations between images and texts at different abstraction levels. Through this multi-level text perception mechanism, the feature extraction network 6 of the present application can better handle image fusion tasks in complex scenarios and adverse weather conditions.

[0118] In some embodiments, see Figure 5, the residual space state module 12 includes a first regularization layer 1201, a visual state space module 1202, a first scaling layer 1203 (the scaling layer is the Scale layer), a second regularization layer 1204, a seventh convolutional layer 1205, a channel attention module 1206, and a second scaling layer 1207;

[0119] The first regularization layer 1201 and the visual state space module 1202 are connected in sequence, and the second regularization layer 1204, the seventh convolutional layer 1205, and the channel attention module 1206 are connected in sequence; the input end of the first regularization layer 1201 and the input end of the first scaling layer 1203 are both connected to the input end of the residual space state module 12, and the output of the visual state space module 1202 and the output of the first scaling layer 1203 are added and then input into the second regularization layer 1204 and the second scaling layer 1207; the output of the channel attention module 1206 and the output of the second scaling layer 1207 are added and then output from the output end of the residual space state module 12.

[0120] The residual space state module 12 is ingeniously designed to combine technologies such as residual connection, regularization, attention mechanism, and multi-scale feature fusion, effectively solving the problems of feature extraction and fusion. Residual connection helps to alleviate the gradient vanishing problem, enabling the network to learn deep features more effectively. Multi-scale feature fusion can capture image information at different scales, improving the model's understanding ability of complex scenes. The introduction of the channel attention mechanism further enhances the model's attention to important features, improving the efficiency and accuracy of feature extraction.

[0121] This design can not only effectively extract and fuse image features, but also maintain good performance when dealing with complex scenes, especially when dealing with images containing multiple weather conditions. In this way, the module provides the entire multi-modal image fusion model with powerful feature extraction and fusion capabilities, laying a solid foundation for subsequent image fusion tasks.

[0122] Among them, the first regularization layer 1201 and the second regularization layer 1204 help to stabilize the network training process and improve the model's convergence speed and generalization ability by normalizing the input data.

[0123] Among them, the visual state space module 1202 is responsible for extracting and processing complex visual features. This module can adopt a variety of advanced visual feature extraction technologies, such as self-attention mechanism, convolutional neural network, or transformer structure.

[0124] For example, in some embodiments, see Figure 6, the visual state space module 1202 includes a first linear layer 13, a depthwise separable convolutional layer 14, a first Swish activation function layer 15 (the Swish activation function layer is the SiLU layer), a 2D selective scanning module 16 (i.e., the 2D-SSM module), a layer normalization module 17, a second linear layer 18, a second Swish activation function layer 19, and a third linear layer 20;

[0125] The first linear layer 13, the depthwise separable convolutional layer 14, the first Swish activation function layer 15, the 2D selective scanning module 16, and the layer normalization module 17 are connected in sequence, and the second linear layer 18 and the second Swish activation function layer 19 are connected in sequence; the input ends of the first linear layer 13 and the second linear layer 18 are both connected to the input end of the visual state space module 1202, the output of the layer normalization module 17 and the output of the second Swish activation function layer 19 are multiplied and then input into the third linear layer 20, and the output end of the third linear layer 20 is connected to the output end of the visual state space module 1202.

[0126] Through multi-level feature extraction and non-linear transformation, the visual state space module 1202 can effectively extract and integrate local and global features of images. The use of depthwise separable convolution and 2D selective scanning improves the model's perception ability of spatial information, while the multiple uses of activation functions and normalization operations enhance the model's non-linear expression ability and training stability. This structural design helps to improve the quality of feature extraction and the overall performance of the model while maintaining computational efficiency.

[0127] Among them, the first scaling layer 1203 and the second scaling layer 1207 achieve effective fusion of features at different levels by adjusting the scales of features. These scaling operations can adaptively adjust the importance of features through learnable parameters.

[0128] Among them, the seventh convolutional layer 1205 further extracts and transforms features, enhancing the model's expression ability. This layer can use a 3×3 convolutional kernel.

[0129] Among them, the channel attention module 1206 highlights important features and suppresses irrelevant information by assigning different weights to different channels. This can be achieved through techniques such as Squeeze-and-Excitation (SE) blocks or ECA (EfficientChannelAttention).

[0130] Furthermore, see Figure 7 , the local text perception module 9 includes several squeeze-and-excitation modules 901, an embedding module 902, a third cross-attention module 903, several dilated convolutional layers 904, and a fifth convolutional layer 905;

[0131] A number of squeeze-and-excitation modules 901, embedding modules 902, and third cross-attention modules 903 are connected in sequence. The input end of the first-layer squeeze-and-excitation module 901 is connected to the input end of the local text perception module 9;

[0132] The squeeze-and-excitation module 901 is used to enhance important features and suppress redundant information of the image features input into the squeeze-and-excitation module 901 through an adaptive channel attention mechanism;

[0133] The embedding module 902 is used to map the CLIP image features (i.e., O1) to the semantic space of the image features input into the embedding module 902 through a multi-layer perceptron and then embed the image features input into the embedding module 902;

[0134] The third cross-attention module 903 uses the BLIP text features (i.e., O3) as query vectors and the image features input into the third cross-attention module 903 as key vectors and value vectors to perform cross-attention calculations, and inputs the calculation results into each dilated convolutional layer 904 respectively;

[0135] The dilation rates of each dilated convolutional layer 904 are different. After the outputs of all dilated convolutional layers 904 are connected in channels, they are input into the fifth convolutional layer 905; the output end of the fifth convolutional layer 905 is connected to the output end of the local text perception module 9.

[0136] As the number of network layers in the feature extraction network 6 increases, the complexity and dimension of the feature maps will continuously increase, which makes it increasingly important to effectively focus on the feature responses of different channels. Therefore, a squeeze-and-excitation module 901 is introduced in the local text perception module 9 to process the fused features. By explicitly modeling the dependencies between channels, the squeeze-and-excitation module 901 can enhance important features and suppress redundant information through an adaptive channel attention mechanism, ensuring that the input features have good expressiveness at the local level.

[0137] As a preferred implementation manner, see Figure 7 , the squeeze-and-excitation module 901 includes an SE attention module 906. The input end of the SE attention module 906 is connected to the input end of the squeeze-and-excitation module 901. The input of the squeeze-and-excitation module 901 is multiplied by the output of the SE attention module 906 and then output from the output end of the squeeze-and-excitation module 901. This design can further enhance the module's ability to identify important features and improve the efficiency of feature extraction.

[0138] The number of layers of the squeeze-and-excitation module 901 can be set according to actual needs. For example, it can be set to two layers, but it is not limited to this.

[0139] Among them, the embedding module 902 uses a multi-layer perceptron to map the CLIP image features to the semantic space of the input image features. This step realizes the semantic alignment and fusion of features, enabling the global image information to be effectively embedded into the fused features. This design helps to enhance the semantic expression ability of the features and provides a richer information basis for subsequent processing.

[0140] Specifically, the embedding module 902 performs a mapping operation on the CLIP image features through a multi-layer perceptron to obtain a scale adjustment parameter and a bias control parameter for performing the following operations:

[0141] ;

[0142] Among them, is the output feature of the embedding module 902, is the scale adjustment parameter, is the bias control parameter, is the image feature input to the embedding module 902, represents the element-wise multiplication operation. This design allows the module to dynamically adjust the fusion process according to the characteristics of the input features, improving the flexibility and adaptability of feature fusion.

[0143] Among them, the third cross-attention module 903 uses the BLIP text features as the query vector, the input image features as the key vector and the value vector, and performs cross-attention calculation. This step realizes text-guided feature enhancement and further integrates multi-modal information from vision and text. In this way, the semantic richness of the fused features is improved, including more comprehensive scene and detail information.

[0144] The cross-attention calculation performed by the third cross-attention module 903 can be expressed as:

[0145] ;

[0146] Among them, is the image feature input to the third cross-attention module 903, is the key vector corresponding to , is the value vector corresponding to , is the dimension of, is the query vector corresponding to O3.

[0147] Among them, the calculation results of the cross-attention calculation are input into several dilated convolutional layers 904. These dilated convolutional layers 904 have different dilation rates and can capture multi-scale context information. By using dilated convolutions with different dilation rates, the module can sparsely sample the image information at a specific stride, effectively removing redundant information while retaining key details. This design significantly enhances the feature expression ability, enabling the model to better understand the local structure and global semantics of the image.

[0148] The number of dilated convolutional layers 904 and the dilation rate of each dilated convolutional layer 904 can be set according to actual needs. For example, there are three dilated convolutional layers 904, and the dilation rates of the three dilated convolutional layers 904 are 1, 2, and 3 respectively; but it is not limited to this.

[0149] Among them, the outputs of all dilated convolutional layers 904 are concatenated in channels and then input into the fifth convolutional layer 905. The fifth convolutional layer 905 finally integrates the features extracted by the previous components to generate the output features of the local text perception module 9. This step ensures that the fused features have stronger discrimination ability and robustness. The kernel size of the fifth convolutional layer 905 can be set to 1×1.

[0150] Through this design, the local text perception module 9 can achieve text-guided feature enhancement and fusion in the local area. The module makes full use of text information to enhance image features, and at the same time improves the feature expression ability and robustness through multi-scale feature extraction and adaptive attention mechanism. This method can not only capture the local detail information of the image, but also effectively integrate the text semantic information into the feature representation, thus providing a richer and more accurate feature input for the subsequent image fusion task.

[0151] Specifically, see Figure 8 , the decoder 7 includes several decoding modules 701 connected in series in sequence. The decoding module 701 includes a wavelet convolutional layer 702, a third ReLU activation function layer 703, and a sixth convolutional layer 704 connected in sequence; the input end of the first decoding module 701 is connected to the input end of the decoder 7, and the output end of the last decoding module 701 is connected to the output end of the decoder 7.

[0152] The wavelet convolutional layer 702 can effectively extract multi-scale features, the third ReLU activation function layer 703 introduces non-linearity, and the sixth convolutional layer 704 performs feature fusion. This combination can gradually restore the image details while maintaining the grasp of global information. Among them, the number of decoding modules 701 can be set according to actual needs, for example, it is 3, but it is not limited to this. Among them, the sixth convolutional layer 704 can use a 3×3 kernel.

[0153] Refer to Figure 9, this application provides a multi-modal image fusion method based on global and local text perception, including the steps:

[0154] A1. Construct an initial multi-modal image fusion model based on global and local text perception; the initial multi-modal image fusion model based on global and local text perception is the multi-modal image fusion model based on global and local text perception mentioned above.

[0155] A2. Train the initial multi-modal image fusion model based on global and local text perception to obtain a trained multi-modal image fusion model based on global and local text perception.

[0156] A3. Input the source images to be fused (including the infrared source image to be fused and the visible light source image to be fused that are mutually registered) into the trained multi-modal image fusion model based on global and local text perception to obtain the fused image output by the trained multi-modal image fusion model based on global and local text perception.

[0157] By introducing the global and local text perception mechanism, this method can make more comprehensive use of text information to guide the image fusion process. Global text perception helps to capture the overall semantic information of the image, while local text perception can focus on the detailed features of the image. This dual perception mechanism enables the fusion process to better balance global and local information, thereby generating higher-quality fused images.

[0158] Preferably, step A2 includes:

[0159] A201. Obtain multiple training samples; the training samples include sample source images and corresponding clean source images; the sample source images include sample infrared source images and sample visible light source images, the clean source images include clean infrared source images and clean visible light source images, the sample source images contain degradation information, the clean source images do not contain degradation information, and the degradation information includes weather interference information (such as at least one of rain, snow, fog, etc.).

[0160] A202. Select training samples, input the corresponding sample source images into the initial multi-modal image fusion model based on global and local text perception to obtain the fused image output by the initial multi-modal image fusion model based on global and local text perception, denoted as the first fused image.

[0161] A203. Input the first fused image into the CLIP image encoder 3 to obtain the CLIP image features corresponding to the first fused image, denoted as the first CLIP image features.

[0162] A204. Input the clean source images of the selected training samples into the CLIP text encoder 4 to obtain the CLIP text features corresponding to the clean source images, denoted as the first CLIP text features.

[0163] A205. Align the first CLIP text feature and the first CLIP image feature in the same feature space;

[0164] A206. Calculate the loss function according to the first fused image, the aligned first CLIP text feature, and the clean source image of the selected training sample;

[0165] A207. Optimize the model parameters of the initial multi-modal image fusion model based on global and local text perception according to the loss function;

[0166] A208. Repeat steps A202 - A207 until the loss function is less than the preset loss function threshold or the number of iterations reaches the preset number threshold.

[0167] It should be understood that among them, the two images in the sample source image are registered with each other, and the clean source image is the ideal image obtained by removing the degradation information from the sample source image.

[0168] Among them, in step A202, the training samples can be selected sequentially or randomly.

[0169] Among them, in step A205, the first CLIP text feature and the first CLIP image feature are aligned in the same feature space. This step ensures the comparability of the text feature and the image feature and lays a foundation for the subsequent loss calculation.

[0170] Preferably, the loss function is:

[0171] ;

[0172] ;

[0173] ;

[0174] Among them, is the loss function value, is the pixel-level loss, is the driving loss, is the first fused image, is the aligned first CLIP text feature, represents the norm, is the clean infrared source image of the selected training sample, is the clean visible light source image of the selected training sample, represents the 1-norm.

[0175] The loss function includes a pixel-level loss and a driving loss, comprehensively considering the image reconstruction quality and the consistency of the feature space, which can ensure enhancing the semantic features of the fused image and improving the model's expression ability for natural scenes.

[0176] Among them, in step A207, the model parameters of the initial multi-modal image fusion model based on global and local text perception can be optimized using the Adam optimizer according to the loss function. The initial learning efficiency can be set to 0.0001, the iteration threshold can be set to 350, and the loss function threshold can be set to 0.01, but it is not limited to this.

[0177] Figures 10 - 15 Shows the fusion results of the source images containing different degradation information in the present invention. In Figures 10 - 15 , a represents the infrared source image, b represents the visible light source image, and c represents the fused image; Figure 10 and Figure 11 the degradation information contained in the source image is fog, Figure 12 and Figure 13 the degradation information contained in the source image is rain, Figure 14 and Figure 15 the degradation information contained in the source image is snow. It can be seen from the figure that in each degradation scenario, the infrared source image generally has the problem of low contrast, which indicates that the present invention not only needs to effectively cope with degradation when processing degraded images, but also needs to enhance the image contrast to highlight the details of significant objects. At the same time, it can be observed that degradation factors such as fog, rain, and snow pollute the visible light image to varying degrees, obscuring most of the useful information. In contrast, the fusion results of the present invention show excellent performance: on the basis of effectively removing degraded pixels, it successfully realizes the deep interaction of multi-modal features, fully captures complementary information from different modalities, and significantly improves the contrast and clarity of the scene.

[0178] In summary, this application has at least the following advantages:

[0179] (1) The present invention proposes a novel multi-modal image fusion method for bad weather. This method uses global and local text perception within a unified architecture and can effectively process multiple degradations simultaneously;

[0180] (2) The present invention proposes a novel global and local text perception module, which guides feature extraction from both macroscopic and microscopic perspectives, enhancing generality while maintaining high-fidelity fusion;

[0181] (3) The present invention introduces a novel visual language model-driven loss function, which uses the feature space of CLIP to guide the generation of fused images based on text information; this loss function comprehensively evaluates the alignment of image-text features through semantic features;

[0182] (4) The present invention has been extensively experimented under different adverse weather conditions, demonstrating its excellent performance in dealing with extreme environments and its applicability to many real-world application scenarios.

[0183] In the embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be electrical, mechanical or other forms.

[0184] Additionally, the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0185] Furthermore, in each embodiment of the present application, the various functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0186] In this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.

[0187] The above description is only for the embodiments of the present application and is not intended to limit the protection scope of the present application. For those skilled in the art, the present application can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A multi-modal image fusion model based on global and local text perception, characterized in that, Including: An input layer (1), a global text perception module (2), a CLIP image encoder (3), a CLIP text encoder (4), a BLIP text encoder (5), a feature extraction network (6), a decoder (7), and an output layer (8); The input layer (1) is used to input a source image and input the source image into the global text perception module (2), the CLIP image encoder (3), the CLIP text encoder (4), and the BLIP text encoder (5) respectively; the source image includes a registered infrared source image and a visible light source image; The CLIP image encoder (3) is used to generate CLIP image features of the source image and input them into the global text perception module (2) and the feature extraction network (6) respectively; the CLIP text encoder (4) is used to generate CLIP text features of the source image and input them into the global text perception module (2); the BLIP text encoder (5) is used to generate BLIP text features of the source image and input them into the feature extraction network (6); The global text perception module (2) is used to fuse the source image and integrate the CLIP image features into the fusion result according to the CLIP text features to obtain a preliminary fusion result, so as to enrich the global information of the preliminary fusion result, and input the preliminary fusion result into the feature extraction network (6); At least one local text perception module (9) is embedded in the feature extraction network (6). The feature extraction network (6) is used to extract features from the preliminary fusion result, and integrate the CLIP image features into the feature extraction result according to the BLIP text features during the feature extraction process, so as to enrich the local information of the feature extraction result, and input the feature extraction result into the decoder (7); The decoder (7) is used to decode the feature extraction result to generate a final fusion image, and output the final fusion image through the output layer (8).

2. The multimodal image fusion model based on global and local text perception according to claim 1, wherein The global text perception module (2) includes a first convolutional layer (201), a first separation layer (202), a first max pooling layer (203), a first residual module (204), a first cross-attention module (205), a second convolutional layer (206), a second separation layer (207), a second max pooling layer (208), a second residual module (209), and a second cross-attention module (210); The first convolutional layer (201) is configured to receive the visible light source image and output a visible light feature map to the first separation layer (202). The first separation layer (202) is configured to divide the visible light feature map into two parts. The second convolutional layer (206) is configured to receive the infrared source image and output an infrared feature map to the second separation layer (207). The second separation layer (207) is configured to divide the infrared feature map into two parts. After corresponding parts of the visible light feature map and the infrared feature map are respectively subjected to max pooling by the first max pooling layer (203) and the second max pooling layer (208), they are concatenated along the channel dimension and input into the first residual module (204). The corresponding other parts of the visible light feature map and the infrared feature map are concatenated along the channel dimension and then input into the second residual module (209). The output feature of the first residual module (204) is added to the CLIP image feature and then input into the first cross-attention module (205). The output feature of the second residual module (209) is added to the CLIP image feature and then input into the second cross-attention module (210). The first cross-attention module (205) uses the CLIP text feature as the query vector and the image feature input into the first cross-attention module (205) as the key vector and the value vector to perform cross-attention calculation to obtain a first fusion feature. The second cross-attention module (210) uses the CLIP text feature as the query vector and the image feature input into the second cross-attention module (210) as the key vector and the value vector to perform cross-attention calculation to obtain a second fusion feature. The first fusion feature and the second fusion feature are added together to obtain the preliminary fusion result.

3. The multimodal image fusion model based on global and local text perception according to claim 2, wherein Both the first residual module (204) and the second residual module (209) include a plurality of first extraction layers (10) and a second extraction layer (11). The plurality of first extraction layers (10) are connected in series in sequence. The input end of the first first extraction layer (10) is connected to the input end of the corresponding residual module. The output feature of the last first extraction layer (10) is added to the input of the corresponding residual module and then input into the second extraction layer (11). The output end of the second extraction layer (11) is connected to the output end of the corresponding residual module. The first extraction layer (10) includes a third convolutional layer (1001) and a first ReLU activation function layer (1002). The second extraction layer (11) includes a fourth convolutional layer (1101) and a second ReLU activation function layer (1102).

4. The multimodal image fusion model based on global and local text perception according to claim 1, wherein The feature extraction network (6) includes multiple layers of residual spatial state modules (12), where a local text perception module (9) is embedded after every N layers of the residual spatial state modules (12), and N is a positive integer; the CLIP image encoder (3) inputs the CLIP image features into each local text perception module (9); the BLIP text encoder (5) inputs the BLIP text features into each local text perception module (9).

5. The multimodal image fusion model based on global and local text perception according to claim 4, wherein The residual spatial state module (12) includes a first regularization layer (1201), a visual state space module (1202), a first scaling layer (1203), a second regularization layer (1204), a seventh convolutional layer (1205), a channel attention module (1206), and a second scaling layer (1207); The first regularization layer (1201) and the visual state space module (1202) are connected in sequence, and the second regularization layer (1204), the seventh convolutional layer (1205), and the channel attention module (1206) are connected in sequence; the input ends of the first regularization layer (1201) and the first scaling layer (1203) are both connected to the input end of the residual spatial state module (12), and the output of the visual state space module (1202) and the output of the first scaling layer (1203) are added and then input into the second regularization layer (1204) and the second scaling layer (1207); the output of the channel attention module (1206) and the output of the second scaling layer (1207) are added and then output from the output end of the residual spatial state module (12).

6. The multimodal image fusion model based on global and local text perception according to claim 5, characterized in that The visual state space module (1202) includes a first linear layer (13), a depthwise separable convolutional layer (14), a first Swish activation function layer (15), a 2D selective scanning module (16), a layer normalization module (17), a second linear layer (18), a second Swish activation function layer (19), and a third linear layer (20); The first linear layer (13), the depthwise separable convolutional layer (14), the first Swish activation function layer (15), the 2D selective scanning module (16), and the layer normalization module (17) are connected in sequence, and the second linear layer (18) and the second Swish activation function layer (19) are connected in sequence; the input ends of the first linear layer (13) and the second linear layer (18) are both connected to the input end of the visual state space module (1202), the output of the layer normalization module (17) and the output of the second Swish activation function layer (19) are multiplied and then input into the third linear layer (20), and the output end of the third linear layer (20) is connected to the output end of the visual state space module (1202).

7. The multimodal image fusion model based on global and local text perception according to claim 1, wherein The local text perception module (9) includes several layers of squeeze-and-excitation modules (901), an embedding module (902), a third cross-attention module (903), several dilated convolutional layers (904), and a fifth convolutional layer (905); A plurality of said squeeze-and-excitation modules (901), said embedding module (902) and said third cross-attention module (903) are connected in sequence, and the input end of the first-layer squeeze-and-excitation module (901) is connected to the input end of the local text perception module (9); The squeeze-and-excitation module (901) is used to enhance important features and suppress redundant information of the image features input to the squeeze-and-excitation module (901) through an adaptive channel attention mechanism; The embedding module (902) is used to map the CLIP image features into the semantic space of the image features input to the embedding module (902) through a multi-layer perceptron and then embed the image features input to the embedding module (902); The third cross-attention module (903) uses the BLIP text features as query vectors and the image features input to the third cross-attention module (903) as key vectors and value vectors to perform cross-attention calculations, and inputs the calculation results into each of the dilated convolutional layers (904); The dilation rates of each of the dilated convolutional layers (904) are different, and the outputs of all the dilated convolutional layers (904) are connected in channels and then input to the fifth convolutional layer (905); the output end of the fifth convolutional layer (905) is connected to the output end of the local text perception module (9).

8. A multi-modal image fusion method based on global and local text perception, characterized in that, It includes the steps of: A1. Construct an initial multi-modal image fusion model based on global and local text perception; the initial multi-modal image fusion model based on global and local text perception is the multi-modal image fusion model based on global and local text perception according to any one of claims 1-7; A2. Train the initial multi-modal image fusion model based on global and local text perception to obtain a trained multi-modal image fusion model based on global and local text perception; A3. Input the source image to be fused into the trained multi-modal image fusion model based on global and local text perception to obtain a fused image output by the trained multi-modal image fusion model based on global and local text perception.

9. The multimodal image fusion method based on global and local text perception according to claim 8, wherein Step A2 includes: A201. Obtain a plurality of training samples; the training samples include sample source images and corresponding clean source images; the sample source images include sample infrared source images and sample visible light source images, the clean source images include clean infrared source images and clean visible light source images, the sample source images contain degradation information, the clean source images do not contain degradation information, and the degradation information includes weather interference information; A202. Select the training samples, input the corresponding sample source images into the initial multi-modal image fusion model based on global and local text perception to obtain a fused image output by the initial multi-modal image fusion model based on global and local text perception, denoted as the first fused image; A203. Input the first fused image into the CLIP image encoder (3) to obtain CLIP image features corresponding to the first fused image, denoted as the first CLIP image features; Input the clean source image of the selected training sample into the CLIP text encoder (4) to obtain CLIP text features corresponding to the clean source image, denoted as the first CLIP text features; Align the first CLIP text features with the first CLIP image features in the same feature space; Calculate the loss function according to the first fused image, the aligned first CLIP text features, and the clean source image of the selected training sample; Optimize the model parameters of the initial multi-modal image fusion model based on global and local text perception according to the loss function; Repeat steps A202 - A207 until the loss function is less than a preset loss function threshold or the number of iterations reaches a preset number threshold.

10. The multimodal image fusion method based on global and local text perception according to claim 9, wherein, The loss function is: ; ; ; Among them, is the loss function value, is the pixel-level loss, is the driving loss, is the first fused image, is the first CLIP text feature after alignment, represents the norm, is the clean infrared source image of the selected training sample, is the clean visible light source image of the selected training sample, represents the 1-norm.

Citation Information

Patent Citations

  • Lightweight complex scene image fusion model and real-time image fusion method

    CN118918428A

  • Vision-language pre-training general framework for realizing multi-granularity cross-modal alignment

    CN119206697A