A multimodal image fusion method and system based on cross-modal interactive perception

By introducing a method of cross-modal interaction perception in multi-modal image fusion, a fusion model including encoder, channel-level correction, dynamic cross-modal interaction and decoder modules is constructed, which solves the problems of information alignment, information loss and insufficient perceptual interaction among modals in the prior art, and achieves higher quality and efficient image fusion.

CN119810606BActive Publication Date: 2025-06-06HUNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510280466.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-06
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The existing multimodal image fusion method has insufficient information alignment between modals, information loss and detail retention, and inter-modal perception interaction, resulting in poor quality of the fusion image.

Method used

A multimodal image fusion method based on cross-modal interaction perception is proposed. By constructing a multimodal image fusion model, the model includes an encoder module, a channel-level correction module, a dynamic cross-modal interaction module and a decoder module. The end-to-end training method is adopted to achieve more refined inter-modal information interaction and deep-aware modeling.

Benefits of technology

This method can effectively interact with relevant information between multimodal data, improve the quality and efficiency of image fusion, enhance the shared representation learning and information complementarity between modals, and provide a more comprehensive and coordinated fusion feature representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119810606B_ABST
    Figure CN119810606B_ABST
Patent Text Reader

Abstract

The present invention discloses a multimodal image fusion method and system based on cross-modal interactive perception, constructs a multimodal image fusion model and trains it, inputs the multimodal image to be fused into the trained multimodal image fusion model for processing, an encoder module receives the multimodal image to be fused and performs layer-by-layer encoding processing, outputs several layers of feature maps of different scales, a channel-level correction module receives several layers of feature maps of different scales and performs weighted correction, outputs several layers of weight-corrected modal features, a dynamic cross-modal interaction module receives several layers of weighted-corrected modal features and processes them to obtain several layers of fusion features, a decoder module receives several layers of fusion features and performs layer-by-layer decoding and fusion processing, and outputs a fusion image corresponding to the multimodal image to be fused. The method can effectively interact the relevant information between multimodal data through the channel-level correction module and the dynamic cross-modal interaction module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image data processing, and in particular to a multimodal image fusion method and system based on cross-modal interactive perception. Background Art

[0002] As an important means of multi-source information integration, image fusion technology has been widely used in remote sensing, medical imaging, autonomous driving, intelligent monitoring and other fields in recent years. The core goal of image fusion is to combine information from different modalities (such as visible light images, infrared images, depth images, etc.) to obtain fused image features, which can effectively overcome the limitations of single-modality images, provide richer target information and stronger image understanding capabilities, thereby improving the performance of tasks such as target detection and image recognition. For example, in the field of medical imaging, there are image data of different modalities, including computed tomography (CT), positron emission tomography (PET), nuclear magnetic resonance imaging (MRI), etc. Each imaging modality has a different imaging method for tissues, lesions or organs. CT images can provide detailed anatomical information, PET images mainly highlight areas with increased metabolic activity, and MRI can display structural images inside the human body. By fusing these images of different modalities, the accuracy of disease detection and early diagnosis capabilities can be effectively improved.

[0003] In recent years, the research on multimodal image fusion has gradually developed from the traditional simple fusion methods based on pixel level or feature level to more complex deep learning methods. The pixel-level fusion method generates a fused image by performing pixel-by-pixel weighted summation or splicing of multiple input modal images. This method is simple and computationally efficient, but due to the large differences in the perceptual characteristics of different modal images, pixel-level fusion often cannot well preserve the unique information of each modality, and is easily affected by factors such as noise and resolution differences, resulting in unstable quality of the fused image; the feature-level fusion method extracts high-level features of the image (such as edges, textures, shapes, etc.) for fusion, which can usually better preserve the structural information of the image and improve the fusion effect. However, when the features between the modalities are inconsistent or misaligned, information loss and mismatching are prone to occur, especially for multimodal images in complex scenes. In recent years, deep learning technology has been widely used, especially the success of convolutional neural network (CNN) in the field of image processing, making multimodal image fusion methods based on deep learning a research hotspot. Through convolutional neural networks, we can automatically learn the correlation features between images of different modalities from the data, and jointly optimize the images of different modalities, thereby overcoming the impact of inter-modal differences to a certain extent. Although deep learning-based image fusion methods have made significant progress in many tasks, they usually do not consider the deep feature mapping relationship between different modalities, making it difficult to achieve information complementarity between modalities. Existing multimodal image fusion methods still have the following problems: (1) Difficulty in aligning information between modalities: There are significant differences in the perceptual characteristics between images of different modalities, including image resolution, color, texture, etc. These differences make it difficult to accurately align information between modalities. Traditional fusion methods often find it difficult to effectively handle such differences, resulting in poor fusion image quality; (2) Information loss and detail retention problems: Many current fusion methods mainly focus on retaining information of certain specific modalities, which will lose some detail information and make it difficult to fully integrate the advantages of each modality; (3) Insufficient interactivity in cross-modal perception: Although some existing methods attempt to achieve feature mapping between modalities through deep learning, most methods still fail to fully explore the interactive perception mechanism between different modalities. Summary of the invention

[0004] In response to the above problems, the present invention proposes a multimodal image fusion method and system based on cross-modal interactive perception. Through more sophisticated inter-modal information interaction and depth perception modeling, it can effectively interact the relevant information between multimodal data to achieve higher quality and more efficient image fusion.

[0005] On the one hand, the present invention provides a multimodal image fusion method based on cross-modal interactive perception, comprising the following steps:

[0006] S1, obtaining multimodal images for training and true value images corresponding to the multimodal images and constructing a training set;

[0007] S2. construct a multimodal image fusion model, which includes an encoder module, a channel-level correction module, a dynamic cross-modal interaction module, and a decoder module connected in sequence;

[0008] S3, using the training set to perform end-to-end training on the multimodal image fusion model to obtain the prediction result of the multimodal image, calculating the deviation between the prediction result of the multimodal image and the corresponding true value image through a preset loss function, returning the gradient and updating the parameters of the multimodal image fusion model to obtain the trained multimodal image fusion model;

[0009] S4, obtaining a multimodal image to be fused, inputting the multimodal image to be fused into a trained multimodal image fusion model for processing, an encoder module receiving the multimodal image to be fused and performing layer-by-layer encoding processing, and outputting several layers of feature maps of different scales;

[0010] S5, the channel-level correction module receives several layers of feature maps of different scales and performs weighted correction, and outputs several layers of weighted corrected modal features;

[0011] S6, the dynamic cross-modal interaction module receives and processes several layers of weighted corrected modal features to obtain several layers of fusion features;

[0012] S7. The decoder module receives several layers of fusion features and performs layer-by-layer decoding and fusion processing, and outputs a fused image corresponding to the multimodal image to be fused.

[0013] Preferably, the encoder module in S2 includes a first encoder module and a second encoder module that share weights, the first encoder module and the second encoder module each include several layers of encoding blocks connected in sequence, the channel-level correction module includes several channel correction blocks, the dynamic cross-modal interaction module includes several dynamic cross-modal interaction blocks, the decoder module includes several layers of decoding blocks connected in sequence, the number of decoding blocks, channel correction blocks and dynamic cross-modal interaction blocks is the same as the number of layers of encoding blocks, several layers of encoding blocks are respectively connected to several channel correction blocks, and several channel correction blocks are connected to several layers of decoding blocks through several dynamic cross-modal interaction blocks.

[0014] Preferably, each of the plurality of channel correction blocks comprises a channel dimension splicing layer, a normalization layer, a feature transformation layer, a Selective SSM layer, a linear transformation layer and a weighted correction layer which are connected in sequence.

[0015] Preferably, each of the several dynamic cross-modal interaction blocks includes a first convolution preprocessing unit, a second convolution preprocessing unit, a regional feature unit, a local feature unit, a regional Mamba block, a local Mamba block and a dynamic interaction Mamba block, wherein the first convolution preprocessing unit, the regional feature unit and the regional Mamba block are connected in sequence, the second convolution preprocessing unit, the local feature unit and the local Mamba block are connected in sequence, and the regional Mamba block and the local Mamba block are respectively connected to the dynamic interaction Mamba block.

[0016] Preferably, the multimodal image includes a modality one image and a modality two image. The encoder module in S4 receives the input multimodal image and performs layer-by-layer encoding processing to output feature maps of several layers with different scales. The specific process is as follows:

[0017] S41, a first encoder module and a second encoder module of the encoder module receive a modality 1 image and a modality 2 image respectively;

[0018] S42, the encoding blocks of the first encoder module connected in sequence encode the modality one image layer by layer, and correspondingly output feature maps of the modality one image at several layers with different scales;

[0019] S43, the encoding blocks of the several layers of the second encoder module connected in sequence encode the modal two image layer by layer, and correspondingly output feature maps of several layers of different scales of the modal two image.

[0020] Preferably, the specific process of S5 is as follows:

[0021] S51, each channel correction block in the channel-level correction module receives the feature maps of the modality 1 image and the modality 2 image at the corresponding layer;

[0022] S52, a channel dimension splicing layer splices the feature maps of the modality 1 image and the modality 2 image on the corresponding layer along the channel direction to obtain a spliced ​​feature map;

[0023] S53, the normalization layer normalizes the concatenated feature map to obtain normalized features;

[0024] S54, the feature transformation layer transforms the normalized features from the original spatial arrangement into a channel sequence view;

[0025] S55, Selective SSM layer receives channel sequence views and models them, extracting important feature sequences;

[0026] S56, the linear transformation layer maps the important feature sequence into a channel weight vector using linear transformation and activation function;

[0027] S57, the weighted correction layer multiplies the channel weight vector with the modal one feature map and the modal two feature map channel by channel to obtain the weighted corrected modal one feature and modal two feature of the corresponding layer.

[0028] Preferably, the specific process of S6 is as follows:

[0029] S61, each dynamic cross-modal interaction block in the dynamic cross-modal interaction module receives the weighted corrected modality one feature and modality two feature of the corresponding layer;

[0030] S62, the first convolution preprocessing unit receives the weighted corrected modal one feature and performs convolution processing, the regional feature unit divides the convolution processed feature into a plurality of regions, and regards the feature of each region as a regional-level feature marker, thereby obtaining a plurality of regional feature markers corresponding to the weighted corrected modal one feature;

[0031] S63, the second convolution preprocessing unit receives the weighted corrected modal 2 feature and performs convolution processing, the local feature unit divides the convolution processed feature into a plurality of regions, and further divides each region in the plurality of regions into a plurality of local feature units, thereby obtaining a plurality of local feature units corresponding to the weighted corrected modal 2 feature;

[0032] S64, the regional Mamba block receives several regional-level feature tags corresponding to the weight-corrected modal-one feature and performs parallel modeling and information exchange to obtain a modal-one regional feature;

[0033] S65, the local Mamba block receives and processes a plurality of local feature units corresponding to the weighted corrected modal 2 features to obtain modal 2 local features;

[0034] S66, the dynamic interaction Mamba block receives the regional features of modality one and the local features of modality two and performs dynamic fusion to obtain the fused features of the corresponding layer.

[0035] Preferably, the specific process of S7 is as follows:

[0036] S71, the lowest level decoding block in the decoder module receives and processes the lowest level fusion features to obtain the lowest level decoding features;

[0037] S72, other non-bottom-layer decoding blocks in the decoder module respectively receive the decoding features output by the adjacent next-layer decoding blocks, and receive and process the fusion features output by the corresponding dynamic cross-modal interaction blocks to obtain the decoding features of the corresponding layers;

[0038] S73. Use the decoding features output by the top-level decoding block in the decoder module as the fused image corresponding to the modality one image and the modality two image.

[0039] Preferably, the coding block is specifically a visual state space block.

[0040] Another aspect of the present invention provides a multimodal image fusion system, the fusion system comprising an image acquisition module, a computer system and a multimodal image fusion model, the image acquisition module is connected to the computer system, and the multimodal image fusion model is arranged in the computer system, wherein:

[0041] The image acquisition module is used to acquire multimodal images and send them to the computer system;

[0042] The multimodal image fusion model in the computer system uses the above-mentioned multimodal image fusion method based on cross-modal interactive perception to process the multimodal image and output a fused image.

[0043] The above-mentioned multimodal image fusion method and system based on cross-modal interactive perception construct a multimodal image fusion model, which includes an encoder module, a channel-level correction module, a dynamic cross-modal interaction module and a decoder module connected in sequence. The multimodal image fusion model is trained using a training set and the loss is calculated through a preset loss function to obtain a trained multimodal image fusion model. The multimodal image to be fused is input into the trained multimodal image fusion model for processing. The encoder module receives the multimodal image to be fused and performs layer-by-layer encoding processing, and outputs several layers of feature maps of different scales. The channel-level correction module receives several layers of feature maps of different scales and performs weighted correction, and outputs several layers of weighted corrected modal features. The dynamic cross-modal interaction module receives several layers of weighted corrected modal features and processes them to obtain several layers of fused features. The decoder module receives several layers of fused features and performs layer-by-layer decoding and fusion processing to output a fused image corresponding to the multimodal image to be fused. On the one hand, this method builds channel state space blocks between multimodal features through a channel-level correction module to enhance shared representation learning between modalities, and helps filter out noise from a specific modality by emphasizing the relevant features between multimodal features. On the other hand, considering that images of different modalities can reflect different forms of information, a dynamic cross-modal interaction module is designed to effectively integrate position information and contextual information, and different visual state space blocks are used to learn high-order semantic information and detailed texture information of images of different modalities. Through dynamic interaction, the information between different modalities is complemented, thereby providing a more comprehensive and coordinated fusion feature representation. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a flow chart of a multimodal image fusion method based on cross-modal interactive perception in one embodiment of the present invention;

[0045] Figure 2 is a schematic diagram of the network structure of a multimodal image fusion model in one embodiment of the present invention;

[0046] Figure 3 is a schematic diagram of the network structure of a channel correction block in one embodiment of the present invention;

[0047] Figure 4 is a schematic diagram of a network structure of a dynamic cross-modal interaction block in an embodiment of the present invention;

[0048] Figure 5 It is a schematic diagram of the system structure of a multimodal image fusion system in one embodiment of the present invention. DETAILED DESCRIPTION

[0049] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings.

[0050] See also Figure 1 , a multimodal image fusion method based on cross-modal interactive perception, the method specifically includes:

[0051] S1. Obtain multimodal images for training and true value images corresponding to the multimodal images and construct a training set.

[0052] Specifically, multimodal image data is obtained, and the multimodal image data can be PET (Positron Emission Tomography) images and CT (Computed Tomography) images in medical images. Since PET images show high brightness and high contrast characteristics in metabolically active areas (such as tumors), and CT images have higher spatial resolution and structural detail information, the two modalities have complementary characteristics. PET images and CT images are also used as true value images, and the PET images, CT images, and the true value images corresponding to the PET images and CT images constitute the training set.

[0053] S2. Construct a multimodal image fusion model, which includes an encoder module, a channel-level correction module, a dynamic cross-modal interaction module, and a decoder module connected in sequence.

[0054] See also Figure 2The multimodal image fusion model includes an encoder module, a channel-level correction module, a dynamic cross-modal interaction module and a decoder module. The encoder module is used to encode the input multimodal image layer by layer to obtain several layers of feature maps of different sizes. The channel-level correction module is used to receive several layers of feature maps of different scales and perform weighted correction, and output several layers of weighted corrected modal features. The dynamic cross-modal interaction module receives several layers of weighted corrected modal features and processes them to obtain several layers of fusion features. The decoder module receives several layers of fusion features and performs layer-by-layer decoding and fusion processing, and outputs a fusion image corresponding to the multimodal image to be fused.

[0055] Furthermore, the encoder module in S2 includes a first encoder module and a second encoder module that share weights, the first encoder module and the second encoder module each include several layers of encoding blocks connected in sequence, the channel-level correction module includes several channel correction blocks, the dynamic cross-modal interaction module includes several dynamic cross-modal interaction blocks, the decoder module includes several layers of decoding blocks connected in sequence, the number of decoding blocks, channel correction blocks and dynamic cross-modal interaction blocks is the same as the number of layers of encoding blocks, several layers of encoding blocks are respectively connected to several channel correction blocks, and several channel correction blocks are connected to several layers of decoding blocks through several dynamic cross-modal interaction blocks.

[0056] See also Figure 2 , the encoder module includes two parallel branches: a first encoder module and a second encoder module. The first encoder module and the second encoder module are used to process image features of different modalities respectively. The first encoder module and the second encoder module are both composed of four layers of encoding blocks. Among them, the first encoder module includes encoding block 1, encoding block 2, encoding block 3 and encoding block 4, and the second encoder module includes encoding block 5, encoding block 6, encoding block 7 and encoding block 8. The encoding block can be a convolutional neural network, and several layers of feature maps of different scales are extracted layer by layer through the convolutional neural network. In this application, VSS (Visual State Space) is used as an encoding block. The visual state space block is an implementation of extending the state space model (SSM) to visual tasks, which captures global context information by scanning image features (for example, mapping two-dimensional features to several row and column sequences). Each visual state space block can realize multi-level downsampling of input features to obtain feature maps of different scales.

[0057] Further, see Figure 3 Each of the several channel correction blocks includes a channel dimension splicing layer, a normalization layer, a feature transformation layer, a Selective SSM layer, a linear transformation layer and a weighted correction layer connected in sequence.

[0058] Specifically, each channel correction block in the channel-level correction module aims to model the correlation and weighted correction of the input feature map in the channel dimension, so as to achieve effective fusion and enhancement of multimodal features. The multimodal image fusion model uses SSM (State Space Models) and its extensions, such as Selective SSM, to achieve efficient modeling of multimodal image features. SSM is a sequence-to-sequence modeling method with LTI (Linear Time-Invariant) characteristics, which maps a one-dimensional input sequence to an output sequence and captures the inherent dynamic characteristics of the system in the hidden state. The parameter data of traditional SSM is independent and time-invariant, and its adaptability to input complexity is limited. To this end, Selective SSM is introduced to make the parameters sensitive to the input data, and adaptively adjust the input features through linear mapping, thereby improving the feature selection ability of SSM for the input sequence.

[0059] Further, see Figure 4 Each of the several dynamic cross-modal interaction blocks includes a first convolution preprocessing unit, a second convolution preprocessing unit, a regional feature unit, a local feature unit, a regional Mamba block, a local Mamba block and a dynamic interaction Mamba block, wherein the first convolution preprocessing unit, the regional feature unit and the regional Mamba block are connected in sequence, the second convolution preprocessing unit, the local feature unit and the local Mamba block are connected in sequence, and the regional Mamba block and the local Mamba block are respectively connected to the dynamic interaction Mamba block.

[0060] Specifically, each dynamic cross-modal interaction block in the dynamic cross-modal interaction module aims to collaboratively process and fuse the two modalities at different spatial scales (regional level and local level), thereby making full use of the advantages of different modal images in spatial resolution and structural details to achieve accurate identification and segmentation of the lesion area.

[0061] S3. Use the training set to perform end-to-end training on the multimodal image fusion model to obtain the prediction result of the multimodal image, calculate the deviation between the prediction result of the multimodal image and the corresponding true value image through a preset loss function, return the gradient and update the parameters of the multimodal image fusion model to obtain the trained multimodal image fusion model.

[0062] Specifically, the multimodal image fusion model is trained end-to-end using a training set (including training multimodal images and corresponding true value images), and the fusion and target recognition results of the training multimodal images, that is, the prediction results, are predicted. The deviation between the prediction results and the true value images is calculated through a preset loss function, and the gradient is returned to update the parameters of the multimodal image fusion model. Through iterative optimization, the multimodal image fusion model can fully extract and fuse information in the multimodal image feature space, thereby correctly identifying and locating the target area.

[0063] S4. Obtain the multimodal image to be fused, input the multimodal image to be fused into the trained multimodal image fusion model for processing, the encoder module receives the multimodal image to be fused and performs layer-by-layer encoding processing, and outputs several layers of feature maps of different scales.

[0064] In one embodiment, the multimodal image includes a modality 1 image and a modality 2 image. The encoder module in S4 receives the input multimodal image and performs layer-by-layer encoding processing to output feature maps of several layers at different scales. The specific process is as follows:

[0065] S41, a first encoder module and a second encoder module of the encoder module receive a modality 1 image and a modality 2 image respectively;

[0066] S42, the encoding blocks of the first encoder module connected in sequence encode the modality one image layer by layer, and correspondingly output feature maps of the modality one image at several layers with different scales;

[0067] S43, the encoding blocks of the several layers of the second encoder module connected in sequence encode the modal two image layer by layer, and correspondingly output feature maps of several layers of different scales of the modal two image.

[0068] Specifically, see Figure 2 The first encoder module includes encoding block 1, encoding block 2, encoding block 3 and encoding block 4, and the second encoder module includes encoding block 5, encoding block 6, encoding block 7 and encoding block 8. The first encoder module receives the modality one image, and encoding block 1, encoding block 2, encoding block 3 and encoding block 4 respectively encode the modality one image layer by layer, and correspondingly obtain the first feature map, second feature map, third feature map and fourth feature map of the modality one image; the second encoder module receives the modality two image, and encoding block 5, encoding block 6, encoding block 7 and encoding block 8 respectively encode the modality two image layer by layer, and correspondingly obtain the first feature map, second feature map, third feature map and fourth feature map of the modality two image.

[0069] S5. The channel-level correction module receives several layers of feature maps of different scales and performs weighted correction, and outputs several layers of weighted corrected modal features.

[0070] In one embodiment, the specific process of S5 is as follows:

[0071] S51, each channel correction block in the channel-level correction module receives feature maps of the modality one image and the modality two image on the corresponding layer;

[0072] S52, a channel dimension splicing layer splices the feature maps of the modality one image and the modality two image on the corresponding layer along the channel direction to obtain a spliced ​​feature map;

[0073] S53, the normalization layer normalizes the concatenated feature map to obtain normalized features;

[0074] S54, the feature transformation layer transforms the normalized features from the original spatial arrangement into a channel sequence view;

[0075] S55, SSM layer receives channel sequence view and models it, extracting important feature sequences;

[0076] S56, the linear transformation layer maps the important feature sequence into a channel weight vector using linear transformation and activation function;

[0077] S57, the weighted correction layer multiplies the channel weight vector with the modal one feature map and the modal two feature map channel by channel to obtain the weighted corrected modal one feature and modal two feature of the corresponding layer.

[0078] Specifically, see Figure 2 and Figure 3 The channel-level correction module includes four channel correction blocks, namely channel correction block 1, channel correction block 2, channel correction block 3 and channel correction block 4. Each channel correction block includes a channel dimension splicing layer, a normalization layer, a feature transformation layer, a Selective SSM layer, a linear transformation layer and a weighted correction layer connected in sequence. Each channel correction block receives and processes the feature maps of the modality 1 image and the modality 2 image on the corresponding layer, and obtains the weighted corrected modality 1 features and modality 2 features of the corresponding layer.

[0079] From the output of the encoder module, two sets of feature maps with the same spatial size are obtained. The height and width of the feature maps are H and W respectively, and the number of channels is C. This ensures that the feature dimensions are aligned in the subsequent channel splicing step. Take channel correction block 1 as an example:

[0080] 1) Channel correction block 1 receives the first layer feature map of modality 1 image and the first layer feature map of modality 2 image respectively.

[0081] 2) The channel dimension splicing layer of the channel correction block 1 splices the first layer feature map of the modality one image and the first layer feature map of the modality two image in the channel direction to form a joint feature map containing information of the two modalities, that is, the spliced ​​feature map. At this time, the spliced ​​feature map has 2C channels in the channel dimension.

[0082] 3) In order to reduce the impact of the difference in numerical scale between modalities, the normalization layer of the channel correction block 1 normalizes the concatenated feature map (such as using layer normalization) to obtain normalized features. Normalization can make subsequent feature processing more stable.

[0083] 4) The feature transformation layer of the channel correction block 1 transforms the normalized feature map from the original spatial arrangement (H×W×2C) to a channel sequence view. For example, the normalized feature map can be flattened in space and the features in each channel direction can be regarded as a sequence.

[0084] 5) After extracting the important feature sequence, the Selective SSM layer of the channel correction block 1 models the converted channel sequence, thereby fully mining the complementary information of the bimodal features at the channel level and extracting the important feature sequence.

[0085] Unlike ordinary state-space models that only have fixed parameters, some parameters of the selective state-space model can be dynamically adjusted according to the input data, thereby making more flexible selection and memory of the input feature sequence. Through this process, the channel-level correction module can capture the long-range correlation and global information of bimodal features in the channel dimension, allowing the network to selectively emphasize important features and filter redundant components within the channel range.

[0086] 6) The linear transformation layer of channel correction block 1 uses linear transformation and activation function to map the important feature sequence into a channel weight vector. The weight vector will be divided into two according to the channel order, corresponding to the channel weights of the bimodal features. By assigning specific weights to each channel, the channel-level correction module can highlight the more useful and prominent feature channels in different modalities and relatively suppress irrelevant or noisy channels.

[0087] 7) The weighted correction layer of the channel correction block 1 multiplies the channel weight vector by the first layer feature map of the modality one image and the first layer feature map of the modality two image channel by channel to obtain the weighted corrected modality one feature and modality two feature of the first layer.

[0088] The weighted modal one and modal two features can be regarded as filtered and enhanced features, in which the complementarity of multimodal information has been fully utilized and strengthened at the channel level.

[0089] Using the method described above, channel correction block 2, channel correction block 3 and channel correction block 4 respectively receive and process the second layer feature map, third layer feature map and fourth layer feature map corresponding to the modal one image and the modal two image, respectively, and correspondingly obtain the weighted corrected modal one feature and modal two feature of the second layer, the weighted corrected modal one feature and modal two feature of the third layer, and the weighted corrected modal one feature and modal two feature of the fourth layer. The specific process will not be repeated here.

[0090] S6. The dynamic cross-modal interaction module receives and processes several layers of weighted corrected modal features to obtain several layers of fusion features.

[0091] In one embodiment, the specific process of S6 is as follows:

[0092] S61. Each dynamic cross-modal interaction block in the dynamic cross-modal interaction module receives the weighted corrected modality one feature and modality two feature of the corresponding layer.

[0093] S62. The first convolution preprocessing unit receives the weighted corrected modal feature and performs convolution processing. The regional feature unit divides the convolution-processed feature into several regions, and regards the feature of each region as a regional-level feature label, thereby obtaining several regional feature labels corresponding to the weighted corrected modal feature.

[0094] S63, the second convolution preprocessing unit receives the weighted corrected modal two features and performs convolution processing, the local feature unit divides the convolution processed features into a number of regions, and further divides each of the several regions into a number of local feature units, thereby obtaining a number of local feature units corresponding to the weighted corrected modal two features.

[0095] S64, the regional Mamba block receives several regional-level feature tags corresponding to the weight-corrected modal-one feature and performs parallel modeling and information exchange to obtain the modal-one regional feature.

[0096] S65, the local Mamba block receives and processes a number of local feature units corresponding to the weighted corrected modal 2 features to obtain modal 2 local features.

[0097] S66, the dynamic interaction Mamba block receives the regional features of modality one and the local features of modality two and performs dynamic fusion to obtain the fused features of the corresponding layer.

[0098] Specifically, after the channel-level correction module, the weighted corrected modality 1 features and modality 2 features that have been enhanced and filtered in the channel dimension can be obtained. These features still exist in the form of two-dimensional images and contain spatial distribution information. Figure 2 and Figure 4The dynamic cross-modal interaction module includes four dynamic cross-modal interaction blocks, namely, dynamic cross-modal interaction block 1, dynamic cross-modal interaction block 2, dynamic cross-modal interaction block 3 and dynamic cross-modal interaction block 4. Each dynamic cross-modal interaction block includes a first convolution preprocessing unit, a second convolution preprocessing unit, a regional feature unit, a local feature unit, a regional Mamba block, a local Mamba block and a dynamic interaction Mamba block, wherein the first convolution preprocessing unit, the regional feature unit and the regional Mamba block are connected in sequence, the second convolution preprocessing unit, the local feature unit and the local Mamba block are connected in sequence, and the regional Mamba block and the local Mamba block are connected to the dynamic interaction Mamba block respectively. Each dynamic cross-modal interaction block receives and processes the weighted corrected modality one feature and modality two feature of the corresponding layer to obtain the fusion feature of the corresponding layer.

[0099] Take dynamic cross-modal interaction block 1 as an example to illustrate:

[0100] 1) The first convolution preprocessing unit of the dynamic cross-modal interaction block 1 receives the modal one feature after the first layer of weighted correction and performs convolution processing. The regional feature unit of the dynamic cross-modal interaction block 1 divides the convolution processed feature into several regions. The feature of each region is regarded as a regional feature label, thereby obtaining several regional feature labels corresponding to the modal one feature after the first layer of weighted correction;

[0101] 2) The second convolution preprocessing unit of the dynamic cross-modal interaction block 1 receives the weighted corrected modal 2 features and performs convolution processing. The local feature unit of the dynamic cross-modal interaction block 1 processes the convolution processed features in the same area division and divides them into several areas, so that the modal 1 and modal 2 features correspond to each other at the regional level. At this time, the weighted corrected modal 2 features still retain a high resolution and details. Each area of ​​the modal 2 features can be further divided into several local feature units, thereby obtaining several local feature units corresponding to the weighted corrected modal 2 features;

[0102] Through the above preprocessing process, the several regional feature markers corresponding to the first-layer weighted correction modal one feature represent a larger-scale spatial regional information, highlighting the target positioning characteristics at the global scale; and each of the several local feature units corresponding to the corresponding weighted correction modal two feature contains detailed structural features. In this process, the first-layer weighted correction modal one feature reflects the global positioning information, and the first-layer weighted correction modal two feature shows the local structural details.

[0103] 3) The regional Mamba block of the dynamic cross-modal interaction block 1 takes several regional feature tags corresponding to the first-layer weighted correction of the modal one feature as input, and uses the data-dependent state space modeling idea to parallel model and exchange information for all regional feature tags to obtain the modal one regional feature. Through the regional Mamba block, the first-layer weighted correction of the modal one feature establishes a global association at the regional level, so as to better identify the regional location of the potential target.

[0104] 4) The local Mamba block of the dynamic cross-modal interaction block 1 receives and processes several local feature units corresponding to the modal 2 features after the first layer of weighted correction to obtain the modal 2 local features. The local Mamba block focuses on processing the local feature units in the modal 2 features, inputs several local feature units corresponding to the modal 2 features after the first layer of weighted correction into the local Mamba block, and establishes fine feature associations in a smaller spatial range. The local Mamba block can capture the structural information at the detail level in the modal 2 features, such as the tissue details, texture features, and edge contours of the target in the medical image. This fine-grained structural modeling helps to clarify the target boundaries and morphological features, so that the modal 2 features can better make up for the shortcomings of the modal 1 features at the detail level.

[0105] 5) The dynamic interaction Mamba block of dynamic cross-modal interaction block 1 receives the regional features of modality 1 and the local features of modality 2 and dynamically fuses them to obtain the first layer of fused features.

[0106] After completing the independent processing of the first-layer weighted corrected modal one features and the weighted corrected modal two features, the modal one regional features and modal two local features are obtained accordingly. The dynamic interaction Mamba block dynamically fuses these two types of features to obtain the first-layer fused features. Specifically, each modal one regional marker is associated with all local feature units under the corresponding modal two region. By integrating the modal one regional marker into each local feature unit of modal two, the modal two features obtain a global spatial reference while retaining their own details, thereby realizing the dynamic interaction and fusion of multi-modal and multi-scale information.

[0107] Using the method described above, dynamic cross-modal interaction block 2, dynamic cross-modal interaction block 3 and dynamic cross-modal interaction block 4 respectively receive and process the modal one feature and modal two feature after the second layer weighted correction, the modal one feature and modal two feature after the third layer weighted correction, and the modal one feature and modal two feature after the fourth layer weighted correction, and obtain the second layer fusion feature, the third layer fusion feature, and the fourth layer fusion feature accordingly. The specific process will not be repeated here.

[0108] S7. The decoder module receives several layers of fusion features and performs layer-by-layer decoding and fusion processing, and outputs a fused image corresponding to the multimodal image to be fused.

[0109] In one embodiment, the specific process of S7 is as follows:

[0110] S71, the lowest level decoding block in the decoder module receives and processes the lowest level fusion features to obtain the lowest level decoding features;

[0111] S72, other non-bottom-layer decoding blocks in the decoder module respectively receive the decoding features output by the adjacent next-layer decoding blocks, and receive and process the fusion features output by the corresponding dynamic cross-modal interaction blocks to obtain the decoding features of the corresponding layers;

[0112] S73. Use the decoding features output by the top-level decoding block in the decoder module as the fused image corresponding to the modality one image and the modality two image.

[0113] Specifically, see Figure 2 The decoder module includes four decoder blocks connected in sequence, namely decoder block 1, decoder block 2, decoder block 3 and decoder block 4. Decoder block 1, decoder block 2, decoder block 3 and decoder block 4 are also connected to dynamic cross-modal interaction block 1, dynamic cross-modal interaction block 2, dynamic cross-modal interaction block 3 and dynamic cross-modal interaction block 4 respectively. The bottom layer decoder block (corresponding to Figure 2 The decoding block 4 in the dynamic cross-modal interaction block 4 receives and processes the fourth-layer fusion features output by the dynamic cross-modal interaction block 4 to obtain the bottom-layer decoding features (that is, the fourth-layer decoding features), and the top-layer decoding block (corresponding to Figure 2 The decoding block 1 in the dynamic cross-modal interaction block 1 receives the first layer fusion features output by the dynamic cross-modal interaction block 1, and also receives the second layer decoding features output by the decoding block 2. The first layer fusion features and the second layer decoding features are fused to obtain the first layer decoding features output by the decoding block 1, and the first decoding features are used as the fused image corresponding to the modality one image and the modality two image.

[0114] In one embodiment, a multimodal image fusion system includes an image acquisition module, a computer system, and a multimodal image fusion model, wherein the image acquisition module is connected to the computer system, and the multimodal image fusion model is set in the computer system, wherein:

[0115] The image acquisition module is used to acquire multimodal images and send them to the computer system;

[0116] The multimodal image fusion model in the computer system uses the above-mentioned multimodal image fusion method based on cross-modal interactive perception to process the multimodal image and output a fused image.

[0117] Specifically, see Figure 5 , Figure 5 The figure is a schematic diagram of the system structure of a multimodal image fusion system in one embodiment of the present invention.

[0118] Figure 5 A multimodal image fusion system is shown, comprising an image acquisition module, a computer system and a multimodal image fusion model, wherein the image acquisition module is connected to the computer system, the multimodal image fusion model is arranged in the computer system, the image acquisition module is used to acquire the multimodal images to be fused, and input the multimodal images to be fused into the computer system, the multimodal image fusion model is used to receive and process the multimodal images to be fused, and output a fused image corresponding to the multimodal images to be fused.

[0119] For the specific definition of a multimodal image fusion system, please refer to the definition of a multimodal image fusion method based on cross-modal interactive perception in the above text, which will not be repeated here.

[0120] Furthermore, the effect of a multimodal image fusion method based on cross-modal interactive perception in the present invention is verified through experiments.

[0121] In the ablation experiment, the present invention first constructed a multimodal medical image fusion and target recognition dataset (PET-CT lung tumor segmentation dataset, PCLT20K), and conducted experiments on a multimodal image fusion method based on cross-modal interactive perception proposed by the present invention on the PCLT20k dataset. First, a channel-level correction module and a dynamic cross-modal interaction module were added to the baseline model respectively, and then a channel-level correction module and a dynamic cross-modal interaction module (corresponding to the solution proposed by the present invention) were added to the baseline model at the same time, thereby obtaining the experimental results of multimodal image target recognition under different methods, as shown in Table 1:

[0122] Table 1 Experimental results of multimodal image target recognition under different methods

[0123] ;

[0124] The benchmark model in Table 1 is a model that only uses the encoder and decoder in the present invention. As can be seen from Table 1, after adding the channel-level correction module to the benchmark model, the accuracy is improved by 1.14%; after adding the dynamic cross-modal interaction module to the benchmark model, the accuracy is improved by 1.82%; after adding the channel-level correction module and the dynamic cross-modal interaction module to the benchmark model at the same time, the accuracy is improved by 4.06% compared with the benchmark model. This verifies that the multimodal image fusion method based on cross-modal interaction perception proposed in the present invention can significantly improve the performance of multimodal image target recognition.

[0125] Table 2 Comparative experimental results

[0126] ;

[0127] In addition, in order to further verify the effectiveness of the multimodal image fusion method based on cross-modal interactive perception proposed in the present invention, the method proposed in the present invention is compared with some previous methods with better performance, such as AsymFormer, DFormer, Sigma, etc. The comparison results are shown in Table 2.

[0128] It can be seen from Table 2 that compared with the previous better methods, the multimodal image fusion method based on cross-modal interactive perception proposed in the present invention (corresponding to the method of the present invention in Table 2) has more significant advantages.

[0129] The above-mentioned multimodal image fusion method and system based on cross-modal interactive perception construct a multimodal image fusion model, which includes an encoder module, a channel-level correction module, a dynamic cross-modal interaction module and a decoder module connected in sequence. The multimodal image fusion model is trained using a training set and the loss is calculated through a preset loss function to obtain a trained multimodal image fusion model. The multimodal image to be fused is input into the trained multimodal image fusion model for processing. The encoder module receives the multimodal image to be fused and performs layer-by-layer encoding processing, and outputs several layers of feature maps of different scales. The channel-level correction module receives several layers of feature maps of different scales and performs weighted correction, and outputs several layers of weighted corrected modal features. The dynamic cross-modal interaction module receives several layers of weighted corrected modal features and processes them to obtain several layers of fused features. The decoder module receives several layers of fused features and performs layer-by-layer decoding and fusion processing to output a fused image corresponding to the multimodal image to be fused. On the one hand, this method builds channel state space blocks between multimodal features through a channel-level correction module to enhance shared representation learning between modalities, and helps filter out noise from a specific modality by emphasizing the relevant features between multimodal features. On the other hand, considering that images of different modalities can reflect different forms of information, a dynamic cross-modal interaction module is designed to effectively integrate position information and contextual information, and different visual state space blocks are used to learn high-order semantic information and detailed texture information of images of different modalities. Through dynamic interaction, the information between different modalities is complemented, thereby providing a more comprehensive and coordinated fusion feature representation.

[0130] The above is a detailed introduction to a multimodal image fusion method and system based on cross-modal interactive perception provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the core idea of ​​the present invention. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.

Claims

1. A multimodal image fusion method based on cross-modal interactive perception, characterized in that: The method comprises: S1, obtaining multimodal images for training and true value images corresponding to the multimodal images and constructing a training set; S2. construct a multimodal image fusion model, which includes an encoder module, a channel-level correction module, a dynamic cross-modal interaction module, and a decoder module connected in sequence; S3, using the training set to perform end-to-end training on the multimodal image fusion model to obtain the prediction result of the multimodal image, calculating the deviation between the prediction result of the multimodal image and the corresponding true value image through a preset loss function, returning the gradient and updating the parameters of the multimodal image fusion model to obtain the trained multimodal image fusion model; S4, obtaining a multimodal image to be fused, inputting the multimodal image to be fused into a trained multimodal image fusion model for processing, an encoder module receiving the multimodal image to be fused and performing layer-by-layer encoding processing, and outputting several layers of feature maps of different scales; S5, the channel-level correction module receives several layers of feature maps of different scales and performs weighted correction, and outputs several layers of weighted corrected modal features; S6, the dynamic cross-modal interaction module receives and processes several layers of weighted corrected modal features to obtain several layers of fusion features; S7, the decoder module receives several layers of fusion features and performs layer-by-layer decoding and fusion processing, and outputs a fused image corresponding to the multimodal image to be fused; The specific process of S6 is as follows: S61, each dynamic cross-modal interaction block in the dynamic cross-modal interaction module receives the weighted corrected modality one feature and modality two feature of the corresponding layer; S62, the first convolution preprocessing unit receives the weighted corrected modal one feature and performs convolution processing, the regional feature unit divides the convolution processed feature into a plurality of regions, and regards the feature of each region as a regional-level feature marker, thereby obtaining a plurality of regional feature markers corresponding to the weighted corrected modal one feature; S63, the second convolution preprocessing unit receives the weighted corrected modal 2 feature and performs convolution processing, the local feature unit divides the convolution processed feature into a plurality of regions, and further divides each region in the plurality of regions into a plurality of local feature units, thereby obtaining a plurality of local feature units corresponding to the weighted corrected modal 2 feature; S64, the regional Mamba block receives several regional-level feature tags corresponding to the weight-corrected modal-one feature and performs parallel modeling and information exchange to obtain a modal-one regional feature; S65, the local Mamba block receives and processes a plurality of local feature units corresponding to the weighted corrected modal 2 features to obtain modal 2 local features; S66, the dynamic interaction Mamba block receives the regional features of modality one and the local features of modality two and performs dynamic fusion to obtain the fused features of the corresponding layer.

2. The multimodal image fusion method based on cross-modal interactive perception as claimed in claim 1, characterized in that: The encoder module in S2 includes a first encoder module and a second encoder module that share weights, the first encoder module and the second encoder module each include several layers of encoding blocks connected in sequence, the channel-level correction module includes several channel correction blocks, the dynamic cross-modal interaction module includes several dynamic cross-modal interaction blocks, the decoder module includes several layers of decoding blocks connected in sequence, the number of decoding blocks, channel correction blocks and dynamic cross-modal interaction blocks is the same as the number of layers of encoding blocks, several layers of encoding blocks are respectively connected to several channel correction blocks, and several channel correction blocks are connected to several layers of decoding blocks through several dynamic cross-modal interaction blocks.

3. The multimodal image fusion method based on cross-modal interactive perception as claimed in claim 2, characterized in that: Each of the plurality of channel correction blocks includes a channel dimension splicing layer, a normalization layer, a feature transformation layer, a Selective SSM layer, a linear transformation layer, and a weighted correction layer which are connected in sequence.

4. The multimodal image fusion method based on cross-modal interactive perception as claimed in claim 3, characterized in that: Each of the multiple dynamic cross-modal interaction blocks includes a first convolution preprocessing unit, a second convolution preprocessing unit, a regional feature unit, a local feature unit, a regional Mamba block, a local Mamba block and a dynamic interaction Mamba block, wherein the first convolution preprocessing unit, the regional feature unit and the regional Mamba block are connected in sequence, the second convolution preprocessing unit, the local feature unit and the local Mamba block are connected in sequence, and the regional Mamba block and the local Mamba block are respectively connected to the dynamic interaction Mamba block.

5. The multimodal image fusion method based on cross-modal interactive perception as claimed in claim 4, characterized in that: The multimodal image includes a modality 1 image and a modality 2 image. The encoder module in S4 receives the input multimodal image and performs layer-by-layer encoding processing, and outputs several layers of feature maps of different scales. The specific process is as follows: S41, a first encoder module and a second encoder module of the encoder module receive a modality 1 image and a modality 2 image respectively; S42, the encoding blocks of the first encoder module connected in sequence encode the modality one image layer by layer, and correspondingly output feature maps of the modality one image at several layers with different scales; S43, the encoding blocks of the several layers of the second encoder module connected in sequence encode the modal two image layer by layer, and correspondingly output feature maps of several layers of different scales of the modal two image.

6. The multimodal image fusion method based on cross-modal interactive perception as claimed in claim 5, characterized in that: The specific process of S5 is as follows: S51, each channel correction block in the channel-level correction module receives the feature maps of the modality 1 image and the modality 2 image at the corresponding layer; S52, a channel dimension splicing layer splices the feature maps of the modality 1 image and the modality 2 image on the corresponding layer along the channel direction to obtain a spliced ​​feature map; S53, the normalization layer normalizes the concatenated feature map to obtain normalized features; S54, the feature transformation layer transforms the normalized features from the original spatial arrangement into a channel sequence view; S55, Selective SSM layer receives channel sequence views and models them, extracting important feature sequences; S56, the linear transformation layer maps the important feature sequence into a channel weight vector using linear transformation and activation function; S57, the weighted correction layer multiplies the channel weight vector with the modal one feature map and the modal two feature map channel by channel to obtain the weighted corrected modal one feature and modal two feature of the corresponding layer.

7. The multimodal image fusion method based on cross-modal interactive perception as claimed in claim 6, characterized in that: The specific process of S7 is as follows: S71, the lowest level decoding block in the decoder module receives and processes the lowest level fusion features to obtain the lowest level decoding features; S72, other non-bottom-layer decoding blocks in the decoder module respectively receive the decoding features output by the adjacent next-layer decoding blocks, and receive and process the fusion features output by the corresponding dynamic cross-modal interaction blocks to obtain the decoding features of the corresponding layers; S73. Use the decoding features output by the top-level decoding block in the decoder module as the fused image corresponding to the modality one image and the modality two image.

8. The multimodal image fusion method based on cross-modal interactive perception as claimed in claim 2, characterized in that: The encoding block is specifically a visual state space block.

9. A multimodal image fusion system, characterized in that: The fusion system includes an image acquisition module, a computer system and a multimodal image fusion model. The image acquisition module is connected to the computer system, and the multimodal image fusion model is set in the computer system, wherein: The image acquisition module is used to acquire multimodal images and send them to the computer system; The multimodal image fusion model in the computer system processes the multimodal image using a multimodal image fusion method based on cross-modal interactive perception as described in any one of claims 1 to 8 to output a fused image.

Citation Information

Patent Citations

  • Named entity recognition method based on comparative learning and multi-modal semantic interaction

    CN117574904A

  • Multi-modal fusion segmentation method and device based on prior information and Mama hybrid model

    CN118691820A