Gated multi-modal image fusion method and device, electronic equipment and storage medium

By employing a gated multimodal image fusion method, which utilizes multi-layer feature extraction and differential information compensation, combined with feature fusion and secondary gated fusion, the problems of modal differences and information imbalance in image fusion are solved, thereby improving the accuracy and information richness of image fusion.

CN119963430BActive Publication Date: 2026-07-24GANSU AGRI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GANSU AGRI UNIV
Filing Date
2025-02-06
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing image fusion technologies suffer from problems such as modal differences, spatial alignment, information richness imbalance, noise and distortion, and semantic differences, which affect the accuracy of image fusion.

Method used

A gated multimodal image fusion method is adopted, which combines multi-layer feature extraction and differential information compensation processing with feature fusion and secondary gated fusion to achieve feature extraction and fusion of multimodal images, reduce modal differences and enhance information richness.

Benefits of technology

It improves the accuracy and information balance of image fusion, and enhances the visual effect and information integrity of the fused image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963430B_ABST
    Figure CN119963430B_ABST
Patent Text Reader

Abstract

The application provides a gated multi-modal image fusion method and device, electronic equipment and storage medium. The method comprises: acquiring a plurality of modal images corresponding to a target object; inputting each modal image into a feature extraction module of a multi-modal image fusion model for multi-layer feature extraction processing and differential information compensation processing to obtain a plurality of layer feature maps corresponding to each modal image; inputting the third layer feature map corresponding to each modal image into a feature fusion module in the multi-modal image fusion model for feature fusion to obtain an initial fusion feature map; inputting the multi-layer feature map corresponding to the first modal image, the multi-layer feature map corresponding to the second modal image and the initial fusion feature map into a secondary gating fusion module in the multi-modal image fusion model for multi-layer fusion to obtain a target fusion image. Through secondary fusion, the fusion of different modal images can be realized, and the target fusion image information is more balanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and more specifically, to a gated multimodal image fusion method, apparatus, electronic device, and storage medium. Background Technology

[0002] Image fusion technology plays a crucial role in the field of digital image processing. It combines image data from multiple sources to generate a single image containing information from all the input images. This technology has a wide range of applications, from remote sensing imaging and medical diagnosis to video surveillance and military reconnaissance. The significance of image fusion lies in its ability to increase the amount of information in an image, enhance visual effects, support expert decision-making, improve the performance of automated systems, and expand the scope of applications.

[0003] Existing technologies offer a variety of image fusion techniques, but they suffer from problems such as modal differences, spatial alignment, information richness imbalance, noise and distortion, and semantic differences. Summary of the Invention

[0004] The purpose of this application is to address the shortcomings of the prior art by providing a gated multimodal image fusion method, apparatus, electronic device, and storage medium to improve the accuracy of image fusion.

[0005] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows:

[0006] In a first aspect, embodiments of this application provide a gated multimodal image fusion method, the method comprising:

[0007] Acquire multiple modal images corresponding to the target object, the multiple modal images including: a first modal image, a second modal image, and a third modal image;

[0008] Each modal image is input into the feature extraction module of the multimodal image fusion model. The feature extraction module performs multi-layer feature extraction processing and differential information compensation processing to obtain multi-layer feature maps corresponding to each modal image. The multi-layer feature maps include a first-layer feature map, a second-layer feature map, and a third-layer feature map.

[0009] The third-layer feature map corresponding to each modal image is input into the feature fusion module in the multimodal image fusion model for feature fusion to obtain an initial fused feature map;

[0010] The multi-layer feature map corresponding to the first modality image, the multi-layer feature map corresponding to the second modality image, and the initial fusion feature map are input into the secondary gated fusion module in the multi-modal image fusion model for multi-layer fusion to obtain the target fused image.

[0011] Secondly, embodiments of this application also provide an image fusion apparatus, the apparatus comprising:

[0012] The acquisition module is used to acquire multiple modal images corresponding to the target object, the multiple modal images including: a first modal image, a second modal image and a third modal image;

[0013] The feature extraction module is used to input each of the modal images into the feature extraction module of the multimodal image fusion model. The feature extraction module performs multi-layer feature extraction processing and differential information compensation processing to obtain multi-layer feature maps corresponding to each of the modal images. The multi-layer feature maps include a first-layer feature map, a second-layer feature map, and a third-layer feature map.

[0014] The first fusion module is used to input the third-layer feature map corresponding to each modal image into the feature fusion module in the multimodal image fusion model for feature fusion to obtain an initial fused feature map;

[0015] The second fusion module is used to input the multi-layer feature map corresponding to the first modality image, the multi-layer feature map corresponding to the second modality image, and the initial fusion feature map into the secondary gated fusion module in the multi-modal image fusion model for multi-layer fusion to obtain the target fused image.

[0016] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a storage medium, and a bus, wherein the storage medium stores program instructions executable by the processor, and when the application runs, the processor communicates with the storage medium via the bus, and the processor executes the program instructions to perform the steps of the gated multimodal image fusion method described in the first aspect above.

[0017] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which is read and executes the steps of the gated multimodal image fusion method described in the first aspect.

[0018] The beneficial effects of this application are:

[0019] This application provides a gated multimodal image fusion method, apparatus, electronic device, and storage medium. By inputting each modal image into the feature extraction module of a multimodal image fusion model, the feature extraction module performs multi-layer feature extraction and differential information compensation processing to obtain multi-layer feature maps corresponding to each modal image. By performing differential information compensation processing on the feature maps, the feature information in the obtained feature maps can be made more complete, and complementary feature maps of each modality image can be extracted, reducing the differences between modalities. The third-layer feature maps corresponding to each modal image are input into the feature fusion module of the multimodal image fusion model for first-stage feature fusion to obtain an initial fused feature map. Then, the multi-layer feature maps corresponding to the first modality image, the multi-layer feature maps corresponding to the second modality image, and the initial fused feature map are input into the secondary gated fusion module of the multimodal image fusion model for multi-layer fusion to obtain the target fused image. Through secondary fusion, heterogeneous images of different modalities can be fused, resulting in a more balanced information richness in the target fused image. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of the structure of a multimodal image fusion model provided in an embodiment of this application;

[0022] Figure 2 A flowchart illustrating the first gated multimodal image fusion method provided in this application embodiment;

[0023] Figure 3 A flowchart illustrating the second gated multimodal image fusion method provided in this application embodiment;

[0024] Figure 4 A flowchart illustrating the third gated multimodal image fusion method provided in this application embodiment;

[0025] Figure 5 A structural diagram of a cross-modal differential sensing unit provided in an embodiment of this application;

[0026] Figure 6 A flowchart illustrating the fourth gated multimodal image fusion method provided in this application embodiment;

[0027] Figure 7A flowchart illustrating the fifth gated multimodal image fusion method provided in this application embodiment;

[0028] Figure 8 This is a schematic diagram of the structure of a feature fusion module provided in an embodiment of this application;

[0029] Figure 9 A flowchart illustrating the sixth gated multimodal image fusion method provided in this application embodiment;

[0030] Figure 10 A flowchart illustrating the seventh gated multimodal image fusion method provided in this application embodiment;

[0031] Figure 11 A schematic diagram of an apparatus for a gated multimodal image fusion method provided in an embodiment of this application;

[0032] Figure 12 This is a structural block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0033] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.

[0034] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0035] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.

[0036] Optionally, the gated multimodal image fusion method provided in this application embodiment can be applied to electronic devices, such as mobile phones, tablets, laptops, PDAs, desktop computers, and other terminal devices with computing power and display functions, or it can be a server. Specifically, it can be applied to applications in terminal devices, such as mobile phone apps (APPs) and computer application systems.

[0037] Figure 1 This is a schematic diagram of the structure of a multimodal image fusion model provided in an embodiment of this application, as shown below. Figure 1 As shown, the multimodal image fusion model may include: a feature extraction module, a feature fusion module, and a secondary gated fusion module.

[0038] The feature extraction module includes multi-layer convolutional units and a cross-modality differential aware fusion (CMDAF) unit. The multi-layer convolutional units can extract features from each modality image to obtain the feature map corresponding to each modality image. CMDAF can perform differential information compensation on the feature map output by the multi-layer convolutional units to obtain the feature map after differential information compensation.

[0039] Optionally, the feature fusion module adds an L2-norm strategy to the three-branch gated fusion module (GFM-T). By performing L2-norm processing on the feature map after compensation of the difference information output from the feature extraction module and then performing feature fusion, an initial fused feature map is output.

[0040] Optionally, the secondary gated fusion module can fuse the feature map corresponding to the source image (i.e., the feature map output from the feature extraction module) with the fusion features of the previous layer to obtain the target fused image. The source image can be either the first modality image or the second modality image as described below. Preprocessing the source image can yield the third modality image.

[0041] The following section will explain in detail the specific implementation process of image fusion provided in the embodiments of this application.

[0042] Figure 2 This is a flowchart illustrating the first gated multimodal image fusion method provided in this application embodiment. The execution subject of this method is as described above: electronic device. Figure 2 As shown, the method includes:

[0043] S101. Obtain multiple modal images corresponding to the target object.

[0044] The plurality of modal images may include: a first modal image, a second modal image, and a third modal image. The first modal image is as follows: Figure 1 The IMG_A and second modality images in the image are as follows: Figure 1 The IMG_B and third modality images in the image are as follows: Figure 1 IMG_F in the middle.

[0045] Specifically, for a target object, image data of the target object can be acquired using different imaging technologies or devices. This image data reflects information about the target object in different physical quantities or properties. Multimodal images can include: visible light images, infrared light images, medical images, remote sensing images, spectral images, and radar images, etc. The first modal image, the second modal image, and the third modal image are three different modal images.

[0046] S102. Input each modal image into the feature extraction module of the multimodal image fusion model. The feature extraction module performs multi-layer feature extraction processing and differential information compensation processing to obtain the multi-layer feature map corresponding to each modal image.

[0047] The multi-layer feature map may include: a first-layer feature map, a second-layer feature map, and a third-layer feature map.

[0048] For example, by performing multi-layer feature extraction and differential information compensation on each modal image through the feature extraction module, the first layer feature map, the second layer feature map and the third layer feature map corresponding to the first modal image can be obtained respectively; the first layer feature map, the second layer feature map and the third layer feature map corresponding to the second modal image; and the first layer feature map, the second layer feature map and the third layer feature map corresponding to the third modal image.

[0049] Optionally, differential information compensation can make the information in the extracted feature maps of each modality more complete.

[0050] S103. Input the third-layer feature map corresponding to each modality image into the feature fusion module in the multimodal image fusion model for feature fusion to obtain the initial fused feature map.

[0051] Optionally, in the feature fusion module, the third-layer feature maps corresponding to each modality image can be fused. Specifically, the third-layer feature maps corresponding to the first modality image, the second modality image, and the third modality image can be fused to obtain an initial fused feature map, which contains the features of the third-layer feature maps of each modality image.

[0052] S104. Input the multi-layer feature map corresponding to the first modality image, the multi-layer feature map corresponding to the second modality image, and the initial fusion feature map into the secondary gated fusion module in the multi-modal image fusion model for multi-layer fusion to obtain the target fused image.

[0053] Optionally, the secondary gated fusion module may contain three fusion layers. Multi-layer fusion can be an upsampling process, and the aforementioned feature extraction can be a downsampling process. Multi-layer fusion can refer to further fusing the feature map obtained from the convolutional layer in the same layer as the fusion layer and the fusion feature map obtained from the previous layer to obtain the target fused image. The fusion feature map obtained from the layer above the first fusion layer refers to the initial fusion feature map obtained from the feature fusion module in S103. The layer above the second fusion layer is the first fusion layer, and the layer above the third fusion layer is the second fusion layer. The target fused image is then output from the third fusion layer.

[0054] Specifically, the first-layer feature map, the second-layer feature map, and the third-layer feature map corresponding to the first modality image, the first-layer feature map, the second-layer feature map, and the third-layer feature map corresponding to the second modality image, and the initial fusion feature map can be input into the secondary gated fusion module for multi-layer fusion to obtain the target fused image.

[0055] In this embodiment, by inputting each modal image into the feature extraction module of the multimodal image fusion model, the feature extraction module performs multi-layer feature extraction processing and differential information compensation processing to obtain multi-layer feature maps corresponding to each modal image. By performing differential information compensation processing on the feature maps, the feature information in the obtained feature maps can be made more complete, and complementary feature maps of each modal image can be extracted, reducing the differences between modalities. The third-layer feature map corresponding to each modal image is input into the feature fusion module in the multimodal image fusion model for first-stage feature fusion to obtain an initial fused feature map. Then, the multi-layer feature map corresponding to the first modal image, the multi-layer feature map corresponding to the second modal image, and the initial fused feature map are input into the secondary gated fusion module in the multimodal image fusion model for multi-layer fusion to obtain the target fused image. Through secondary fusion, heterogeneous images of different modalities can be fused, and the information richness of the obtained target fused image can be more balanced.

[0056] Figure 3 This is a flowchart illustrating the second gated multimodal image fusion method provided in the embodiments of this application, as shown below. Figure 3 As shown, obtaining multiple modal images corresponding to the target object in S101 above may include:

[0057] S201. Obtain a first initial modal image and a second initial modal image of the target object, and use the first initial modal image as the first modal image and the second initial modal image as the second modal image.

[0058] Optionally, if only a first initial modal image and a second initial modal image are available for the target object, then the first initial modal image is used as the first modal image, and the second initial modal image is used as the second modal image.

[0059] S202. Input the first modal image and the second modal image into the dual-branch gated fusion module. The dual-branch gated fusion module extracts the global features of the first modal image and the global features of the second modal image, and generates the third modal image based on the global features of the first modal image and the global features of the second modal image.

[0060] Specifically, such as Figure 1 As shown in the preprocessing block diagram, the first modality image can be input into one branch of the dual-branch gated fusion module (GFM_D) to extract global features of the first modality image. The second modality image is input into the other branch of the dual-branch gated fusion module to extract global features of the second modality image. Then, the global features of the first modality image IMG_A and the global features of the second modality image IMG_B are fused to generate a third modality image IMG_F. The first modality image IMG_A, the second modality image IMG_B, and the generated third modality image IMG_F are then input into the multimodal image fusion model.

[0061] Optionally, if there are a first initial modal image, a second initial modal image, and a third initial modal image for the target object, the first initial modal image is used as the first modal image, the second initial modal image is used as the second modal image, and the third initial modal image is used as the third modal image and input into the multimodal image fusion model, so that the third modal image does not need to be generated through the dual-branch gating fusion module.

[0062] Figure 4 A flowchart illustrating the third gated multimodal image fusion method provided in this application embodiment is shown below. Figure 4 As shown, in S102 above, the feature extraction module performs multi-layer feature extraction processing and differential information compensation processing to obtain multi-layer feature maps corresponding to each modality image, which may include:

[0063] S301. Input each modality image into a multi-layer convolutional unit for feature extraction to obtain the first initial feature map, the second initial feature map, and the third initial feature map corresponding to each modality image.

[0064] The multi-layer convolutional unit contains a first convolutional layer, a second convolutional layer, and a third convolutional layer. These three convolutional layers extract features from various modalities of the image, and the size of the feature maps extracted by each convolutional layer is different. The processing of the multi-layer convolutional unit is a downsampling process. The output of the first convolutional layer is input into the second convolutional layer, and the output of the second convolutional layer is input into the third convolutional layer.

[0065] Specifically, the first convolutional layer extracts features from each modality image to obtain a first-layer initial feature map corresponding to each modality image, and then outputs the first-layer initial feature map corresponding to each modality image. The first-layer initial feature map corresponding to each modality image is then input into the second convolutional layer for further feature extraction to obtain a second-layer initial feature map corresponding to each modality image, and then outputs the second-layer initial feature map corresponding to each modality image. The second-layer initial feature map corresponding to each modality image is then input into the third convolutional layer for feature extraction to obtain a third-layer initial feature map corresponding to each modality image, and then outputs the third-layer initial feature map corresponding to each modality image. The first, second, and third-layer initial feature maps have different sizes.

[0066] S302. Input the first-layer initial feature map corresponding to each modal image into the cross-modal differential sensing unit for differential information compensation to obtain the first-layer feature map corresponding to each modal image.

[0067] Specifically, the cross-modal differential sensing unit (CMDAF) can use a preset method to compensate for the differential information of the first-layer initial feature map corresponding to each modality. Specifically, it can compensate for the first-layer initial feature map corresponding to each modality image based on the information differences between each pair of the first-layer initial feature maps corresponding to each modality image.

[0068] S303. Input the second-layer initial feature map corresponding to each modal image into the cross-modal differential sensing unit for differential information compensation to obtain the second-layer feature map corresponding to each modal image.

[0069] Specifically, the cross-modal differential sensing unit can use a preset method to compensate for the differential information of the second-layer initial feature map corresponding to each modality. Specifically, it can compensate for the second initial feature map corresponding to each modality image based on the information differences between each pair of the second-layer initial feature maps corresponding to each modality image.

[0070] S304. Input the initial feature map of the third layer corresponding to each modal image into the cross-modal differential sensing unit for differential information compensation to obtain the feature map of the third layer corresponding to each modal image.

[0071] Specifically, the cross-modal differential sensing unit can use a preset method to compensate for the differential information of the third-layer initial feature map corresponding to each modality. Specifically, it can compensate for the third-layer initial feature map corresponding to each modality image based on the information differences between each pair of the third-layer initial feature maps corresponding to each modality image.

[0072] In this embodiment, the feature information between the third-layer feature maps corresponding to each modality image can be complementary, making the feature information of the third-layer feature maps corresponding to each modality more complete.

[0073] Figure 5 A structural diagram of a cross-modal differential sensing unit provided in an embodiment of this application is shown below. Figure 5 As shown, the cross-modal differential sensing unit can include three channels: a first channel, a second channel, and a third channel. Differential information compensation is performed on the third initial feature maps corresponding to the modal images input to the three channels. The original CMDAF is a dual-channel mode; this application modifies the original CMDAF to a three-channel mode.

[0074] It is worth noting that the process of performing differential information compensation on the first-layer initial feature map and the second-layer initial feature map is the same as the process of performing differential information compensation on the third-layer initial feature map. This can be referred to as... Figure 5 The execution process is not described in detail here.

[0075] Figure 6 A flowchart illustrating the fourth gated multimodal image fusion method provided in this application embodiment is shown below. Figure 6 As shown, in S304 above, the initial feature map of the third layer corresponding to each modal image is input to the cross-modal differential sensing unit for differential information compensation to obtain the third layer feature map corresponding to each modal image, which may include:

[0076] S401. Input the third-layer initial feature map corresponding to each modal image into the first channel, determine the first difference information between the third-layer initial feature map corresponding to the first modal image and the third-layer initial feature map corresponding to other modal images, and fuse the third-layer initial feature map corresponding to the first modal image according to the first difference information to obtain the third-layer feature map corresponding to the first modal image.

[0077] Specifically, if the third initial feature map corresponding to the first mode is The initial feature map of the third layer corresponding to the second modality image is: The initial feature map of the third layer corresponding to the third modality image is: In the first channel, the first difference information between the third-layer initial feature map corresponding to the first modality image and the third-layer initial feature map corresponding to the second modality image can be determined. And the first difference information between the third layer feature map corresponding to the first modality image and the third layer initial feature map corresponding to the third modality image.

[0078] Optionally, based on the first difference information First difference information and the third layer initial feature map corresponding to the first modality image The fusion process is performed to obtain the third-layer feature map corresponding to the first modality image. Specifically, it can be obtained according to the following formula (I).

[0079]

[0080] Where ⊕ represents element-wise summation, For channel-level multiplication, δ(·) is the sigmoid function, and GAP(·) is the global average pooling. The resulting third-layer feature map corresponds to the first-modality image. The first modality image and the second modality image are feature maps after mutual compensation between the first modality image and the third modality image.

[0081] S402. Input the third-layer initial feature map corresponding to each modal image into the second channel, determine the second difference information between the third-layer initial feature map corresponding to the second modal image and the third-layer initial feature map corresponding to other modal images, and fuse the third-layer initial feature map corresponding to the second modal image according to the second difference information to obtain the third-layer feature map corresponding to the second modal image.

[0082] Specifically, in the second channel, the second difference information between the third-layer initial feature map corresponding to the second modality image and the third-layer initial feature map corresponding to the first modality image can be determined. And the second difference information between the third layer feature map corresponding to the second modality image and the third layer initial feature map corresponding to the third modality image.

[0083] Optionally, based on the second difference information Second difference information and the third layer initial feature map corresponding to the second modality image The fusion process is performed to obtain the third-layer feature map corresponding to the second modality image. Specifically, it can be obtained according to the following formula (II).

[0084]

[0085] Where ⊕ represents element-wise summation, For channel-level multiplication, δ(·) is the sigmoid function, and GAP(·) is the global average pooling. The resulting third-layer feature map corresponds to the second-modality image. The feature map is the result of mutual compensation between the second modality image, the first modality image, and the third modality image.

[0086] S403. Input the third-layer initial feature map corresponding to each modal image into the third channel, determine the third difference information between the third-layer initial feature map corresponding to the third modal image and the third-layer initial feature map corresponding to other modal images, and fuse each third difference information with the third-layer initial feature map corresponding to the third modal image to obtain the third-layer feature map corresponding to the third modal image.

[0087] Specifically, in the third channel, the third difference information between the third-layer initial feature map corresponding to the third modality image and the third-layer initial feature map corresponding to the first modality image can be determined. And the third difference information between the third layer feature map corresponding to the third modality image and the third layer initial feature map corresponding to the second modality image.

[0088] Optionally, based on the third difference information Third difference information and the initial feature map of the third layer corresponding to the third modality image. The fusion process is performed to obtain the third-layer feature map corresponding to the third modality image. Specifically, it can be obtained according to the following formula (iii).

[0089]

[0090] Where ⊕ represents element-wise summation, For channel-level multiplication, δ(·) is the sigmoid function, and GAP(·) is the global average pooling. The resulting third-modality image corresponds to the third-layer feature map. The feature map is the result of mutual compensation between the third modality image and the first modality image, and between the third modality image and the second modality image.

[0091] In this embodiment, the improved cross-modal differential sensing fusion unit is used to compensate for the differential information between different modalities. The third-layer feature map corresponding to each modal image takes into account the differences between the modalities, making the information in the feature map output by the feature extraction module more comprehensive and complete.

[0092] Figure 7 A flowchart illustrating the fifth gated multimodal image fusion method provided in this application embodiment is shown below. Figure 7As shown, in S103 above, the third-layer feature map corresponding to each modality image is input into the feature fusion module in the multimodal image fusion model for feature fusion to obtain the initial fused feature map, which may include:

[0093] S501. Input the third-layer feature map corresponding to each modal image into the feature fusion module in the multimodal image fusion model. The feature fusion module squares the pixel values ​​in the third-layer feature map corresponding to each modal image to obtain the processed third-layer feature map.

[0094] Specifically, the third-layer feature map corresponding to each modality image can be squared using the following formula (iv).

[0095]

[0096] Where S∈(A,B,F), then S is the identifier of each modality image, * represents convolution processing, L2(·) indicates the calculation of the mean square value of L2, and K s This refers to the third-layer feature map after processing, i.e., the third-layer feature map corresponding to each processed modal image. The third-layer feature map corresponding to each modal image can be processed using formula (iv), i.e., the L2-norm strategy.

[0097] S502. Based on each processed third-layer feature map and each third-layer feature map, obtain each enhanced third-layer feature map.

[0098] Specifically, it can be obtained from the following formulas (V), (VI) and (VII).

[0099]

[0100] in, This is the enhanced third-layer feature map corresponding to the first modality image. This is the enhanced third-layer feature map corresponding to the second modality image. This is the enhanced third-layer feature map corresponding to the third modality image.

[0101] S503. Perform feature fusion on each enhanced third-layer feature map to obtain an initial fused feature map.

[0102] Specifically, the enhanced third-layer feature map corresponding to the first modality image can be used first. Enhanced third-layer features corresponding to the second modality image Figure 2 Feature fusion is performed to obtain the fusion result, and then the fusion result is compared with the enhanced third-layer feature map corresponding to the third modality image. Feature fusion is performed to obtain the initial fused feature map F. cSpecifically, such as Figure 8 The fusion diagram shown is as follows: Figure 8 This is a schematic diagram of the structure of a feature fusion module provided in an embodiment of this application.

[0103] In this embodiment, by first performing L2-norm processing on the third feature map corresponding to each modality image, and then performing feature fusion on each feature map after L2-norm processing, the intensity information of pixels can be enhanced, the contrast of pixel intensity can be effectively enhanced, and thus the visual effect of the fused image can be improved.

[0104] Optionally, the secondary gated fusion module may include a first fusion layer, a second fusion layer, and a third fusion layer. This secondary gated fusion module is a three-branch input module (GFM_T). Figure 10 This is a schematic diagram of the structure of a secondary gating fusion module provided in an embodiment of this application.

[0105] Figure 9 A flowchart illustrating the sixth gated multimodal image fusion method provided in this application embodiment is shown below. Figure 9 As shown, in S104 above, the multi-layer feature map corresponding to the first modality image, the multi-layer feature map corresponding to the second modality image, and the initial fusion feature map are input into the secondary gated fusion module in the multi-modal image fusion model for multi-layer fusion to obtain the target fused image. This can include:

[0106] S601. Input the third-layer feature map corresponding to the first modality image, the third-layer feature map corresponding to the second modality image, and the initial fusion feature map into the first fusion layer for feature fusion to obtain the first fusion feature map.

[0107] Specifically, the third layer feature map corresponding to the first modality image The third layer feature map corresponding to the second modality image and the initial fusion feature map F c All inputs are fed into the first fusion layer and feature fusion is performed using a preset method to obtain the first fused feature map.

[0108] S602. Input the second feature map corresponding to the first modality image, the second feature map corresponding to the second modality image, and the first fusion feature map into the second fusion layer for feature fusion to obtain the second fusion feature map.

[0109] Specifically, the second layer feature map corresponding to the first modality image The second feature map corresponding to the second modality image Both the first fused feature map F1 and the first fused feature map F2 are input into the first fusion layer and feature fusion is performed using a preset method to obtain the second fused feature map F2.

[0110] S603. Input the first layer feature map corresponding to the first modality image, the first feature map corresponding to the second modality image, and the second fusion feature map into the third fusion layer for feature fusion to obtain the target fusion image.

[0111] Specifically, the first layer feature map corresponding to the first modality image The first feature map corresponding to the second modality image Both the second fused feature map F2 and the first fusion layer are input to perform feature fusion using a preset method to obtain the target fused image F3.

[0112] In this embodiment, in each fusion layer, the feature maps of the first modality image and the second modality image extracted by the feature extraction layer of the downsampling process in the same layer as the fusion layer are fused together, and then further fused with the fusion feature map obtained from the previous layer. This can make the information richness in the target fusion image more balanced, and combine the semantic differences between the images of different modalities, so that the images of different modalities can be effectively fused, making the target fusion image more consistent with the real features of the target object.

[0113] Figure 10 A flowchart illustrating the seventh gated multimodal image fusion method provided in this application embodiment is shown below. Figure 10 As shown, in S601, the third-layer feature map corresponding to the first modality image, the third-layer feature map corresponding to the second modality image, and the initial fused feature map are input into the first fusion layer for feature fusion to obtain the first fused feature map, which may include:

[0114] S701. Determine the first weight value and the second weight value based on the third layer feature map corresponding to the first modality image and the third feature map corresponding to the second modality image.

[0115] Specifically, the first weight value The second weight value is 1-G. Here, the first weight value G is the third-layer feature map corresponding to the first modality image. The weight values ​​are: the second weight value 1-G is the third layer feature map corresponding to the second modality image. The weight values, Sig(·) is the Sigmoid function, Con 1×1 It is a convolution with a 1×1 kernel.

[0116] S702. Perform depthwise separable convolution on the initial fused feature map to obtain the first depthwise convolution map.

[0117] Specifically, the first depthwise convolutional map is DS_C(F c ).

[0118] S703. Calculate the product of the third feature map corresponding to the first modality image and the first weight value to obtain the first product map.

[0119] Specifically, the first product graph is

[0120] S704. Calculate the product of the third feature map corresponding to the second modality image and the second weight value to obtain the second product map.

[0121] Specifically, the second product graph is

[0122] S705. Merge the first product graph and the second product graph to obtain the first merged graph.

[0123] Specifically, the first fusion diagram is as follows:

[0124] S706. The first fused image and the first depthwise convolutional image are fused to obtain the second fused image, and the second fused image is subjected to two depthwise separable convolutional processes to obtain the first fused feature map.

[0125] Specifically, the second fusion diagram is as follows: The first fusion diagram is specifically shown in the following formula (eight).

[0126]

[0127] Where DS_C(·) is depthwise separable convolution, Cat(·) is fusion processing, and F1 is the first fused feature map.

[0128] Optionally, the process of obtaining the second fusion feature map and the target fusion image is similar to the process of obtaining the first fusion feature map. Specifically, the second fusion feature map can be obtained according to the following formula (ix).

[0129]

[0130] in, G represents the weight value of the second-layer feature map corresponding to the first modality image, 1-G represents the weight value of the second-layer feature map corresponding to the second modality image, and F2 represents the second fused feature map. This is the second layer feature map corresponding to the first modality image. F1 is the second feature map corresponding to the second modality image, and F1 is the first fused feature map.

[0131] Alternatively, the target fused image can be obtained according to the following formula (x).

[0132]

[0133] in, G represents the weight value of the first layer feature map corresponding to the first modality image, 1-G represents the weight value of the first layer feature map corresponding to the second modality image, F2 represents the second fused feature map, and F3 represents the target fused image. The first layer feature map corresponding to the first modality image is as follows: This is the first layer feature map corresponding to the second modality image.

[0134] Figure 11 A schematic diagram of an apparatus for a gated multimodal image fusion method provided in an embodiment of this application is shown below. Figure 11 As shown, the device includes:

[0135] The acquisition module 801 is used to acquire multiple modal images corresponding to the target object, the multiple modal images including: a first modal image, a second modal image and a third modal image;

[0136] The feature extraction module 802 is used to input each of the modal images into the feature extraction module of the multimodal image fusion model. The feature extraction module performs multi-layer feature extraction processing and differential information compensation processing to obtain multi-layer feature maps corresponding to each of the modal images. The multi-layer feature maps include a first-layer feature map, a second-layer feature map, and a third-layer feature map.

[0137] The first fusion module 803 is used to input the third-layer feature map corresponding to each modal image into the feature fusion module in the multimodal image fusion model for feature fusion to obtain an initial fusion feature map;

[0138] The second fusion module 804 is used to input the multi-layer feature map corresponding to the first modality image, the multi-layer feature map corresponding to the second modality image, and the initial fusion feature map into the secondary gated fusion module in the multi-modal image fusion model for multi-layer fusion to obtain the target fused image.

[0139] Optionally, the acquisition module 801 is specifically used for:

[0140] Acquire a first initial modal image and a second initial modal image of the target object;

[0141] The first initial modal image is used as the first modal image, and the second initial modal image is used as the second modal image;

[0142] The first modal image and the second modal image are input into the dual-branch gated fusion module, which extracts the global features of the first modal image and the global features of the second modal image, and generates the third modal image based on the global features of the first modal image and the global features of the second modal image.

[0143] Optionally, the feature extraction module includes: a multi-layer convolutional unit and a cross-modal differential sensing fusion unit;

[0144] The feature extraction module 802 is specifically used for:

[0145] Each modal image is input into the multi-layer convolutional unit for feature extraction, resulting in a first-layer initial feature map, a second-layer initial feature map, and a third-layer initial feature map corresponding to each modal image.

[0146] The first-layer initial feature map corresponding to each modal image is input into the cross-modal differential sensing fusion unit for differential information compensation to obtain the first-layer feature map corresponding to each modal image;

[0147] The second-layer initial feature map corresponding to each modal image is input into the cross-modal differential sensing fusion unit for differential information compensation to obtain the second-layer feature map corresponding to each modal image;

[0148] The initial third-layer feature map corresponding to each modal image is input into the cross-modal differential sensing fusion unit for differential information compensation to obtain the third-layer feature map corresponding to each modal image.

[0149] Optionally, the cross-modal differential sensing fusion unit includes: a first channel, a second channel, and a third channel;

[0150] The feature extraction module 802 is specifically used for:

[0151] The third-layer initial feature map corresponding to each modal image is input into the first channel to determine the first difference information between the third-layer initial feature map corresponding to the first modal image and the third-layer initial feature map corresponding to other modal images. The first difference information is then fused with the third-layer initial feature map corresponding to the first modal image to obtain the third-layer feature map corresponding to the first modal image.

[0152] The third-layer initial feature map corresponding to each modal image is input into the second channel to determine the second difference information between the third-layer initial feature map corresponding to the second modal image and the third-layer initial feature map corresponding to other modal images. The second difference information is then fused with the third-layer initial feature map corresponding to the second modal image to obtain the third-layer feature map corresponding to the second modal image.

[0153] The third-layer initial feature map corresponding to each modal image is input into the third channel to determine the third difference information between the third-layer initial feature map corresponding to the third modal image and the third-layer initial feature map corresponding to other modal images. The third difference information is then fused with the third-layer initial feature map corresponding to the third modal image to obtain the third-layer feature map corresponding to the third modal image.

[0154] Optionally, the first fusion module 803 is specifically used for:

[0155] The third-layer feature map corresponding to each modal image is input into the feature fusion module in the multimodal image fusion model. The feature fusion module squares the pixel values ​​in the third-layer feature map corresponding to each modal image to obtain the processed third-layer feature map.

[0156] Based on the processed third-layer feature maps and the third-layer feature maps described above, each enhanced third-layer feature map is obtained.

[0157] The enhanced third-layer feature maps are fused to obtain the initial fused feature map.

[0158] Optionally, the secondary gating fusion module includes: a first fusion layer, a second fusion layer, and a third fusion layer;

[0159] The second fusion module 804 is specifically used for:

[0160] The third-layer feature map corresponding to the first modality image, the third-layer feature map corresponding to the second modality image, and the initial fused feature map are input into the first fusion layer for feature fusion to obtain the first fused feature map;

[0161] The second feature map corresponding to the first modality image, the second feature map corresponding to the second modality image, and the first fused feature map are input into the second fusion layer for feature fusion to obtain the second fused feature map;

[0162] The first feature map corresponding to the first modality image, the first feature map corresponding to the second modality image, and the second fused feature map are input into the third fusion layer for feature fusion to obtain the target fused image.

[0163] Optionally, the second fusion module 804 is specifically used for:

[0164] Based on the third-layer feature map corresponding to the first modality image and the third feature map corresponding to the second modality image, determine the first weight value and the second weight value;

[0165] The initial fused feature map is subjected to depthwise separable convolution to obtain a first depthwise convolution map;

[0166] Calculate the product of the third feature map corresponding to the first modality image and the first weight value to obtain the first product map;

[0167] Calculate the product of the third feature map corresponding to the second modality image and the second weight value to obtain the second product map;

[0168] The first product graph and the second product graph are merged to obtain the first fused graph;

[0169] The first fused image and the first depthwise convolutional image are fused to obtain a second fused image, and the second fused image is subjected to two depthwise separable convolutional processes to obtain the first fused feature map.

[0170] Figure 12 This is a structural block diagram of an electronic device 900 provided in an embodiment of this application. (See diagram below.) Figure 12 As shown, the electronic device may include: a processor 901 and a memory 902.

[0171] Optionally, a bus 903 may also be included, wherein the memory 902 is used to store machine-readable instructions executable by the processor 901. When the electronic device 900 is running, the processor 901 and the memory 902 communicate via the bus 903. When the machine-readable instructions are executed by the processor 901, the method steps in the above method embodiments are performed.

[0172] This application also provides a computer-readable storage medium storing a computer program, which, when run by a processor, executes the method steps described in the above-described gated multimodal image fusion method embodiments.

[0173] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces; the indirect coupling or communication connection of devices or modules can be electrical, mechanical, or other forms.

[0174] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.

[0175] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A gated multimodal image fusion method, characterized in that, The method includes: Acquire multiple modal images corresponding to the target object, the multiple modal images including: a first modal image, a second modal image, and a third modal image; Each modal image is input into the feature extraction module of the multimodal image fusion model. The feature extraction module performs multi-layer feature extraction processing and differential information compensation processing to obtain multi-layer feature maps corresponding to each modal image. The multi-layer feature maps include a first-layer feature map, a second-layer feature map, and a third-layer feature map. The third-layer feature map corresponding to each modal image is input into the feature fusion module in the multimodal image fusion model for feature fusion to obtain an initial fused feature map; The multi-layer feature map corresponding to the first modality image, the multi-layer feature map corresponding to the second modality image, and the initial fusion feature map are input into the secondary gated fusion module in the multi-modal image fusion model for multi-layer fusion to obtain the target fused image; The secondary gating fusion module includes: a first fusion layer, a second fusion layer, and a third fusion layer; The step of inputting the multi-layer feature map corresponding to the first modality image, the multi-layer feature map corresponding to the second modality image, and the initial fusion feature map into the secondary gated fusion module in the multi-modal image fusion model for multi-layer fusion to obtain the target fused image includes: The third-layer feature map corresponding to the first modality image, the third-layer feature map corresponding to the second modality image, and the initial fused feature map are input into the first fusion layer for feature fusion to obtain the first fused feature map; The second feature map corresponding to the first modality image, the second feature map corresponding to the second modality image, and the first fused feature map are input into the second fusion layer for feature fusion to obtain the second fused feature map; The first feature map corresponding to the first modality image, the first feature map corresponding to the second modality image, and the second fused feature map are input into the third fusion layer for feature fusion to obtain the target fused image.

2. The gated multimodal image fusion method according to claim 1, characterized in that, The acquisition of multiple modal images corresponding to the target object includes: Acquire a first initial modal image and a second initial modal image of the target object; The first initial modal image is used as the first modal image, and the second initial modal image is used as the second modal image; The first modal image and the second modal image are input into the dual-branch gated fusion module, which extracts the global features of the first modal image and the global features of the second modal image, and generates the third modal image based on the global features of the first modal image and the global features of the second modal image.

3. The gated multimodal image fusion method according to claim 1, characterized in that, The feature extraction module includes: a multi-layer convolutional unit and a cross-modal differential sensing fusion unit; The process of performing multi-layer feature extraction and differential information compensation by the feature extraction module to obtain multi-layer feature maps corresponding to each modality image includes: Each modal image is input into the multi-layer convolutional unit for feature extraction, resulting in a first-layer initial feature map, a second-layer initial feature map, and a third-layer initial feature map corresponding to each modal image. The first-layer initial feature map corresponding to each modal image is input into the cross-modal differential sensing fusion unit for differential information compensation to obtain the first-layer feature map corresponding to each modal image; The second-layer initial feature map corresponding to each modal image is input into the cross-modal differential sensing fusion unit for differential information compensation to obtain the second-layer feature map corresponding to each modal image; The initial third-layer feature map corresponding to each modal image is input into the cross-modal differential sensing fusion unit for differential information compensation to obtain the third-layer feature map corresponding to each modal image.

4. The gated multimodal image fusion method according to claim 3, characterized in that, The cross-modal differential sensing fusion unit includes: a first channel, a second channel, and a third channel; The step of inputting the third-layer initial feature map corresponding to each modal image into the cross-modal differential sensing fusion unit for differential information compensation to obtain the third-layer feature map corresponding to each modal image includes: The third-layer initial feature map corresponding to each modal image is input into the first channel to determine the first difference information between the third-layer initial feature map corresponding to the first modal image and the third-layer initial feature map corresponding to other modal images. The first difference information is then fused with the third-layer initial feature map corresponding to the first modal image to obtain the third-layer feature map corresponding to the first modal image. The third-layer initial feature map corresponding to each modal image is input into the second channel to determine the second difference information between the third-layer initial feature map corresponding to the second modal image and the third-layer initial feature map corresponding to other modal images. The second difference information is then fused with the third-layer initial feature map corresponding to the second modal image to obtain the third-layer feature map corresponding to the second modal image. The third-layer initial feature map corresponding to each modal image is input into the third channel to determine the third difference information between the third-layer initial feature map corresponding to the third modal image and the third-layer initial feature map corresponding to other modal images. The third difference information is then fused with the third-layer initial feature map corresponding to the third modal image to obtain the third-layer feature map corresponding to the third modal image.

5. The gated multimodal image fusion method according to claim 1, characterized in that, The step of inputting the third-layer feature map corresponding to each modal image into the feature fusion module of the multimodal image fusion model for feature fusion to obtain an initial fused feature map includes: The third-layer feature map corresponding to each modal image is input into the feature fusion module in the multimodal image fusion model. The feature fusion module squares the pixel values ​​in the third-layer feature map corresponding to each modal image to obtain the processed third-layer feature map. Based on the processed third-layer feature maps and the third-layer feature maps described above, each enhanced third-layer feature map is obtained. The enhanced third-layer feature maps are fused to obtain the initial fused feature map.

6. The gated multimodal image fusion method according to claim 5, characterized in that, The step of inputting the third-layer feature map corresponding to the first modality image, the third-layer feature map corresponding to the second modality image, and the initial fused feature map into the first fusion layer for feature fusion to obtain the first fused feature map includes: Based on the third-layer feature map corresponding to the first modality image and the third feature map corresponding to the second modality image, determine the first weight value and the second weight value; The initial fused feature map is subjected to depthwise separable convolution to obtain a first depthwise convolution map; Calculate the product of the third feature map corresponding to the first modality image and the first weight value to obtain the first product map; Calculate the product of the third feature map corresponding to the second modality image and the second weight value to obtain the second product map; The first product graph and the second product graph are merged to obtain the first fused graph; The first fused image and the first depthwise convolutional image are fused to obtain a second fused image, and the second fused image is subjected to two depthwise separable convolutional processes to obtain the first fused feature map.

7. An image fusion apparatus, characterized in that, include: The acquisition module is used to acquire multiple modal images corresponding to the target object, the multiple modal images including: a first modal image, a second modal image and a third modal image; The feature extraction module is used to input each of the modal images into the feature extraction module of the multimodal image fusion model. The feature extraction module performs multi-layer feature extraction processing and differential information compensation processing to obtain multi-layer feature maps corresponding to each of the modal images. The multi-layer feature maps include a first-layer feature map, a second-layer feature map, and a third-layer feature map. The first fusion module is used to input the third-layer feature map corresponding to each modal image into the feature fusion module in the multimodal image fusion model for feature fusion to obtain an initial fused feature map; The second fusion module is used to input the multi-layer feature map corresponding to the first modality image, the multi-layer feature map corresponding to the second modality image, and the initial fusion feature map into the secondary gated fusion module in the multi-modal image fusion model for multi-layer fusion to obtain the target fused image; The secondary gating fusion module includes: a first fusion layer, a second fusion layer, and a third fusion layer; The second fusion module is specifically used for: The third-layer feature map corresponding to the first modality image, the third-layer feature map corresponding to the second modality image, and the initial fused feature map are input into the first fusion layer for feature fusion to obtain the first fused feature map; The second feature map corresponding to the first modality image, the second feature map corresponding to the second modality image, and the first fused feature map are input into the second fusion layer for feature fusion to obtain the second fused feature map; The first feature map corresponding to the first modality image, the first feature map corresponding to the second modality image, and the second fused feature map are input into the third fusion layer for feature fusion to obtain the target fused image.

8. An electronic device, characterized in that, The method includes a memory and a processor, wherein the memory stores a computer program executable by the processor, and the processor executes the computer program to implement the steps of the gated multimodal image fusion method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the gated multimodal image fusion method as described in any one of claims 1-6.