A multimodal image fusion method based on joint learning of multiple scene features
By constructing a cross-modal knowledge-enhancing network, spectrum and spatial attention optimization network, and edge-guided learning network, the shortcomings of the existing multimodal image fusion methods in feature extraction and fusion results are solved, and the effective combination of multimodal image features is achieved and significant information enhancement of the results of the fusion.
Patent Information
- Application Number
- CN202310991670.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-08-08
AI Technical Summary
The existing multimodal image fusion method ignores the differences in sensor imaging mechanisms. General feature extraction cannot effectively guide the network to obtain the advantages of multi-fusion tasks, and the fusion results are difficult to preserve the foreground significant object and background texture details of the source image at the same time.
A multimodal image fusion method for joint learning of multi-scene features is proposed. By constructing a cross-modal knowledge-enhancing network, an optimized network based on spectrum attention and spatial attention, and an edge-guided learning network, it realizes effective combination of different modal features and significant information enhancement of the results of fusion.
The effective combination of different modal features is achieved, significant information of the fusion result is enhanced, the foreground targets and background details of the source image can be preserved at the same time, and the effect of multimodal image fusion is improved.
Smart Images

Figure CN117011208B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image fusion, and in particular to a multi-modal image fusion method for joint learning of multi-scene features. Background Art
[0002] Multimodal image fusion MMIF can fuse the information obtained by different sensors into one image to fully reflect the real information of the scene. Typical MMIF includes infrared visible light image fusion IVIF and medical image fusion MIF. IVIF focuses on the advantages of the source images, and the fusion results not only highlight the significant targets, but also contain rich texture details. Similarly, the goal of MIF is to combine structural and functional information from multiple different types of images to help people diagnose diseases quickly and accurately. Common medical fusion tasks include CT and MRI image fusion, PET and MRI image fusion, and SPECT and MRI image fusion.
[0003] Multimodal image fusion (MMIF) methods are generally divided into three categories, namely, autoencoders (AE), convolutional neural networks (CNN), and generative adversarial networks (GAN); AE-based methods require training autoencoders on large image datasets and then using the trained encoders for feature extraction. In most cases, the feature combination of the encoder relies on a manually designed fusion strategy. To avoid this limitation, researchers have developed an end-to-end image fusion framework based on CNN, which uses an implicit feature learning model to build a complex network structure and loss function. In addition, GAN is introduced into the image fusion task, which can effectively learn the data distribution from the source image. However, these methods all have some obvious shortcomings: (1) Most unified fusion methods ignore the differences in sensor imaging mechanisms and do not distinguish between single-modal and multi-modal tasks; (2) General feature extraction is used to process information of different modalities and cannot effectively guide the network to obtain the advantages of multi-fusion tasks; (3) The fusion result cannot simultaneously retain the foreground salient objects and background texture details of the source image. Summary of the invention
[0004] The purpose of the present invention is to propose a multimodal image fusion method for joint learning of multi-scene features, which realizes the effective combination of different modal features and enhances the significant information of the fusion result.
[0005] To achieve the above objectives, the present application proposes a multi-modal image fusion method for joint learning of multi-scene features, including:
[0006] Construct a cross-modal knowledge enhancement network to enhance the consistency and difference features of multimodal images;
[0007] Construct an optimized network based on spectral attention and spatial attention to enhance the important contextual information in the fusion results;
[0008] Construct an edge-guided learning network to learn contour information from source images;
[0009] The multimodal image fusion model is obtained through the cross-modal knowledge enhancement network, the optimization network based on spectral attention and spatial attention, and the edge-guided learning network:
[0010]
[0011] Stu Ψ =Ψ(x,y;ω Ψ ),u Φ =Φ(x;ω Φ ),u γ =γ(y;ω γ )
[0012] Among them, L f is the structural similarity loss, and F is a loss with learnable parameters ω f image reconstruction network; λ is a trade-off parameter; x, y are input images of different modalities; u Ψ ,u Φ ,u γ Represents the output results of the three networks; the model T consists of three networks, namely the cross-modal knowledge reinforcement network Ψ, the optimization network γ based on spectral attention and spatial attention, and the edge-guided learning network Φ; L t is a composite loss function, including L Ψ , L γ and L Φ ;ω t are trainable parameters of the model T.
[0013] Furthermore, the cross-modal knowledge enhancement network makes full use of the complementary information of the source image to achieve feature aggregation. The optimization process of the network is defined as:
[0014]
[0015] Among them, ω Ψ Represents the trainable parameters of the network; L Ψ is the loss function.
[0016] Furthermore, the cross-modal knowledge reinforcement network constructs a multi-path correction strategy to strengthen the feature relationship between source images, specifically: and are initial feature maps of the same dimension, which are multiplied element by element to generate intermediate feature maps Therefore, F MDefined as: In order to maintain the consistency of features, the intermediate feature map F M To enhance and The process is defined as follows:
[0017]
[0018] Among them, Conv3 is a convolution operation, and its activation function is PReLU; F a The information aggregation in the global feature relationship is realized; for the difference of features, the subtraction operation of feature mapping is introduced to obtain and The joint differential feature F s :
[0019]
[0020] Finally, F a and F s Adding together, the consistency and differences of cross-modal information are unified:
[0021] u Ψ =Conv3(F a +F s ).
[0022] Furthermore, the initial feature map and The acquisition method is: put two source images x and y of different modalities into convolution blocks with the same structure, respectively, and the convolution blocks include 3×3 convolution layers, BatchNorm layers and RELU activation functions; obtain the initial feature map through the convolution blocks and
[0023] Furthermore, the optimization network includes a spectral attention optimization subnetwork and a spatial attention optimization subnetwork, specifically:
[0024]
[0025] Among them, ω γ is the trainable parameter of the network γ; y represents the input image, which includes infrared images, CT images, PET images and SPECT images; after entering the optimization network, the source image is first processed by the convolution block to obtain the input features of the sub-network Then, the feature F is fine-tuned from the spectral and spatial perspectives. o .
[0026] Furthermore, in the spectral attention optimization subnetwork, feature F oIt is copied into three feature sequences k, q and v, whose resolutions are all c×h×w; q and k use 1×1 convolution layer Conv1 to change the spatial resolution, and get and Use with Attention remap of the same dimension, with spectral resolution of Therefore, the above F o The conversion process is as follows:
[0027]
[0028]
[0029]
[0030] Among them, R s (·) is the remapping function used to facilitate size matching; Represents a matrix multiplication operation; after channel adjustment and Sigmoid activation, Get the weight factor for each channel; then, recalibrate the feature F by multiplying the channel elements o ,get Defined as:
[0031]
[0032] Furthermore, in the spatial attention optimization subnetwork, feature F o It is copied into three feature sequences k, q and v, and the number of channels of k and q is changed by convolution operation to obtain In addition, Put it into the global average pooling GAP, the above process is as follows:
[0033]
[0034]
[0035]
[0036] The resolution is It uses the Sigmoid function to derive activation values that represent the weight factors for each coordinate of the spatial feature; the initial feature map F is recalibrated according to these values. o :
[0037]
[0038] Furthermore, the edge guided learning network is defined as:
[0039]
[0040] Where x is a visible light image and a magnetic resonance imaging medical image, ω Φ represents a learnable parameter.
[0041] As a further step, in the edge-guided learning network, image x performs 3×3 convolution and 1×1 convolution to modify the channel size to obtain features And use BatchNorm layer and RELU activation function to accelerate convergence; then use two methods to process feature F d ; In the first method, we first use a convolutional layer with a stride of 2 to reduce the spatial dimension, and its convolution kernel size is 7×7; then we use a convolution group consisting of three 3×3 convolutional layers and two RELU activation functions to obtain the feature F u ; Then the spatial resolution is restored by upsampling; For the second method, the feature F d The feature F is obtained through a 1×1 convolution layer g ; Finally, the values of the two paths are combined by addition; the process definition is as follows:
[0042] F u =Ups(Conv g (MP(Conv s (F d ))))
[0043] F g =Conv1(F g )
[0044] u Φ =Sigmoid(Conv1(F g +F u )).
[0045] As a further step, the structural similarity loss L f Defined as:
[0046]
[0047] Among them, SSIM f,x is the structural similarity loss between the fused image f and the input image x, SSIM f,y It is the structural similarity loss between the fused image f and the input image y;
[0048] For the cross-modal knowledge enhancement network, the mean square error loss function L is introduced Ψ , which limits the pixel difference between input and output:
[0049]
[0050] For the optimized network based on spectral attention and spatial attention, the saliency weight map m of the source image is calculated and L is defined γ :
[0051] L γ =||m⊙u γ -m⊙y||1
[0052] For the edge-guided learning network, the loss function L is constructed Φ , the loss function is based on the Sober gradient operator
[0053]
[0054] Therefore, the multimodal image fusion model loss is as follows:
[0055] L total =L f +αL Ψ +βL γ +γL Φ
[0056] where α, β and γ are equilibrium L total Hyperparameters of .
[0057] Compared with the prior art, the above technical solution adopted by the present invention has the following advantages: the present invention completely separates the single-mode task and the multi-mode task, and proposes a multi-scene feature joint learning structure to unify the infrared and visible light image fusion and medical image fusion tasks; in order to break down the fusion barriers between different modalities, the fusion process is decomposed into three tasks, and a multi-task learning strategy is applied to guide the model to learn different feature representations more accurately; the cross-modal knowledge reinforcement network enhances the communication between features, thereby achieving effective feature aggregation. In addition, the spatial-spectral domain optimization network and the edge-guided learning network enable the multi-modal image fusion model to simultaneously maintain the foreground target and background details of the source image. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 Flowchart of the multimodal image fusion method for joint learning of multi-scene features;
[0059] Figure 2 Schematic diagram of network structure for cross-modal knowledge enhancement;
[0060] Figure 3 Schematic diagram of the optimized network structure based on spectral attention and spatial attention;
[0061] Figure 4 Schematic diagram of the edge-guided learning network structure;
[0062] Figure 5 For MSRS and M 3 Qualitative comparison of this method and other advanced fusion methods on the FD image dataset;
[0063] Figure 6 The figure is a qualitative comparison between this method and other advanced fusion methods on the Harvard medical image dataset. Specific implementation methods
[0064] In order to make the purpose, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application, that is, the embodiments described are only part of the embodiments of the present application, not all of the embodiments.
[0065] like Figure 1 As shown, this embodiment provides a multimodal image fusion method for joint learning of multi-scene features, including:
[0066] Construct a cross-modal knowledge enhancement network to enhance the consistency and difference features of multimodal images;
[0067] Specifically, visible light images often capture texture details that match human vision, while infrared images can highlight significant targets in harsh environments and obstacles. CT (X-ray computed tomography) images reflect dense structures such as bones. PET (positron emission tomography) images show the function and metabolism of tumors. SPECT (single photon emission computed tomography) images depict tissues, organs, and blood flow. MRI (magnetic resonance imaging) images can provide information about soft tissues. These source images are collected from multiple sensors using various imaging mechanisms. In the cross-modal knowledge reinforcement network, the multi-path correction strategy is used to enhance the consistency features and difference features of multi-modal images, thereby achieving effective feature fusion.
[0068] Multimodal images are captured by sensors with different imaging principles, which can more comprehensively describe real-world scenes. However, how to achieve high-quality and efficient information fusion between images of different modes has always been the goal pursued by researchers. This paper proposes a cross-modal knowledge enhancement network Ψ based on the consistency and difference features of the source images, such as Figure 2 As shown in Figure 2, the network can make full use of the complementary information of the source images to achieve feature aggregation. Specifically, the two source images x and y indicating different modalities are placed in a convolutional block with the same structure, which includes a 3×3 convolutional layer, a BatchNorm layer, and a RELU activation function; the cross-modal knowledge reinforcement network Ψ requires the output of the above convolutional block and The present invention constructs a multi-path correction strategy from the perspective of feature consistency and feature difference to strengthen the feature relationship between source images, thereby reliably enhancing the information exchange between multiple channels. Specifically, and are initial feature maps of the same dimension, which are multiplied element by element to generate intermediate feature maps F M Defined as: In order to maintain the consistency of features, the intermediate feature map F M To enhance and Related information in .
[0069] Construct an optimized network based on spectral attention and spatial attention to enhance the important contextual information in the fusion results;
[0070] Specifically, infrared images have a common feature with CT images, PET images and SPECT images, that is, there is significant information. To this end, an optimization network is constructed, which consists of two sub-networks: spectral attention optimization sub-network and spatial attention optimization sub-network to enhance the important contextual information in the fusion result. Figure 3 The structure is shown;
[0071] In the spectral attention optimization sub-network, feature F o It is copied into three feature sequences k, q and v, whose resolutions are all c×h×w; q and k use 1×1 convolution layer Conv1 to change the spatial resolution, and get and Use with Attention remap of the same dimension, with spectral resolution of
[0072] In the spatial attention optimization subnetwork, feature F o It is copied into three feature sequences k, q and v, and the number of channels of k and q is changed by convolution operation to obtain In addition, Put it into the global average pooling GAP to obtain more spatial dimension information.
[0073] Finally, the outputs of the two sub-networks are combined through an addition operation to aggregate all the attention features.
[0074] Construct an edge-guided learning network to learn contour information from source images;
[0075] Specifically, visible images and MRI images can provide rich texture details; in order to effectively extract texture features from input images, an edge-guided learning network is developed, such as Figure 4 Specifically, the input image x first performs two convolution operations—3×3 convolution and 1×1 convolution to modify the channel size to obtain Use the BatchNorm layer and RELU activation function to accelerate convergence. Next, use two methods to process F d In the first method, a convolution layer with a stride of 2 is used to reduce the spatial dimension. In order to further expand the receptive field, the convolution kernel size is 7×7. Then a convolution group consisting of three 3×3 convolution layers and two RELU activation functions is used to mine more feature information F. u , and then restore the spatial resolution through upsampling operation; for the second method, F d F is obtained through a 1×1 convolutional layer. g ; Finally, the values of the two paths are combined through addition.
[0076] The multimodal image fusion model is obtained through the cross-modal knowledge enhancement network, the optimization network based on spectral attention and spatial attention, and the edge-guided learning network:
[0077] Specifically, the multimodal image fusion model is:
[0078]
[0079] Stu Ψ =Ψ(x,y;ω Ψ ),u Φ =Φ(x;ω Φ ),u γ =γ(y;ω γ )
[0080] Among them, L f is the structural similarity loss, and F is a loss with learnable parameters ω f image reconstruction network; λ is a trade-off parameter; x, y are input images of different modalities; u Ψ ,u Φ ,u γ Represents the output results of the three networks; the model T consists of three networks, namely the cross-modal knowledge reinforcement network Ψ, the optimization network γ based on spectral attention and spatial attention, and the edge-guided learning network Φ; L t is a composite loss function, including L Ψ , L γ and L Φ ;ω t are trainable parameters of the model T.
[0081] The structural similarity loss L f , used to optimize the entire network, it can balance the learning ability of the network and help the model learn structural elements from the source image. f Defined as:
[0082]
[0083] Synthetic loss function L t , by L Ψ , L γ and L Φ Three optimization loss components (corresponding to three networks). For the cross-modal knowledge enhancement network, the mean square error loss function L is introduced Ψ , which limits the pixel difference between input and output:
[0084]
[0085] In addition, an optimized network based on spectral attention and spatial attention is needed to learn as much meaningful information as possible. Therefore, the saliency weight map m of the source image is calculated and L is defined γ :
[0086] L γ =||m⊙u γ -m⊙y||1
[0087] In order to enable the edge-guided learning network to learn more texture details from the source image, a loss function L is constructed. Φ , the loss function is based on the Sober gradient operator
[0088]
[0089] The multimodal image fusion model loss is as follows:
[0090] L total =L f +αL Ψ +βL γ +γL Φ
[0091] where α, β and γ are equilibrium L total Hyperparameters of .
[0092] This method is aimed at two types of multimodal image fusion tasks, including infrared and visible light image fusion and medical image fusion. 3 Test image sequences were selected from the FD dataset and the Harvard medical image dataset for comparison with seven state-of-the-art image fusion methods.
[0093] For the infrared and visible light image fusion task, Figure 5 It shows the overall effect and local feature details. It can be observed that this method can capture the details of the branches under daytime conditions, but the performance of other methods is not very good. Obviously, the contrast of ReCoNet is too large, especially the texture in the shadows of the trees is almost lost. DDcGAN and DIVFusion have good light intensity, but poor clarity. The fusion results produced by the remaining methods are closer to infrared images, and the overall brightness is quite dim. Under dark conditions, DDcGAN, PMGI and DIVFusion can enhance the brightness of the source image, and the details of the scene can be roughly observed. However, this method also changes the information of the source image, causing certain information distortion. U2Fusion, SDNet, ReCoNet and SDDGAN are not good at capturing salient targets in dark areas. In contrast, whether it is the outline of the scene or the outline of the foreground object, the multimodal image fusion model can learn well and get the expected results.
[0094] Table 1 shows the difference between MSRS and M 3 Quantitative comparison between this patent and other advanced fusion methods on FD image dataset
[0095]
[0096] In order to comprehensively compare the performance of the other seven methods, six different metrics are introduced, including visual perception VIF, image structure SSIM, correlation (CC and SCD), information entropy (Q AB / F Table 1 shows the objective quality analysis of the two test datasets MSRS and M 3 Quantitative results on FD; in addition, the average score of all models is provided, and the subscript of each value indicates the standard deviation. ↑ indicates a positive sign. The higher the value, the better. 3 On FD, the scores obtained by this method are almost always the best. On MSRS, this method has the best performance in SCD, Q AB / F , FMI_pixel and SSIM. PMGI's CC and DIVFusion's VIF scores are the highest. The comparative analysis shows that the proposed method has obvious generalization, the fusion result has a strong spatial structure with the source image, and maintains a consistent spatial structure and human visual effect.
[0097] For medical image fusion tasks, Figure 6The overall effect and local feature details are shown. In the first group of CT-MRI datasets, Zero-LMF, DDcGAN and PMGI cannot restore the bright areas in the source image well. Both DDcGAN and PMGI will have information jumps, resulting in poor matching between the fusion results and the source image. U2Fusion, SDNet and ReCoNet lose a lot of structural information during the fusion process, and the entire image becomes blurred. In addition, the results of MSRPAN seem acceptable, but there are jagged local edges. In the second group (PET-MRI dataset) and the third group (SPECT-MRI dataset), the fusion results of DDcGAN and PMGI make it difficult to observe the structural information of MRI. The fused images of U2Fusion and SDNet are generally darker and cannot restore important information well. Although ReCoNet and Zero-LMF can capture important areas from the source image, the local colors are not bright enough. The edges of MSRPAN are not smooth enough. Overall, the fusion results of this method overcome the above defects and can restore structural information and salient information well.
[0098] Table 2 shows the qualitative comparison between this patent and other advanced fusion methods on the Harvard medical image dataset.
[0099]
[0100] After visual evaluation, the quantitative analysis results are given in Table 2. On the CT-MRI test set, the proposed method scores very high in VIF, SSIM, and FMI_pixel, ranking first, indicating that the proposed model has good capabilities in image structure learning and feature preservation. In addition, PMGI's SCD, Zero-LMF's Q AB / F The CC of and U2Fusion are the highest scores, but the corresponding visual effects cannot match them well. In other words, a good model needs to strike a balance between indicators and visual effects. On the PET-MRI test set, U2Fusion has the best CC value, but this method ranks first or second on the other five indicators. On the SPECT-MRI test set, the metrics of this model are excellent. In summary, the fusion results of this method are not only consistent with human vision, but also perform well in various indicators, achieving the unity of subjective and objective evaluation.
[0101] The present invention proposes a multi-scene feature joint learning architecture, which decomposes the fusion process into three tasks, each of which constitutes a learnable network. Specifically, the cross-modal knowledge reinforcement network realizes the effective combination of different modal features. The optimization network based on spectral attention and spatial attention enhances the significant information of the fusion result. The edge-guided learning network can improve the texture capture ability of the model. In addition, three corresponding loss functions are used to optimize the network. Finally, the image reconstruction network produces the desired fusion result. Sufficient experiments have proved that the model proposed in the present invention can generate vivid fusion results in visual perception while ensuring quantitative indicators. Therefore, this method contributes to the development of infrared and visible light image fusion and medical image fusion.
[0102] In this application, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements that are inherent to such process, method, article, or apparatus.
[0103] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A multimodal image fusion method for joint learning of multi-scene features, characterized in that: include: Construct a cross-modal knowledge enhancement network to enhance the consistency and difference features of multimodal images; Construct an optimized network based on spectral attention and spatial attention to enhance the important contextual information in the fusion results; Construct an edge-guided learning network to learn contour information from source images; The multimodal image fusion model is obtained through the cross-modal knowledge enhancement network, the optimization network based on spectral attention and spatial attention, and the edge-guided learning network: student Ψ =Ψ(x,y;ω Ψ ),u Φ =Φ(x;ω Φ ),u γ =γ(y;ω γ ) Among them, L f is the structural similarity loss, and F is a loss with learnable parameters ω f image reconstruction network; λ is a trade-off parameter; x, y are input images of different modalities; u Ψ ,u Φ ,u γ Represents the output results of the three networks; the model T consists of three networks, namely the cross-modal knowledge reinforcement network Ψ, the optimization network γ based on spectral attention and spatial attention, and the edge-guided learning network Φ; L t is a composite loss function, including L Ψ , L γ and L Φ ;ω t are the trainable parameters of model T; The cross-modal knowledge enhancement network constructs a multi-path correction strategy to strengthen the feature relationship between source images. Specifically: and are initial feature maps of the same dimension, which are multiplied element by element to generate intermediate feature maps Therefore, F M Defined as: In order to maintain the consistency of features, the intermediate feature map F M To enhance and The process is defined as follows: Among them, Conv3 is a convolution operation, and its activation function is PReLU; F a The information aggregation in the global feature relationship is realized; for the difference of features, the subtraction operation of feature mapping is introduced to obtain and The joint differential feature F s : Finally, F a and F s Adding together, the consistency and differences of cross-modal information are unified: and Ψ =Conv3(F a +F s ) The optimization network includes a spectral attention optimization subnetwork and a spatial attention optimization subnetwork, specifically: Among them, ω γ is the trainable parameter of the network γ; y represents the input image, which includes infrared images, CT images, PET images and SPECT images; after entering the optimization network, the source image is first processed by the convolution block to obtain the input features of the sub-network Then, the feature F is fine-tuned from the spectral and spatial perspectives. o ; In the edge-guided learning network, the image performs 3×3 convolution and 1×1 convolution to modify the channel size to obtain features. And use BatchNorm layer and RELU activation function to accelerate convergence; then use two methods to process feature F d ; In the first method, we first use a convolutional layer with a stride of 2 to reduce the spatial dimension, and its convolution kernel size is 7×7; then we use a convolution group consisting of three 3×3 convolutional layers and two RELU activation functions to obtain the feature F u ; Then the spatial resolution is restored by upsampling; For the second method, the feature F d The feature F is obtained through a 1×1 convolution layer g ; Finally, the values of the two paths are combined by addition; the process definition is as follows: F u =Ups(Conv g (MP(Conv s (F d )))) F g =Conv1(F g ) u Φ =Sigmoid(Conv1(F g +F u ))。 2. According to claim 1, a multi-modal image fusion method for joint learning of multi-scene features is characterized in that: The cross-modal knowledge enhancement network makes full use of the complementary information of the source image to achieve feature aggregation. The optimization process of the network is defined as: Among them, ω Ψ Represents the trainable parameters of the network; L Ψ is the loss function.
3. The multimodal image fusion method for joint learning of multi-scene features according to claim 1, characterized in that: The initial feature map and The acquisition method is: put two source images x and y of different modalities into convolution blocks with the same structure, respectively, and the convolution blocks include 3×3 convolution layers, BatchNorm layers and RELU activation functions; obtain the initial feature map through the convolution blocks and 4. The multimodal image fusion method for joint learning of multi-scene features according to claim 1, characterized in that: In the spectral attention optimization sub-network, feature F o It is copied into three feature sequences k, q and v, whose resolutions are all c×h×w; q and k use 1×1 convolution layer Conv1 to change the spatial resolution, and get and Use with Attention remap of the same dimension, with spectral resolution of Therefore, the above F o The conversion process is as follows: Among them, R s (·) is the remapping function used to facilitate size matching; Represents a matrix multiplication operation; after channel adjustment and Sigmoid activation, Get the weight factor for each channel; then, recalibrate the feature F by multiplying the channel elements o ,get Defined as:
5. The multimodal image fusion method for joint learning of multi-scene features according to claim 1, characterized in that: In the spatial attention optimization subnetwork, feature F o It is copied into three feature sequences k, q and v, and the number of channels of k and q is changed by convolution operation to obtain In addition, Put it into the global average pooling GAP, the above process is as follows: The resolution is It uses the Sigmoid function to derive activation values that represent the weight factors for each coordinate of the spatial feature; the initial feature map F is recalibrated according to these values. o :
6. The multimodal image fusion method for joint learning of multi-scene features according to claim 1, characterized in that: The edge-guided learning network is defined as: Where x is a visible light image and a magnetic resonance imaging medical image, ω Φ represents a learnable parameter.
7. The multimodal image fusion method for joint learning of multi-scene features according to claim 1, characterized in that: The structural similarity loss L f Defined as: Among them, SSIM f,x is the structural similarity loss between the fused image f and the input image x, SSIM f,y It is the structural similarity loss between the fused image f and the input image y; For the cross-modal knowledge enhancement network, the mean square error loss function L is introduced Ψ , which limits the pixel difference between input and output: For the optimized network based on spectral attention and spatial attention, the saliency weight map m of the source image is calculated and L is defined γ : L γ =||m⊙u γ -m⊙y||1 For the edge-guided learning network, the loss function L is constructed Φ , the loss function is based on the Sober gradient operator Therefore, the multimodal image fusion model loss is as follows: L total =L f +αL Ψ +βL γ +γL Φ where α, β and γ are equilibrium L total Hyperparameters of .
Citation Information
Cited By
Multi-modal image fusion method based on dynamic pseudo supervision and semantic guidance
CN121788981A
A multi-modal image fusion method based on dynamic pseudo-supervision and semantic guidance
CN121788981B