Medical image fusion method based on multi-scale cavity convolution and channel attention mechanism

Through the residual network of multi-scale void convolution and channel attention mechanism, combined with the adaptive weight mechanism and joint loss function, the problem of insufficient extraction of structural and semantic information in multimodal medical image fusion is solved, and the flexibility and quality of image fusion are improved.

CN120635642APending Publication Date: 2025-09-12CHANGZHOU UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510709369.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing multimodal medical image fusion methods have problems in feature extraction and fusion strategies, such as insufficient extraction of structural and semantic information, inflexible fusion weight distribution, and limited optimization of output image quality.

Method used

A residual network with multi-scale dilated convolution and channel attention mechanism is used to extract feature maps, which are then combined with an adaptive weight mechanism for weighted fusion. The model parameters are optimized through a joint loss function to generate high-quality fused images.

Benefits of technology

It improves the flexibility and information retention capability of multimodal medical image fusion, improves the structural and semantic quality of the fused image, has good generalization capability and end-to-end trainability, and is suitable for multimodal medical image fusion scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635642A_ABST
    Figure CN120635642A_ABST
Patent Text Reader

Abstract

The invention discloses a medical image fusion method based on multi-scale cavity convolution and a channel attention mechanism, and the method comprises the steps: obtaining a sample set of fMRI and MRI-T1 image pairs based on an Alzheimer's disease neuroimaging plan database; constructing a medical image fusion model; adopting a residual network combining multi-scale cavity convolution and a channel attention mechanism to extract each modal feature pattern of the sample set; performing weighted fusion on the modal feature map through an adaptive weight mechanism to generate a fused feature map; reconstructing the fusion feature map into a single-channel medical image by adopting a multi-layer convolutional network, and outputting a fusion result map; constructing a joint loss function training fusion model, and optimizing parameters; and performing reasoning processing on the fMRI and MRI-T1 image pair based on the optimized parameters to generate a fused image. The method is suitable for a multi-modal medical image fusion task, and the overall performance of the fusion image in the aspects of structure restoration, detail maintenance and semantic expression can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image processing, and in particular to a medical image fusion method based on multi-scale dilated convolution and channel attention mechanism. Background Art

[0002] With the advancement of medical imaging technology, functional magnetic resonance imaging (fMRI) and T1-weighted magnetic resonance imaging (MRI-T1) have gained widespread application in neuroscience and medical analysis. fMRI can reveal functional connectivity between brain regions, while MRI-T1 provides clear anatomical information. The integration of these two approaches enhances the integrity of image representation and enables multidimensional interpretation.

[0003] Existing multimodal medical image fusion methods still have shortcomings in feature extraction and fusion strategies. Some methods lack the ability to jointly model the multi-scale structure of images and key semantic channels, limiting the effective expression of deep structural and semantic information during the feature extraction process. For example, CN117788314A, "CT and MRI Image Fusion Method Based on Multi-branch Multi-scale Features and Semantic Information," uses a multi-branch structure to extract features of different scales. However, because this method does not incorporate a channel attention mechanism, it is unable to automatically distinguish and assign weights to different channels based on information content during the multimodal feature fusion process, failing to suppress redundant features or enhance the expressiveness of significant features. As a result, this method is prone to information loss or redundant accumulation in the fusion results, resulting in insufficient discriminability, expressiveness, and adaptability of the fused image, affecting the final image quality and the accuracy of downstream tasks. CN117788313A, "A Medical Image Fusion Method Based on Dual Contrast Learning and Gradient Channel Attention Mechanism," although this method introduces a gradient channel attention mechanism, it is not combined with a multi-scale dilated convolutional structure, resulting in an inability to simultaneously account for information expression of features at different scales, limiting the utilization of features at a single scale. Furthermore, its fusion strategy relies solely on static gradient thresholds and pooling rules, failing to adaptively learn the importance of each modality's features based on image content. This results in inflexible fusion weight allocation, making it impossible to dynamically adjust contributions for different modalities and content, limiting further improvements in fusion performance. Summary of the Invention

[0004] The purpose of the present invention is to provide a medical image fusion method based on multi-scale dilated convolution and channel attention mechanism, which solves the problems of insufficient structural and semantic information extraction, inflexible fusion weight distribution and limited output image quality optimization in multimodal medical image fusion.

[0005] To solve the above technical problems, the technical solution of the present invention is: a medical image fusion method based on multi-scale dilated convolution and channel attention mechanism, comprising the following steps:

[0006] Step 1: Based on the Alzheimer's Disease Neuroimaging Project database, obtain a sample set of fMRI and MRI-T1 image pairs;

[0007] Step 2: Build a medical image fusion model and use the sample set as input;

[0008] Step 3: Using a residual network that combines multi-scale dilated convolution with a channel attention mechanism to extract feature maps of each modality of the sample set;

[0009] Step 4: Perform weighted fusion on the modal feature maps through an adaptive weight mechanism to generate a fused feature map;

[0010] Step 5: Use a multi-layer convolutional network to reconstruct the fused feature map into a single-channel medical image and output the fusion result map;

[0011] Step 6: Construct a joint loss function to train the fusion model and optimize the parameters of the fusion model; Step 7: Perform inference processing on the fMRI and MRI-T1 image pairs based on the optimized parameters to generate a fusion image.

[0012] Preferably, the step 1 of obtaining a sample set of fMRI and MRI-T1 image pairs based on the Alzheimer's Disease Neuroimaging Project database is specifically implemented as follows:

[0013] Step 1.1: Acquire paired fMRI and MRI-T1 image data from the Alzheimer's Disease Neuroimaging Project database and align the fMRI images to the reference space of the MRI-T1 images using a fusion of rigid and non-rigid spatial registration.

[0014] Step 1.2: After performing noise reduction processing on the image data, Z-score intensity normalization processing is performed;

[0015] Step 1.3: Based on the normalized fMRI image data from step 1.2, select the middle time frame in the time series, extract the functional area slice sequence at equal intervals along the z-axis of the image, and construct the functional feature representation;

[0016] Step 1.4: Based on the normalized MRI-T1 image data from step 1.2, extract a continuous structural slice sequence containing the target brain region along the middle z-axis of the image to construct a structural feature representation;

[0017] Step 1.5: Resize and standardize the slice sequences selected in steps 1.3 and 1.4 to obtain fMRI and MRI-T1 images with consistent input dimensions.

[0018] Step 1.6: Based on the image edge distribution and the proportion of non-zero pixels, images with background dominance or insufficient structural information are screened out to obtain a sample set of fMRI and MRI-T1 image pairs with complementary functions and structures.

[0019] Preferably, in step 3, a residual network combining multi-scale dilated convolution and channel attention mechanism is used to extract feature maps of each modality of the sample set, which is specifically implemented as follows:

[0020] Step 3.1: Use the residual connection structure to enhance the effective propagation of gradients in the convolutional network and output the intermediate feature map. The expression is as follows:

[0021] h i =Conv(x i )+x i

[0022] Among them, x i represents the input feature map of the i-th modality, h i Represents the intermediate feature map of the i-th modality, Conv(x i ) represents the input feature map x i The feature map obtained after the convolution operation;

[0023] Step 3.2: Use three parallel paths of dilated convolution to fuse the multi-scale feature information of each modality under different receptive fields to obtain a multi-scale feature map. The fusion expression is as follows:

[0024]

[0025] Among them, Conv k (h i ) represents the k-th convolution path to the intermediate feature map h i The feature map obtained after the convolution operation, f i Represents the multi-scale feature map of the i-th modality;

[0026] Step 3.3: Through the channel attention mechanism, the channel-level features of the multi-scale feature map obtained in step 3.2 are adaptively weighted and adjusted to output the enhanced feature maps of each modality. The calculation formula is as follows:

[0027] y i =f i ·σ(W2·ReLU(W1·Pool(f i )))

[0028] Among them, f i Represents the multi-scale feature map of the i-th modality, y i Represents the modal feature map after the i-th modal enhancement, Pool(f i ) indicates the value of fi All channels are globally average pooled, W1 and W2 represent the weight parameters of the first and second fully connected layers, respectively, ReLU(·) represents the activation function, and σ represents the Sigmoid function.

[0029] Preferably, the weighted fusion of the modal feature maps by the adaptive weight mechanism in step 4 to generate a fused feature map is specifically implemented as follows:

[0030] Step 4.1: Calculate the global strength of each modal feature map using the Frobenius norm. The calculation formula is as follows:

[0031]

[0032] Among them, ‖f i ‖ F represents the Frobenius norm of the i-th modality multi-scale feature map, f i Represents the multi-scale feature map of the i-th modality, f i (j) represents the j-th pixel value in the multi-scale feature map of the i-th modality.

[0033] Step 4.2: Introduce a learnable scaling factor and bias term, normalize the Frobenius norm obtained in step 4.1 through the Softmax function, and obtain the dynamic fusion weight of the modal feature map. The calculation formula is as follows:

[0034]

[0035] Among them, w i represents the dynamic fusion weight of the i-th modal feature map, α represents a learnable scaling factor, which is used to control the sensitivity of Softmax to the importance of modal features; β represents a learnable bias term, which is used to adjust the balance of weight distribution, e represents a natural constant, which is the base of the exponential function and is used to calculate the exponential term in the Softmax function, ‖f i ‖ F represents the Frobenius norm of the i-th modal multi-scale feature map;

[0036] Step 4.3: Use the dynamic fusion weights to perform weighted fusion on the modal feature maps to generate a fused feature map, which is expressed as follows:

[0037]

[0038] Among them, f fusion Represents the fusion feature map, w i represents the dynamic fusion weight of the i-th modality feature map, y i Represents the modal feature map after the i-th modality is enhanced.

[0039] Preferably, in step 5, a multi-layer convolutional network is used to reconstruct the fused feature map into a single-channel medical image, and a fusion result map is output. The reconstruction process is as follows:

[0040]

[0041] in, Represents the fusion result graph, f fusion Represents the fusion feature map, Conv1 represents the first layer of convolution operation, Conv n Indicates the nth layer convolution operation, Conv1 to Conv n Represents the convolution operations from the 1st to the nth layer in the reconstruction process.

[0042] Preferably, the joint loss function in step 6 includes a pixel loss term, a gradient loss term and a perceptual loss term.

[0043] Preferably, the pixel loss term is defined by the mean square error, and the calculation formula is as follows:

[0044]

[0045] Among them, L MSE Represents the pixel loss term, N represents the total number of pixels in the image, Represents the i-th pixel value in the fusion result image, and Represent the corresponding i-th pixel value in the fMRI image and MRI-T1 image respectively;

[0046] The gradient loss term is calculated as follows:

[0047]

[0048] Among them, L Gradient represents the gradient loss term, N represents the total number of pixels in the image, Indicates the position of the i-th pixel in the fusion result image, represents the i-th pixel position in the fMRI image, represents the i-th pixel position in the MRI-T1 image, Represents the image gradient of the i-th pixel position in the fusion result image, represents the image gradient at the i-th pixel position in the fMRI image, represents the image gradient at the i-th pixel position in the MRI-T1 image;

[0049] The perceptual loss term is calculated as follows:

[0050]

[0051] Among them, L Perceptual Represents the perceptual loss term, M represents the number of feature dimensions involved in the perceptual loss calculation, Represents the fusion result graph, I (f) and I (s) Represent the FMRI image and MRI-T1 image respectively, which serve as the reference images of the fusion result image in terms of functional information and structural information. represents the feature representation of the i-th feature dimension through the pre-trained neural network, and Represent the feature representation of the fusion result image, fMRI image and MRI-T1 image in the i-th feature dimension respectively;

[0052] The joint loss function L is constructed by weighting the pixel loss term, the gradient loss term, and the perceptual loss term, and is expressed as follows:

[0053] L=λ1L MSE +λ2L Gradient +λ3L Perceptual

[0054] Among them, λ1, λ2, and λ3 represent the weighted coefficients of the pixel loss term, gradient loss term, and perceptual loss term, respectively, which are used to optimize the weights of the corresponding loss terms in the neural network.

[0055] Preferably, the step 7 of performing inference processing on the fMRI and MRI-T1 image sample pairs based on the optimized parameters to generate a fused image further includes the following steps:

[0056] Step 7.1: Post-process the fused image including noise suppression, brightness adjustment and boundary cropping;

[0057] Step 7.2: Use performance evaluation indicators to evaluate the quality of the fused image that has been post-processed in step 7.1 and output the quality evaluation results.

[0058] Preferably, the performance evaluation indicators in step 7.2 include structural similarity index, peak signal-to-noise ratio, mutual information, visual information fidelity and entropy value.

[0059] Compared with the existing technology, the beneficial effects of the present invention are as follows: (1) in the feature extraction stage, the expression ability of key semantic channels is enhanced through multi-scale void convolution and channel attention mechanism, and hierarchical structure and semantic information are extracted under multiple receptive fields, which provides support for the full expression of fusion features; (2) in the feature fusion stage, learnable parameters are introduced for dynamic weighting, so that the multimodal image features can adaptively adjust the weights during the fusion process, thereby improving the flexibility of fusion and the ability to retain information; (3) in the training stage, a multi-dimensional loss term including pixel error, structural edge and semantic perception is adopted, so that the optimization process can simultaneously constrain the output image quality from multiple levels, avoid information loss or structural distortion caused by single indicator optimization, and help improve the structural and semantic quality of the fused image; (4) the present invention has good generalization ability and end-to-end trainable characteristics, is suitable for multimodal medical image fusion scenarios, and can provide a high-quality fusion foundation for subsequent image analysis and auxiliary processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0061] Figure 1 It is a schematic diagram of the implementation process of the present invention. DETAILED DESCRIPTION

[0062] The following is a further detailed description of the present invention in conjunction with the accompanying drawings. The terminal technical solutions of the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0063] like Figure 1 As shown, the present invention provides a medical image fusion method based on multi-scale dilated convolution and channel attention mechanism, comprising the following steps:

[0064] Step 1: Based on the Alzheimer's Disease Neuroimaging Project database, obtain a sample set of fMRI and MRI-T1 image pairs. The specific implementation is as follows:

[0065] Step 1.1: Paired fMRI and MRI-T1 brain image data were collected from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database. The fMRI images were aligned to the reference space of the MRI-T1 images using a fusion of rigid and nonrigid spatial registration to achieve multimodal anatomical alignment.

[0066] Step 1.2: After the image data is subjected to noise reduction processing, Z-score intensity normalization processing is performed to improve image clarity and inter-modality comparability; the noise reduction processing uses Gaussian and median filtering to suppress non-structural noise in the image data.

[0067] Step 1.3: Based on the normalized fMRI image data from step 1.2, select the middle time frame in the time series, extract the functional area slice sequence covering different anatomical depths at equal intervals along the z-axis of the image, and construct the functional feature representation.

[0068] Step 1.4: Based on the normalized MRI-T1 image data from step 1.2, extract a continuous structural slice sequence containing the target brain region space along the middle of the image z-axis to construct a structural feature representation.

[0069] Step 1.5: Resize and standardize the slice sequences selected in steps 1.3 and 1.4 to obtain fMRI and MRI-T1 images with consistent input dimensions; the slice sequences are uniformly scaled to 256×256 pixels.

[0070] Step 1.6: Based on the image edge distribution and the proportion of non-zero pixels, images with background dominance or insufficient structural information are screened out to obtain a sample set of fMRI and MRI-T1 image pairs with complementary functions and structures. Only valid medical content is retained in the sample set for subsequent processing.

[0071] Through modality registration and image screening, the quality of the sample set and the expression ability of structure-function complementarity are improved in the image data preparation stage. Compared with the average slice sequence used in existing methods or the sample construction method without effective screening, the present invention can provide more representative and effective input data for subsequent processing.

[0072] Step 2: Build a medical image fusion model and take the sample set as input.

[0073] Step 3: A residual network combining multi-scale dilated convolution and channel attention mechanism is used to extract the modal feature maps of the sample set. The modal feature maps include deep structure feature maps and semantic feature maps, which are used to enhance the expression of key semantic channels. The specific implementation is as follows:

[0074] Step 3.1: Use the residual connection structure to enhance the effective propagation of gradients in the convolutional network to alleviate the gradient attenuation problem in the deep network, enhance the expression continuity and training convergence ability in the feature extraction process, and output the intermediate feature map. The expression is as follows:

[0075] h i =Conv(x i )+x i

[0076] Among them, x i represents the input feature map of the i-th modality, h i Represents the intermediate feature map of the i-th modality, Conv(x i ) represents the input feature map x i The feature map obtained after the convolution operation;

[0077] Step 3.2: Use three parallel dilated convolution paths to fuse the multi-scale feature information of each modality in different receptive fields to obtain a multi-scale feature map. Each path uses a different dilation rate to extract local features and global context features.

[0078] The fusion expression is as follows:

[0079]

[0080] Among them, Conv k (h i ) represents the k-th convolution path to the intermediate feature map h i The feature map obtained after the convolution operation, f i Represents the multi-scale feature map of the i-th modality.

[0081] Step 3.3: Adaptively weight the channel-level features of the multi-scale feature maps obtained in Step 3.2 using the channel attention mechanism, outputting enhanced feature maps for each modality. The introduction of the channel attention mechanism further enhances the semantic responsiveness of key channels. Global average pooling compresses channel information, and then a two-layer fully connected network calculates channel weights, ultimately achieving weighted adjustment of the multi-scale feature maps.

[0082] The output feature map is calculated as follows:

[0083] y i =f i ·σ(W2·ReLU(W1·Pool(f i )))

[0084] Among them, f i Represents the multi-scale feature map of the i-th modality, y iRepresents the modal feature map after the i-th modal enhancement, Pool(f i ) indicates the value of f i All channels are globally average pooled, W1 and W2 represent the weight parameters of the first and second fully connected layers, respectively, ReLU(·) represents the activation function, and σ represents the Sigmoid function.

[0085] The above feature extraction structure ensures smooth gradient propagation through the residual connection structure, enhances feature diversity through the degree-scale spatial convolutional network, and improves the expression ability of key semantic information through the channel attention mechanism, laying the foundation for subsequent fusion processing.

[0086] Step 4: Perform weighted fusion on the modal feature maps through an adaptive weight mechanism to generate a fused feature map, dynamically weight and fuse the features of each modality, and generate a unified feature representation. The specific implementation is as follows:

[0087] Step 4.1: Calculate the global strength of each modal feature map using the Frobenius norm. The calculation formula is as follows:

[0088]

[0089] Among them, ‖f i ‖ F represents the Frobenius norm of the i-th modality multi-scale feature map, f i Represents the multi-scale feature map of the i-th modality, f i (j) represents the j-th pixel value in the multi-scale feature map of the i-th modality.

[0090] Step 4.2: Introduce a learnable scaling factor and bias term, normalize the Frobenius norm obtained in step 4.1 through the Softmax function, and obtain the dynamic fusion weight of the modal feature map. The calculation formula is as follows:

[0091]

[0092] Among them, w i represents the dynamic fusion weight of the i-th modal feature map, α represents a learnable scaling factor, which is used to control the sensitivity of Softmax to the importance of modal features; β represents a learnable bias term, which is used to adjust the balance of weight distribution, e represents a natural constant, which is the base of the exponential function and is used to calculate the exponential term in the Softmax function, ‖f i ‖ F represents the Frobenius norm of the multi-scale feature map of the i-th modality.

[0093] By dynamically optimizing α and β through training, adaptive adjustment of the contribution of different modal features is achieved, enhancing the flexibility and information retention ability of the fusion process.

[0094] Step 4.3: Use the dynamic fusion weights to perform weighted fusion on the modal feature maps to generate a fused feature map, which is expressed as follows:

[0095]

[0096] Among them, f fusion Represents the fusion feature map, w i represents the dynamic fusion weight of the i-th modality feature map, y i Represents the modal feature map after the i-th modality is enhanced.

[0097] Compared with the traditional static splicing or average weighting method, the above fusion process has stronger expressiveness and adaptability in terms of fusion flexibility and information retention.

[0098] Step 5: Use a multi-layer convolutional network to reconstruct the fused feature map into a single-channel medical image and output the fusion result map. The reconstruction process is as follows:

[0099]

[0100] in, Represents the fusion result graph, f fusion Represents the fusion feature map, Conv1 represents the first layer of convolution operation, Conv n Indicates the nth layer convolution operation, Conv1 to Conv n Represents the convolution operations from the 1st to the nth layer in the reconstruction process.

[0101] Step 6: Construct a joint loss function to train the fusion model and optimize the parameters of the fusion model.

[0102] The joint loss function includes a pixel loss term, a gradient loss term, and a perceptual loss term. The pixel loss term is used to measure the error in pixel space between the fusion result image and the reference image, which are the FMRI image and the MRI-T1 image. The gradient loss term focuses on the edge structure information of the fusion result image, which can improve the structural expression ability of the fusion result image by enhancing the contour continuity and clarity of the fusion result image. The perceptual loss term is based on the high-level semantic features extracted by the pre-trained convolutional network, and measures the semantic similarity between the fusion result image and the reference image in the deep semantic feature space.

[0103] The pixel loss term is defined by the mean square error, which is calculated as follows:

[0104]

[0105] Among them, L MSE Represents the pixel loss term, N represents the total number of pixels in the image, Represents the i-th pixel value in the fusion result image, and Represent the corresponding i-th pixel value in the fMRI image and MRI-T1 image respectively;

[0106] The gradient loss term is calculated as follows:

[0107]

[0108] Among them, L Gradient represents the gradient loss term, N represents the total number of pixels in the image, Indicates the position of the i-th pixel in the fusion result image, represents the i-th pixel position in the fMRI image, represents the i-th pixel position in the MRI-T1 image, Represents the image gradient of the i-th pixel position in the fusion result image, represents the image gradient at the i-th pixel position in the fMRI image, represents the image gradient at the i-th pixel position in the MRI-T1 image.

[0109] The perceptual loss term is calculated as follows:

[0110]

[0111] Among them, L Perceptual Represents the perceptual loss term, M represents the number of feature dimensions involved in the perceptual loss calculation, Represents the fusion result graph, I (f) and I (s) Represent the FMRI image and MRI-T1 image respectively, which serve as the reference images of the fusion result image in terms of functional information and structural information. represents the feature representation of the i-th feature dimension through the pre-trained neural network, and Represent the feature representation of the fusion result image, fMRI image and MRI-T1 image in the i-th feature dimension respectively.

[0112] The joint loss function L is constructed by weighting the pixel loss term, the gradient loss term, and the perceptual loss term, and is expressed as follows:

[0113] L=λ1L MSE +λ2L Gradient +λ3L Perceptual

[0114] Among them, λ1, λ2, and λ3 represent the weighted coefficients of the pixel loss term, gradient loss term, and perceptual loss term, respectively, which are used to optimize the weights of the corresponding loss terms in the neural network.

[0115] During the training process of the fusion model, the joint loss function is used to guide the collaborative optimization of three levels: pixel accuracy, structural edges, and semantic features. Compared with the existing training method that only relies on a single loss term, the present invention helps to improve the comprehensive performance of the fused image in terms of structural clarity and semantic consistency.

[0116] Step 7: Based on the optimized parameters, perform inference processing on the fMRI and MRI-T1 image pairs to generate a fused image. Combined with image enhancement and evaluation strategies, improve the output quality and verify the fusion effect. The specific implementation is as follows:

[0117] Step 7.1: Post-process the fused image, including noise suppression, brightness adjustment, and boundary cropping, to further improve the visual performance and applicability of the fused image to downstream tasks.

[0118] (1) Noise suppression: Gaussian filter or non-local mean denoising method is introduced to suppress low-amplitude and high-frequency noise in the fused image.

[0119] (2) Brightness adjustment: Perform a brightness averaging operation based on histogram equalization to enhance the contrast and visual structure of the fused image.

[0120] (3) Boundary cropping: Based on the edge information of the fused image, the invalid background area is cropped and the main part of the image containing the effective structure of the brain area is retained.

[0121] Through the above post-processing process, the quality of the fused image can be improved in terms of structural integrity, grayscale consistency and subjective clarity, providing strong support for clinical readability and subsequent analysis.

[0122] Step 7.2: To objectively measure the performance of the fusion method in terms of structural restoration and modal information preservation, the quality of the post-processed fusion image in step 7.1 is evaluated using performance evaluation indicators, and the quality evaluation results are output.

[0123] The performance evaluation indicators include structural similarity index, peak signal-to-noise ratio, mutual information, visual information fidelity and entropy value.

[0124] (1) Structural Similarity Index (SSIM): used to measure the consistency of the structure and texture between the fused image and the reference image.

[0125] (2) Peak signal-to-noise ratio (PSNR): used to measure the reconstruction accuracy of the image at the pixel level.

[0126] (3) Mutual information (MI): reflects the degree to which the fused image preserves the information of the source image.

[0127] (4) Visual Information Fidelity (VIFF): evaluates the authenticity of the fused image at the visual perception level.

[0128] (5) Entropy (EN): as a reference indicator of information richness.

[0129] The above performance evaluation indicators reflect the advantages of the fusion of multi-scale dilated convolution and channel attention mechanism adopted by this invention in terms of structural alignment and information retention. Furthermore, the Frobenius norm of the modal feature map is used to introduce a learnable scaling factor and bias term, enabling dynamic weighting and fusion of the features of each modality to generate a unified feature representation. This further enhances the flexibility and information retention of the fusion process, thereby improving the overall performance of the fused image.

[0130] The parts not involved in the present invention are the same as the existing technology or are implemented by using the existing technology.

[0131] The above is a further detailed description of the present invention in conjunction with specific embodiments, and the specific implementation of the present invention cannot be considered to be limited to these descriptions. For those skilled in the art of the present invention, several simple deductions or substitutions can be made without departing from the concept of the present invention, and all of these should be considered to fall within the scope of protection of the present invention.

Claims

1. A medical image fusion method based on multi-scale dilated convolution and channel attention mechanism, characterized by: The steps include: Step 1: Based on the Alzheimer's Disease Neuroimaging Project database, obtain a sample set of fMRI and MRI-T1 image pairs; Step 2: Build a medical image fusion model and use the sample set as input; Step 3: Using a residual network that combines multi-scale dilated convolution with a channel attention mechanism to extract feature maps of each modality of the sample set; Step 4: Perform weighted fusion on the modal feature maps through an adaptive weight mechanism to generate a fused feature map; Step 5: Use a multi-layer convolutional network to reconstruct the fused feature map into a single-channel medical image and output the fusion result map; Step 6: Construct a joint loss function to train the fusion model and optimize the parameters of the fusion model; Step 7: Perform inference processing on the fMRI and MRI-T1 image pairs based on the optimized parameters to generate a fused image.

2. The medical image fusion method based on multi-scale dilated convolution and channel attention mechanism according to claim 1, characterized in that: In step 1, a sample set of fMRI and MRI-T1 image pairs is obtained based on the Alzheimer's Disease Neuroimaging Project database. The specific implementation is as follows: Step 1.1: Acquire paired fMRI and MRI-T1 image data from the Alzheimer's Disease Neuroimaging Project database and align the fMRI images to the reference space of the MRI-T1 images using a fusion of rigid and non-rigid spatial registration. Step 1.2: After performing noise reduction processing on the image data, Z-score intensity normalization processing is performed; Step 1.3: Based on the fMRI image data normalized in step 1.2, the middle time frame in the time series is selected, and a functional area slice sequence is extracted at equal intervals along the z-axis of the image to construct a functional feature representation; Step 1.4: Based on the normalized MRI-T1 image data from step 1.2, extract a continuous structural slice sequence containing the target brain region along the middle z-axis of the image to construct a structural feature representation; Step 1.5: Resize and standardize the slice sequences selected in steps 1.3 and 1.4 to obtain fMRI and MRI-T1 images with consistent input dimensions. Step 1.6: Based on the image edge distribution and the proportion of non-zero pixels, images with background dominance or insufficient structural information are screened out to obtain a sample set of fMRI and MRI-T1 image pairs with complementary functions and structures.

3. The medical image fusion method based on multi-scale dilated convolution and channel attention mechanism according to claim 1, characterized in that: In step 3, a residual network combining multi-scale dilated convolution and channel attention mechanism is used to extract the feature maps of each modality of the sample set. The specific implementation is as follows: Step 3.1: Use the residual connection structure to enhance the effective propagation of gradients in the convolutional network and output the intermediate feature map. The expression is as follows: h i =Conv(x i )+x i Among them, x i represents the input feature map of the i-th modality, h i Represents the intermediate feature map of the i-th modality, Conv(x i ) represents the input feature map x i The feature map obtained after the convolution operation; Step 3.2: Use three parallel paths of dilated convolution to fuse the multi-scale feature information of each modality under different receptive fields to obtain a multi-scale feature map. The fusion expression is as follows: Among them, Conv k (h i ) represents the k-th convolution path to the intermediate feature map h i The feature map obtained after the convolution operation, f i Represents the multi-scale feature map of the i-th modality; Step 3.3: Through the channel attention mechanism, the channel-level features of the multi-scale feature map obtained in step 3.2 are adaptively weighted and adjusted to output the enhanced feature maps of each modality. The calculation formula is as follows: yes i =f i ·σ(W2·ReLU(W1·Pool(f i ))) Among them, f i Represents the multi-scale feature map of the i-th modality, y i Represents the modal feature map after the i-th modal enhancement, Pool(f i ) indicates the value of f i All channels are globally average pooled, W1 and W2 represent the weight parameters of the first and second fully connected layers, respectively, ReLU(·) represents the activation function, and σ represents the Sigmoid function.

4. The medical image fusion method based on multi-scale dilated convolution and channel attention mechanism according to claim 1, characterized in that: In step 4, the modal feature maps are weighted and fused through the adaptive weight mechanism to generate a fused feature map. The specific implementation is as follows: Step 4.1: Calculate the global strength of each modal feature map using the Frobenius norm. The calculation formula is as follows: Among them, ‖f i ‖ F represents the Frobenius norm of the i-th modality multi-scale feature map, f i Represents the multi-scale feature map of the i-th modality, f i (j) represents the jth pixel value in the multi-scale feature map of the i-th modality; Step 4.2: Introduce a learnable scaling factor and bias term, normalize the Frobenius norm obtained in step 4.1 through the Softmax function, and obtain the dynamic fusion weight of the modal feature map. The calculation formula is as follows: Among them, w i represents the dynamic fusion weight of the i-th modal feature map, α represents a learnable scaling factor, which is used to control the sensitivity of Softmax to the importance of modal features; β represents a learnable bias term, which is used to adjust the balance of weight distribution, e represents a natural constant, which is the base of the exponential function and is used to calculate the exponential term in the Softmax function, ‖f i ‖ F represents the Frobenius norm of the i-th modal multi-scale feature map; Step 4.3: Use the dynamic fusion weights to perform weighted fusion on the modal feature maps to generate a fused feature map, which is expressed as follows: Among them, f fusion Represents the fusion feature map, w i represents the dynamic fusion weight of the i-th modality feature map, y i Represents the modal feature map after the i-th modality is enhanced.

5. The medical image fusion method based on multi-scale dilated convolution and channel attention mechanism according to claim 1, characterized in that: As described in step 5, a multi-layer convolutional network is used to reconstruct the fused feature map into a single-channel medical image and output the fusion result map. The reconstruction process is as follows: in, Represents the fusion result graph, f fusion Represents the fusion feature map, Conv1 represents the first layer of convolution operation, Conv n Indicates the nth layer convolution operation, Conv1 to Conv n Represents the convolution operations from the 1st to the nth layer in the reconstruction process.

6. The medical image fusion method based on multi-scale dilated convolution and channel attention mechanism according to claim 1, characterized in that: The joint loss function described in step 6 includes pixel loss term, gradient loss term and perceptual loss term.

7. The medical image fusion method based on multi-scale dilated convolution and channel attention mechanism according to claim 6, characterized in that: The pixel loss term is defined by the mean square error, which is calculated as follows: Among them, L MSE Represents the pixel loss term, N represents the total number of pixels in the image, Represents the i-th pixel value in the fusion result image, and Represent the corresponding i-th pixel value in the fMRI image and MRI-T1 image respectively; the gradient loss term is calculated as follows: Among them, L Gradient represents the gradient loss term, N represents the total number of pixels in the image, Indicates the position of the i-th pixel in the fusion result image, represents the i-th pixel position in the fMRI image, represents the i-th pixel position in the MRI-T1 image, Represents the image gradient of the i-th pixel position in the fusion result image, represents the image gradient at the i-th pixel position in the fMRI image, represents the image gradient at the i-th pixel position in the MRI-T1 image; The perceptual loss term is calculated as follows: Among them, L Perceptual Represents the perceptual loss term, M represents the number of feature dimensions involved in the perceptual loss calculation, Represents the fusion result graph, I (f) and I (s) Represent the FMRI image and MRI-T1 image respectively, which serve as the reference images of the fusion result image in terms of functional information and structural information. represents the feature representation of the i-th feature dimension through the pre-trained neural network, and Represent the feature representation of the fusion result image, fMRI image and MRI-T1 image in the i-th feature dimension respectively; The joint loss function L is constructed by weighting the pixel loss term, the gradient loss term, and the perceptual loss term, and is expressed as follows: L=λ1L MSE +λ2L Gradient +λ3L Perceptual Among them, λ1, λ2, and λ3 represent the weighted coefficients of the pixel loss term, gradient loss term, and perceptual loss term, respectively, which are used to optimize the weights of the corresponding loss terms in the neural network.

8. The medical image fusion method based on multi-scale dilated convolution and channel attention mechanism according to claim 1, characterized in that: The inference processing of the fMRI and MRI-T1 image sample pairs based on the optimized parameters in step 7 to generate a fused image also includes the following steps: Step 7.1: Post-process the fused image including noise suppression, brightness adjustment and boundary cropping; Step 7.2: Use performance evaluation indicators to evaluate the quality of the fused image that has been post-processed in step 7.1 and output the quality evaluation results.

9. The medical image fusion method based on multi-scale dilated convolution and channel attention mechanism according to claim 8, characterized in that: The performance evaluation indicators described in step 7.2 include structural similarity index, peak signal-to-noise ratio, mutual information, visual information fidelity and entropy value.

Citation Information

Patent Citations

  • Medical image fusion method based on double contrast learning and gradient channel attention mechanism

    CN117788313A