Multi-modal three-dimensional medical image segmentation method based on attention mechanism
By introducing multi-head recursive spectral attention, progressive spectral attention and multi-scale feedforward network feature extraction methods, combined with the feature fusion and enhancement of SE attention and spatial attention mechanism, the problems of limited receptive fields, insufficient fusion of cross-modal feature and noise sensitivity in multimodal medical image segmentation are solved, and efficient and robust medical image segmentation is achieved.
Patent Information
- Application Number
- CN202510240739.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-07-08
AI Technical Summary
The existing medical image segmentation methods have problems such as limited receptive fields, insufficient cross-modal feature fusion, high computing resource requirements and noise sensitivity when processing multimodal medical images, resulting in limited segmentation accuracy and efficiency.
The multimodal three-dimensional medical image segmentation method based on attention mechanism is adopted, and through data preprocessing, feature extraction, feature fusion and feature enhancement modules, combined with multi-head recursive spectral attention, progressive spectral attention and multi-scale feedforward network, effective denoising and feature extraction of multimodal images is achieved, and feature fusion and enhancement are performed through SE attention and spatial attention mechanisms.
It improves the accuracy and efficiency of medical image segmentation, enhances the ability to understand complex structures, reduces computational complexity and resource requirements, and improves the robustness of the model.
Smart Images

Figure CN120279265A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image segmentation, and specifically to a multimodal three-dimensional medical image segmentation method based on an attention mechanism. Background Art
[0002] Medical image segmentation is one of the most challenging and valuable research areas in deep learning and artificial intelligence in recent years. In this regard, medical image segmentation plays a vital role as the basis of medical image analysis, especially in guiding lesion localization, lesion area analysis and surgical planning.
[0003] Competitive results have been achieved in segmentation accuracy due to the recent development of attention-based neural networks. These algorithms are usually trained on a single medical imaging modality, such as magnetic resonance images (MRI) or computed tomography (CT) images. Therefore, they often suffer from data variability when tested on images different from those seen during training. Data variability is a common problem in the medical field and is caused by some uncontrollable factors, such as different imaging methods. Some improved U-Net variants such as UNETR further improve the performance of image segmentation by introducing different attention mechanisms for multimodal feature extraction and multimodal feature fusion.
[0004] At present, the framework based on the attention mechanism has performed well in computer vision tasks. However, since medical image data usually has the characteristics of blur, noise, low contrast, etc., it is more difficult to extract the region of interest from medical images than normal RGB images.
[0005] Researchers have begun to combine CNNs with various attention mechanisms and proposed some hybrid architectures, such as the MI-Seg model, which uses a dual-branch encoder of CNN and Transformer for feature extraction, and the Transformer-based encoder is used to guide the training of another encoder and decoder to improve the accuracy and efficiency of medical image segmentation.
[0006] But in general, there are still the following technical problems:
[0007] 1. Limited receptive field, difficult to capture global context information. Segmentation methods based on traditional convolutional neural networks (such as UNet) usually rely on local convolution operations. Although they can effectively extract local features, their receptive field is limited and it is difficult to capture the global contextual relationship in multimodal medical images. This limitation is particularly unfavorable for the accurate segmentation of complex structures, and may cause the model to perform poorly when dealing with targets with complex boundaries or long-distance dependencies.
[0008] 2. Insufficient cross-modal feature fusion. Multimodal medical image segmentation requires the fusion of information from different modalities (such as CT and MRI), but existing methods still face challenges in the alignment and fusion of cross-modal features. Some methods rely too much on simple splicing or weighting strategies and cannot fully explore the complementarity between modalities, resulting in limited effectiveness of fused information and segmentation performance.
[0009] 3. Increased computing and resource requirements. With the increase in model complexity, especially after the introduction of new architectures such as Transformer, the model's computing cost and memory requirements have increased significantly. This poses a great challenge to actual application scenarios with limited hardware resources (such as clinical environments), limiting the actual implementation of the method.
[0010] 4. Sensitivity to noise in medical image data. Due to the particularity of medical image data, it generally has problems of noise and low contrast. Existing methods have poor robustness to medical image data and cannot effectively remove noise.
[0011] Therefore, a new solution to the above problems needs to be proposed. Summary of the invention
[0012] The purpose of the present invention is to provide a multimodal three-dimensional medical image segmentation method based on an attention mechanism to solve the technical problems raised in the background technology.
[0013] To achieve the above object, the present invention provides the following technical solution: a multimodal three-dimensional medical image segmentation method based on an attention mechanism, comprising at least the following steps:
[0014] S1: data preprocessing, including but not limited to data enhancement, modality selection, cross-validation, bias field correction, and resampling and normalization;
[0015] S2: Building a feature extraction module, the feature extraction module includes a potential feature extractor and a UNet-based encoder, the potential feature extractor is used to denoise the image and guide the encoder to extract features, and the UNet-based encoder is a dual-branch encoder;
[0016] S3: Building a feature fusion module through the SE attention mechanism and the spatial attention mechanism, the feature fusion module is used to fuse the denoising features with the encoder generated features;
[0017] S4: Build the dfe feature enhancement module to enhance the target features based on the attention mechanism;
[0018] S5: Output the enhanced image.
[0019] Furthermore, the S1 at least comprises the following steps:
[0020] S1.1: First, perform data augmentation by applying augmentation techniques to improve the robustness of the model. The augmentation techniques include, but are not limited to, cropping, rotation, scaling, and correction;
[0021] S1.2: Second, perform modality selection. According to existing cross-modal segmentation research, select MRI as the auxiliary modality and CT as the target modality. Among them, MRI has better soft tissue contrast and can provide more effective information for cardiac substructure segmentation;
[0022] S1.3: Then, perform cross-validation. Uniformly and randomly split the CT data, implement double-fold cross-validation, and train the conditional model. Use 20 MRI and 10 CT samples in each training to simulate the lack of data in the target modality;
[0023] S1.4: Then, perform bias field correction. Apply the N4 bias field correction algorithm to process the MRI data to correct the low-frequency intensity inhomogeneity in the image;
[0024] S1.5: Finally, perform resampling and normalization. Resample all 3D images to an isotropic space and normalize the pixel values to the range of 0-1 to meet the requirements of subsequent processing.
[0025] Furthermore, one of the encoders of the dual-branch encoder is designed based on the convolutional structure of the baseline, and the other encoder performs denoising processing on the input multi-modal image by introducing a spectral attention module, that is, the MAB module;
[0026] The core of the MAB module includes a multi-head recursive spectral attention module, a progressive spectral attention module, and a multi-scale feed-forward network module. The hybrid attention module composed of the multi-head recursive spectral attention module and the progressive spectral attention module is used to explore the inter-spectral and intra-spectral correlations, and the multi-scale feed-forward network module is used for the fusion of the multi-head recursive spectral attention module and the progressive spectral attention module.
[0027] Furthermore, the multi-head recursive spectral attention module calculates the inter-spectral correlation inside the image by introducing a spectral attention mechanism. The multi-head recursive spectral attention module dynamically calculates the weights of the pixels along the spectral direction for each band and averages them to use for denoising;
[0028] For the input, the multi-head recursive spectral attention module passes it through two multi-layer perceptron modules (MLP) respectively, and then uses two different activation functions to transform the input features to generate different weights;
[0029] Then, the two transformed input features are cumulatively merged to form a hybrid spectral attention. The merging process can fuse the features of all previous bands, enabling the multi-head recursive spectral attention module to correlate the spectral features between images and potentially utilize the information from cleaner bands for denoising;
[0030] Specifically, for the input intermediate feature map After passing through two multi-layer perceptrons and different activation functions respectively, two tensors Z and W' will be generated, as shown in the following formula:
[0031] Z = tanh(MLP(F)) (1)
[0032] W' = sigmoid(MLP(F)) (2)
[0033] MLP(X) = W'1·(tanh(W'2·X)) (3)
[0034] Where, In C represents the number of channels, D represents the spectral dimension, and H and W are the height and width of the image respectively;
[0035] In formula (1), the tanh activation function is used to perform a non-linear transformation on the result obtained through MLP(F) to generate the tensor Z;
[0036] In formula (2), the sigmoid activation function is used to process the same input F to generate the tensor W';
[0037] In formula (3), W'1, W'1 and W'2 are weight matrices, and a non-linear output is obtained through the tanh function after linear transformation;
[0038] The processes in formulas (1) to (3) can be equivalently considered as the query, key, and value projections in the self-attention mechanism;
[0039] After that, an attention operation is performed through a cyclic merging step, which only requires linear memory and time complexity instead of quadratic, and is more suitable for three-dimensional data. Specifically, the cyclic merging step of spectral mixing is achieved by cumulatively merging the weight W' and Z, and this process is defined as:
[0040] O i = (1 - W' i ) ⊙ Z i + W' ⊙ O i-1 (4)
[0041] Where, O i , W' i and Z iThey are the output feature, candidate feature, and merging weight for the i-th merging respectively.
[0042] Furthermore, the progressive spectral attention module focuses on exploring the correlation between images and thus ignores the intra-image correlation. The progressive spectral channel attention module introduces the channel attention in the compression excitation SENet and improves it into a pixel-level operation with a progressive attention pipeline.
[0043] The intermediate feature map processed by the self-multi-head recursive spectral attention module is subjected to three groups of convolution operations.
[0044] The first group of convolution is used to calculate the channel attention, the second group of convolution is used for channel dilation, and the third group of convolution is used for channel compression. This process is defined as:
[0045] F1 = Conv3d(x)·x + x (5)
[0046] F2 = GELU(Conv3d(F1)) (6)
[0047] F3 = Conv3d(F2) (7)
[0048] where GELU represents the activation function;
[0049] Formula (5) represents convolution and fusion, enhancing the model's judgment of the importance of features;
[0050] Formula (6) represents that the second group of convolution performs a convolution operation on the output of the first group and undergoes a non-linear change through the activation function, which can provide a stable training process while maintaining information flow;
[0051] Formula (7) represents that the third group of convolution processes the output of the second group, reducing the channel features after the convolution operation to a certain extent, which helps to overcome computational complexity and reduce memory consumption.
[0052] Furthermore, the multi-scale feed-forward network module can process features at different scales. The multi-scale feed-forward network module includes 3 parallel convolutional layers. By performing three groups of convolutions on the output from the progressive spectral attention module, adding them respectively, and finally generating the final multi-scale output through a group of convolution operations.
[0053] Furthermore, the feature fusion module in S3 fuses the features denoised by the self-latent feature extractor and the features extracted by the encoder through the SE attention mechanism and the spatial attention mechanism.
[0054] The SE attention mechanism is used to capture the interdependencies between channels, and the subsequent spatial attention mechanism is used to model the semantic dependencies in the spatial dimension;
[0055] The SE attention mechanism includes global average pooling, a fully connected layer, and a Sigmoid activation function. The weights generated by the SE attention mechanism will be divided into two parts and respectively assigned to the two input features from the latent feature extractor and the encoder. This process can ensure that each feature channel obtains the corresponding weight, thereby reflecting the importance of different features during the fusion process;
[0056] The spatial attention mechanism mainly captures spatial dependencies by using parallel operations and concatenation operations;
[0057] F′1 / F′2 = Conv3d(LeakyReLU(Conv3d(x)))(8)
[0058] F′3,F′4 = split(softmax(Concat(F′1,F′2)))(9)
[0059] F′ = F′3×F′1 + F′4×F′2(10)
[0060] Among them, formula (8) means that the input feature map x is processed through a 3D convolutional layer and the Leaky ReLU activation function is used. At this time, F′1 and F′2 are the features extracted from the input feature map and obtained after two convolutional operations;
[0061] Formula (9) means that F′1 and F′2 are combined through a concatenation operation, and then the softmax function is applied to the combined features to generate a new feature weight distribution. The split operation reduces the number of channels to the original Thereby obtaining two branch features F′3 and F′4;
[0062] Formula (10) means that the final fused feature F′ is obtained by multiplying F′1 and F′2 with the corresponding weights F′3 and F′4 and then adding them. This weighted fusion method can effectively integrate features from different sources and improve the overall performance of the model.
[0063] Furthermore, the S4 at least includes the following steps:
[0064] Set the denoised feature from the latent feature extractor as F l and the feature output by the encoder as F c , and the dfe feature enhancement module will enhance the target feature based on the attention mechanism. The process is defined as follows:
[0065] F ″1 = Flatten(Conv3d(F l )) × Flatten(F c ))(11)
[0066] F ″ 2 = Softmax(Max(F ″ 1) - F ″ 1) × Flatten(F c ))(12)
[0067] F ″ = Conv3d(Reshape(F ″ 2)))(13)
[0068] Among them, Flatten represents the flattening operation, Max represents the maximum value operation, and Reshape represents the dimension reshaping operation.
[0069] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0070] 1. High - efficiency denoising and feature extraction: By introducing a denoising and feature extraction method based on the spectral attention mechanism, using the multi - head recursive spectral attention module, progressive spectral attention, and multi - scale feed - forward network, the present invention effectively captures the long - distance dependencies in spectral information, improves the understanding ability of spectral features, and thus enhances the accuracy and effectiveness of feature extraction.
[0071] 2. Optimized feature fusion: The feature fusion module of the present invention based on channel and spatial attention mechanisms, using SE attention and spatial attention strategies, realizes the effective fusion of denoised features and features extracted by the encoder; this fusion strategy enhances the model's ability to capture key features by establishing mutual dependence relationships between feature channels and spatial dimensions.
[0072] 3. Enhanced segmentation efficiency: The improved feature enhancement module based on the attention mechanism of the present invention significantly improves the segmentation accuracy by performing feature interaction between fine - grained features and coarse - grained features; this module uses attention maps to strengthen the correlation between features and achieves a more accurate segmentation result. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0074] Figure 1 It is the overall framework diagram of the MMSegNet network of the present invention;
[0075] Figure 2 Schematic diagram of the multi-head recursive spectral attention module of the present invention;
[0076] Figure 3 Schematic diagram of the progressive spectral channel attention module of the present invention;
[0077] Figure 4 Schematic diagram of the feature fusion module of the present invention;
[0078] Figure 5 Schematic diagram of the feature enhancement module of the present invention. Detailed implementation manners
[0079] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0080] Please refer to Figure 1 , the entire medical image segmentation model of the present invention adopts the MMSegNet model, and a multi-modal three-dimensional medical image segmentation method based on the attention mechanism is proposed based on the MMSegNet model, which at least includes the following steps:
[0081] S1: Data preprocessing, which includes but is not limited to data augmentation, modality selection, cross-validation, bias field correction, and resampling and normalization;
[0082] S2: Build a feature extraction module, which includes a latent feature extractor and a UNet-based encoder. The latent feature extractor is used to denoise the image and guide the encoder to extract features. The UNet-based encoder is a dual-branch encoder;
[0083] S3: Build a feature fusion module through the SE attention mechanism and the spatial attention mechanism. The feature fusion module is used to fuse the denoised features and the features generated by the encoder;
[0084] S4: Build a dfe feature enhancement module to enhance the target features based on the attention mechanism;
[0085] S5: Output the enhanced image.
[0086] S1 at least includes the following steps:
[0087] S1.1: First, perform data augmentation, apply augmentation techniques to improve the robustness of the model. The augmentation techniques include but are not limited to cropping, rotation, scaling, and correction;
[0088] S1.2: Secondly, perform modality selection. According to existing cross-modal segmentation research, select MRI as the auxiliary modality and CT as the target modality. Among them, MRI has better soft tissue contrast and can provide more effective information for cardiac substructure segmentation;
[0089] S1.3: Connect the wires to perform cross-validation. Uniformly and randomly segment the CT data, implement double-fold cross-validation, and train the conditional model. Use 20 MRI and 10 CT samples in each training to simulate the lack of data in the target modality;
[0090] S1.4: Then perform bias field correction. Apply the N4 bias field correction algorithm to process the MRI data to correct the low-frequency intensity inhomogeneity in the image;
[0091] S1.5: Finally, perform resampling and normalization. Resample all 3D images into an isotropic space and normalize the pixel values to the range of 0-1 to meet the requirements of subsequent processing.
[0092] One of the encoders of the double-branch encoder is designed based on the convolutional structure of the baseline, and the other encoder performs denoising on the input multi-modal image by introducing a spectral attention module, that is, the MAB module;
[0093] The core of the MAB module includes a multi-head recursive spectral attention module, a progressive spectral attention module, and a multi-scale feed-forward network module. The hybrid attention module composed of the multi-head recursive spectral attention module and the progressive spectral attention module is used to explore the inter-spectral and intra-spectral correlations, and the multi-scale feed-forward network module is used for the fusion of the multi-head recursive spectral attention module and the progressive spectral attention module.
[0094] Refer to Figure 2 , the multi-head recursive spectral attention module calculates the inter-spectral correlation inside the image by introducing a spectral attention mechanism. The multi-head recursive spectral attention module dynamically calculates the weights of the pixels along the spectral direction for each band and averages them to use it for denoising;
[0095] For the input, the multi-head recursive spectral attention module passes it through two multi-layer perceptron modules (MLP) respectively, and then uses two different activation functions to transform the input features to generate different weights;
[0096] Then, the two transformed input features are cumulatively merged to form a mixed spectral attention. The merging process can fuse the features of all previous bands, enabling the multi-head recursive spectral attention module to associate the spectral features between images and potentially use the information from cleaner bands for denoising;
[0097] Specifically, for the input intermediate feature map After passing through two multi-layer perceptrons and different activation functions respectively, two tensors Z and W' will be generated, as shown in the following formula:
[0098] Z = tanh(MLP(F))(1)
[0099] W' = sigmoid(MLP(F))(2)
[0100] MLP(X) = W'1·(tanh(W'2·X))(3)
[0101] Where In it, C represents the number of channels, D represents the spectral dimension, and H and W are the height and width of the image respectively;
[0102] In formula (1), the tanh activation function is used to perform a non-linear transformation on the result obtained through MLP(F) to generate the tensor Z;
[0103] In formula (2), the sigmoid activation function is used to process the same input F to generate the tensor W';
[0104] In formula (3), W'1 W'1 and W'2 are weight matrices, and a non-linear output is obtained through the tanh function after linear transformation;
[0105] The processes in formulas (1) to (3) can be equivalently considered as the query, key, and value projections in the self-attention mechanism;
[0106] After that, an attention operation is performed through a loop merging step, which only requires linear memory and time complexity instead of quadratic, and is more suitable for three-dimensional data. Specifically, the loop merging step of spectral mixing is achieved by accumulating and merging the weights W' and Z, and this process is defined as:
[0107] O i = (1 - W' i )⊙Z i + W'⊙O i-1 (4)
[0108] Where, O i 、W' i and Z i are the output feature, candidate feature, and merging weight of the i-th merge respectively.
[0109] See Figure 3, the progressive spectral attention module focuses on exploring the correlation between images and thus ignores the intra-image correlation. The progressive spectral channel attention module introduces the channel attention in the squeeze-and-excitation SENet and improves it into a pixel-level operation with a progressive attention pipeline;
[0110] The intermediate feature map processed by the self-multi-head recursive spectral attention module is subjected to three groups of convolution operations;
[0111] The first group of convolutions is used to calculate the channel attention, the second group of convolutions is used for channel dilation, and the third group of convolutions is used for channel compression. This process is defined as:
[0112] F1 = Conv3d(x)·x + x (5)
[0113] F2 = GELU(Conv3d(F1)) (6)
[0114] F3 = Conv3d(F2) (7)
[0115] where GELU represents the activation function;
[0116] Equation (5) represents convolution and fusion, enhancing the model's judgment of the importance of features;
[0117] Equation (6) represents that the second group of convolutions performs a convolution operation on the output of the first group and undergoes a non-linear transformation through the activation function, which can provide a stable training process while maintaining information flow;
[0118] Equation (7) represents that the third group of convolutions processes the output of the second group, reducing the channel features after the convolution operation to a certain extent, which helps to overcome computational complexity and reduce memory consumption.
[0119] The multi-scale feed-forward network module can process features at different scales. The multi-scale feed-forward network module includes 3 parallel convolutional layers. By performing three groups of convolutions on the output from the progressive spectral attention module, adding them respectively, and finally generating the final multi-scale output through a group of convolution operations.
[0120] Refer to Figure 4 , the feature fusion module in S3 fuses the features after denoising by the self-latent feature extractor and the features extracted by the encoder through the SE attention mechanism and the spatial attention mechanism;
[0121] The SE attention mechanism is used to capture the interdependencies between channels, and the subsequent spatial attention mechanism is used to model the semantic dependencies in the spatial dimension;
[0122] The SE attention mechanism includes global average pooling, a fully connected layer, and a Sigmoid activation function. The weights generated by the SE attention mechanism will be divided into two parts and respectively assigned to the two input features from the latent feature extractor and the encoder. This process can ensure that each feature channel obtains corresponding weights, thereby reflecting the importance of different features during the fusion process;
[0123] The spatial attention mechanism mainly uses parallel operations and cascaded operations to capture spatial dependencies;
[0124] F′1 / F′2 = Conv3d(LeakyReLU(Conv3d(x)))(8)
[0125] F′3,F′4 = split(softmax(Concat(F′1,F′2)))(9)
[0126] F′ = F′3×F′1 + F′4×F′2(10)
[0127] Among them, formula (8) means that the input feature map x is processed through a 3D convolutional layer, and the Leaky ReLU activation function is used. At this time, F′1 and F′2 are the features extracted from the input feature map and obtained after two convolutional operations;
[0128] Formula (9) means that F′1 and F′2 are combined through a concatenation operation, and then the softmax function is applied to the combined features to generate a new feature weight distribution. The split operation reduces the number of channels to the original so as to obtain two branch features F′3 and F′4;
[0129] Formula (10) means that the final fused feature F′ is obtained by multiplying F′1 and F′2 by the corresponding weights F′3 and F′4 and then adding them. This weighted fusion method can effectively integrate features from different sources and improve the overall performance of the model.
[0130] Refer to Figure 5 , S4 at least includes the following steps:
[0131] Set the denoised feature from the latent feature extractor as F l and the feature output by the encoder as F c , the dfe feature enhancement module will enhance the target feature based on the attention mechanism, and the process is defined as follows:
[0132] F ″ 1 = Flatten(Conv3d(F l )) × Flatten(F c ))(11)
[0133] F ″ 2 = Softmax(Max(F ″ 1) - F ″ 1) × Flatten(F c ))(12)
[0134] F ″ = Conv3d(Reshape(F ″ 2)))(13)
[0135] Among them, Flatten represents the flattening operation, Max represents the maximum value operation, and Reshape represents the dimension reshaping operation.
[0136] 1) Performance evaluation
[0137] We evaluated the proposed method on the Multimodal Whole Heart Segmentation Challenge 2017 (MMWHS2017) dataset, which contains 20 non-registered MRIs and 20 CT 3D images for training, as well as ground truth (GT) annotations of 7 cardiac substructures including the left ventricle (LV), right ventricle (RV), left atrium (LA), right atrium (RA), left ventricular myocardium (Myo), ascending aorta (AA), and pulmonary artery (PA). This invention also follows the previous work and reports the results of the segmentation targets and the average results on the validation set.
[0138] This invention compared the proposed method with the current state-of-the-art convolution- and Transformer-based methods. The evaluation metric is the Dice score, which is expressed as a percentage. Table 1 shows the results of the segmentation on the MMWHS validation set.
[0139] Table 1 Comparative experimental results on the MMWHS dataset
[0140]
[0141]
[0142] As shown in Table 1, in the comparative experimental results on the MMWHS dataset, our method performed relatively well, with an average Dice score of 88.77, second only to the MOSMOS method. Specifically, in the segmentation tasks of the seven cardiac substructures, our method achieved relatively high Dice scores. Among them, the segmentation effect of AA was particularly prominent, with a Dice score as high as 95.05%, ranking first among all methods. In addition, our method also achieved relatively high accuracy in the segmentation of LV, RV, LA, and RA.
[0143] Although the Ours method is slightly lower than the MOSMOS method in terms of the average Dice score, with its unique innovative points, the Ours method demonstrates high accuracy and robustness when facing the complex heart structure segmentation task. The MOSMOS method adopted contrastive learning in the pre-training stage and was optimized in aspects such as learning pixel-label attention maps based on the Transformer decoder in the fine-tuning stage. However, our method realizes effective denoising of the input data and deep feature extraction by introducing innovative modules such as multi-head recursive spectral attention, progressive spectral attention, and multi-scale feed-forward networks. These modules not only enhance the model's ability to capture spectral information but also improve the model's understanding ability of complex heart structures. In addition, the Ours method also uses a feature fusion module, which can intelligently integrate feature information from different modules and levels to form a richer and more comprehensive feature representation. This comprehensive processing strategy enables the Ours method to better handle various challenges and maintain high accuracy and robustness when facing the complex heart structure segmentation task.
[0144] We demonstrated the effectiveness of each innovative module of the model by visualizing the segmentation performance through comparing the segmentation results of different modules on the dataset. Table 2 shows the results of segmentation on the MMWHS validation set.
[0145] Table 2 Results of ablation experiments on the MMWHS dataset
[0146]
[0147]
[0148] The results of the ablation experiments show that with the gradual addition of the MAB module, the feature fusion strategy, and the dfe module proposed in the present invention, the heart segmentation performance of the model on the MMWHS dataset shows a significant and continuous improvement. Specifically, from the average Dice coefficient of 82.44% when only using the baseline model (MI-seg), it is improved to 84.89% after adding the MAB module, and further increased to 87.56% through the feature fusion strategy. Finally, after integrating the dfe module proposed in the present invention, the average Dice coefficient reaches 88.77%. This series of data not only clearly demonstrates the positive impact of each module on the model performance but also strongly verifies the effectiveness of each module and the significant superiority of the method proposed in the present invention. This gradual performance improvement not only reflects the progress of the model's accuracy in the heart segmentation task but also highlights the synergistic effect of each component in improving the overall performance.
[0149] To sum up:
[0150] The present invention uses the MAB module, which consists of three parts: multi-head recursive spectral attention, progressive spectral attention, and multi-scale feed-forward network, in the encoder to achieve effective denoising and feature extraction of the input data. To effectively fuse the high-level and low-level features of the dual-branch encoder, the present invention introduces a feature fusion module with channel attention and spatial attention. At the same time, to improve the segmentation efficiency, the dfe feature enhancement module improved based on the attention mechanism is used to replace the bottleneck layer in the encoder-decoder structure, enabling it to effectively fuse the input features and the rough segmentation results, and enhancing the model's ability to capture key features. The experimental results of the present invention show that the proposed method performs better than the state-of-the-art methods on the MMWHS multi-modal dataset, highlighting potential future research directions and improving the accuracy of medical image segmentation.
[0151] Among them, the feature extraction method based on spectral attention, the feature fusion module based on channel and spatial attention mechanisms, and the dfe feature enhancement module improved based on the attention mechanism are the key technologies of the present invention.
[0152] 1) Denoising and feature extraction method based on spectral attention mechanism
[0153] The overall operation process of the denoising and feature extraction method based on spectral attention mechanism (as shown in Figure 2 and Figure 3 ): The multi-head recursive spectral attention module can deeply capture the long-range dependencies in the spectral information in parallel on multiple feature heads, enhancing the model's ability to understand spectral features; the progressive spectral attention extracts spectral features in a step-by-step refinement manner, capturing key information from coarse to fine; the multi-scale feed-forward network module can perform convolution operations on the input data at different scales, fusing multi-scale feature information.
[0154] 2) Feature fusion module based on channel and spatial attention mechanisms
[0155] This module adopts a fusion strategy based on SE attention and spatial attention for the features denoised by the potential feature extractor and the features extracted by the encoder in the previous stage. The SE attention mechanism is mainly responsible for capturing the interdependencies between feature channels, while the spatial attention mechanism models the semantic dependencies of features in the spatial dimension. Specifically, the SE module consists of global average pooling, fully connected layers, and a Sigmoid activation function, and can generate weights representing the importance of channels. These weights are then divided into two parts, respectively used to adjust the channel weights of the denoised features and the features extracted by the encoder in the previous stage, thus realizing the effective fusion of features.
[0156] 3) Feature enhancement module improved based on attention mechanism
[0157] This module mainly uses the attention mechanism to enhance the feature interaction between the fine-grained features and the coarse-grained features of the input, thereby improving the accuracy of segmentation. Specifically, the fine-grained features and the coarse-grained features are flattened respectively, and the attention map between the fine-grained features and the coarse-grained features is obtained through matrix multiplication and normalization, and then matrix multiplication is performed to obtain the attention-enhanced features. Finally, the features are reshaped.
[0158] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
Claims
1. A multi-modal three-dimensional medical image segmentation method based on an attention mechanism, characterized in that: At least include the following steps: S1: Data preprocessing, which includes but is not limited to data augmentation, modality selection, cross-validation, bias field correction, and resampling and normalization; S2: Build a feature extraction module, which includes a latent feature extractor and a UNet-based encoder. The latent feature extractor is used to denoise the image and guide the encoder to extract features. The UNet-based encoder is a dual-branch encoder; S3: Build a feature fusion module through the SE attention mechanism and the spatial attention mechanism. The feature fusion module is used to fuse the denoised features and the features generated by the encoder; S4: Build a dfe feature enhancement module to enhance the target features based on the attention mechanism; S5: Output the enhanced image.
2. The multi-modal three-dimensional medical image segmentation method based on the attention mechanism according to claim 1, wherein: The S1 at least includes the following steps: S1.1: First, perform data augmentation and apply augmentation techniques to improve the robustness of the model. The augmentation techniques include but are not limited to cropping, rotation, scaling, and correction; S1.2: Secondly, perform modality selection. According to the existing cross-modal segmentation research, select MRI as the auxiliary modality and CT as the target modality. Among them, MRI has better soft tissue contrast and can provide more effective information for cardiac substructure segmentation; S1.3: Then perform cross-validation. Uniformly and randomly split the CT data, implement double-fold cross-validation, and train the conditional model. Use 20 MRIs and 10 CT samples in each training to simulate the lack of data in the target modality; S1.4: Then perform bias field correction and apply the N4 bias field correction algorithm to process the MRI data to correct the low-frequency intensity inhomogeneity in the image; S1.5: Finally, perform resampling and normalization. Resample all 3D images to an isotropic space and normalize the pixel values to the range of 0-1 to meet the requirements of subsequent processing.
3. The multimodal three-dimensional medical image segmentation method based on the attention mechanism according to claim 1, wherein: One of the encoders of the dual-branch encoder is designed based on the convolutional structure of the baseline, and the other encoder performs denoising processing on the input multi-modal image by introducing a spectral attention module, that is, the MAB module; The core of the MAB module includes a multi-head recursive spectral attention module, a progressive spectral attention module, and a multi-scale feed-forward network module. The hybrid attention module composed of the multi-head recursive spectral attention module and the progressive spectral attention module is used to explore the inter-spectral and intra-spectral correlations. The multi-scale feed-forward network module is used for the fusion of the multi-head recursive spectral attention module and the progressive spectral attention module.
4. The multimodal three-dimensional medical image segmentation method based on the attention mechanism according to claim 3, wherein: The multi-head recursive spectral attention module calculates the inter-spectral correlation inside the image by introducing the spectral attention mechanism. The multi-head recursive spectral attention module dynamically calculates the weights of the pixels along the spectral direction for each band and averages them to use for denoising; For the input, the multi-head recursive spectral attention module passes it through two multi-layer perceptron modules respectively, and then uses two different activation functions to transform the input features to generate different weights; Then, the two transformed input features are cumulatively merged to form a hybrid spectral attention. The merging process can fuse the features of all previous bands, enabling the multi-head recursive spectral attention module to correlate the spectral features between images and potentially utilize the information from cleaner bands for denoising; Specifically, for the input intermediate feature map After passing through two multi-layer perceptrons and different activation functions respectively, two tensors Z and W' will be generated, as shown in the following formula: Z = tanh(MLP(F)) (1) W′ = sigmoid(MLP(F)) (2) MLP(X) = W′1·(tanh(W′2·X)) (3) Among them, where C represents the number of channels, D represents the spectral dimension, and H and W are the height and width of the image respectively; In formula (1), the tanh activation function is used to perform a non-linear transformation on the result obtained through MLP(F) to generate the tensor Z; In formula (2), the sigmoid activation function is used to process the same input F to generate the tensor W′; In formula (3) W′1 and W′2 are weight matrices, and after linear transformation, a non-linear output is obtained through the tanh function; The processes in formulas (1) to (3) can be equivalently regarded as the query, key, and value projections in the self-attention mechanism; Subsequently, an attention operation is performed through a cyclic merging step, which only requires linear memory and time complexity instead of quadratic, and is more suitable for three-dimensional data. Specifically, the cyclic merging step of spectral mixing is achieved by cumulatively merging the weight W′ and Z, and this process is defined as: O i = (1 - W′ i ) ⊙ Z i + W′ ⊙ O i-1 (4) Among them, O i , W′ i and Z i are the output feature, candidate feature, and merging weight of the i-th merging respectively.
5. The multi-modal three-dimensional medical image segmentation method based on the attention mechanism according to claim 4, wherein: The progressive spectral attention module focuses on exploring the inter-image correlation, thus ignoring the intra-image correlation. The progressive spectral channel attention module introduces the channel attention in the compression excitation SENet and improves it to a pixel-level operation with a progressive attention pipeline; The intermediate feature map processed by the self-multi-head recursive spectral attention module Perform three groups of convolution operations; The first group of convolutions is used to calculate the channel attention, the second group of convolutions is used for channel dilation, and the third group of convolutions is used for channel compression. This process is defined as: F1 = Conv3d(x)·x + x (5) F2 = GELU(Conv3d(F1)) (6) F3 = Conv3d(F2) (7) Among them, GELU represents the activation function; Formula (5) represents convolution and fusion, enhancing the model's judgment of the importance of features; Formula (6) represents that the second group of convolutions performs a convolution operation on the output of the first group and undergoes a non-linear change through the activation function, which can provide a stable training process while maintaining the information flow; Formula (7) represents that the third group of convolutions processes the output of the second group, reducing the channel features after the convolution operation to a certain extent, which helps to overcome the computational complexity and reduce the memory consumption.
6. The multi-modal three-dimensional medical image segmentation method based on the attention mechanism according to claim 5, characterized in that: The multi-scale feed-forward network module can process features at different scales. The multi-scale feed-forward network module includes 3 parallel convolutional layers. By performing three groups of convolutions on the output from the progressive spectral attention module, adding them respectively, and finally generating the final multi-scale output through a group of convolution operations.
7. The multi-modal three-dimensional medical image segmentation method based on the attention mechanism according to claim 6, wherein: The feature fusion module in S3 fuses the features denoised by the self-potential feature extractor and the features extracted by the encoder through the SE attention mechanism and the spatial attention mechanism; The SE attention mechanism is used to capture the interdependencies between channels, and the subsequent spatial attention mechanism is used to model the semantic dependencies in the spatial dimension; The SE attention mechanism includes global average pooling, a fully connected layer, and a Sigmoid activation function. The weights generated by the SE attention mechanism are divided into two parts and respectively assigned to the two input features from the latent feature extractor and the encoder. This process ensures that each feature channel obtains corresponding weights, thereby reflecting the importance of different features during the fusion process; The spatial attention mechanism mainly captures spatial dependency relationships by using parallel operations and cascading operations; F′1 / F′2 = Conv3d(LeakyReLU(Conv3d(x))) (8) F′3,F′4 = split(softmax(Concat(F′1,F′2))) (9) F′ = F′3×F′1 + F′4×F′2 (10) Among them, formula (8) indicates that the input feature map x is processed through a 3D convolutional layer and the Leaky ReLU activation function is used. At this time, F′1 and F′2 are the features extracted from the input feature map and obtained after two convolutional operations; Equation (9) indicates that F′1 and F′2 are combined through a concatenation operation, and then the softmax function is applied to the combined features to generate a new feature weight distribution. The split operation reduces the number of channels to the original so as to obtain two branch features F′3 and F′4; Formula (10) indicates that the final fused feature F′ is obtained by multiplying F′1 and F′2 with the corresponding weights F′3 and F′4 and then adding them. This weighted fusion method can effectively integrate features from different sources and improve the overall performance of the model.
8. The multimodal three-dimensional medical image segmentation method based on the attention mechanism according to claim 7, wherein: S4 described above at least includes the following steps: Set the denoised features from the potential feature extractor as F l and the features output by the encoder as F c , the dfe feature enhancement module will enhance the target features based on the attention mechanism, and the process is defined as follows: F″1 = Flatten(Conv3d(F l )) × Flatten(F c )) (11) F″2 = Softmax(Max(F″1) - F″1) × Flatten(F c )) (12) F″ = Conv3d(Reshape(F″2))) (13) Among them, Flatten represents the flattening operation, Max represents the maximum value operation, and Reshape represents the dimension reshaping operation.