MRI (Magnetic Resonance Imaging) brain tumor image segmentation method based on Conv + Swin and CSWin Transformer
By introducing Conv+Swin and CSWin Transformer structures into the TransUNet model, combining SENetV2 backbone network and BCEDiceLoss function, the feature extraction and computing efficiency problems in MRI brain tumor image segmentation are solved, and higher-precision image segmentation is achieved, supporting medical diagnosis and treatment.
Patent Information
- Application Number
- CN202510286393.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-04
AI Technical Summary
The existing MRI brain tumor image segmentation method has limitations in feature extraction diversity and comprehensiveness. The traditional Transformer encoder lacks efficiency and accuracy when processing local features, and has high computational cost, making it difficult to meet the high-precision and high-efficiency segmentation needs.
The Conv+Swin structure is used to replace the CNN structure in the TransUNet encoder, and the original Transformer encoder is replaced with the CSWin Transformer structure, combined with the SENetV2 backbone network for feature extraction, and the model is trained using the BCEDiceLoss function to optimize the loss function to improve segmentation accuracy.
The segmentation accuracy of MRI brain tumor images is significantly improved. By combining local and global feature extraction capabilities, the computing efficiency is optimized, and more efficient image segmentation effect is achieved and more reliable medical diagnostic support is provided.
Smart Images

Figure CN120259198A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image processing, and particularly relates to an MRI brain tumor image segmentation method based on Conv+Swin and CSWin Transformer. Background Art
[0002] In the field of medical image processing, medical image segmentation is a crucial task. Its purpose is to accurately divide different tissues, organs, or lesion areas in medical images, providing intuitive and detailed information for doctors to assist in diagnostic decisions, formulate treatment plans, and evaluate treatment effects, etc. Among many medical imaging technologies, magnetic resonance imaging (MRI) has been widely used in the diagnosis and treatment of brain tumors due to its excellent soft tissue resolution ability. Through MRI technology, doctors can obtain high-resolution brain images, clearly presenting the anatomical structure and lesion conditions of the brain, which plays a crucial role in the early detection, accurate diagnosis, and subsequent treatment of brain tumors.
[0003] However, the complexity of brain tumors makes MRI brain tumor image segmentation face many challenges. On the one hand, the morphology, size, location of brain tumors, and the boundaries with surrounding normal tissues show great diversity, making it extremely difficult to accurately define the tumor area. On the other hand, problems such as noise, artifacts, and partial volume effects in MRI images further increase the difficulty of image segmentation. In addition, although different modalities of MRI images (such as T1-weighted imaging, T1-weighted enhanced imaging, T2-weighted imaging, FLAIR imaging, etc.) can provide complementary information, how to effectively fuse this multi-modal information and give full play to its role in brain tumor segmentation is also one of the current research focuses and difficulties.
[0004] In recent years, deep learning technology has made remarkable progress in the field of medical image segmentation, and segmentation methods based on deep learning have gradually become the mainstream. Among them, TransUNet, as a model that combines Transformer and U-Net, has shown certain advantages in medical image segmentation tasks. Through the self-attention mechanism of Transformer, it can capture long-range dependencies in images, thus better extracting global features. However, the traditional TransUNet only uses the Conv structure in the CNN structure of the encoder, and there are certain limitations in the diversity and comprehensiveness of feature extraction, making it difficult to fully mine complex features in images. At the same time, the original Transformer as an encoder has room for improvement in the efficiency and accuracy of processing local features. Its attention calculation method has a high computational cost when dealing with large-scale image data and insufficient attention to local detail information.
[0005] In summary, the existing MRI brain tumor image segmentation methods still have many deficiencies and are difficult to meet the clinical requirements for high-precision and high-efficiency segmentation. Therefore, developing a more effective MRI brain tumor image segmentation method has important practical significance and clinical application value. Based on Conv+Swin and CSWin Transformer, the present invention improves TransUNet, aiming to overcome the defects of existing methods and enhance the accuracy and performance of MRI brain tumor image segmentation. Summary of the Invention
[0006] Based on TransUNet as the basic model, the present invention adopts the Conv+Swin structure in the encoder and uses the SENetV2 backbone network to extract features. The original Transformer encoder is replaced with the CSWin Transformer structure. After preprocessing the three-dimensional MRI brain tumor image dataset, the model is trained. The BCEDiceLoss function is used as the loss function, and the parameters are updated through backpropagation. Finally, the image to be segmented is input into the trained model to obtain the segmentation result.
[0007] To solve the above technical problems, the technical solution of the present invention is as follows:
[0008] An MRI brain tumor image segmentation method based on Conv+Swin and CSWin Transformer, comprising the following steps.
[0009] Data preprocessing: The three-dimensional MRI brain tumor image dataset is processed into two-dimensional slices to obtain a two-dimensional picture sequence of multiple feature channels (such as channels corresponding to T1-weighted imaging, T1-weighted enhanced imaging, T2-weighted imaging, and FLAIR imaging), and the pictures with pixel values of zero are deleted; then the two-dimensional picture sequence is normalized, and the formula is
[0010] b_{i}=(a_{i}-u) / s
[0011] where u is the average pixel value, a_{i} is the pixel value before normalization, and b_{i} is the pixel value after normalization; finally, the normalized two-dimensional picture sequence is centrally cropped to obtain the preprocessed dataset.
[0012] Constructing an improved model: Based on TransUNet as the basic model, in its encoder, the original CNN structure is replaced with the Conv+Swin structure, and SENetV2 is used as the backbone network for feature extraction; at the same time, the original Transformer encoder in the original model is replaced with the CSWin Transformer structure, and the self-attention is calculated in parallel on the horizontal and vertical stripes to expand the attention area, and the stripe width can be adjusted to balance the calculation cost and modeling ability.
[0013] Model training: Use the preprocessed dataset to train the improved segmentation model. Set the loss function as the BCEDiceLoss function, and its calculation formula is
[0014] L'(x,y) = L'_{dice}(x,y) + 0.5 * L'_{bce}(x,y),
[0015] L(x,y) = L_{whole}(x,y) + L_{core}(x,y) + L_{enh}(x,y)
[0016] where L'_{dice} is the dice loss function, L'_{bce} is the cross-entropy loss function, L_{whole} is the loss function for the entire tumor region, L_{core} is the loss function for the tumor core region, and L_{enh} is the loss function for the tumor enhancement region; update the model parameters through the backpropagation algorithm until the loss function converges to obtain the trained model.
[0017] Image segmentation: Input the MRI brain tumor image to be segmented into the trained model to obtain the image segmentation results of different regions such as the entire region containing the tumor, the tumor core region, and the tumor enhancement region, thereby realizing the accurate segmentation of the MRI brain tumor image.
[0018] Compared with the prior art, the advantages of the present invention are as follows:
[0019] Innovative model architecture: Current MRI brain tumor image segmentation technologies are mostly based on traditional models, such as the classic TransUNet. The present invention innovatively improves the basic model. In the encoder, a Conv+Swin structure is adopted, replacing the original single Conv structure, and at the same time, the original Transformer encoder is replaced with a CSWin Transformer structure. This unique architecture design is a breakthrough in traditional models and provides more powerful technical support for image segmentation.
[0020] More powerful feature extraction ability: Existing methods have limitations in feature extraction and are difficult to comprehensively capture image information. The present invention combines the local feature extraction advantages of the convolutional neural network with the global feature capture ability of Swin Transformer through the Conv+Swin structure, and can obtain richer and more comprehensive features from MRI brain tumor images. For example, when dealing with complex tumor boundaries and subtle features, the Conv part can accurately focus on local textures, while Swin Transformer can grasp the relationship between the tumor and surrounding tissues from a macroscopic perspective, which is difficult to achieve by the prior art.
[0021] Optimized computational efficiency: Traditional Transformer encoders have high computational costs and low efficiency in processing local features when dealing with large-scale image data. The CSWin Transformer structure of the present invention expands the attention area through parallel self-attention calculation, enhancing the modeling ability for local and global features. At the same time, by adjusting the stripe width, it can flexibly balance the computational cost and modeling ability in different network layers, significantly improving the computational efficiency while ensuring the segmentation accuracy.
[0022] Higher segmentation accuracy: Experimental results show that the present invention is significantly superior to the prior art in evaluation metrics such as the Dice coefficient. In model training, the BCEDiceLoss function adopted by the present invention comprehensively considers the losses of the entire tumor region, the tumor core region, and the tumor enhancement region, guiding the model to more accurately segment each tumor region and providing more reliable and accurate image segmentation results for medical diagnosis and treatment. Description of the Drawings
[0023] Figure 1 It is a schematic flowchart of the MRI brain tumor image segmentation method based on Conv+Swin and CSWin Transformer.
[0024] Figure 2 It is a structural diagram of the Swin-Transformer module in the MRI brain tumor image segmentation model.
[0025] Figure 3 It is a structural diagram of the SENetV2 module in the MRI brain tumor image segmentation model.
[0026] Figure 4 It is a structural diagram of the CSWin Transformer module in the MRI brain tumor image segmentation model.
[0027] Figure 5 It is a display diagram of the segmentation results of the MRI brain tumor image segmentation model. Detailed Embodiment
[0028] The following further elaborates on the present invention in conjunction with the drawings and embodiments, but does not limit the present invention.
[0029] In order to more clearly elaborate the objectives, technical solutions, and advantages of the present invention, the following will further detail the present invention through specific embodiments in conjunction with the drawings.
[0030] Figure 1 It shows the overall process of the MRI brain tumor image segmentation method based on Conv+Swin and CSWin Transformer, mainly including steps such as data preprocessing, constructing an improved model, model training, and image segmentation.
[0031] (1) Data preprocessing
[0032] Data processing steps: Perform two-dimensional slicing on the three-dimensional MRI brain tumor image dataset to obtain a sequence of two-dimensional images with multiple feature channels, and delete the images with pixel values of zero. Normalize the sequence of two-dimensional images to convert the image gray values into a distribution with the same mean and variance. Finally, perform central cropping to obtain the preprocessed dataset.
[0033] Function: These processing steps help reduce noise and interference in the data, enhance the consistency and stability of the data, and provide a better basis for subsequent model training and image segmentation.
[0034] (2) Construct an improved model
[0035] Model architecture: Based on TransUNet as the basic model, adopt the Conv+Swin structure in the encoder, use SENetV2 as the backbone network to extract features, and replace the original Transformer encoder with the CSWin Transformer structure.
[0036] Innovation: Multi-structure fusion. By combining different structures such as Conv, Swin Transformer, and CSWin Transformer, give full play to their respective advantages, and improve the model's feature extraction ability and modeling ability for MRI brain tumor images.
[0037] Feature extraction optimization: The SENetV2 backbone network can enhance the relationship modeling between channels, and the parallel computing self-attention mechanism of CSWinTransformer expands the attention area, which helps to capture global and local features, thereby improving the segmentation accuracy of the model.
[0038] (3) Model training
[0039] Loss function: Use the BCEDiceLoss function as the loss function. Its calculation formula comprehensively considers the Dice loss function and the cross loss function, and optimizes the losses of the entire tumor region, tumor core region, and tumor enhancement region at the same time.
[0040] Training process: Use the preprocessed dataset to train the improved segmentation model, update the model parameters through the backpropagation algorithm until the loss function converges, and obtain the trained model.
[0041] Optimization strategy: Adopt appropriate parameters such as optimizer, learning rate, total number of training epochs, and batch size to improve training efficiency and model performance.
[0042] (4) Image segmentation
[0043] Output result: The MRI brain tumor image to be segmented is input into the trained model, and the image segmentation results of different regions such as the entire region containing the tumor, the tumor core region, and the tumor enhancement region are obtained, realizing the precise segmentation of the MRI brain tumor image.
[0044] Application value: It provides detailed tumor information for doctors, which helps in auxiliary diagnosis decision-making, treatment plan formulation, and treatment effect evaluation.
[0045] Figure 2 The following is the structural diagram of the Swin-Transformer module in the present invention. This diagram shows the module structure of Swin Transformer, including parts such as the slicing process of the input image, patch embedding, and multiple Transformer blocks.
[0046] (1) Key components
[0047] Patch segmentation: First, the input image is segmented into non-overlapping patches, and each patch is regarded as a "token", and its features are projected into any dimension through a linear embedding layer.
[0048] Transformer block shifted window self-attention: The self-attention mechanism with shifted windows is adopted. By alternately using different window partitioning configurations in consecutive Transformer blocks, the connection between windows is realized, enhancing the modeling ability of the model.
[0049] Multi-layer perceptron: A module composed of two-layer MLP is connected after each Transformer block, which is used for further processing and non-linear transformation of the features.
[0050] Hierarchical structure: Through processing in multiple stages, the resolution of the image is gradually reduced, generating representations with different hierarchical features, and finally realizing multi-scale feature extraction of the MRI brain tumor image.
[0051] (2) Advantages
[0052] Efficient feature extraction: The shifted window self-attention mechanism can effectively calculate self-attention without losing too much information, improving the computational efficiency of the model.
[0053] Multi-scale modeling: The hierarchical structure enables the model to model the image at different scales, capturing different hierarchical features, thereby improving the segmentation accuracy.
[0054] Figure 3 The following is the structural diagram of the SENetV2 module in the present invention. This diagram shows the module structure of SENetV2, including parts such as the squeeze operation, excitation operation, and multi-branch dense layer.
[0055] (1) Key steps
[0056] Squeeze operation: Compress the input features through a global average pooling layer, extract channel-level statistical information, and then reduce the dimension through a fully connected layer to obtain the compressed features.
[0057] Excitation operation: Use a fully connected layer to restore the compressed features to the original feature dimension, and then through a scaling operation, perform element-wise multiplication of the output and the input features to achieve re-weighting of the features.
[0058] Multi-branch dense layer: After the squeeze operation, introduce a multi-branch dense layer, and further process and fuse the features through fully connected layers with multiple branches, enhancing the model's ability to learn global features.
[0059] (2) Advantages
[0060] Enhance feature expression ability: SENetV2 can effectively capture the relationships between channels through squeeze and excitation operations, enhance the feature expression ability, and thus improve the performance of the model.
[0061] Optimize the network structure: The introduction of the multi-branch dense layer increases the flexibility and diversity of the network, helps to optimize the network structure, and improves the training efficiency and generalization ability of the model.
[0062] Figure 4 Shown is the structural diagram of the CSWin Transformer module in the present invention. This diagram shows the module structure of CSWin Transformer, including the division of input features, horizontal and vertical stripe self-attention calculations, multi-head grouping, etc.
[0063] (1) Key components
[0064] Cross-shaped window self-attention: Divide the input features into horizontal and vertical stripes, and perform self-attention calculations separately, improving the calculation efficiency of self-attention through parallel computing. At the same time, by adjusting the stripe width, a balance between calculation cost and modeling ability is achieved.
[0065] Multi-head grouping: Divide the multi-heads into different groups, apply different self-attention operations to different groups of heads, expand the attention area of each token within the Transformer block, and enhance the model's ability to model features.
[0066] Local Enhanced Position Encoding: Local Enhanced Position Encoding (LePE) is introduced, adding position information as a parallel module to the self-attention operation, directly operating on the projected values, and enhancing the model's ability to process local position information.
[0067] (2) Advantages
[0068] Efficient self-attention calculation: The cross-shaped window self-attention mechanism can effectively reduce the computational cost while maintaining a strong modeling ability, improving the running efficiency of the model.
[0069] Powerful feature modeling ability: The combination of multi-head grouping and local enhanced position encoding enables the model to better capture local and global features of the image, improving the segmentation accuracy.
[0070] Figure 5 Shown is a display diagram of the MRI brain tumor image segmentation result. This diagram shows the result of segmenting the MRI brain tumor image using the segmentation method of the present invention, including the annotation of different regions such as the entire tumor region, the tumor core region, and the tumor enhancement region.
Claims
1. A method for MRI brain tumor image segmentation based on Conv+Swin and CSWin Transformer, characterized in that, It includes the following steps: (1) Preprocess the three-dimensional MRI brain tumor image dataset, including slicing the three-dimensional MRI data into two-dimensional slices to obtain a two-dimensional picture sequence with multiple feature channels, and deleting the pictures with pixel values of zero; perform normalization processing on the two-dimensional picture sequence; Perform central cropping on the normalized two-dimensional picture sequence to obtain the preprocessed dataset. (2) Construct an improved segmentation model, using TransUNet as the basic model. In its encoder, replace the original CNN structure with a Conv+Swin structure, and use SENetV2 as the backbone network for feature extraction; Replace the original version of the Transformer encoder in the original model with a CSWin Transformer structure. (3) Use the preprocessed dataset to train the improved segmentation model, set the loss function, and update the model parameters through the backpropagation algorithm until the loss function converges to obtain the trained model. (4) Input the MRI brain tumor image to be segmented into the trained model to obtain the image segmentation result.
2. A method for MRI brain tumor image segmentation based on Conv+Swin and CSWin Transformer according to claim 1, characterized in that: In step 1, the feature channels obtained by slicing the three-dimensional MRI data into two-dimensional slices include the channels corresponding to T1-weighted imaging, T1-weighted enhanced imaging, T2-weighted imaging, and FLAIR imaging; the formula used for normalization processing is b_{i}=(a_{i}-u) / s where u is the average value of the pixel points, a_{i} is the value of each pixel point before normalization, and b_{i} is the value of each pixel point after normalization; central cropping is to crop the picture size from the original size to the specified size.
3. A method for segmenting MRI brain tumor images based on Conv+Swin and CSWin Transformer according to claim 1, characterized in that: In step 2, the Conv+Swin structure extracts image features at different scales through the combination of convolutional operations and Swin Transformer; the SENetV2 backbone network enhances the network's ability to model the relationship between channels by introducing an aggregated multi-layer perceptron, thereby improving the feature extraction effect.
4. A method for segmenting MRI brain tumor images based on Conv+Swin and CSWin Transformer according to claim 1, characterized in that: In step 2, the CSWin Transformer structure expands the attention area by calculating self-attention in parallel on horizontal and vertical stripes; by adjusting the stripe width, a balance between computational cost and modeling ability is achieved in different network layers.
5. A method for MRI brain tumor image segmentation based on Conv+Swin and CSWin Transformer according to claim 1, characterized in that: In step 3, the loss function uses the BCEDiceLoss function, and its calculation formula is L′(x,y)=L′_{dice}(x,y)+0.5*L'_{bce}(x,y) L(x,y)=L_{whole}(x,y)+L_{core}(x,y)+L_{enh}(x,y) Among them, \(L'_{dice}\) represents the dice loss function, \(L'_{bce}\) represents the cross-entropy loss function, and their combination is BCEDiceLoss; \(L_{whole}\) represents the loss function of the entire tumor region, \(L_{core}\) represents the loss function of the tumor core region, and \(L_{enh}\) represents the loss function of the tumor enhancement region; \(x\) represents the predicted value, \(y\) represents the true label value, \(L'(x,y)\) represents the loss value of a single region, and \(L(x,y)\) represents the combination of the loss values of the three regions.
6. A method for MRI brain tumor image segmentation based on Conv+Swin and CSWin Transformer according to claim 1, characterized in that: In the step 3, the optimizer used in the training process is the Adam optimizer, the learning rate is set to 0.003, the momentum parameter is set to 0.9, and the weight decay is set to 0.0001.
7. A method for segmenting MRI brain tumor images based on Conv+Swin and CSWin Transformer according to claim 1, characterized in that: In the step 4, the obtained image segmentation results include different regions of the tumor, such as the entire tumor region, the tumor core region, and the tumor enhancement region. Through the segmentation of these regions, the precise segmentation of the MRI brain tumor image is achieved.