Multi-modal MRI (Magnetic Resonance Imaging) brain tumor image segmentation method based on fusion Transform and U-Net

Through the multimodal MRI brain tumor image segmentation method that fuses Transformer and U-Net, the coordinate attention mechanism and inverse residual module are introduced, the complexity and accuracy of MRI brain tumor segmentation are solved, and the efficient tumor segmentation effect is achieved.

CN120259342APending Publication Date: 2025-07-04GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510418053.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing MRI brain tumor segmentation methods are complex and time-consuming and susceptible to professional knowledge and subjective factors. The existing segmentation networks are not effective in capturing small-target tumors, and the complex model limits its application feasibility.

Method used

A multimodal MRI brain tumor image segmentation method based on fusion Transformer and U-Net is adopted. By introducing coordinate attention mechanism and inverse residual module, a lightweight network model is built to improve segmentation accuracy.

Benefits of technology

It improves the accuracy and speed of MRI brain tumor image segmentation, enhances the generalization ability of the model, and reduces the computational complexity and parameter quantity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259342A_ABST
    Figure CN120259342A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal MRI (Magnetic Resonance Imaging) brain tumor image segmentation method based on fusion of Transform and U-Net, belongs to the field of semantic segmentation, and aims to improve the speed and precision of MRI brain tumor image segmentation. The invention provides a novel method, existing knowledge and experience are utilized, Transform and U-Net models are fused and applied to an MRI brain tumor image segmentation task, an attention mechanism and an inverse residual module are introduced at the same time, and the model performance and generalization ability are remarkably improved. The MRI brain tumor image segmentation method is superior to a traditional segmentation method in the aspect of MRI brain tumor image segmentation, the accuracy of MRI brain tumor segmentation is improved, and important practical significance and guidance are provided for development of medical auxiliary diagnosis and treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of semantic segmentation, and particularly to the research on MRI brain tumor image segmentation based on deep learning. Background Art

[0002] Brain cancer is the most common cancer in the world and one of the main causes of cancer death globally. Brain tumors refer to intracranial tumors. As brain tumors continue to expand, they will compress the cranial nerves, causing serious consequences such as headache, nausea, mental confusion, and even death, and will cause irreversible damage to the brain. According to their different causes, brain tumors can be divided into two types: primary and secondary.

[0003] In recent years, with significant progress in the field of medical imaging, a variety of imaging technologies have emerged. Among them, magnetic resonance imaging (MRI) is of great significance in the diagnosis and treatment of brain tumors because it does not cause radiation or damage to the human body and can effectively eliminate the interference of bone image artifacts existing in the imaging process, and has become one of the currently widely used brain tumor detection technologies. Magnetic resonance imaging technology can make accurate judgments on pathological and tissue morphological lesions by using the detailed information in the brain, further helping doctors more intuitively understand the morphological structure of the lesion, thereby increasing the success rate of surgery.

[0004] Due to the complexity of brain tumors, including small volume, diverse shapes, and heterogeneity of structure and function, current MRI brain tumor segmentation still faces some challenges: First, the tasks of manual segmentation and classification of brain tumors are complex and time-consuming, and are easily affected by professional knowledge and subjective factors; second, the existing segmentation networks perform poorly in capturing small target tumors; in addition, complex models contain a large number of parameters, which limits their feasibility in applications. Therefore, using an MRI brain tumor segmentation model based on deep learning for auxiliary diagnosis and constructing a network model with good segmentation effect and fast segmentation speed is the main direction of the present invention. Summary of the Invention

[0005] The purpose of the present invention is to provide a multi-modal MRI brain tumor image segmentation method based on the fusion of Transformer and U-Net, using the fusion structure of Transformer and CNN as the encoding stage, introducing a new attention mechanism at the skip connection, and adopting the lightweight idea to replace the ordinary convolution module with an inverted residual module to improve the accuracy of MRI brain tumor image segmentation.

[0006] To achieve the above purpose, the present invention provides a multi-modal MRI brain tumor image segmentation method based on the fusion of Transformer and U-Net, including:

[0007] Step 1. Preprocess the multi-institutional, multi-parameter, multi-modal magnetic resonance imaging dataset BraTS2021 provided by the International Society for Medical Image Computing and Computer-Assisted Intervention, and divide it into a training set and a test set;

[0008] Step 2. Construct a multi-modal MRI brain tumor image segmentation model for brain tumor segmentation based on the fusion of Transformer and U-Net;

[0009] Step 3. Set the pre-training parameters, perform iterative training on the constructed network model, and save the optimal model;

[0010] Step 4. Use the optimal model to predict the data in the test set, obtain the predicted segmentation results, and compare the results with the existing methods by calculating various evaluation metrics to verify the superiority of this method in MRI brain tumor image segmentation.

[0011] In Step 1, in the dataset, each case is divided into image data of four MRI modalities: T1, T1ce, T2, and FLAIR. The annotation content mainly includes background, enhancing tumor (ET), peritumoral edema / invasive tissue (ED), and necrotic tumor core and non-enhancing tumor (NCR / NET). The size of all images in the dataset is (240, 240, 155). And slice operations are performed on the three-dimensional images of each modality, and preprocessing is carried out on them. The core steps include central cropping, intensity normalization, noise removal, and data augmentation, etc. The dataset is divided into a training set and a test set at a ratio of 8:2.

[0012] In Step 2, construct a multi-modal MRI brain tumor image segmentation model for brain tumor segmentation based on the fusion of Transformer and U-Net. Take the downsampling process after the fusion of CNN and Transformer as the encoder stage. Use the Coordinate Attention mechanism in the skip connections between the encoder and the decoder, and use the Inverted Residual Block to replace the convolutional module in the network model.

[0013] In Step 3, set relevant hyperparameters, including selecting the Adam optimizer, using the CosineAnnealing strategy to dynamically adjust the learning rate. After setting the pre-training parameters, input the BraTS2021 dataset into the constructed network model. Use the cross-entropy loss function to calculate the predicted values of the network output and the true label values, and use backpropagation to update the various parameters of the network and save the optimal weight information.

[0014] In step 4, evaluation index calculations are performed on the segmentation results of the whole tumor (WT), tumor core (TC), and enhanced tumor (ET) in the predicted image. The evaluation indexes include the following index parameters: Dice Similarity Coefficient (DSC), Sensitivity, Hausdorff Distance (HD), and Specificity. Brief Description of the Drawings

[0015] Figure 1 is a schematic flow chart of the multi-modal MRI brain tumor image segmentation method based on the fusion of Transformer and U-Net proposed by the present invention;

[0016] Figure 2 is a schematic flow chart of the preprocessing of the MRI brain tumor image proposed by the present invention;

[0017] Figure 3 is a schematic diagram of the encoder stage proposed by the present invention;

[0018] Figure 4 is a schematic diagram of the structure of introducing the coordinate attention mechanism (CA) at the skip connection proposed by the present invention;

[0019] Figure 5 is a schematic diagram of the structure of replacing all convolutional modules with inverted residual modules proposed by the present invention.

[0020] Figure 6 is the model of the multi-modal MRI brain tumor image segmentation method based on the fusion of Transformer and U-Net constructed by the present invention Detailed Embodiment

[0021] The following specific examples illustrate the embodiments of the present invention.

[0022] Please refer to Figure 1 , the schematic flow chart of the multi-modal MRI brain tumor image segmentation method based on the fusion of Transformer and U-Net proposed by the present invention.

[0023] As Figure 2As shown, specifically, the preprocessing in the process of the multimodal MRI brain tumor image segmentation method based on the fusion of Transformer and U-Net includes: Since there are many black edges around the image, the image is cropped to 160×160×128. Then, the images of the four modalities are merged into a 4D image (C×W×N×D, C = 4), and saved together with the segmentation label as a.h5 file. Then, data augmentation is performed on the image, specifically including cropping, rotation, flipping, Gaussian noise, contrast transformation, and brightness enhancement. Finally, the dataset is divided into a training set and a test set in a ratio of 8:2.

[0024] As Figure 3As shown, the constructed network model mainly consists of an encoding path and a decoding path, forming an asymmetric U-shaped structure. Among them, the downsampling process after fusing CNN and Transformer is used as the encoder stage. First, convolutional downsampling is performed on the input image. Through these three layers of convolutional operations, for the three layers of convolutional operations, sequential 3×3 convolutional layers with a stride of 1, zero-padding, and batch normalization (BN) layers are used to obtain feature maps at different levels, providing rich local features for the subsequent Transformer module. Then, image serialization and patch embedding are performed on the feature maps obtained from the CNN operation: First, the input image is reshaped into a series of flattened 2D patches, each patch with a size of P×P, and the vectorized patches are mapped to the latent D-dimensional embedding space through a trainable linear projection. At the same time, position embeddings are learned and added to the patch embeddings to retain position information. After that, the image patch embeddings are input into a 12-layer Transformer structure. The Transformer structure includes: composed of L layers of multi-head self-attention (MSA) and multi-layer perceptron (MLP) blocks. In each layer, first, the multi-head self-attention mechanism is calculated, and the global context information is captured by modeling the correlation between each patch in the input sequence and other patches. Then, through layer normalization and the processing of the multi-layer perceptron, the features are further transformed and extracted. The self-attention mechanism in Transformer endows the encoder with powerful feature representation learning ability. It can dynamically focus on information at different positions of the input, adjust the focus of attention according to the characteristics of the data itself, and through layer-by-layer encoding, the original data can be transformed into more discriminative, more abstract, and higher-level feature representations, which helps to accurately extract and encode key features when facing data of different categories and forms, improving the overall performance of the model. The feature maps passing through the Transformer encoder enter the decoder of the U-Net, and then their resolution is gradually restored to a size close to that of the input image through upsampling operations. After the encoding part passes through the convolutional layer operations and the Transformer encoding structure, feature maps of different sizes are generated respectively. From the top layer to the bottom layer, they are: feature maps of 48×256×256, feature maps of 96×128×128, feature maps of 192×64×64, feature maps of 384×32×32, and feature maps of 768×16×16.

[0025] As Figure 4As shown, the output features after a series of convolutional operations and image patch embedding in the encoding stage, after adding positional encoding and entering the Transformer structure, are used as the input to the Coordinate Attention mechanism. The main principle of the Coordinate Attention mechanism: First, perform global average pooling operations along the spatial dimensions (horizontal and vertical) respectively. For the input feature map X ∈ R C×H×W (where (W, H) is the size of the Feature map and C is the number of channels), perform average pooling along the width W direction to obtain z h , whose dimension is R C×H×1 , perform average pooling along the height H direction to obtain z w , with the dimension of R C×1×W The specific calculation is as follows:

[0026]

[0027] Then, perform feature transformation on z h and z w respectively using 1×1 convolutions. The purpose is to further encode the pooled features and compress their dimensions to R C / r×H×1 and R C / r×1×W (r is the dimensionality reduction ratio, used to reduce the computational amount while controlling the feature dimension), obtaining δ h and δ w . Then, restore the dimensions to R C×H×1 and R C×1×W through 1×1 convolutions, and obtain γ h and γ w through the Sigmoid function. They represent the attention weights in the vertical and horizontal directions respectively. Finally, multiply the attention weights with the original high-resolution feature map channel by channel to obtain the feature map weighted by CA attention. The CA attention mechanism encodes the channel relationship and long-range dependence relationship through two steps: coordinate information embedding and coordinate attention generation. It can more accurately locate the position of features in space and effectively utilize the position information to enhance the feature representation. The specific implementation process: Perform operations such as concatenation or addition on the high-resolution feature map processed by the CA attention mechanism and the feature map with the corresponding resolution in the decoder.

[0028] Such as Figure 5As shown in the figure, considering that ordinary convolution will occupy a large amount of memory space and have a high computational complexity as the network deepens, the inverted residual block is used to replace all convolution modules in the network model. The specific process includes: the expansion stage, the depthwise separable convolution stage, and the projection stage. That is, the input features first go through a 1×1 convolution layer for dimensionality increase operation, increasing from the input C channels to tC channels (t is the expansion factor). The purpose of increasing the number of channels of the features is to increase the expression dimension of the features, so that the subsequent depthwise separable convolution has more channels to capture rich feature information. The features after dimensionality increase then enter the depthwise separable convolution layer. The depthwise separable convolution effectively reduces the computational complexity and the number of model parameters. First is the depth convolution operation. It performs convolution on each channel separately, that is, uses C 3×3 convolution kernels to perform convolution on each of the tC channels respectively, so as to effectively capture the spatial features of each channel itself. After completing the depth convolution, pointwise convolution is performed, which is a 1×1 convolution operation whose role is to fuse the features extracted from each channel after depth convolution to achieve information interaction between channels. The features after depthwise separable convolution are then reduced in dimension through a 1×1 convolution layer, restoring the number of channels from tC channels to C channels close to the original input, completing a structural change process similar to a "bottleneck" shape. The Linear linear activation function is selected to retain more features, and such a structure is called a linear bottleneck layer. After completing the main operations of the inverted residual block, a residual connection is introduced, adding the original input feature map of the module to the output feature map after the above series of operations. The purpose is to make it easier for the network to learn the identity mapping, alleviate problems such as gradient disappearance, and help improve the training effect and performance of the model. Its structure can conveniently stack multiple inverted residual blocks to build a deeper network. As the network depth increases, the model can learn more advanced and abstract features, and due to the advantages of the inverted residual block itself, the problem of network degradation is avoided to a certain extent, and the performance of the model can be continuously improved.

[0029] As Figure 6 shown in the figure, it combines the advantages of Transformer and U-Net, aiming to improve the performance and efficiency of image segmentation tasks. The design idea of the network structure is to use the self-attention mechanism of Transformer to replace the encoder structure in U-Net. Among them, the inverted residual block (IR) is used to replace all convolution modules, and the coordinate attention mechanism (CA) is introduced at the skip connection to better capture the global information of the image and the relationship between features.

Claims

1. A multi-modal MRI brain tumor image segmentation method based on the fusion of Transformer and U-Net, characterized in that, It includes the following steps: Step 1. Data acquisition and data preprocessing; Step 2. Construct a multi-modal MRI brain tumor image segmentation brain tumor segmentation model based on the fused Transformer and U-Net; Step 3. Set the pre-training parameters, perform iterative training on the constructed network model, and save the optimal model; Step 4. Use the optimal model to predict the data in the test set, obtain the predicted segmentation results, and compare the results with existing methods by calculating various evaluation metrics to verify the superiority of this method in MRI brain tumor image segmentation.

2. A multimodal MRI brain tumor image segmentation method based on the fusion of Transformer and U-Net according to claim 1, characterized in that, In Step 1, the experimental dataset uses the multi-institutional, multi-parameter multi-modal nuclear magnetic resonance imaging dataset BraTS2021 provided by the International Society for Medical Image Computing and Computer-Assisted Intervention. Each case is divided into image data of four MRI modalities: T1, T1ce, T2, and FLAIR. The annotation content mainly includes background, enhanced tumor (ET), peritumoral edema / invasive tissue (ED), and necrotic tumor core and non-enhanced tumor (NCR / NET). The size of all images is (240, 240, 155). Slice operations are performed on the three-dimensional images of each modality, and preprocessing is performed on them. The core steps include central cropping, intensity normalization, noise removal, and data augmentation operations. The dataset is randomly divided into a training set and a test set according to 8:

2.

3. A multi-modal MRI brain tumor image segmentation method based on the fusion of Transformer and U-Net according to claim 1, characterized in that, In Step 2, the segmentation model that fuses Transformer and U-Net is constructed, mainly including using the downsampling process after fusing CNN and Transformer as the encoder stage, using coordinate attention as the skip connection hub, and replacing all convolutional operations in the network with inverted residual modules. Specifically, it includes the following sub-steps: 3-1. The constructed network model mainly consists of an encoding path and a decoding path to form an asymmetric U-shaped structure. In the encoding stage, first, the downsampling process after fusing CNN and Transformer is used as the encoder stage. Obtain the local fine-grained features in the brain tumor, perform downsampling, gradually reduce the spatial size of the feature map, and at the same time achieve hierarchical extraction of the input brain tumor information. The decoding stage is responsible for restoring the abstract feature map obtained by the encoder to the resolution of the original input image through upsampling operations, mainly including upsampling operations and a series of convolutional layers, so as to gradually restore the extracted detailed information to the size of the original picture. 3-2. Use the coordinate attention mechanism (CoordinateAttention) in the skip connection between the encoder and the decoder, and then perform channel splicing on the output feature matrix and the feature matrix of the same size after upsampling at each level in the decoding stage; 3-3. Use the inverted residual block (Inverted Residual Block) to replace the convolutional module in the network model to lightweight the network model, achieving the purpose of reducing the network model parameters and improving the training speed.

4. A multimodal MRI brain tumor image segmentation method based on the fusion of Transformer and U-Net according to claim 1, characterized in that, Specifically, step 3 includes: setting relevant hyperparameters, including selecting the Adam optimizer and using the cosine annealing strategy to dynamically adjust the learning rate. The specific formula is as follows: Among them is the learning rate corresponding to the current round T cur The corresponding learning rate is the initial learning rate is the minimum learning rate, T max is the set maximum number of training rounds, T cur is the current training round. After setting the pre-training parameters, the BraTS2021 dataset is input into the constructed multi-modal MRI brain tumor image segmentation brain tumor segmentation network model based on the fusion of Transformer and U-Net. The cross-entropy loss function is used to calculate the predicted values of the network output and the true label values, and the backpropagation is used to update the parameters of the network and save the optimal weight information.

5. A multimodal MRI brain tumor image segmentation method based on the fusion of Transformer and U-Net according to claim 1, characterized in that: Specifically, step 4 includes: calculating evaluation indicators for the segmentation results of the whole tumor (WT), tumor core (TC), and enhanced tumor (ET) in the predicted image. The evaluation indicators include the following parameter indicators: Dice Similarity Coefficient (DSC), Sensitivity, Hausdorff Distance (HD), and Specificity.

6. A multimodal MRI brain tumor image segmentation method based on the fusion of Transformer and U-Net according to claim 3, characterized in that In step 3-1, first, convolution downsampling is performed on the input image. Through three-layer convolution operations, feature maps at different levels are obtained, providing rich local features for the subsequent Transformer module. For the three-layer convolution operations, sequential 3×3 convolution layers with a stride of 1, zero-padding, and batch normalization BN layers are used. Then, image serialization and patch embedding are performed on the feature maps obtained from the CNN operation: First, the input image is reshaped into a series of flattened 2D patches, each patch with a size of P×P, and the vectorized patches are mapped to a latent D-dimensional embedding space through a trainable linear projection. At the same time, position embeddings are learned and added to the patch embeddings to preserve position information. After that, the image patch embeddings are input into a 12-layer Transformer structure. The Transformer structure includes: composed of L layers of multi-head self-attention (MSA) and multi-layer perceptron (MLP) blocks. In each layer, first, the multi-head self-attention mechanism is calculated to capture global context information by modeling the correlation between each patch in the input sequence and other patches. Then, through layer normalization and the multi-layer perceptron, the features are further transformed and extracted. After the feature maps pass through the Transformer encoder, they enter the U-Net decoder. To facilitate effective feature stitching between the feature maps at each level in the decoding part and the feature maps at different levels in the encoder, before the upsampling operation, 3×3 convolution is used to change the number of channels of the output feature maps to ensure that it is the same as the output dimension of the encoder. After the encoding part goes through the convolution layer operation and the Transformer encoding structure, feature maps of different sizes are generated respectively. From the top layer to the bottom layer, they are: feature maps of 48×256×256, feature maps of 96×128×128, feature maps of 192×64×64, feature maps of 384×32×32, and feature maps of 768×16×16.

7. A multi-modal MRI brain tumor image segmentation method based on the fusion of Transformer and U-Net according to claim 3, characterized in that, In step 3-2, the output features after a series of convolutional operations and image patch embedding in the encoding stage, after adding positional encoding and entering the Transformer structure, are used as the input to the Coordinate Attention mechanism. The main principle of the Coordinate Attention mechanism: First, global average pooling operations are performed separately along the spatial dimensions (horizontal and vertical). For the input feature map X ∈ R C×H×W (where (W, H) is the size of the Feature map and C is the number of channels), average pooling is performed along the width W direction to obtain z h , whose dimension is R C×H×1 , and average pooling is performed along the height H direction to obtain Z w , with the dimension of R C×1×W . The specific calculation is as follows: Next, perform feature transformation on z h and z w respectively using 1×1 convolutions. The purpose is to further encode the features after pooling, compressing their dimensions to R C / r×H×1 and R C / r×1×W (r is the dimensionality reduction ratio, used to reduce the computational amount while controlling the feature dimensions), obtaining δ h and δ w . Then, restore the dimensions to R C×H×1 and R C×1×W through 1×1 convolutions, and obtain γ h and γ w through the Sigmoid function. They represent the vertical and horizontal attention weights respectively. Finally, multiply the attention weights with the original high-resolution feature map channel by channel to obtain the feature map weighted by CA attention. The CA attention mechanism encodes the channel relationship and long-range dependence relationship through two steps: coordinate information embedding and coordinate attention generation. It can more accurately locate the position of features in space and effectively utilize the position information to enhance the feature representation. Specific implementation process: Concatenate or add the high-resolution feature map processed by the CA attention mechanism with the feature map of the corresponding resolution in the decoder, etc.

8. A multi-modal MRI brain tumor image segmentation method based on the fusion of Transformer and U-Net according to claim 3, characterized in that, In step 3-3, considering that ordinary convolution will occupy a large amount of memory space and have a high computational complexity as the network deepens, an inverted residual block is used to replace all convolution modules in the network model. The specific process includes: an expansion stage, a depthwise separable convolution stage, and a projection stage. That is, the input features first go through a 1×1 convolutional layer for dimensionality expansion, increasing from the input C channels to tC channels (t is the expansion factor). The purpose of increasing the number of channels of the features is to increase the expression dimension of the features, so that the subsequent depthwise separable convolution has more channels to capture rich feature information. The features after dimensionality expansion then enter the depthwise separable convolutional layer, starting with the depth convolution operation. It performs convolution on each channel separately, that is, uses C 3×3 convolutional kernels to perform convolution on each of the tC channels respectively, so as to effectively capture the spatial features of each channel itself. After completing the depth convolution, pointwise convolution is performed, which is a 1×1 convolutional operation whose role is to fuse the features extracted from each channel after depth convolution and achieve information interaction between channels. The features after depthwise separable convolution are then reduced in dimension through a 1×1 convolutional layer, restoring the number of channels from tC channels to C channels close to the original input, completing a structural change process similar to a "bottleneck" shape. After completing the main operations of the inverted residual block, a residual connection is introduced, adding the original input feature map of the module to the output feature map after the above series of operations. The purpose is to make it easier for the network to learn the identity mapping, alleviate problems such as gradient disappearance, and help improve the training effect and performance of the model.

Citation Information

Cited By

  • MRI brain tumor segmentation method

    CN120451568A

  • Brain glioma non-invasive grade diagnosis and IDH typing method for multi-mode MRI data missing

    CN122048844A