A method for MRI brain tumor image segmentation based on multi-modal feature fusion with attention mechanism

By constructing a multimodal feature fusion model based on attention mechanism, BraTSegNet solves the problems of time-consuming and high professional knowledge requirements in traditional brain tumor segmentation methods, and realizes efficient and accurate segmentation of multimodal MRI images.

CN114782350BActive Publication Date: 2025-07-08ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210393464.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-14
Publication Date
2025-07-08
Estimated Expiration
2042-04-14

AI Technical Summary

Technical Problem

The existing brain tumor segmentation method relies on traditional machine learning models, and has problems such as time-consuming, high professional knowledge requirements and inconsistency. Multimodal MRI image segmentation is difficult, and single modal information is insufficient or redundant, resulting in insufficient segmentation accuracy.

Method used

The multimodal feature fusion method based on attention mechanism is adopted, and the ResNet backbone network, a hybrid context perception module and a global attention fusion module are used to build a BraTSegNet model through data augmentation and deep supervision strategies to realize multi-layer feature extraction and fusion, combining channel and spatial attention mechanisms to improve segmentation accuracy.

Benefits of technology

It improves the accuracy and consistency of brain tumor segmentation, solves the problems of time-consuming and professional knowledge requirements in traditional methods, enhances the utilization of multimodal information, and improves the segmentation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114782350B_ABST
    Figure CN114782350B_ABST
Patent Text Reader

Abstract

The present invention discloses an MRI brain tumor image segmentation method based on an attention mechanism for multi-modal feature fusion, which relates to the field of deep learning. First, the present invention performs data preprocessing and data augmentation on the data set, and then constructs a network model. The network model includes a backbone network, a hybrid context awareness module, and a global attention fusion module. The image enters the network model, first undergoes encoding through the backbone network, then perceives global and local information through the hybrid context awareness module, and finally fuses multi-modal features through the attention fusion module and outputs the image. After passing through the trained network model, the two-dimensional magnetic resonance brain tumor image to be segmented is input into the trained model, and the segmentation result of the image is output. This patent can train an effective network model for automatically segmenting MRI brain tumor images, fuse multi-modal features, improve the segmentation accuracy, and has high application value and application prospects in clinical treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning and is applied to medical image segmentation. Specifically, it relates to an MRI brain tumor image segmentation method based on attention mechanism and multi-modal feature fusion. Background Art

[0002] Brain tumor segmentation is crucial for the diagnosis and prognosis of glioma patients. Segmenting brain tumors from magnetic resonance images is a necessary procedure for brain tumor treatment, enabling clinicians to identify the location, extent, and type of tumors. This not only helps with the initial diagnosis but also with managing and monitoring treatment progress. Given the importance of this task, the precise delineation of tumors and their sub-regions is usually done manually by experienced neuroradiologists. This is a cumbersome and time-consuming process that requires a great deal of time and expertise, especially when segmenting images of patients with large tumor volumes, multi-modal images, and heterogeneous tumors. The labeling process is also affected by different perceptions among different labelers, and a consensus on label and segmentation interpretation is required, adding additional complexity.

[0003] Computer-aided segmentation algorithms have the potential to address these drawbacks as they can reduce the labor intensity of the labeling process and maintain consistency across different scenarios. Automatic brain tumor segmentation initially relied on traditional machine learning methods such as atlas-based, decision forest, conditional random field-based methods, etc. With the development of deep learning, traditional machine learning methods have been slowly replaced by deep neural networks. How to better optimize previous models and apply them to medical images to make the segmented medical images more accurate has become an important area of current research.

[0004] Researchers have completed the task of brain tumor segmentation by using two-dimensional slices and three-dimensional volumes as inputs. Although 3D models naturally utilize the three-dimensional structural information inherent in brain anatomy, such models do not necessarily produce better results. In addition, they are often computationally more expensive and thus slower during the inference process. The 3D-based whole-volume methods also require a predefined number of slices through the brain volume as input, and in practice, the number of slices varies depending on the protocol, making these models potentially less versatile.

[0005] MRI images have multiple modalities, which are divided into modalities such as T1-weighted imaging, T2-weighted imaging, T1ce imaging, and free water suppression sequence (FLAIR). Each modality has its own characteristics. A single modality often leads to failure or inadequacy because it cannot fully subdivide tumors in relevant regions. By using different nuclear magnetic resonance imaging modalities, the above weaknesses can be effectively compensated. The image information of multiple modalities can effectively complement each other, which can effectively improve the accuracy of segmentation. However, to a certain extent, it also increases the difficulty of the segmentation problem. The input multi-modal image information increases the necessary information for segmentation but also adds a large amount of unnecessary information, thus increasing the difficulty of the segmentation problem. Summary of the Invention

[0006] The present invention aims to overcome the above problems of the prior art and proposes a method for segmenting MRI brain tumor images based on attention mechanism multi-modal feature fusion, which is used to accurately segment brain tumor images from MRI images.

[0007] A method for segmenting MRI brain tumor images based on attention mechanism multi-modal feature fusion is characterized by comprising the following steps:

[0008] Step 1) Input the data set;

[0009] Input the MRI brain tumor image data set BraTS2021. The Brain Tumor Segmentation Challenge (BraTS) is an annual international competition held since 2012. A large number of fully annotated, multi-institutional, multi-modal nuclear magnetic resonance images of glioma patients at different levels are provided for participants. The magnetic resonance image modalities in the BraTS2021 data set are four modalities: T1-weighted imaging, T2-weighted imaging, T1ce imaging, and free water suppression sequence (FLAIR).

[0010] Input the multi-modal two-dimensional MRI brain tumor image to be segmented.

[0011] Step 2) Data preprocessing and data augmentation;

[0012] By slicing the coronal plane of the three-dimensional images in the data set BraTS2021, each slice should simultaneously obtain the slices corresponding to the other three modalities and the segmentation slice at the corresponding position. The sliced images are changed to 4 channels, corresponding to T1-weighted imaging, T2-weighted imaging, T1ce imaging, and free water suppression sequence (FLAIR) in sequence. The obtained two-dimensional image data set is denoted as 2DBraTS2021. By cropping, flipping, rotating, scaling, shifting, etc. the images in the data set 2DBraTS2021 to expand the data set, this operation is called data augmentation. Data augmentation can increase the amount of training data and improve the generalization ability of the deep neural network model. Finally, all data are normalized to limit the image intensity values within a certain range to avoid adverse effects on training caused by certain abnormal samples.

[0013] Step 3) Construct the network model;

[0014] Construct the segmentation model BraTSegNet of our invention. Our segmentation model mainly consists of a backbone network and two key modules, namely the ResNet backbone network, the Hybrid Context-Aware (HCA) module, and the Dual Attention Fusion (DAF) module. The backbone network extracts multi-level features from the input CT images. Then, the HCA module enhances the features and then inputs them into the DAF module to predict the segmentation map.

[0015] First, extract multi-level features from different levels of the backbone network. Then, both low-level and high-level features are input into the HCA module and enhanced by expanding the receptive field. It should be noted that the low / high-level features represent the features closer to the start / end (i.e., input / output) of the backbone network. Then, we use three DAF modules for feature fusion to predict the segmentation map. In addition, we adopt a deep supervision strategy to supervise the outputs of the three DAF modules and the output of the last HCA module. We use the first four layers of the pre-trained ResNet50 as the encoder of BraTSegNet. The size of the feature map is halved, and the number of channels between two adjacent residual blocks (RBs) is doubled.

[0016] 3.1. Construct the HCA module:

[0017] This module utilizes an enlarged receptive field to utilize more informative features. An HCA module consists of 4 parallel branches, and each branch consists of different convolutional layers. In particular, the third branch utilizes dilated convolutional layers with different dilation rates in series, i.e., hybrid dilated convolution, which provides rich multi-scale features from different receptive fields. After fusing the multi-scale features, we obtain more informative features, providing rich image information features. Mathematically, the HCA module is defined as

[0018] f HCA = ReLU(Conv 3x3 (Cat(Conv 1×1 (f RB ), Conv 3×3 (f RB ), f HDC )) + Conv 1×1 (f RB )) (1)

[0019] f HDC = f3(f2(f1(f RB ))) (2)

[0020] where f i represents an atrous convolution unit with a dilation rate of i and a 3×3 convolutional kernel; Cat(x) represents a concatenation operation; Conv 1×1 (x) and Conv 3×3 (x) represent convolutional units with convolutional kernel sizes of 1×1 and 3×3 respectively; f RB represents the features extracted from the backbone.

[0021] 3.2. Constructing the DAF Module:

[0022] To fuse the rich features of the HCA module, we propose a new DAF module. This module uses the attention weight map generated by the high-level features to enhance the low-level features, and then fuses the enhanced low-level features with the high-level features. We consider both channel attention and spatial attention mechanisms, and connect the channel attention (CA, Channel Attention) module and the spatial attention (SA, Spatial Attention) module in series. We use average pooling in the CA module and max pooling in the SA module. As shown in the figure, the high-level features pass through the CA module and the SA module to generate the attention weight map, and then enhance the low-level features. The sum of the upsampled high-level features and the enhanced low-level features serves as the fused features. Mathematically, we define the DAF module as:

[0023]

[0024]

[0025]

[0026] and represent the features provided by the k-th (low-level) and the (k + 1)-th (high-level) HCA modules, where k = 1, 2, 3. The symbol * represents the Hadamard product, i.e., element-wise multiplication. Deconv 4×4 (x) represents a deconvolution operation with a kernel size of 4×4, which enlarges the size of the feature map. W CA is the attention weight matrix after the features pass through the CA module, and W SA (x) is the operation of the SA module. ArgPool(x) represents the average pooling operation, and MaxPool(x) represents the max pooling operation. σ(x) represents the Sigmoid activation function.

[0027] 3.3 Constructing the Loss Function

[0028] We consider two losses, namely the binary cross entropy (BCE) loss and the Dice loss.

[0029] Therefore, the overall loss is designed to be

[0030] Loss = L BCE + L Dice (6)

[0031] Step 4) Training strategy;

[0032] The pre-processed dataset is sequentially divided into a training set, a test set, and a validation set in a ratio of 6:2:2. Random initialization and the Adam optimization algorithm are adopted. Set BatchSize (the number of samples selected for one training), epoch (which means a round, and training all the data represents one round), and appropriate initial learning rate and the value of learning rate decay at each update. In the BraTSegNet network model, the backpropagation algorithm (BP) is used to update the weights and biases in the network. During the training iteration process, the loss function described in Step 3.3 is used to update the parameters.

[0033] The BraTSegNet network model is trained according to the set training strategy. First, the pre-trained ResNet block parameters on ImageNet are loaded into the corresponding residual blocks of our model. Then, our model is trained using the 2D BraTS2021 dataset. The training segments the whole tumor (WT: whole tumor), the tumor core (TC: tumor core), and the enhancing tumor region (ET, enhancing tumor).

[0034] Step 5) Evaluation metrics;

[0035] The evaluation metrics are as follows:

[0036] Dice similarity coefficient (DSC): DSC is used to measure the similarity between the predicted brain tumor region and the actual brain tumor region. The DSC is defined as follows:

[0037]

[0038] where V S represents the dataset segmented by the model, and V T represents the actual segmented data. |x| represents the operation of cardinality calculation, which provides the number of elements in a set. According to this formula, the Dice similarity coefficients of the whole tumor (WT: whole tumor), the tumor core (TC: tumor core), and the enhancing tumor region (ET, enhancing tumor) are calculated respectively.

[0039] Step 6) Using the trained network model;

[0040] Save the trained network model, perform semantic segmentation on the multi-modal MRI brain tumor images to be segmented, and finally obtain the segmented images.

[0041] The advantages of the present invention are as follows:

[0042] 1. A new context-aware module is designed, and a hybrid dilated convolutional network is used to expand the receptive field. The convolutional neural network will cause problems such as the loss of internal data structure and the loss of spatial hierarchical information. The dilated convolutional network can solve these problems. However, continuously using the dilated convolutional network with the same dilation rate will result in the loss of information continuity, and the hybrid dilated convolutional network is used to solve this problem.

[0043] 2. The attention mechanism is used for multi-modal feature fusion to fully segment the tumors in the relevant regions, solve the problem of low boundary contrast, and effectively improve the accuracy of segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is the specific flowchart of the method proposed by the present invention;

[0045] Figure 2 is the network structure diagram of BraTSegNet proposed by the present invention;

[0046] Figure 3 is the structure diagram of the HCA module proposed by the present invention;

[0047] Figure 4 is the structure diagram of the DAF module proposed by the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0048] The following further describes the present invention with reference to the accompanying drawings:

[0049] As Figure 1 shown, a new method for segmenting lung CT images based on transfer learning and attention mechanism of the present invention specifically includes the following steps:

[0050] Step 1) Input the data set;

[0051] Input the MRI brain tumor image data set BraTS2021. The Brain Tumor Segmentation Challenge (BraTS) is an annual international competition held since 2012. A large number of fully annotated, multi-institutional, multi-modal magnetic resonance images of glioma patients with different degrees are provided for participants. The magnetic resonance image modalities in the BraTS2021 data set are T1-weighted imaging, T2-weighted imaging, T1ce imaging, and free water suppression sequence (FLAIR).

[0052] Input the two-dimensional multi-modal MRI brain tumor images to be segmented.

[0053] Step 2) Data augmentation and data preprocessing;

[0054] By slicing the coronal plane of the three-dimensional images in the BraTS2021 dataset, for each slice, the corresponding slices of the other three modalities and the segmentation slice are obtained simultaneously. The sliced images are changed to 4 channels, corresponding to T1-weighted imaging, T2-weighted imaging, T1ce imaging, and free water suppression sequence (FLAIR) in sequence. The resulting two-dimensional image dataset is denoted as 2DBraTS2021. The dataset is enlarged by cropping, flipping, rotating, scaling, shifting, etc. on the images in the 2DBraTS2021 dataset. This operation is called data augmentation. Data augmentation can increase the amount of training data and improve the generalization ability of the deep neural network model. Finally, all data are normalized to limit the image intensity values within a certain range to avoid adverse effects on training caused by certain abnormal samples.

[0055] Step 3) Construct a network model;

[0056] Construct the segmentation model BraTSegNet invented by us. As Figure 2 shown, our segmentation model mainly consists of a backbone network and two key modules, namely the ResNet backbone network, the Hybrid Context-Aware (HCA) module, and the Dual Attention Fusion (DAF) module. The backbone network extracts multi-level features from the input CT images. Then, the HCA module enhances the features and then inputs them into the SAF module to predict the segmentation map.

[0057] As Figure 1 shown, first, multi-level features are extracted from different levels of the backbone network. Then, both low-level and high-level features are input into the HCA module and enhanced by expanding the receptive field. It should be noted that low / high-level features represent features closer to the start / end (i.e., input / output) of the backbone network. Then, we use three DAF modules for feature fusion to predict the segmentation map. In addition, we adopt a deep supervision strategy to supervise the outputs of the three DAF modules and the output of the last HCA module. We use the first four layers of the pre-trained ResNet50 as the encoder of BraTSegNet. The size of the feature map is halved, and the number of channels between two adjacent residual blocks (RBs) is doubled.

[0058] 3.1. The HCA module proposed by the present invention:

[0059] This module utilizes more information features by using an enlarged receptive field. As Figure 3As shown, an HCA module consists of 4 parallel branches, each branch being composed of different convolutional layers. In particular, the third branch utilizes dilated convolutional layers with different dilation rates in series, namely Hybrid Dilated Convolution (HDC), providing rich multi-scale features from different receptive fields. After fusing the multi-scale features, we obtain more informative features, providing rich image information features. Mathematically, the HCA module is defined as

[0060] f HCA = ReLU(Conv 3×3 (Cat(Conv 1×1 (f RB ),Conv 3×3 (f RB ),f HDC )) + Conv 1×1 (f RB )) (1)

[0061] f HDC = f3(f2(f1(f RB ))) (2)

[0062] where f i represents a dilated convolutional unit with a dilation rate of i and a convolutional kernel of 3×3; Cat(x) represents the concatenation operation; Conv 1×1 (x) and Conv 3×3 (x) represent convolutional units with convolutional kernel sizes of 1×1 and 3×3 respectively; f RB represents the features extracted from the backbone.

[0063] 3.2. The DAF module proposed by the present invention:

[0064] To fuse the rich features of the HCA module, we propose a new DAF module. As Figure 4 shown, this module utilizes the attention weight map generated by high-level features to enhance low-level features, and then fuses the enhanced low-level features with high-level features. We simultaneously consider the channel attention and spatial attention mechanisms, and concatenate the channel attention (CA) module and the spatial attention (SA) module. We use average pooling in the CA module and max pooling in the SA module. As shown in the figure, the high-level features pass through the CA module and the SA module to generate attention weight maps, and then enhance the low-level features. The sum of the upsampled high-level features and the enhanced low-level features serves as the fused features. Mathematically, we define the DAF module as:

[0065]

[0066]

[0067]

[0068] and represent the features provided by the k-th (lower level) and (k + 1)-th (higher level) HCA modules, where k = 1, 2, 3. The symbol * represents the Hadamard product, i.e., element-wise multiplication. Deconv 4×4 (x) represents a deconvolution operation with a kernel size of 4×4, which enlarges the size of the feature map. W CA is the attention weight matrix after the features pass through the CA module, and W SA (x) is the operation of the SA module. ArgPool(x) represents the average pooling operation, and MaxPool(x) represents the max pooling operation. σ(x) represents the Sigmoid activation function.

[0069] 3.3 Loss Function

[0070] We consider two losses, namely, the binary cross-entropy (BCE) loss and the Dice loss.

[0071] Therefore, the overall loss is designed as

[0072] Loss = L BCE + L Dice (6)

[0073] Step 4) Training Strategy;

[0074] The preprocessed dataset is sequentially divided into a training set, a test set, and a validation set in a ratio of 6:2:2. The training set does not contain slices without lesions to alleviate the problem of class imbalance. Random initialization and the Adam optimization algorithm are adopted. Set BatchSize (the number of samples selected for one training), epoch (which represents one round, and training all the data represents one round), and appropriate initial learning rates and the value of the learning rate decay at each update. The backpropagation (BP) algorithm is used to update the weights and biases in the network in the BraTSegNet network model. During the training iteration process, the loss function described in Step 3.3 is used to update the parameters.

[0075] The BraTSegNet network model is trained according to the set training strategy. First, the parameters of the ResNet blocks pre-trained on ImageNet are loaded into the corresponding residual blocks of our model.

[0076] Then, the 2D BraTS2021 dataset is regarded as the source data to pre-train the model. The settings are as follows: 100 epochs, an initial learning rate of 1e-4, a batch size of 10, and an image size of 240×240. The Adam optimizer is used for optimization.

[0077] The number of epochs is set to 100, and an early stopping strategy is adopted to prevent overfitting.

[0078] Step 5) Evaluation metrics;

[0079] The evaluation metrics are as follows:

[0080] Dice Similarity Coefficient (DSC): DSC is used to measure the similarity between the predicted pulmonary infection and the ground truth. The DSC is defined as follows:

[0081]

[0082] where V S represents the dataset after being segmented by the model, and V T represents the ground truth segmentation data. |x| represents the operation of cardinality calculation, which provides the number of elements in a set. According to this formula, the Dice Similarity Coefficient of the whole tumor (WT: whole tumor), tumor core (TC: tumor core), and enhancing tumor region (ET, enhancing tumor) are calculated respectively.

[0083] Step 6) Use the trained network model;

[0084] Save the trained network model, input the two-dimensional multimodal MRI brain tumor image to be segmented for semantic segmentation, and finally obtain the segmented image.

[0085] In the attached drawings of the disclosed embodiments of the present invention, only the structures related to the disclosed embodiments are involved. However, the above is only the preferred embodiment of the present invention, and it can be fully applied to various fields suitable for the present invention. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the illustrated and described examples here.

Claims

1. A method for segmenting MRI brain tumor images with multi-modal feature fusion based on the attention mechanism, characterized in that, It includes the following steps: Step 1) Input the dataset; Input the MRI brain tumor image dataset BraTS2021; The magnetic resonance image modalities in the BraTS2021 dataset include four modalities: T1-weighted imaging, T2-weighted imaging, T1ce imaging, and free water suppression sequence FLAIR; Input the two-dimensional multimodal MRI brain tumor image to be segmented; Step 2) Data augmentation and data preprocessing; By slicing the coronal plane of the three-dimensional images in the dataset BraTS2021, each slice should simultaneously obtain the slices at the corresponding positions of the other three modalities and the segmentation slice, and change the sliced images to 4 channels, corresponding to T1-weighted imaging, T2-weighted imaging, T1ce imaging, and free water suppression sequence FLAIR in sequence. The obtained two-dimensional image dataset is denoted as 2DBraTS2021; By cropping, flipping, rotating, scaling, shifting, etc. the images in the dataset 2DBraTS2021 to expand the dataset, this operation is called data augmentation. Data augmentation can increase the amount of training data and improve the generalization ability of the deep neural network model. Finally, all data is normalized to limit the image intensity value within a certain range to avoid adverse effects on training caused by certain abnormal samples; Step 3) Build a network model; Build the segmentation model BraTSegNet; The segmentation model consists of a backbone network and two key modules, namely the ResNet backbone network, the Hybrid Context-Aware module, abbreviated as the HCA module, and the Dual Attention Fusion module, abbreviated as the DAF module; The backbone network extracts multi-level features from the input CT images; Then, the HCA module enhances the features and then inputs them into the DAF module to predict the segmentation map; First, multi-level features are extracted from different levels of the backbone network; Then, both low-level and high-level features are input into the HCA module and enhanced by expanding the receptive field; It should be noted that low / high-level features represent features closer to the start / end of the backbone network. The starting feature is the input feature, and the ending feature is the output feature; Then, three DAF modules are used for feature fusion to predict the segmentation map; In addition, a deep supervision strategy is adopted to supervise the outputs of the three DAF modules and the output of the last HCA module; The first four layers of the pre-trained ResNet50 are used as the encoder of BraTSegNet; The size of the feature map is halved, and the number of channels between two adjacent residual blocks RB is doubled; 3.

1. Build the HCA module: The HCA module utilizes an enlarged receptive field to exploit more information features; an HCA module consists of 4 parallel branches, each branch being composed of different convolutional layers; in particular, the third branch utilizes dilated convolutional layers with different dilation rates in series, namely, hybrid dilated convolution, which provides rich multi-scale features from different receptive fields; after fusing the multi-scale features, we obtain more information features, providing rich image information features; mathematically, the HCA module is defined as f HCA = ReLU(Conv 3x3 (Cat(Conv 1×1 (f RB ), Conv 3×3 (f RB ), f HDC )) + Conv 1×1 (f RB )) (1) f HDC = f3(f2(f1(f RB ))) (2) Among them, f i represents an atrous convolution unit with a dilation rate of i and a 3×3 convolutional kernel; Cat(x) represents a concatenation operation; Conv 1×1 (x) and Conv 3×3 (x) respectively represent convolutional units with 1×1 and 3×3 convolutional kernels; f RB represents the features extracted from the backbone; 3.

2. Construction of the DAF module: To fuse the rich features of the HCA module, a new DAF module is proposed; the DAF module utilizes the attention weight map generated by high-level features to enhance low-level features, and then fuses the enhanced low-level features with high-level features; we consider both channel attention and spatial attention mechanisms, connecting the channel attention ChannelAttention module, abbreviated as the CA module, and the spatial attention SpatialAttention module in series, using average pooling in the CA module and max pooling in the SA module; The high-level features pass through the CA module and the SA module to generate the attention weight map, and then enhance the low-level features; the sum of the upsampled high-level features and the enhanced low-level features serves as the fused features; Mathematically, the DAF module is defined as: and represent the features provided by the k-th and (k + 1)-th HCA modules, where k = 1, 2, 3; the symbol * represents the Hadamard product, i.e., element-wise multiplication; Deconv 4×4 (x) represents a deconvolution operation with a kernel size of 4×4, which enlarges the size of the feature map; W CA is the attention weight matrix after the features pass through the CA module, W SA (x) is the operation of the SA module; ArgPool(x) represents an average pooling operation, and MaxPool(x) represents a max pooling operation; σ(x) represents the Sigmoid activation function; 3.3 Construction of the loss function; A deep supervision strategy is adopted to design the loss function; specifically, supervision is added in each DAF module and the last HCA module, a total of 4, allowing for better gradient flow and more effective network training; for each supervision, two losses are considered, namely, the binary cross entropy loss, abbreviated as the BCE loss, and the Dice loss; therefore, the overall loss is designed as Loss=L BCE +L Dice (6) Step 4) Training strategy; The preprocessed dataset is sequentially divided into a training set, a test set, and a validation set in a ratio of 6:2:2; random initialization and the Adam optimization algorithm are adopted; the BatchSize, that is, the number of samples selected for one training, the epoch, which means the number of rounds, where training all the data represents one round, and appropriate initial learning rates and the value by which the learning rate decreases at each update are set; the backpropagation algorithm BP algorithm is used in the BraTSegNet network model to update the weights and biases in the network; During the training iteration process, the loss function described in step 3.3 is used to update the parameters; The BraTSegNet network model is trained according to the set training strategy; first, the parameters of the ResNet blocks pre-trained on ImageNet are loaded into the corresponding residual blocks of the BraTSegNet network model; then, the BraTSegNet network model is trained using the 2D BraTS2021 dataset; the training segments the whole tumor wholetumor, abbreviated as WT, the tumor core tumor core, abbreviated as TC, and the enhancing tumor region enhancingtumor, abbreviated as ET; Step 5) Evaluation metrics; The evaluation metrics are as follows: Dice Similarity Coefficient DSC: DSC is used to measure the similarity between the predicted brain tumor region and the actual brain tumor region; the DSC is defined as follows: Among them, V S represents the dataset after model segmentation, and V T represents the segmented data of the facts; x represents the operation of cardinality calculation, which provides the number of elements in a set; according to this formula, the Dice similarity coefficients of the whole tumor WT, tumor core TC, and enhanced tumor region ET are calculated respectively; Step 6) Use the trained network model; Save the trained network model, perform semantic segmentation on the two-dimensional multi-modal MRI brain tumor image to be segmented, and finally obtain the segmented image.

Citation Information

Patent Citations

  • Brain glioma medical image segmentation method based on U-Net network

    CN112446891A

  • Brain tumor detection method based on attention mechanism and MRI (Magnetic Resonance Imaging) multi-modal fusion

    CN114119515A