A deep learning-based 3D MRI brain tumor segmentation method

By introducing multi-scale convolution joint module and global context aggregation module in deep convolution neural networks, the problem that traditional methods are difficult to capture multi-scale and global context information in brain tumor segmentation is solved, and a more accurate brain tumor segmentation effect is achieved.

CN114359293BActive Publication Date: 2025-05-16NANCHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111516472.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-10
Publication Date
2025-05-16
Estimated Expiration
2041-12-10

AI Technical Summary

Technical Problem

Traditional image segmentation algorithms are difficult to effectively segment complex shapes and changeable brain tumors, and lack multi-scale and global context information, resulting in poor segmentation effect.

Method used

A three-dimensional MRI brain tumor segmentation method based on deep learning is proposed. By building a deep convolutional neural network, combining multi-scale convolutional joint module and global context aggregation module, local and global feature information are integrated to enhance the segmentation capability of the network.

Benefits of technology

The network segmentation ability of the target area is improved, the redundant characteristics are reduced, and the segmentation effect of tumor areas is improved, especially in more difficult-to-predict enhancing tumor areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114359293B_ABST
    Figure CN114359293B_ABST
Patent Text Reader

Abstract

The present invention discloses a three-dimensional MRI brain tumor segmentation method based on deep learning, comprising the following steps: S1: pre-processing the three-dimensional MRI brain data and dividing the data set to meet the input conditions of the model. S2: constructing and training a deep convolutional neural network, the network framework adopts the form of an encoder and a decoder, and adds a multi-scale convolution joint module and a global context aggregation module. S3: post-processing the obtained predicted data to further improve the segmentation effect. The segmentation method proposed in the present invention combines the low-level features and high-level features of the segmentation object, effectively integrates multi-scale information and global context information, and reduces the influence of learned redundant features, thereby improving the segmentation results of brain tumors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image segmentation, and in particular to a three-dimensional MRI brain tumor segmentation method based on deep learning. Background Art

[0002] Predicting a complete 3D point cloud is a core task in many computer vision applications. Brain tumors have a very high mortality and morbidity rate, but if they are discovered in time, early diagnosis and early treatment can increase the possibility of cure. Brain tumor image segmentation is a very important step in the clinical diagnosis and treatment of brain tumors. By segmenting the tumor in MRI images, doctors can locate the location of the tumor and obtain the size of the tumor, and then formulate relevant treatment and rehabilitation strategies. However, due to the complex structure, variable shape and extremely unbalanced categories of brain tumors, traditional image segmentation algorithms such as region growing and threshold methods often find it difficult to obtain satisfactory segmentation results. Therefore, developing robust and accurate automatic segmentation methods to achieve effective and objective segmentation is a very challenging research field.

[0003] In recent years, image segmentation methods based on deep learning technology, especially deep convolutional neural networks (DCNNs), have developed rapidly. The most popular method is to use a U-shaped architecture to segment medical images. The architecture includes an encoder path to capture high-level semantics related to segmentation, and a symmetric decoder with jump connections from the encoder to generate segmentation results, so that low-level information and high-level information are integrated with each other. However, all convolution modules are composed of 2 stacked 3 convolution layers, which leads to a relatively single scale of extracted features and a lack of multi-scale and global contextual information in the captured semantic features, which makes it less effective in the more challenging brain tumor segmentation. Summary of the invention

[0004] In order to solve the above problems, the present invention proposes a 3D MRI brain tumor segmentation method based on deep learning. The technical solution is as follows:

[0005] S1. Preprocess the 3D MRI brain data and divide the data set;

[0006] MRI images have four different modalities including T1, T1ce, T2 and FLAIR. We spliced ​​the four data together to form four input channels, cropped the original 3D brain MRI image of 155*240*240 to 150*192*192, and removed the redundant background pixels. The data was normalized before input into the network to reduce the learning difficulty of the network, and the training data was diversified online, including random scaling, random flipping and random cropping along the three-dimensional direction. The final 3D image size fed into the network was 96*144*144. Finally, the brain dataset was divided into training set, validation set and test set in a ratio of 8:1:1.

[0007] S2. Build and train a deep convolutional neural network model;

[0008] The deep convolutional neural network is trained using the training set and the trained network is verified at any time using the verification set. The deep convolutional neural network has an encoder and a corresponding decoder, and the decoder obtains the features of the encoder through a skip connection.

[0009] The preprocessed data is first input into an encoder consisting of three groups of downsampling convolution modules and one group of multi-scale convolution joint modules for encoding. After that, the encoded feature map is input into a decoder consisting of three global context aggregation modules for decoding, and finally the segmentation result is output.

[0010] Specifically, the downsampling convolution module includes two 3*3*3 convolutions, each of which is followed by a group normalization layer with a group number of 8 and a ReLu unit for increasing nonlinearity, and then a 2*2*2 maximum pooling layer with a stride of 2 in each dimension.

[0011] The multi-scale convolutional joint module contains two groups of convolutions with different dilation rates. The two groups of convolutions are combined in a cascaded manner, and each group of convolutions consists of three dilated convolutions with a kernel size of 3 and different dilation rates and one 1*1*1 convolution superimposed in parallel. The dilation rates of the first group of dilated convolutions are 1, 2, and 4; the dilation rates of the second group of dilated convolutions are 1, 2, and 5. In addition, each convolution with a kernel of 3*3*3 is followed by a group normalization layer with a group size of 8 and a ReLu linear unit. Each convolution in the group produces the same number of output dimensions. In the multi-scale convolutional joint module, we not only incorporate dilated convolutions with different dilation rates to extract features of objects of different sizes, but also add an additional 1*1*1 convolution to each group, thereby realizing the linear combination of multiple feature maps, and realizing cross-channel interaction and information integration.

[0012] After receiving the features from the encoder, the global context aggregation module first performs upsampling. The upsampling operation is implemented by deconvolution with a stride of 2 in each dimension and a convolution kernel of 2*2*2. Then, the feature information with the same resolution in the encoding path and the decoding path is fused by element-wise summation (implemented by jump connections). The fused features are subjected to two 1*1*1 convolutions to obtain two feature maps, denoted as feature map A and feature map B. At the same time, two branches are formed, which we denote as branch Z1 and branch Z2. In order to efficiently collect background information at each spatial position, we first perform a global average pooling operation in the Z1 branch, generate a global context feature representation, and then add it to A. Then, the obtained features are applied with Sigmoid layer to obtain feature weight map S. Finally, S is element-wise multiplied with feature map A after convolution operation to obtain feature map D. Subsequently, feature map D is fed into a 3D convolution with a kernel size of 3*3*3, followed by a group convolution with a group size of 8 and a nonlinear activation function ReLu, and finally the recalibrated feature Y1 is obtained. In the Z2 branch, we let feature map B undergo a 3D convolution operation with a step size of 2 in each direction and a kernel size of 3 to obtain feature map Y2. Finally, we splice the outputs Y1 and Y2 of the two branches together along the channel direction to form the global context aggregation module output Y. In the global context aggregation module, the Z1 branch not only enhances the position features by modeling the spatial context information, but also obtains the dependency between channels using the global pooling layer. In addition, the features of the Z1 branch and the Z2 branch are fused to establish long-distance semantic dependencies.

[0013] S3, post-processing the obtained prediction data;

[0014] The test data is fed into the trained deep convolutional neural network model for prediction, and the output feature map is post-processed to obtain the final tumor. Usually, the enhanced tumor area is difficult to predict and is prone to false positive prediction results, so when the predicted enhanced tumor area is too small, we replace the enhanced tumor area with the necrotic / edema area.

[0015] Furthermore, in step S2, in the deep convolutional neural network, we introduce deep supervision at each stage of the decoding path to add auxiliary outputs from the decoding layer to help the model better learn high-level semantic features and low-level location information. Each deep supervision subnet uses a 1*1*1 convolution to normalize the output channel, then uses a trilinear upsampling operation to restore the image space dimension, and finally applies a sigmoid function to obtain a predicted probability map representing the tumor area. The overall loss function is:

[0016]

[0017] Among them, L g1, L g2 , L g3 They are the loss functions output by the three global context aggregation modules. Specifically, we use the Dice loss, which is good at mining foreground areas, as the loss function to alleviate the adverse effects of severe imbalance between positive and negative samples. The specific calculation method of the Dice loss is as follows:

[0018]

[0019] Among them, T is the real tumor pixel manually annotated, and S is the tumor pixel predicted by the model.

[0020] The beneficial effects of the present invention are:

[0021] The present invention proposes a deep convolutional network that combines multi-scale context and global context, which can simultaneously capture spatial information of different scales and long-distance feature dependencies. The multi-scale convolutional joint module and the global context aggregation module can integrate global and local feature information, combine low-level features and high-level features of the segmented object, effectively fuse global context information and reduce the influence of learned redundant features, ultimately improving the network's ability to segment the target area. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a flow chart for implementing the method of the present invention;

[0023] Figure 2 A schematic diagram of the deep convolutional neural network structure of the present invention;

[0024] Figure 3 Schematic diagram of the multi-scale convolutional joint module of the present invention;

[0025] Figure 4 is a schematic diagram of a global context aggregation module of the present invention; DETAILED DESCRIPTION

[0026] The present invention will be further described below with reference to the accompanying drawings.

[0027] A deep learning-based 3D MRI brain tumor segmentation method, such as Figure 1 As shown, the following steps are included:

[0028] S1, preprocessing the 3D MRI brain data and dividing the data set to meet the input conditions of the model;

[0029] MRI images have 4 different modalities including T1, T1ce, T2 and FLAIR. We splice the 4 data together to form 4 input channels. Usually, the background information accounts for a large proportion of the entire image, and the tumor area accounts for a very small proportion, which will lead to serious data imbalance, and the background is not helpful for segmentation. Therefore, we choose to remove the background information around the brain area and crop the original size of 155*240*240 3D brain MRI image to 150*192*192. The final size of the 3D image sent to the network is 96*144*144.

[0030] In addition, data standardization can reduce the redundant information of the data, thereby reducing the difficulty of network learning. This experiment uses z-score normalization, which is defined as:

[0031]

[0032] where μ is the mean value of the pixel-level MRI sequence and ρ is the standard deviation of the pixel-level MRI sequence.

[0033] In addition, in order to effectively prevent overfitting in the training phase due to the small size of the medical dataset, we used online data augmentation techniques, including random scaling, random flipping in three dimensions, and random cropping. Finally, the brain dataset was divided into training set, validation set, and test set in a ratio of 8:1:1.

[0034] S2. Build and train a deep convolutional neural network model;

[0035] The deep convolutional neural network is trained using the training set and the trained network is verified at any time using the verification set. The deep convolutional neural network has an encoder and a corresponding decoder, and the decoder obtains the features of the encoder through a skip connection.

[0036] The structure of deep convolutional neural network is as follows Figure 2 As shown in the figure, the preprocessed data is first input into an encoder comprising three groups of downsampling convolution modules and one group of multi-scale convolution joint modules for encoding. After that, the encoded feature map is input into a decoder consisting of three global context aggregation modules for decoding, and finally the segmentation result is output.

[0037] Specifically, the downsampling convolution module includes two 3*3*3 convolutions, each of which is followed by a group normalization layer with a group number of 8 and a ReLu unit for increasing nonlinearity, and then a 2*2*2 maximum pooling layer with a stride of 2 in each dimension.

[0038] The multi-scale convolutional joint module contains two sets of convolutions with different dilation rates. In three-dimensional space, the convolution kernel sizes are I, J, and K respectively, and the input signal is x(p, q, s). The dilated convolution can be defined as:

[0039]

[0040] Among them, y(q,r,s) is the output signal of the three-dimensional dilated convolution, w(i,j,k) represents the convolution filter, and r is the dilation rate of the three-dimensional dilated convolution. The main idea of ​​dilated convolution is to insert holes (fill with zeros) between the pixels of the convolution kernel. Assuming I=J=K, then for the dilated convolution with the original convolution kernel length of I and the dilation rate of r, the actual convolution kernel length is i'=i+(i-1)·(r-1). If we set r to 1, then this convolution becomes an ordinary 3D convolution. According to the principle of dilated convolution, dilated convolution can change the range of the receptive field by inserting holes between kernels, that is, dilated convolution can obtain multi-scale information without increasing parameters.

[0041] The two sets of convolutions are combined in a cascaded manner, such as Figure 3 As shown. Each group of convolutions consists of 3 dilated convolutions with kernel size 3 and different dilation rates and 1 1*1*1 convolution stacked in parallel. The dilation rates of the first group of dilated convolutions are 1, 2, and 4; the dilation rates of the second group of dilated convolutions are 1, 2, and 5. In addition, each convolution with a kernel size of 3*3*3 is followed by a group normalization layer with a group size of 8 and a ReLu linear unit. Each convolution in the group produces the same number of output dimensions. In the multi-scale convolutional joint module, we not only incorporate dilated convolutions with different dilation rates to extract features of objects of different sizes, but also add an additional 1*1*1 convolution to each group, thereby realizing the linear combination of multiple feature maps and achieving cross-channel interaction and information integration.

[0042] like Figure 4 As shown in the figure, after receiving the features from the encoder, the global context aggregation module first performs upsampling. The upsampling operation is implemented by deconvolution with a stride of 2 in each dimension and a convolution kernel of 2*2*2. Then, the feature information with the same resolution in the encoding path and the decoding path is fused by element summation (implemented by skip connection). After two 1*1*1 convolutions, two feature maps are obtained and recorded as feature maps and feature map At the same time, two branches are also formed, which we record as branch Z1 and branch Z2.

[0043] A l =ω θ I l +bθ ;

[0044] B m =ω σ I m +b σ ;

[0045] where ω θ ,ω σ are the weights of two 1*1*1 convolutions, b θ , b σ They are the biases of two 1*1*1 convolutions respectively.

[0046] In order to efficiently collect background information at each spatial position, we first perform a global average pooling operation in the Z1 branch to generate a global context feature representation and then add it to A. Then, the obtained features are applied to the Sigmoid layer to obtain the feature weight map Finally, S is element-wise multiplied with the convolutional feature map A to obtain the feature map Subsequently, the feature map D is fed into a 3D convolution with a kernel size of 3*3*3, followed by a group convolution with a group size of 8 and a nonlinear activation function ReLu, and finally the recalibrated feature Y1 is obtained; in the Z2 branch, we let the feature map B undergo a 3D convolution operation with a step size of 2 in each direction and a convolution kernel size of 3 to obtain the feature map Y2. Finally, we concatenate the outputs Y1 and Y2 of the two branches along the channel direction to form the output of the global context aggregation module. (C o =C l +C m ).

[0047] In the global context aggregation module, the Z1 branch not only enhances the position features by modeling the spatial context information, but also obtains the inter-channel dependencies using the global pooling layer. In addition, the features of the Z1 branch and the Z2 branch are fused to establish long-distance semantic dependencies.

[0048] Furthermore, in step S2, in the deep convolutional neural network, we introduce deep supervision at each stage of the decoding path to add auxiliary outputs from the decoding layer to help the model better learn high-level semantic features and low-level location information. Each deep supervision subnet uses a 1*1*1 convolution to normalize the output channel, then uses a trilinear upsampling operation to restore the image space dimension, and finally applies a sigmoid function to obtain a predicted probability map representing the tumor area. The overall loss function is:

[0049]

[0050] Among them, L g1 , Lg2 , L g3 They are the loss functions output by the three global context aggregation modules. Specifically, we use the Dice loss, which is good at mining foreground areas, as the loss function to alleviate the adverse effects of severe imbalance between positive and negative samples. The specific calculation method of the Dice loss is as follows:

[0051]

[0052] Among them, T is the real tumor pixel manually annotated, and S is the tumor pixel predicted by the model.

[0053] Furthermore, we use four evaluation indicators to comprehensively evaluate the segmentation effect of brain tumors, including Dice similarity coefficient, sensitivity, and specificity. The Dice similarity coefficient measures the spatial overlap between automatic segmentation and labeling. It is defined as:

[0054]

[0055] Where FP, FN, and TP are false positive, false negative, and true positive, respectively. Sensitivity, also known as true positive rate or detection probability, measures the proportion of correctly identified positives:

[0056]

[0057] Finally, specificity, also known as the true negative rate, measures the proportion of negatives that are correctly identified. It is defined as:

[0058]

[0059] TN is true negative.

[0060] S3, post-processing the obtained prediction data;

[0061] The test data is fed into the trained deep convolutional neural network model for prediction, and the output feature map is post-processed to obtain the final tumor. Usually, the enhanced tumor area is difficult to predict and is prone to false positive prediction results, so when the predicted enhanced tumor area is too small, we replace the enhanced tumor area with the necrotic / edema area.

[0062] In summary, the present invention provides a three-dimensional MRI brain tumor segmentation method based on deep learning, which can realize brain tumor segmentation in an end-to-end training manner without multiple trainings, thus reducing the training time; it introduces a multi-scale convolution joint module and a global context aggregation module, integrates local features and long-distance global features, fuses high-level semantic features and low-level visual features, and reduces the influence of redundant features; it improves the segmentation effect of the tumor area, especially the enhanced tumor area that is difficult to predict.

[0063] It should be understood that the above implementation modes are not limitations of the present invention, and as long as they conform to the basic concept of the present invention, they should belong to the protection scope of the present invention.

Claims

1. A deep learning-based 3D MRI brain tumor segmentation method, characterized in that: The following steps are involved: S1. Preprocess the 3D MRI brain data and divide the data set; MRI images have four different modalities including T1, T1ce, T2 and FLAIR. The four data are spliced ​​together to form four input channels, and some necessary data preprocessing is performed and the data set is divided into training set, validation set and test set. S2. Build and train a deep convolutional neural network model; The deep convolutional neural network is trained using the training set and the trained network is verified at any time using the verification set; wherein the deep convolutional neural network has an encoder and a corresponding decoder, and the decoder obtains the features of the encoder through a skip connection; S3, post-processing the obtained prediction data; The test data is sent to the trained deep convolutional neural network model for prediction. After post-processing, the output feature map is used to obtain the final tumor. Usually, the enhanced tumor area is difficult to predict and is prone to false positive prediction results. Therefore, when the predicted enhanced tumor area is too small, the enhanced tumor area is replaced by the necrotic / edema area. The encoder described in S2 includes 3 groups of downsampling convolution modules and 1 group of multi-scale convolution joint modules, and the corresponding decoder includes 3 global context aggregation modules; The preprocessed data is input into the downsampling convolution module and then into the multi-scale convolution joint module to complete the encoding process; The downsampling convolution module includes two 3*3*3 convolutions, each followed by a group normalization layer with 8 groups and a ReLu unit for increasing nonlinearity, followed by a 2*2*2 max pooling layer with a stride of 2 in each dimension; The multi-scale convolutional joint module includes two groups of convolutions with different dilation rates; the two groups of convolutions are combined in a cascaded manner, and each group of convolutions consists of three dilated convolutions with a kernel size of 3 and different dilation rates and one 1*1*1 convolution superimposed in parallel; the dilation rates of the first group of dilated convolutions are 1, 2, and 4; the dilation rates of the second group of dilated convolutions are 1, 2, and 5; in addition, each convolution with a kernel of 3*3*3 is followed by a group normalization layer with a group size of 8 and a ReLu linear unit; each convolution in the group produces the same number of output dimensions; The encoded feature map is input into a decoder consisting of three global context aggregation modules for decoding, and the segmentation result is finally output; The feature map is first upsampled by deconvolution with a stride of 2 in each dimension and a convolution kernel of 2 * 2 * 2, and then the feature information with the same resolution in the encoding path and the decoding path is fused by element-wise summation; the fused features are subjected to two 1*1*1 convolutions to obtain two feature maps, denoted as feature map A and feature map B, and two branches are also formed, denoted as branch Z1 and branch Z2; in order to efficiently collect the background information of each spatial position, the global average pooling operation is first performed in the Z1 branch to generate a global context feature representation and then add it to A; then, the obtained features are applied to the Sigmoid layer to obtain the feature weight map S Finally, S is element-wise multiplied with the feature map A after the convolution operation to obtain the feature map D. Subsequently, the feature map D is sent to a three-dimensional convolution with a convolution kernel size of 3*3*3, followed by a group convolution with a group size of 8 and a nonlinear activation function ReLu, and finally the recalibrated feature Y1 is obtained; in the Z2 branch, the feature map B is subjected to a 3D convolution operation with a step size of 2 in each direction and a convolution kernel size of 3 to obtain the feature map Y2. Finally, the outputs Y1 and Y2 of the two branches are spliced ​​together along the direction of the channel to form the global context aggregation module output Y.

2. The deep learning-based 3D MRI brain tumor segmentation method according to claim 1, characterized in that: The preprocessing described in S1 is to crop the original three-dimensional brain MRI image of 155*240*240 to 150*192*192 and remove redundant background pixels; normalize the data before inputting into the network to reduce the learning difficulty of the network, and perform diversified operations on the training data in an online manner, including random scaling, random flipping and random cropping along the three-dimensional direction. The final three-dimensional image size sent to the network is 96*144*144; finally, the brain data set is divided into training set, validation set and test set in a ratio of 8:1:

1.

3. The deep learning-based 3D MRI brain tumor segmentation method according to claim 1, characterized in that: In step S2, in the deep convolutional neural network, the loss function is set to: ; in, , , These are the loss functions output by the three global context aggregation modules, all of which are Dice losses. The specific calculation method is as follows: ; in, are the real tumor pixels manually annotated. are the tumor pixels predicted by the model.

Citation Information

Patent Citations

  • Unsupervised cross-domain self-adaptive medical image segmentation method based on deep adversarial learning

    AU2020103905A4

  • Brain tumor segmentation method and system based on multi-scale cavity convolutional neural network

    CN110910405A