A Brain Tumor Segmentation Method Based on Uncertainty Estimation
By introducing an uncertainty estimation mechanism into the brain tumor segmentation method, combining Monte Carlo simulation and image processing modules, multiple uncertain images are output, which solves the problem of lack of uncertainty estimation and high computational complexity in the existing methods, and achieves more accurate and stable brain tumor segmentation, providing a more reliable diagnostic basis.
Patent Information
- Application Number
- CN202510336226.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-03-21
AI Technical Summary
The existing brain tumor segmentation methods lack uncertainty estimation mechanism, which is difficult to accurately reflect the uncertainty of multimodal data, and the calculation complexity is high, making it difficult to achieve real-time application.
Using a brain tumor segmentation method based on uncertainty estimation, by inputting T1, T1c, T2 and Flair images into the trained image segmentation model, the mean image, random uncertain image, cognitive uncertain image, entropy uncertain image and mutual information uncertain image are output, and combined with the Monte Carlo simulation and image processing module, the model's attention to high uncertainty areas is dynamically adjusted.
It improves the accuracy and stability of segmentation results, provides doctors with a more reliable diagnostic basis, improves the segmentation accuracy of tumor boundaries and complex areas, and is suitable for a variety of medical image analysis tasks.
Smart Images

Figure CN119863624B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image analysis, and in particular relates to a brain tumor segmentation method based on uncertainty estimation. Background Art
[0002] In the vast field of medical diagnosis and treatment planning, brain tumor segmentation plays a vital role. It is not only the basis for doctors to assess the size, location and morphology of tumors, but also the key basis for formulating personalized treatment plans, predicting patient prognosis and monitoring treatment effects. Accurate brain tumor segmentation can significantly improve the accuracy of surgery, reduce unnecessary tissue damage, and optimize the formulation of radiotherapy plans to ensure that the radiation dose accurately covers the tumor area while maximizing the protection of surrounding normal tissues. Therefore, improving the accuracy and reliability of brain tumor segmentation is of immeasurable value in improving the quality of life of patients and prolonging their survival.
[0003] Traditional single-modality segmentation methods, such as segmentation techniques based on MRI modalities such as T1, T2, and FLAIR, have shown certain application value in specific situations, but their inherent limitations cannot be ignored. The diversity and complexity of brain tumors require segmentation methods to be able to capture subtle characteristic changes of tumors in different modalities. However, single-modality segmentation methods only rely on information from a single modality and are difficult to fully reflect the overall picture of the tumor, especially when the tumor boundaries are blurred, the morphology is irregular, or the modality information is missing. The segmentation results are often unsatisfactory. This may not only lead to the omission or misjudgment of the tumor area, but may also affect the accuracy and effectiveness of subsequent treatment plans, increasing the risk of clinical decision-making.
[0004] In order to overcome the limitations of single-modality segmentation methods, researchers have begun to explore multimodal data fusion methods. By fusing images of multiple MRI modalities such as T1, T2, and FLAIR, multimodal segmentation methods can comprehensively utilize the information provided by different modalities to more comprehensively capture the characteristics of the tumor and improve the accuracy and stability of segmentation. This fusion strategy can not only enhance the ability to identify tumor boundaries, but also reveal subtle structural changes inside the tumor, providing doctors with richer diagnostic information.
[0005] Although multimodal segmentation methods have made significant progress in improving segmentation accuracy, existing methods generally lack uncertainty estimation mechanisms. Uncertainty estimation is a method for evaluating the reliability of model prediction results. It can provide additional information about the confidence of the prediction results, which helps doctors understand the credibility of the segmentation results more comprehensively. Existing uncertainty estimation techniques are mostly combined with single-modal segmentation methods, lack effective processing and fusion of multimodal features, and are difficult to accurately reflect the uncertainty of multimodal data. In addition, some uncertainty estimation methods have high computational complexity and are difficult to implement in real time in actual clinical environments.
[0006] Therefore, it is of great significance to develop a brain tumor segmentation method based on uncertainty estimation, which not only helps to improve the accuracy and stability of the segmentation results, but also provides more reliable diagnostic basis for doctors, further promoting the progress of the field of brain tumor treatment. Summary of the Invention
[0007] The purpose of the present invention is to solve the above problems and provide a brain tumor segmentation method based on uncertainty estimation.
[0008] In order to achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0009] A brain tumor segmentation method based on uncertainty estimation inputs the T1 image, T1c image, T2 image and Flair image of the same brain tumor into a trained image segmentation model, and outputs the mean image, stochastic uncertainty image, epistemic uncertainty image, entropy uncertainty image and mutual information uncertainty image;
[0010] The image segmentation model includes 4 encoder modules, 1 decoder module, 1 Monte Carlo simulation module and 1 image processing module;
[0011] Both the encoder module and the decoder module contain dropout layers;
[0012] The 4 encoder modules respectively downsample the T1 image, T1c image, T2 image and Flair image, and 1 decoder module simultaneously upsamples the downsampled T1 image, T1c image, T2 image and Flair image to obtain image X;
[0013] The Monte Carlo simulation module repeatedly calls the 4 encoder modules and 1 decoder module to perform times of Monte Carlo sampling to obtain images X, = 10 - 50;
[0014] The image processing module processes the images X to obtain the mean image, stochastic uncertainty image, epistemic uncertainty image, entropy uncertainty image and mutual information uncertainty image;
[0015] The calculation formula of the mean image is as follows:
[0016] ;
[0017] In the formula, is the images X obtained by times of Monte Carlo sampling;
[0018] Random uncertainty image The calculation formula is as follows:
[0019] ;
[0020] In the formula, represents the expectation of the posterior distribution (under the condition of given data D) with respect to the model parameter , and represents the variance of the prediction under the condition of given model parameter ;
[0021] Epistemic uncertainty image The calculation formula is as follows:
[0022] ;
[0023] In the formula, represents the variance of the posterior distribution (under the condition of given data D) with respect to the model parameter , and represents the expected value of the prediction under the condition of given model parameter ;
[0024] Entropy uncertainty image The calculation formula is as follows:
[0025] ;
[0026] In the formula, is the probability mass function of the random variable X at x;
[0027] Mutual information uncertainty image The calculation formula is as follows:
[0028] ;
[0029] In the formula, P(x) and P(y) are the marginal probability mass functions of the random variables X and Y respectively, and P(x,y) is the joint probability mass function of X and Y evaluated at x and y;
[0030] The mean image is the predicted segmentation image, which is the brain tumor MRI image with 3 segmentation regions obtained by prediction. The 3 segmentation regions are the whole tumor region, the tumor core region, and the enhanced tumor region respectively;
[0031] The confidence level of the segmentation result in the predicted segmentation image can be seen from the random uncertainty image, the epistemic uncertainty image, the entropy uncertainty image, and the mutual information uncertainty image.
[0032] The present invention innovatively combines an encoder module and a decoder module with an uncertainty estimation module (composed of a Monte Carlo simulation module and an image processing module), introducing uncertainty information during the encoding and decoding processes to achieve the synchronization of multi-scale feature extraction and uncertainty evaluation. Such a combination can achieve unexpected results, not only dynamically adjusting the model's attention to high-uncertainty regions and improving the segmentation accuracy of boundaries and complex regions.
[0033] Monte Carlo sampling is applied at the bottom layer of the network to generate multiple segmentation results, and the mean and variance of these segmentation results are calculated. Then, these segmentation results are used to generate different types of uncertainty maps, including the aleatoric uncertainty image, the epistemic uncertainty image, the entropy uncertainty image, and the mutual information uncertainty image. These uncertainty maps can provide insights into the credibility of the segmentation results, making the model more effective when highly reliable predictions are required.
[0034] As a preferred technical solution:
[0035] For a brain tumor segmentation method based on uncertainty estimation as described above, the dropout rate of each dropout layer is 0.3.
[0036] For a brain tumor segmentation method based on uncertainty estimation as described above, the decoder module includes a multi-modal fusion module a and a multi-modal fusion module b, and their working processes are as follows:
[0037] (a) Connect the T2 image and the T1 image to form a student modality, and connect the Flair image and the T1c image to form a teacher modality. ;
[0038] (b) Apply a convolutional block operation with a size of 3×3×3 to in sequence, and operate to obtain the student attention weight ;
[0039] Apply a convolutional block operation with a size of 5×5×5 to in sequence, and operate to obtain the student attention weight ;
[0040] Multiply element-wise with , and at the same time multiply element-wise with , and add the two products element-wise to obtain the student weighted feature representation ;
[0041] Apply the convolution block operation of size 3×3×3 to sequentially, and obtain the teacher attention weights through the operation ;
[0042] Apply the convolution block operation of size 5×5×5 to sequentially, and obtain the teacher attention weights through the operation ;
[0043] Multiply element-wise with , and at the same time multiply element-wise with . Then add the two products element-wise to obtain the teacher weighted feature representation ;
[0044] Connect and to form the spatial feature representation ;
[0045] Apply average pooling operations to and respectively, connect the results, and then apply multi-layer perceptron operations and to the connected results sequentially to obtain the attention weights ;
[0046] Apply max pooling operations to and respectively, connect the results, and then apply multi-layer perceptron operations and to the connected results sequentially to obtain the attention weights ;
[0047] Multiply element-wise with , and at the same time multiply element-wise with . Then add the two products element-wise to obtain the student weighted feature representation ;
[0048] Multiply element-wise with , and at the same time multiply element-wise with . Then add the two products element-wise to obtain the teacher weighted feature representation ;
[0049] Connect and to form the channel attention feature ;
[0050] (c) Connect and to form a fused feature representation .
[0051] To learn cross-modal complementary information, this method proposes a multi-modal teacher-student fusion (MTSF) strategy, including a spatial attention fusion module (SAFM) and a channel attention fusion module (CAFM), as Figure 2 shown. The former is used to learn the fused feature representation from a spatial perspective, and the latter is used to extract the feature representation from a channel perspective. Initially, Flair and T1c are connected to form the teacher modality. Similarly, T2 and T1 are connected to obtain the student modality.
[0052] In the spatial attention fusion module, convolutional blocks with sizes of 3×3×3 and 5×5×5 are applied to the student modality and the teacher modality respectively, and each convolutional block is followed by a operation. This process can obtain the student attention weights , , as well as the teacher attention weights , . Then, these weights are multiplied element-wise and added to the original student modality and the teacher modality respectively, and the weighted feature representations and can be obtained. Subsequently, the student spatial feature representation and the teacher spatial feature representation are connected to produce the final spatial feature representation . In this way, the spatial attention fusion module ensures that the fused feature representation captures the most informative spatial details in each modality.
[0053] In the channel attention module, global average pooling (GAP) and global max pooling (GMP) operations are applied to the student modality and the teacher modality respectively. The results of each pooling operation are connected, and subsequently, a multi-layer perceptron and a function are used to obtain the attention weights and . Then, these weights are multiplied by the original student modality and the teacher modality and added together to obtain the student and teacher weighted feature representations and . Subsequently, these feature representations are connected to achieve the final channel attention feature Therefore, the channel attention fusion module promotes the extraction of complementary information by emphasizing important channels and suppressing irrelevant channels.
[0054] Finally, the spatial feature representation and the channel attention feature are concatenated to achieve the fused feature representation By combining these two attention modules, the proposed multi-modal teacher-student fusion strategy can effectively learn comprehensive and complementary feature representations. This dual attention method takes advantage of the spatial and channel information, improving the overall performance of the multi-modal fusion process.
[0055] As described above, in a brain tumor segmentation method based on uncertainty estimation, the structures of the 4 encoder modules are the same, each including a downsampling initial block, two downsampling convolutional normalization activation modules b, and two convolutional normalization activation dropout blocks connected in sequence;
[0056] The downsampling initial block consists of a standard convolutional normalization activation module and a downsampling convolutional normalization activation module a connected in sequence;
[0057] Both of the two convolutional normalization activation dropout blocks consist of a downsampling convolutional normalization activation module a and a dropout layer connected in sequence;
[0058] The output of the last convolutional normalization activation dropout block is the output of the encoder module;
[0059] The decoder module also includes a first upsampling fusion convolutional dropout block, a second upsampling fusion convolutional dropout block, a first upsampling fusion convolutional block, a second upsampling fusion convolutional block, a third upsampling fusion convolutional block, a first convolutional block, a second convolutional block, a third convolutional block, a fourth convolutional block, a fifth convolutional block, a first upsampling module b, a second upsampling module b, a third upsampling module b, a fourth upsampling module b, a first element-wise addition module, a second element-wise addition module, a third element-wise addition module, and a fourth element-wise addition module;
[0060] The multi-modal fusion module b, the first upsampling fusion convolutional dropout block, the second upsampling fusion convolutional dropout block, the first upsampling fusion convolutional block, the second upsampling fusion convolutional block, the third upsampling fusion convolutional block, the fifth convolutional block, and the fourth element-wise addition module are connected in sequence;
[0061] The first upsampling fusion convolutional dropout block, the first convolutional block, the first upsampling module b, the first element-wise addition module, the second upsampling module b, the second element-wise addition module, the third upsampling module b, the third element-wise addition module, the fourth upsampling module b, and the fourth element-wise addition module are connected in sequence;
[0062] The second upsampling fusion convolutional dropout block, the second convolutional block, and the first element-wise addition module are connected in sequence;
[0063] The first upsampling fusion convolutional block, the third convolutional block, and the second element-wise addition module are connected in sequence;
[0064] The second upsampling fusion convolutional block, the fourth convolutional block, and the third element-wise addition module are connected in sequence;
[0065] Both the first upsampling fusion convolutional dropout block and the second upsampling fusion convolutional dropout block are composed of an upsampling module a, a multimodal fusion module a, a standard convolutional normalization activation module, and a dropout layer connected in sequence;
[0066] Both the first upsampling fusion convolutional block, the second upsampling fusion convolutional block, and the third upsampling fusion convolutional block are composed of an upsampling module a, a multimodal fusion module a, and a standard convolutional normalization activation module connected in sequence;
[0067] The output of the fourth element-wise addition module is the output of the decoder;
[0068] The decoder part is composed of multiple upsampling and convolutional modules, aiming to gradually restore the feature maps compressed by the encoder; the feature maps are processed through upsampling (green modules), convolution (blue modules), and multimodal fusion modules (purple modules) to gradually increase the resolution of the feature maps and finally reconstruct them to the original size; the decoder path also integrates a deep supervision mechanism (yellow and green modules) to promote the generation of multi-scale segmentation results and enhance the learning ability of the model during training;
[0069] The standard convolutional normalization activation module is composed of a convolutional layer, an instance normalization layer, and a LeakyReLU activation function layer connected in sequence;
[0070] Both the downsampling convolutional normalization activation module a and the downsampling convolutional normalization activation module b are composed of a convolutional layer with a stride of 2, an instance normalization layer, and a LeakyReLU activation function layer connected in sequence;
[0071] All the connections in sequence mean connecting in sequence from front to back according to the data flow direction;
[0072] In the present invention, "a" and "b" are used to distinguish whether a certain module is an independent module or jointly forms a module with other modules. For example, the downsampling convolutional normalization activation module a refers to the downsampling convolutional normalization activation module that jointly forms a module with other modules, while the downsampling convolutional normalization activation module b refers to the downsampling convolutional normalization activation module that is an independent module.
[0073] As described above, for a brain tumor segmentation method based on uncertainty estimation, the loss function of the image segmentation model The calculation formula is as follows:
[0074] ;
[0075] ;
[0076] ;
[0077] ;
[0078] ;
[0079] In the formula, , are hyperparameters, = = 1, represents the number of voxels in the image, represents the voxel predicted value, represents the voxel true value, is a constant used to prevent division by zero, is 10 -5 , represents the -th Monte Carlo sampling voxel predicted value. Other unexplained symbols can be eliminated during the calculation process, and they should not affect the accuracy and effectiveness of the results;
[0080] This is a hybrid loss function that combines the traditional Dice loss (L Dice ) with an uncertainty-based penalty term (L Uncertainty ). By adding uncertainty information to the loss function, it further enhances the performance of the network in the segmentation task.
[0081] For a brain tumor segmentation method based on uncertainty estimation as described in any of the above items, before the T1 image, T1c image, T2 image, and Flair image of the same brain tumor are jointly input into the trained image segmentation model, preprocessing is also performed, and the preprocessing is sequentially performed with gray normalization processing, pixel adjustment, and quantization.
[0082] For a brain tumor segmentation method based on uncertainty estimation as described above, the training steps of the image segmentation model are as follows:
[0083] (a) Collect brain tumor cases, ≥285, and each brain tumor case has a T1 image, a T1c image, a T2 image, and a Flair image;
[0084] (b) Preprocess the T1 image, T1c image, T2 image, and Flair image; obtain the true segmentation image of each brain tumor case, where the true segmentation image is the brain tumor MRI image with the three segmentation regions obtained through manual annotation;
[0085] (c) Use the T1 image, T1c image, T2 image, Flair image, and true segmentation image corresponding to a number of brain tumor cases to construct a training set and a test set;
[0086] (d) Train the image segmentation model using the training set. During training, use the T1 image, T1c image, T2 image, and Flair image as the input of the image segmentation model, and use the true segmentation image as the theoretical output of the image segmentation model. Continuously adjust the weight parameters of the image segmentation model (the weights and bias parameters of the convolutional layer, the instance normalization layer parameters) until the image segmentation model converges (i.e., the loss function value gradually decreases and tends to a stable value);
[0087] (e) Test the trained image segmentation model using the test set.
[0088] Beneficial effects:
[0089] A brain tumor segmentation method based on uncertainty estimation of the present invention improves the segmentation accuracy by introducing uncertainty estimation: existing brain tumor segmentation methods usually ignore the uncertainty of the model output, resulting in errors easily occurring when dealing with complex or blurred boundary regions; in contrast, by introducing uncertainty estimation, the present invention can not only accurately capture the uncertain regions of the model, but also improve the overall segmentation accuracy, especially showing better performance in the segmentation of tumor boundaries and irregular shapes.
[0090] A brain tumor segmentation method based on uncertainty estimation of the present invention improves the reliability of diagnosis by quantifying uncertainty: traditional segmentation methods only output the result image and lack a quantitative evaluation of the segmentation accuracy; by introducing uncertainty estimation, the present invention provides an uncertainty measure for each pixel, which can help doctors better understand the credibility level of the model in different regions, and thus more accurately judge the tumor location and size, providing reliable auxiliary information for clinical diagnosis.
[0091] A brain tumor segmentation method based on uncertainty estimation of the present invention has strong scalability and is applicable to various medical image analysis tasks: the uncertainty estimation method adopted by the present invention is not limited to brain tumor segmentation, but also has strong generality and can be extended to other medical image analysis tasks, such as automatic segmentation and diagnosis of organs such as the heart, liver, and lungs, and has broad clinical application prospects. Description of the Drawings
[0092] Figure 1 and Figure 2 is a schematic diagram of the overall network architecture of a brain tumor segmentation method based on uncertainty estimation in Embodiment 1;
[0093] Figure 3 is Figure 2 a schematic diagram of the network architecture of the purple module in;
[0094] Figure 4 is a schematic diagram of the overall network architecture of a brain tumor segmentation method in Comparative Example 1;
[0095] Figure 5 is a schematic diagram of the overall network architecture of a brain tumor segmentation method in Comparative Example 2;
[0096] Figure 6 is a schematic diagram of the overall network architecture of a brain tumor segmentation method in Comparative Example 3;
[0097] Figure 7 is a schematic diagram of the overall network architecture of a brain tumor segmentation method in Comparative Example 4;
[0098] Figure 8 is the random uncertainty image output in Embodiment 1;
[0099] Figure 9 is the epistemic uncertainty image output in Embodiment 1;
[0100] Figure 10 is the entropy uncertainty image output in Embodiment 1;
[0101] Figure 11 is the mutual information uncertainty image output in Embodiment 1. Detailed implementation manners
[0102] The present invention will be further described below in conjunction with the detailed implementation manners. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0103] Embodiment 1
[0104] A brain tumor segmentation method based on uncertainty estimation, as Figures 1 - 3 shown, the specific steps are as follows:
[0105] Step 1, construct an image segmentation model;
[0106] The image segmentation model includes 4 encoder modules, 1 decoder module, 1 Monte Carlo simulation module, and 1 image processing module;
[0107] The 4 encoder modules respectively perform downsampling on T1 images, T1c images, T2 images, and Flair images;
[0108] 1 decoder module simultaneously performs upsampling on the downsampled T1 images, T1c images, T2 images, and Flair images to obtain image X;
[0109] The Monte Carlo simulation module repeatedly calls the 4 encoder modules and 1 decoder module to perform times of Monte Carlo sampling to obtain images X, = 10 - 50;
[0110] The image processing module processes images X to obtain a mean image, a random uncertainty image, an epistemic uncertainty image, an entropy uncertainty image, and a mutual information uncertainty image; The mean image is the predicted segmentation image, and the predicted segmentation image is the brain tumor MRI image with 3 segmentation regions obtained through prediction. The 3 segmentation regions are the whole tumor region, the tumor core region, and the enhanced tumor region;
[0111] The structures of the 4 encoder modules are the same, and each includes a downsampling initial block, two downsampling convolution normalization activation modules b, and two convolution normalization activation dropout blocks connected in sequence;
[0112] The downsampling initial block consists of a standard convolution normalization activation module and a downsampling convolution normalization activation module a connected in sequence; The standard convolution normalization activation module consists of a convolution layer, an instance normalization layer, and a LeakyReLU activation function layer connected in sequence;
[0113] Both the downsampling convolution normalization activation module a and the downsampling convolution normalization activation module b consist of a convolution layer with a stride of 2, an instance normalization layer, and a LeakyReLU activation function layer connected in sequence;
[0114] Both of the two convolution normalization activation dropout blocks consist of a downsampling convolution normalization activation module a and a dropout layer connected in sequence. The dropout rate of the dropout layer is 0.3;
[0115] The output of the last convolution normalization activation dropout block is the output of the encoder module;
[0116] The decoder module includes a multi-modal fusion module a, a multi-modal fusion module b, a first upsampling fusion convolution dropout block, a second upsampling fusion convolution dropout block, a first upsampling fusion convolution block, a second upsampling fusion convolution block, a third upsampling fusion convolution block, a first convolution block, a second convolution block, a third convolution block, a fourth convolution block, a fifth convolution block, a first upsampling module b, a second upsampling module b, a third upsampling module b, a fourth upsampling module b, a first element-wise addition module, a second element-wise addition module, a third element-wise addition module, and a fourth element-wise addition module;
[0117] The multi-modal fusion module b, the first upsampling fusion convolution dropout block, the second upsampling fusion convolution dropout block, the first upsampling fusion convolution block, the second upsampling fusion convolution block, the third upsampling fusion convolution block, the fifth convolution block, and the fourth element-wise addition module are connected in sequence;
[0118] The first upsampling fusion convolution dropout block, the first convolution block, the first upsampling module b, the first element-wise addition module, the second upsampling module b, the second element-wise addition module, the third upsampling module b, the third element-wise addition module, the fourth upsampling module b, and the fourth element-wise addition module are connected in sequence;
[0119] The second upsampling fusion convolution dropout block, the second convolution block, and the first element-wise addition module are connected in sequence;
[0120] The first upsampling fusion convolution block, the third convolution block, and the second element-wise addition module are connected in sequence;
[0121] The second upsampling fusion convolution block, the fourth convolution block, and the third element-wise addition module are connected in sequence;
[0122] Both the first upsampling fusion convolution dropout block and the second upsampling fusion convolution dropout block are composed of an upsampling module a, a multi-modal fusion module a, a standard convolution normalization activation module, and a dropout layer connected in sequence. The dropout rate of the dropout layer is 0.3;
[0123] The first upsampling fusion convolution block, the second upsampling fusion convolution block, and the third upsampling fusion convolution block are all composed of an upsampling module a, a multi-modal fusion module a, and a standard convolution normalization activation module connected in sequence;
[0124] The output of the fourth element-wise addition module is the output of the decoder;
[0125] All the above-mentioned sequential connections refer to sequential connections from front to back according to the data flow direction;
[0126] The working processes of both the multi-modal fusion module a and the multi-modal fusion module b are as follows:
[0127] (i)Connect the T2 image and the T1 image to form a student modality Connect the Flair image and the T1c image to form the teacher modality ;
[0128] (ii) Apply a convolutional block operation of size 3×3×3 to and operations to obtain the student attention weights ;
[0129] Apply a convolutional block operation of size 5×5×5 to and operations to obtain the student attention weights ;
[0130] Multiply element-wise with , and at the same time multiply element-wise with . Add the two products element-wise to obtain the student weighted feature representation ;
[0131] Apply a convolutional block operation of size 3×3×3 to and operations to obtain the teacher attention weights ;
[0132] Apply a convolutional block operation of size 5×5×5 to and operations to obtain the teacher attention weights ;
[0133] Multiply element-wise with , and at the same time multiply element-wise with . Add the two products element-wise to obtain the teacher weighted feature representation ;
[0134] Connect and to form the spatial feature representation ;
[0135] Apply average pooling operations to and respectively and connect the results. Then, apply a multi-layer perceptron operation and operations to the connected result in sequence to obtain the attention weights ;
[0136] Apply average pooling operations to and After applying the max pooling operation and concatenating the results, successively apply the multi-layer perceptron operation and operation to obtain the attention weights ;
[0137] Multiply element-wise with , and at the same time multiply element-wise with . Add the two products element-wise to obtain the student weighted feature representation ;
[0138] Multiply element-wise with , and at the same time multiply element-wise with . Add the two products element-wise to obtain the teacher weighted feature representation ;
[0139] Connect and to form the channel attention feature ;
[0140] (iii) Connect and to form the fused feature representation ;
[0141] The module that performs the operation to obtain from the T2 image, T1 image, Flair image, and T1c image is denoted as the spatial attention fusion module (SAFM); the module that performs the operation to obtain from the T2 image, T1 image, Flair image, and T1c image is denoted as the channel attention fusion module (CAFM);
[0142] Step 2: Determine the loss function of the image segmentation model;
[0143] The loss function of the image segmentation model is calculated as follows:
[0144] ;
[0145] ;
[0146] ;
[0147] ;
[0148] ;
[0149] In the formula, = = 1, representing the number of voxels in the image, representing the voxel predicted value, representing the voxel true value, is 10 -5 , representing the predicted value of the voxel during the th Monte Carlo sampling; The loss function of this embodiment is the uncertainty loss function (UL);
[0150] Step three, training the image segmentation model;
[0151] (a) Collect 285 brain tumor cases (from the dataset of the Brain Tumor Segmentation Challenge (BraTS2018), which is widely used in brain tumor segmentation tasks), and each brain tumor case has T1 image, T1c image, T2 image, and Flair image;
[0152] (b) Preprocess the T1 image, T1c image, T2 image, and Flair image (i.e., each image is sequentially subjected to gray normalization, pixel adjustment, and quantization. The size of the image before preprocessing is 240×240×155 pixels. Gray normalization means adjusting the gray value of the image to a normal distribution with a mean of 0 and a standard deviation of 1, and pixel adjustment means adjusting the image to 128×128×128 pixels); Obtain the true segmentation image of each brain tumor case, and the true segmentation image is the brain tumor MRI image with the three segmentation regions obtained through manual annotation;
[0153] (c) Construct a training set and a test set using the T1 image, T1c image, T2 image, Flair image, and true segmentation image corresponding to 285 brain tumor cases. The training set corresponds to 228 brain tumor cases, and the test set corresponds to 57 brain tumor cases;
[0154] (d) Train the image segmentation model using the training set. During training, use the T1 image, T1c image, T2 image, and Flair image as the input of the image segmentation model, and use the true segmentation image as the theoretical output of the image segmentation model. Continuously adjust the weight parameters of the image segmentation model until the image segmentation model converges;
[0155] (e) Test the trained image segmentation model using the test set;
[0156] (f)
[0157] The dice similarity coefficient (DSC) and Hausdorff distance (HD) are used to characterize the segmentation accuracy of the trained image segmentation model. The calculation formulas are as follows:
[0158] ;
[0159] In the formula, P and G respectively represent the predicted region Q (the whole tumor region WT, the tumor core region TC, or the enhanced tumor region ET) and the ground truth region Q. respectively represent the size of the intersection of the predicted region Q and the ground truth region Q. and respectively represent the total sizes of the predicted region Q and the ground truth region Q.
[0160] ;
[0161] In the formula, X and Y respectively represent the boundaries of the predicted region Q and the ground truth region Q. is the Euclidean distance. is the maximum value of the minimum distances from each point on the boundary of the predicted region Q to the boundary of the ground truth region Q. is the maximum value of the minimum distances from each point on the boundary of the ground truth region Q to the boundary of the predicted region Q.
[0162] The test results are shown in Table 1.
[0163] Table 1
[0164]
[0165] Step 4: Preprocess the T1 image, T1c image, T2 image, and Flair image of the same brain tumor (perform gray normalization, pixel adjustment, and quantization in sequence). The preprocessed images are jointly input into the trained image segmentation model, which outputs the mean image, random uncertainty image (as shown in Figure 8 ), cognitive uncertainty image (as shown in Figure 9 ), entropy uncertainty image (as shown in Figure 10 ), and mutual information uncertainty image (as shown in Figure 11 ).
[0166] The traditional teacher-student mechanism realizes model compression by transferring the knowledge of a large model (teacher model) to a smaller model (student model) through knowledge distillation, enabling the small model to achieve high performance in resource-constrained environments, so as to solve the balance between model complexity and inference efficiency. The present invention innovatively applies the teacher-student mechanism to multi-modal feature fusion, designs SAFM and CAFM, and integrates the feature information of the teacher and student modalities respectively, thereby improving the performance of the model in tumor segmentation tasks. This fusion strategy not only enhances the model's ability to capture complex structures, but also significantly improves the segmentation accuracy and robustness. This fusion strategy overcomes the defects of traditional multi-modal fusion methods, that is, traditional methods usually adopt a direct splicing method, ignoring the differences and importance of different modalities in the feature space, resulting in the loss of feature effectiveness in the fusion process, being unable to make full use of the complementary information between multi-modalities, and thus affecting the segmentation accuracy.
[0167] Comparative Example 1
[0168] A brain tumor segmentation method, as Figure 4 shown, is basically the same as Example 1, except that:
[0169] In step one, the image segmentation model does not include a Monte Carlo simulation module and an image processing module, and the image X output by the decoder module is the predicted segmentation image;
[0170] In step two, the loss function of the image segmentation model has the following calculation formula:
[0171] ;
[0172] ;
[0173] In the formula, represents the number of voxels in the image, represents the predicted value of the voxel , represents the voxel 's true value, is 10 -5 .
[0174] The test results are shown in Table 2.
[0175] Table 2
[0176]
[0177] Comparative Example 2
[0178] A brain tumor segmentation method, as Figure 5 shown, is basically the same as Comparative Example 1, except that:
[0179] In step one, the working processes of multimodal fusion module a and multimodal fusion module b are as follows:
[0180] (i) Connect the T2 image and the T1 image to form the student modality , and connect the Flair image and the T1c image to form the teacher modality ;
[0181] (ii) Apply a convolutional block operation with a size of 3×3×3 to in sequence, and perform operations to obtain the student attention weight ;
[0182] Apply a convolutional block operation with a size of 5×5×5 to in sequence, and perform operations to obtain the student attention weight ;
[0183] Multiply element-wise with , and at the same time multiply element-wise with , and add the two products element-wise to obtain the student weighted feature representation ;
[0184] Apply a convolutional block operation with a size of 3×3×3 to in sequence, and perform operations to obtain the teacher attention weight ;
[0185] Apply a convolutional block operation with a size of 5×5×5 to in sequence, and perform operations to obtain the teacher attention weight ;
[0186] Multiply element-wise with , and at the same time multiply element-wise with , and add the two products element-wise to obtain the teacher weighted feature representation ;
[0187] Connect and to form the spatial feature representation .
[0188] The test results are shown in Table 3
[0189] Table 3
[0190]
[0191] Comparative Example 3
[0192] A brain tumor segmentation method, as Figure 6 shown, is basically the same as Comparative Example 1, except that:
[0193] In Step 1, the working processes of both the multi-modal fusion module a and the multi-modal fusion module b are as follows:
[0194] (i) Connect the T2 image and the T1 image to form the student modality , and connect the Flair image and the T1c image to form the teacher modality ;
[0195] (ii) Apply average pooling operations to and respectively, connect the results, and then apply a multi-layer perceptron operation and operations to the connected results in sequence to obtain the attention weight ;
[0196] Apply max pooling operations to and respectively, connect the results, and then apply a multi-layer perceptron operation and operations to the connected results in sequence to obtain the attention weight ;
[0197] Multiply element-wise with , and at the same time multiply element-wise with , and add the two products element-wise to obtain the student weighted feature representation ;
[0198] Multiply element-wise with , and at the same time multiply element-wise with , and add the two products element-wise to obtain the teacher weighted feature representation ;
[0199] Connect and to form the channel attention feature .
[0200] The test results are shown in Table 4.
[0201] Table 4
[0202]
[0203] Comparative Example 4
[0204] A brain tumor segmentation method (using the Baseline model as the basic model), as Figure 7 shown, is basically the same as Comparative Example 1, except that:
[0205] Both the multimodal fusion module a and the multimodal fusion module b are replaced with a series module.
[0206] The test results are shown in Table 5.
[0207] Table 5
[0208]
[0209] It can be seen from Tables 1 to 5 that on the Baseline, without adding CAFM, SAFM, and UL, the average DSC is 82.3% and the average HD is 5.4; on the Baseline, when CAFM is added, the average DSC is 82.7% and the average HD is 4.8; on the Baseline, when SAFM is added, the average DSC is 83.2% and the average HD is 4.9; on the Baseline, when CAFM and SAFM are added, the average DSC is 84.1% and the average HD is 4.5; on the Baseline, when CAFM, SAFM, and UL are added, the average DSC is 84.3% and the average HD is 4.2. By comparison, it can be seen that the addition of CAFM, SAFM, and UL can effectively improve the accuracy of the segmentation model, and the segmentation accuracy and robustness of the model have been significantly improved, making it suitable for brain tumor segmentation tasks with high accuracy requirements.
Claims
1. A brain tumor segmentation method based on uncertainty estimation, characterized in that: The T1 image, T1c image, T2 image and Flair image of the same brain tumor are input into the trained image segmentation model, which outputs the mean image, random uncertainty image, cognitive uncertainty image, entropy uncertainty image and mutual information uncertainty image; The image segmentation model includes 4 encoder modules, 1 decoder module, 1 Monte Carlo simulation module, and 1 image processing module; Both the encoder and decoder modules contain dropout layers; The four encoder modules downsample the T1 image, T1c image, T2 image, and Flair image respectively, and the one decoder module upsamples the downsampled T1 image, T1c image, T2 image, and Flair image to obtain image X; The Monte Carlo simulation module repeatedly calls four encoder modules and one decoder module to perform Monte Carlo sampling image X, =10-50; Image processing module The image X is processed to obtain the mean image, random uncertainty image, cognitive uncertainty image, entropy uncertainty image and mutual information uncertainty image; The mean image is the predicted segmentation image, which is a brain tumor MRI image with three segmentation regions obtained through prediction. The three segmentation regions are the complete tumor region, the tumor core region, and the enhancement tumor region.
2. The brain tumor segmentation method based on uncertainty estimation according to claim 1, characterized in that: The dropout rate of each dropout layer is 0.
3.
3. The brain tumor segmentation method based on uncertainty estimation according to claim 1, characterized in that: The decoder module includes multimodal fusion module a and multimodal fusion module b, and the workflows of both are as follows: (a) Concatenate T2 and T1 images to form the student mode , connect the Flair image and T1c image to form the teacher modality ; (b) Apply a convolutional block operation of size 3×3×3, Operation to obtain student attention weight ; Sequentially Apply a convolutional block operation of size 5×5×5, Operation to obtain student attention weight ; Will and Perform element-by-element multiplication and and Perform element-by-element multiplication and add the two products element-by-element to obtain the student weighted feature representation ; Sequentially Apply a convolutional block operation of size 3×3×3, Operation to get the teacher's attention weight ; Sequentially Apply a convolutional block operation of size 5×5×5, Operation to get the teacher's attention weight ; Will and Perform element-by-element multiplication and and Perform element-by-element multiplication and add the two products element-by-element to obtain the teacher weighted feature representation ; Will and Concatenate to form spatial feature representation ; Respectively and After applying the average pooling operation and concatenating the results, the multi-layer perceptron operation and Operation to get attention weight ; Respectively and After applying the max pooling operation and concatenating the results, the multi-layer perceptron operation and Operation to get attention weight ; Will and Perform element-by-element multiplication and and Perform element-by-element multiplication and add the two products element-by-element to obtain the student weighted feature representation ; Will and Perform element-by-element multiplication and and Perform element-by-element multiplication and add the two products element-by-element to obtain the teacher weighted feature representation ; Will and Concatenate to form channel attention features ; (c) and Concatenate to form fused feature representation .
4. The brain tumor segmentation method based on uncertainty estimation according to claim 3, characterized in that: The structures of the four encoder modules are the same, which include a down-sampled initial block, two down-sampled convolutional normalized activation modules b, and two convolutional normalized activation discard blocks connected in sequence; The downsampling initial block consists of a standard convolutional normalized activation module and a downsampling convolutional normalized activation module a connected in sequence; Both convolution-normalized activation dropout blocks consist of a downsampled convolution-normalized activation module a and a dropout layer connected in sequence; The output of the last convolution normalization activation drop block is the output of the encoder module; The decoder module also includes a first up-sampling fusion convolution discard block, a second up-sampling fusion convolution discard block, a first up-sampling fusion convolution block, a second up-sampling fusion convolution block, a third up-sampling fusion convolution block, a first convolution block, a second convolution block, a third convolution block, a fourth convolution block, a fifth convolution block, a first up-sampling module b, a second up-sampling module b, a third up-sampling module b, a fourth up-sampling module b, a first element-by-element addition module, a second element-by-element addition module, a third element-by-element addition module, and a fourth element-by-element addition module; The multimodal fusion module b, the first upsampling fusion convolution discard block, the second upsampling fusion convolution discard block, the first upsampling fusion convolution block, the second upsampling fusion convolution block, the third upsampling fusion convolution block, the fifth convolution block, and the fourth element-by-element addition module are connected in sequence; The first upsampling fusion convolution discarding block, the first convolution block, the first upsampling module b, the first element-by-element addition module, the second upsampling module b, the second element-by-element addition module, the third upsampling module b, the third element-by-element addition module, the fourth upsampling module b, and the fourth element-by-element addition module are connected in sequence; The second upsampling fusion convolution discard block, the second convolution block, and the first element-by-element addition module are connected in sequence; The first upsampling fusion convolution block, the third convolution block, and the second element-by-element addition module are connected in sequence; The second upsampling fusion convolution block, the fourth convolution block, and the third element-by-element addition module are connected in sequence; The first upsampling fusion convolution discard block and the second upsampling fusion convolution discard block are both composed of an upsampling module a, a multimodal fusion module a, a standard convolution normalization activation module, and a discard layer connected in sequence; The first upsampling fusion convolution block, the second upsampling fusion convolution block and the third upsampling fusion convolution block are all composed of an upsampling module a, a multimodal fusion module a, and a standard convolution normalization activation module connected in sequence; The output of the fourth element-by-element addition module is the output of the decoder; The standard convolutional normalization activation module consists of a convolutional layer, an instance normalization layer, and a LeakyReLU activation function layer connected in sequence; The downsampling convolution normalization activation module a and the downsampling convolution normalization activation module b are composed of a convolution layer with a step size of 2, an instance normalization layer, and a LeakyReLU activation function layer connected in sequence; All sequential connections refer to connections from front to back according to the data flow direction.
5. The brain tumor segmentation method based on uncertainty estimation according to claim 1, characterized in that: Loss function for image segmentation model The calculation formula is as follows: ; ; ; ; ; In the formula, = =1, represents the number of voxels in the image, Representing voxels The predicted value of Representing voxels The true value of For 10 -5 , Indicates Monte Carlo sampling of voxels The predicted value of .
6. A brain tumor segmentation method based on uncertainty estimation according to any one of claims 1 to 5, characterized in that: The T1 image, T1c image, T2 image and Flair image of the same brain tumor are preprocessed before being input into the trained image segmentation model. The preprocessing includes grayscale normalization, pixel adjustment and tensor quantization.
7. The brain tumor segmentation method based on uncertainty estimation according to claim 6, characterized in that: The training steps of the image segmentation model are as follows: (a) Collection Brain tumor cases, ≥285, each brain tumor case had T1 images, T1c images, T2 images, and Flair images; (b) performing the above-mentioned preprocessing on the T1 image, T1c image, T2 image and Flair image; obtaining a true segmentation image of each brain tumor case, where the true segmentation image is a brain tumor MRI image with the above-mentioned three segmentation regions obtained by manual annotation; (c) Adoption The training and test sets were constructed using T1 images, T1c images, T2 images, Flair images, and true segmentation images corresponding to each brain tumor case. (d) Using the training set to train the image segmentation model, during training, the T1 image, T1c image, T2 image, and Flair image are used as the input of the image segmentation model, and the real segmentation image is used as the theoretical output of the image segmentation model. The weight parameters of the image segmentation model are continuously adjusted until the image segmentation model converges; (e) Use the test set to test the trained image segmentation model.
Citation Information
Patent Citations
Brain tumor segmentation quality evaluation method and device based on deep learning, and medium
CN114155195A
Polyp image segmentation system of uncertainty enhanced context attention network
CN116563536A
Cited By
A brain tumor segmentation method based on three-axis context encoding
CN121353673B