Brain tumor segmentation method based on boundary enhancement
Through the segmentation network of the six-level encoder-decoder architecture, combined with the boundary extraction and guidance module, the accuracy and robustness of brain tumor segmentation in the existing technology are solved, and more efficient tumor boundary capture and segmentation effects are achieved.
Patent Information
- Application Number
- CN202510347638.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
AI Technical Summary
The existing brain tumor segmentation methods are insufficiently accurate and robust when dealing with complex tumor boundaries and heterogeneous tumors. Traditional methods are susceptible to noise and artifacts. Deep learning methods require a large amount of computing resources and rely on manual annotation.
A segmentation network with a six-level encoder-decoder architecture is adopted, combined with the boundary extraction module, the boundary guidance module and the cross feature fusion module, the tumor boundary features are extracted through multi-scale convolution and Sobel filter, and the boundary is automatically identified during the training process, and the segmentation effect is optimized using a mixed loss function.
It significantly improves the accuracy and robustness of brain tumor segmentation, reduces dependence on data preprocessing, improves the flexibility and applicability of the model, and can better adapt to complex morphological and heterogeneous tumors.
Smart Images

Figure CN120298440A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of brain tumor segmentation and relates to a brain tumor segmentation method based on boundary enhancement. Background Art
[0002] Accurate segmentation of brain tumors is crucial for clinical diagnosis and treatment planning. With the continuous progress of medical imaging technology, the detection and analysis of brain tumors have become increasingly dependent on image segmentation technology. The traditional segmentation methods mainly include the following three types:
[0003] (1) Methods based on traditional image processing: such as threshold segmentation, region growing and other methods. These techniques attempt to segment the tumor region by manually setting parameters, but when dealing with tumors with complex shapes and blurred boundaries, they have poor adaptability to complex tumor boundaries and are easily affected by noise and artifacts, resulting in unstable and inaccurate segmentation results.
[0004] (2) Deep learning segmentation networks: such as U-Net and FCN (Fully Convolutional Network). These networks use convolutional neural networks (CNNs) to segment tumors and have strong feature extraction capabilities, but when dealing with complex tumor boundaries and heterogeneous manifestations, problems such as mis-segmentation and missed-segmentation may occur.
[0005] (3) Ensemble learning techniques: improve the segmentation accuracy by combining the prediction results of multiple models. However, this type of method requires training multiple models simultaneously, often consuming a large amount of computing resources. The complexity and time cost of training multiple models are relatively high, and the integration between models may introduce new uncertainties.
[0006] For example, the literature (Edge U-Net: Brain tumor segmentation using MRI based on deep U-Net model with boundary information[J]. Expert Systems with Applications, 2023, 213: 118833) proposed a brain tumor segmentation model called Edge U-Net. By using the brain MRI boundary image and the brain MRI image together as input images to train the network, the tumor can be located. This method requires pre-acquiring the boundary image and relies on manual annotation and boundary extraction, which increases the complexity of data preprocessing and at the same time reduces the flexibility and application scope of the model.
[0007] The above traditional segmentation methods and existing deep learning algorithms often face problems of insufficient accuracy and poor robustness when dealing with complex tumor boundaries and heterogeneous tumors.
[0008] Therefore, it is of great significance to study a brain tumor segmentation method based on boundary enhancement to solve the problems existing in the prior art. Summary of the Invention
[0009] The purpose of the present invention is to solve the problems existing in the prior art and provide a brain tumor segmentation method based on boundary enhancement.
[0010] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0011] A brain tumor segmentation method based on boundary enhancement includes the following steps:
[0012] (1) Randomly select 80% from the collected brain tumor segmentation dataset as the training set, and the remaining 20% as the test set;
[0013] (2) Construct a segmentation network;
[0014] The segmentation network adopts a six-level encoder-decoder architecture, including an encoder, a boundary supervision module (BSM), a decoder, and a segmentation supervision module;
[0015] The encoder includes a boundary extraction module and a boundary guidance module, and the decoder includes a cross-feature fusion module;
[0016] The boundary extraction module is used to extract the rough boundary and fine edges of the tumor, enhancing the ability to capture complex tumor morphology and heterogeneity;
[0017] The boundary guidance module is used to generate an attention map by combining boundary features and modal features, improving the segmentation performance of modal features;
[0018] The cross-feature fusion module is used to fuse global context information and initial features, enhancing the feature expression ability and optimizing the segmentation result;
[0019] (3) Use the training set to train the constructed segmentation network to obtain a trained segmentation network;
[0020] Through iterative training of the deep neural network (i.e., the segmentation network of the present invention) using training data and a defined loss function until the model converges, a trained deep neural network model is obtained. Specifically, the process of iterative training of the deep neural network until convergence includes the following steps:
[0021] (Ⅰ) Adopt the forward propagation algorithm to extract features from the preprocessed MRI data pairs to obtain the segmentation result of the brain tumor;
[0022] (Ⅱ) Use the loss function to evaluate the difference between the segmentation result generated by the deep neural network and the true segmentation label, and quantify this difference into a specific loss value;
[0023] (Ⅲ) Update the gradients of the weight parameters in the deep neural network by combining the backpropagation algorithm with the chain rule of differentiation and the Adam optimization algorithm;
[0024] (Ⅳ) Use the updated weight parameters to repeat the iterative training of the deep neural network until the model converges;
[0025] (Ⅴ) During the process of updating the weight parameters, record the loss value and generate a change curve, and save the weight parameters by setting corresponding thresholds, so as to obtain the final deep neural network;
[0026] (4) Input the test set into the encoder of the trained segmentation network, and finally obtain the tumor segmentation results through the output of the segmentation supervision module. The output tumor segmentation results include whole tumor, tumor core, and enhanced tumor.
[0027] As a preferred technical solution:
[0028] For a brain tumor segmentation method based on boundary enhancement as described above, the encoder is divided into six levels, and the first to sixth level encoders are connected in sequence. The first to third level encoders respectively include a convolutional block, a boundary extraction module, and a boundary guidance module, and the fourth to sixth level encoders respectively include only a convolutional block;
[0029] The boundary supervision module includes four convolutional blocks, where three convolutional blocks are respectively connected in series with the boundary extraction modules in the first to third level encoders. After the three convolutional blocks are connected in sequence through upsampling, they are connected to the remaining one convolutional block;
[0030] The decoder is divided into six levels, and the channels of each level of the decoder are connected in series with the corresponding level of the encoder; The first to fifth level decoders respectively include a cross feature fusion module and a convolutional block, and the sixth level decoder includes only a cross feature fusion module;
[0031] The segmentation supervision module includes five convolutional blocks, and the five convolutional blocks are respectively connected to the first to fifth level decoders. The bottom convolutional block (i.e., the convolutional block connected to the fifth level encoder) is combined with the convolutional block in the upper layer through element-wise addition after upsampling, and the result after addition is continued to be combined with the convolutional block in the upper layer through element-wise addition after upsampling until the final segmentation result is output;
[0032] The boundary supervision module extracts the boundary information of the segmentation target, uses it as an additional supervision signal, and constrains the performance of the network in the boundary region through an additional boundary loss, so as to better guide the network to segment in the boundary region and improve the segmentation fineness; The segmentation supervision module fuses the segmentation results at different scales of each level as the final output, which can improve the robustness and segmentation accuracy of the model;
[0033] The Boundary Extraction Module (BEM) applies convolutional layers with three different dilation rates to the input features to capture multi-scale features, ensuring that the boundaries of tumors are widely detected; after convolution, global average pooling (GAP) is performed, which compresses each feature into a single value by averaging the spatial dimensions, effectively capturing the global context; the pooled features can distill high-level spatial information and provide a more abstract understanding of the tumor boundaries; subsequently, the output of global average pooling is subtracted from the respective convolutional outputs; next, the resulting features are concatenated along the channels and channel adjustment is performed through a convolutional layer; this stage focuses on coarse boundary detection, capturing the broad features of the tumor. To further obtain detailed boundaries, finally, the Sobel filter is used to highlight regions with significant intensity changes by calculating the gradients in the spatial dimensions, which are indicators of the boundaries, and this step is crucial for isolating the edges of the tumor from the surrounding tissues, thus achieving more precise segmentation; by combining multi-scale dilated convolutions and the Sobel filter, complex tumor boundaries and heterogeneous tumors can be processed more effectively. The multi-scale dilated convolution captures the rough tumor boundaries by increasing the receptive field, while the Sobel filter finely detects the tumor edges; the combination of the two can simultaneously extract the overall morphology and details of the tumor, improving the accuracy of tumor segmentation, especially in cases where the tumor morphology is complex and the internal heterogeneity is strong.
[0034] The input of the Boundary Guidance Module (BGM) includes the boundary features from the Boundary Extraction Module (BEM) and the modality features extracted from each encoder; first, the modality features and the boundary features are combined by element-wise multiplication and then the Softmax activation is applied to generate an attention map; at the same time, the modality features and the boundary features are combined additively and then the Sigmoid activation is applied to generate another attention map; then, the two attention maps obtained through the Softmax and Sigmoid operations respectively are combined to generate the final attention map; then the final attention map is multiplied by the modality features to generate enhanced attention features; finally, the enhanced attention features are combined with the modality features and the boundary features to generate the final feature representation, and the resulting feature representation combines both modality-specific information and boundary context, thus improving the segmentation performance in the subsequent layers of the network;
[0035] The Cross Feature Fusion Module (CFF) performs global average pooling (GAP) and global max pooling (GMP) on the initial input features (the modal features obtained by each encoder at six levels and the features upsampled by the decoder) respectively to capture the global context information across channels. The pooled features obtained after global average pooling and global max pooling are first processed by a multi-layer perceptron (MLP) layer respectively to learn non-linear transformations, so as to generate more expressive attention weights. Subsequently, the attention weights of the global average pooling features and the global max pooling features are generated through Sigmoid. Then, the attention weights of the global average pooling features and the global max pooling features are multiplied element-wise with the initial input features respectively to obtain two recalibrated input features. Finally, the initial input features are fused with the two recalibrated input features to obtain the fused features.
[0036] A brain tumor segmentation method based on boundary enhancement as described above. Each sample in the brain tumor segmentation dataset contains a brain tumor MRI image and the corresponding segmentation map label. The segmentation map label is the whole tumor segmentation image, the tumor core segmentation image, and the enhanced tumor segmentation image corresponding to the brain tumor MRI image.
[0037] A brain tumor segmentation method based on boundary enhancement as described above. The brain tumor MRI image includes four modalities: Flair, T1c, T2, and T1.
[0038] The number of encoders is equal to the number of modalities, that is, there are four encoders in total.
[0039] A brain tumor segmentation method based on boundary enhancement as described above. The first-level encoder includes a first convolutional block, a first boundary extraction module, and a first boundary guidance module connected in sequence, and the output of the first convolutional block is directly input to the first boundary guidance module at the same time. The second-level encoder includes a second convolutional block, a second boundary extraction module, and a second boundary guidance module, and the output of the second convolutional block is directly input to the second boundary guidance module at the same time. The third-level encoder includes a third convolutional block, a third boundary extraction module, and a third boundary guidance module, and the output of the third convolutional block is directly input to the third boundary guidance module at the same time. The fourth-level encoder includes a fourth convolutional block, the fifth-level encoder includes a fifth convolutional block, and the sixth-level encoder includes a sixth convolutional block. The first boundary guidance module is connected to the second convolutional block, the second boundary guidance module is connected to the third convolutional block, the third boundary guidance module is connected to the fourth convolutional block, and the fourth convolutional block, the fifth convolutional block, and the sixth convolutional block are connected in sequence.
[0040] A brain tumor segmentation method based on boundary enhancement as described above, where the boundary supervision module includes a seventh convolutional block, an eighth convolutional block, a ninth convolutional block, and a tenth convolutional block; after the third boundary extraction module obtains the boundary information of the modality, this boundary information is sent to the seventh convolutional block through a first concatenation operation; after the second boundary extraction module obtains the boundary information of the modality, this boundary information is sent to the eighth convolutional block through a second concatenation operation; after the first boundary extraction module obtains the boundary information of the modality, this boundary information is sent to the ninth convolutional block through a third concatenation operation; the features of the seventh convolutional block are upsampled to the eighth convolutional block, the features of the eighth convolutional block are upsampled to the ninth convolutional block, and the ninth convolutional block is connected to the tenth convolutional block.
[0041] A brain tumor segmentation method based on boundary enhancement as described above, where the first-level decoder includes a connected first cross-feature fusion module and an eleventh convolutional block, and the first cross-feature fusion module is connected to the first boundary guidance module; the second-level decoder includes a connected second cross-feature fusion module and a twelfth convolutional block, and the second cross-feature fusion module is connected to the second boundary guidance module; the third-level decoder includes a connected third cross-feature fusion module and a thirteenth convolutional block, and the third cross-feature fusion module is connected to the third boundary guidance module; the fourth-level decoder includes a connected fourth cross-feature fusion module and a fourteenth convolutional block, and the fourth cross-feature fusion module is connected to the fourth convolutional block; the fifth-level decoder includes a connected fifth cross-feature fusion module and a fifteenth convolutional block, and the fifth cross-feature fusion module is connected to the fifth convolutional block; the sixth decoder includes a sixth cross-feature fusion module, and the sixth cross-feature fusion module is connected to the sixth convolutional block; the features of the sixth cross-feature fusion module are upsampled to the fifth cross-feature fusion module, the features of the fifteenth convolutional block are upsampled to the fourth cross-feature fusion module, the features of the fourteenth convolutional block are upsampled to the third cross-feature fusion module, the features of the thirteenth convolutional block are upsampled to the second cross-feature fusion module, and the features of the twelfth convolutional block are upsampled to the first cross-feature fusion module.
[0042] A brain tumor segmentation method based on boundary enhancement as described above, where the segmentation supervision module includes a sixteenth convolutional block, a seventeenth convolutional block, an eighteenth convolutional block, a nineteenth convolutional block, and a twentieth convolutional block. The sixteenth convolutional block is connected to the eleventh convolutional block, the seventeenth convolutional block is connected to the twelfth convolutional block, the eighteenth convolutional block is connected to the thirteenth convolutional block, the nineteenth convolutional block is connected to the fourteenth convolutional block, and the twentieth convolutional block is connected to the fifteenth convolutional block. The features of the twentieth convolutional block are upsampled and combined with the nineteenth convolutional block through element-wise addition, and the result after addition is upsampled and combined with the eighteenth convolutional block through element-wise addition. Then, the result after addition is upsampled and combined with the seventeenth convolutional block through element-wise addition, and the result after addition is further upsampled and combined with the sixteenth convolutional block through element-wise addition to output the final segmentation result.
[0043] A brain tumor segmentation method based on boundary enhancement as described above with three convolutional layers having different dilation rates of 1, 2, and 3 in sequence.
[0044] The three different dilation rates (1, 2, 3) play different roles in the convolutional operation, mainly obtaining multi-level feature information by adjusting the receptive field of the convolutional kernel. Specifically:
[0045] 1. Dilation rate of 1: This is a standard convolution with no dilation effect. The receptive field at this time is the size of the convolutional kernel. For example, for a 3×3×3 convolutional kernel, the receptive field is also 3×3×3. It is mainly used to capture detailed local features and is suitable for dealing with the correlation between adjacent pixels.
[0046] 2. Dilation rate of 2: When the dilation rate is 2, it is equivalent to inserting a pixel interval between the elements in the convolutional kernel. Therefore, the receptive field of the convolutional kernel doubles. For example, the equivalent receptive field of a 3×3×3 convolutional kernel with a dilation rate of 2 is 5×5×5. This setting can capture a wider range of features than the standard convolution while retaining some details.
[0047] 3. Dilation rate of 3: When the dilation rate is 3, two pixel intervals are inserted between each convolutional kernel element. Therefore, the equivalent receptive field of the convolutional kernel further increases. For example, the equivalent receptive field of a 3×3×3 convolutional kernel with a dilation rate of 3 is 7×7×7. This configuration is suitable for capturing more global features and helps the network extract features in a larger range.
[0048] Quantitative representation: In the convolutional operation, the relationship between the size of the receptive field and the size of the convolutional kernel and the dilation rate d can be expressed as:
[0049] Equivalent receptive field = (k - 1)×d + 1;
[0050] Where k is the size of the convolutional kernel and d is the dilation rate.
[0051] Taking a 3×3×3 convolutional kernel as an example:
[0052] Dilation rate 1: Equivalent receptive field = (3 - 1)×1 + 1 = 3;
[0053] Dilation rate 2: Equivalent receptive field = (3 - 1)×2 + 1 = 5;
[0054] Dilation rate 3: Equivalent receptive field = (3 - 1)×3 + 1 = 7;
[0055] Therefore, these three dilation rates can help the model extract features at different scales, gradually from detailed information to global features, forming a multi-scale feature representation and improving the accuracy of the segmentation result.
[0056] A brain tumor segmentation method based on boundary enhancement as described above uses a hybrid loss function for network training. The hybrid loss function is obtained by combining the Dice loss function and the boundary loss function, and the formula is as follows:
[0057] Loss = L seg + λ·L boundary ;
[0058]
[0059] where Loss represents the hybrid loss function, L seg represents the Dice loss function, and L bounary represents the boundary loss function. λ = 0.1, p i and g i respectively represent the predicted probability and the true value of voxel i, y i represents the true result, represents the predicted result, and N is the total number of voxels in the image.
[0060] Advantages:
[0061] (1) The brain tumor segmentation method based on boundary enhancement of the present invention can achieve multi-level boundary extraction, while the existing methods usually rely on a single boundary extraction technology and are easily affected by noise and boundary blurring. By combining a coarse boundary extraction module and a fine boundary extraction module, the present invention can capture tumor boundaries more effectively and adapt to complex morphological changes.
[0062] (2) Most traditional methods fail to fully utilize tumor boundary information. In the present invention, the boundary feature is fused with the modality feature through the boundary guidance module (BGM), and the boundary guidance mechanism is integrated, which strengthens the network's attention to the tumor boundary and significantly improves the segmentation accuracy.
[0063] (3) The brain tumor segmentation method based on boundary enhancement of the present invention does not rely on obtaining boundary images in advance, but automatically identifies boundaries during network training, avoiding the dependence on manual annotation and boundary extraction, thereby reducing the need for data preprocessing and improving the flexibility and applicability of the model.
[0064] (4) The brain tumor segmentation method based on boundary enhancement of the present invention can effectively capture and utilize tumor boundary information by introducing a boundary extraction module and a boundary guidance module, thereby significantly improving the segmentation accuracy; it has significant advantages in terms of accuracy and clinical applicability, providing an innovative solution for the diagnosis and treatment of brain tumors. Description of the Drawings
[0065] Figure 1For the overall framework of the segmentation network; among them, the gray and blue are convolutional blocks; the orange is the boundary extraction module; the green is the boundary guidance module; the yellow is the cross-feature fusion module; the purple is the segmentation supervision module; the red arrow is upsampling; the yellow arrow is concatenation operation;
[0066] Figure 2 For the structural diagram of the boundary extraction module, where F in is the input feature, DConv is the dilated convolution, and Conv is the standard convolution;
[0067] Figure 3 For the structural diagram of the boundary guidance module, where F m is the modal feature, F b is the boundary feature, F out is the output feature;
[0068] Figure 4 For the cross-feature fusion module, where F in is the input feature, F out is the output feature. Specific implementation manners
[0069] The present invention will be further described below in conjunction with specific implementation manners. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0070] A brain tumor segmentation method based on boundary enhancement, the specific steps are as follows:
[0071] (1) Randomly select 80% from the collected BraTS2018 brain tumor segmentation dataset as the training set, and the remaining 20% as the test set; the BraTS2018 brain tumor segmentation dataset contains 285 cases;
[0072] Each sample in the brain tumor segmentation dataset contains a brain tumor MRI image and the corresponding segmentation map label. The segmentation map label is the whole tumor segmentation image, tumor core segmentation image, and enhanced tumor segmentation image corresponding to the brain tumor MRI image. The brain tumor MRI image includes four modalities: Flair, T1c, T2, and T1. The original size of the MRI image is 255×255×155. Each MRI image is subjected to gray normalization, and the image gray value is adjusted to a normal distribution with a mean of 0 and a standard deviation of 1. The training image is adjusted to 128×128×128 pixels, and the normalized image is converted into a tensor, and then input into the following constructed segmentation network;
[0073] (2) Construct a segmentation network;
[0074] As shown Figure 1 in the figure, the segmentation network adopts a six-level encoder-decoder architecture, including an encoder, a boundary supervision module (BSM), a decoder, and a segmentation supervision module;
[0075] The encoder is divided into six levels. The first-level encoder includes a first convolutional block, a first boundary extraction module, and a first boundary guidance module connected in sequence, and the output of the first convolutional block is directly input to the first boundary guidance module at the same time; the second-level encoder includes a second convolutional block, a second boundary extraction module, and a second boundary guidance module, and the output of the second convolutional block is directly input to the second boundary guidance module at the same time; the third-level encoder includes a third convolutional block, a third boundary extraction module, and a third boundary guidance module, and the output of the third convolutional block is directly input to the third boundary guidance module at the same time; the fourth-level encoder includes a fourth convolutional block, the fifth-level encoder includes a fifth convolutional block, and the sixth-level encoder includes a sixth convolutional block; the first boundary guidance module is connected to the second convolutional block, the second boundary guidance module is connected to the third convolutional block, the third boundary guidance module is connected to the fourth convolutional block, and the fourth convolutional block, the fifth convolutional block, and the sixth convolutional block are connected in sequence;
[0076] The boundary supervision module includes a seventh convolutional block, an eighth convolutional block, a ninth convolutional block, and a tenth convolutional block; after the third boundary extraction module obtains the boundary information of the modality, the boundary information is sent to the seventh convolutional block through a first concatenation operation; after the second boundary extraction module obtains the boundary information of the modality, the boundary information is sent to the eighth convolutional block through a second concatenation operation; after the first boundary extraction module obtains the boundary information of the modality, the boundary information is sent to the ninth convolutional block through a third concatenation operation; the features of the seventh convolutional block are upsampled to the eighth convolutional block, the features of the eighth convolutional block are upsampled to the ninth convolutional block, and the ninth convolutional block is connected to the tenth convolutional block;
[0077] The decoder is divided into six levels. The first-level decoder includes a connected first cross-feature fusion module and an eleventh convolutional block, and the first cross-feature fusion module is connected to the first boundary guidance module; the second-level decoder includes a connected second cross-feature fusion module and a twelfth convolutional block, and the second cross-feature fusion module is connected to the second boundary guidance module; the third-level decoder includes a connected third cross-feature fusion module and a thirteenth convolutional block, and the third cross-feature fusion module is connected to the third boundary guidance module; the fourth-level decoder includes a connected fourth cross-feature fusion module and a fourteenth convolutional block, and the fourth cross-feature fusion module is connected to the fourth convolutional block; the fifth-level decoder includes a connected fifth cross-feature fusion module and a fifteenth convolutional block, and the fifth cross-feature fusion module is connected to the fifth convolutional block; the sixth decoder includes a sixth cross-feature fusion module, and the sixth cross-feature fusion module is connected to the sixth convolutional block; the sixth cross-feature fusion module is upsampled to the fifth cross-feature fusion module, the fifteenth convolutional block is upsampled to the fourth cross-feature fusion module, the fourteenth convolutional block is upsampled to the third cross-feature fusion module, the thirteenth convolutional block is upsampled to the second cross-feature fusion module, and the twelfth convolutional block is upsampled to the first cross-feature fusion module;
[0078] The segmentation supervision module includes a sixteenth convolutional block, a seventeenth convolutional block, an eighteenth convolutional block, a nineteenth convolutional block, and a twentieth convolutional block. The sixteenth convolutional block is connected to the eleventh convolutional block, the seventeenth convolutional block is connected to the twelfth convolutional block, the eighteenth convolutional block is connected to the thirteenth convolutional block, the nineteenth convolutional block is connected to the fourteenth convolutional block, and the twentieth convolutional block is connected to the fifteenth convolutional block. The output of the twentieth convolutional block is upsampled and combined with the nineteenth convolutional block through element-wise addition. The result after addition is upsampled and combined with the eighteenth convolutional block through element-wise addition. Then, the result after addition is upsampled and combined with the seventeenth convolutional block through element-wise addition. The result after addition is further upsampled and combined with the sixteenth convolutional block through element-wise addition to output the final segmentation result;
[0079] The boundary extraction module is used to extract the rough boundary and fine edge of the tumor, enhancing the ability to capture complex tumor morphology and heterogeneity; as Figure 2 shown, the boundary extraction module (BEM) applies three convolutional layers with dilation rates of 1, 2, and 3 to the input features to capture multi-scale features; after convolution, global average pooling (GAP) is performed, and each feature is compressed into a single value by averaging over the spatial dimension; subsequently, the output of global average pooling is subtracted from the respective convolutional output; next, the resulting features are concatenated along the channel dimension and the channel is adjusted through a convolutional layer; finally, the Sobel filter is used to highlight the regions with obvious intensity changes by calculating the gradient in the spatial dimension, thereby achieving more accurate segmentation;
[0080] The boundary guidance module is used to generate an attention map by combining boundary features and modality features, improving the segmentation performance of modality features; as Figure 3 shown, the input of the boundary guidance module (BGM) includes boundary features from the boundary extraction module (BEM) and modality features extracted from each encoder; first, the modality features and boundary features are combined by element-wise multiplication, and then the Softmax activation is applied to generate an attention map; at the same time, the modality features and boundary features are combined by addition, and then the Sigmoid activation is applied to generate another attention map; then, the two attention maps obtained by the Softmax and Sigmoid operations respectively are combined to generate the final attention map; then the final attention map is multiplied by the modality features to generate enhanced attention features; finally, the enhanced attention features are combined with the modality features and boundary features to generate the final feature representation;
[0081] The cross-feature fusion module is used to fuse global context information and initial features, enhancing the feature expression ability and optimizing the segmentation result; as Figure 4 shown, the cross-feature fusion module (CFF) performs global average pooling (GAP) and global max pooling (GMP) on the initial input features respectively to capture the global context information across channels; the pooled features obtained after global average pooling and global max pooling are first processed by multi-layer perceptron (MLP) layers respectively to learn non-linear transformations, so as to generate more expressive attention weights, and then the Sigmoid is used to generate the attention weights of the global average pooling feature and the global max pooling feature; then the attention weights of the global average pooling feature and the global max pooling feature are multiplied by the initial input features element-wise to obtain two recalibrated input features; finally, the initial input features are fused with the two recalibrated input features to obtain the fused features; this fusion integrates the information of average pooling and max pooling with the original features, creating a more informative output for subsequent tasks. By combining cross-feature attention, the CFF module ensures that the most relevant features are highlighted, thus improving the discriminative ability of the features.
[0082] For simplicity, Figure 1 only the detailed part of one encoding is shown. In each encoder, the input image (e.g., Figure 1 Modality1 with size: 8×128×128×128) extracts modality features through convolutional layers, instance normalization, and the LeakyReLU activation function. After feature extraction at each level, the downsampling layer (i.e., Figure 1The convolution block with a medium step size of 2 is used to reduce the spatial dimension. In the first three levels, a boundary extraction module is introduced. Initially, the Boundary Extraction Module (BEM) is used to extract boundary features. These features are then combined with the main modality features and input into the Boundary Guidance Module (BGM), which refines the modality features to guide the network to focus on the tumor boundary. In the decoder, the features are first upsampled to restore their spatial resolution and then concatenated with the corresponding encoder-level features. Next, the Cross Feature Fusion Module (CFF) is introduced to integrate the modality features to enhance the segmentation effect. Finally, a convolutional layer is applied to adjust the number of feature channels. In addition, a segmentation supervision module is used to generate predicted segmentation results at each stage, and these results are then combined to obtain the final segmentation result. In addition, a Boundary Supervision Module (BSM) is used to generate predicted boundary results.
[0083] (3) The constructed segmentation network is trained using the training set to obtain a trained segmentation network;
[0084] The network is trained using a hybrid loss function, which is obtained by combining the Dice loss function and the boundary loss function. The formula is as follows:
[0085] Loss = L seg + λ·L boundary ;
[0086]
[0087] where Loss represents the hybrid loss function, L seg represents the Dice loss function, L bounary represents the boundary loss function, λ = 0.1, p i and g i respectively represent the predicted probability and the true value of voxel i, y i represents the true result, represents the predicted result, and N is the total number of voxels in the image;
[0088] (4) The test set is input into the encoder of the trained segmentation network, and finally, the tumor segmentation result (i.e., the segmentation image in Figure 1 ) is output through the segmentation supervision module. The output tumor segmentation results include whole tumor, tumor core, and enhanced tumor.
[0089] In this embodiment, the dataset of the Brain Tumor Segmentation Challenge (BraTS2018) is used in the experiment. This dataset is widely used in brain tumor segmentation tasks and contains 285 cases. Each case has four different imaging modalities, and the goal is to segment three tumor regions.
[0090] Table 1
[0091]
[0092] Table 1 shows the comparison results of ablation experiments conducted on the BraTS2018 dataset, with the evaluation metrics being the Dice similarity coefficient (DSC) and Hausdorff distance (HD) for the whole tumor (WT), tumor core (TC), and enhanced tumor (ET). "Baseline method" in Table 1 refers to the case where the network structure proposed in the present invention is not added. A tick (√) in Table 1 represents the addition of the strategy, and a cross (×) represents the non - inclusion of the strategy. The best experimental results are marked in bold in Table 1. It can be seen from the experimental results that: on the basic model, that is, without adding BEM, BGM, BSM, and CFF, the average DSC is 82.1% and the average HD is 7.1; on the Baseline, with the addition of BEM, the average DSC is 83% and the average HD is 4.3; on the Baseline, with the addition of BEM and BGM, the average DSC is 83.4% and the average HD is 4.4; on the Baseline, with the addition of BEM, BGM, and BSM, the average DSC is 83.8% and the average HD is 3.7; on the Baseline, when the present invention adds BEM, BGM, BSM, and CFF, the average DSC is 84.1% and the average HD is 3.8, which can effectively improve the accuracy of the segmentation model. By comparison, it can be seen that the addition of BEM, BGM, BSM, and CFF can effectively improve the accuracy of the segmentation model, and the segmentation accuracy and robustness of the model have been significantly improved, making it suitable for brain tumor segmentation tasks with high accuracy requirements.
[0093] Table 2
[0094]
[0095] Table 2 shows the comparison results between the present invention and the prior art, and it can be seen that the present algorithm can achieve better segmentation accuracy.
[0096] The specific information of the prior art documents [1] - [7] is as follows:
[0097] [1]Analyzing the quality and challenges of uncertainty estimations for brain tumor segmentation, Frontiers in neuroscience 14(2020)501743.
[0098] [2]Latent correlation representation learning for brain tumor segmentation with missing MRI modalities, IEEE Transactions on Image Processing 30(2021)4263–4274.
[0099] [3]Attention gate resu-net for automatic MRI brain tumor segmentation, IEEE Access 8(2020)58533–58545.
[0100] [4]Httu-net: Hybrid two track u-net for automatic brain tumor segmentation, IEEE Access 8(2020)101406–101415.
[0101] [5]A tri-attention fusion guided multimodal segmentation network, Pattern Recognition 124(2022)108417.
[0102] [6]Raagr2-net: A brain tumor segmentation network using parallel processing of multiple spatial frames, Computers in Biology and Medicine 152(2023)106426.
[0103] [7]Multi-modal graph convolution network for precise brain tumor segmentation across multiple MRI sequences, IEEE Transactions on Image Processing(2024).
Claims
1. A brain tumor segmentation method based on boundary enhancement, characterized in that It includes the following steps: (1) Randomly select 80% of the collected brain tumor segmentation dataset as the training set, and the remaining 20% as the test set; (2) Construct a segmentation network; The segmentation network adopts a six-level encoder-decoder architecture, including an encoder, a boundary supervision module, a decoder, and a segmentation supervision module; The encoder includes a boundary extraction module and a boundary guidance module, and the decoder includes a cross-feature fusion module; The boundary extraction module is used to extract the rough boundary and fine edges of the tumor, enhancing the ability to capture complex tumor morphology and heterogeneity; The boundary guidance module is used to generate an attention map by combining boundary features and modality features, improving the segmentation performance of modality features; The cross-feature fusion module is used to fuse global context information and initial features, enhancing the feature expression ability and optimizing the segmentation result; (3) Use the training set to train the constructed segmentation network to obtain a trained segmentation network; (4) Input the test set into the encoder of the trained segmentation network, and finally output the tumor segmentation result through the segmentation supervision module. The output tumor segmentation results include whole tumor, tumor core, and enhanced tumor.
2. The method for brain tumor segmentation based on boundary enhancement according to claim 1, wherein The encoder is divided into six levels, and the first to sixth level encoders are connected in sequence. The first to third level encoders each include a convolutional block, a boundary extraction module, and a boundary guidance module, and the fourth to sixth level encoders each only include a convolutional block; The boundary supervision module includes four convolutional blocks, where three convolutional blocks are respectively connected in series with the boundary extraction modules in the first to third level encoders. After the three convolutional blocks are connected in sequence through upsampling, they are connected to the remaining one convolutional block; The decoder is divided into six levels, and each level of the decoder is connected in series with the channels of the corresponding level of the encoder. The first to fifth level decoders each include a cross-feature fusion module and a convolutional block, and the sixth level decoder only includes a cross-feature fusion module; The segmentation supervision module includes five convolutional blocks, and the five convolutional blocks are respectively connected to the first to fifth level decoders. The bottom convolutional block is combined with the upper convolutional block through element-wise addition after upsampling, and the result after addition is continuously combined with the convolutional block of the upper layer through element-wise addition after upsampling until the final segmentation result is output; The boundary extraction module applies convolutional layers with three different dilation rates to the input features to capture multi-scale features; after convolution, global average pooling is performed, and each feature is compressed into a single value by averaging the spatial dimension; subsequently, the output of global average pooling is subtracted from the respective convolutional output; next, the obtained features are concatenated in channels and the channels are adjusted through a convolutional layer; finally, the Sobel filter is used to highlight the regions with obvious intensity changes by calculating the gradient in the spatial dimension, thereby achieving more accurate segmentation; The input of the boundary guidance module includes the boundary features from the boundary extraction module and the modality features extracted from each encoder; first, the modality features and the boundary features are combined by element-wise multiplication, and then the Softmax activation is applied to generate an attention map; meanwhile, the modality features and the boundary features are combined by addition, and then the Sigmoid activation is applied to generate another attention map; then, the two attention maps obtained through the Softmax and Sigmoid operations respectively are combined to generate the final attention map; then the final attention map is multiplied by the modality features to generate enhanced attention features; finally, the enhanced attention features are combined with the modality features and the boundary features to generate the final feature representation; The cross-feature fusion module performs global average pooling and global max pooling on the initial input features respectively to capture the global context information across channels; the pooling features obtained after global average pooling and global max pooling are first processed through multi-layer perceptron layers respectively to learn non-linear transformations, so as to generate more expressive attention weights, and then the Sigmoid is used to generate the attention weights of the global average pooling feature and the global max pooling feature; Then the attention weights of the global average pooling feature and the global max pooling feature are respectively multiplied by the initial input features element-wise to obtain two recalibrated input features; finally, the initial input features are fused with the two recalibrated input features to obtain the fused features.
3. A brain tumor segmentation method based on boundary enhancement according to claim 1, characterized in that, Each sample in the brain tumor segmentation dataset contains a brain tumor MRI image and the corresponding segmentation map label, and the segmentation map label is the whole tumor segmentation image, tumor core segmentation image, and enhanced tumor segmentation image corresponding to the brain tumor MRI image.
4. A brain tumor segmentation method based on boundary enhancement according to claim 3, characterized in that The brain tumor MRI images include four modalities: Flair, T1c, T2, and T1; The number of encoders is equal to the number of modalities.
5. A method for brain tumor segmentation based on boundary enhancement according to claim 2, characterized in that, The first-level encoder includes a first convolutional block, a first boundary extraction module, and a first boundary guidance module connected in sequence, and the output of the first convolutional block is directly input to the first boundary guidance module at the same time; the second-level encoder includes a second convolutional block, a second boundary extraction module, and a second boundary guidance module, and the output of the second convolutional block is directly input to the second boundary guidance module at the same time; the third-level encoder includes a third convolutional block, a third boundary extraction module, and a third boundary guidance module, and the output of the third convolutional block is directly input to the third boundary guidance module at the same time; the fourth-level encoder includes a fourth convolutional block, the fifth-level encoder includes a fifth convolutional block, and the sixth-level encoder includes a sixth convolutional block; the first boundary guidance module is connected to the second convolutional block, the second boundary guidance module is connected to the third convolutional block, the third boundary guidance module is connected to the fourth convolutional block, and the fourth convolutional block, the fifth convolutional block, and the sixth convolutional block are connected in sequence.
6. The method for segmenting brain tumors based on boundary enhancement according to claim 5, characterized in that The boundary supervision module includes a seventh convolutional block, an eighth convolutional block, a ninth convolutional block, and a tenth convolutional block; after the third boundary extraction module obtains the boundary information of the modality, the boundary information is sent to the seventh convolutional block through a first concatenation operation; after the second boundary extraction module obtains the boundary information of the modality, the boundary information is sent to the eighth convolutional block through a second concatenation operation; after the first boundary extraction module obtains the boundary information of the modality, the boundary information is sent to the ninth convolutional block through a third concatenation operation; the features of the seventh convolutional block are upsampled to the eighth convolutional block, the features of the eighth convolutional block are upsampled to the ninth convolutional block, and the ninth convolutional block is connected to the tenth convolutional block.
7. A method for brain tumor segmentation based on boundary enhancement according to claim 6, characterized in that, The first-level decoder includes a connected first cross-feature fusion module and an eleventh convolutional block, and the first cross-feature fusion module is connected to the first boundary guidance module; the second-level decoder includes a connected second cross-feature fusion module and a twelfth convolutional block, and the second cross-feature fusion module is connected to the second boundary guidance module; the third-level decoder includes a connected third cross-feature fusion module and a thirteenth convolutional block, and the third cross-feature fusion module is connected to the third boundary guidance module; the fourth-level decoder includes a connected fourth cross-feature fusion module and a fourteenth convolutional block, and the fourth cross-feature fusion module is connected to the fourth convolutional block; the fifth-level decoder includes a connected fifth cross-feature fusion module and a fifteenth convolutional block, and the fifth cross-feature fusion module is connected to the fifth convolutional block; the sixth decoder includes a sixth cross-feature fusion module, and the sixth cross-feature fusion module is connected to the sixth convolutional block; the sixth cross-feature fusion module is upsampled to the fifth cross-feature fusion module, the fifteenth convolutional block is upsampled to the fourth cross-feature fusion module, the fourteenth convolutional block is upsampled to the third cross-feature fusion module, the thirteenth convolutional block is upsampled to the second cross-feature fusion module, and the twelfth convolutional block is upsampled to the first cross-feature fusion module.
8. A brain tumor segmentation method based on boundary enhancement according to claim 7, characterized in that, The segmentation supervision module includes a sixteenth convolutional block, a seventeenth convolutional block, an eighteenth convolutional block, a nineteenth convolutional block, and a twentieth convolutional block. The sixteenth convolutional block is connected to the eleventh convolutional block, the seventeenth convolutional block is connected to the twelfth convolutional block, the eighteenth convolutional block is connected to the thirteenth convolutional block, the nineteenth convolutional block is connected to the fourteenth convolutional block, and the twentieth convolutional block is connected to the fifteenth convolutional block. The twentieth convolutional block is upsampled and combined with the nineteenth convolutional block through element-wise addition, and the resulting sum is upsampled and combined with the eighteenth convolutional block through element-wise addition. Then, the resulting sum is upsampled and combined with the seventeenth convolutional block through element-wise addition, and the resulting sum is further upsampled and combined with the sixteenth convolutional block through element-wise addition to output the final segmentation result.
9. The method for segmenting brain tumors based on boundary enhancement according to claim 2, wherein, Three convolutional layers with different dilation rates, and the dilation rates are 1, 2, and 3 in sequence.
10. The method for brain tumor segmentation based on boundary enhancement according to claim 1, wherein, The network is trained using a hybrid loss function, which is obtained by combining the Dice loss function and the boundary loss function. The formula is as follows: Loss=L seg +λ·L boundary ; Among them, Loss represents the hybrid loss function, L seg represents the Dice loss function, L bounary represents the boundary loss function, λ = 0.1, p i and g i respectively represent the predicted probability and the true value of voxel i, y i represents the true result, represents the predicted result, and N is the total number of voxels in the image.
Citation Information
Cited By
Brain tumor segmentation method based on boundary perception mechanism
CN120976222A