A lightweight brain tumor segmentation network based on CNN-Mamba parallel encoding

Through the combination of U-shaped network architecture and PMamba and MDA modules, the existing brain tumor segmentation network has been solved, and the effective brain tumor segmentation effect has been achieved.

CN120088278BActive Publication Date: 2025-07-22BRITE SEMICON SHANGHAI CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510526389.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-22
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The existing brain tumor segmentation network is expensive to calculate and has limited modeling capabilities for remote dependencies, making it difficult to effectively improve segmentation accuracy.

Method used

Using a U-shaped network architecture, combined with the PMamba module and the MDA module, through parallel encoding and multi-view decoupling of attention, the remote dependence relationship and local features of brain tumor images are enhanced, and the network complexity is reduced.

Benefits of technology

While maintaining low computational costs, the accuracy and robustness of brain tumor segmentation are significantly improved, especially in the segmentation effect of detailed areas such as the edges of brain tumors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088278B_ABST
    Figure CN120088278B_ABST
Patent Text Reader

Abstract

The present invention discloses a lightweight brain tumor segmentation network based on CNN-Mamba parallel encoding, belonging to the technical fields of neural networks and Mamba architecture. The lightweight network adopts a U-shaped network architecture, and the main body of the network consists of four layers of encoders, decoders, a bottleneck layer, and a skip connection part; among them, each layer of encoder and decoder is composed of a PMamba module and a corresponding sampling module, and the skip connection part is composed of an MDA module. The lightweight network efficiently extracts local features by using the PMamba module, enhances the network's ability to model the long-range dependence relationship between brain tumor images, and effectively improves the utilization efficiency of global features. At the same time, through the MDA module, from three view directions of 3D brain tumor images, the associated feature information of the same layer of the encoder and decoder is focused, and redundant background information is filtered. While maintaining a low computational cost, the network effectively improves the segmentation accuracy of each region of the brain tumor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of neural networks and Mamba architecture, and specifically relates to a lightweight brain tumor segmentation network based on CNN-Mamba parallel coding. Background Art

[0002] Brain tumors are tumor tissues formed by abnormal growth of intracranial brain cells, seriously threatening the life and health of patients. With the continuous development of medical imaging technology and computer technology, using deep learning methods to quickly and accurately separate brain tumors from normal tissues in multi-modal brain tumor MRI images is of great significance for the clinical treatment of patients. In the brain tumor segmentation task, networks based on CNN and Transformer have been widely used.

[0003] However, there are still the following disadvantages: (1) The disadvantage of high network complexity: Since most existing brain tumor segmentation datasets are 3D MRI images, CNN networks based on 3D convolution have expensive computational costs, and for networks based on Transformer, the computational cost will increase quadratically when processing large window or long sequence data. Therefore, for complex-structured networks, higher-performance hardware resources are required; (2) The disadvantage of limited ability of the network to model long-range dependence relationships: Due to the limited size of the convolutional kernel of CNN, there are obvious deficiencies in extracting long-range dependence relationships between long-distance features. However, capturing the long-range dependence relationships between long-distance voxels in brain tumor MRI images is crucial for improving the accuracy of brain tumor segmentation. Summary of the Invention

[0004] The purpose of the present invention is to provide a lightweight brain tumor segmentation network based on CNN-Mamba parallel coding. This network efficiently extracts local features using a parallel Mamba feature extraction module, and enhances the network's ability to model long-range dependence relationships between brain tumor images. While maintaining a low computational cost, it improves the utilization efficiency of global features and enhances the segmentation accuracy of brain tumors, helping to improve several aspects of existing brain tumor segmentation networks: reducing model complexity, enhancing the ability to model long-range dependence relationships between long-distance features, and improving the local feature extraction effect of the Mamba architecture.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A lightweight brain tumor segmentation network based on CNN-Mamba parallel coding, which adopts a U-shaped network architecture, and the main body of the network consists of an encoder, a decoder, a bottleneck layer, and a skip connection part;

[0006] Among them, the encoder and the decoder are divided into four layers in total. Each layer of the encoder consists of a PMamba module and a downsampling module, each layer of the decoder consists of a PMamba module and an upsampling module, the bottleneck layer consists of a PMamba module, and the skip connection part is input from each layer of the encoder and each layer of the decoder and connected to the MDA module, and the output of the MDA module is used as the input of the decoder of the same layer;

[0007] During operation, the MRI brain tumor images of four modalities are combined into a four-channel network input. The PMamba module captures local features and long-range dependencies in the brain tumor images, and the downsampling module uses max pooling with a stride of 2 to reduce the image size; in each layer of the encoder, the size of the brain tumor image is halved and the number of channels is doubled; while in each layer of the decoder, the PMamab module and trilinear interpolation upsampling with a multiple of 2 are used to restore the size and channels of the brain tumor image.

[0008] Preferably, the PMamba module is based on the Mamba architecture. By adding a path with multi-level dilated convolutions, while keeping the network lightweight, it enhances the ability of the Mamba architecture to extract local detailed features, and combines the state space model to capture long-range dependencies to extract global multi-scale features.

[0009] Preferably, when the PMamba module extracts global multi-scale features, it includes:

[0010] Separate the features of the upper layer by channels and send them into two branches of the PMamba module. One branch is a Mamba block with two branches, and the other branch is a multi-level dilated convolution with a residual structure, and the dilation rates are 1, 2, and 4 respectively;

[0011] In the first branch of the Mamba block, the input features are dimensionally expanded through a linear layer, and then after a one-dimensional convolutional layer, a depthwise separable convolutional layer, a SiLU activation function, an SSM layer, and a LayerNorm normalization layer, the output features are obtained; while in the second branch, the input features are passed through a linear layer and a SiLU activation function to obtain the output features, and then the outputs of the two branches are subjected to a Hadamard product operation and a linear layer and then output; finally, the output features of the two branches of the PMamba module are concatenated, and a learnable adjustment factor Scale is set at the residual connection to control the proportional residual input of the original feature map to realize the utilization of global feature information.

[0012] Preferably, after fusing the two input features, the MDA module enters three pseudo-3D convolutions with a dilation rate of 1 in parallel. While increasing the receptive field, it focuses on the detailed features from the coronal, sagittal, and horizontal view directions respectively. After passing through the BatchNorm normalization layer and the ReLU activation function, feature fusion is performed. Then, an attention map is generated using the sigmoid function. Finally, the two original input feature maps are multiplied element-wise with the generated attention map and fused with themselves through residual connections to achieve precise localization of the detailed features.

[0013] Preferably, the MDA module is also located in the skip connection part of the network, taking the outputs of the decoder and encoder at the same layer as inputs, extracting the correlation between the two, enhancing the intercommunication of multi-scale information and detailed information between the encoder and the decoder, and enhancing the robustness of the network.

[0014] Compared with the prior art, the beneficial effects of the present invention are:

[0015] In order to overcome the problems that the limited size of 3D convolution leads to insufficient ability of the network to capture long-range dependence relationships, the high computational cost of the Transformer architecture, and the insufficient extraction of local detailed features by the Mamba architecture, the present invention proposes a lightweight brain tumor segmentation network DP-Net based on CNN-Mamba parallel encoding. DP-Net uses a lightweight parallel Mamba feature extraction module (PMamba module) to reduce the network complexity while enhancing the network's ability to extract long-range dependence relationships between distant features of brain tumors and local features, realizing the utilization of global feature information. In addition, DP-Net focuses on the associated feature information of the same layer of the encoder and decoder from three view directions of 3D brain tumor images through a multi-view decoupled attention module (MDA module), strengthens the information intercommunication between the same layers, and filters redundant background information. Experimental results on the BraTS2020 dataset show that DP-Net uses 1.48M parameters and 54.85G floating-point operation counts, significantly improving the segmentation performance of brain tumor images while maintaining a small network complexity. Brief Description of the Drawings

[0016] Figure 1 It is a framework diagram of the DP-Net network.

[0017] Figure 2 It is a structural diagram of the PMamba module.

[0018] Figure 3 It is a structural diagram of the MDA module.

[0019] Figure 4Visualization images of ablation experiment results. Among them, (a) FLAIR modality; (b) predicted segmentation result of Baseline; (c) predicted segmentation result of Baseline+MDA; (d) predicted segmentation result of Baseline+PMamba; (e) predicted segmentation result of Baseline+MDA+PMamba (DP-Net); (f) ground truth label. Detailed implementation mode

[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0021] In order to overcome the problems that the limited size of 3D convolution leads to insufficient ability of the network to capture long-range dependence relationships, the high computational cost of the Transformer architecture, and the insufficient extraction of local detail features by the Mamba architecture, the present invention proposes a lightweight brain tumor segmentation network DP-Net based on CNN-Mamba parallel encoding.

[0022] The network structure of this DP-Net is as Figure 1 shown. The network DP-Net of the present invention adopts a U-shaped network architecture. The main body of the network consists of an encoder, a decoder, a bottleneck layer, and skip connections. Among them, the encoder and the decoder are divided into four layers in total. Each layer of the encoder is composed of a PMamba module and a downsampling module, and each layer of the decoder is composed of a PMamba module and an upsampling module. The bottleneck layer is composed of a PMamba module, and the skip connections are input from each layer of the encoder and each layer of the decoder and connected to the MDA module. The output of the MDA module is used as the input of the same layer of the decoder. First, MRI brain tumor images of four modalities are combined into a four-channel network input. The PMamba module captures local features and long-range dependence relationships in the brain tumor images, and the downsampling module uses max pooling with a stride of 2 to reduce the image size. In each layer of the encoder, the size of the brain tumor image is halved and the number of channels is doubled. In each layer of the decoder, the PMamab module and trilinear interpolation upsampling with a factor of 2 are used to restore the size and channels of the brain tumor image. The specific network parameters of DP-Net are shown in Table 1.

[0023] Table 1 shows the module names of each layer of the DP-Net network, as well as the sizes of the input images and output images.

[0024] Table 1 DP-Net network parameters

[0025]

[0026] The structure diagram of the PMamba module is as Figure 2 shown. By adding a path with multi-level dilated convolutions on the basis of the Mamba architecture, the PMamba module enhances the ability of the Mamba architecture to extract local detailed features while maintaining the network lightweight, and combines the State Space Model (SSM) to capture long-range dependencies and efficiently extract global multi-scale features. Specifically, first, the features from the upper layer are channel-separated and fed into two branches of the PMamba module respectively. One branch is a Mamba block with two branches, and the other branch is a multi-level dilated convolution with a residual structure, and the dilation rates are 1, 2, and 4 respectively. In the first branch of the Mamba block, the input features are dimension-expanded through a linear layer, and then the output features are obtained after passing through a one-dimensional convolutional layer, a depthwise separable convolutional layer, a SiLU activation function, an SSM layer, and a LayerNorm normalization layer. In the second branch, the input features are passed through a linear layer and a SiLU activation function to obtain the output features, and then the outputs of the two branches are subjected to Hadamard product operation and a linear layer and then output. Finally, the output features of the two branches of the PMamba module are concatenated, and a learnable adjustment factor Scale is set at the residual connection to control the residual input of the original feature map proportionally, further strengthening the effective utilization of global features.

[0027] The structure diagram of the MDA module is as Figure 3 shown. First, after the two input features are fused, they enter three pseudo-3D convolutions with a dilation rate of 1 in parallel, increasing the receptive field and focusing on the detailed features from the coronal, sagittal, and horizontal view directions respectively. After passing through the BatchNorm normalization layer and the ReLU activation function, feature fusion is performed. Then, an attention map is generated using the sigmoid function. Finally, the two original input feature maps are multiplied element-wise with the generated attention map respectively, and are fused with themselves through residual connections to achieve precise localization of detailed features such as the tumor edge.

[0028] The overall framework of the DP-Net network proposed by the present invention is as Figure 1 shown, including two important parts: (1) A lightweight parallel Mamba feature extraction module (Parallel Mamba Feature Extraction Module, PMamba) is proposed, and the structure of this module is as Figure 2As shown. Inspired by the visual state space model, on the basis of the Mamba architecture, the PMamba module adds dilated convolutional branch paths with different dilation rates and learnable adjustment factors, while maintaining the low computational cost of the Mamba architecture and enhancing the network's ability to model long-range dependencies, making up for the shortcoming of the Mamba architecture in extracting local detailed features and realizing the utilization of global feature information. (2) A Multiview Decoupled Attention Module (MDA) is proposed as the network skip connection part, and the specific structure of this module is as Figure 3 shown. This module decouples 3D convolution by combining three parallel pseudo-3D dilated convolutions. While reducing the network's computational complexity and parameter quantity and increasing the receptive field, it focuses on local detailed features from different views, filters redundant background information, and improves the segmentation accuracy of detail regions such as the brain tumor edge. In addition, this module is located in the network's skip connection part, takes the outputs of the decoder and encoder at the same layer as inputs, extracts the correlation between the two, enhances the multi-scale information and detail information exchange between the encoder and decoder, and enhances the network's robustness.

[0029] As an implementation manner of the present invention, the present invention selects to conduct experiments on the BraTS2020 dataset to verify the performance of the DP-Net. The training set of each dataset includes multiple cases, each case includes four-modal MRI brain tumor images and a ground truth label, while the validation set of each dataset does not include the ground truth label. The ground truth label is manually annotated by professional radiologists. Each brain tumor is divided into the following three parts, as Figure 4For example: in the regions of necrotic and non-enhancing tumors (NCR & NET), the label value is 1; in the regions of peritumoral edema (ED), the label value is 2; in the regions of enhancing tumors (ET), the label value is 4. The brain tumor segmentation task is to accurately segment the enhancing tumor (ET), which is composed of label 4; the tumor core (TC), which is composed of labels 1 and 4; and the whole tumor (WT), which is composed of labels 1, 2, and 4. To comprehensively and intuitively evaluate the performance of DP-Net, different evaluation metrics are selected for network complexity and segmentation accuracy. For segmentation accuracy, the Dice coefficient (Dice Similarity Coefficient, Dice) and the Hausdorff95 distance (Hausdorff95 Distance, HD) are selected as evaluation metrics. Among them, the Dice coefficient represents the degree of difference between the predicted segmentation result and the ground truth label, and the Hausdorff95 distance measures the distance between the predicted segmentation result and the ground truth label. For network complexity, the number of network parameters and the number of floating-point operations are selected to quantitatively analyze the network complexity, which are represented by the following formulas respectively: the number of network parameters , the number of floating-point operations . In the formulas, , , are the depth, height, and width of the convolutional kernel respectively; , are the number of input and output channels; , , are the depth, height, and width of the image respectively. The experiments of the present invention use the PyTorch deep learning framework. The server operating system is Ubuntu, the server CPU is an Intel Core i9-9900x processor, and the server GPU uses 4 parallel NVIDIA GTX2080Ti graphics cards with a video memory of 11GB. In the experiments of the present invention, the Adam optimizer is selected, the batch size is set to 8, the initial learning rate is set to 1x10 -3 , the weight decay coefficient is set to 1x10 -5 , and the maximum number of iterations is set to 1000.

[0030] In the preprocessing stage of the training set images, first, the MRI images with a size of 240×240×155 are cropped to 128×128×128 to adapt to the input image size of the network. In addition, images are randomly selected with a probability of 50%, flipped in different spatial axes, and randomly rotated within the range of -10° to 10°. Through the above data augmentation method, the amount of data is increased, the generalization of the network is improved, and the problem of overfitting that may occur during the training process is prevented.

[0031] The visualization results of the ablation experiment of the present invention are as Figure 4 shown. As can be seen from Figure (c), compared with the baseline network, by adding the MDA module, the network focuses more on the tumor boundary region, improving the segmentation effect of the tumor boundary shape. As can be seen from Figure (d), by adding the PMamba module, the network's ability to capture the long-range dependence relationship and local detail features between distant voxels in brain tumor images is improved, making the network more accurate in segmenting the sizes of the tumor core and edema regions. As can be seen from Figure (e), when both the MDA and PMamba modules are added, that is, the network DP-Net of the present invention. In the three views, compared with the Baseline, DP-Net reduces the problem of segmentation label errors, and the segmentation results of local detail regions such as the tumor boundary are closer to the ground truth labels, indicating the effectiveness of the modules proposed in the present invention.

[0032] Table 2 shows the ablation experiment results of DP-Net. From the results, it can be seen that taking the Residual U-Net as the baseline network Baseline, adding the MDA module of the present invention can help the network filter redundant background information. Although the network complexity increases slightly, the brain tumor segmentation accuracy is improved. When the basic module of Baseline is replaced with the lightweight module PMamba module of the present invention, the number of network parameters and the amount of computation are greatly reduced, the segmentation accuracy is significantly improved, and the Hausdroff95 distance is reduced, indicating that the introduction of the PMamba module improves the utilization efficiency of global features. When both the PMamba module and the MDA module are added, that is, the network DP-Net of the present invention, the overall performance of DP-Net reaches the best in terms of the Dice coefficient and the Hausdroff95 distance.

[0033] Table 2 Overall network ablation experiment on the BraTS2020 dataset

[0034]

[0035] Table 3 shows the comparison experiment results of DP-Net with other networks on the BraTS2020 dataset respectively. The experimental results show that DP-Net significantly improves the segmentation effect of brain tumor images while maintaining a low network complexity.

[0036] Table 3 Comparison results with different brain tumor segmentation networks on the BraTS2020 dataset

[0037]

[0038] The present invention discloses a lightweight brain tumor segmentation network DP-Net based on CNN-Mamba parallel encoding. The DP-Net uses a lightweight parallel Mamba feature extraction module to reduce the network complexity while enhancing the network's ability to extract long-range dependence relationships between distant features of brain tumors and local features, thereby realizing the utilization of global feature information. In addition, the DP-Net uses a multi-view decoupled attention module to focus on the associated feature information of the same layer between the encoder and the decoder from three view directions of the 3D brain tumor image, strengthening the information interaction between the same layers and filtering redundant background information. The experimental results on the BraTS2020 dataset show that the DP-Net uses 1.48M parameters and 54.85G floating-point operation counts, and significantly improves the segmentation performance of brain tumor images while maintaining a small network complexity.

[0039] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0040] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A lightweight brain tumor segmentation network based on CNN-Mamba parallel encoding, characterized in that, The lightweight brain tumor segmentation network adopts a U-shaped network architecture, and the main body of the network consists of an encoder, a decoder, a bottleneck layer, and a skip connection part; Among them, the encoder and the decoder are divided into four layers in total. Each layer of the encoder consists of a PMamba module and a downsampling module, each layer of the decoder consists of a PMamba module and an upsampling module, the bottleneck layer consists of a PMamba module, and the skip connection part is input by each layer of the encoder and each layer of the decoder and connected to the MDA module, and the output of the MDA module is used as the input of the same layer of the decoder; During operation, the MRI brain tumor images of four modalities are combined into a four-channel network input. The PMamba module captures the local features and long-range dependencies in the brain tumor images, and the downsampling module uses max pooling with a stride of 2 to reduce the image size; in each layer of the encoder, the size of the brain tumor image is halved and the number of channels is doubled; while in each layer of the decoder, the PMamab module and trilinear interpolation upsampling with a multiple of 2 are used to restore the size and channels of the brain tumor image; The PMamba module is based on the Mamba architecture. By adding a path with multi-level dilated convolutions, while keeping the network lightweight, it enhances the ability of the Mamba architecture to extract local detailed features, and combines a state space model to capture long-range dependencies to extract global multi-scale features; When the PMamba module extracts global multi-scale features, it includes: Separating the features of the upper layer by channels and sending them into two branches of the PMamba module respectively. One branch is a Mamba block with two branches, and the other branch is a multi-level dilated convolution with a residual structure, and the dilation rates are 1, 2, and 4 respectively; In the first branch of the Mamba block, the input features are dimensionally expanded through a linear layer, and then after a one-dimensional convolutional layer, a depthwise separable convolutional layer, a SiLU activation function, an SSM layer, and a LayerNorm normalization layer, the output features are obtained; while in the second branch, the input features are passed through a linear layer and a SiLU activation function to obtain the output features, and then the outputs of the two branches are subjected to Hadamard product operation and a linear layer and then output; finally, the output features of the two branches of the PMamba module are concatenated, and a learnable adjustment factor Scale is set at the residual connection to control the residual input of the original feature map proportionally to realize the utilization of global feature information; After fusing the two input features, the MDA module enters three pseudo-3D convolutions with a dilation rate of 1 in parallel. While increasing the receptive field, it focuses on the detailed features from the three view directions of coronal, sagittal, and horizontal respectively, and after passing through the BatchNorm normalization layer and the ReLU activation function, feature fusion is performed; then, an attention map is generated using the sigmoid function; finally, the two original input feature maps are multiplied element-wise with the generated attention map respectively, and fused with themselves through a residual connection to realize the precise localization of the detailed features.

2. The lightweight brain tumor segmentation network based on CNN-Mamba parallel encoding according to claim 1, wherein The MDA module is also located in the skip connection part of the network, taking the outputs of the decoder and encoder at the same layer as inputs, extracting the correlation between the two, and enhancing the intercommunication of multi-scale information and detailed information between the encoder and the decoder.

Citation Information

Patent Citations

  • Multi-context brain tumor segmentation system based on scale fusion guidance

    CN117237320A

  • CNN and attention-based MRI brain tumor segmentation method

    CN117764953A