Lightweight brain tumor segmentation network based on CNN-Mama parallel coding

By introducing parallel Mamba feature extraction module and multi-view decoupling attention module in the brain tumor segmentation network, the problems of high computing costs and insufficient remote dependency capture capabilities in the prior art are solved, and efficient brain tumor image segmentation is achieved.

CN120088278AActive Publication Date: 2025-06-03BRITE SEMICON SHANGHAI CORP

Patent Information

Application Number
CN202510526389.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-06-03
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The existing brain tumor segmentation network has high computational cost when processing 3D MRI images, making it difficult to effectively capture remote dependencies in brain tumor images, and the Mamba architecture does not extract local detail features sufficiently.

Method used

The lightweight brain tumor segmentation network DP-Net based on CNN-Mamba parallel encoding is adopted to improve the network's modeling ability of remote dependencies and local features through the parallel Mamba feature extraction module (PMamba module) and the multi-view decoupling attention module (MDA module).

Benefits of technology

While maintaining low computational costs, the segmentation accuracy of brain tumor images is significantly improved, the utilization of global feature information is enhanced, and the network complexity is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088278A_ABST
    Figure CN120088278A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight brain tumor segmentation network based on CNN-Mama parallel coding, and belongs to the technical field of neural networks and Mama architectures. The lightweight network adopts a U-shaped network architecture, and a network main body consists of four layers of encoders and decoders, a bottleneck layer and a jump connection part; wherein each layer of encoder and decoder is composed of a PMamba module and a corresponding sampling module, and the jump connection part is composed of an MDA module. According to the lightweight network, local features are efficiently extracted by using a PMamba module, the modeling capability of the network on a remote dependency relationship between brain tumor images is enhanced, and the utilization efficiency of global features is effectively improved. Meanwhile, redundant background information is filtered by focusing associated feature information of the same layer of the encoder and the decoder from three view directions of the 3D brain tumor image through an MDA module. According to the network, the segmentation precision of each region of the brain tumor is effectively improved while relatively low calculation cost is maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of neural networks and Mamba architecture, and specifically relates to a lightweight brain tumor segmentation network based on CNN-Mamba parallel encoding. Background Art

[0002] Brain tumors are tumor tissues formed by abnormal growth of intracranial brain cells, which seriously threaten the life and health of patients. With the continuous development of medical imaging technology and computer technology, using deep learning methods to quickly and accurately separate brain tumors from normal tissues in multi-modal brain tumor MRI images is of great significance for the clinical treatment of patients. In the brain tumor segmentation task, networks based on CNN and Transformer have been widely used.

[0003] However, there are still the following disadvantages: (1) The disadvantage of high network complexity: Since most existing brain tumor segmentation datasets are 3D MRI images, CNN networks based on 3D convolution have expensive computational costs, and for networks based on Transformer, when dealing with large window or long sequence data, the computational cost will increase quadratically. Therefore, for complex networks, higher-performance hardware resources are required; (2) The disadvantage of limited ability of the network to model long-range dependencies: Due to the limited size of the convolutional kernel of CNN, there are obvious deficiencies in extracting long-range dependencies between long-distance features. However, capturing long-range dependencies between long-distance voxels in brain tumor MRI images is crucial for improving the accuracy of brain tumor segmentation. Summary of the Invention

[0004] The purpose of the present invention is to provide a lightweight brain tumor segmentation network based on CNN-Mamba parallel encoding. This network efficiently extracts local features using a parallel Mamba feature extraction module, and enhances the network's ability to model long-range dependencies between brain tumor images. While maintaining a low computational cost, it improves the utilization efficiency of global features, enhances the segmentation accuracy of brain tumors, and helps to improve several aspects of the existing brain tumor segmentation network: reducing the model complexity, enhancing the ability to model long-range dependencies between long-distance features, and improving the local feature extraction effect of the Mamba architecture.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A lightweight brain tumor segmentation network based on CNN-Mamba parallel encoding, which adopts a U-shaped network architecture, and the main body of the network consists of an encoder, a decoder, a bottleneck layer, and a skip connection part;

[0006] Among them, the encoder and decoder are divided into four layers in total. Each layer of the encoder consists of a PMamba module and a downsampling module, each layer of the decoder consists of a PMamba module and an upsampling module, the bottleneck layer consists of a PMamba module, and the skip connection part is input from each layer of the encoder and each layer of the decoder and connected to the MDA module, and the output of the MDA module is used as the input of the decoder of the same layer;

[0007] During operation, the MRI brain tumor images of four modalities are combined into a four-channel network input. The PMamba module captures local features and long-range dependencies in the brain tumor images, and the downsampling module uses max pooling with a stride of 2 to reduce the image size; in each layer of the encoder, the size of the brain tumor image is halved and the number of channels is doubled; while in each layer of the decoder, the PMamab module and trilinear interpolation upsampling with a factor of 2 are used to restore the size and channels of the brain tumor image.

[0008] Preferably, the PMamba module is based on the Mamba architecture. By adding a path with multi-level dilated convolutions, while keeping the network lightweight, it enhances the ability of the Mamba architecture to extract local detailed features, and combines the state space model to capture long-range dependencies to extract global multi-scale features.

[0009] Preferably, when the PMamba module extracts global multi-scale features, it includes:

[0010] Separate the features of the upper layer by channels and send them into the two branches of the PMamba module. One branch is a Mamba block with two branches, and the other branch is a multi-level dilated convolution with a residual structure, and the dilation rates are 1, 2, and 4 respectively;

[0011] In the first branch of the Mamba block, the input features are dimensionally expanded through a linear layer, and then pass through a one-dimensional convolutional layer, a depthwise separable convolutional layer, a SiLU activation function, an SSM layer, and a LayerNorm normalization layer to obtain output features; while in the second branch, the input features pass through a linear layer and a SiLU activation function to obtain output features, and then the outputs of the two branches are subjected to Hadamard product operation and a linear layer and then output; finally, the output features of the two branches of the PMamba module are concatenated, and a learnable adjustment factor Scale is set at the residual connection to control the proportional residual input of the original feature map to realize the utilization of global feature information.

[0012] Preferably, after fusing the two input features, the MDA module enters three pseudo-3D convolutions with a dilation rate of 1 in parallel. While increasing the receptive field, it focuses on the detailed features from the coronal, sagittal, and horizontal view directions respectively. After passing through the BatchNorm normalization layer and the ReLU activation function, feature fusion is performed; then, an attention map is generated using the sigmoid function; finally, the two original input feature maps are multiplied element-wise with the generated attention map and fused with themselves through residual connections to achieve precise localization of the detailed features.

[0013] Preferably, the MDA module is also located in the skip connection part of the network, taking the outputs of the decoder and encoder at the same layer as inputs, extracting the correlation between the two, enhancing the intercommunication of multi-scale information and detailed information between the encoder and decoder, and enhancing the robustness of the network.

[0014] Compared with the prior art, the beneficial effects of the present invention are:

[0015] In order to overcome the problems that the limited size of 3D convolution leads to insufficient ability of the network to capture long-range dependence relationships, the high computational cost of the Transformer architecture, and the insufficient extraction of local detailed features by the Mamba architecture, the present invention proposes a lightweight brain tumor segmentation network DP-Net based on CNN-Mamba parallel encoding. DP-Net uses a lightweight parallel Mamba feature extraction module (PMamba module) to reduce the network complexity while enhancing the network's ability to extract long-range dependence relationships between distant features of brain tumors and local features, and to realize the utilization of global feature information. In addition, DP-Net focuses on the associated feature information of the same layer of the encoder and decoder from three view directions of 3D brain tumor images through a multi-view decoupled attention module (MDA module), strengthens the information intercommunication between the same layers, and filters redundant background information. Experimental results on the BraTS2020 dataset show that DP-Net uses 1.48M parameters and 54.85G floating-point operation counts, and while maintaining a small network complexity, significantly improves the segmentation performance of brain tumor images. Description of the Drawings

[0016] Figure 1 It is the network framework diagram of DP-Net.

[0017] Figure 2 It is the structure diagram of the PMamba module.

[0018] Figure 3 It is the structure diagram of the MDA module.

[0019] Figure 4Visualization images of ablation experiment results. Among them, (a) FLAIR modality; (b) predicted segmentation results of Baseline; (c) predicted segmentation results of Baseline+MDA; (d) predicted segmentation results of Baseline+PMamba; (e) predicted segmentation results of Baseline+MDA+PMamba (DP-Net); (f) ground truth labels. Detailed implementation manners

[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0021] In order to overcome the problems that the limited size of 3D convolution leads to insufficient ability of the network to capture long-range dependency relationships, the high computational cost of the Transformer architecture, and the insufficient extraction of local detailed features by the Mamba architecture, the present invention proposes a lightweight brain tumor segmentation network DP-Net based on CNN-Mamba parallel encoding.

[0022] The network structure of this DP-Net is as Figure 1 shown. The network DP-Net of the present invention adopts a U-shaped network architecture. The main body of the network consists of an encoder, a decoder, a bottleneck layer, and skip connections. Among them, the encoder and the decoder are divided into four layers in total. Each layer of the encoder is composed of a PMamba module and a downsampling module, and each layer of the decoder is composed of a PMamba module and an upsampling module. The bottleneck layer is composed of a PMamba module. The skip connections are input from each layer of the encoder and each layer of the decoder and connected to the MDA module. The output of the MDA module is used as the input of the same layer of the decoder. First, MRI brain tumor images of four modalities are combined into a four-channel network input. The PMamba module captures local features and long-range dependency relationships in the brain tumor images, and the downsampling module uses max pooling with a stride of 2 to reduce the image size. In each layer of the encoder, the size of the brain tumor image is halved and the number of channels is doubled. In each layer of the decoder, the PMamab module and trilinear interpolation upsampling with a factor of 2 are used to restore the size and channels of the brain tumor image. The specific network parameters of DP-Net are shown in Table 1.

[0023] Table 1 shows the module names of each layer of the DP-Net network, as well as the sizes of the input images and output images.

[0024] Table 1 DP-Net network parameters

[0025]

[0026] The structural diagram of the PMamba module is as follows Figure 2 shown. The PMamba module adds a path with multi-level dilated convolutions on the basis of the Mamba architecture. While maintaining the lightweight of the network, it enhances the ability of the Mamba architecture to extract local detailed features, and combines the State Space Model (SSM) to capture long-range dependencies and efficiently extract global multi-scale features. Specifically, first, the features of the upper layer are separated by channels and sent into two branches of the PMamba module respectively. One branch is a Mamba block with two branches, and the other branch is a multi-level dilated convolution with a residual structure, and the dilation rates are 1, 2, and 4 respectively. In the first branch of the Mamba block, the input features are dimensionally expanded through a linear layer, and then the output features are obtained after passing through a one-dimensional convolutional layer, a depthwise separable convolutional layer, a SiLU activation function, an SSM layer, and a LayerNorm normalization layer. In the second branch, the input features are passed through a linear layer and a SiLU activation function to obtain the output features, and then the outputs of the two branches are subjected to Hadamard product operation and a linear layer and then output. Finally, the output features of the two branches of the PMamba module are concatenated, and a learnable adjustment factor Scale is set at the residual connection to control the proportional residual input of the original feature map and further strengthen the effective utilization of global features.

[0027] The structural diagram of the MDA module is as follows Figure 3 shown. First, after two input features are fused, they enter three pseudo-3D convolutions with a dilation rate of 1 in parallel, increasing the receptive field while focusing on detailed features from the coronal, sagittal, and horizontal view directions respectively. After passing through the BatchNorm normalization layer and the ReLU activation function, feature fusion is performed. Then, an attention map is generated using the sigmoid function. Finally, the two original input feature maps are multiplied element-wise with the generated attention map respectively, and fused with themselves through residual connections to achieve precise localization of detailed features such as tumor edges.

[0028] The overall framework of the DP-Net network proposed by the present invention is as follows Figure 1 shown, including two important parts: (1) A lightweight parallel Mamba feature extraction module (Parallel Mamba Feature Extraction Module, PMamba) is proposed, and the structure of this module is as follows Figure 2As shown. Inspired by the visual state space model, the PMamba module is based on the Mamba architecture. By adding dilated convolution branch paths with different dilation rates and learnable adjustment factors, while maintaining the low computational cost of the Mamba architecture and enhancing the network's ability to model long-range dependencies, it makes up for the shortcoming of the Mamba architecture in extracting local detailed features and realizes the utilization of global feature information. (2) A Multiview Decoupled Attention Module (MDA) is proposed as the network skip connection part. The specific structure of this module is as Figure 3 shown. This module decouples 3D convolution using a combination of three parallel pseudo-3D dilated convolutions. While reducing the network's computational load and parameter quantity and increasing the receptive field, it focuses on local detailed features from different views, filters redundant background information, and improves the segmentation accuracy of detail regions such as the brain tumor edge. In addition, this module is located in the network's skip connection part, takes the outputs of the decoder and encoder at the same layer as inputs, extracts the correlation between them, enhances the intercommunication of multi-scale information and detail information between the encoder and decoder, and enhances the robustness of the network.

[0029] As an implementation of the present invention, the present invention selects to conduct experiments on the BraTS2020 dataset to verify the performance of the DP-Net. The training set of each dataset includes multiple cases, each case includes four-modal MRI brain tumor images and a ground truth label, while the validation set of each dataset does not include the ground truth label. The ground truth label is manually annotated by professional radiologists. Each brain tumor is divided into the following three parts, as Figure 4For example: in the regions of necrotic and non-enhancing tumors (NCR & NET), the label value is 1; in the regions of peritumoral edema (ED), the label value is 2; in the regions of enhancing tumors (ET), the label value is 4. The brain tumor segmentation task is to accurately segment the enhancing tumor (ET), which is composed of label 4; the tumor core (TC), which is composed of labels 1 and 4; and the whole tumor (WT), which is composed of labels 1, 2, and 4. To comprehensively and intuitively evaluate the performance of DP-Net, different evaluation metrics are selected for network complexity and segmentation accuracy. For segmentation accuracy, the Dice coefficient (Dice Similarity Coefficient, Dice) and the Hausdorff95 distance (Hausdorff95 Distance, HD) are selected as evaluation metrics. Among them, the Dice coefficient represents the degree of difference between the predicted segmentation result and the ground truth label, and the Hausdorff95 distance measures the distance between the predicted segmentation result and the ground truth label. For network complexity, the number of network parameters and the number of floating-point operations are used to quantitatively analyze the network complexity, which are expressed by the following formulas respectively: the number of network parameters , the number of floating-point operations . In the formulas, , , are the depth, height, and width of the convolutional kernel respectively; , are the number of input and output channels; , , are the depth, height, and width of the image respectively. The experiments of the present invention use the PyTorch deep learning framework. The server operating system is Ubuntu, the server CPU is an Intel Core i9-9900x processor, and the server GPU uses 4 parallel NVIDIA GTX2080Ti graphics cards with a video memory of 11GB. In the experiments of the present invention, the Adam optimizer is selected, the batch size is set to 8, the initial learning rate is set to 1x10 -3 , the weight decay coefficient is set to 1x10 -5 , and the maximum number of iterations is set to 1000.

[0030] In the preprocessing stage of the training set images, first, the MRI images with a size of 240×240×155 are cropped to 128×128×128 to adapt to the input image size of the network. In addition, images are randomly selected with a probability of 50%, flipped in different spatial axes, and randomly rotated within the range of -10° to 10°. Through the above data augmentation method, the amount of data is increased, the generalization of the network is improved, and the problem of overfitting that may occur during the training process is prevented.

[0031] The visualization results of the ablation experiment of the present invention are as Figure 4 shown. It can be seen from Figure (c) that compared with the baseline network, adding the MDA module makes the network more focused on the tumor boundary region and improves the segmentation effect of the tumor boundary shape. It can be seen from Figure (d) that by adding the PMamba module, the network's ability to capture the long-range dependence relationship and local detail features between distant voxels in brain tumor images is improved, making the network more accurate in segmenting the sizes of the tumor core and edema regions. It can be seen from Figure (e) that when both the MDA and PMamba modules are added, that is, the network DP-Net of the present invention. In the three views, compared with the Baseline, DP-Net reduces the problem of segmentation label errors, and the segmentation results of local detail regions such as the tumor boundary are closer to the ground truth label, indicating the effectiveness of the modules proposed in the present invention.

[0032] Table 2 shows the ablation experiment results of DP-Net. From the results, it can be seen that taking Residual U-Net as the baseline network Baseline, adding the MDA module of the present invention can help the network filter redundant background information. Although the network complexity increases slightly, the brain tumor segmentation accuracy is improved. When the basic module of Baseline is replaced with the lightweight module PMamba module of the present invention, the number of network parameters and the amount of calculation are greatly reduced, the segmentation accuracy is significantly improved, and the Hausdroff95 distance is reduced, indicating that the introduction of the PMamba module improves the utilization efficiency of global features. When both the PMamba module and the MDA module are added, that is, the network DP-Net of the present invention, the overall performance of DP-Net reaches the best in terms of the Dice coefficient and the Hausdroff95 distance.

[0033] Table 2 Overall network ablation experiment on the BraTS2020 dataset

[0034]

[0035] Table 3 shows the comparison experiment results of DP-Net with other networks on the BraTS2020 dataset respectively. The experimental results show that DP-Net significantly improves the segmentation effect of brain tumor images while maintaining a low network complexity.

[0036] Table 3 Comparison results with different brain tumor segmentation networks on the BraTS2020 dataset

[0037]

[0038] The present invention discloses a lightweight brain tumor segmentation network DP-Net based on CNN-Mamba parallel encoding. DP-Net reduces the network complexity while enhancing the network's ability to extract long-range dependencies and local features between distant features of brain tumors through a lightweight parallel Mamba feature extraction module, realizing the utilization of global feature information. In addition, DP-Net focuses on the associated feature information of the same layer between the encoder and the decoder from three view directions of 3D brain tumor images through a multi-view decoupled attention module, strengthening the information exchange between the same layers and filtering redundant background information. Experimental results on the BraTS2020 dataset show that DP-Net uses 1.48M parameters and 54.85G floating-point operations, significantly improving the segmentation performance of brain tumor images while maintaining a small network complexity.

[0039] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0040] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these modifications and variations.

Claims

1. A lightweight brain tumor segmentation network based on CNN-Mamba parallel coding, characterized in that: The lightweight brain tumor segmentation network adopts a U-shaped network architecture, and the main body of the network consists of an encoder, a decoder, a bottleneck layer, and a jump connection part; The encoder and decoder are divided into four layers. Each encoder layer consists of a PMamba module and a downsampling module. Each decoder layer consists of a PMamba module and an upsampling module. The bottleneck layer consists of a PMamba module. The skip connection part is input by each encoder layer and each decoder layer, and connected to the MDA module. The output of the MDA module is used as the input of the decoder at the same layer. During operation, MRI brain tumor images of four modalities are combined into a four-channel network input. The PMamba module captures local features and long-range dependencies in brain tumor images, and the downsampling module uses maximum pooling with a step size of 2 to reduce the image size. In each encoder layer, the size of the brain tumor image is halved and the channel is doubled. In each decoder layer, the PMamab module and trilinear interpolation upsampling with a multiple of 2 are used to achieve the recovery of the brain tumor image size and channel.

2. The lightweight brain tumor segmentation network based on CNN-Mamba parallel coding according to claim 1, characterized in that: The PMamba module is based on the Mamba architecture. By adding a path with multi-level hole convolution, the Mamba architecture enhances the ability to extract local detail features while keeping the network lightweight, and combines the state space model to capture long-range dependencies to extract global multi-scale features.

3. The lightweight brain tumor segmentation network based on CNN-Mamba parallel coding according to claim 2, characterized in that: When extracting global multi-scale features, the PMamba module includes: The features of the upper layer are channel-separated and sent to two branches of the PMamba module. One branch is a Mamba block with two branches, and the other branch is a multi-level dilated convolution with a residual structure, with dilation rates of 1, 2, and 4 respectively. In the first branch of the Mamba block, the input features are dimensionally expanded through a linear layer, and then the output features are obtained after passing through a one-dimensional convolutional layer, a depth-wise separable convolutional layer, a SiLU activation function, an SSM layer, and a LayerNorm normalization layer; in the second branch, the input features are obtained after passing through a linear layer and a SiLU activation function, and then the outputs of the two branches are subjected to Hadamard product operations and linear layer outputs; finally, the output features of the two branches of the PMamba module are concatenated, and a learnable adjustment factor Scale is set at the residual connection to control the original feature map to perform residual input proportionally, thereby realizing the utilization of global feature information.

4. The lightweight brain tumor segmentation network based on CNN-Mamba parallel coding according to claim 3, characterized in that: After fusing the two input features, the MDA module enters three pseudo 3D convolutions with a void rate of 1 in parallel, while increasing the receptive field, focusing on detail features from the coronal, sagittal and horizontal viewing directions, and performing feature fusion after passing through the BatchNorm normalization layer and the ReLU activation function; Then, the sigmoid function is used to generate the attention map. Finally, the two feature maps of the original input are element-wise multiplied with the generated attention map and fused with themselves through residual connection to achieve accurate positioning of detail features.

5. The lightweight brain tumor segmentation network based on CNN-Mamba parallel coding according to claim 4, characterized in that: The MDA module is also located in the skip connection part of the network, taking the output of the decoder and encoder at the same layer as input, extracting the correlation between the two, and enhancing the communication of multi-scale information and detail information between the encoder and decoder.

Citation Information

Patent Citations

  • Multi-context brain tumor segmentation system based on scale fusion guidance

    CN117237320A

  • CNN and attention-based MRI brain tumor segmentation method

    CN117764953A

  • MRI (Magnetic Resonance Imaging) brain tumor segmentation method based on multi-scale feature fusion of improved U-Net

    CN117876399A

  • Brain tumor image segmentation method based on multi-scale convolution and Mama structure

    CN118447244A

  • Image segmentation method based on medical hyperspectral image segmentation network

    CN118967706A

Cited By

  • Automobile part defect detection method and system based on multi-scale feature fusion

    CN121563974A

  • A method and system for detecting defects of automobile parts based on multi-scale feature fusion

    CN121563974B

  • Medical image segmentation method and system based on Mama and CNN parallel double-flow fusion

    CN121811417A