Three-dimensional brain tumor image segmentation method based on multi-level Mama architecture

By adopting a multi-level Mamba architecture in three-dimensional medical image segmentation technology, combining Mamba Block and SS3D modules, the problems of memory limitations, computing bottlenecks, overfitting and gradient disappearance in the existing technology are solved, and a more efficient and accurate image segmentation effect is achieved.

CN120107577AActive Publication Date: 2025-06-06HANGZHOU DIANZI UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510064616.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-06-06
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

Existing three-dimensional medical image segmentation technology is prone to memory limitations and computing bottlenecks when processing large-size data, and may have problems such as overfitting and gradient disappearance.

Method used

Using a three-dimensional brain tumor image segmentation method based on a multi-level Mamba architecture, the segmentation speed and accuracy are improved by introducing Mamba Block and SS3D modules into the encoder, combined with a lightweight decoder design.

Benefits of technology

It significantly improves the speed and accuracy of three-dimensional medical image segmentation, reduces network complexity, avoids overfitting, and solves the problem of gradient vanishing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107577A_ABST
    Figure CN120107577A_ABST
Patent Text Reader

Abstract

The invention provides a three-dimensional brain tumor image segmentation method based on a multilevel Mama framework, and the method comprises the steps: constructing a SegMama 3D network, inputting an original three-dimensional brain tumor MRI image into the network, carrying out the coding and decoding operation, and outputting a predicted brain tumor segmentation mask. And the SegMamba3D network comprises an encoder and a decoder. The encoder performs multi-layer encoding operation on an input sample, first performs down-sampling operation once on each layer, and then outputs an encoding feature map through the Mama Block. The decoder firstly unifies the size and channel number of coding feature maps of different levels output by the encoder through convolution operation, and splices the coding feature maps. And inputting the spliced feature map into a multi-layer perceptron, and restoring the spliced feature map to be consistent with the input sample in size. And learning a real segmentation mask image through a training network, and outputting a tumor segmentation result of the three-dimensional brain tumor image by using the trained network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of MRI image segmentation, and is a three-dimensional brain tumor image segmentation method based on a multi-level Mamba architecture. Background Art

[0002] Computer-aided medical image analysis plays an important role in assisting doctors in making clinical diagnoses and determining the best treatment options. Medical image segmentation, in particular, can help doctors quickly locate the target area, which not only improves the efficiency and accuracy of diagnosis, but also promotes the deep integration of clinical practice and scientific research.

[0003] In the field of medical image segmentation, models based on deep learning have been widely explored. Traditional convolutional neural networks (CNNs) extract features of 3D images through 3D convolutions, and complete the segmentation of 3D images with pooling layers, upsampling layers, and activation functions, but may suffer from overfitting and gradient vanishing problems. The Transformer architecture has demonstrated extraordinary capabilities in modeling global relationships. The architecture based on Vision Transformers (ViTs) can effectively capture global information in images through the self-attention mechanism, thereby improving the ability of image segmentation, making significant progress in 3D medical image segmentation. Correspondingly, it also generates huge computational and memory overheads, especially when processing large-size 3D data, it is easy to encounter memory limitations and computational bottlenecks. Summary of the invention

[0004] In view of the shortcomings of the existing technology, the present invention proposes a three-dimensional brain tumor image segmentation method based on the multi-level Mamba architecture, performs a lightweight design on the decoder, and introduces the SS3D module to improve the speed and accuracy of three-dimensional medical image segmentation.

[0005] The three-dimensional brain tumor image segmentation method based on the multi-level Mamba architecture includes the following steps:

[0006] Step 1: Use the original 3D brain tumor MRI image as a sample, annotate the true segmentation mask image as a label, and construct a training set.

[0007] Step 2: Build a SegMamba3D network, input the sample obtained in step 1 into the network, and after encoding and decoding operations, output the predicted brain tumor segmentation mask.

[0008] The SegMamba3D network includes an encoder and a decoder. The encoder performs multi-layer encoding operations on the input sample, performs a downsampling operation at each layer, and then outputs a coding feature map by introducing the Mamba Block of the SS3D layer. The decoder first unifies the size and number of channels of the coding feature maps of different levels output by the encoder through convolution operations and splices them. The spliced ​​feature map is then input into a multi-layer perceptron to restore it to the same size as the input sample, and a predicted brain tumor segmentation mask is obtained.

[0009] The Mamba Block includes a normalization layer, a linear layer, a convolutional layer, and an SS3D layer, and through residual connections, the network can more easily transfer gradients during back propagation, thereby solving the gradient vanishing problem in deep network training:

[0010] F 1 =Linear(LayerNorm(F))

[0011] F 2 =SiLU(DWConv(F 1 ))

[0012] F 3 =SS3D(F 2 )

[0013] F 4 =Linear(LayerNorm(F 3 ))

[0014] F 5 =F 4 +F

[0015] F 6 =Linear(LayerNorm(F 5 ))

[0016] F 7 =Linear(GELU(DWConv(F 6 )))

[0017] M=F 5 +F 7

[0018] Among them, LayerNorm() represents the normalization layer, Linear() represents the linear layer, DWConv() represents the channel-by-channel convolution layer, SS3D() represents the SS3D layer, SiLU() represents the SiLU activation function, and GELU() represents the GELU activation function. F represents the input feature of Mamba Block, M represents the output feature of Mamba Block, and F represents the input feature of Mamba Block. 1 ~F7 Indicates the intermediate features of the Mamba Block.

[0019] The multi-layer perceptron performs convolution, normalization and upsampling operations on the concatenated feature maps:

[0020] Q 1 =Conv3d(4C,C)

[0021] Q 2 =ReLU(BatchNorm(Q 1 ))

[0022] Q 3 =Conv3d(C,N CLS )(Q 2 )

[0023] Q 4 =Upsample(D×H×W)(Q 3 )

[0024] Among them, Conv3d() represents three-dimensional convolution, C represents the channel dimension size of the input sample, ReLU() represents the ReLU activation function, BatchNorm() represents batch normalization, and N CLS is the number of categories of the real segmentation mask image, Upsample() is the upsampling operation, and D, H, and W represent the depth, height, and width of the input sample, respectively.

[0025] Preferably, the upsampling operation is trilinear interpolation, and the image is interpolated in three dimensions: width, height and depth.

[0026] Step 3: Compare the brain tumor segmentation mask predicted in step 2 with the real segmentation mask image annotated in step 1, calculate the loss value, and adjust the model parameters according to the loss value.

[0027] Step 4: Input the three-dimensional brain tumor MRI image into the SegMamba3D network trained in step 3 to obtain the tumor segmentation result in the image.

[0028] The present invention has the following beneficial effects:

[0029] Using the above technical solution, the overall network framework is designed with a classic encoder-decoder. Mamba is integrated into the encoder as a state space model (SSM) to model the long-range dependencies of sequence data. Compared with Transformer, it has significant storage efficiency and computing speed. The hierarchical structure of the encoder can capture multi-level features and improve the performance of semantic segmentation. The lightweight design of the decoder reduces the complexity of the network framework and avoids overfitting to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a flow chart of the present invention.

[0031] Figure 2 This is a network model framework diagram of the present invention.

[0032] Figure 3 It is a framework diagram of the Mamba module of the present invention.

[0033] Figure 4 It is a schematic diagram of the path of SS3D scanning three-dimensional images in the Mamba module of the present invention.

[0034] Figure 5 It is a framework diagram of the MLP module of the present invention. DETAILED DESCRIPTION

[0035] The present invention will be further explained below with reference to the accompanying drawings;

[0036] The present invention covers any substitution, modification, equivalent method and scheme made on the essence and scope of the present invention as defined by the claims. Further, in order to make the public have a better understanding of the present invention, some specific details are described in detail in the following detailed description of the present invention. Those skilled in the art can fully understand the present invention without the description of these details.

[0037] The present invention provides a three-dimensional brain tumor image segmentation method based on a multi-level Mamba architecture, such as Figure 1 As shown, the specific steps include:

[0038] Step 1. This embodiment uses the BraTs tumor dataset as the original data source, merges the four modalities of FLAIR, T1w, T1gd and T2w of each original three-dimensional brain tumor MRI image into the same three-dimensional image after preprocessing, and divides it into a training set, a validation set and a test set together with the corresponding true segmentation mask image; the steps of data preprocessing are as follows:

[0039] s1.1. Read four modalities of FLAIR, T1w, T1gd and T2w images, and separate the data information and spatial conversion information.

[0040] s1.2. The sizes of the original images of the four modalities are all (240, 240, 155). Each dimension of the image data information of each modality is normalized and the channel dimension is added. The final size of the image data information is (1, 240, 240, 155).

[0041] s1.3. Create a MetaTensor object containing the final image data information and the corresponding spatial transformation information; then rearrange it to conform to the direction of the RAS coordinate system, and finally separate the data information from the above MetaTensor object containing data and spatial information and convert it into a normal Numpy data.

[0042] s1.4. Crop the area with pixel value 0 and crop the image from the original (1, 240, 240, 155) to (1, 128, 128, 128).

[0043] s1.5. The four modes of the image processed by the above steps are spliced ​​into a three-dimensional image along the first dimension, and then all the processed data are divided into a training set, a validation set and a test set according to a certain ratio.

[0044] Step 2: Build a Figure 2 The SegMamba3D network shown in the figure inputs the sample obtained in step 1 into the network, and after encoding and decoding operations, outputs the predicted brain tumor segmentation mask. The specific steps are as follows:

[0045] s2.1. The SegMamba3D network performs a 4-layer encoding operation on the input sample, performs a downsampling operation on each layer, and then outputs a coding feature map through two consecutive Mamba Blocks.

[0046] For an input sample of size D×H×W×C, a downsampling operation is first performed through a three-dimensional convolution with a stride of 4, a padding of 3, and a convolution kernel size of 7, and then feature extraction is performed through two consecutive Mamba blocks to obtain a sample of size Repeat the above steps three times to obtain the encoding feature map of size These features provide high-resolution coarse features and low-resolution fine-grained features, which can improve the performance of semantic segmentation. 1 =32, C 2 =64, C 3 =160, C 4 =256.

[0047] Figure 3 The structure diagram of the Mamba Block includes a normalization layer, a linear layer, a convolutional layer, a SS3D layer, and an activation function:

[0048] F 1 =Linear(LayerNorm(F))

[0049] F 2 =SiLU(DWConv(F 1 ))

[0050] F 3 =SS3D(F 2 )

[0051] F 4 =Linear(LayerNorm(F 3 ))

[0052] F 5 =F 4 +F

[0053] F 6 =Linear(LayerNorm(F 5 ))

[0054] F 7 =Linear(GELU(DWConv(F 6 )))

[0055] M=F 5 +F 7

[0056] Among them, LayerNorm() represents the normalization layer, Linear() represents the linear layer, DWConv() represents the channel-by-channel convolution layer, SS3D() represents the SS3D layer, SiLU() represents the SiLU activation function, and GELU() represents the GELU activation function. F represents the input feature of Mamba Block, M represents the output feature of Mamba Block, and F represents the input feature of Mamba Block. 1 ~F 7 Indicates the intermediate features of the Mamba Block.

[0057] In Mamba Block, the normalization layer uses Layer Norm to normalize all features of each sample, which eliminates the size relationship between different samples, but retains the size relationship between different features within a sample. The convolution layer uses channel-by-channel convolution, which has lower parameters and computational cost than conventional convolution. The SS3D module specifies the scanning path of the input 3D image, such as Figure 4 As shown in the figure, the entire 3D image is scanned along the specified route in the figure, making full use of every possible unexpanded sequence, so as to have an in-depth analysis of the entire 3D image and achieve a better segmentation effect. The activation function uses GELU and SiLU. In addition, the design of the module also uses the residual connection idea of ​​ResNet, which makes it easier for the network to transfer gradients during back propagation, thereby solving the gradient disappearance problem in deep network training.

[0058] s2.2, M i Represents the output features of the 2ith Mamba Block, for the 4-level encoding feature map M output by the encoder i , processed uniformly through the multi-layer perceptron layer The size is then concatenated in the fourth dimension to a size of The feature map P is:

[0059]

[0060] Where i = 1, 2, 3, 4, C i Represents the encoded feature map M i The number of channels, Concat() is a concatenation operation.

[0061] s2.3, such as Figure 5 As shown, for the concatenated feature map P, through convolution, normalization and upsampling operations, the output size is D×H×W×N CLS The predicted brain tumor segmentation mask:

[0062] Q 1 =Conv3d(4C,C)

[0063] Q 2 =ReLU(BatchNorm(Q 1 ))

[0064] Q 3 =Conv3d(C,N CLS )(Q 2 )

[0065] Q 4 =Upsample(D×H×W)(Q 3 )

[0066] Among them, Conv3d() represents three-dimensional convolution, C represents the channel dimension size of the input sample, ReLU() represents the ReLU activation function, BatchNorm() represents batch normalization, and N CLSis the number of categories of the real segmentation mask image, Upsample() is the upsampling operation, and D, H, and W represent the depth, height, and width of the input sample, respectively.

[0067] Step 3: Compare the predicted brain tumor segmentation mask output by the SegMamba3D network with the actual segmentation mask image corresponding to the sample, calculate the loss value DiceLoss, adjust the model parameters, and then perform the segmentation processing of the next image. After all images are processed, perform the next round of prediction. After all rounds are completed, the final prediction model is obtained:

[0068]

[0069] Among them, N is the number of pixels of the input sample, y n For labels, p n is the probability predicted by the model.

[0070] After training for 800 rounds on the BraTs2017 and BraTs2021 datasets, respectively, it can be seen that the average Dice scores are approximately 81.74 and 86.87, respectively, indicating that there is an average of 81.74% and 86.87% overlap between the predicted images and the real images on the two datasets.

[0071] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, various changes, modifications, substitutions and variations of these embodiments including components are made without departing from the principles and spirit of the present invention, and still fall within the scope of protection of the present invention.

Claims

1. A 3D brain tumor image segmentation method based on a multi-level Mamba architecture, characterized by: The method is used to segment the location of a brain tumor from a three-dimensional MRI image in which a brain tumor exists, and specifically includes the following steps: Step 1: Use the original 3D brain tumor MRI image as a sample, annotate the true segmentation mask image as a label, and construct a training set; Step 2: Build a SegMamba3D network including an encoder and a decoder, input the sample obtained in step 1 into the network, and output the predicted brain tumor segmentation mask; The encoder performs multi-layer encoding operations on the input sample, performs a downsampling operation at each layer, and then outputs a coding feature map by introducing the Mamba Block of the SS3D layer; the decoder first unifies the size and number of channels of the coding feature maps of different layers output by the encoder through a convolution operation, and splices them; then the spliced ​​feature map is input into a multi-layer perceptron to restore it to the same size as the input sample, and a predicted brain tumor segmentation mask is obtained; Step 3: Compare the brain tumor segmentation mask predicted in step 2 with the real segmentation mask image marked in step 1, calculate the loss value, and adjust the model parameters according to the loss value; Step 4: Input the three-dimensional brain tumor MRI image into the SegMamba3D network trained in step 3 to obtain the tumor segmentation result in the image.

2. The three-dimensional brain tumor image segmentation method based on the multi-level Mamba architecture as claimed in claim 1, characterized in that: For the images of multiple modalities of the original three-dimensional brain tumor MRI image, the data information and spatial transformation information are extracted, the images of multiple modalities are merged into the same three-dimensional image, and the edge areas with pixel values ​​of 0 are cut off as samples.

3. The three-dimensional brain tumor image segmentation method based on the multi-level Mamba architecture as claimed in claim 1, characterized in that: The encoder performs a 4-layer encoding operation on the input sample, performs a downsampling operation on each layer, and then outputs a coded feature map M through two consecutive Mamba Blocks introduced into the SS3D layer. i .

4. The three-dimensional brain tumor image segmentation method based on the multi-level Mamba architecture as claimed in claim 1 or 3, characterized in that: The Mamba Block that introduces the SS3D layer includes a normalization layer, a linear layer, a convolutional layer, a SS3D layer, and a residual connection operation: F1=Linear(LayerNorm(F)) F2=SiLU(DWConv(F1)) F3=SS3D(F2) F4=Linear(LayerNorm(F3)) F5=F4+F F6=Linear(LayerNorm(F5)) F7=Linear(GELU(DWConv(F6))) M=F5+F7 Among them, LayerNorm() represents the normalization layer, Linear() represents the linear layer, DWConv() represents the channel-by-channel convolution layer, SS3D() represents the SS3D layer, SiLU() represents the SiLU activation function, and GELU() represents the GELU activation function; F represents the input feature of Mamba Block, M represents the output feature of Mamba Block, and F1~F7 represent the intermediate features of Mamba Block.

5. The three-dimensional brain tumor image segmentation method based on the multi-level Mamba architecture as claimed in claim 1, characterized in that: The multi-layer perceptron performs convolution, normalization and upsampling operations on the concatenated feature maps: Q1=Conv3d(4C,C) Q2 = ReLU(BatchNorm(Q1)) <h2 style=";text-align:left;direction:ltr">Q3=Conv3d(C,N<h2 style=";text-align:left;direction:ltr"> CLS <h2 style=";text-align:left;direction:ltr"> (Q2) Q4=Upsample(D×H×W)(Q3) Among them, Conv3d() represents three-dimensional convolution, C represents the channel dimension size of the input sample, ReLU() represents the ReLU activation function, BatchNorm() represents batch normalization, and N CLS is the number of categories of the real segmentation mask image, Upsample() is the upsampling operation, and D, H, and W represent the depth, height, and width of the input sample, respectively.

6. The three-dimensional brain tumor image segmentation method based on the multi-level Mamba architecture as claimed in claim 5, characterized in that: The upsampling operation is trilinear interpolation, which interpolates the image in three dimensions: width, height, and depth.

7. The three-dimensional brain tumor image segmentation method based on the multi-level Mamba architecture as claimed in claim 1, characterized in that: Use DiceLoss as the loss value: Among them, N is the number of pixels of the input sample, y n is the label of the input sample, p n Predict probabilities for the output of the SegMamba3D network.

8. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 3 and 5 to 7.

Citation Information

Patent Citations

  • HL-UNet image segmentation model and heart dynamic magnetic resonance imaging segmentation method

    CN118470036A

  • Mandibular neural tube CBCT panoramic image segmentation network training method based on Mmba architecture

    CN119152212A

  • Video sequence segmentation method based on selective scanning visual state space model

    CN119206568A

  • Multi-modal brain tumor image segmentation method based on self-supervised learning

    WO2024108522A1

  • Transitive and commutative multimodal models and uses

    WO2024223621A1