Brain tumor segmentation algorithm based on section anisotropy CASVim and uncertainty gating UCG
By integrating the multi-section visual Mamba feature extraction and uncertainty gating module into the brain tumor segmentation algorithm, the problems of blurred boundaries and complex background in 3D brain tumor image segmentation are solved, achieving high-precision and high-robustness segmentation effects.
Patent Information
- Application Number
- CN202510795698.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-23
AI Technical Summary
When existing algorithms process 3D brain tumor images, the problems of blurred segmentation boundaries and complex backgrounds are difficult to solve, resulting in limited diagnostic accuracy and stability.
A brain tumor segmentation algorithm based on slice anisotropy CASVim and uncertainty gating UCG is adopted. Through the multi-slice visual Mamba feature extraction module CAS, combined with the uncertainty gating module UCG and the multi-view fusion module MPFB, the cross entropy fusion Dice loss function and the uncertainty loss function are used to achieve efficient feature extraction and segmentation.
It significantly improves the accuracy and reliability of brain tumor image segmentation, reduces computational complexity, enhances robustness to complex backgrounds and noise, and improves the accuracy and stability of segmentation results.
Smart Images

Figure BDA0005449660140000041 
Figure BDA0005449660140000051 
Figure BDA0005449660140000071
Abstract
Description
Technical Field
[0001] The present invention relates to an image data processing method, and in particular to an algorithm for processing fuzzy boundary and complex background of 3D brain tumor image segmentation by using section anisotropy and uncertainty gating, belonging to the technical field of medical image analysis. Background Art
[0002] Brain tumors are one of the most common and serious diseases of the nervous system. Early screening and accurate diagnosis of brain tumors have a vital impact on the overall health and survival rate of patients. The growth characteristics of brain tumors are complex and diverse, including abnormal proliferation of tumor cells, invasion of surrounding normal brain tissue, induction of cerebral edema, and destruction of the blood-brain barrier. These changes present diverse characteristics in medical imaging, such as unclear tumor boundaries, uneven internal structures, and complex enhancement patterns after enhanced scanning. These complex imaging features make the early identification and accurate diagnosis of brain tumors a huge challenge, and traditional imaging diagnostic methods often have certain limitations when dealing with these complex situations. In addition, there are many types of brain tumors. Different types of brain tumors have different imaging manifestations, such as gliomas, meningiomas, metastatic tumors, etc. Each type has its own unique biological behavior and treatment response, which further increases the complexity of brain tumor diagnosis and treatment.
[0003] With the rapid development of modern medical imaging technologies, such as magnetic resonance imaging (MRI) and computed tomography (CT), medical institutions have accumulated vast amounts of brain imaging data. In particular, high-resolution, multimodal brain tumor imaging data continues to emerge, providing a rich information resource for in-depth research and precise diagnosis of brain tumors. However, radiologists face immense workloads, and their limited resources struggle to meet the growing demand for diagnostic imaging. Manual observation and analysis of this complex brain tumor imaging data is not only time-consuming and laborious, but also susceptible to factors such as physician fatigue and experience differences, limiting the accuracy and stability of diagnostic results. Against this backdrop, the rapid development of artificial intelligence (AI), particularly deep learning, has brought new hope and transformative opportunities to brain tumor medical imaging analysis. Deep learning, as a powerful machine learning method, is highly adaptable and can automatically learn rich feature representations from large amounts of brain tumor imaging data, eliminating the need for manual feature extraction algorithms and significantly improving the efficiency and accuracy of feature extraction. It also possesses powerful large-scale data processing capabilities, enabling efficient processing and analysis of massive amounts of multimodal brain tumor imaging data to uncover valuable potential insights. Furthermore, deep learning models can learn deep, abstract features within the data. These features can often capture subtle lesions and complex patterns in brain tumor images that are imperceptible to the naked eye, providing stronger support for accurate diagnosis. Furthermore, after sufficient training, deep learning models demonstrate excellent generalization capabilities, maintaining relatively stable performance across diverse datasets and clinical scenarios. This gives them greater potential and value in real-world clinical applications.
[0004] Deep learning technology demonstrates unique advantages in brain tumor image segmentation, effectively addressing the numerous challenges faced in this area. For example, the morphological diversity, anatomical complexity, and individual variability of brain tumors have long been bottlenecks that traditional segmentation methods have struggled to overcome. However, deep learning models, leveraging their powerful feature learning and nonlinear mapping capabilities, can automatically learn the various morphological characteristics of brain tumors and their boundaries with surrounding brain tissue, enabling precise segmentation.
[0005] U-Mamba innovatively combines the advantages of convolutional neural networks (CNNs) and state-space sequence models (SSMs). It can not only extract local fine-grained features in images through CNNs, but also capture long-range dependencies with the help of SSMs, thereby more comprehensively understanding image information and providing richer feature representations for biomedical image segmentation. Compared with Transformer-based architectures, U-Mamba has the ability to scale linearly in feature size, rather than the quadratic complexity of traditional Transformer architectures. This enables U-Mamba to more efficiently utilize computing resources when processing large-scale image data, reducing computing costs and time overhead, while avoiding performance bottlenecks caused by high computational complexity, allowing medical image segmentation tasks to be performed more quickly. Summary of the Invention
[0006] The present invention aims to provide a brain tumor segmentation algorithm based on slice anisotropy CASVim and uncertainty gated UCG to solve the problems of fuzzy segmentation boundaries and complex background in existing algorithms for processing 3D images.
[0007] The brain tumor segmentation algorithm based on slice anisotropy CASVim and uncertainty gated UCG in this scheme includes the following steps:
[0008] S1: Obtain original medical images and perform preprocessing;
[0009] S2: Feature extraction is performed on the image using CASVim (CAS), a visual feature extraction module that integrates multiple slices. This module captures and segments the image at different scales. The original 3D image is re-segmented into axial, coronal, and sagittal slices, which are orthogonal to each other. Different spatial features are obtained through multi-structure decomposition. Downsampling is also performed using the CAS module.
[0010] S3: The uncertainty gating module (UCG) cooperates with the lightweight residual module in the jump connection to decode features at different levels, ensuring the quality of features obtained by the model and reducing the large number of parameters brought by the fusion of multi-faceted spatial features; it calculates channel uncertainty and uses the standard deviation in the channel dimension to distinguish high-uncertainty features from low-uncertainty features.
[0011] S4: While decoding shallow to deep features, low-uncertainty features are input into the multi-view fusion module (MPFB). For high-uncertainty features, multiple levels of feature decoding and fusion are used.
[0012] S5: Cross entropy fusion Dice loss function combined with uncertainty loss function is used to ensure segmentation results and pay attention to the uncertainty of model prediction.
[0013] The beneficial effects of this program are:
[0014] The algorithm proposed in this scheme has significant beneficial effects. By integrating the multi-section visual Mamba feature extraction module CAS, it is possible to fully capture the features of brain tumor images at different scales and spatial directions. The original 3D image is re-divided into three orthogonal 2D sections: axial, coronal, and sagittal, realizing multi-structure decomposition to obtain different spatial features. At the same time, combined with shallow downsampling operations, it effectively solves the problem of blurred section boundaries when existing algorithms process 3D images, making the segmentation results more accurately fit the actual boundaries of brain tumors. CAS cooperates with deep downsampling operations to ensure the effectiveness of subsequent processing of complex background problems. The uncertainty gating module (UCG) quantifies the uncertainty of features by calculating the standard deviation on the channel dimension, distinguishing high-uncertainty features from low-uncertainty features. In the process of decoding features at different levels, UCG cooperates with the lightweight residual module in the jump connection to ensure that the model obtains high-quality features while reducing the large number of parameters brought by the fusion of multi-section spatial features. For low-uncertainty features, their weight in segmentation decisions is increased to ensure the model fully utilizes reliable features. For high-uncertainty features, the impact of uncertainty is gradually reduced through multiple levels of feature decoding and fusion, effectively reducing the interference of complex backgrounds on segmentation results and improving segmentation accuracy. Furthermore, the algorithm uses a slice consistency loss function fused with an uncertainty loss function to evaluate the difference between segmentation results and true segmentation results. This comprehensive evaluation method more comprehensively and objectively reflects the algorithm's segmentation performance, providing more instructive feedback for model optimization.
[0015] Furthermore, in S1, irrelevant regions in the original 3D brain tumor image are cropped, and the 3D image is Z-Score normalized (Norm).
[0016] The beneficial effect is that cropping irrelevant areas in the original 3D brain tumor image can effectively reduce the amount of data and remove background information and unnecessary anatomical structures in the image that are irrelevant to the brain tumor segmentation task, thereby reducing the computational complexity of subsequent processing and improving the operating efficiency of the algorithm. At the same time, this operation helps to reduce background noise and interference, allowing the model to focus more on key areas containing brain tumors, thereby improving the accuracy of segmentation. In addition, Z-Score standardization of 3D images can convert the image data into a distribution with a mean of 0 and a standard deviation of 1, eliminating the skewed distribution and dimensional differences that may exist in the original data, making brain tumor image data of different modalities or different patients comparable and consistent. This not only helps to accelerate the convergence process of the model and improve training efficiency, but also enhances the robustness of the model to the grayscale features of the image, ensuring the stability and reliability of the subsequent feature extraction and segmentation process.
[0017] Furthermore, in S2, feature extraction is performed on the image using the multi-slice-fused visual Mamba feature extraction module CASVim (CAS). This module captures and segments the image at different scales. The original 3D image is re-segmented into three mutually orthogonal 2D slices: axial, coronal, and sagittal. Different spatial features are then extracted through multi-structure decomposition. The CAS inference process can be summarized as follows:
[0018] F Coronal =f coronal (X),F Sagittal =f sagittal (X)
[0019] F Shallow =f shallow (F Coronal ,F Sagittal )
[0020] F Axial =f axial (X),F Deep =f deep (F Axial )
[0021] Among them F Coronal ∈R C×H×W and F Sagittal ∈R C×H×W The superficial features of the coronal and sagittal sections, F Axial ∈R C×H×W is the deep feature of the axial section, where C represents the number of channels, and H and W represent the height and width, respectively.
[0022] The beneficial effect is that CAS can extract features from different perspectives by segmenting the 3D image in three directions: axial, coronal, and sagittal, thereby capturing the structure and semantic information of the 3D image more comprehensively. This multi-perspective feature extraction method helps to better understand the spatial relationship in the image and the three-dimensional form of the object. It fuses the coronal and sagittal sections for extraction and inputs them into the shallow features, and extracts the axial section features into the deep features. The feature fusion of the coronal and sagittal sections can provide richer local structural information, which helps the model better understand the details and boundary information of the image at the shallow stage. The axial section feature extraction can pay more attention to the key semantic information of the image, thereby improving the model's ability to understand the complex background of the image, thereby ensuring the subsequent image segmentation effect.
[0023] Furthermore, in S3, the uncertainty gating module (UCG) cooperates with the lightweight residual module in the jump connection to decode features at different levels, ensuring the quality of features obtained by the model and reducing the large number of parameters brought by the fusion of multi-faceted spatial features; channel uncertainty is calculated, and the standard deviation in the channel dimension is used to distinguish high-uncertainty features from low-uncertainty features. The reasoning process can be summarized as follows:
[0024]
[0025] U S(h,w) =Fusion(U Cor(h,w) ,U Sag(h,w) ), U D(h,w) =U Axial(h,w)
[0026] (U low ,U high )=UCG(Reshape(U D(h,w) ,U S(h,w) ))
[0027] Among them, U Cor(h,w) ,U Sag(h,w) ,U Axial(h,w) Represents the uncertainty quantification results of the coronal and sagittal sections and the axial section respectively; U S(h,w) ,U D(h,w) ,U low ,U high Represent the shallow layer, deep layer and the results of distinguishing high uncertainty and low uncertainty respectively; θ (h,w) Represents the average value of all channels at this spatial position.
[0028] The beneficial effects are as follows: using the UCG module to extract deep features from medical image features of different sections, and quantifying uncertainty by calculating the standard deviation of the channel dimension for the shallow feature fusion of the coronal and sagittal sections and the deep features of the axial section. The larger the standard deviation, the greater the variation of the feature in the channel dimension and the higher the uncertainty; the smaller the standard deviation, the smaller the variation of the feature in the channel dimension and the lower the uncertainty. By setting a threshold, the features are divided into high uncertainty and low uncertainty categories. Based on the quantified channel uncertainty, the feature weights are dynamically adjusted. For high uncertainty features, their weights are reduced to reduce reliance on these uncertain features; for low uncertainty features, the weights are maintained to enhance the utilization of these reliable features.
[0029] Furthermore, in S4, while decoding the shallow to deep features, the low-uncertainty features are input into the multi-view fusion module (MPFB), and the high-uncertainty features are fused through multiple levels of feature decoding.
[0030] The beneficial effect is that the low-uncertainty features can be fused by the MPFB module to fully integrate reliable information from different perspectives (axial, coronal and sagittal), thereby generating a more accurate and robust feature representation, which is crucial for enhancing the model's ability to accurately segment brain tumor boundaries. When processing high-uncertainty features, through multi-level feature decoding fusion, feature information from different levels can be effectively integrated, so that high-level semantic information and low-level spatial details complement each other, and the feature representation is gradually optimized. This feature fusion strategy not only improves the model's efficiency in utilizing features, but also enhances the model's robustness to complex backgrounds and noise, thereby improving the accuracy and reliability of brain tumor image segmentation as a whole. This step provides a high-quality feature basis for the final segmentation result through reasonable feature processing and fusion, and is a key link in achieving high-precision brain tumor image segmentation.
[0031] Furthermore, in S5, the cross entropy fusion Dice loss function is combined with the uncertainty loss function to ensure the segmentation result and pay attention to the uncertainty of the model prediction. The specific reasoning process is:
[0032]
[0033] in is the result of cross entropy fusion Dice loss function, is the definition of the uncertain loss function, is the result of the joint loss function, and α, β, γ, and δ are the hyperparameters for balancing the components.
[0034] The beneficial effect is that the cross entropy fusion Dice loss function can ensure the sensitivity of the segmentation results to boundaries and details in different sections (axial, coronal and sagittal), thereby improving the accuracy of the segmentation results. At the same time, the uncertainty loss function quantifies the uncertainty of the segmentation results, allowing the model to pay more attention to those areas with higher uncertainty, thereby further optimizing the segmentation results. The application of this joint loss function not only fully considers the spatial consistency of the segmentation results, but also fully considers the uncertainty of the model prediction, making the evaluation process more comprehensive and objective. In this way, the model can learn key features more effectively during the training process, improve the segmentation accuracy of blurred boundaries and complex background structures, and thus improve the performance and reliability of brain tumor image segmentation as a whole. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 Schematic diagram of the overall architecture of an embodiment of a brain tumor segmentation algorithm based on slice anisotropy CASVim and uncertainty gated UCG of the present invention;
[0036] Figure 2This is a network architecture diagram of a multi-slice-fused visual Mamba feature extraction module in accordance with an embodiment of a brain tumor segmentation algorithm based on slice anisotropy CASVim and uncertainty gated UCG according to the present invention;
[0037] Figure 3 This is a network architecture diagram of the uncertainty gated UCG module of an embodiment of the brain tumor segmentation algorithm based on section anisotropy CASVim and uncertainty gated UCG of the present invention. DETAILED DESCRIPTION
[0038] The following is further explained in detail through specific implementation methods.
[0039] To address the problem of blurred slice boundaries and complex backgrounds in existing algorithms for 3D images, the present invention provides a fusion of multi-slice feature extraction (CASVim) and uncertainty gated (UCG) modules to improve the segmentation accuracy for blurred slice boundaries and complex background structures, thereby improving the overall performance and reliability of brain tumor image segmentation. The specific solutions of the present invention are as follows:
[0040] S1: Obtain original medical images and perform preprocessing;
[0041] Cropping irrelevant areas in the original 3D brain tumor image can effectively reduce the amount of data, remove background information and unnecessary anatomical structures in the image that are irrelevant to the brain tumor segmentation task, and convert the image data into a distribution with a mean of 0 and a standard deviation of 1.
[0042] S2: Perform feature extraction on the image, using the visual Mamba feature extraction module CASVim (CAS) that integrates multiple slices to capture features of different scales of the image and segment it; re-segment the original 3D image into axial, coronal, and sagittal slices, with the three 2D slices being orthogonal to each other, and using multi-structure decomposition to obtain different spatial features; and use the CAS module for downsampling. Specifically, perform feature extraction on the image, using the visual Mamba feature extraction module CASVim (CAS) that integrates multiple slices to capture features of different scales of the image and segment it; re-segment the original 3D image into axial, coronal, and sagittal slices, with the three 2D slices being orthogonal to each other, and using multi-structure decomposition to obtain different spatial features. The CAS reasoning process here can be summarized as follows:
[0043] F Coronal =f coronal (X),F Sagittal =f sagittal (X)
[0044] F Shallow =f shallow(F Coronal ,F Sagittal )
[0045] F Axial =f axial (X),F Deep =f deep (F Axial )
[0046] Among them F Coronal ∈R C×H×W and F Sagittal ∈R C×H×W The superficial features of the coronal and sagittal sections, F Axial ∈R C×H×W is the deep feature of the axial section, where C represents the number of channels, and H and W represent the height and width, respectively.
[0047] S3: The Uncertainty Gating Module (UCG) works with the lightweight residual module in the skip connection to decode features at different levels, ensuring the quality of features acquired by the model and reducing the large number of parameters required by multi-faceted spatial feature fusion. Channel uncertainty is calculated, and the standard deviation in the channel dimension is used to distinguish high-uncertainty features from low-uncertainty features. The reasoning process can be summarized as follows:
[0048]
[0049] U S(j,w) =Fusion(U Cor(h,w) ,U Sag(h,w) ), U D(h,w) =U Axial(h,w)
[0050] (U low ,U high )=UCG(Reshape(U D(h,w) ,U S(h,w) ))
[0051] Among them, U Cor(h,w) ,U Sag(h,w) ,U Axial(h,w) Represents the uncertainty quantification results of the coronal and sagittal sections and the axial section respectively; U S(h,w) ,U D(h,w) ,U low ,U high Represent the shallow layer, deep layer and the results of distinguishing high uncertainty and low uncertainty respectively; θ (h,w) Represents the average value of all channels at this spatial position.
[0052] S4: While decoding shallow to deep features, low-uncertainty features are input into the multi-view fusion module (MPFB). For high-uncertainty features, multiple levels of feature decoding and fusion are used.
[0053] S5: Use cross entropy fusion Dice loss function combined with uncertainty loss function to ensure segmentation results and pay attention to the uncertainty of model prediction. The specific reasoning process is:
[0054]
[0055] in is the result of cross entropy fusion Dice loss function, is the definition of the uncertain loss function, is the result of the joint loss function, and α, β, γ, and δ are the hyperparameters for balancing the components.
Claims
1. A brain tumor segmentation algorithm based on section anisotropy CASVim and uncertainty gated UCG, characterized by: The following steps are involved: S1: Obtain original medical images and perform preprocessing; S2: Feature extraction is performed on the image using CASVim (CAS), a visual feature extraction module that integrates multiple slices. This module captures and segments the image at different scales. The original 3D image is re-segmented into axial, coronal, and sagittal slices, which are orthogonal to each other. Different spatial features are obtained through multi-structure decomposition. Downsampling is also performed using the CAS module. S3: The uncertainty gating module (UCG) cooperates with the lightweight residual module in the jump connection to decode features at different levels, ensuring the quality of features obtained by the model and reducing the large number of parameters brought by the fusion of multi-faceted spatial features; it calculates channel uncertainty and uses the standard deviation in the channel dimension to distinguish high-uncertainty features from low-uncertainty features. S4: While decoding shallow to deep features, low-uncertainty features are input into the multi-view fusion module (MPFB). For high-uncertainty features, multiple levels of feature decoding and fusion are used. S5: The slice consistency loss function is integrated with the uncertainty loss function to ensure the consistency of segmentation results across different slices and focus on the uncertainty of model prediction. In step S1, irrelevant regions in the original 3D brain tumor image are cropped, and the 3D image is Z-Score normalized (Norm); In step S2, feature extraction is performed on the image using CASVim (CAS), a visual feature extraction module that integrates multiple slices. Features at different scales of the image are captured and segmented. The original 3D image is re-segmented into axial, coronal, and sagittal slices. These three 2D slices are orthogonal to each other, and different spatial features are obtained through multi-structure decomposition.
2. The brain tumor segmentation algorithm based on slice anisotropy CASVim and uncertainty gated UCG according to claim 1 is characterized by: In step S2, the coronal and sagittal section features segmented by the CAS module are input into the superficial processing, and the axial section features are input into the deep processing. CAS is an enhanced feature extraction module for 3D image processing that performs data-dependent global visual context modeling. By segmenting the 3D image in the axial, coronal, and sagittal directions, it can extract features from different perspectives, thereby more comprehensively capturing the structure and semantic information of the 3D image. This multi-perspective feature extraction method helps to better understand the spatial relationships in the image and the three-dimensional morphology of the object. It fuses the coronal and sagittal sections for extraction and inputs them into the shallow features, and extracts the axial section features into the deep features. The feature fusion of the coronal and sagittal sections can provide richer local structural information, which helps the model better understand the details and boundary information of the image at the shallow stage. The axial section feature extraction can pay more attention to the key semantic information of the image, thereby improving the model's ability to understand the complex background of the image, thereby ensuring the subsequent image segmentation effect. The CAS reasoning process here can be summarized as follows: F Coronal =f coronal (X),F Sagittal =f sagittal (X) F Shallow =f shallow (F Coronal ,F Sagittal ) F Axial =f axial (X),F Deep =f deep (F Axial ) Among them F Coronal ∈R C×H×W and F Sagittal ∈R C×H×W The superficial features of the coronal and sagittal sections, F Axial ∈R C ×H×W is the deep feature of the axial section, where C represents the number of channels, and H and W represent the height and width, respectively.
3. The brain tumor segmentation algorithm based on slice anisotropy CASVim and uncertainty gated UCG according to claim 1 is characterized by: In step S3, the UCG module is used to perform deep feature extraction on medical image features of different sections. Its main function is to quantify uncertainty by calculating the standard deviation of the shallow feature fusion of the coronal and sagittal sections and the deep features of the axial section in the channel dimension. The larger the standard deviation, the greater the variation of the feature in the channel dimension and the higher the uncertainty; the smaller the standard deviation, the smaller the variation of the feature in the channel dimension and the lower the uncertainty. By setting a threshold, features are divided into high uncertainty and low uncertainty categories. Based on the quantified channel uncertainty, the feature weights are dynamically adjusted. For high uncertainty features, their weights are reduced to reduce reliance on these uncertain features; for low uncertainty features, their weights are maintained to enhance the utilization of these reliable features. The main reasoning processes of shallow and deep layers can be summarized as follows: IN S(h,w) =Fusion(U Cor(h,w) ,IN Sag(h,w) ),IN D(h,w) =U Axial(h,w) (U low ,U high )=UCG(Reshape(U D(h,w) ,U S(h,w) )) in, U Cor(h,w) ,U Sag(h,w) ,U Axial(h,w) represent the uncertainty quantification results of the coronal and sagittal sections and the axial section, respectively. S(h,w) ,U D(h,w) ,U low ,U high Represent the shallow layer, deep layer and the results of distinguishing high uncertainty and low uncertainty respectively; θ (h,w) Represents the average value of all channels at this spatial position.