MRI brain tumor image segmentation method based on state space model and frequency domain

By introducing state space model and 3D frequency domain fusion module in MRI brain tumor image segmentation, the problems of insufficient attention to important information and high frequency noise are solved, and the tumor segmentation effect with higher accuracy and stability is achieved.

CN120147332APending Publication Date: 2025-06-13DALIAN NATIONALITIES UNIVERSITY
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510086400.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

There are problems of insufficient important information attention and high frequency noise in MRI brain tumor image segmentation, resulting in low segmentation accuracy and accuracy.

Method used

The MRI brain tumor image segmentation method based on state space model and frequency domain is adopted to enhance global modeling capabilities and attention to important information by introducing state space models, and a 3D frequency domain fusion module is introduced into the jump connection to reduce high-frequency noise interference.

Benefits of technology

More precise tumor segmentation is achieved, attention to key areas of brain tumors is improved, and the fusion effect of low-frequency and high-frequency characteristics is optimized, improving the stability and accuracy of segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147332A_ABST
    Figure CN120147332A_ABST
Patent Text Reader

Abstract

The invention provides an MRI brain tumor image segmentation method based on a state space model and a frequency domain, and relates to the field of image processing. The method comprises the following steps: acquiring a brain tumor MRI image to be segmented and preprocessing the brain tumor MRI image; and inputting the preprocessed brain tumor MRI image into a segmentation network based on a state space model and a frequency domain. The network is based on a U-Net architecture, and global context modeling is performed on feature information by introducing a state space model into a bottleneck layer. In order to optimize feature extraction, a dynamic weight mechanism is designed aiming at features input into a state space model, the importance of feature sequences in three directions is dynamically adjusted according to input data, and the attention to a brain tumor key area is enhanced. By adding a 3D frequency domain fusion module in jump connection, high-frequency noise is reduced, and perception of the overall structure of a brain tumor image is enhanced. According to the method, the state space model and the 3D frequency domain fusion module are introduced, and the importance of the feature sequence is dynamically adjusted, so that brain tumor MRI image segmentation with higher precision is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and more particularly, to an MRI brain tumor image segmentation method based on a state space model and the frequency domain. Background Art

[0002] Magnetic Resonance Imaging (MRI) technology has become the main means for diagnosing brain tumors. Compared with Computed Tomography (CT), MRI can non-invasively display the three-dimensional structure of the brain with higher resolution and generate multi-dimensional images through different imaging modalities, providing comprehensive information for the accurate diagnosis of brain tumors. MRI technology can accurately locate tumors, estimate tumor size, and formulate personalized treatment plans.

[0003] However, the application of MRI images also faces many challenges. Due to factors such as bias fields, noise, and differences in brain gray-scale distributions among patients, the quality of MRI images may be unstable. In addition, the complexity of brain tissues and the ambiguity and irregularity of tumor boundaries often result in gray-scale overlap between the tumor region and the surrounding normal tissues, greatly reducing the accuracy of brain tumor segmentation. In clinical practice, manual delineation of the tumor region or using labels to mark the lesion area is usually relied on. However, traditional manual segmentation methods are not only time-consuming and laborious but also easily affected by subjective experience, with differences in the calibration results of the same tumor region reaching up to 28%, and there may also be a 20% deviation in the calibration results at different times. These problems seriously affect the efficiency and accuracy of brain tumor image analysis. Therefore, developing a high-precision and automated MRI segmentation algorithm is of great significance for providing accurate quantitative analysis results and assisting doctors in diagnosis and decision-making.

[0004] In recent years, deep learning techniques have made remarkable progress in the task of brain tumor segmentation, and among them, the automatic segmentation method based on convolutional neural network (CNN) has become the mainstream. The U-shaped structure network represented by U-Net can achieve ideal segmentation results even on small-sample data due to its simple structure and skip connection design. However, CNN has limitations in capturing long-range dependencies in images, which poses a challenge to improving the accuracy of brain tumor segmentation. The Transformer architecture has rapidly expanded from the field of natural language processing to the field of computer vision due to its excellent long-range dependency modeling ability. Compared with CNN, Transformer performs better in modeling global information, but its high computational complexity limits its application in high-resolution medical images. At the same time, the method based on the state space model (SSM) has shown significant advantages in the field of brain tumor segmentation. As the first attempt to apply SSM to medical image segmentation, SegMamba can effectively capture long-range dependencies and exhibit excellent computational performance, setting a new technical benchmark for brain tumor segmentation. Swin-UMamba significantly improves the segmentation performance and reduces the consumption of computing resources by combining the ImageNet pre-trained model. VM-Unet combines the Mamba and U-Net architectures and designs the visual state space (VSS) block, greatly enhancing the ability to model the details of medical images. SSM breaks through the limitation of the local receptive field and can learn long-range dependencies globally, performing well in processing global features and significantly improving the computational efficiency. However, due to SSM emphasizing global information, it is often difficult to fully focus on local or strongly directional important information when processing three-dimensional images, especially when scanning sequences in different directions, thus having certain limitations. Summary of the Invention

[0005] According to the above-mentioned technical problems that the existing networks pay insufficient attention to important information and have high-frequency noise in brain tumor MRI images, a method for segmenting MRI brain tumor images based on the state space model and frequency domain is provided. The present invention is based on a brain tumor segmentation network, which introduces a state space model to improve the global modeling ability and the attention to important information, and realizes more accurate tumor segmentation; a 3D frequency domain fusion module is introduced in the skip connection to reduce the interference of high-frequency noise through frequency domain information.

[0006] The technical means adopted by the present invention are as follows:

[0007] A method for segmenting MRI brain tumor images based on the state space model and frequency domain, comprising:

[0008] S1. Collect multi-modal MR images of brain tumors to form a brain tumor image dataset, and preprocess the brain tumor image dataset;

[0009] S2. Based on the state space model and frequency domain, construct a brain tumor MRI segmentation network model;

[0010] S3. Use the preprocessed brain tumor image dataset to train the constructed brain tumor MRI segmentation network model and save the weights;

[0011] S4. Load the weights of the trained brain tumor MRI segmentation network model, perform inference and prediction on the test set, obtain the brain tumor segmentation results, and generate the final segmentation results through post-processing.

[0012] Further, step S1 specifically includes:

[0013] S11. Assume that the brain tumor image dataset includes MR images of four modalities, label the dataset into four types: background and healthy part, gangrene part, edema part, and enhanced tumor part, and crop the useless background information in the images;

[0014] S12. Use the Min-Max Normalization method to perform normalization operations on each modality image to reduce the contrast difference between different modalities;

[0015] S13. Create a new channel, perform one-hot encoding on the foreground voxels (tumor regions), and distinguish the background voxels (non-tumor regions) and the voxels with normalized values close to 0;

[0016] S14. Divide the brain tumor image dataset into a training set and a test set according to a certain ratio, and use two data augmentation methods, random intensity variation and random flipping, to perform data augmentation on the training set.

[0017] Further, in step S2, the constructed brain tumor MRI segmentation network model is based on the 4-layer original U-Net encoder structure, and also includes a state space model module and a 3D frequency domain fusion module, where:

[0018] Each layer of the encoder uses 2-layer convolutional blocks to extract brain tumor feature information and uses one convolutional layer for downsampling;

[0019] The state space model is set at the bottleneck layer to enhance the network's feature extraction ability and the ability to model long-range dependencies;

[0020] The 3D frequency domain fusion module is set in the lateral connection to reduce the interference of high-frequency noise and effectively fuse high-frequency and low-frequency information by introducing the fast Fourier transform.

[0021] Further, the state space model module enhances the attention to key regions through a weighting mechanism, enhances the network's ability to process complex tumor morphologies, and consists of multiple dynamic weighted Mamba modules; where:

[0022] Each dynamic weighted Mamba module includes a gated spatial convolutional layer, a layer normalization layer, a dynamic weighted tri-directional Mamba layer, and a multi-layer perceptron; the gated spatial convolutional layer, the layer normalization layer, the dynamic weighted tri-directional Mamba layer, and the multi-layer perceptron are connected in sequence and are integrally connected through residual connections.

[0023] Further, the gated spatial convolutional layer includes a plurality of normalization convolutional modules and residual units, where:

[0024] There are three normalization convolutional modules, each of which consists of a normalization layer, a convolutional kernel, and a non-linear activation layer. The convolutional kernels are two 3×3 convolutional kernels and one 1×1 convolutional kernel; the plurality of normalization convolutional modules are integrally connected through residual connections.

[0025] Further, the dynamic weighted tri-directional Mamba layer sequentially includes a flattening unit, a layer normalization layer, a state space model (SSM), a dynamic weighted fusion module, and a linear reshaping unit, where:

[0026] The flattening unit flattens the dimension of the input feature map from (B, C, H, W), converts (H, W) into the sequence length L, and transposes it into the form of (B, L, C) (where L = H×W);

[0027] The layer normalization layer normalizes the features;

[0028] Extract feature sequences through initial projection and one-dimensional convolution, perform non-linear processing using the SiLU activation function, and capture temporal dependencies through the state space model to calculate dynamic time steps;

[0029] Perform tri-directional feature processing on the forward time sequence, reverse sequence, and spatial sequence through convolutional layers and linear layers respectively;

[0030] The dynamic weighted fusion module performs weighted fusion on the tri-directional features to generate comprehensive features;

[0031] The linear reshaping unit rearranges the fused features from (B, L, C) to (B, C, H, W) to generate a spatial feature map processed through the time sequence idea for the image segmentation task.

[0032] Further, the dynamic weighted fusion module fully extracts global information and local information in 3D data through forward, reverse, and spatial feature combination methods; after each group of feature combinations, it passes through a gated layer to learn dynamic weights; after each gated layer, the Mish activation function is applied to increase non-linear expression ability and help the model capture complex relationships.

[0033] Further, the 3D frequency domain fusion module includes a point convolution layer, an activation function, a frequency domain conversion layer, a convolution layer, and a point convolution layer connected in sequence, which is used to reduce high-frequency noise interference, enhance the perception of the overall structure of brain tumor images, and optimize the information transmission and fusion effect of skip connections.

[0034] Further, step S3 specifically includes:

[0035] S31. Set the hyperparameters of the network, select the Adam optimizer as the optimization tool, adopt the cosine annealing strategy to dynamically adjust the learning rate, use automatic mixed precision to reduce the computational amount, save the optimal training weights, and during the verification process, adopt the test-time augmentation and voxel clipping post-processing methods for the prediction results;

[0036] S32. Input the training set into the MRI brain tumor image segmentation model based on the state space model and frequency domain for training, use the Dice loss function as the main loss function of the network model, and the KL divergence loss as the auxiliary loss function, and take the three lesion regions of all tumors, tumor cores, and enhanced tumors as the main segmentation targets.

[0037] Further, step S4 specifically includes:

[0038] S41. Load the optimal model weights saved during the training process, and predict the segmentation results of all tumors, tumor cores, and enhanced tumors in the test set;

[0039] S42. Apply the test-time augmentation (TTA) method to flip the input brain tumor image along the x-axis, y-axis, z-axis, xy-axis, xz-axis, yz-axis, and xyz-axis, and make predictions respectively, and take the average of the prediction results in each direction to improve the stability and accuracy of the segmentation;

[0040] S43. Use the voxel clipping method to post-process the prediction results, and replace the prediction results with the corresponding labels with a certain probability to generate the final segmentation results.

[0041] Compared with the prior art, the present invention has the following advantages:

[0042] 1. A method for segmenting MRI brain tumor images based on a state space model and frequency domain provided by the present invention is based on the U-Net architecture, introduces a state space model module in the bottleneck layer, performs global context modeling on the features extracted by the encoder, enhances the network's ability to model long-range dependencies and attention to important information, and realizes more accurate tumor segmentation.

[0043] 2. A method for segmenting MRI brain tumor images based on a state space model and frequency domain provided by the present invention designs a dynamic weight mechanism for the features input into the state space model to further optimize feature extraction, dynamically adjusts the importance of the three-direction feature sequences according to the input data, enhances the attention to the key regions of the brain tumor; and introduces a 3D frequency domain fusion module in the skip connection to reduce the interference of high-frequency noise and optimize the fusion effect of low-frequency and high-frequency features.

[0044] 3. A method for segmenting MRI brain tumor images based on a state space model and frequency domain provided by the present invention realizes higher-precision segmentation of MRI brain tumor images by introducing a state space model, a 3D frequency domain fusion module, and dynamically adjusting the importance of the feature sequences.

[0045] For the above reasons, the present invention can be widely promoted in the fields of image processing and the like. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0047] Figure 1 It is a flowchart of the method of the present invention.

[0048] Figure 2 It is a structural diagram of the MRI brain tumor image segmentation model provided by the embodiment of the present invention.

[0049] Figure 3 It is a structural diagram of the gated spatial convolutional layer provided by the embodiment of the present invention.

[0050] Figure 4 It is a structural diagram of the dynamically weighted three-way Mamba layer provided by the embodiment of the present invention.

[0051] Figure 5 It is a structural diagram of the dynamically weighted fusion module provided by the embodiment of the present invention.

[0052] Figure 6 It is a structural diagram of the 3D frequency domain fusion module provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0054] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0055] As Figure 1 shown, the present invention provides an MRI brain tumor image segmentation method based on a state space model and frequency domain, including:

[0056] S1. Collect multi-modal MR images of brain tumors to form a brain tumor image data set, and preprocess the brain tumor image data set;

[0057] S2. Based on the state space model and frequency domain, construct a brain tumor MRI segmentation network model;

[0058] S3. Use the preprocessed brain tumor image data set to train the constructed brain tumor MRI segmentation network model and save the weights;

[0059] S4. Load the weights of the trained brain tumor MRI segmentation network model, perform inference and prediction on the test set, obtain the brain tumor segmentation result, and generate the final segmentation result through post-processing.

[0060] Specifically, as a preferred implementation manner of the present invention, step S1 specifically includes:

[0061] S11. Assume that the brain tumor image dataset includes MR images of four modalities. Label the dataset into four types: background and healthy parts, gangrenous parts, edematous parts, and enhanced tumor parts. Crop the useless background information in the images. In this embodiment, the brain tumor image dataset uses the BraTS2020 dataset, which has MR images of four modalities. The size of each modality image is 240mm×240mm×155mm. Experts manually label the dataset, where the background and healthy parts are labeled 0, gangrene is labeled 1, edema is labeled 2, and enhanced tumor is labeled 4. Since there is a large amount of redundant background information in the images, these useless background information are cropped first.

[0062] S12. Use the Min-Max Normalization method to perform normalization operations on each modality image to reduce the contrast difference between different modalities. The formula is as follows:

[0063]

[0064] Among them, X represents the voxel value of the input MR image, X norm represents the voxel value of the normalized MR image, X min represents the maximum voxel value of the input MR image, X max represents the minimum voxel value of the input MR image;

[0065] S13. Create a new channel, perform one-hot encoding on the foreground voxels (tumor regions), and distinguish the background voxels (non-tumor regions) and the voxels with normalized values close to 0. In this embodiment, the size of the image input to the network during the training phase is 4×128×128×128.

[0066] S14. Divide the brain tumor image dataset into a training set and a test set according to a certain ratio, and use two data augmentation methods, random intensity variation and random flipping, to perform data augmentation processing on the training set.

[0067] Specifically, as a preferred embodiment of the present invention, in step S2, as Figure 2 shown, the constructed brain tumor MRI segmentation network model is based on the 4-layer original U-Net encoder structure, and also includes a state space model module and a 3D frequency domain fusion module, where:

[0068] Each layer of the encoder uses 2-layer convolutional blocks to extract brain tumor feature information and uses one convolutional layer for downsampling;

[0069] The state space model is set at the bottleneck layer to enhance the network's feature extraction ability and the ability to model long-range dependencies;

[0070] The 3D frequency domain fusion module is set in the horizontal connection to reduce the interference of high-frequency noise and effectively fuse high-frequency and low-frequency information by introducing the fast Fourier transform.

[0071] In specific implementation, as a preferred implementation manner of the present invention, the state space model module integrates multi-scale information at the spatial and semantic levels of the brain tumor image to enhance the perception ability of brain tumor features and is composed of multiple dynamic weighted Mamba modules; where:

[0072] Each dynamic weighted Mamba module includes a gated spatial convolutional layer, a layer normalization layer, a dynamic weighted three-way Mamba layer, and a multi-layer perceptron; the gated spatial convolutional layer, the layer normalization layer, the dynamic weighted three-way Mamba layer, and the multi-layer perceptron are connected in sequence, and the overall connection is achieved through a residual connection.

[0073] In specific implementation, as a preferred implementation manner of the present invention, as Figure 3 shown, the gated spatial convolutional layer includes multiple normalized convolutional modules and residual units, where:

[0074] There are three normalized convolutional modules, each of which is composed of a normalization layer, a convolutional kernel, and a non-linear activation layer. The convolutional kernels are two 3×3 convolutional kernels and one 1×1 convolutional kernel; the multiple normalized convolutional modules are connected as a whole through a residual connection.

[0075] In this embodiment, the gated spatial convolutional module combines multi-path feature extraction with a gated mechanism to achieve multi-scale fusion and enhancement of 3D features. The input features are first processed through two feature extraction paths respectively: in the first path, the features are first standardized through a normalization layer, then local spatial features are extracted through a 3×3 convolution, and the expression ability is further enhanced through a non-linear activation function; in the second path, after the features are also normalized, channel features are captured through a 1×1 convolution, and the feature expression is further optimized through a Relu activation function. The features extracted by the two paths are then adaptively fused through element-wise multiplication, and this gated mechanism adaptively adjusts the importance of the features of the two paths according to the input features. The fused features are then further subjected to deep feature extraction through a 3×3 convolution, and the expression ability is strengthened through normalization and a Relu activation function. The original input features are added to the processed features element-wise. While retaining the original information through a residual connection, the expression ability of the deep features is improved, and the final output features with multi-scale information are generated. This design effectively improves the ability to capture global and local features.

[0076] In specific implementation, as a preferred implementation manner of the present invention, by normalizing the value of each data in the feature dimension, the feature distribution can be stabilized, and the problems of gradient disappearance or gradient explosion can be alleviated. Layer normalization calculates the mean and variance of the features of each data, normalizes the features to zero mean and unit variance, and scales and translates the normalized features through learnable parameters. The formula is as follows:

[0077]

[0078] Where: h′ i represents the feature after layer normalization, h i represents the input feature, μ represents the feature mean, σ 2 represents the feature variance, γ and β represent learnable parameters, which are used for scaling and translation operations respectively. ∈ represents a very small positive number, which is used to avoid the denominator being zero.

[0079] In specific implementation, as a preferred implementation manner of the present invention, as Figure 4 shown, the dynamic weighted three-way Mamba layer sequentially includes a flattening unit, a layer normalization layer, a state space model (SSM), a dynamic weighted fusion module, and a linear reshaping unit, where:

[0080] The flattening unit flattens the dimension of the input feature map from (B, C, H, W), converts (H, W) into the sequence length L, and transposes it into the form of (B, L, C) (where L = H×W);

[0081] The layer normalization layer normalizes the features;

[0082] Extract the feature sequence through initial projection and one-dimensional convolution, perform non-linear processing using the SiLU activation function, and capture the time dependence through the state space model to calculate the dynamic time step. The state update formula is as follows:

[0083] h t+1 = A·h t + B·X t

[0084] Where, A represents the state transition equation, B represents the input projection matrix, X t is the input feature at time step t, X t = C·h t ;

[0085] Perform three-way feature processing on the forward time series, reverse sequence, and spatial sequence through the convolutional layer and the linear layer respectively;

[0086] The dynamic weighted fusion module performs weighted fusion on the three-way features to generate comprehensive features;

[0087] The linear reshaping unit rearranges the fused features from (B, L, C) to (B, C, H, W), generating a spatial feature map processed by the time series idea for the image segmentation task.

[0088] In specific implementation, as a preferred implementation manner of the present invention, as Figure 5 shown, the dynamic weighted fusion module fully extracts the global information and local information in the 3D data through forward, backward, and spatial feature combination methods; where:

[0089] After each group of feature combinations passes through the gating layer to learn the dynamic weights; as follows:

[0090] W forward = σ(W f ·X forward + b f )

[0091] W reverse = σ(W r ·X reverse + b r )

[0092] W spatial = σ(W s ·X spatial + b r )

[0093] Among them, W f , W r , W s represent the weight matrices of the gating layer, b f , b r , b s represent the bias terms, and σ represents the activation function, which is used to constrain the range of the dynamic weights.

[0094] After each group of features is weighted by the corresponding dynamic weights, the Mish activation function is applied to increase the non-linear expression ability and help the model capture complex relationships. The formula is as follows:

[0095] X′ forward = Mish(W forward ⊙X forward )

[0096] X′ reverse = Mish(W reverse ⊙X reverse )

[0097] X′ spatial = Mish(W spatial ⊙X spatial )

[0098] Among them, the formula of the Mish activation function is as follows:

[0099] Mish(x) = x · tanh(ln(1 + e x ))

[0100] The non-linearly processed features are fused in a dynamic weighting manner to generate a comprehensive feature, and the formula is as follows:

[0101] X fuesd = α · X' forward + β · X' reverse + γ · X' spatial

[0102] Wherein, X fuesd represents the final fused feature output, and α, β, and γ represent the importance weights dynamically adjusted according to the input features.

[0103] In specific implementation, as a preferred implementation manner of the present invention, as Figure 6 shown, the 3D frequency domain fusion module includes a point convolution layer, an activation function, a frequency domain conversion layer, a convolution layer, and a point convolution layer connected in sequence, and is used to simultaneously enhance local details and global feature modeling, and optimize the information transmission and fusion effect of the skip connection. The specific process is as follows:

[0104] First, the input 3D feature X is expanded in the channel dimension through a point-wise convolution to obtain X pw1 ; then the expanded feature introduces non-linear capabilities through a custom activation function to obtain X act1 , and the formula is as follows:

[0105] X act1 = s · ReLU(x) 2 + b

[0106] Wherein, s represents a learnable scaling parameter, and b represents a learnable bias parameter. ReLU(x) = max(0, x), representing the standard ReLU activation function.

[0107] Subsequently, to capture global information, the feature X act1 is mapped to the frequency domain, and is converted to the frequency domain through a 3D fast Fourier transform, and the high-frequency components are suppressed and the low-frequency components are enhanced according to the high-frequency ratio parameter. Among them, the fast Fourier transform formula is as follows:

[0108]

[0109] Specifically, the high-frequency components are multiplied by 0.3, and the low-frequency components are multiplied by 0.7. This process can be expressed by the following formula:

[0110] X'(k) = M(k) · X(k)

[0111] Among them, X(k) is the original frequency-domain feature after fast Fourier transform, X′(k) is the frequency-domain feature after mask processing, and M(k) is the mask function assigned according to the frequency region (high frequency or low frequency). The definition of M(k) is as follows:

[0112]

[0113] Among them, k represents a certain position in the frequency domain, and M(k) assigns corresponding weights according to the frequency region where k is located.

[0114] Then, the inverse fast Fourier transform is used to restore the frequency-domain feature X′(k) after mask processing to the spatial domain. The formula for the inverse fast Fourier transform is as follows:

[0115]

[0116] After frequency-domain modeling, the features are successively passed through a non-linear activation function, a convolutional layer, and a point convolutional layer to obtain the output 3D features, which fuse global and local features, significantly improve the feature expression ability, and achieve more efficient information transmission and fusion in skip connections.

[0117] In specific implementation, as a preferred implementation manner of the present invention, step S3 specifically includes:

[0118] S31. Set the hyperparameters of the network, select the Adam optimizer as the optimization tool, adopt the cosine annealing strategy to dynamically adjust the learning rate, use automatic mixed precision to reduce the computational amount, save the optimal training weights, and in the verification process, use test-time augmentation and voxel cropping post-processing methods for the prediction results;

[0119] S32. Input the training set into the MRI brain tumor image segmentation model based on the state space model and frequency domain for training, use the Dice loss function as the main loss function of the network model, and the KL divergence loss and mean square error loss as auxiliary loss functions, and use the three lesion regions of all tumors, tumor cores, and enhanced tumors as the main segmentation targets.

[0120] In specific implementation, as a preferred implementation manner of the present invention, step S4 specifically includes:

[0121] S41. Load the optimal model weights saved during the training process, and predict the segmentation results of all tumors, tumor cores, and enhanced tumors in the test set;

[0122] S42. Apply the Test Time Augmentation (TTA) method to flip the input brain tumor image along the x-axis, y-axis, z-axis, xy-axis, xz-axis, yz-axis, and xyz-axis, and perform predictions separately. Take the average of the prediction results in each direction to improve the stability and accuracy of segmentation;

[0123] S43. Use the voxel clipping method to post-process the prediction results. Replace the prediction results with the corresponding labels with a certain probability to generate the final segmentation result. In this embodiment, specifically: replace the voxels belonging to all tumors with label 0 with a certain probability, replace the voxels belonging to the tumor core with label 2 with a certain probability, and replace the voxels belonging to the enhanced tumor with label 1 with a certain probability. In addition, for independent enhanced tumor blocks and overall enhanced tumor blocks with the number of voxels and average probability lower than a certain threshold, uniformly replace their labels with 1 to further optimize the accuracy of the segmentation result.

[0124] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for segmenting MRI brain tumor images based on state space model and frequency domain, characterized in that: include: S1, collecting multimodal MR images of brain tumors to form a brain tumor image dataset, and preprocessing the brain tumor image dataset; S2. Construct a brain tumor MRI segmentation network model based on the state space model and frequency domain; S3, using the preprocessed brain tumor image dataset to train the constructed brain tumor MRI segmentation network model and save the weights; S4. Load the trained brain tumor MRI segmentation network model weights, perform inference and prediction on the test set, obtain the brain tumor segmentation results, and generate the final segmentation results through post-processing.

2. The MRI brain tumor image segmentation method based on state space model and frequency domain according to claim 1, characterized in that: Step S1 specifically includes: S11. Assume that the brain tumor image dataset includes MR images of four modalities. The dataset is annotated into four types: background and healthy part, gangrene part, edema part and enhanced tumor part. The useless background information in the image is trimmed. S12, using the Min-Max Normalization method to normalize each modality image to reduce the difference in contrast between different modalities; S13, create a new channel, perform one-hot encoding on the foreground voxels, and distinguish between background voxels and voxels whose values ​​are close to 0 after normalization; S14. Divide the brain tumor image dataset into a training set and a test set according to a certain ratio, and use two data enhancement methods, random intensity variation and random flipping, to perform data enhancement on the training set.

3. The MRI brain tumor image segmentation method based on state space model and frequency domain according to claim 1, characterized in that: In step S2, the constructed brain tumor MRI segmentation network model is based on the original 4-layer U-Net encoder structure, and also includes a state space model module and a 3D frequency domain fusion module, where: Each encoder layer uses 2 layers of convolutional blocks to extract brain tumor feature information and uses one convolutional layer for downsampling; The state space model is set at the bottleneck layer to enhance the network's feature extraction capabilities and ability to model long-range dependencies; The 3D frequency domain fusion module is set in the lateral connection to reduce the interference of high-frequency noise and effectively fuse high-frequency and low-frequency information by introducing fast Fourier transform.

4. The MRI brain tumor image segmentation method based on state space model and frequency domain according to claim 3, characterized in that: The state-space model module enhances the focus on key areas through a weighted mechanism, and enhances the network's ability to handle complex tumor morphology, and is composed of multiple dynamically weighted Mamba modules; wherein: Each dynamic weighted Mamba module includes a gated spatial convolution layer, a layer normalization layer, a dynamic weighted three-way Mamba layer and a multi-layer perceptron; the gated spatial convolution layer, the layer normalization layer, the dynamic weighted three-way Mamba layer and the multi-layer perceptron are connected in sequence and are overall connected through residual connections.

5. The MRI brain tumor image segmentation method based on state space model and frequency domain according to claim 4, characterized in that: The gated spatial convolution layer includes a plurality of normalized convolution modules and residual units, wherein: There are three normalized convolution modules, each consisting of a normalization layer, a convolution kernel, and a nonlinear activation layer. The convolution kernels are two 3×3 convolution kernels and one 1×1 convolution kernel. Multiple normalized convolution modules are connected as a whole through residual connections.

6. The MRI brain tumor image segmentation method based on state space model and frequency domain according to claim 4, characterized in that: The dynamic weighted three-way Mamba layer includes a flattening unit, a layer normalization layer, a state space model, a dynamic weighted fusion module and a linear reshaping unit in sequence, wherein: The flattening unit flattens the dimension of the input feature map from (B, C, H, W), converts (H, W) to a sequence length of L, and transposes it to the form of (B, L, C) (where L = H × W); The layer normalization layer normalizes the features; The feature sequence is extracted by initialization projection and one-dimensional convolution, nonlinear processing is performed using SiLU activation function, and the time dependency is captured by the state space model to calculate the dynamic time step; The three-way feature processing of the forward time series, reverse series and spatial series is performed through the convolution layer and the linear layer respectively; The dynamic weighted fusion module performs weighted fusion on the three-way features to generate comprehensive features; The linear reshaping unit rearranges the fused features from (B, L, C) to (B, C, H, W), generating a spatial feature map processed by the time series idea for image segmentation tasks.

7. The MRI brain tumor image segmentation method based on state space model and frequency domain according to claim 6, characterized in that: The dynamic weighted fusion module fully extracts global and local information from 3D data through forward, backward and spatial feature combinations; each set of features is combined through a gating layer to learn dynamic weights; and a Mish activation function is applied after each gating layer to increase nonlinear expression capabilities and help the model capture complex relationships.

8. The MRI brain tumor image segmentation method based on state space model and frequency domain according to claim 3, characterized in that: The 3D frequency domain fusion module includes a point convolution layer, an activation function, a frequency domain conversion layer, a convolution layer, and a point convolution layer connected in sequence, which is used to reduce high-frequency noise interference, enhance the perception of the overall structure of the brain tumor image, and optimize the information transmission and fusion effect of the jump connection.

9. The MRI brain tumor image segmentation method based on state space model and frequency domain according to claim 3, characterized in that: Step S3 specifically includes: S31, set the hyperparameters of the network, select Adam optimizer as the optimization tool, use cosine annealing strategy to dynamically adjust the learning rate, use automatic mixed precision to reduce the amount of calculation, save the optimal training weights, and use test time enhancement and voxel clipping post-processing methods for the prediction results during the verification process; S32. The training set is input into the MRI brain tumor image segmentation model based on the state space model and the frequency domain for training. The Dice loss function is used as the main loss function of the network model, the KL divergence loss is used as the auxiliary loss function, and the three lesion areas of the entire tumor, the tumor core and the enhanced tumor are used as the main segmentation targets.

10. The MRI brain tumor image segmentation method based on state space model and frequency domain according to claim 3, characterized in that: Step S4 specifically includes: S41, loading the optimal model weights saved during the training process, and predicting the segmentation results of all tumors, tumor cores, and enhanced tumors in the test set; S42, applying a test time enhancement method, flipping the input brain tumor image along the x-axis, y-axis, z-axis, xy-axis, xz-axis, yz-axis and xyz-axis, and performing predictions respectively, and averaging the prediction results in each direction to improve the stability and accuracy of the segmentation; S43. Use the voxel clipping method to post-process the prediction results, replace the prediction results with corresponding labels according to a certain probability, and generate the final segmentation results.

Citation Information

Cited By

  • Remote sensing image building extraction method and system based on visual Mama model

    CN120599504A

  • Colorectal cancer pathological image segmentation method and device based on frequency domain characteristics and readable storage medium thereof

    CN120635104A

  • Colorectal cancer pathological image segmentation method and device based on frequency domain features and readable storage medium thereof

    CN120635104B

  • Remote sensing image ground object segmentation method based on Fourier space-channel interaction

    CN120765932A

  • Remote sensing image segmentation method and device

    CN120997496A