Multi-modal brain tumor segmentation method and system based on channel adaptive structure feedback
By introducing a multimodal brain tumor segmentation method with multimodal pyramid feature encoding, spatial shift attention and channel adaptive fusion modules, the difficulties of traditional methods in balancing global context modeling and retaining local details of weak tumors are solved, and high-precision segmentation of multimodal brain tumor images is achieved.
Patent Information
- Application Number
- CN202510877559.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-24
AI Technical Summary
Existing brain tumor segmentation methods have difficulty balancing global context modeling and retaining local details of weak tumors. In particular, when tumor structures in multimodal brain tumor images exhibit different scales and fuzzy boundaries, traditional convolution operations are constrained by a fixed receptive field, resulting in insufficient segmentation accuracy.
A multimodal brain tumor segmentation method based on channel-adaptive structural feedback is adopted. By introducing a multimodal pyramid feature encoding module, a spatial shift attention mechanism module and a channel-adaptive fusion module, a multimodal brain tumor segmentation model is constructed. The multi-path convolution structure is used to capture boundary, local and global features, and the channel-adaptive fusion module is used to realize the dynamic fusion of structural and semantic information, and finally image segmentation is performed.
It effectively improves the segmentation accuracy of multimodal brain tumor images, especially the segmentation ability of weak tumor boundary information, and improves the segmentation effect.
Smart Images

Figure CN120833345A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical image processing, in particular to a multi-modal brain tumor segmentation method and system based on channel adaptive structure feedback. BACKGROUND
[0002] Brain tumor is an abnormal proliferation of brain tissue, and malignant tumor poses a serious threat to human health. Therefore, accurate segmentation of brain tumor is of great significance for clinicians to analyze brain tumor conditions and clinical application research. Due to the limitations of MRI imaging technology, there are often noise and artifacts in the image, and the tumor boundary is not clear and fuzzy, which poses a great challenge to accurate tumor segmentation. Multi-modal brain tumor segmentation technology can effectively segment tumor shape, size and location and other information, and has developed rapidly in recent years and is mainly applied to clinical diagnosis, treatment planning, surgical assistance, case study and functional image data processing.
[0003] The brain tumor segmentation method can be generally divided into a deep learning (DL) based method and a traditional image segmentation algorithm. The DL based method is widely used in the field of medical image segmentation and has achieved good results. For example, an existing method proposes to extend convolution to obtain image features from different receptive fields, which can obtain more global information but ignores part of the local features. Or a 3D multi-ray unit is proposed, that is, a group of lightweight 3D CNN is used to solve dense volume segmentation. Further, it emphasizes capturing sparse spatial dynamic correlation and multi-level edge features in 3D volume data, and has stronger modeling ability for spatial boundary information. For the combination of an improved Transformer module and a U-shaped structure, a multi-head attention module and an augmented shortcut are embedded in the encoder and the bottleneck layer, respectively, which significantly improves the global feature expression ability. Although existing research has made great progress in brain tumor segmentation, due to the different scales and fuzzy boundaries of tumor structures, the traditional convolution operation is constrained by a fixed receptive field, and it still faces difficulties in balancing global context modeling and preserving small tumor local details. SUMMARY
[0004] To solve the above technical problems, the purpose of the present application is to provide a multi-modal brain tumor segmentation method and system based on channel adaptive structure feedback, which can effectively segment multi-modal brain tumor images with weak and small tumor boundary information and improve the segmentation accuracy of multi-modal brain tumor images.
[0005] The first technical solution adopted by the present application is: a multi-modal brain tumor segmentation method based on channel adaptive structure feedback, comprising the following steps:
[0006] Obtain multi-modal brain tumor image data and perform three-dimensional data preprocessing to obtain preprocessed multi-modal brain tumor image data;
[0007] The multi-modal pyramid feature encoding module, the spatial shift attention mechanism module and the channel adaptive fusion module are introduced to construct the multi-modal brain tumor segmentation model.
[0008] Based on the multi-modal brain tumor segmentation model, the preprocessed multi-modal brain tumor image data is subjected to image segmentation processing to obtain segmented brain tumor images.
[0009] Further, the step of obtaining multi-modal brain tumor image data and performing three-dimensional data preprocessing to obtain preprocessed multi-modal brain tumor image data specifically includes:
[0010] The preset modal brain tumor image data and the corresponding label image are stacked to obtain multi-modal brain tumor image data, and the preset modal includes T1, T1ce, T2 and Flair.
[0011] The multi-modal brain tumor image data is sequentially subjected to data normalization processing and image enhancement processing to obtain preliminary preprocessed multi-modal brain tumor image data.
[0012] The preliminary preprocessed multi-modal brain tumor image data is subjected to random cropping and center cropping processing and image tensor conversion processing to obtain preprocessed multi-modal brain tumor image data.
[0013] Further, the multi-modal brain tumor segmentation model specifically includes a multi-modal pyramid feature encoding module, a spatial shift attention mechanism module, a channel adaptive fusion module, a detail feedback mechanism module, a decoding module and a convolution module, the multi-modal pyramid feature encoding module includes a first pyramid feature extraction module, a second pyramid feature extraction module, a third pyramid feature extraction module and a fourth pyramid feature extraction module, the spatial shift attention mechanism module includes a first spatial shift attention mechanism, a second spatial shift attention mechanism, a third spatial shift attention mechanism and a fourth spatial shift attention mechanism, the detail feedback mechanism module includes a first detail feedback mechanism and a second detail feedback mechanism, and the decoding module includes a first decoding layer, a second decoding layer and a third decoding layer.
[0014] Further, the step of performing image segmentation processing on the preprocessed multi-modal brain tumor image data based on the multi-modal brain tumor segmentation model to obtain segmented brain tumor images specifically includes:
[0015] The preprocessed multi-modal brain tumor image data is input into the multi-modal brain tumor segmentation model.
[0016] The multi-modal pyramid feature encoding module based on the multi-modal brain tumor segmentation model performs cross-scale semantic feature extraction processing on the preprocessed multi-modal brain tumor image data to obtain multi-modal brain tumor cross-scale semantic features.
[0017] The spatial shift attention mechanism module based on the multi-modal brain tumor segmentation model performs feature enhancement processing on the multi-modal brain tumor cross-scale semantic features to obtain enhanced multi-modal brain tumor cross-scale semantic features.
[0018] The channel adaptive fusion module based on the multi-modal brain tumor segmentation model performs feature fusion processing on the enhanced multi-modal brain tumor cross-scale semantic features to obtain fused multi-modal brain tumor cross-scale semantic features.
[0019] The detail feedback mechanism module based on the multi-modal brain tumor segmentation model performs feature optimization processing on the fused multi-modal brain tumor cross-scale semantic features and the enhanced multi-modal brain tumor cross-scale semantic features to obtain optimized multi-modal brain tumor cross-scale semantic features.
[0020] The decoding module and the convolution module based on the multi-modal brain tumor segmentation model perform multi-modal information segmentation prediction on the optimized multi-modal brain tumor cross-scale semantic features to obtain segmented brain tumor images.
[0021] Further, the multi-modal pyramid feature encoding module based on the multi-modal brain tumor segmentation model performs cross-scale semantic feature extraction processing on the preprocessed multi-modal brain tumor image data to obtain multi-modal brain tumor cross-scale semantic features, which specifically includes:
[0022] The preprocessed multi-modal brain tumor image data is input into the multi-modal pyramid feature encoding module of the multi-modal brain tumor segmentation model;
[0023] The first convolution kernel based on the multi-modal pyramid feature encoding module performs compression channel dimension and non-linear transformation processing on the preprocessed multi-modal brain tumor image data to obtain first multi-modal brain tumor features;
[0024] The second convolution kernel based on the multi-modal pyramid feature encoding module extracts tumor edge structure features from the preprocessed multi-modal brain tumor image data to obtain second multi-modal brain tumor features;
[0025] The third convolution kernel based on the multi-modal pyramid feature encoding module performs global context information extraction processing on the preprocessed multi-modal brain tumor image data to obtain third multi-modal brain tumor features;
[0026] The second multi-modal brain tumor features and the third multi-modal brain tumor features are fused and compressed for redundant information processing to obtain fourth multi-modal brain tumor features;
[0027] The fourth multi-modal brain tumor features are added to the first multi-modal brain tumor features to obtain multi-modal brain tumor cross-scale semantic features.
[0028] Further, the spatial shift attention mechanism module based on the multi-modal brain tumor segmentation model performs feature enhancement processing on the multi-modal brain tumor cross-scale semantic features to obtain enhanced multi-modal brain tumor cross-scale semantic features, and the specific steps include:
[0029] The multi-modal brain tumor cross-scale semantic features are input into the spatial shift attention mechanism module of the multi-modal brain tumor segmentation model.
[0030] The feature channel grouping and spatial shift processing module based on the spatial shift attention mechanism module performs feature channel division and spatial shift operation on the multi-modal brain tumor cross-scale semantic features to obtain shifted multi-modal brain tumor cross-scale semantic features.
[0031] The cross-branch feature fusion module based on the spatial shift attention mechanism module sequentially performs channel concatenation and global average pooling processing on the shifted multi-modal brain tumor cross-scale semantic features, constructs an attention weight vector, and obtains a multi-modal brain tumor cross-scale semantic feature set.
[0032] The dynamic weight regulation mechanism based on the spatial shift attention mechanism module re-fuses the multi-modal brain tumor cross-scale semantic feature set and the attention weight vector in the channel dimension to obtain enhanced multi-modal brain tumor cross-scale semantic features.
[0033] Further, the channel adaptive fusion module based on the multi-modal brain tumor segmentation model performs feature fusion processing on the enhanced multi-modal brain tumor cross-scale semantic features to obtain fused multi-modal brain tumor cross-scale semantic features, and the specific steps include:
[0034] The enhanced multi-modal brain tumor cross-scale semantic features are input into the channel adaptive fusion module of the multi-modal brain tumor segmentation model.
[0035] The enhanced multi-modal brain tumor cross-scale semantic features are divided into shallow features and deep features.
[0036] The dual-path mechanism based on the channel adaptive fusion module sequentially performs channel global average pooling processing and dimension compression and recovery processing on the shallow features and the deep features to obtain weighted shallow features and weighted deep features.
[0037] The adaptive convolution fusion structure based on the channel adaptive fusion module splices the channel dimension of the weighted shallow layer feature and the weighted deep layer feature to obtain the fused multi-modal brain tumor cross-scale semantic feature.
[0038] Further, the detail feedback mechanism module based on the multi-modal brain tumor segmentation model performs feature optimization processing on the fused multi-modal brain tumor cross-scale semantic feature and the enhanced multi-modal brain tumor cross-scale semantic feature to obtain the optimized multi-modal brain tumor cross-scale semantic feature, and the step specifically includes:
[0039] The enhanced multi-modal brain tumor cross-scale semantic feature output by the fourth spatial shift attention mechanism is input into the detail feedback mechanism module of the multi-modal brain tumor segmentation model.
[0040] The first group of 3D transpose convolutions based on the detail feedback mechanism module performs up-sampling processing on the enhanced multi-modal brain tumor cross-scale semantic feature output by the fourth spatial shift attention mechanism to obtain an intermediate feature map.
[0041] The second group of 3D transpose convolutions based on the detail feedback mechanism module performs secondary up-sampling processing on the intermediate feature map to obtain a feature map.
[0042] The third group of 3D transpose convolutions based on the detail feedback mechanism module up-sample the feature map to the spatial scale required by the first decoding layer to obtain a final feedback feature map.
[0043] The final feedback feature map and the fused multi-modal brain tumor cross-scale semantic feature are fused to obtain the optimized multi-modal brain tumor cross-scale semantic feature.
[0044] Further, the decoding module and the convolution module based on the multi-modal brain tumor segmentation model perform multi-modal information segmentation prediction on the optimized multi-modal brain tumor cross-scale semantic feature to obtain the segmented brain tumor image, and the step specifically includes:
[0045] The decoding module and the convolution module based on the multi-modal brain tumor segmentation model;
[0046] The fourth spatial shift attention mechanism outputs the enhanced multi-modal brain tumor cross-scale semantic feature, which is up-sampled by the first three-dimensional transpose convolution module to obtain an up-sampled multi-modal brain tumor cross-scale semantic feature.
[0047] The up-sampled multi-modal brain tumor cross-scale semantic feature and the final feedback feature map are fused to obtain a first fusion feature.
[0048] The first fusion feature is subjected to feature enhancement processing by two layers of three-dimensional convolution modules to obtain an output feature map of the first decoding layer.
[0049] Based on the second decoding layer, upsampling the output feature map of the first decoding layer through a second set of transposed convolutions to obtain the output feature map of the second decoding layer;
[0050] The enhanced multimodal brain tumor cross-scale semantic features output by the third spatial shift attention mechanism are fused with the output feature map of the second decoding layer to obtain the second fused features;
[0051] Perform long sampling on the second fusion feature to obtain the output feature map of the third decoding layer;
[0052] Through the convolution module, the number of channels of the output feature map of the third decoding layer is mapped to the dimension of the number of segmentation categories, and the activation function is connected to obtain the segmented brain tumor image.
[0053] The second technical solution adopted by the present invention is: a multimodal brain tumor segmentation system based on channel adaptive structure feedback, comprising:
[0054] The first module is used to acquire multimodal brain tumor image data and perform three-dimensional data preprocessing to obtain preprocessed multimodal brain tumor image data;
[0055] The second module is used to introduce a multimodal pyramid feature encoding module, a spatial shift attention mechanism module, and a channel adaptive fusion module to build a multimodal brain tumor segmentation model;
[0056] The third module is used to perform image segmentation processing on the preprocessed multimodal brain tumor image data based on the multimodal brain tumor segmentation model to obtain a segmented brain tumor image.
[0057] The beneficial effects of the method and system of the present invention are as follows: the present invention obtains multimodal brain tumor image data and performs three-dimensional data preprocessing to obtain preprocessed multimodal brain tumor image data, further introduces a multimodal pyramid feature encoding module, a spatial shift attention mechanism module and a channel adaptive fusion module to construct a multimodal brain tumor segmentation model, adopts a pyramid feature extraction module (MPFE) to encode structural information at different scales, and uses a multi-path convolutional structure to capture boundary, local and global features. After each level of encoding, a spatial shift attention mechanism is introduced to improve the model's modeling ability of directional and spatially heterogeneous structures. To address the fusion gap between shallow structures and deep semantics, a channel adaptive fusion module (CAFM) is used to achieve dynamic fusion of structural and semantic information through a dual-channel attention mechanism and a convolutional recoding strategy. Finally, based on the multimodal brain tumor segmentation model, image segmentation processing is performed on the preprocessed multimodal brain tumor image data to obtain a segmented brain tumor image. The multimodal brain tumor image can effectively segment multimodal brain tumor images with weak tumor boundary information, thereby improving the segmentation accuracy of multimodal brain tumor images. BRIEF DESCRIPTION OF DRAWINGS
[0058] Figure 1 is a step flow chart of the multi-modal brain tumor segmentation method based on channel adaptive structure feedback of the present application;
[0059] Figure 2 is a structural block diagram of the multi-modal brain tumor segmentation system based on channel adaptive structure feedback of the present application;
[0060] Figure 3 is a structural schematic diagram of the multi-modal brain tumor segmentation model provided by the specific embodiment of the present application;
[0061] Figure 4 is a BraTS2021 two patient brain tumor data slice image schematic diagram provided by the specific embodiment of the present application;
[0062] Figure 5 is a subjective segmentation result schematic diagram of the BraTS2021 data set by different methods provided by the specific embodiment of the present application. DETAILED DESCRIPTION
[0063] The present application will be further described in detail below in combination with the drawings and specific embodiments. For the step numbers in the following embodiments, only the setting is for the convenience of description, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.
[0064] Referring to Figure 1 , the present application provides a multi-modal brain tumor segmentation method based on channel adaptive structure feedback, which comprises the following steps:
[0065] S100, acquiring multi-modal brain tumor image data and performing three-dimensional data preprocessing to obtain preprocessed multi-modal brain tumor image data;
[0066] First of all, it needs to be pointed out that before performing the brain tumor segmentation task, the present application embodiment constructs a complete multi-modal MRI image preprocessing procedure, which ensures that the model input data quality is high, consistency is strong, and can adapt to the training requirements of complex model structure. The data preprocessing procedure adapts to the BraTS2021 public data set.
[0067] S110, acquiring pre-set modality brain tumor image data and corresponding label image for modality stacking to obtain multi-modal brain tumor image data, the pre-set modality including T1, T1ce, T2 and Flair;
[0068] Specifically, first, four-mode brain tumor images T1, T1ce, T2, Flair and corresponding label images (seg.nii.gz) are acquired. The.nii.gz image files of each mode are loaded using SimpleITK, and the image dimension is three-dimensional (H, W, D). After loading, it is uniformly transformed to (channel, H, W, D) order, four modes are stacked to form a shape 4x240x240x155 image tensor. The label image is read according to the original label value and converted into a three-dimensional array. The label categories include: 0 represents background, 1 represents tumor necrosis / non-enhanced core (NCR / NET), 2 represents edema (ED), and 4 represents enhanced tumor area (ET). The part with a value of 4 in the label image is remapped to 3, and the uniform category is 4 (【0, 1, 2, 3】).
[0069] S120, sequentially performing data normalization processing and image enhancement processing on the multi-modal brain tumor image data to obtain preliminary preprocessed multi-modal brain tumor image data;
[0070] Specifically, in order to improve the training stability and reduce the influence of inter-modal differences, the present application performs background-free area normalization on each modal image, and performs mean zero and standard deviation normalization on the voxel values outside the background area. This special normalization method is as follows:
[0071]
[0072] In the formula, i represents the number of channels, x i represents all non-background points in the i-th channel, u i is the mean value on channel i, and sigma i is the variance on channel i. g represents random superimposed noise, which is derived from a Gaussian distribution with a mean of 0, and the variance is a mean distribution between 0 and 0.1.
[0073] In order to enhance the robustness of the model to different spatial transformations and noise, the present application randomly rotates and flips the source data, randomly selects 90°, 180° and 270° rotation angles for two-dimensional rotation of the image and its label, and randomly performs axial (x, y, z) flipping to randomly disturb the space and improve the generalization ability of the model.
[0074] S130, performing random cropping and center cropping processing and image tensor conversion processing on the preliminary preprocessed multi-modal brain tumor image data to obtain preprocessed multi-modal brain tumor image data.
[0075] Specifically, the embodiment of the present application is to adapt the input limit of the 3D segmentation model, introduce a cropping strategy, randomly select a cropping starting point in the spatial dimension (W, H, D) of the image, and crop a fixed size region from the starting point, while cropping the image and the label to maintain it. In the validation set, the center of the image is taken as the reference point, and a volume region of a specified size is cropped. According to the actual application scenario, the axial 64 key slice range is selected as the training target region to reduce the consumption of computing resources.
[0076] After preprocessing is completed, all images and labels are converted into PyTorch tensors, the image is of the float32 type, and the label is of the int64 type. In the final data structure, the image tensor is [B, 4, H, W, D], and the label tensor is [B, H, W, D].
[0077] S200, introduce a multi-modal pyramid feature encoding module, a spatial shift attention mechanism module, and a channel adaptive fusion module to construct a multi-modal brain tumor segmentation model;
[0078] Specifically, as shown in Figure 3 the multi-modal brain tumor segmentation model specifically includes a multi-modal pyramid feature encoding module, a spatial shift attention mechanism module, a channel adaptive fusion module, a detail feedback mechanism module, a decoding module, and a convolution module. The multi-modal pyramid feature encoding module includes a first pyramid feature extraction module, a second pyramid feature extraction module, a third pyramid feature extraction module, and a fourth pyramid feature extraction module. The spatial shift attention mechanism module includes a first spatial shift attention mechanism, a second spatial shift attention mechanism, a third spatial shift attention mechanism, and a fourth spatial shift attention mechanism. The detail feedback mechanism module includes a first detail feedback mechanism and a second detail feedback mechanism. The decoding module includes a first decoding layer, a second decoding layer, and a third decoding layer.
[0079] S300, based on the multi-modal brain tumor segmentation model, performing image segmentation processing on the preprocessed multi-modal brain tumor image data to obtain segmented brain tumor images.
[0080] S310, inputting the preprocessed multi-modal brain tumor image data to the multi-modal brain tumor segmentation model;
[0081] S320, based on the multi-modal pyramid feature encoding module of the multi-modal brain tumor segmentation model, performing cross-scale semantic feature extraction processing on the preprocessed multi-modal brain tumor image data to obtain multi-modal brain tumor cross-scale semantic features;
[0082] Specifically, the preprocessed multi-modal brain tumor image data is input to a multi-modal pyramid feature encoding module of the multi-modal brain tumor segmentation model; based on a first convolution kernel of the multi-modal pyramid feature encoding module, the preprocessed multi-modal brain tumor image data is subjected to compression channel dimension and nonlinear transformation processing, to obtain first multi-modal brain tumor features; based on a second convolution kernel of the multi-modal pyramid feature encoding module, the preprocessed multi-modal brain tumor image data is subjected to extraction of tumor edge structure features, to obtain second multi-modal brain tumor features; based on a third convolution kernel of the multi-modal pyramid feature encoding module, the preprocessed multi-modal brain tumor image data is subjected to global context information extraction processing, to obtain third multi-modal brain tumor features; the second multi-modal brain tumor features and the third multi-modal brain tumor features are fused and subjected to compression of redundant information, to obtain fourth multi-modal brain tumor features; the fourth multi-modal brain tumor features and the first multi-modal brain tumor features are added, to obtain multi-modal brain tumor cross-scale semantic features.
[0083] In the embodiment, in order to better extract multi-scale structure information in the brain tumor image, the application introduces a multi-scale feature extraction module, i.e., a pyramid feature encoding module (MPFE), in the encoder, which adopts three-dimensional convolution kernel structures of three different receptive fields, and introduces a normalization and activation mechanism, to effectively enhance the capturing ability of local details, edge structures and large-scale semantic regions of the image.
[0084] The embodiment of the application fully considers the structural complexity, size difference and modality heterogeneity of the tumor region in three-dimensional space, and therefore adopts three kinds of convolution structure combinations to extract different types of features:
[0085] The first path: 1x1x1 convolution kernel.
[0086] The first path: 1x1x1 convolution kernel.
[0087] F1=GN(ReLU(Conv 1×1×1 (X)))
[0088] The second path: 3x3x3 convolution kernel.
[0089] The receptive field is a local mesoscale region, which can extract typical structures such as tumor edge contour and tissue boundary, and enhance the boundary perception ability of the network, and its expression is:
[0090] F2=GN(ReLU(Conv 3×3×3 (X)))
[0091] The third path: 5x5x5 convolution kernel.
[0092] The receptive field covers a larger area, is suitable for perceiving global context information of a large tumor or edema area in the space, and helps to model the whole tumor structure (Whole Tumor), and the expression is:
[0093] F3=GN(ReLU(Conv 5×5×5 (X)))
[0094] Considering the complementarity of the features extracted by different convolution paths in terms of rescaling and representation ability, the application designs a multi-path fusion strategy to fuse the features of the three paths. The specific fusion process is as follows.
[0095] First, F2 and F3 are added and fused to form a medium-large receptive field information combination, and the expression is:
[0096] F 23 =F2+F3
[0097] Then, a layer of 1x1x1 convolution is used to unify the channel dimension and compress redundant information, and the expression is:
[0098] F fused =Conv 1×1×1 (F 23 )
[0099] Finally, the first path output F1 and the fused feature F fused are added to form the output feature of the current layer, and the expression is:
[0100] F out =F1+F fused
[0101] The structure maintains a light amount of calculation, realizes comprehensive modeling capability across sizes and boundaries through parallel multi-scale fusion in structure, and effectively improves the overall perception effect of the tumor core area, enhanced edge and edema area.
[0102] S330, a spatial shift attention mechanism module based on the multi-modal brain tumor segmentation model, performs feature enhancement processing on the multi-modal brain tumor cross-scale semantic features to obtain enhanced multi-modal brain tumor cross-scale semantic features;
[0103] Specifically, the multi-modal brain tumor cross-scale semantic features are input to a spatial shift attention mechanism module of a multi-modal brain tumor segmentation model; based on a feature channel grouping and spatial shift processing module of the spatial shift attention mechanism module, feature channel division and spatial shift operation are performed on the multi-modal brain tumor cross-scale semantic features to obtain the shifted multi-modal brain tumor cross-scale semantic features; based on a cross-branch feature fusion module of the spatial shift attention mechanism module, the shifted multi-modal brain tumor cross-scale semantic features are sequentially subjected to channel splicing and global average pooling processing to construct an attention weight vector, and a multi-modal brain tumor cross-scale semantic feature set is obtained; based on a dynamic weight regulation mechanism of the spatial shift attention mechanism module, the multi-modal brain tumor cross-scale semantic feature set and the attention weight vector are re-fused in the channel dimension to obtain enhanced multi-modal brain tumor cross-scale semantic features.
[0104] In the embodiment, to further improve the spatial modeling ability and multi-directional perception ability of the encoder output feature map, the application introduces a spatial-channel joint attention mechanism module, namely a spatial shift attention mechanism, after each MPFE. The module effectively establishes spatial anisotropic correlation and enhances the perception ability of the model to directional, edge structure and non-local features when segmenting complex tumor structures by introducing a shift operation in the spatial dimension and a branch attention mechanism.
[0105] The Spatial Attention Mechanism module mainly consists of three parts: feature channel grouping and spatial shift processing, cross-branch feature fusion, and dynamic weight regulation mechanism, and the specific process is as follows:
[0106] 1) Feature channel division and spatial shift operation.
[0107] Let the input feature map be: X ∈ R B×C×D×H×W where B represents the batch size, C is the number of channels, D, H, and W are the dimensions of the depth, height, and width directions respectively. The channel dimension is equally divided into three sub-tensors in order: Different direction spatial shift operations are applied to X1, X2, and X3 to capture the directional structure in three-dimensional space. X1 is shifted up by one step in the height dimension, X2 is shifted left by one step in the width dimension, and X3 is shifted forward by one step in the depth dimension. The shift is realized by a tensor space moving function, such as the roll operation in PyTorch. The shifted result is defined as:
[0108]
[0109] The spatial displacement enables each branch feature to perceive the position information in space, thereby enhancing the modeling capability of tumor edges, strip structures and directional features.
[0110] 2) Cross-branch channel fusion and attention construction.
[0111] The three-way offset features are spliced again in the channel dimension to form a three-branch feature set, and its expression is:
[0112]
[0113] The spliced results are subjected to a global average pooling operation to obtain global description information of each branch, and its expression is:
[0114]
[0115] The above global features are sent into two layers of multi-layer perceptron (MLP) to construct an attention weight vector, which is specifically as follows:
[0116] a = Softmax (W2·GELU (W1·Z))
[0117] wherein are the weight matrices of the two fully connected layers in the MLP, and r is the channel compression ratio. The Softmax function ensures the normalization of the three branch weights, and the sum is 1.
[0118] 3) Weighted fusion output enhanced features.
[0119] The three channel groups are multiplied by their corresponding weights, and finally re-fused in the channel dimension to obtain the final output of the module, and its expression is:
[0120]
[0121] In the above formula, X out represents the enhanced multi-modal brain tumor cross-scale semantic features.
[0122] S340, based on the channel adaptive fusion module of the multi-modal brain tumor segmentation model, the enhanced multi-modal brain tumor cross-scale semantic features are subjected to feature fusion processing to obtain fused multi-modal brain tumor cross-scale semantic features;
[0123] Specifically, the enhanced multi-modal brain tumor cross-scale semantic features are input to a channel adaptive fusion module of a multi-modal brain tumor segmentation model; the enhanced multi-modal brain tumor cross-scale semantic features are subjected to feature division processing to obtain shallow features and deep features; based on a two-way mechanism of the channel adaptive fusion module, the shallow features and the deep features are subjected to channel global average pooling processing and dimension compression and recovery processing in sequence to obtain weighted shallow features and weighted deep features; based on an adaptive convolution fusion structure of the channel adaptive fusion module, the weighted shallow features and the weighted deep features are subjected to channel dimension splicing to obtain fused multi-modal brain tumor cross-scale semantic features.
[0124] In the embodiment, in a conventional encoding-decoding segmentation network, there is a semantic gap between shallow features focusing on spatial structure and deep features focusing on semantic abstraction, and if direct splicing is performed, semantic conflicts are easily introduced, affecting segmentation performance.
[0125] The application proposes a channel adaptive fusion module (CAFM), which fuses network shallow structure detail features and deep high semantic features, introduces a two-way Squeeze-and-Excitation (SE) mechanism and an adaptive convolution fusion structure, realizes channel selectivity regulation in the feature fusion process, and effectively alleviates the heterogeneity difference between multi-layer features.
[0126] For double-input feature construction, high and low layer semantics are fused, and first, F low ∈R B×C×D×H×W is a shallow feature, and F high ∈R B×C×D×H×W is a feedback feature after upsampling from a deep layer, and the two features are sent to independent channel attention modules for channel importance regulation, and are respectively represented as:
[0127] F’ low =SE(F low )
[0128] F’ high =SE(F high )
[0129] wherein the SE module is a three-dimensional Squeeze-and-Excitation module, and its calculation process is as follows:
[0130] 1) Channel global average pooling, and its expression is as follows:
[0131]
[0132] 2) Two layers of 1x1x1 convolution to realize dimension compression and recovery (i.e. MLP), whose expression is:
[0133] s = σ(W2·δ(W1·z))
[0134] where δ is ReLU activation, σ is Sigmoid activation, and W1, W2 are weight matrices of 1x1x1 convolution.
[0135] The final channel weighting weight, whose expression is:
[0136]
[0137] Feature splicing and 1x1x1 convolution fusion, which splices the two weighted feature maps in the channel dimension, whose expression is:
[0138]
[0139] Then use a 1x1x1 three-bit convolution kernel to compress it to the target channel number C t , whose expression is
[0140]
[0141] This fusion process not only realizes channel dimension compression, but also can be regarded as re-encoding the fused features to obtain discriminative structural semantic information.
[0142] S350, based on the multi-modal brain tumor segmentation model, the detail feedback mechanism module is used to optimize the fused multi-modal brain tumor cross-scale semantic features and the enhanced multi-modal brain tumor cross-scale semantic features, and obtain the optimized multi-modal brain tumor cross-scale semantic features;
[0143] Specifically, the enhanced multi-modal brain tumor cross-scale semantic features output by the fourth spatial shift attention mechanism are input to the detail feedback mechanism module of the multi-modal brain tumor segmentation model; the first group of 3D transpose convolutions based on the detail feedback mechanism module are used to upsample the enhanced multi-modal brain tumor cross-scale semantic features output by the fourth spatial shift attention mechanism, to obtain an intermediate feature map; the second group of 3D transpose convolutions based on the detail feedback mechanism module are used to perform secondary upsample processing on the intermediate feature map, to obtain a feature map; the third group of 3D transpose convolutions based on the detail feedback mechanism module are used to upsample the feature map to the spatial scale required by the first decoding layer, to obtain a final feedback feature map; the final feedback feature map is fused with the fused multi-modal brain tumor cross-scale semantic features, to obtain the optimized multi-modal brain tumor cross-scale semantic features.
[0144] In this embodiment, in order to further improve the segmentation accuracy of small tumor regions, fuzzy boundaries and multi-modal lesion structures by the model, a deep semantic feature feedback mechanism is introduced in the decoder path, which guides the high semantic information of the deepest layer of the encoder to the first layer of the decoder through layer-by-layer upsampling, realizes the direct perception of high-dimensional abstract information in the initial stage of decoding, and achieves the cooperative optimization of structure and semantics.
[0145] First, the deep semantic feature map output by the fourth layer of the encoder is obtained. The feature map has the strongest semantic abstraction ability, and its tensor representation is: Where B represents the batch size, C4 represents the number of channels of this layer, D4, H4 and W4 represent the depth, height and width respectively. Then a step-by-step upsampling path is constructed using three-dimensional transposed convolution (Transposed 3D Convolution), and F deep is expanded to the same spatial feature as the first layer of the decoder. This process includes the following three steps in turn:
[0146] 1) The first set of 3D transposed convolution is used to perform the first upsampling on F deep , and obtain the intermediate feature map F up1 The tensor form is as follows:
[0147]
[0148] 2) The second set of 3D transposed convolution is used to perform the second upsampling on F up1 , and generate the feature map F up2 , whose expression is:
[0149]
[0150] 3) Finally, the third set of 3D transposed convolution is used to upsample F up2 to the spatial scale required by the first layer of the decoder, and obtain the final feedback feature map F feedback , whose expression is:
[0151]
[0152] Then the feedback feature map F feedback is fused with the structural feature map generated in the first layer of the decoder. The two have been aligned in the spatial dimension, so they can be directly involved in the fusion. The specific fusion method is to call the channel adaptive fusion module (CAFM) defined in the fourth step, and form the fusion feature F fusion after the two feature maps are compressed by double-channel attention weighting and 1x1x1 convolution. The process is described as follows:
[0153] F fusion = CAFM(Ffeedback F decode1 )
[0154] The fusion feature F fusion is input to the subsequent convolution layer and prediction head to realize the final segmentation output. This path effectively realizes the closed-loop transmission of feature information from the highest semantic layer of the encoding end to the lowest layer of the decoder, provides high-level semantic support for the initial decoding, thereby improving the recovery quality of the tumor boundary and the structural consistency of the multi-region.
[0155] Through the above detail feedback mechanism, the application introduces deep semantic guidance while maintaining structural information, realizes the bidirectional collaborative enhancement of structure-semantics, effectively alleviates the problem of excessive dependence on low-level information in the structure reconstruction stage of the traditional decoder, and improves the overall segmentation performance of the model.
[0156] S360, the decoding module based on the multi-modal brain tumor segmentation model and the convolution module, perform multi-modal information segmentation prediction on the optimized multi-modal brain tumor cross-scale semantic features to obtain a segmented brain tumor image.
[0157] Specifically, the decoding module based on the multi-modal brain tumor segmentation model and the convolution module; the enhanced multi-modal brain tumor cross-scale semantic features output by the fourth spatial shift attention mechanism are up-sampled by a first three-dimensional transposed convolution module to obtain up-sampled multi-modal brain tumor cross-scale semantic features; the up-sampled multi-modal brain tumor cross-scale semantic features are fused with the final feedback feature map to obtain a first fusion feature; the first fusion feature is subjected to feature enhancement processing by two layers of three-dimensional convolution modules to obtain an output feature map of a first decoding layer; based on a second decoding layer, the output feature map of the first decoding layer is up-sampled by a second group of transposed convolution to obtain an output feature map of the second decoding layer; the enhanced multi-modal brain tumor cross-scale semantic features output by the third spatial shift attention mechanism are fused with the output feature map of the second decoding layer to obtain a second fusion feature; the second fusion feature is up-sampled to obtain an output feature map of a third decoding layer; the channel number of the output feature map of the third decoding layer is mapped to the dimension of the segmentation class number by a convolution module, and an activation function is connected to obtain a segmented brain tumor image.
[0158] In this embodiment, to realize the effective restoration of the deep semantic features extracted in the encoding stage to the original spatial resolution, the application constructs a symmetrical decoder module containing three layers of structures. The decoder as a whole realizes the gradual recovery of spatial resolution by multiple three-dimensional transposed convolution (Transposed 3D Convolution), and at the same time, in the initial decoding stage, the detail feedback feature map formed by up-sampling from the deepest layer of the encoder is fused, the structure-semantics coupling is realized through a channel adaptive fusion module (CAFM), and the boundary expression and small target structure recovery ability are significantly improved.
[0159] The layer receives the high semantic feature map output by the 4th layer of the encoder First, it is up-sampled by a set of 2x2x2 three-dimensional transpose convolution modules, whose expression is:
[0160]
[0161] At the same time, the deep semantic features of the 4th layer of the encoder also form structural features F through the detail feedback mechanism described in the fifth step feedback , whose spatial resolution is consistent with F up1 . Both are input into the channel adaptive fusion module for fusion, whose expression is:
[0162] F fusion1 = CAFM(F feedback , F up1 )
[0163] The fusion feature F fusion1 is input into two layers of three-dimensional convolution modules (each layer contains 3x3x convolution + GroupNorm + ReLU) for feature enhancement, and the expression of the output feature map of the first layer of the decoder is:
[0164] F decode1 = ConvUnit(F fusion1 )
[0165] The second layer of encoding uses the second set of transpose convolution to up-sample the output feature of the first layer of the decoder, whose expression is:
[0166] F up2 = UpConv2(F decode1 )
[0167] Channel concatenation is performed with the feature map F enc3 of the 3rd layer of the encoder, whose expression is:
[0168] F cat2 = Concat(F up2 , F enc3 )
[0169] The fusion feature is input into two layers of three-dimensional convolution structure for feature extraction and fusion, whose expression is:
[0170] F decode2 = ConvUnit(F cat2 )
[0171] The third layer coding layer repeats the operation of the second layer, performs third upsampling on the coding layer second layer output feature map, splices the feature map obtained by the encoder jump connection, and inputs two convolution modules for final spatial feature restoration. The output feature F decode3 An input 1x1x1 three-dimensional convolution module maps the channel number to the segmentation class number dimension, and its expression is:
[0172] F logit =Conv 1×1×1 (F decode3 )
[0173] An access activation function obtains a voxel-level segmentation prediction result, and its expression is:
[0174]
[0175] In the above formula, denotes the segmented brain tumor image.
[0176] In summary, the embodiment of the application first uniformly loads, normalizes, data enhances and crops the T1, T1ce, T2 and FLAIR four-mode MRI images, constructs a standardized three-dimensional multi-modal input data block. Then, a pyramid feature extraction module (MPFE) is used to encode the structural information at different scales, and a multi-path convolution structure is used to capture boundary, local and global features. After each level of coding, a spatial shift attention mechanism is introduced to improve the modeling ability of the model for directionality and spatial heterogeneous structure. In view of the fusion gap existing between the shallow structure and the deep semantic, the application designs a channel adaptive fusion module (CAFM), which realizes dynamic fusion of structure and semantic information through a double-channel attention mechanism and a convolution re-encoding strategy. In addition, in order to further improve the weak target boundary perception ability, a detail feedback mechanism is proposed, which feeds back the deepest semantic features of the encoder to the first layer of the decoder through multi-level upsampling, realizes structure-semantic closed-loop enhancement. Finally, through three layers of decoding structure, the spatial resolution is gradually restored to generate a tumor segmentation map. Experimental results show that the application achieves excellent Dice and HD95 indicators on multiple tumor regions (WT, TC, ET) of the BraTS2021 dataset, significantly better than existing mainstream methods, and has good medical application prospects.
[0177] Further, the embodiments of the present application evaluate the proposed application method on the BraTS2021 dataset, which includes 1251 cases collected from the Kaggle competition platform. The dataset is divided into 1000 training cases, 125 validation cases and 126 test cases. Each case includes four MRI modes (T1, T1ce, FLAIR, T2) and the corresponding expert annotation mask, and the segmentation label covers background, necrotic tissue, edema and enhanced tumor. The spatial resolution of all scans is 240x240x150, and the axial plane is used for slice processing. The focus of the evaluation is on three tumor subregions: enhanced tumor (ET), whole tumor (WT) and tumor core (TC), as shown in Figure 2
[0178] As shown in Figure 4 To further demonstrate the advantages and effectiveness of the present application, the embodiments of the present application perform comparative experiments on the BraTS2021 dataset with several representative brain tumor segmentation methods, including 3D U-Net for learning dense collective segmentation from sparse annotations, TransBTS and 3D PSwinBTS for multi-modal brain tumor segmentation based on Transformer, QT-UNet for 3D segmentation based on self-supervised self-query full transformer, Ultralight-VM-UNet for skin lesion segmentation with significantly reduced parameters of parallel visual mamba, DAUnet for 2D human pose estimation with detail perception U-shaped network, and Yaru3DFPN for lightweight improved 3D UNet. As shown in Table 1, on the BraTS2021 validation set, the Dice scores of the present application on the whole tumor (WT), tumor core (TC) and enhanced tumor (ET) are 91.349%, 87.089% and 84.976% respectively. Compared with the classic 3D U-Net, the present application improves by 3.329%, 10.919% and 8.776% on WT, TC and ET respectively. In addition, the present application is superior to transformer-based methods such as TransBTS and 3D PSwinBTS, and mamba-based models, achieving the highest Dice score on ET and TC, and showing better Hausdorff95 distance. Overall, the present application achieves the best performance of Dice and HD95 indicators in all tumor subregions, proving its strong generalization and segmentation ability.
[0179] Table 1 Performance comparison of different methods on BraTS2021 dataset
[0180]
[0181]
[0182] As Figure 5 shown, to intuitively compare the segmentation performance of the method of the present application with other methods, we give two examples of visualization results. In the image, yellow represents the enhanced tumor; the combination of yellow and red represents the tumor core; and the entire area composed of yellow, red and green corresponds to the entire tumor, respectively. Figure 5 The qualitative segmentation results of randomly selected samples 00016_068 and 00071_062 on the BraTS2021 dataset are shown. As shown in the figure, Yaru3DFPN roughly captures the tumor contour, but due to the presence of small target points, the enhanced tumor area overlaps with the tumor core area. DAUnet cannot accurately segment the entire tumor, losing a lot of boundary information. In contrast, thanks to its attention mechanism, the present application performs excellent performance in image segmentation.
[0183] With reference to Figure 2 , the multi-modal brain tumor segmentation system based on channel adaptive structure feedback comprises:
[0184] The first module 201 is configured to acquire multi-modal brain tumor image data and perform three-dimensional data preprocessing to obtain preprocessed multi-modal brain tumor image data.
[0185] The second module 202 is configured to introduce a multi-modal pyramid feature encoding module, a spatial shift attention mechanism module and a channel adaptive fusion module to construct a multi-modal brain tumor segmentation model.
[0186] The third module 203 is configured to perform image segmentation processing on the preprocessed multi-modal brain tumor image data based on the multi-modal brain tumor segmentation model to obtain segmented brain tumor images.
[0187] The contents in the above method embodiments are all applicable to the system embodiments, the system embodiments specifically implement the same functions as the above method embodiments, and achieve the same beneficial effects as the above method embodiments.
[0188] The above is a specific description of the preferred embodiments of the present application, but the present application is not limited to the described embodiments, and those skilled in the art can make various equivalent modifications or replacements without deviating from the spirit of the present application, and these equivalent modifications or replacements are all included in the scope defined by the claims of the present application.
Claims
1. A multi-modal brain tumor segmentation method based on channel-adaptive structure feedback, characterized in that, The following steps are involved: Acquiring multimodal brain tumor image data and performing three-dimensional data preprocessing to obtain preprocessed multimodal brain tumor image data; A multimodal pyramid feature encoding module, a spatial shift attention mechanism module, and a channel adaptive fusion module are introduced to build a multimodal brain tumor segmentation model. Based on the multimodal brain tumor segmentation model, image segmentation processing is performed on the preprocessed multimodal brain tumor image data to obtain a segmented brain tumor image.
2. The multi-modal brain tumor segmentation method based on channel-adaptive structure feedback according to claim 1, characterized in that, The step of acquiring multimodal brain tumor image data and performing three-dimensional data preprocessing to obtain preprocessed multimodal brain tumor image data specifically includes: Acquire preset modality brain tumor image data and corresponding annotation images for modality stacking to obtain multimodal brain tumor image data, wherein the preset modalities include T1, T1ce, T2, and Flair; The multimodal brain tumor image data is subjected to data normalization and image enhancement processing in sequence to obtain preliminary pre-processed multimodal brain tumor image data; The pre-processed multimodal brain tumor image data is subjected to random cropping, center cropping and image tensor conversion to obtain pre-processed multimodal brain tumor image data.
3. The multi-modality brain tumor segmentation method based on channel-adaptive structure feedback according to claim 2, characterized in that, The multimodal brain tumor segmentation model specifically includes a multimodal pyramid feature encoding module, a spatial shift attention mechanism module, a channel adaptive fusion module, a detail feedback mechanism module, a decoding module and a convolution module. The multimodal pyramid feature encoding module includes a first pyramid feature extraction module, a second pyramid feature extraction module, a third pyramid feature extraction module and a fourth pyramid feature extraction module. The spatial shift attention mechanism module includes a first spatial shift attention mechanism, a second spatial shift attention mechanism, a third spatial shift attention mechanism and a fourth spatial shift attention mechanism. The detail feedback mechanism module includes a first detail feedback mechanism and a second detail feedback mechanism. The decoding module includes a first decoding layer, a second decoding layer and a third decoding layer.
4. The multi-modality brain tumor segmentation method based on channel-adaptive structure feedback according to claim 3, characterized in that, The step of performing image segmentation processing on the pre-processed multimodal brain tumor image data based on the multimodal brain tumor segmentation model to obtain a segmented brain tumor image specifically includes: Inputting the preprocessed multimodal brain tumor image data into the multimodal brain tumor segmentation model; Based on the multimodal pyramid feature encoding module of the multimodal brain tumor segmentation model, cross-scale semantic feature extraction is performed on the preprocessed multimodal brain tumor image data to obtain multimodal brain tumor cross-scale semantic features; Based on the spatial shift attention mechanism module of the multimodal brain tumor segmentation model, feature enhancement processing is performed on the multimodal brain tumor cross-scale semantic features to obtain enhanced multimodal brain tumor cross-scale semantic features; Based on the channel-adaptive fusion module of the multimodal brain tumor segmentation model, the enhanced multimodal brain tumor cross-scale semantic features are fused to obtain the fused multimodal brain tumor cross-scale semantic features. The detail feedback mechanism module based on the multi-modal brain tumor segmentation model performs feature optimization processing on the fused multi-modal brain tumor cross-scale semantic features and the enhanced multi-modal brain tumor cross-scale semantic features, to obtain optimized multi-modal brain tumor cross-scale semantic features; The decoding module and the convolution module based on the multi-modal brain tumor segmentation model perform multi-modal information segmentation prediction on the optimized multi-modal brain tumor cross-scale semantic features, to obtain segmented brain tumor images.
5. The multi-modal brain tumor segmentation method based on channel-adaptive structure feedback according to claim 4, characterized in that, The multi-modal pyramid feature encoding module based on the multi-modal brain tumor segmentation model performs cross-scale semantic feature extraction processing on the preprocessed multi-modal brain tumor image data, to obtain multi-modal brain tumor cross-scale semantic features, which specifically includes: The preprocessed multi-modal brain tumor image data is input into the multi-modal pyramid feature encoding module of the multi-modal brain tumor segmentation model; The first convolution kernel based on the multi-modal pyramid feature encoding module performs compression channel dimension and nonlinear transformation processing on the preprocessed multi-modal brain tumor image data, to obtain first multi-modal brain tumor features; The second convolution kernel based on the multi-modal pyramid feature encoding module extracts tumor edge structure features from the preprocessed multi-modal brain tumor image data, to obtain second multi-modal brain tumor features; The third convolution kernel based on the multi-modal pyramid feature encoding module performs global context information extraction processing on the preprocessed multi-modal brain tumor image data, to obtain third multi-modal brain tumor features; The second multi-modal brain tumor features and the third multi-modal brain tumor features are fused and compressed for redundant information processing, to obtain fourth multi-modal brain tumor features; The fourth multi-modal brain tumor features and the first multi-modal brain tumor features are added, to obtain multi-modal brain tumor cross-scale semantic features.
6. The multi-modality brain tumor segmentation method based on channel-adaptive structure feedback according to claim 5, characterized in that, The spatial shift attention mechanism module based on the multi-modal brain tumor segmentation model performs feature enhancement processing on the multi-modal brain tumor cross-scale semantic features, to obtain enhanced multi-modal brain tumor cross-scale semantic features, which specifically includes: The multi-modal brain tumor cross-scale semantic features are input into the spatial shift attention mechanism module of the multi-modal brain tumor segmentation model; The feature channel grouping and spatial shift processing module based on the spatial shift attention mechanism module performs feature channel division and spatial displacement operation on the multi-modal brain tumor cross-scale semantic features, to obtain displaced multi-modal brain tumor cross-scale semantic features; The cross-branch feature fusion module based on the spatial shift attention mechanism module sequentially performs channel splicing and global average pooling processing on the displaced multi-modal brain tumor cross-scale semantic features, constructs an attention weight vector, and obtains a multi-modal brain tumor cross-scale semantic feature set; The dynamic weight regulation mechanism based on the spatial shift attention mechanism module re-fuses the multi-modal brain tumor cross-scale semantic feature set and the attention weight vector in the channel dimension, to obtain enhanced multi-modal brain tumor cross-scale semantic features.
7. The multi-modality brain tumor segmentation method based on channel-adaptive structure feedback according to claim 6, characterized in that, The channel adaptive fusion module based on the multi-modal brain tumor segmentation model performs feature fusion processing on the enhanced multi-modal brain tumor cross-scale semantic features, and obtains fused multi-modal brain tumor cross-scale semantic features. The specific steps include: The enhanced multi-modal brain tumor cross-scale semantic features are input into the channel adaptive fusion module of the multi-modal brain tumor segmentation model. The enhanced multi-modal brain tumor cross-scale semantic features are subjected to feature division processing to obtain shallow features and deep features. Based on the two-way mechanism of the channel adaptive fusion module, the shallow features and the deep features are sequentially subjected to channel global average pooling processing and dimension compression and recovery processing to obtain weighted shallow features and weighted deep features. Based on the adaptive convolution fusion structure of the channel adaptive fusion module, the weighted shallow features and the weighted deep features are subjected to channel dimension splicing to obtain fused multi-modal brain tumor cross-scale semantic features.
8. The multi-modality brain tumor segmentation method based on channel-adaptive structure feedback according to claim 7, characterized in that, The detail feedback mechanism module based on the multi-modal brain tumor segmentation model performs feature optimization processing on the fused multi-modal brain tumor cross-scale semantic features and the enhanced multi-modal brain tumor cross-scale semantic features to obtain optimized multi-modal brain tumor cross-scale semantic features. The specific steps include: The enhanced multi-modal brain tumor cross-scale semantic features output by the fourth spatial shift attention mechanism are input into the detail feedback mechanism module of the multi-modal brain tumor segmentation model. Based on the first group of 3D transpose convolution of the detail feedback mechanism module, the enhanced multi-modal brain tumor cross-scale semantic features output by the fourth spatial shift attention mechanism are subjected to up-sampling processing to obtain intermediate feature maps. Based on the second group of 3D transpose convolution of the detail feedback mechanism module, the intermediate feature maps are subjected to secondary up-sampling processing to obtain feature maps. Based on the third group of 3D transpose convolution of the detail feedback mechanism module, the feature maps are up-sampled to the spatial scale required by the first decoding layer to obtain the final feedback feature maps. The final feedback feature maps and the fused multi-modal brain tumor cross-scale semantic features are fused to obtain the optimized multi-modal brain tumor cross-scale semantic features.
9. The multi-modality brain tumor segmentation method based on channel-adaptive structure feedback according to claim 8, characterized in that, The decoding module and the convolution module based on the multi-modal brain tumor segmentation model perform multi-modal information segmentation prediction on the optimized multi-modal brain tumor cross-scale semantic features to obtain segmented brain tumor images. The specific steps include: The decoding module and the convolution module based on the multi-modal brain tumor segmentation model; The enhanced multi-modal brain tumor cross-scale semantic features output by the fourth spatial shift attention mechanism are up-sampled by the first three-dimensional transpose convolution module to obtain up-sampled multi-modal brain tumor cross-scale semantic features. The up-sampled multi-modal brain tumor cross-scale semantic features and the final feedback feature maps are fused to obtain first fusion features. The first fusion features are subjected to feature enhancement processing by two layers of three-dimensional convolution modules to obtain output feature maps of the first decoding layer. Based on the second decoding layer, the output feature maps of the first decoding layer are up-sampled by the second group of transpose convolution to obtain output feature maps of the second decoding layer. The enhanced multi-modal brain tumor cross-scale semantic features output by the third spatial shift attention mechanism are fused with the output feature maps of the second decoding layer to obtain second fusion features; The second fusion features are down-sampled to obtain output feature maps of a third decoding layer; The output feature maps of the third decoding layer are mapped to a dimension of a segmentation class number by a convolution module, and an activation function is connected to obtain a segmented brain tumor image.
10. A multi-modal brain tumor segmentation system based on channel-adaptive structure feedback, characterized in that, The method comprises the following modules: A first module is configured to obtain multi-modal brain tumor image data and perform three-dimensional data preprocessing to obtain preprocessed multi-modal brain tumor image data; A second module is configured to introduce a multi-modal pyramid feature encoding module, a spatial shift attention mechanism module and a channel adaptive fusion module to construct a multi-modal brain tumor segmentation model; A third module is configured to perform image segmentation processing on the preprocessed multi-modal brain tumor image data based on the multi-modal brain tumor segmentation model to obtain a segmented brain tumor image.
Citation Information
Cited By
Brain tumor image segmentation method based on dynamic token processing and multi-scale pyramid attention and related equipment
CN121482072A