Deleted modal brain tumor image segmentation method based on multi-axis shift MLP modal mask
By introducing a multi-level cross-modal interaction module and a spatial weight attention module in the brain tumor image segmentation model, the multi-axis shifted MLP modal mask is used to solve the problem of unsatisfactory feature interaction between modes in missing modal brain tumor image segmentation, and higher segmentation accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510253823.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to effectively capture complementary features between modes in the segmentation of missing modal brain tumor images, resulting in low segmentation accuracy, especially when the modality is incomplete or missing.
Using a method based on multi-axis shifted MLP modal mask, a brain tumor image segmentation model consisting of four encoders and a shared decoder is constructed. Through a multi-level cross-modal interaction module and a spatial weight attention module, complementary features between modals are captured and reweighted.
More accurate image segmentation of missing modal brain tumors is achieved, which enhances the cross-modal feature extraction ability and robustness of the model, especially in the case of incomplete or missing modalities, and improves the accuracy and robustness of the segmentation.
Smart Images

Figure CN120147637A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image segmentation, and more specifically, to a method for segmenting missing-modal brain tumor images based on a multi-axis shifted MLP modal mask. Background Art
[0002] Given the great threat posed by malignant brain tumors to health, timely diagnosis and treatment are necessary. Accurate segmentation of brain tumors is crucial for diagnosis, prognosis, and clinical evaluation. Magnetic resonance imaging (MRI) provides various tissue contrast views and spatial resolutions for brain diagnosis. Specifically, four magnetic resonance imaging (MRI) sequences are widely used for brain tumor segmentation, including T1-weighted (T1), contrast-enhanced T1-weighted (T1ce), T2-weighted (T2), and fluid-attenuated inversion recovery (FLAIR) modalities. Recently, many multi-modal methods have utilized these four MRI modalities to provide complementary information for brain tumor segmentation. However, in clinical practice, missing modalities are a common phenomenon due to image corruption, artifacts, and different acquisition protocols.
[0003] In recent years, deep learning techniques have made great progress in missing-modal brain tumor segmentation, and various strategies have been proposed to address missing modalities in different scenarios. One approach involves training dedicated networks for each potential combination of available modalities, which results in high training costs and large deployment space requirements. Some researchers have attempted to synthesize missing modalities to create a complete multi-modal set; however, the segmentation accuracy is limited by the quality of the generated modalities, which may introduce noise and artifacts. In incomplete modality learning, effective feature interaction is crucial for capturing complementary features across available modalities.
[0004] Due to the high computational cost of transformers, many recent studies have focused on attention mechanisms to improve the efficiency of transformers. Alternative solutions using masked attention have been proposed to make self-attention more focused on specific regions and features. However, for incomplete multi-modal brain tumor segmentation, existing inter-modal feature interaction strategies are still far from ideal.
[0005] Therefore, how to more accurately achieve missing-modal brain tumor image segmentation is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0006] In view of this, the present invention provides a method for segmenting missing-modal brain tumor images based on a multi-axis shifted MLP modal mask, which can more accurately achieve missing-modal brain tumor image segmentation.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] In a first aspect, the present invention provides a method for segmenting missing-modal brain tumor images based on a multi-axis shifted MLP modal mask, comprising the following steps:
[0009] Construct a brain tumor image segmentation model, which includes four encoders and a shared decoder, and introduce a multi-level cross-modal interaction module and a spatial weight attention module between the encoder and the shared encoder;
[0010] Obtain brain tumor images containing four modalities as a training set, and train the brain tumor image segmentation model; in the training stage, input the brain tumor images of the four modalities into the corresponding encoders respectively, and output the encoded features of different levels and different resolutions of each modality image; the multi-level cross-modal interaction module combines the global context information of the encoded features to capture complementary features between modalities; the spatial weight attention module re-weights the four modality features from different angles and inputs them into the shared decoder for decoding;
[0011] After training is completed, input the missing-modal brain tumor image to be segmented into the trained brain tumor image segmentation model for image segmentation.
[0012] Further, each encoder has multiple layers, and each layer has different scale perception capabilities.
[0013] Further, the working process of the multi-level cross-modal interaction module includes:
[0014] Downsample the output features of different levels in the four encoders to the same resolution through downsampling operations;
[0015] Concatenate the downsampled features through concatenation operations, and adaptively adjust the weights of each channel through a channel attention mechanism to obtain fused features;
[0016] Generate a local probability map after passing the fused features through a convolutional block and a spatial attention module;
[0017] Input the fused features into a spatial axial shift MLP, and use the local probability map to constrain and guide the spatial axial shift MLP to model the long-range dependence relationship of the fused features.
[0018] Further, the multi-level cross-modal interaction module is represented by the following formula:
[0019] X fuse = CA(Concat(DownSampling(x 1 ),
[0020] DownSampling(x 2),
[0021] DownSampling(x 3 ),
[0022] DownSampling(x 4 ),
[0023] DownSampling(x 5 )))
[0024] Y local =SA(CB(X fuse ))
[0025] Z=X fuse +Y local *VAS-MLP(X fuse )
[0026] Among them, X fuse represents the output feature map after the fusion of the corresponding layers of different encoders, and x 1 , x 2 , x 3 , x 4 , x 5 respectively represent the corresponding coding features of 5 different layers in the encoder; VAS-MLP represents the spatial axial shift MLP; CA represents the channel attention; Concat represents the splicing operation; CB represents the convolutional block; SA represents the spatial attention; Y local represents the local constraint weight generated by continuous convolution, that is, the local probability map; Z represents the output feature map of the multi-level cross-modal interaction module.
[0027] Furthermore, the shift operation of the spatial axial shift MLP on X fuse is expressed as:
[0028] X n =Norm(X fuse )
[0029] X lr =Proj(Shift lr (Proj(X n )))
[0030] X ld =Proj(Shift ld (Proj(X lr )))
[0031] X td =Proj(Shift td (Proj(X ld )))
[0032] X rd = Proj(Shift rd (Proj(X td )))
[0033] X = X lr + X ld + X td + X rd
[0034] X t = Norm(X)
[0035] X m = Proj(X t )
[0036] X y = X m + X fuse
[0037] Among them, X fuse represents the input feature map; X n represents the mapped map after feature map normalization; X lr , X ld , X td and X rd represent the feature maps after shift operations in different directions; X represents the fused feature map containing different position dependencies; X t represents the mapped result after normalizing the feature map X; X m represents the feature X t after linear projection; X y represents the output feature map of the spatial axial shift MLP.
[0038] Furthermore, the working process of the spatial weight attention module includes:
[0039] Dividing the input features into two categories: specific modality tokens and fused tokens;
[0040] Using the self-attention mechanism to directly concatenate and project the features of different modalities into Q, K, and V;
[0041] Constructing a binary attention mask M ∈ {0, 1} 5N×5N , controlling the interaction rules between different tokens;
[0042] Given any position (i, j) in M, and q i ∈ Q, k j ∈ K; if q i and k j belong to the same and existing specific modality tokens, then set M i,j = 1; if qi Belonging to the fusion marker, k j Coming from the existing specific modality marker, then set M i,j = 1; If q i or k j Both come from the missing modality, then set M i,j = 0; When M i,j = 1, allow q i and k j There is an interaction relationship between them, and the attention weights at the corresponding positions are retained; when M i,j = 0, do not allow q i and k j There is an interaction relationship between them, and the attention weights at the corresponding positions are filtered out;
[0043] Dynamically adjust the attention weight matrix through the binary attention mask M, and calculate the weights of each modality along the spatial dimension to re-weight the existing modalities.
[0044] Furthermore, calculate the weights of all modalities by summing column vectors:
[0045]
[0046] Among them, H represents the number of attention heads; N = H×W×D, representing the total number of voxels in the single-layer encoded feature map, H is the height of the single-layer encoded feature map, W is the width of the single-layer encoded feature map, and D is the depth of the single-layer encoded feature map; MA(i,j) represents the modality-specific feature weight, and (i,j) represents different positions in the feature dimension 1×N;
[0047] Multiply the weight of each modality marker by the input feature to obtain the output feature of the spatial weight attention module, expressed as:
[0048] Calculate the output feature of the spatial weight attention module according to the following formula:
[0049] Y 1 = V*Col
[0050] Y = Z*Y 1
[0051] Among them, V is equal to Q, Y 1 represents the output feature weight after passing through the spatial weight attention module, Y represents the output feature after passing through the spatial weight attention module, and Z represents the input feature of the spatial weight attention module.
[0052] Furthermore, in the training stage, the loss function adopted by the brain tumor image segmentation model is:
[0053]
[0054] Among them, represents the regularization loss, which is used to balance the contributions of different modalities; represents the segmentation loss, which is used to evaluate the segmentation effect of the final fused features; represents the multi-scale deep supervision loss, which is used to supervise the segmentation results at different levels.
[0055] Furthermore, and The calculation formulas of are as follows:
[0056]
[0057] Among them, represents the Dice loss, D reg represents the shared encoder, E m represents the encoder of the m-th modality, x m represents the input data of the encoder of the m-th modality, y represents the ground truth label, and M represents the set of modalities; represents the weighted cross loss; D fusion represents the fusion decoder, and Concat represents the feature concatenation operation; represents 2 l-1 upsampling operations, represents the fused feature at the l-th layer, and l represents the number of layers of the encoder.
[0058] In a second aspect, the present invention provides a missing modality brain tumor image segmentation system based on a multi-axis shifted MLP modality mask, including:
[0059] An image acquisition module, which is used to acquire the missing modality brain tumor image to be segmented;
[0060] An image segmentation module, which is used to segment the missing modality brain tumor image by using the trained brain tumor image segmentation model as described above to obtain a segmentation result.
[0061] Through the above technical solutions, it can be seen that compared with the prior art, the present invention has the following beneficial effects:
[0062] 1. Through the multi-level cross-modal interaction module, the present invention combines multi-scale and multi-level context information with global context information, effectively captures the complementary features between modalities, avoids the over-preference of high-level features for global context, thereby realizing multi-scale global perception, and enhancing the cross-modal feature extraction ability of the model.
[0063] 2. The present invention re - weights the modalities from different perspectives through the spatial weight attention module, effectively reducing the possibility that the fused features are dominated by a specific modality, enhancing its robustness in incomplete multi - modal situations, especially in the case of the absence of the dominant modality, and improving the model's ability to capture context information.
[0064] 3. The present invention uses a modality mask to learn the probability distribution of brain tumor images, effectively fusing cross - modal features for brain tumor segmentation, performing effective feature extraction and image segmentation on brain tumor images, thereby improving the accuracy of segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.
[0066] Figure 1 It is a flowchart of the method for segmenting brain tumor images with missing modalities based on multi - axis shifted MLP modality mask provided by the present invention;
[0067] Figure 2 It is a flowchart of training and testing the brain tumor image segmentation model provided by the present invention;
[0068] Figure 3 It is a schematic structural diagram of the brain tumor image segmentation model provided by the present invention;
[0069] Figure 4 It is a schematic structural diagram of the multi - level cross - modal interaction module provided by the present invention;
[0070] Figure 5 It is an operation schematic diagram of the spatial axial shifted MLP provided by the present invention;
[0071] Figure 6 It is a schematic workflow diagram of the spatial weight attention module provided by the present invention;
[0072] Figure 7 It is a schematic diagram of the channel - level fusion process provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0074] As Figure 1 shown, an embodiment of the present invention discloses a method for segmenting missing modality brain tumor images based on a multi-axis shifted MLP modality mask, including the following steps:
[0075] S1. Construct a brain tumor image segmentation model. The brain tumor image segmentation model includes four encoders and a shared decoder, and a multi-level cross-modal interaction module and a spatial weight attention module are introduced between the encoders and the shared encoder;
[0076] S2. Obtain brain tumor images containing four modalities as a training set and train the brain tumor image segmentation model; in the training stage, input the brain tumor images of the four modalities into the corresponding encoders respectively, and output the encoded features of different levels and different resolutions of each modality image; the multi-level cross-modal interaction module combines the global context information of the encoded features to capture complementary features between modalities; the spatial weight attention module re-weights the four modality features from different angles and inputs them into the shared decoder for decoding;
[0077] S3. After the training is completed, input the brain tumor image of the missing modality to be segmented into the trained brain tumor image segmentation model for image segmentation.
[0078] As Figure 2 shown, it is the training and testing process of the brain tumor image segmentation model.
[0079] Specifically, the brain tumor image segmentation model is a multi-axis shifted MLP modality mask model (M 3 ASTrans). As Figure 3 shown, E Flair 、E t1c 、E t1 、E t2 respectively represent the encoders of four different brain tumor modalities (T1, T1c, T2, Flair). Each encoder has multiple layers, and each layer has different scale perception capabilities. D reg represents the shared decoder. The shared decoder performs regularization so that modality-specific features are projected into a shared latent space, and each modality is trained separately for segmentation to reduce the negative impact of missing modalities.
[0080] The multi-level cross-modal interaction module (MCMIM) based on multi-axis shifted MLP and attention mechanism combines multi-scale and multi-level context information with global context, makes full use of the complementary features between modalities, adaptively fuses features in the spatial and channel dimensions, effectively uses the auxiliary local branch to generate local probability maps to constrain and guide the long-range modeling of the model, effectively avoids the over-preference of high-level features for global context, thus achieving multi-scale global perception, enhancing the cross-modal feature extraction ability of the model, ensuring high accuracy and reliability in brain tumor image segmentation, especially in the case of incomplete or missing modalities, and showing excellent robustness and precision.
[0081] By introducing the Spatial Weight Attention (SWA) module, the modalities are reweighted from different perspectives, effectively reducing the possibility of the fused features being dominated by a specific modality, enhancing its robustness in the case of incomplete multi-modalities, and improving the model's ability to capture context information. Specifically, in the embodiments of the present invention, the spatial weight attention mechanism is used for weighted cross-attention between the fused tokens and the specific modality tokens, and important tokens in each modality are gradually mined along the spatial dimension, so as to reweight the tokens from the accessible modalities and improve the feature fusion ability of the model.
[0082] In this embodiment, the modality mask model is combined with the fully supervised learning framework to improve the performance of the brain tumor image segmentation model with missing modalities, aiming to solve the problem of performance degradation caused by modality missing in multi-modal brain tumor image segmentation. This method can effectively fuse cross-modal features by using modality mask to learn the probability distribution of brain tumor images, and still maintain a high segmentation accuracy in the case of missing modalities. The introduction of the modality mask not only improves the effectiveness of feature extraction, but also helps the model better capture the tumor information under different modalities, and then optimizes the segmentation results. To further enhance the robustness of the model in the case of incomplete data, the spatial weight attention mechanism is introduced in the feature fusion process. This mechanism effectively avoids the over-dominance of a certain modality in the final features during the fusion process by adaptively assigning different weights to different modality features. Especially when the dominant modality is missing, the model can still maintain a high segmentation accuracy.
[0083] The working processes of the multi-level cross-modal interaction module and the spatial weight attention module are specifically described below.
[0084] As Figure 4As shown in the figure, the multi-level cross-modal interaction module integrates multi-level and multi-resolution features to achieve multi-scale global perception. This design effectively utilizes cascade operations and channel attention for channel weight adaptation, and inputs the fused features into the spatial axial shift MLP (VAS-MLP) for long-range dependency modeling. The specific working process includes:
[0085] (1) Scale the output features of different levels in the four encoders to the same resolution through downsampling operations; the encoded features corresponding to the five different layers of the four encoders are: Scaled to the same resolution through different convolution kernel sizes and strides where C represents the number of channels, H is the height, W is the width, and D is the depth.
[0086] (2) Concatenate the downsampled features through a concatenation operation, and adaptively adjust the weights of each channel through a channel attention mechanism to obtain fused features.
[0087] (3) Generate a local probability map after passing the fused features through a convolutional block and a spatial attention module.
[0088] (4) Input the fused features into the spatial axial shift MLP, and use the local probability map to constrain and guide the spatial axial shift MLP to model the long-range dependencies of the fused features, thus effectively avoiding the over-preference of high-level features for the global context.
[0089] The multi-level cross-modal interaction module is represented by the following formula:
[0090] X fuse = CA(Concat(DownSampling(x 1 ),
[0091] DownSampling(x 2 ),
[0092] DownSampling(x 3 ),
[0093] DownSampling(x 4 ),
[0094] DownSampling(x 5 )))
[0095] Y local = SA(CB(X fuse ))
[0096] Z = X fuse + Y local*VAS-MLP(X fuse )
[0097] where X fuse represents the output feature map after fusing the features of different corresponding layers in the four encoders, and x 1 , x 2 , x 3 , x 4 , x 5 respectively represent the corresponding encoded features of 5 different layers in the encoder; VAS-MLP represents the Spatial Axial Shift MLP; CA represents Channel Attention; Concat represents the concatenation operation; CB represents the Convolutional Block (Conv + IN + PReLU); SA represents Spatial Attention; Y local represents the local constraint weight generated by continuous convolution, that is, the local probability map; Z represents the output feature map of the multi-level cross-modal interaction module.
[0098] In a specific embodiment, as Figure 5 shown, for each target pixel, its neighboring pixels are aligned to the same channel through 4 different-direction shift operations, and then an MLP operation is performed along the channel dimension to achieve spatial information interaction. The feature maps containing different position dependencies are fused through matrix addition operations, and the fused features are aggregated through the channel MLP, and rich shift operations are performed to expand the actual effective receptive field. Through the effective receptive field of the Spatial Axial Shift MLP square, the dependence of the central anchor point on all surrounding neighboring pixels can be effectively captured, and equivalent global spatial information interaction is achieved through the iterative update of the central anchor point. The shift operation of the Spatial Axial Shift MLP on X fuse is expressed as:
[0099] X n = Norm(X fuse )
[0100] X lr = Proj(Shift lr (Proj(X n )))
[0101] X ld = Proj(Shift ld (Proj(X lr )))
[0102] X td = Proj(Shift td (Proj(X ld )))
[0103] X rd = Proj(Shift rd (Proj(Xtd )))
[0104] X = X lr + X ld + X td + X rd
[0105] X t = Norm(X)
[0106] X m = Proj(X t )
[0107] X y = X m + X fuse
[0108] Wherein, X fuse represents the input feature map; X n represents the mapped map after feature mapping normalization; X lr , X ld , X td and X rd represent the feature maps after shift operations in different directions; X represents the fused feature map containing different position dependencies; X t represents the mapped result after normalizing the feature map X; X m represents the feature X t after the mapped result of linear projection (Linearprojection); X y represents the output feature map of the spatial axial shift MLP.
[0109] Next, the working process of the spatial weight attention module is described as Figure 6 shown, specifically including:
[0110] The input features are divided into two categories: specific modality tokens and fused tokens;
[0111] The self-attention mechanism is used to directly concatenate and project the features of different modalities into Q, K, and V;
[0112] Construct a binary attention mask M ∈ {0, 1} 5N×5N to control the interaction rules between different tokens;
[0113] Given any position (i, j) in M, and q i ∈ Q, k j ∈ K; if q i and k j belong to the same and existing specific modality tokens, then set M i,j = 1; if q i belongs to the fused token, kj If it comes from the existing specific modality markers, then set M i,j = 1; if q i or k j both come from the missing modality, then set M i,j = 0; when M i,j = 1, allow an interaction relationship between q i and k j , and retain the attention weights at the corresponding positions; when M i,j = 0, do not allow an interaction relationship between q i and k j , and filter out the attention weights at the corresponding positions;
[0114] Dynamically adjust the attention weight matrix through the binary attention mask M, and calculate the weights of each modality along the spatial dimension to re-weight the existing modalities. Specifically, calculate the similarity through the dot product of Q and K, then obtain the weight A through softmax, and then filter out the invalid interactions through the binary attention mask M to obtain the attention weight matrix MA, which can be expressed as MA = A ⊙ M, where ⊙ represents element-wise multiplication, and retain the attention connections allowed by the mask.
[0115] In this embodiment, only operations on Q and K are involved, which is to assist in the construction of the binary attention mask and adjust the attention weight matrix.
[0116] Specifically, calculate the weights of all modalities through column vector summation:
[0117]
[0118] where H represents the number of attention heads; N = H × W × D, representing the total number of voxels in the single-layer encoded feature map, H is the height of the single-layer encoded feature map, W is the width of the single-layer encoded feature map, and D is the depth of the single-layer encoded feature map; MA(i,j) represents the modality-specific feature weights, and (i,j) represents different positions in the feature dimension 1 × N;
[0119] As Figure 7 shown, after re-weighting the modality-specific features, perform channel-level fusion of cross-modal features along the channel dimension for the accessible modality features, where V is equal to Q, multiply the modality-specific feature weights of each modality marker by V to obtain the output feature weights of the spatial weight attention module, and then multiply it by the input features to obtain the output features of the spatial weight attention module, which is expressed as:
[0120] Y 1 = V * Col
[0121] Y = Z * Y 1
[0122] Y represents the output features after passing through the spatial weight attention module, Z represents the input features of the spatial weight attention module, and Y 1 represents the output feature weights after passing through the spatial weight attention module.
[0123] The token weights for each modality m ∈ M can be obtained by slicing represents the feature space dimension, and Split represents the slicing operation.
[0124] In addition, in the multi-modal framework, the features of different modalities are directly fused and fed into the decoder for segmentation. The decoder tends to select the most discriminative modality as the main modality for brain tumor segmentation. Especially when the main modality is missing, it will lead to a serious performance decline. To avoid modality bias and balance different modalities, the embodiment of the present invention introduces a shared decoder D reg for regularization. In this way, the modality-specific features are projected into a shared latent space, and each modality is separately trained for segmentation to reduce the negative impact of the missing modality.
[0125] In the training phase, the loss function adopted by the brain tumor image segmentation model is:
[0126]
[0127] where represents the regularization loss, which is used to balance the contributions of different modalities; represents the segmentation loss, which is used to evaluate the segmentation effect of the final fused features; represents the multi-scale deep supervision loss, which is used to supervise the segmentation results at different levels.
[0128] and The calculation formulas of are respectively:
[0129]
[0130] where In the formula, each modality separately extracts features through the encoder E m and the shared decoder D reg maps the features to the shared space, calculates the sum of the Dice and weighted cross-entropy losses with the ground truth labels, and finally sums the losses of all modalities to balance the contributions of different modalities.
[0131] In the formula, the encoded features of each modality are concatenated and input into the fusion decoder D fusion to generate the segmentation result, and calculate the sum of the Dice and weighted cross-entropy losses with the ground truth labels. This loss ensures that the model can effectively fuse multi-modal information for final segmentation.
[0132] In the formula, fused features F are extracted at different levels of the encoder l fusion , after upsampling to the original resolution, the segmentation loss (Dice + weighted cross-entropy) is calculated layer by layer to achieve multi-scale supervision. This design enhances the model's learning ability for local details and global context.
[0133] Specifically, represents the Dice loss, D reg represents the shared encoder, E m represents the encoder of the m-th modality, x m represents the input data of the encoder of the m-th modality, y represents the ground truth label, and M represents the set of modalities; represents the weighted cross loss; D fusion represents the fusion decoder, and Concat represents the feature concatenation operation; represents 2 l-1 upsampling operations, represents the fused features at the l-th layer, and l represents the number of layers of the encoder.
[0134] In the present invention, through the method of fully supervised learning, a large amount of labeled data is used for the segmentation of missing modality brain tumor images. This method makes full use of the detailed information in the labeled data, thereby improving the accuracy of the model in the task of missing modality brain tumor segmentation. Fully supervised learning relies on accurate labeled data to ensure that the model can learn the complex features of brain tumors and exhibit excellent performance in image segmentation. Especially in medical image analysis, accurate segmentation results are crucial for the formulation of diagnosis and treatment plans. Therefore, segmenting missing modality brain tumor images through fully supervised learning not only improves the segmentation accuracy but also provides reliable support for medical decision-making.
[0135] In addition, by adopting a multi-level cross-modal interaction module based on multi-axis shifted MLP and attention mechanism, the model can fuse multi-scale and multi-level context information, and effectively combine global context information. It not only optimizes the complementary feature capture ability between modalities but also enhances the feature integration in the spatial and channel dimensions through adaptive cross-modal interaction, improving the model's feature extraction ability. Introducing a spatial weight attention mechanism during the fusion process enables the model to re-weight the features of each modality from different angles, thereby effectively reducing the dominant influence of a specific modality on the fused features. This strategy significantly enhances the robustness of the model when multi-modal data is incomplete or the dominant modality is missing, ensuring the balance and accuracy of feature fusion, further improving the accuracy and robustness of missing modality brain tumor image segmentation, and providing a more efficient and reliable solution for multi-modal data processing in medical image analysis.
[0136] In other embodiments, the present invention further provides a missing modality brain tumor image segmentation system based on a multi-axis shifted MLP modality mask, including:
[0137] An image acquisition module, configured to acquire a missing modality brain tumor image to be segmented;
[0138] An image segmentation module, configured to segment the missing modality brain tumor image by using the trained brain tumor image segmentation model as described above to obtain a segmentation result.
[0139] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part.
[0140] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A missing modality brain tumor image segmentation method based on multi-axis shift MLP modality mask, characterized in that: The following steps are involved: Constructing a brain tumor image segmentation model, the brain tumor image segmentation model includes four encoders and one shared decoder, and introducing a multi-level cross-modal interaction module and a spatial weight attention module between the encoders and the shared encoder; Acquire brain tumor images containing four modalities as a training set, and train the brain tumor image segmentation model; In the training phase, the brain tumor images of the four modalities are respectively input into the corresponding encoders, and the encoding features of different levels and resolutions of each modality image are output; the multi-level cross-modal interaction module combines the global context information of the encoding features to capture the complementary features between the modalities; the spatial weight attention module re-weights the four modality features from different angles and inputs them into the shared decoder for decoding; After the training is completed, the brain tumor image with the missing modality to be segmented is input into the trained brain tumor image segmentation model to perform image segmentation.
2. The missing modality brain tumor image segmentation method based on multi-axis shift MLP modality mask according to claim 1 is characterized in that: Each of the encoders has multiple layers, each layer having different scale perception capabilities.
3. The missing modality brain tumor image segmentation method based on multi-axis shift MLP modality mask according to claim 1 is characterized in that: The working process of the multi-level cross-modal interaction module includes: Scaling the output features of different levels in the four encoders to the same resolution through a downsampling operation; The downsampled features are concatenated through cascade operations, and the weights of each channel are adaptively adjusted through the channel attention mechanism to obtain fused features; After passing the fused features through the convolution block and the spatial attention module, a local probability map is generated; The fused features are input into the spatial axial shift MLP, and the local probability map is used to constrain and guide the spatial axial shift MLP to model the long-range dependency of the fused features.
4. The method for missing modality brain tumor image segmentation based on multi-axis shift MLP modality mask according to claim 3, characterized in that: The multi-level cross-modal interaction module is represented by the following formula: X fuse =CA(Concat(DownSampling(x1), DownSampling(x2), DownSampling(x3), DownSampling(x4), DownSampling(x5))) Y local =SA(CB(X fuse )) Z=X fuse +Y local *VAS-MLP(X fuse ) Among them, X fuse represents the output feature map after the fusion of the features of different corresponding layers in the four encoders, x1, x2, x3, x4, x5 represent the corresponding encoding features of 5 different layers in the encoder respectively; VAS-MLP represents spatial axial shift MLP; CA represents channel attention; Concat represents concatenation operation; CB represents convolution block; SA represents spatial attention; Y local Represents the local constraint weight generated by continuous convolution, that is, the local probability map; Z represents the output feature map of the multi-level cross-modal interaction module.
5. The missing modality brain tumor image segmentation method based on multi-axis shift MLP modality mask according to claim 4 is characterized in that: The spatial axial shift of the MLP to X fuse The shift operation is expressed as: X n =Norm(X fuse ) X lr =Proj(Shift lr (Proj(X n ))) X ld =Proj(Shift ld (Proj(X lr ))) X td =Proj(Shift td (Proj(X ld ))) X rd =Proj(Shift rd (Proj(X td ))) X=X lr +X ld +X td +X rd X t =Norm(X) X m =Proj(X t ) X y =X m +X fuse Among them, X fuse Represents the input feature map; X n represents the normalized feature map; X lr , X ld , X td and X rd represents the feature map after the shift operation in different directions; X represents the fusion of feature maps containing different position dependencies; X t Represents the mapping result after the feature map X is normalized; X m Represents feature X t The mapping result after linear projection; X y Represents the output feature map of the spatial axial shift MLP.
6. The missing modality brain tumor image segmentation method based on multi-axis shift MLP modality mask according to claim 1, characterized in that: The working process of the spatial weight attention module includes: The input features are divided into two categories: modality-specific tags and fusion tags; The self-attention mechanism is used to directly concatenate the features of different modalities to Q, K, and V; Construct a binary attention mask M∈{0,1} 5N×5N , which controls the interaction rules between different tags; Given any position (i,j) in M, and q i ∈Q,k j ∈K; if q i and k j belongs to the same specific modal tag and exists, then set M i,j =1; if q i Belongs to the fusion marker, k j From the existence of a specific modal marker, set M i,j =1; if q i or k j All come from the missing mode, then set M i,j =0; when M i,j =1, allowing q i and k j There is an interactive relationship between them, and the attention weights of the corresponding positions are retained; when M i,j = 0, q is not allowed i and k j There is an interactive relationship between them, and the attention weights of the corresponding positions are filtered out; The attention weight matrix is dynamically adjusted through the binary attention mask M, and the weight of each modality is calculated along the spatial dimension to re-weight the existing modalities.
7. The missing modality brain tumor image segmentation method based on multi-axis shift MLP modality mask according to claim 6 is characterized in that: The weights of all modes are calculated by summing the column vectors: Where H represents the number of attention heads; N = H × W × D, represents the total number of voxels in a single-layer encoding feature map, H is the height of a single-layer encoding feature map, W is the width of a single-layer encoding feature map, and D is the depth of a single-layer encoding feature map; MA(i, j) represents the modality-specific feature weight, and (i, j) represents different positions of the feature dimension 1×N; The output features of the spatial weighted attention module are calculated as follows: Y1=V*Col Y=Z*Y1 Among them, V is equal to Q, Y1 represents the output feature weight after the spatial weight attention module, Y represents the output feature after the spatial weight attention module, and Z represents the input feature of the spatial weight attention module.
8. The missing modality brain tumor image segmentation method based on multi-axis shift MLP modality mask according to claim 1, characterized in that: During the training phase, the loss function used by the brain tumor image segmentation model is: in, represents the regularization loss, which is used to balance the contribution of different modes; Represents the segmentation loss, which is used to evaluate the segmentation effect of the final fusion feature; Represents a multi-scale deep supervision loss, which is used to supervise the segmentation results at different levels.
9. The method for missing modality brain tumor image segmentation based on multi-axis shift MLP modality mask according to claim 8, characterized in that: and The calculation formulas are: in, represents the Dice loss, D reg represents the shared encoder, E m represents the encoder of the mth mode, x m represents the input data of the encoder of the mth mode, y represents the true label, and M represents the mode set; represents the weighted cross loss; D fusion represents a fusion decoder, and Concat represents a feature concatenation operation; Representation 2 l-1 Upsampling operation, represents the fused features of the lth layer, and l represents the number of encoder layers.
10. A missing modality brain tumor image segmentation system based on multi-axis shift MLP modality mask, characterized in that: include: An image acquisition module, used for acquiring a missing modality brain tumor image to be segmented; An image segmentation module is used to segment the missing modality brain tumor image using the trained brain tumor image segmentation model as described in any one of claims 1 to 9 to obtain a segmentation result.
Citation Information
Cited By
Brain tumor segmentation method and system based on anatomical perception symmetric comparison and cross-modal migration
CN121147242A
A brain tumor segmentation method and system based on anatomical perception symmetry comparison and cross-modal transfer
CN121147242B
Adapter-based interactive camouflage target segmentation method, electronic equipment and storage medium
CN121353310A
Adapter-based interactive camouflage target segmentation method, electronic device and storage medium
CN121353310B