Glioma deletion mode segmentation method based on MST-KDNet neural network

The MST-KDNet neural network addresses the challenge of brain tumor segmentation with missing MRI modalities by using multi-scale Transformer knowledge distillation and dual-mode Logit distillation to enhance feature learning and adaptability, achieving high accuracy and efficiency in brain tumor segmentation.

CN120318241APending Publication Date: 2025-07-15HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510330444.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the prior art, when dealing with MRI missing modality, model performance deteriorates and student models cannot accurately learn key features of teacher models, and lacks training stability and flexibility.

Method used

The MST-KDNet neural network is adopted, combining the multi-scale Transformer knowledge distillation module (MS-TKD), global style matching module (GSME) and dual-mode Logit distillation module (DMLD), and the attention weights of different resolutions are extracted through multi-scale Transformer knowledge distillation, and the key features are accurately learned by extreme distillation, and the knowledge adaptability is improved through Logit alignment and standardized KL distillation, and the model expression ability is enhanced by combining feature matching and adversarial learning.

Benefits of technology

It significantly improves the segmentation accuracy and stability of the student model under the missing mode, and achieves efficient deletion mode segmentation of brain glioma.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318241A_ABST
    Figure CN120318241A_ABST
Patent Text Reader

Abstract

The invention discloses a glioma deletion mode segmentation method based on an MST-KDNet neural network, and the method comprises the steps: firstly constructing an MST-KDNet neural network model composed of a multi-scale Transform knowledge distillation module, a global style matching module and a dual-mode Logit distillation module, enabling the MST-KDNet neural network model to be used for the segmentation of a glioma deletion mode, and outputting a segmentation result graph; secondly, establishing a magnetic resonance imaging data set related to glioma, and dividing the magnetic resonance imaging data set into a training set and a test set; and finally, training the MST-KDNet neural network model by using the training set, and performing parameter tuning and testing in combination with the characteristics of the multi-modal glioma image. According to the method, the problems of insufficient training stability and flexibility of the student model are solved, and the segmentation accuracy is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of computer vision deep learning and medical image processing, and particularly relates to a method for segmenting missing modalities of brain gliomas based on the MST-KDNet neural network. Background Art

[0002] Brain tumor segmentation is a key task in medical neuroimaging analysis and is of great significance for diagnosis, treatment planning, and prognosis assessment. Among brain tumors, gliomas are particularly invasive, and their complex biological behavior and heterogeneity pose great challenges to clinical treatment.

[0003] Magnetic resonance imaging (MRI) containing multiple modalities is the preferred method for visualizing and segmenting brain tumors. In brain tumor segmentation, different modalities contain unique information. For example, T1 and T2 modalities help identify vasogenic edema in subacute strokes; the T1-Gd (contrast-enhanced) modality reveals details of vascular structures and the blood-brain barrier; the FLAIR modality provides general information about stroke lesions. The mutual complementarity of different modalities can provide detailed information about the tumor size, location, and morphology. However, due to scan damage, artifacts, incorrect machine settings, allergies to certain contrast agents, and limited available scan time, missing modalities are common in the actual clinical environment. Moreover, in clinical practice, MRI may be obtained from different machines under different acquisition protocols and parameters, and there is no guarantee of the number of available modalities, nor can it guarantee that any modality will be correctly labeled for algorithm use, and the segmentation accuracy is often limited. To address the challenges brought by missing modalities, Ding et al. proposed a region-aware fusion module (RFM), Zhao et al. proposed modality-adaptive feature interaction (MFI) with multi-modal codes, Zhang et al. proposed an unlabeled multi-modal adversarial generation network, and Dai et al. calibrated the missing modality representation to the global full-modal anchor through scaled dot-product cross-attention. However, it is common that the contribution of some modalities is relatively greater, and if these important modalities are missing, the model performance will decrease significantly. To solve this problem, Wang et al. proposed a learnable cross-modal knowledge distillation (LCKD) model, Liu et al. proposed M3AE; Huo et al. gradually transmitted cross-modal knowledge through bidirectional distillation; furthermore, Xing et al. designed local topology-preserving and difference-eliminating contrast distillation (TDC-Distill). The above methods highlight the effectiveness of knowledge distillation in dealing with missing modalities and improving model performance. However, the above methods still have problems such as the student model being unable to accurately learn the key features of the teacher model, and the training stability and flexibility of the student model being insufficient. Summary of the Invention

[0004] To address the above problems, the present invention provides a method for segmenting missing modalities of brain gliomas based on the MST-KDNet neural network. The MST-KDNet neural network model consists of a multi-scale Transformer knowledge distillation module (MS-TKD), a global style matching module (GSME), and a dual-mode Logit distillation module (DMLD), which can achieve automatic and intelligent segmentation of missing modalities of brain gliomas, and has high classification accuracy and efficiency.

[0005] The present invention extracts attention weights at different resolutions through multi-scale Transformer knowledge distillation (MS-TKD), and uses the designed extreme value distillation method to enable the student model to more accurately learn the key features of the teacher model. At the same time, the innovatively proposed dual-mode Logit distillation (DMLD) greatly improves the knowledge adaptation ability through Logit alignment and normalized KL distillation. The innovatively designed global style matching module (GSME) combines feature matching and adversarial learning well, enhancing the model's expression ability in the case of missing modalities. The method includes the following steps:

[0006] Step a: Construct an MST-KDNet neural network model for glioma segmentation, which is used for segmenting missing modalities of gliomas and outputs a segmentation result map. The MST-KDNet neural network model consists of a multi-scale Transformer knowledge distillation module (MS-TKD), a global style matching module (GSME), and a dual-mode Logit distillation module (DMLD).

[0007] The multi-scale Transformer knowledge distillation module (MS-TKD) consists of a 3D convolutional encoder, a multi-scale Transformer unit, and its extreme value distillation subunit (EVD).

[0008] The global style matching module (GSME) performs max-pooling and concatenation on the outputs of multiple convolutional layers, and then extracts features through a Transformer block to calculate the adversarial loss and MSE loss. By minimizing the loss, the output features of high and low convolutional layers are maintained, realizing accurate segmentation of gliomas.

[0009] The dual-mode Logit distillation module (DMLD) consists of a Logit divergence distillation unit and a Logit normalized KL loss distillation unit.

[0010] Step b: Establish a magnetic resonance imaging dataset for gliomas, which is divided into a training set and a test set.

[0011] Step c: Use the training set to train the MST-KDNet neural network model, and optimize the parameters of the MST-KDNet neural network model in combination with the characteristics of multimodal glioma images to achieve the best segmentation effect of the model.

[0012] Step d: Use the test set to test and evaluate the obtained MST-KDNet network model, and finally realize the automatic and intelligent brain glioma missing modality segmentation function.

[0013] Preferably, the multi-scale Transformer knowledge distillation module (MS-TKD) in step a first respectively passes the complete modality data and the missing modality data after data processing through a 3D convolution layer to obtain the corresponding initial features and inputs them to the teacher model and the student model, and completes the downsampling process through two 3D convolution encoders of the corresponding models. Then, the two feature maps after downsampling are converted into one-dimensional sequences: First, divide the input x ′ into flattened uniform non-overlapping patches. Then project these patches into the embedding space of K dimensions through a linear layer and add learnable position encoding to obtain two feature sequences. Then, pass the feature sequences in the teacher model and the student model through a series of Transformer blocks in the multi-scale Transformer unit. The MSA sub-layer in the Transformer block consists of n parallel self-attention (SA) blocks. Specifically, the self-attention (SA) block is a parameterized function that learns the mapping relationship between the query (q), key (k), and value (v) representations in the sequence z.

[0014] Subsequently, the attention weights A of each self-attention block in each layer are taken out through the extreme value distillation sub-unit (EVD) of the Transformer unit as the sequence As for extreme value distillation, and the maximum value A max and the minimum value A min and the mean value A mean of the attention weights at each pixel point are obtained:

[0015] A max = max(As, axis = 1)

[0016] A min = min(As, axis = 1)

[0017]

[0018] Next, multiply the weights of the teacher model and the student model by their corresponding pixel points through the Transformer block respectively to obtain three sequences output by the teacher model and three sequences output by the student model And calculate the mean squared error (MSE) loss L for the sequences corresponding to the two models EVD .

[0019] Then extract the sequence representation z output by every three Transformer blocks in the multi-scale Transformer unit i (i ∈ {3, 6, 9, 12}, where i is the number of Transformer blocks), with a size of And reshape it into a tensor of . This module calculates the loss L using MSE for the sequence representations z of the teacher model and the student model i respectively z :

[0020] Then, the feature maps output by the last layer of Transformer blocks are respectively subjected to skip connection operations with the feature maps output by the previous layer of Transformer blocks after two consecutive 3D convolutions combined with normalization operations, and the output is upsampled using a deconvolution layer to restore to the original shape before inputting into the Transformer block; the final output is input into a 1×1×1 convolution layer to obtain the feature f t , and finally, the total loss of this module is obtained by multiplying these two MSE losses by the weights α and β:

[0021] L MS-TKD = αL EVD + βL z

[0022] Preferably, the global style matching module (GSME) described in step a combines the MSE loss with the distillation method of adversarial learning. This module will perform max pooling on the features output by the penultimate layer convolution of the 3D convolution encoder in the multi-scale Transformer knowledge distillation module and then concat them with the output of the 3D convolution encoder to obtain f enc , and then concat the output f after passing through the multi-scale Transformer knowledge distillation module t with f enc and input the concatenated result into a layer of 3D convolution to obtain f enc&t , and input it into the feature discriminator D to obtain the adversarial loss:

[0023]

[0024] Subsequently, f enc&t is input into the 3D convolution decoder, and at the same time, the features obtained from the first layer convolution and the second layer convolution of the 3D convolution decoder are also concatenated to obtain f dec , and use f enc , f t , f decThe H, W, and D dimensions are flattened to obtain a two-dimensional tensor G enc , G t , G dec Then, a feature fusion operation is performed:

[0025]

[0026] After obtaining the fused feature sequence M ∈ {M enc&dec , M enc&t , M edc&t}, the MSE loss is calculated:

[0027] The total loss of this module is:

[0028] L GSM = εL adv + θL enc&dec&t

[0029] Finally, the feature map obtained by the 3D convolutional decoder for f enc&t is input into the dual-mode Logit distillation module.

[0030] Preferably, the dual-mode Logit distillation module (DMLD) described in step a adjusts the number of channels of the input feature map through a 1×1×1 convolution to obtain a feature map I. Then, the feature maps l of the student model and the teacher model are logit-normalized, that is, l is input into the Z-score normalization function and then passed through the softmax function to obtain the features q(l f ), q(l m ):

[0031]

[0032] q(l f ) = softmax[Z(l f ; τ)]

[0033] q(l m ) = softmax[Z(l m ; τ)]

[0034] where μ is the mean of l, σ is the variance of l, and τ is the temperature coefficient. Then, the mean squared error L mse and the KL loss L KL are calculated for the outputs respectively;

[0035] Then, the sigmoid activation function is used to generate the segmentation result, and the total loss of the dual-mode Logit distillation module is calculated through these two losses. The formula is as follows:

[0036] L logit = λ mse L mse + λKD τ 2 L KL

[0037] Among them, λ mse and λ KD are corresponding weight coefficients.

[0038] Then, the segmentation results obtained by the sigmoid activation functions of the student model and the teacher model are used to calculate the Dice loss L dice with the true segmentation map. Finally, three modules use L dice to calculate the total loss for knowledge distillation, and by minimizing the loss, the performance of the student model in the case of missing modalities is significantly improved.

[0039] L joint = λ1λ MS-TKD + λ2L logit + λ3L GSM + L dice

[0040] Among them, λ1, λ2, and λ3 are corresponding weight coefficients.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] The main contributions of the present invention are in three aspects: extracting attention weights of different resolutions through multi-scale Transformer knowledge distillation (MS-TKD), and using the designed extreme value distillation method to enable the student model to more accurately learn the key features of the teacher model; at the same time, the innovatively proposed dual-mode Logit distillation (DMLD) greatly improves the knowledge adaptation ability through Logit alignment and normalized KL distillation; while the innovatively designed global style matching module (GSME) combines feature matching and adversarial learning well, enhancing the expression ability of the model in the case of missing modalities. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is the overall structure diagram of the MST-KDNet neural network model of the present invention;

[0044] Figure 2 is the structure diagram of the multi-scale Transformer unit in the multi-scale Transformer knowledge distillation module (MS-TKD) of the MST-KDNet neural network model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below by combining the accompanying drawings and specific embodiments. The following embodiments are only descriptive and not restrictive, and the protection scope of the present invention cannot be limited thereby.

[0046] The present invention provides a method for segmenting missing modalities of brain gliomas based on the MST-KDNet neural network, as Figure 1 and Figure 2 shown. The MST-KDNet neural network model consists of a multi-scale Transformer knowledge distillation module (MS-TKD), a global style matching module (GSME), and a dual-mode Logit distillation module (DMLD). It can realize the automatic and intelligent function of segmenting missing modalities of brain gliomas, and has a high classification accuracy and efficiency.

[0047] The method includes the following steps:

[0048] Step a: Construct an MST-KDNet neural network model for glioma segmentation, which is used for segmenting missing modalities of gliomas and outputs a segmentation result map. The MST-KDNet neural network model consists of a multi-scale Transformer knowledge distillation module (MS-TKD), a global style matching module (GSME), and a dual-mode Logit distillation module (DMLD).

[0049] The multi-scale Transformer knowledge distillation module (MS-TKD) consists of a 3D convolutional encoder, a multi-scale Transformer unit, and its extreme value distillation subunit (EVD).

[0050] The global style matching module (GSME) performs max pooling and concatenation on the outputs of multiple convolutional layers, and then extracts features through a Transformer block to calculate the adversarial loss and MSE loss. By minimizing the loss, the output features of the high and low convolutional layers are maintained, and the precise segmentation of gliomas is achieved.

[0051] The dual-mode Logit distillation module (DMLD) consists of a Logit divergence distillation unit and a Logit-normalized KL loss distillation unit.

[0052] Step b: Establish a magnetic resonance imaging dataset for gliomas, which is divided into a training set and a test set.

[0053] Step c: Use the training set to train the MST-KDNet neural network model, and combine the characteristics of multi-modal glioma images to optimize the parameters of the MST-KDNet neural network model to make the segmentation effect of the model reach the best.

[0054] Step d: Use the test set to test and evaluate the obtained MST-KDNet network model, and finally realize the automatic and intelligent function of segmenting missing modalities of brain gliomas.

[0055] Preferably, in step a, the multi-scale Transformer knowledge distillation module (MS-TKD) first respectively inputs the complete modal data and the missing modal data after data processing into a teacher model and a student model through a 3D convolution layer to obtain corresponding initial features, and completes downsampling through two 3D convolution encoders of the corresponding models. Then, the two feature maps after downsampling are converted into one-dimensional sequences: First, the input x ′ is divided into flattened uniform non-overlapping patches. Then, these patches are projected into an embedding space of dimension K through a linear layer and learnable position encoding is added to obtain two feature sequences. Then, the feature sequences in the teacher model and the student model pass through a series of Transformer blocks in the multi-scale Transformer unit. The MSA sub-layer in the Transformer block consists of n parallel self-attention (SA) blocks. Specifically, the self-attention (SA) block is a parameterized function that learns the mapping relationship between the query (q), key (k), and value (v) representations in the sequence z. The attention weight (A) is calculated based on the similarity between two elements in z and their key-value pairs:

[0056]

[0057] where K h = K / n is a scaling factor. Using the calculated attention weights, the result of the value v in the sequence z after passing through SA is calculated:

[0058] SA(z) = Av

[0059] Subsequently, the attention weights A of each self-attention block in each layer are taken out through the extreme value distillation sub-unit (EVD) of the Transformer unit as the sequence As for extreme value distillation, and the maximum value A max of the attention weights at each pixel point, the minimum value A min , and the average value A mean are obtained:

[0060] A max = max(As, axis = 1)

[0061] A min = min(As, axis = 1)

[0062]

[0063] Next, the weights of the teacher model and the student model are respectively multiplied by their corresponding pixel points through the Transformer block to obtain three sequences output by the teacher model and three sequences output by the student model And calculate the mean squared error (MSE) loss L for the sequences corresponding to the two models EVD as follows:

[0064]

[0065] Then extract the sequence representation z of the output of every three Transformer blocks in the multi-scale Transformer unit i (i ∈ {3, 6, 9, 12}, where i is the number of Transformer blocks), with a size of and reshape it into a tensor of This module calculates the loss L for the sequence representations z of the teacher model and the student model respectively using MSE i as follows: z :

[0066]

[0067] Then, the feature maps of the output of the last Transformer block are respectively subjected to skip connection operations with the feature maps of the output of the previous Transformer block after two consecutive 3D convolutions with normalization operations, and the output is upsampled using a deconvolution layer to restore to the original shape before inputting into the Transformer block; the final output is input into a 1×1×1 convolution layer to obtain the feature f t , and finally the total loss of this module is obtained by multiplying these two MSE losses by the weights α and β:

[0068] L MS-TKD = αL EVD + βL z

[0069] Preferably, the global style matching module (GSME) described in step a combines the MSE loss with the distillation method of adversarial learning. This module will perform max pooling on the features of the output of the second-to-last convolution of the 3D convolution encoder in the multi-scale Transformer knowledge distillation module and then concat it with the output of the 3D convolution encoder to obtain f enc , and then concatenate the output f t after passing through the multi-scale Transformer knowledge distillation module with f enc and input the concatenated result into a 3D convolution layer to obtain f enc&t , and input it into the feature discriminator D to obtain the adversarial loss:

[0070]

[0071] Subsequently, f enc&tInput it into the 3D convolutional decoder, and at the same time, splice the features obtained from the first and second convolutions of the 3D convolutional decoder to obtain f dec , and for f enc , f t , f dec , flatten the H, W, and D dimensions of f to obtain the two-dimensional tensor G enc , G t , G dec Then perform a feature fusion operation:

[0072]

[0073] After obtaining the fused feature sequence M ∈ {M enc&dec , M enc&t , M edc&t}, calculate the MSE loss:

[0074]

[0075] The total loss of this module is:

[0076] L GSM = εL adv + θL enc&dec&t

[0077] Finally, input the feature map obtained by passing f enc&t through the 3D convolutional decoder into the dual-mode Logit distillation module.

[0078] Preferably, the dual-mode Logit distillation module (DMLD) described in step a adjusts the number of channels of the input feature map through 1×1×1 convolution to obtain the feature map I. Then, normalize the feature maps l of the student model and the teacher model by logit, that is, input l into the Z-score normalization function and then pass through the softmax function to obtain the features q(l f ), q(l m ):

[0079]

[0080] q(l f ) = softmax[Z(l f ; τ)]

[0081] q(l m ) = softmax[Z(l m ; τ)]

[0082] where μ is the mean of l, σ is the variance of l, and τ is the temperature coefficient. Then calculate the mean square error and KL loss for the outputs respectively:

[0083]

[0084] Then, the sigmoid activation function is adopted to generate the segmentation result, and the total loss of the dual-mode Logit distillation module is calculated through these two losses. The formula is as follows:

[0085] L logit = λ mse L mse + λ KD τ 2 L KL

[0086] where λ mse and λ KD are the corresponding weight coefficients.

[0087] Then, the segmentation results obtained from the sigmoid activation functions of the student model and the teacher model are used to calculate the Dice loss L dice with the ground truth segmentation map. Finally, the total loss is calculated by these three modules with L dice for knowledge distillation, and the performance of the student model in the case of missing modalities is significantly improved by minimizing the loss.

[0088] L joint = λ1L MS-TKD + λ2L logit + λ3L GSM + L dice

[0089] where λ1, λ2, and λ3 are the corresponding weight coefficients.

[0090] The glioma BraTS2024 dataset described in step b is the Multimodal Brain Tumor Segmentation Challenge (BraTS) 2024 dataset. This dataset provides 3D multimodal magnetic resonance imaging of the brain and the corresponding ground truth annotations. The dataset contains a total of 1080 cases, and each case consists of magnetic resonance images of four modalities: T1CE, T1, T2, and FLAIR. The image regions are divided into three categories: enhancing tumor (ET), core tumor (CT), and whole tumor (WT). Randomly select 80% of the data in the dataset as the training set, and the remaining 20% as the test set.

[0091] The glioma FeTS2024 dataset described in step b is the Multimodal Brain Tumor Segmentation Challenge (FeTS) 2024 dataset. This dataset is used to study federated learning through tumor segmentation tasks and is jointly composed of the BraTS dataset and data provided by FeTS federated learning task collaborators. Each sample is associated with a classification label for its acquisition institution. The dataset contains a total of 1250 cases, providing 3D multimodal brain magnetic resonance images (T1, T2, T1C, and FLAIR) of the corresponding cases and their true annotations (ET, CT, WT). To ensure that the segmentation and preprocessing methods of this dataset are consistent with the BraTS2024 dataset.

[0092] Preferably, in step c, the MST-KDNet neural network model is trained using the training set. Combining the characteristics of glioma images, it has been implemented through the PyTorch framework and trained using NVIDIA Tesla V100. During the training process, the batch size is set to 1, and the parameters in the model are updated through the Adam optimization algorithm, where the initial learning rate is 0.0001, the β1 value is 0.9, and the β2 value is 0.99.

[0093] Preferably, in step d, the obtained MST-KDNet network model is tested and evaluated using the test set, and finally an automatic and intelligent multimodal glioma missing modality segmentation function is achieved.

[0094] To evaluate the classification performance of the model on images, MST-KDNet is compared with other classification models on the glioma Brats2024 and FeTS2024 datasets. Two standard metrics are used to evaluate the classification performance. They are: Dice scores, which measure the degree of overlap of regions, and HD95 scores, which measure the degree of overlap of boundaries. Their formulas are as follows:

[0095]

[0096] Where |X∩Y| is the intersection between X and Y, |X| and |Y| respectively represent the number of elements in X and Y, x and y are elements in X and Y, d represents the distance, and HD95 is calculated based on the 95th percentile of the distances between boundary points in X and Y.

[0097] Randomly discard input modalities on the BraTS2024 dataset to simulate comparative experiments with different combinations such as single-modal, dual-modal, triple-modal, and quadruple-modal, and compare the MST-KDNet with methods such as smunet, acn, mmcformer, RAHVED, RMBTS, and mmformer. In terms of the HD95 metric, MST-KDNet achieved the best and the second-best or tied-best results in the whole tumor region (WT), tumor core region (TC), and enhanced tumor region (ET) respectively, reflecting the precise localization ability of the tumor boundary. At the same time, in terms of the Dice metric, MST-KDNet ranked first in the average performance of the Dice score metric, showing an overall advantage in balancing voxel overlap and spatial distance. It is worth noting that even in extreme situations where only single-modal or dual-modal data is available, MST-KDNet still maintains a relatively high Dice and a low HD95; when all four modalities (FL, T1, T1c, T2) are available, the Dice and HD95 metrics of MST-KDNet perform even better.

[0098] The comparative experiment conducted on the FeTS2024 dataset is the same as above. The average Dice metric of MST-KDNet in the three regions of the whole tumor region (WT), tumor core region (TC), and enhanced tumor region (ET) is better than other methods in most missing modality combinations. It is worth noting that when only a single modality (such as only flair or only t1) remains, MST-KDNet can still outperform the control method by 2-3%, showing the strong robustness of the model in the case of severe modality loss; in the case where all four modalities (flair, t1, t1c, t2) are available, the Dice of WT further exceeds 91%.

[0099] Under different combinations of missing modalities, the present invention conducts ablation by removing multi-scale Transformer knowledge distillation (MS-TKD), global style matching (GSM), and standardized Logit distillation (SLKD) to examine the actual contributions of each module. First, on BRATS2024, removing MS-TKD causes the Dice of WT / Core / Enhance to decrease by approximately 3.4%, 7.4%, and 7.7% respectively, and the HD95 to increase significantly, indicating that its multi-scale attention alignment is crucial for feature encoding; removing GSM results in a Dice decline of approximately 4.5%, 6.6%, and 6.2%, suggesting that global style and low-level texture compensation are indispensable for maintaining structural integrity; eliminating SLKD also causes a Dice loss of approximately 4.2%, 5.2%, and 4.7%, highlighting the importance of flexible teacher-student distribution matching. Similar trends can be observed on FeTS2024: removing MS-TKD causes the Dice of WT / Core / Enhance to decrease by approximately 1.6%, 3.4%, and 3.4% respectively, removing GSM leads to a decline of 2.5%, 2.3%, and 3.6%, and there is a decrease of 1.1%, 3.1%, and 3.3% after removing SLKD. These results show that MS-TKD, GSM, and SLKD each play their own roles and complement each other in multi-scale feature distillation, style and texture compensation, and flexible alignment at the output end, jointly improving the segmentation accuracy and stability of the model under different combinations of missing modalities.

Claims

1. A method for glioma missing modality segmentation based on the MST-KDNet neural network, characterized in that It includes the following steps: Step a: Construct an MST-KDNet neural network model composed of a multi-scale Transformer knowledge distillation module MS-TKD, a global style matching module GSME, and a dual-mode Logit distillation module DMLD for glioma missing modality segmentation, and output a segmentation result map; Step b: Establish a magnetic resonance imaging dataset of gliomas and divide it into a training set and a test set; Step c: Use the training set to train the MST-KDNet neural network model, and optimize the parameters of the MST-KDNet neural network model in combination with the characteristics of multi-modal glioma images; Step d: Use the test set to test and evaluate the obtained MST-KDNet network model, and finally realize the automatic and intelligent brain glioma missing modality segmentation function.

2. The glioma missing modality segmentation method based on the MST-KDNet neural network according to claim 1, characterized in that The multi-scale Transformer knowledge distillation module MS-TKD is composed of a 3D convolutional encoder, a multi-scale Transformer unit, and its extreme value distillation subunit EVD; The global style matching module GSME performs max pooling and splicing on the outputs of multiple convolutional layers, and then extracts features through a Transformer block to calculate the adversarial loss and the MSE loss; The dual-mode Logit distillation module DMLD is composed of a Logit divergence distillation unit and a Logit normalized KL loss distillation unit.

3. The glioma missing modality segmentation method based on the MST-KDNet neural network according to claim 2, wherein The specific implementation process of the multi-scale Transformer knowledge distillation module MS-TKD is as follows: First, the processed complete modality data and missing modality data are respectively passed through a 3D convolution layer to obtain corresponding initial features, which are input into the teacher model and the student model, and downsampling processing is completed through two 3D convolutional encoders of the corresponding models; Then, the two downsampled feature maps are respectively converted into one-dimensional sequences: the input is divided into flattened uniform non-overlapping patches; secondly, the patches are projected into an embedding space of K dimensions through a linear layer and a learnable position encoding is added to obtain two feature sequences; then, the feature sequences in the teacher model and the student model pass through a series of Transformer blocks in the multi-scale Transformer unit, and the MSA sublayer in the Transformer block is composed of n parallel self-attention SA blocks; Subsequently, the attention weights A of each self-attention block in each layer are taken out by the extreme value distillation sub-unit EVD of the Transformer unit for extreme value distillation as the sequence As, and the maximum value A of the attention weights at each pixel point is obtained max , the minimum value A min , and the mean value A mean : A max = max(As, axis = 1) A min = min(As, axis=1) Next, the weights of the teacher model and the student model are multiplied by their corresponding pixel points through the Transformer block respectively to obtain three sequences output by the teacher model and three sequences output by the student model and calculate the mean squared error (MSE) loss L for the corresponding sequences of the two models EVD ; Extract the sequence representation of the output of every three Transformer blocks in the multi-scale Transformer unit, and calculate the loss L using MSE for the sequence representations of the teacher model and the student model respectively z ; Then, the feature maps of the output of the last Transformer block are respectively subjected to skip connection operations with the feature maps of the output of the previous Transformer block after two consecutive 3D convolutions with normalization operations, and the output is upsampled using a transposed convolution layer to restore it to the original shape before being input into the Transformer block; the final output is input into a 1×1×1 convolution layer to obtain the feature f t , and finally, the total loss of this module is obtained by multiplying these two MSE losses by the weights α and β: L MS-TKD = αL EVD + βL z .

4. The glioma missing modality segmentation method based on the MST-KDNet neural network according to claim 3, wherein, The specific implementation process of the global style matching module GSME is as follows: Max-pool the features output from the penultimate convolutional layer of the 3D convolutional encoder in the multi-scale Transformer knowledge distillation module, and then concatenate the result with the output of the 3D convolutional encoder to obtain f enc , and then use the output f t after passing through the multi-scale Transformer knowledge distillation module enc and concatenate it with f enc&t and input the result into a 3D convolutional layer to obtain f adv ; Subsequently, f enc&t is input into the 3D convolutional decoder, and the features obtained from the first and second convolutions of the 3D convolutional decoder are also concatenated to obtain f dec , and f enc , f t , f dec is flattened in dimension to obtain the two-dimensional tensor G enc , G t , G dec Then, a feature fusion operation is performed: Fuse the feature sequence M enc&dec , M enc&t , M dec&t Perform MSE loss calculation L enc&dec&t : The total loss of the global style matching module GSME is L GSM : L GSM = εL adv + θL enc&dec&t where ε and θ are weight coefficients; Input f enc&t Input the feature map obtained by the 3D convolutional decoder into the dual-mode Logit distillation module.

5. The glioma missing modality segmentation method based on the MST-KDNet neural network according to claim 4, characterized in that, The implementation process of the dual-mode Logit distillation module DMLD is as follows: Adjust the number of channels of the input feature map through a 1×1×1 convolution to obtain feature map I; then normalize the feature maps l of the student model and the teacher model through logit, that is, input l into the Z-score normalization function, and then pass through the softmax function to obtain features q(l f ) and q(l m ): q(l f ) = softmax[Z(l f ; τ)], q(l m ) = softmax[Z(l m ; τ)], where μ is the mean of l, σ is the variance of l, τ is the temperature coefficient, and then the mean square error L is calculated for the output respectively mse and the KL loss L KL ; Then, a sigmoid activation function is used to generate a segmentation result, and the total loss of the dual-mode Logit distillation module is calculated through these two losses. The formula is as follows: L logit = λ mse L mse + λ KD τ 2 L KL Among them, λ mse and λ KD are corresponding weight coefficients.

6. The glioma missing modality segmentation method based on the MST-KDNet neural network according to claim 5, characterized in that The total loss of the MST-KDNet neural network model is specifically as follows: Calculate the Dice loss L between the segmentation results obtained from the sigmoid activation functions of the student model and the teacher model and the ground truth segmentation map dice , and finally calculate the total loss L joint Perform knowledge distillation: L joint = λ1L MS-TKD + λ2L logit + λ3L GSM + L dice where λ1, λ2, and λ3 are corresponding weight coefficients.