Multimodal MRI brain tumor segmentation method based on teacher and student learning

Through the multimodal MRI brain tumor segmentation method based on teacher-student learning, the teacher mode guides the feature representation of students' modality, and through the modal enhancement and fusion module and the deep supervision module, the multimodal MRI data integration problem is solved, and the segmentation accuracy and robustness are improved.

CN120510368APending Publication Date: 2025-08-19HANGZHOU NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510473988.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate information from different modes in multimodal MRI brain tumor segmentation, resulting in insufficient segmentation accuracy and robustness, especially poor performance when tumor shape and position changes.

Method used

Using a method based on teacher-student learning, the MRI modality is divided into teacher modality and student modality. The strong feature representation of the teacher modality is used to guide the student modality, and the feature representation is improved through the modal enhancement module and the modal fusion module, and a deep supervision module is introduced for multi-scale feature fusion to build a six-level encoder-decoder architecture to capture the intrinsic relationship of multi-modal data.

Benefits of technology

Maintain high segmentation accuracy when the shape and position change greatly, reduce dependence on a large amount of labeled data, and improve the training efficiency and performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510368A_ABST
    Figure CN120510368A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-modal MRI (Magnetic Resonance Imaging) brain tumor segmentation method based on teacher and student learning, which comprises the following steps: firstly, randomly selecting 80% as a training set and 20% as a test set from a brain tumor segmentation data set, the data comprising four modalities of Flair, T2, T1 and T1c, the Flair and T1c being teacher modalities, and the T2 and T1 being student modalities; then a segmentation network of a six-level encoder-decoder architecture is constructed, four encoders extract four modal features respectively, a teacher modal encoder comprises a standard convolution block, a cavity convolution block and a modal enhancement module, and a student modal encoder comprises a modal fusion module; the decoder comprises a cross-modal fusion module, a standard convolution block, a cavity convolution block, an up-sampling block and a depth supervision module; and finally, training the segmentation network by using the training set, inputting the test set into the trained network, and outputting a segmentation result by a decoder. The method provided by the invention can enhance the adaptability to tumor morphology and position changes, and the segmentation precision is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of brain tumor segmentation and relates to a multimodal MRI brain tumor segmentation method based on teacher-student learning. Background Art

[0002] Accurate brain tumor segmentation plays a crucial role in clinical diagnosis, treatment planning, and efficacy monitoring. To fully understand the morphology and development of brain tumors, the integration of multiple imaging modalities, especially different magnetic resonance imaging (MRI) sequences, has become key in research and practice. However, noise, variability, and image heterogeneity across different MRI modalities pose significant challenges to accurate brain tumor segmentation, further highlighting the need for efficient and accurate segmentation methods.

[0003] Brain tumors and other central nervous system (CNS) tumors pose a serious health threat worldwide and represent a significant challenge within the cancer field. Statistics show that brain tumors are the fifth most common type of cancer worldwide. According to the U.S. Central Brain Tumor Registry, approximately 90,000 people are diagnosed with primary brain tumors each year, nearly one-third of which are malignant. However, the prognosis for adult brain tumor patients is poor, with a five-year survival rate of only 12%. These data reveal the complexity of brain tumors and their profound impact on patient survival, highlighting the urgency of precise diagnosis and treatment of brain tumors.

[0004] Accurate segmentation not only plays a crucial role in surgical navigation and treatment planning, but also helps physicians effectively assess treatment response. However, relying on manual segmentation is both time-consuming and labor-intensive, and prone to inconsistent results due to inter-observer variability. Therefore, developing automated brain tumor segmentation methods that can accurately delineate tumor boundaries and improve diagnostic efficiency is crucial.

[0005] With the rapid development of deep learning, AI-based automated brain tumor segmentation technologies have received widespread attention in recent years. These technologies can not only improve the accuracy of segmentation, but also address the subjectivity of traditional manual segmentation methods. MRI is an indispensable tool in the diagnosis and monitoring of brain tumors, and its different sequences can reveal the size, location, and morphological characteristics of the tumor. By integrating multiple sequences such as fluid-attenuated inversion recovery (FLAIR), T1-weighted, T2-weighted, and T1 contrast-enhanced (T1c), researchers have effectively improved the robustness and accuracy of segmentation by leveraging the complementarity between these modalities. Multimodal image fusion technology has shown great potential and provides important support for accurately characterizing the complex nature of brain tumors and improving the quality of treatment decisions.

[0006] Promoting research into automated brain tumor segmentation methods will not only reduce the burden on medical personnel but also significantly improve the efficiency and consistency of clinical decision-making. With the continuous advancement of technology and optimization of algorithms, multimodal fusion methods are expected to further improve the accuracy and stability of segmentation in the future, providing more personalized treatment options for brain tumor patients and contributing to the overall advancement of brain tumor research and treatment.

[0007] Existing methods for brain tumor segmentation can be roughly divided into three categories. The first category is based on traditional image processing techniques such as thresholding, region growing, and mathematical morphology. These methods segment using simple image features or rules. While they perform well in some scenarios, they are limited in effectiveness when faced with complex medical images.

[0008] The second category of methods employs traditional machine learning models, such as support vector machines (SVMs) and random forests. These models rely on handcrafted features (such as texture, intensity, and shape) extracted from medical images to train classifiers. While these methods are effective under certain conditions, their performance relies heavily on handcrafted features, limiting their applicability to more complex tasks.

[0009] The third category of methods is deep learning-based models, which have made significant progress in brain tumor segmentation tasks in recent years. U-Net and its variants excel in segmentation accuracy and effectively capture both local and global information. Furthermore, transformer-based methods improve model performance by capturing long-term dependencies through attention mechanisms; GANs are used to generate synthetic data, enhancing the diversity of training data and the generalization ability of the model; and graph convolutional networks (GCNs) enhance contextual awareness by modeling dependencies between spaces and channels. These methods not only improve segmentation accuracy but also provide powerful technical support for the automated detection and monitoring of brain tumors.

[0010] While existing methods have made some progress in the field of brain tumor segmentation, many shortcomings remain. Specifically, traditional image processing techniques lack sufficient expressive power to accurately capture tumor boundary details. Segmentation accuracy is often low, especially when tumor shapes are irregular or undergo significant changes. While traditional machine learning models can handle certain specific tasks, their reliance on handcrafted features significantly limits their performance when dealing with variations in tumor morphology and position. While deep learning models have significantly improved segmentation accuracy, they still face challenges when fusing multimodal MRI data. Because data from different modalities typically have different spatial distributions and information characteristics, efficiently integrating this information remains an urgent problem. Specifically, in the paper (Sherman R. Avolumetric convolutional neural network for brain tumor segmentation [J]. arXiv preprint arXiv:1811.02654, 2018), this study employed a multi-stage symmetric encoder and decoder based on a U-Net for semantic segmentation of brain tumors. While achieving some success, this study failed to fully consider the differences in characteristics between modalities and failed to propose a targeted module to effectively integrate multimodal information.

[0011] Therefore, it is of great significance to study a multimodal MRI brain tumor segmentation method based on teacher-student learning to solve the problems existing in the existing technology. Summary of the Invention

[0012] The purpose of the present invention is to solve the problems existing in the prior art and provide a multimodal MRI brain tumor segmentation method based on teacher-student learning.

[0013] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0014] A multimodal MRI brain tumor segmentation method based on teacher-student learning includes the following steps:

[0015] (1) 80% of the collected brain tumor segmentation datasets were randomly selected as the training set, and the remaining 20% were used as the test set; each brain tumor segmentation dataset contains four modalities: Flair, T2, T1, and T1c. Flair and T1c are defined as teacher modalities, and T2 and T1 are defined as the corresponding student modalities;

[0016] By dividing the four MR modalities into a teacher modality and a student modality, the strong feature representation of the teacher modality is used to guide and enhance the feature representation of the student modality, achieving knowledge transfer from the teacher modality to the student modality. The teacher modalities (Flair and T1c) have stronger capabilities in identifying tumor boundaries and distinguishing different tumor regions, providing more accurate initial feature representations. The student modalities (T2 and T1), guided by the teacher modalities, are more robust to complex morphology and noise, thereby maintaining high segmentation accuracy even in the presence of large variations in shape and position.

[0017] (2) Construct a segmentation network;

[0018] The segmentation network adopts a six-stage encoder-decoder architecture, including 4 encoders and 1 decoder;

[0019] The four encoders are used to extract features of four different modalities;

[0020] Corresponding to each encoder of the teacher modality, the first level includes a standard convolution block, a dilated convolution block, and a modality enhancement module (MEM), and the second to sixth levels include a standard convolution block, a dilated convolution block, and a modality enhancement module, respectively;

[0021] The modality enhancement module is used to refine the feature representation of the teacher modality through a series of operations (such as global maximum pooling, convolution and residual connection) to make it richer and more accurate;

[0022] Corresponding to each encoder of the student modality, the first level includes a standard convolution block, a dilated convolution block, and a modality fusion module (MFM), and the second to sixth levels include a standard convolution block, a dilated convolution block, and a modality fusion module, respectively;

[0023] The modality fusion module is used to guide and improve the feature representation of the student modality by utilizing the optimized teacher modality features generated by the modality enhancement module;

[0024] The modality enhancement module (MEM) and the modality fusion module (MFM) together constitute the modality guidance module (MGM); through the synergistic effect of the two sub-modules MEM and MFM, MGM can effectively capture the intrinsic relationship of multimodal data and improve the quality of feature representation; this multi-level feature enhancement and fusion mechanism enables the model to better capture the complex morphology and positional changes of tumors, improving the accuracy and robustness of segmentation.

[0025] The modality fusion module in each encoder corresponding to the student modality also obtains a feature map from the modality enhancement module in each encoder corresponding to the teacher modality, and the modality fusion modules in the first to fifth levels of each encoder corresponding to the student modality also output the obtained feature map to the decoder;

[0026] From the first to the sixth level of each encoder, features are extracted by gradually reducing the size of the feature map and increasing the number of channels;

[0027] The decoder includes a cross-modal fusion module, a standard convolution block, a hole convolution block, an upsampling block, and a deep supervision module. The results of the four encoders are first spliced and passed through a cross-modal fusion module, and then spliced with the output results of the modal fusion module in the fifth level of the encoder through the upsampling block. They then pass through a cross-modal fusion module, a convolution block, and a hole convolution block in sequence, and then the output results of the modal fusion modules of the fourth, third, second, and first levels of the encoder are cycled for four rounds. Two copies of the results obtained in each round are sent out, one is upsampled as the input for the next round, and the other is input to the deep supervision module at this level. Finally, the results input to the deep supervision modules at each level are element-wise added to obtain the output result, i.e., the final segmentation result.

[0028] The Cross-Modal Fusion Module (CMFM) is used to further explore inter-modal relationships and learn information-rich feature representations. It captures complementary information between modalities through bidirectional feature fusion, generates cross-modal weights, and dynamically adjusts the importance of modalities, enabling the model (i.e., the segmentation network) to better utilize useful information and suppress noise and irrelevant information. Specifically, CMFM performs feature fusion at each level of the decoder path, ensuring that features from different modalities are effectively combined at all scales. This dynamic adjustment mechanism makes the model more stable when dealing with complex morphology and noise, and can effectively cope with changes in tumor shape and position.

[0029] The deep supervision module is a convolutional block with a kernel size of 1×1×1. Introduced in the network's decoder path, the deep supervision module adds a 1×1×1 convolutional block to the segmentation results at each layer and generates the final segmentation output through element-wise summation. Deep supervision not only enhances gradient flow and prevents vanishing and exploding gradients, but also promotes the fusion of multi-scale features, improving the model's generalization and segmentation accuracy. This multi-scale feature fusion mechanism enables the model to better capture the local and global characteristics of tumors, improving segmentation accuracy even with large variations in shape and position, and providing a new solution for multimodal medical image analysis.

[0030] (3) Use the training set to train the constructed segmentation network to obtain a trained segmentation network;

[0031] (4) The test set is input into the encoder of the trained segmentation network, and the decoder outputs the tumor segmentation results. The output tumor segmentation results include three types: the whole tumor, the tumor core, and the enhanced tumor.

[0032] As the preferred technical solution:

[0033] The multimodal MRI brain tumor segmentation method based on teacher-student learning as described above, corresponding to each encoder of the teacher modality, includes, in sequence, a first standard convolution block, a first dilated convolution block, a first modality enhancement module, a second standard convolution block, a second dilated convolution block, a second modality enhancement module, a third standard convolution block, a third dilated convolution block, a third modality enhancement module, a fourth standard convolution block, a fourth dilated convolution block, a fourth modality enhancement module, a fifth standard convolution block, a fifth dilated convolution block, a fifth modality enhancement module, a sixth standard convolution block, a sixth dilated convolution block and a sixth modality enhancement module, and each modality enhancement module is connected to an adjacent convolution block;

[0034] Each encoder corresponding to the student modality includes, in sequence, a seventh standard convolution block, a seventh hole convolution block, a first modality fusion module, an eighth standard convolution block, an eighth hole convolution block, a second modality fusion module, a ninth standard convolution block, a ninth hole convolution block, a third modality fusion module, a tenth standard convolution block, a tenth hole convolution block, a fourth modality fusion module, an eleventh standard convolution block, an eleventh hole convolution block, a fifth modality fusion module, a twelfth standard convolution block, a twelfth hole convolution block and a sixth modality fusion module, and each modality fusion module is connected to the adjacent convolution blocks;

[0035] For each group of corresponding teacher modality and student modality, the first modality fusion module also obtains feature maps from the first modality enhancement module, the second modality fusion module also obtains feature maps from the second modality enhancement module, the third modality fusion module also obtains feature maps from the third modality enhancement module, the fourth modality fusion module also obtains feature maps from the fourth modality enhancement module, the fifth modality fusion module also obtains feature maps from the fifth modality enhancement module, and the sixth modality fusion module also obtains feature maps from the sixth modality enhancement module.

[0036] As described above, a multimodal MRI brain tumor segmentation method based on teacher-student learning, the decoder includes a cross-modal fusion module A, a cross-modal fusion module B, a standard convolution block A, a hole convolution block A, an upsampling block A, a deep supervision module A, a cross-modal fusion module C, a standard convolution block B, a hole convolution block B, an upsampling block B, a deep supervision module B, a cross-modal fusion module D, a standard convolution block C, a hole convolution block C, an upsampling block C, a deep supervision module C, a cross-modal fusion module E, a standard convolution block D, a hole convolution block D, an upsampling block D, a deep supervision module D, a cross-modal fusion module F, a standard convolution block E, a hole convolution block E, an upsampling block E and a deep supervision module E;

[0037] The results obtained by the four encoders are first spliced and passed through the cross-modal fusion module A, and then spliced with the output results of the fifth modal fusion module through the upsampling block A, and then passed through the cross-modal fusion module B, standard convolution block A and void convolution block A in sequence. The output results are sent out in two copies, one of which is input into the deep supervision module A, and the other is spliced with the output results of the fourth modal fusion module through the upsampling block B, and then passed through the cross-modal fusion module C, standard convolution block B and void convolution block B in sequence. At this time, the output results are sent out in two copies, one of which is input into the deep supervision module B, and the other is spliced with the output results of the third modal fusion module through the upsampling block C, and then passed through the cross-modal fusion module D, standard convolution block C and void convolution block C in sequence. The output result obtained at this time is sent out in two copies, one is input into the depth supervision module C, and the other is spliced with the output result of the second modal fusion module through the upsampling block D, and then passes through the cross-modal fusion module E, the standard convolution block D and the hole convolution block D in sequence. The output result obtained at this time is sent out in two copies, one is input into the depth supervision module D, and the other is spliced with the output result of the first modal fusion module through the upsampling block E, and then passes through the cross-modal fusion module F, the standard convolution block E and the hole convolution block E in sequence. The output result obtained at this time is input into the depth supervision module E, and finally input into the depth supervision module A, depth supervision module B, depth supervision module C, depth supervision module D and the depth supervision module E. The output result is obtained by element-wise addition.

[0038] As described above, a multimodal MRI brain tumor segmentation method based on teacher-student learning, the feature map T i After the maximum pooling, multi-layer perceptron and Sigmoid activation, the unprocessed feature map T i The corresponding position elements are multiplied, and the result of the multiplication is then multiplied with the unprocessed feature map T i Add the corresponding position elements to get the feature map T i '.

[0039] As described above, a multimodal MRI brain tumor segmentation method based on teacher-student learning, the modality fusion module first transforms the input feature map S j The feature map T output by the modality enhancement module i 'Splice, then perform 3×3×3 convolution to get the feature map S ij ', feature map S ij 'After the maximum pooling, multi-layer perceptron and Sigmoid activation, it is compared with the unprocessed feature map S ij 'Multiply the corresponding position elements, and then multiply the result with the unprocessed feature map S ij 'Add the corresponding position elements to get the feature map S j '.

[0040] As described above, in a multimodal MRI brain tumor segmentation method based on teacher-student learning, the cross-modal fusion module receives the output of the modality fusion module of the four modal corresponding levels in the encoder and uses it as input, and finally outputs a fused tensor; denoted by f F is the Flair modal feature map, f T2 is the T2 modal feature map, f T1c is the T1c modal feature map, f T1 is the T1 modal feature map, f up is the feature map obtained by upsampling the previous layer in the decoder; first, f F Perform two 3×3×3 convolutions to obtain ψ(f F ) and ξ(f F ), f T2 Perform another 3×3×3 convolution to get φ(f T2 ),ξ(f F ) and φ(f T2 ) The corresponding position elements are multiplied point by point and activated by Softmax to obtain W FT2 , W FT2 and ψ(f F ) The corresponding position elements are multiplied point by point and then added to f F and f T2 Splicing to get f FT2 ; Then, swap f F With f T2 The position of f is obtained by following the same steps. T2F ; Then, f T1c Perform two 3×3×3 convolutions to obtain ψ(f T1c ) and ξ(f T1c ), f T1 Perform another 3×3×3 convolution to get φ(f T1 ),ξ(f T1c ) and φ(f T1 ) The corresponding position elements are multiplied point by point and activated by Softmax to obtain W T1cT1 , W T1cT1 and ψ(f T1c ) The corresponding position elements are multiplied point by point and then added to f T1c and f T1 Splicing to get f T1cT1 ; Then, swap f T1c With f T1 The position of f is obtained by following the same steps. T1T1c ; Finally, f FT2 、f T2F 、f T1cT1 、f T1T1c and f up(The bottom layer only has the feature maps of 4 modes as the input of CMFM, and the other 5 layers use the feature maps of 4 modes and the feature maps obtained by upsampling the previous layer in the decoder as the input of CMFM) spliced in the channel dimension to obtain the final module output f out .

[0041] As described above, a multimodal MRI brain tumor segmentation method based on teacher-student learning uses the standard Dice loss function for network training, which is formulated as follows:

[0042]

[0043] Where N is the total number of pixels in the image, C is the number of segmentation categories, and p ij ∈[0,1] and g ij ∈[0,1] represents the predicted value and true value of voxel i belonging to category j, and ∈ represents a constant included to prevent division by zero, with a value of 1×10 -5 .

[0044] Beneficial effects:

[0045] The present invention provides a multimodal MRI brain tumor segmentation method based on teacher-student learning. By dividing the four MR modalities into a teacher modality and a student modality, the strong feature representation of the teacher modality is used to guide and enhance the feature representation of the student modality, thereby realizing knowledge transfer from the teacher modality to the student modality. By introducing a modality enhancement module, a modality fusion module, and a deep supervision module, the present invention can maintain high segmentation accuracy even when the shape and position vary greatly. In addition, it reduces the dependence on a large amount of labeled data, thereby improving the training efficiency and performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is the overall framework of the segmentation network;

[0047] Figure 2 This is the architecture diagram of the modal enhancement module and modal fusion module;

[0048] Figure 3 is the cross-modal fusion module architecture diagram; where f up : Figure 1 The red arrows in the figure refer to the feature maps from the corresponding layer encoder above; ψ is the convolution operation 1; ξ is the convolution operation 2; φ is the convolution operation 3. DETAILED DESCRIPTION

[0049] Below in conjunction with specific embodiment, further set forth the present invention.Should be understood that these embodiments are only used to illustrate the present invention and are not used in limiting the scope of the present invention.In addition, should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall equally within the scope limited by the appended claims of the application.

[0050] A multimodal MRI brain tumor segmentation method based on teacher-student learning includes the following steps:

[0051] (1) 80% of the collected brain tumor segmentation datasets were randomly selected as the training set, and the remaining 20% were used as the test set; each brain tumor segmentation dataset contains four modalities: Flair, T2, T1, and T1c. Flair and T1c are defined as teacher modalities, and T2 and T1 are defined as the corresponding student modalities;

[0052] (2) Construct a segmentation network;

[0053] like Figure 1 As shown, the segmentation network adopts a six-stage encoder-decoder architecture, including 4 encoders and 1 decoder;

[0054] Figure 1 The white in the middle is a standard convolution block (first 3×3×3 convolution, then instance normalization, and finally leakyReLU activation function); the light gray is a standard convolution block (first 3×3×3 convolution with a stride of 2, then instance normalization, and finally leakyReLU activation function); the dark gray is a dilated convolution block with residual connection (first 3×3×3 dilated convolution with a dilation rate of 2, instance normalization, leakyReLU activation function, then 3×3×3 dilated convolution with a dilation rate of 4, instance normalization, leakyReLU activation function, and finally element-wise addition with the unprocessed input feature map); orange is the modal enhancement module (MEM); yellow is the modal fusion module (MFM); green is the cross-modal fusion module (CMFM); purple is the deep supervision module; the red arrow is the concatenation of four modal feature maps; the blue arrow is the upsampling block with a stride of 2;

[0055] The four encoders are used to extract features of four different modalities;

[0056] Corresponding to each encoder of the teacher mode, the first level includes a first standard convolution block, a first hole convolution block and a first modality enhancement module; the second to sixth levels respectively include a second standard convolution block, a second hole convolution block, a second modality enhancement module, a third standard convolution block, a third hole convolution block, a third modality enhancement module, a fourth standard convolution block, a fourth hole convolution block, a fourth modality enhancement module, a fifth standard convolution block, a fifth hole convolution block, a fifth modality enhancement module, a sixth standard convolution block, a sixth hole convolution block and a sixth modality enhancement module; each modality enhancement module is connected to the adjacent convolution blocks;

[0057] like Figure 2 As shown in (a), the feature map T of the modality enhancement module (MEM) input i After the maximum pooling, multi-layer perceptron and Sigmoid activation, the unprocessed feature map T i Multiply the corresponding position elements and then add them to the unprocessed feature map T i Add the corresponding position elements to get the feature map T i ';The modality enhancement module is used to refine the feature representation of the teacher modality;

[0058] Corresponding to each encoder of the student modality, the first level includes the seventh standard convolution block, the seventh void convolution block and the first modal fusion module; the second to sixth levels respectively include the eighth standard convolution block, the eighth void convolution block, the second modal fusion module, the ninth standard convolution block, the ninth void convolution block, the third modal fusion module, the tenth standard convolution block, the tenth void convolution block, the fourth modal fusion module, the eleventh standard convolution block, the eleventh void convolution block, the fifth modal fusion module, the twelfth standard convolution block, the twelfth void convolution block and the sixth modal fusion module; each modal fusion module is connected to the adjacent convolution block;

[0059] like Figure 2 As shown in (b), the modal fusion module (MFM) first transforms the input feature map S j The feature map T output by the modality enhancement module i 'Splice, then perform 3×3×3 standard convolution to obtain feature map S ij ', feature map S ij 'After the maximum pooling, multi-layer perceptron and Sigmoid activation, it is compared with the unprocessed feature map S ij 'Multiply the corresponding position elements and then add them to the unprocessed feature map S ij 'Add the corresponding position elements to get the feature map S j ';The modality fusion module is used to guide and improve the feature representation of the student modality by using the optimized teacher modality features generated by the modality enhancement module;

[0060] The modal enhancement module (MEM) and the modal fusion module (MFM) together constitute the modal guidance module (MGM);

[0061] For each group of corresponding teacher modalities and student modalities, the first modal fusion module further obtains a feature map from the first modal enhancement module, the second modal fusion module further obtains a feature map from the second modal enhancement module, the third modal fusion module further obtains a feature map from the third modal enhancement module, the fourth modal fusion module further obtains a feature map from the fourth modal enhancement module, the fifth modal fusion module further obtains a feature map from the fifth modal enhancement module, and the sixth modal fusion module further obtains a feature map from the sixth modal enhancement module;

[0062] From the first to the sixth level of each encoder, features are extracted by gradually reducing the size of the feature map and increasing the number of channels;

[0063] The decoder includes cross-modal fusion module A, cross-modal fusion module B, standard convolution block A, void convolution block A, upsampling block A, deep supervision module A, cross-modal fusion module C, standard convolution block B, void convolution block B, upsampling block B, deep supervision module B, cross-modal fusion module D, standard convolution block C, void convolution block C, upsampling block C, deep supervision module C, cross-modal fusion module E, standard convolution block D, void convolution block D, upsampling block D, deep supervision module D, cross-modal fusion module F, standard convolution block E, void convolution block E, upsampling block E and deep supervision module E;

[0064] The results obtained by the four encoders are first spliced and passed through the cross-modal fusion module A, and then spliced with the output results of the fifth modal fusion module through the upsampling block A, and then passed through the cross-modal fusion module B, standard convolution block A and void convolution block A in sequence. The output results are sent out in two copies, one of which is input into the deep supervision module A, and the other is spliced with the output results of the fourth modal fusion module through the upsampling block B, and then passed through the cross-modal fusion module C, standard convolution block B and void convolution block B in sequence. At this time, the output results are sent out in two copies, one of which is input into the deep supervision module B, and the other is spliced with the output results of the third modal fusion module through the upsampling block C, and then passed through the cross-modal fusion module D, standard convolution block C and void convolution block C in sequence. The output result obtained at this time is sent out in two copies, one is input into the depth supervision module C, and the other is spliced with the output result of the second modal fusion module through the upsampling block D, and then passes through the cross-modal fusion module E, the standard convolution block D and the hole convolution block D in sequence. The output result obtained at this time is sent out in two copies, one is input into the depth supervision module D, and the other is spliced with the output result of the first modal fusion module through the upsampling block E, and then passes through the cross-modal fusion module F, the standard convolution block E and the hole convolution block E in sequence. The output result obtained at this time is input into the depth supervision module E, and finally input into the depth supervision module A, depth supervision module B, depth supervision module C, depth supervision module D and the depth supervision module E. The output result is obtained by element-wise addition;

[0065] The Cross-Modal Fusion Module (CMFM) captures the complementary information between modalities through bidirectional feature fusion, generates cross-modal weights, and dynamically adjusts the importance of modalities;

[0066] like Figure 3 As shown, the cross-modal fusion module receives the output of the modality fusion module of the four modal corresponding levels in the encoder and takes it as input, and finally outputs a fused tensor; denoted by f F is the Flair modal feature map, f T2 is the T2 modal feature map, f T1c is the T1c modal feature map, f T1 is the T1 modal feature map, f up is the feature map obtained by upsampling the previous layer in the decoder; first, f F Perform two 3×3×3 convolutions to obtain ψ(f F ) and ξ(f F ), f T2 Perform another 3×3×3 convolution to get φ(f T2 ),ξ(f F ) and φ(f T2 ) The corresponding position elements are multiplied point by point and activated by Softmax to obtain W FT2 , W FT2 and ψ(fF ) The corresponding position elements are multiplied point by point and then added to f F and f T2 Splicing to get f FT2 ; Then, swap f F With f T2 The position of f is obtained by following the same steps. T2F ; Then, f T1c Perform two 3×3×3 convolutions to obtain ψ(f T1c ) and ξ(f T1c ), f T1 Perform another 3×3×3 convolution to get φ(f T1 ),ξ(f T1c ) and φ(f T1 ) The corresponding position elements are multiplied point by point and activated by Softmax to obtain W T1cT1 , W T1cT1 and ψ(f T1c ) The corresponding position elements are multiplied point by point and then added to f T1c and f T1 Splicing to get f T1cT1 ; Then, swap f T1c With f T1 The position of f is obtained by following the same steps. T1T1c ; Finally, f FT2 、f T2F 、f T1cT1 、f T1T1c and f up Splicing in the channel dimension to get the final module output f out ; Among them, C, L, W, and H represent the number of channels, length, width, and height respectively; if the input is 4 channels, the output is 8 channels, if the input is 5 channels, the output is 9 channels; if upsampling is performed, one more channel is added;

[0067] The deep supervision module is a convolution block with a convolution kernel size of 1×1×1;

[0068] (3) Use the training set to train the constructed segmentation network to obtain a trained segmentation network;

[0069] The standard Dice loss function is used for network training, and the formula is as follows:

[0070]

[0071] Where N is the total number of pixels in the image, C is the number of segmentation categories, and p ij ∈[0,1] and g ij∈[0,1] represents the predicted value and true value of voxel i belonging to category j, and ∈ represents a constant included to prevent division by zero, with a value of 1×10 -5 ;

[0072] (4) The test set is input into the encoder of the trained segmentation network, and the decoder outputs the tumor segmentation results. The output tumor segmentation results include three types: the whole tumor, the tumor core, and the enhanced tumor.

[0073] MRI offers multiple modalities, each of which provides unique insights into different aspects of brain tissue. The four main MR modalities used for brain tumor segmentation are Flair, T2, T1, and T1c. On the one hand, Flair and T2 play a key role in delineating tumor boundaries and identifying surrounding edema areas; on the other hand, T1c and T1 are crucial in distinguishing between enhancing tumors, necrosis, and non-enhancing tumor areas, helping to identify high-grade tumor components. In particular, Flair is a T2-weighted image that can effectively suppress the signal of cerebrospinal fluid (CSF) and is more sensitive to edema areas than T2. In addition, T1c (enhanced T1-weighted image) is superior to T1 (non-enhanced T1-weighted image) in visualizing enhancing tumors, necrosis, and non-enhancing tumor areas.

[0074] Based on this background, we can view these modalities as a feature hierarchy and propose to use a teacher-student learning approach to learn multimodal features hierarchically. Teacher-student learning is commonly used to address data scarcity and generalization issues. In this paper, we propose for the first time its application to multimodal feature distillation (i.e., extracting information from multimodal data). Specifically, we divide the four MR modalities into teacher modalities (Flair and T1c) and student modalities (T2 and T1). To implement teacher-student learning, we propose a novel modal guidance module (MGM) to learn the unique information provided by each modality. This innovative module aims to leverage the strengths of the teacher modalities (T1c and Flair) to enhance the feature representation of the student modalities (T1 and T2). In addition, we propose a cross-modal fusion module (CMFM) to further integrate multimodal information and improve the overall accuracy of brain tumor segmentation. Starting from analyzing the unique characteristics and synergistic effects of different MR modalities, this paper explores their intrinsic relationship in brain tumor segmentation.

[0075] To evaluate the proposed method, the BraTS (Brain Tumor Segmentation) 2018 dataset was used. The BraTS 2018 dataset contains 285 training samples. Each sample includes four MRI modalities: Flair, T2, T1, and T1c, and is accompanied by expert annotations. These annotations classify brain tumors into three categories: whole tumor, tumor core, and enhancing tumor. It should be noted that the whole tumor includes edema, enhancing tumor, non-enhancing tumor, and necrotic areas; the tumor core includes enhancing tumor, non-enhancing tumor, and necrotic areas.

[0076] During data preprocessing, the images were cropped and resized to 128×128×128, and normalized using the N4ITK method (reference: Advanced normalization tools(ants), Insight j 2(2009)1–35.) and intensity normalization technology.

[0077] To implement the brain tumor segmentation network, we selected Keras as the framework and used an NVIDIA GeForce RTX4090 graphics card (24GB) for model training. The initial learning rate was set to 0.0005, and if the validation loss did not improve for 10 consecutive epochs, the learning rate was halved. To prevent overfitting, an early stopping mechanism was introduced: if the validation loss did not improve for 30 epochs, training was terminated. During model training, the Adam optimizer was used to update parameters. The detailed experimental settings are shown in Table 1.

[0078] Table 1

[0079]

[0080] To evaluate the impact of each module on brain tumor segmentation performance, comprehensive ablation experiments were conducted; the experimental results on the BraTS2018 dataset are shown in Table 2. First, a baseline method was established without the proposed module. Subsequently, deep supervision (DS), modality guidance module (MGM), and cross-modal fusion module (CMFM) were incorporated into the baseline method to evaluate their specific contributions. The results show that the integration of these modules significantly improves the segmentation performance. Integrating the DS-guided segmentation decoder improves the segmentation accuracy by 0.8% in terms of average Dice Similarity Coefficient (DSC) and 21.1% in terms of average 95% Hausdorff Distance (HD) compared to the baseline method.

[0081] Table 2

[0082]

[0083] The "Baseline Method" in Table 2 above refers to the network structure without the proposed DS, MGM, and CMFM strategies. "√" in Table 2 indicates the strategy was added, and "–" indicates the strategy was not included. The best experimental results are marked in bold in Table 2.

[0084] At Baseline, without the addition of DS, MGM, and CMFM, the average DSC was 82.9% and the average HD was 5.7;

[0085] At Baseline, with the addition of DS, the average DSC was 83.6% and the average HD was 4.5;

[0086] On the baseline, with the addition of DS and MGM, the average DSC is 83.8% and the average HD is 4.1; compared with the baseline method, the average DSC and average HD are improved by 1.1% and 28.1% respectively.

[0087] This improvement stems from the exploitation of multimodal correlations, which enables the teacher modality to effectively guide the student modality, resulting in refined feature representations. Compared to the baseline method, the addition of CMFM further improves the performance, with an average improvement of 1.2% in DSC and 28.1% in HD.

[0088] This improvement is attributed to the effective combination of the four modalities, which achieves a more informative feature representation. Furthermore, the combination of these modules achieves state-of-the-art performance for brain tumor segmentation on two public datasets.

Claims

1. A multimodal MRI brain tumor segmentation method based on teacher-student learning, characterized by The steps include: (1) Randomly select 80% of the collected brain tumor segmentation datasets as the training set, and the remaining 20% as the test set; each brain tumor segmentation dataset contains four modalities: Flair, T2, T1, and T1c. Flair and T1c are defined as teacher modalities, and T2 and T1 are defined as the corresponding student modalities; (2) Construct a segmentation network; The segmentation network adopts a six-stage encoder-decoder architecture, including 4 encoders and 1 decoder; The four encoders are used to extract features of four different modalities; Each encoder corresponding to the teacher modality includes a modality enhancement module, which is used to refine the feature representation of the teacher modality; Each encoder corresponding to the student modality includes a modality fusion module, which is used to guide and improve the feature representation of the student modality by using the optimized teacher modality features generated by the modality enhancement module; The decoder includes a cross-modal fusion module, which captures the complementary information between modalities through bidirectional feature fusion, generates cross-modal weights, and dynamically adjusts the importance of modalities; (3) Use the training set to train the constructed segmentation network to obtain a trained segmentation network; (4) The test set is input into the encoder of the trained segmentation network, and the decoder outputs the tumor segmentation results. The output tumor segmentation results include three types: the whole tumor, the tumor core, and the enhanced tumor.

2. The multimodal MRI brain tumor segmentation method based on teacher-student learning according to claim 1, characterized in that: Corresponding to each encoder of the teacher modality, the first level includes a standard convolution block, a dilated convolution block, and a modality enhancement module, and the second to sixth levels include a standard convolution block, a dilated convolution block, and a modality enhancement module respectively; Corresponding to each encoder of the student modality, the first level includes a standard convolution block, a dilated convolution block, and a modality fusion module, and the second to sixth levels include a standard convolution block, a dilated convolution block, and a modality fusion module respectively; The modality fusion module in each encoder corresponding to the student modality also obtains a feature map from the modality enhancement module in each encoder corresponding to the teacher modality, and the modality fusion modules in the first to fifth levels of each encoder corresponding to the student modality also output the obtained feature map to the decoder; From the first to the sixth level of each encoder, features are extracted by gradually reducing the size of the feature map and increasing the number of channels; The decoder also includes a standard convolution block, a hole convolution block, an upsampling block and a deep supervision module; the results obtained by the four encoders are first spliced and passed through a cross-modal fusion module, and then spliced with the output results of the modal fusion module in the fifth level of the encoder through the upsampling block, and then pass through a cross-modal fusion module, a convolution block and a hole convolution block in sequence. After that, the output results of the modal fusion modules of the fourth, third, second and first levels of the encoder are cycled for four rounds respectively; the results obtained in each round are sent out in two copies, one is upsampled as the input of the next round, and the other is input to the deep supervision module of this level; the results input to the deep supervision modules at all levels are finally added element by element to obtain the output result, that is, the final segmentation result; The deep supervision module is a convolution block with a kernel size of 1×1×1.

3. The multimodal MRI brain tumor segmentation method based on teacher-student learning according to claim 2, characterized in that: Each encoder corresponding to the teacher mode includes, in sequence, a first standard convolution block, a first hole convolution block, a first modal enhancement module, a second standard convolution block, a second hole convolution block, a second modal enhancement module, a third standard convolution block, a third hole convolution block, a third modal enhancement module, a fourth standard convolution block, a fourth hole convolution block, a fourth modal enhancement module, a fifth standard convolution block, a fifth hole convolution block, a fifth modal enhancement module, a sixth standard convolution block, a sixth hole convolution block and a sixth modal enhancement module, and each modal enhancement module is connected to the adjacent convolution block; Each encoder corresponding to the student modality includes, in sequence, a seventh standard convolution block, a seventh hole convolution block, a first modality fusion module, an eighth standard convolution block, an eighth hole convolution block, a second modality fusion module, a ninth standard convolution block, a ninth hole convolution block, a third modality fusion module, a tenth standard convolution block, a tenth hole convolution block, a fourth modality fusion module, an eleventh standard convolution block, an eleventh hole convolution block, a fifth modality fusion module, a twelfth standard convolution block, a twelfth hole convolution block and a sixth modality fusion module, and each modality fusion module is connected to the adjacent convolution blocks; For each group of corresponding teacher modality and student modality, the first modality fusion module also obtains feature maps from the first modality enhancement module, the second modality fusion module also obtains feature maps from the second modality enhancement module, the third modality fusion module also obtains feature maps from the third modality enhancement module, the fourth modality fusion module also obtains feature maps from the fourth modality enhancement module, the fifth modality fusion module also obtains feature maps from the fifth modality enhancement module, and the sixth modality fusion module also obtains feature maps from the sixth modality enhancement module.

4. The multimodal MRI brain tumor segmentation method based on teacher-student learning according to claim 3, characterized in that: The decoder includes cross-modal fusion module A, cross-modal fusion module B, standard convolution block A, void convolution block A, upsampling block A, deep supervision module A, cross-modal fusion module C, standard convolution block B, void convolution block B, upsampling block B, deep supervision module B, cross-modal fusion module D, standard convolution block C, void convolution block C, upsampling block C, deep supervision module C, cross-modal fusion module E, standard convolution block D, void convolution block D, upsampling block D, deep supervision module D, cross-modal fusion module F, standard convolution block E, void convolution block E, upsampling block E and deep supervision module E; The results obtained by the four encoders are first spliced and passed through the cross-modal fusion module A, and then spliced with the output results of the fifth modal fusion module through the upsampling block A, and then passed through the cross-modal fusion module B, standard convolution block A and void convolution block A in sequence. The output results are sent out in two copies, one of which is input into the deep supervision module A, and the other is spliced with the output results of the fourth modal fusion module through the upsampling block B, and then passed through the cross-modal fusion module C, standard convolution block B and void convolution block B in sequence. At this time, the output results are sent out in two copies, one of which is input into the deep supervision module B, and the other is spliced with the output results of the third modal fusion module through the upsampling block C, and then passed through the cross-modal fusion module D, standard convolution block C and void convolution block C in sequence. The output result obtained at this time is sent out in two copies, one is input into the depth supervision module C, and the other is spliced with the output result of the second modal fusion module through the upsampling block D, and then passes through the cross-modal fusion module E, the standard convolution block D and the hole convolution block D in sequence. The output result obtained at this time is sent out in two copies, one is input into the depth supervision module D, and the other is spliced with the output result of the first modal fusion module through the upsampling block E, and then passes through the cross-modal fusion module F, the standard convolution block E and the hole convolution block E in sequence. The output result obtained at this time is input into the depth supervision module E, and finally input into the depth supervision module A, depth supervision module B, depth supervision module C, depth supervision module D and the depth supervision module E. The output result is obtained by element-wise addition.

5. The multimodal MRI brain tumor segmentation method based on teacher-student learning according to claim 4, characterized in that: Feature map T input to the modality enhancement module i After the maximum pooling, multi-layer perceptron and Sigmoid activation, the unprocessed feature map T i Multiply the corresponding position elements and then add them to the unprocessed feature map T i Add the corresponding position elements to get the feature map T i '.

6. The multimodal MRI brain tumor segmentation method based on teacher-student learning according to claim 5, characterized in that: The modality fusion module first transforms the input feature map S j The feature map T output by the modality enhancement module i 'Splice, then perform 3×3×3 convolution to get the feature map S ij ', feature map S ij 'After average pooling, multi-layer perceptron and Sigmoid activation, it is compared with the unprocessed feature map S ij 'Multiply the corresponding position elements and then add them to the unprocessed feature map S ij 'Add the corresponding position elements to get the feature map S j '.

7. The multimodal MRI brain tumor segmentation method based on teacher-student learning according to claim 6, characterized in that: The cross-modal fusion module receives the output of the modality fusion module of the four modal corresponding layers in the encoder and takes it as input, and finally outputs a fused tensor; denoted by f F is the Flair modal feature map, f T2 is the T2 modal feature map, f T1c is the T1c modal feature map, f T1 is the T1 modal feature map, f up is the feature map obtained by upsampling the previous layer in the decoder; first, f F Perform two 3×3×3 convolutions to obtain ψ(f F ) and ξ(f F ), f T2 Perform another 3×3×3 convolution to get ϕ(f T2 ),ξ(f F ) and ϕ(f T2 ) The corresponding position elements are multiplied point by point and activated by Softmax to obtain W FT2 , W FT2 and ψ(f F ) The corresponding position elements are multiplied point by point and then added to f F and f T2 Splicing to get f FT2 ; Then, swap f F With f T2 The position of f is obtained by following the same steps. T2F ; Then, f T1c Perform two 3×3×3 convolutions to obtain ψ(f T1c ) and ξ(f T1c ), f T1 Perform another 3×3×3 convolution to get ϕ(f T1 ),ξ(f T1c ) and ϕ(f T1 ) The corresponding position elements are multiplied point by point and activated by Softmax to obtain W T1cT1 , W T1cT1 and ψ(f T1c ) The corresponding position elements are multiplied point by point and then added to f T1c and f T1 Splicing to get f T1cT1 ; Then, swap f T1c With f T1 The position of f is obtained by following the same steps. T1T1c ; Finally, f FT2 、f T2F 、f T1cT1 、f T1T1c and f up Splicing in the channel dimension to get the final module output f out .

8. The multimodal MRI brain tumor segmentation method based on teacher-student learning according to claim 1, characterized in that: The standard Dice loss function is used for network training, and the formula is as follows: ; Where N is the total number of pixels in the image, C is the number of segmentation categories, and p ij ∈[0,1] and g ij ∈[0,1] represents the predicted value and true value of voxel i belonging to category j, Indicates a constant included to prevent division by zero.