Semi-Supervised Segmentation Method Based on Multimodal Teacher-Student Consistency Learning

By constructing a teacher-dual student model framework and multimodal consistency learning strategy, the problem of pseudo-label generation errors in semi-supervised multimodal medical image segmentation is solved, and the efficient segmentation effect is achieved under the limited labeling data, reducing the risk of overfitting.

CN119624997BActive Publication Date: 2025-05-30JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510163504.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-30
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

The existing semi-supervised multimodal medical image segmentation method is susceptible to the initial deviation of the model and insufficient heuristic methods when generating pseudo-labels, resulting in pseudo-label errors and affecting the training process and model performance.

Method used

A semi-supervised segmentation method based on multimodal teacher-student consistent learning is adopted to construct a teacher-dual student model framework, and pseudo-labels are generated through multimodal teacher model and guide single-modal student model training, while maintaining the consistency between multimodal teacher model and single-modal student model. The mixed segmentation loss of cross entropy loss and Dice loss is used, combined with a consistent learning strategy, to reduce the prediction bias between modals.

Benefits of technology

By ensuring semantic consistency and using unlabeled data, we can effectively alleviate the overfitting problems caused by insufficient labeled data, improve segmentation accuracy, and achieve reliable segmentation results in actual medical scenarios, reducing the pressure on medical conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119624997B_ABST
    Figure CN119624997B_ABST
Patent Text Reader

Abstract

The present invention is applicable to the technical field of medical image segmentation, and provides a semi-supervised segmentation method based on multi-modal teacher-student consistency learning, including the following steps: constructing a teacher-dual student model framework, which includes a multi-modal teacher model and two independent single-modal student models with different learning conditions. These two student models use segmentation networks with the same structure to process images of two different modalities respectively; in the data input and training stage, the labeled data is used to train the student models through real labels, while the unlabeled data is trained through the pseudo-labels generated by the teacher model; during the training process, the consistency of predictions between the teacher model and the student models, as well as between the two student models, is ensured. The present invention effectively alleviates the overfitting problem caused by insufficient labeled data. In the actual medical scenario with limited labeled data, the present invention can achieve reliable segmentation results with only a few labels to relieve the pressure of medical conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image segmentation, and particularly relates to a semi-supervised segmentation method based on multi-modal teacher-student consistency learning. Background Art

[0002] Multi-modal medical image segmentation is crucial for improving diagnostic accuracy in various clinical settings. To achieve optimal learning performance, a large amount of annotated data is usually required, which can effectively guide the deep model on how to generate outputs. However, the process of obtaining these annotated data requires the participation of domain experts, who must perform tedious annotation tasks and conduct multiple rounds of consensus building among peers. The preparation of the dataset is not only time-consuming but also potentially quite costly. In addition, unlabeled data cannot provide effective supervision signals for the model. Therefore, how to train a deep model for multi-modal medical image segmentation with limited annotated data has become an urgent problem to be solved. For the problem of lacking segmentation labels, semi-supervised learning methods are usually adopted, that is, using labeled and unlabeled data for training, so as to help the segmentation model learn rich label information from unlabeled images, especially when only a limited number of annotated data are available in medical image segmentation.

[0003] Currently, the state-of-the-art semi-supervised multi-modal medical image segmentation methods mainly rely on two frameworks: the teacher-student framework and the cross pseudo-supervision (CPS) framework. In the teacher-student framework, the student model is first trained using labeled data, and then the teacher model is generated through the exponential moving average (EMA) of the student model. Then, the pseudo-labels generated from the teacher model are used together with the labeled data to further train the student model, and its performance is enhanced through this collaborative feedback loop. In the cross pseudo-supervision framework, the consistency between two perturbed models is utilized, and the pseudo-labels generated by one network are used to train the other network. However, the generation of pseudo-labels in the semi-supervised learning process usually depends on the model's own predictions or some heuristic methods. If the model has biases in the initial stage, or the heuristic methods cannot accurately capture the intrinsic features of the data, the generated pseudo-labels may contain errors, thus affecting the training process and the performance of the final model.

[0004] In view of the above problems, the present invention proposes a semi-supervised segmentation method based on multi-modal teacher-student consistency learning. Summary of the Invention

[0005] The purpose of the present invention is to provide a semi-supervised segmentation method based on multi-modal teacher-student consistency learning, aiming to solve the problems proposed in the above background art.

[0006] The purpose of the present invention is achieved through the following technical solutions:

[0007] Semi-supervised segmentation method based on multi-modal teacher-student consistency learning, comprising the following steps:

[0008] Step 1, constructing a teacher-dual student model framework: The teacher-dual student model framework includes a multi-modal teacher model and two independent single-modal student models with different learning conditions; the two single-modal student models use segmentation networks with the same structure to process two different modalities of images respectively;

[0009] Step 2, data input and training: Input the two different modalities of images in the multi-modal teacher model into the two single-modal student models respectively; for the labeled data, the single-modal student model is trained with its corresponding true label; for the unlabeled data, the multi-modal teacher model generates pseudo-labels to guide the training of the single-modal student model; during the training process, maintain the consistency between the multi-modal teacher model and the single-modal student model, and at the same time generate consistent predictions between the two single-modal student models.

[0010] Further, in the said Step 2, the two single-modal student models use a hybrid segmentation loss combining cross-entropy loss and Dice loss in the medical image segmentation task, which is defined as follows:

[0011]

[0012] where, represents the segmentation loss of the student model of the first modality ; represents the segmentation loss of the student model of the second modality ; represents the corresponding student model of the first modality m 1 ; represents the corresponding student model of the second modality m 2 ; and represent the expectations of sampling from the labeled first modality dataset and the second modality dataset and ; represents the cross-entropy loss; represents the Dice loss, which is used to measure the matching degree between the model output and the true label; and represent the true labels of the first modality and the second modality respectively; represents the predicted output of the first modality student model for the input image ; represents the predicted output of the second modality student model for the input image The predicted output.

[0013] Furthermore, in step 2, when images of two different modalities are respectively input into two single-modal student models, consistent predictions are generated, and a consistency learning strategy is used to reduce the prediction deviation between these two images:

[0014]

[0015] Among them, represents the consistency loss between the two student models; represents the unlabeled first modality data set; represents the unlabeled second modality data set; represents and the symmetric KL divergence between, represents the predicted output of the first modality student model for the input image ; represents the predicted output of the second modality student model for the input image ;

[0016] Furthermore, in step 2, the specific implementation steps for maintaining the consistency between the multi-modal teacher model and the single-modal student model are:

[0017] Configure the multi-modal teacher model to have the same architecture as the single-modal student model, so that the multi-modal teacher model contains all four modalities at the same time, and perform consistency learning by aligning the outputs of the multi-modal teacher model and the two single-modal student models;

[0018]

[0019] Among them, represents the consistency loss between the teacher and student models; T represents the teacher model; represents the multi-modal input data sample formed by splicing all four modalities; m 4 represents the image after splicing all four modalities of the BraTS data set.

[0020] Compared with the prior art, the beneficial effects of the present invention are:

[0021] The present invention proposes a semi-supervised segmentation method based on multi-modal teacher-student consistency learning. This method effectively alleviates the overfitting problem caused by insufficient labeled data by ensuring semantic consistency and utilizing unlabeled data. In addition, the method also introduces multi-modal knowledge distillation technology to promote the consistency learning of pseudo-labels. In practical medical scenarios with limited labeled data, this method can achieve reliable segmentation results with only a few labels, thus helping to relieve the pressure on medical conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a flowchart of the present invention.

[0023] Figure 2 It is a framework diagram of the present invention.

[0024] Figure 3 It is a diagram of multi-modal medical image segmentation results under different models. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] In order to have a clearer understanding of the technical features, objectives, and beneficial effects of the present invention, the technical solutions of the present invention are described in detail below, but it should not be construed as a limitation on the scope of implementation of the present invention. The experimental methods described in the following embodiments are all conventional methods unless otherwise specified; the reagents and materials can be obtained from commercial channels unless otherwise specified.

[0026] The following describes the specific implementation of the present invention in detail with reference to specific embodiments.

[0027] As Figure 1 and Figure 2 shown, a semi-supervised segmentation method based on multi-modal teacher-student consistency learning provided by an embodiment of the present invention includes the following steps:

[0028] Step 1: Construct a teacher-dual student model framework: The teacher-dual student model framework includes a multi-modal teacher model and two independent single-modal student models with different learning conditions; the two single-modal student models use segmentation networks with the same structure and process images of two different modalities respectively;

[0029] Step 2: Data input and training: Input the images of two different modalities in the multi-modal teacher model into the two single-modal student models respectively; for the labeled data, the single-modal student model is trained with its corresponding true label to ensure accurate segmentation of the target area. For the unlabeled data, the multi-modal teacher model generates pseudo-labels to guide the training of the single-modal student model, ensuring that high-quality pseudo-labels provide sufficient supervision information for the student model to make up for the lack of unlabeled data and enabling the network to continuously improve the segmentation performance.

[0030] During the training process, the consistency between the multi-modal teacher model and the single-modal student model is maintained. Specifically, the multi-modal teacher model is configured with the same architecture as the single-modal student model. Through multi-modal information fusion, the prediction results generated by the multi-modal teacher model are aligned with the outputs of the two single-modal student models, thereby achieving cross-modal consistency learning. This consistency mechanism not only improves the robustness of the model but also makes more full use of the complementary information between multi-modal images, thus improving the segmentation accuracy.

[0031] Meanwhile, to ensure information sharing between multi-modalities and enhance the accuracy of segmentation, the two single-modal student models should produce consistent prediction outputs when processing images of their respective modalities. By using the consistency learning strategy, the prediction bias between the two sets of images can be effectively reduced.

[0032] The solution in step 2 includes:

[0033] (1) Dual semi-supervised student segmentation network: The two single-modal student models widely use a hybrid segmentation loss that combines cross-entropy loss and Dice loss, and its loss is defined as follows:

[0034] The two single-modal student models use a hybrid segmentation loss that combines cross-entropy loss and Dice loss in the medical image segmentation task, which is defined as follows:

[0035]

[0036] Among them, represents the segmentation loss of the student model of the first modality ; represents the segmentation loss of the student model of the second modality ; represents the corresponding student model of the first modality m 1 ; represents the corresponding student model of the second modality m 2 ; and represent the expectations of sampling and from the labeled first modality dataset and the second modality dataset ; represents the cross-entropy loss; represents the Dice loss, which is used to measure the matching degree between the model output and the true label; and represent the true labels of the first modality and the second modality respectively; represents the input image of the first modality student model The predicted output; Indicates the predicted output of the second-modal student model for the input image The predicted output.

[0037] In addition, to make the most of the unlabeled data, it is assumed that consistent predictions should be produced when the images of the two modalities are input into the two single-modal student models. At the same time, a consistency learning strategy is used to reduce the prediction bias between these two sets of images:

[0038]

[0039] Among them, Indicates the consistency loss between the two student models; Indicates the unlabeled first modality Dataset; Indicates the unlabeled second modality Dataset; Indicates And The symmetric KL divergence between Indicates the predicted output of the first-modal student model for the input image The predicted output; Indicates the predicted output of the second-modal student model for the input image The predicted output.

[0040] Therefore, these two single-modal student models can not only learn the semantic consistency between modalities, but also use the unlabeled data for semi-supervised training, make up for the deficiency of the supervision signal, and alleviate the overfitting problem caused by insufficient labeling.

[0041] (2) Multi-modal knowledge distillation for pseudo-label consistency learning: Configure the multi-modal teacher model to have the same architecture as the single-modal student model, change the input of the multi-modal teacher model so that it contains four modalities at the same time, and then perform consistency learning by aligning the output of the multi-modal teacher model with the output of the two single-modal student models;

[0042]

[0043] Among them, Indicates the consistency loss between the teacher and student models; T Indicates the teacher model; Indicates the multi-modal input data sample formed by splicing four modalities; m 4 Indicates the image after splicing the four modalities of the BraTS dataset.

[0044] Example 1. The present invention has achieved good segmentation results on the registered BraTS dataset, outperforming other competing methods;

[0045] Specifically, the present invention (consistency learning model) was compared with a baseline model (using the Unet method), Y-shaped model, X-shaped model, Ummkd model, SSUML model, DSAN model, LE-UDA model, and ASTCMSeg model in multi-modal medical image segmentation on MRI (magnetic resonance imaging) and CT (computed tomography) imaging. The results are as Figure 3 shown. Compared with the segmentation results of other models, the segmentation results of the method of the present invention are significantly closer to the ground truth in terms of edges and more segmentation details.

[0046] In the segmentation tasks of the whole tumor (WT), enhancing tumor (ET), and tumor core (TC) regions, the model of the present invention has achieved significant improvements in evaluation metrics such as the Dice Similarity Coefficient (DSC), Sensitivity, and Precision. Specifically: compared with the baseline model, the DSC of the model of the present invention in the whole tumor region has increased by 11.4% and reached 88.7%; the sensitivity in the enhancing tumor region has increased by 7.5%; the precision in the tumor core region has increased by 8.9%. The specific results are shown in Table 1.

[0047]

[0048] The above are only the preferred embodiments of the present invention. It should be noted that for those skilled in the art, without departing from the concept of the present invention, several modifications and improvements can be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent.

Claims

1. A semi-supervised segmentation method based on multimodal teacher-student consistency learning, characterized by: The following steps are involved: Step 1: Construct a teacher-dual student model framework: The teacher-dual student model framework includes a multimodal teacher model and two independent unimodal student models with different learning conditions; the two unimodal student models use the same segmentation network to process images of two different modalities respectively; Step 2, data input and training: The images of two different modalities in the multimodal teacher model are input into the two unimodal student models respectively; for labeled data, the unimodal student model is trained by the corresponding real labels; for unlabeled data, the multimodal teacher model guides the training of the unimodal student model by generating pseudo labels; during the training process, the consistency between the multimodal teacher model and the unimodal student model is maintained, and consistent predictions are generated between the two unimodal student models; In step 2, the two unimodal student models use a hybrid segmentation loss that combines cross entropy loss and Dice loss in the medical image segmentation task, which is defined as follows: in, represents the segmentation loss of the student model of the first modality m1; represents the segmentation loss of the student model of the second modality m2; S 1 represents the student model of the corresponding first mode m1; S 2 represents the student model of the corresponding second mode m2; and Represents the first modality dataset from the labeled And the second modality dataset Extract samples from and Expectations; L CE represents the cross entropy loss; L D ice represents Dice loss, which is used to measure the degree of match between the model output and the true label; and Represent the true labels of the first modality and the second modality respectively; Represents the first modality student model for the input image The predicted output of Represents the second modality student model for the input image The predicted output of In step 2, when two images of different modalities are respectively input into two single-modal student models, consistent predictions are generated, and a consistent learning strategy is used to reduce the prediction deviation between the two images: in, represents the consistency loss between two student models; Represents the unlabeled first modality m1 dataset; represents the unlabeled second modality m2 dataset; L KL express and The symmetric KL divergence between Represents the first modality student model for the input image The predicted output of Represents the second modality student model for the input image The predicted output of In step 2, the specific implementation steps for maintaining the consistency between the multimodal teacher model and the unimodal student model are as follows: The multimodal teacher model is configured with the same architecture as the unimodal student model, so that the multimodal teacher model contains four modalities at the same time, and consistency learning is performed by aligning the outputs of the multimodal teacher model with the two unimodal student models; in, represents the consistency loss between the teacher and student models; T represents the teacher model; Represents a multimodal input data sample composed of four modalities; m4 represents an image composed of four modalities of the BraTS dataset.

Citation Information

Patent Citations

  • Semi-supervised remote sensing image semantic segmentation method based on double consistency

    CN116416618A

  • Full-modal and missing-modal land coverage classification method based on multi-modal online distillation framework

    CN118196649A