A semi-supervised medical image segmentation method and system based on mutual learning

By adopting a semi-supervised learning method based on mutual learning in medical image segmentation, a deep semi-supervised learning network is built, and a multi-level consistency regularity and cross-modal distribution alignment mechanism is introduced, the problems of insufficient utilization of labeled and unlabeled data and insufficient single-modal data in the existing technology are solved, and higher segmentation accuracy and generalization capabilities are achieved.

CN114418954BActive Publication Date: 2025-05-23SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111601008.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-24
Publication Date
2025-05-23
Estimated Expiration
2041-12-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize limited labeled data and a large amount of unlabeled data in medical image segmentation, and single-modal data is difficult to capture the coordinated semantic information between different levels and different factors in the image.

Method used

Using a semi-supervised learning method based on mutual learning, a deep semi-supervised learning network based on mutual learning is constructed by introducing at least two semi-supervised learning models. The network forces the class prediction probability distribution between different subnets to remain consistent through alternating mutual supervision of the teacher network and the student network, as well as imitation loss functions in the student network. At the same time, a multi-level consistency regular constraint and cross-modal distribution alignment mechanism are introduced to enhance the model's processing ability of multi-modal data.

Benefits of technology

It improves the robustness of the network for feature extraction of labeled samples and labelless samples, enhances the segmentation accuracy and generalization ability of multimodal medical images, and effectively utilizes the information in unlabeled data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114418954B_ABST
    Figure CN114418954B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image segmentation and recognition technology, and specifically to a semi-supervised medical image segmentation method based on mutual learning and a system thereof. The method of the present invention is based on the technical idea of ​​the mutual learning algorithm, adopts at least two semi-supervised learning models for dual combination, and constructs a deep semi-supervised learning network based on mutual learning. By alternately supervising different sub-networks (student network, teacher network) in the network during the training process, and forcing the class prediction probabilities outputted by them to remain consistent, the judgment accuracy of unlabeled sample images is effectively improved, and the robustness of feature extraction of labeled samples and unlabeled samples is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image segmentation and recognition, and in particular to a semi-supervised medical image segmentation method based on mutual learning and a system thereof. Background Art

[0002] In the medical field, the high cost of manual labeling and the inevitable ambiguity of labels have limited the ability of AI-based auxiliary diagnosis models to achieve higher accuracy and better generalization. However, it is very expensive and time-consuming to annotate large-scale training data for higher accuracy. For example, medical imaging datasets generate massive amounts of imaging data every day, but due to the professionalism of the labeling itself and the inevitable labeling time cost, they may only contain a small part of the labeled data from experts, leaving a large part unlabeled. This part of unlabeled data cannot be used in the paradigm of supervised learning, but in fact this part of the data contains a large amount of available "unexpressed" information. Combined with the labeling information, effectively mining this part of the sample data has a high clinical application value. Therefore, how to use limited labeled data and a large amount of unlabeled data to build an effective model is one of the key challenges for the generalization of current artificial intelligence models.

[0003] On the other hand, in today's rapidly developing era of big data, a large amount of data is generated from different fields every day. Through these different types / modal data, the model's representation of the same thing has become more specific and comprehensive. In the medical field, when doctors diagnose whether a patient has a certain disease, they can summarize the patient's various medical imaging data, check their medical records, or obtain clinical pathology results, and make a more accurate judgment by integrating this information. However, most of the current image segmentation methods based on artificial intelligence are modeled by statistical analysis or machine learning of single-modal data. In fact, it is difficult to discover the collaborative semantic information contained in different levels and factors in the image by relying solely on single-modal data. For example, when judging the area and category of the sample image, since single-modal images can only specifically reflect certain physiological structures or functional changes, the judgment accuracy is not satisfactory. Different modal medical imaging data, such as MRI, DT1, and fMRI, can provide physiological structure regional status information from their own specific angles, providing auxiliary assistance for the diagnosis of diseases. However, the types and label attributes of the data are quite different, making it difficult to judge by an image segmentation method.

[0004] Existing semi-supervised learning allows the model to integrate some or all of the unlabeled data in its supervised learning, maximizing the model's learning performance through this large amount of unlabeled data while minimizing the cost of labeled data [1][2]. In semi-supervised learning methods, they are generally supported by some assumptions [3]: Smoothness & Low-density assumption and Manifold assumption.

[0005] The smooth-low density hypothesis means that when the distances between sample data are relatively close, they have the same category and the decision boundary should try to pass through areas where the data is relatively sparse, avoiding dividing dense samples on both sides of the decision boundary. Under this assumption, the learning algorithm can use a large amount of unlabeled sample data to analyze the sample distribution in the sample space, thereby guiding the learning algorithm to adjust the classification boundary so that it passes through areas where the sample data layout is relatively sparse.

[0006] The manifold hypothesis is mainly based on the fact that the input space is composed of multiple low-dimensional manifolds, all data points are located on these manifolds, and data points on the same manifold have the same label. Thanks to the development of deep learning technology [4-6], semi-supervised learning methods have developed rapidly in the past few years. For example, Dong-Hyun Lee [7] proposed a very simple and effective pseudo-label utilization scheme. First, a model is trained on a batch of labeled and unlabeled images at the same time. The labeled images are trained in a normal supervised manner. The same model is used to predict a batch of unlabeled images, and the class with the largest confidence is used as the pseudo-label. Ding et al. [8] learned a generative model through adversarial training to approximate the actual data manifold, and proposed a new pseudo-labeling method based on feature similarity. Two possible pseudo-label encodings are used under a unified setting to improve the distinguishability of features. Laine et al. [9] proposed the π-Model. For any given input, two predictions are made using different regularizations. The goal is to reduce the distance between the two predictions and improve the consistency of the model under different perturbations. Temporal Ensembling was further proposed based on the π-Model. Its overall framework is similar to that of the π-Model. The same idea is adopted in the processing of obtaining unlabeled data. The temporal combination model is used in the unsupervised term of the objective function to effectively retain historical information and stabilize the current value. Xie et al.

[10] used automatic enhancement to create an enhanced version of an unlabeled sample. Then the same model was used to predict the labels of the two samples before and after enhancement, and the distribution of the two predictions was constrained to remain consistent. FixMatch

[11] used cross entropy to regularize the consistency of weakly enhanced and strongly enhanced unlabeled data, used weakly enhanced data as pseudo labels, and performed consistency regularization by using cross entropy.

[0007] The main goal of existing multimodal fusion is to reduce the heterogeneity between modalities while maintaining the integrity of the specific semantics of each modality. Currently, multimodal fusion architectures are divided into three categories

[12]

[13] : joint architecture, coordinated architecture, and encoder-decoder architecture. The joint architecture projects the single modality representation into a shared semantic subspace so that multimodal features can be fused; each single modality is mapped to a shared subspace after a separate encoding. Following this strategy, it has shown excellent performance in multimodal classification or regression tasks such as video classification

[14] , emotion recognition

[15] , and visual question answering

[16] .

[0008] At present, joint representation methods based on neural networks have shown superior performance and can pre-train representations without supervision. However, the performance improvement depends on the number of training samples. Collaborative architectures include cross-modal similarity models and canonical correlation analysis, which aim to coordinate the correlation between modalities in the subspace; the mainstream collaborative method is based on the cross-modal similarity method, which aims to learn a shared subspace to maximize the correlation between different modal representation sets

[17] . In recent years, neural networks have become a commonly used method for constructing coordinated representations. Its advantage is that it can jointly learn coordinated representations in an end-to-end manner

[18] . However, its disadvantage is that modal fusion is difficult and cross-modal learning models are not easy to implement. The encoder-decoder architecture mainly utilizes the intermediate representation of modal mapping. It is usually used in multimodal conversion tasks that map one modality to another modality, and mainly consists of two parts: an encoder and a decoder

[19]

[20] . The encoder-decoder architecture mainly focuses on the shared semantic capture and encoding and decoding problems of multimodal sequences. In order to more effectively capture the shared semantics of the two modalities, the encoder usually uses some regularization techniques to maintain the semantic consistency between the modalities. The decoder is responsible for reasoning high-level semantics to ensure the correct understanding of the semantics in the source modality and the generation of new samples in the target modality. Compared with other frameworks, the advantage of the encoder-decoder framework is that it can generate new target modality samples based on the source modality. Its disadvantage is that each encoder and decoder can only encode one of the modalities. There are deficiencies in the existing technology.

[0009] References:

[0010] [1]Van Engelen JE, Hoos H H. "A survey on semi-supervised learning," Machine Learning, 2020, 109(2): 373-440.

[0011] [2] Zhou Zhihua, Wang Qian. Machine Learning and Its Applications, Beijing: Tsinghua University Press, 2007.

[0012] [3] Liu Jianwei, Liu Yuan, Luo Xionglin. “Semi-supervised learning methods,” Chinese Journal of Computers, 2015, 38(8): 1592-1617.

[0013] [4]Olsson V,Tranheden W,Pinto J,Svensson,L.“Classmix:Segmentation-based data augmentation for semi-supervised learning,”in Proceedings of theIEEE Winter Conference on Applications of Computer Vision,2021:1369-1378.

[0014] [5]Zhai X,Oliver A,Kolesnikov A,Beyer,L.“S4l:Self-supervised semi-supervised learning,”in Proceedings of the IEEE International Conference onComputer Vision,2019:1476-1485.

[0015] [6]Wang Q,Li W,Gool L V.“Semi-supervised learning by augmenteddistribution alignment,”in Proceedings of the IEEE International Conferenceon Computer Vision,2019:1466-1475.

[0016] [7]Lee D H.“Pseudo-label:The simple and efficient semi-supervisedlearning method for deep neural networks,”in Workshop on challenges inrepresentation learning,ICML,2013,3(2).

[0017] [8]Ding G,Zhang S,Khan S,Tang,Z.,Zhang,J.,Porikli,F.“Featureaffinity-based pseudo labeling for semi-supervised person re-identification,”IEEE Transactions on Multimedia,2019,21(11):2891-2902.

[0018] [9]Laine S,Aila T.“Temporal ensembling for semi-supervised learning,”in Proceedings of the International Conference on LearningRepresentations.2017.

[0019]

[10] Xie Q,Dai Z,Hovy E,Luong,M.T,Le,Q.V.“Unsupervised dataaugmentation for consistency training,”arXiv preprint arXiv:1904.12848,2019.

[0020]

[11] Sohn K et al.“FixMatch:Simplifying Semi-Supervised Learning withConsistency and Confidence,”in Proceedings of the Advances in NeuralInformation Processing Systems,2020,33.

[0021]

[12] T,Ahuja C,Morency L P.“Multimodal machine learning:A survey and taxonomy,”IEEE Transactions on Pattern Analysis and MachineIntelligence,2018,41(2):423-443.

[0022]

[13] Zhang C,Yang Z,He X,Deng,L.“Multimodal intelligence:Representation learning,information fusion,and applications,”IEEE Journal ofSelected Topics in Signal Processing,2020,14(3):478-493.

[0023]

[14] Qi M,Qin J,Yang Y,Wang,Y.,Luo,J.“Semantics-Aware Spatial-TemporalBinaries for Cross-Modal Video Retrieval,”IEEE Transactions on ImageProcessing,2021,30:2989-3004.

[0024]

[15] Zhang J,Yin Z,Chen P,Nichele,S.“Emotion recognition using multi-modal data and machine learning techniques:A tutorial and review,”InformationFusion,2020,59:103-126.

[0025]

[16] Anderson P,He X,Buehler C,Teney,D.,Johnson,M.,Gould,S.,Zhang,L.“Bottom-up and top-down attention for image captioning and visual questionanswering,”in Proceedings of the IEEE Conference on Computer Vision andPattern Recognition,2018:6077-6086.

[0026]

[17] Zhang J, Peng Y, Yuan M. "Sch-gan: Semi-supervised cross-modalhashing by generative adversarial network," IEEE Transactions on Cybernetics, 2018, 50(2): 489-502.

[0027]

[18] Zhang H, Wang Y, Long Y, Yang, L., Shao, L. "Modality independent adversarial network for generalized zero shot image classification," NeuralNetworks, 2021, 134: 11-22.

[0028]

[19] Zhang P, Zhang B, Chen D, Yuan, L., Wen, F. "Cross-domain correspondencelearning for exemplar-based image translation," in Proceedings of the IEEEConference on Computer Vision and Pattern Recognition, 2020: 5143-5153.

[0029]

[20] Huang X, Liu MY, Belongie S, et al. Multimodal unsupervised image-to-image translation," in Proceedings of the European Conference on ComputerVision, 2018: 172-189. Summary of the invention

[0030] In order to solve at least one of the above technical problems, the present invention is based on the improvement of the supervision mode in the training of the semi-supervised learning model, and improves the robustness of the network for the feature extraction of labeled samples and unlabeled samples under limited annotated data resources. An embodiment of the present invention provides a semi-supervised medical image segmentation method based on mutual learning, comprising the following steps:

[0031] S1. Introduce at least two semi-supervised learning models;

[0032] S2. The two semi-supervised models are connected in a dual form, so that the teacher network and the student network in the two semi-supervised models can realize multi-teacher-student interaction to construct a deep semi-supervised learning network based on mutual learning;

[0033] S3. Input samples to the deep semi-supervised learning network based on mutual learning, and force the category prediction probability distributions between different student networks, between teacher networks, and between student networks and teacher networks to remain consistent through alternating mutual supervision between the teacher network and the student network during the training process and the imitation loss function in the student network;

[0034] The parameters of the teacher network are obtained by moving average of the parameters of the student network during the training process; the samples include labeled sample images and unlabeled sample images.

[0035] The present invention also provides a semi-supervised medical image segmentation system based on mutual learning, comprising: a sample, and also a deep semi-supervised learning network model based on mutual learning;

[0036] The deep semi-supervised learning network model based on mutual learning is composed of at least two semi-supervised learning models;

[0037] The two semi-supervised models are connected in a dual form, so that the teacher network and the student network in the two semi-supervised models can realize multi-teacher-student interaction;

[0038] After the sample is input, the deep semi-supervised learning network based on mutual learning forces the category prediction probability distributions between different student networks, between teacher networks, and between student networks and teacher networks to remain consistent through alternating mutual supervision between the teacher network and the student network during the training process and the imitation loss function in the student network;

[0039] The parameters of the teacher network are obtained by moving average of the parameters of the student network during the training process; the samples include labeled sample images and unlabeled sample images.

[0040] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the medical image segmentation method as described above are executed.

[0041] The semi-supervised medical image segmentation method and system based on mutual learning in the embodiment of the present invention adopts at least two semi-supervised learning models for dual combination to construct a deep semi-supervised learning network based on mutual learning. By alternately supervising different sub-networks (student network, teacher network) in the network during the training process and forcing the output category prediction probability distribution to be consistent, the judgment accuracy of unlabeled sample images is effectively improved, and the robustness of feature extraction of labeled samples and unlabeled samples is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0043] Figure 1 A flowchart of a semi-supervised medical image segmentation method based on mutual learning according to the present invention;

[0044] Figure 2 It is a functional schematic diagram of a deep semi-supervised learning network based on mutual learning of the present invention;

[0045] Figure 3 A functional schematic diagram of cross-modal distribution alignment of the student network of the present invention;

[0046] Figure 4 It is the constraint type of the multi-level consistency constraint of the present invention. DETAILED DESCRIPTION

[0047] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0048] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0049] See also Figure 1-2 According to an embodiment of the present invention, a semi-supervised medical image segmentation method based on mutual learning is provided, comprising the following steps:

[0050] S1. Introduce at least two semi-supervised learning models.

[0051] S2. The two semi-supervised models are connected in a dual form, so that the teacher network and the student network in the two semi-supervised models can achieve multi-teacher-student interaction, so as to construct a deep semi-supervised learning network based on mutual learning and realize feature supervision and correction of the image.

[0052] The main network is constructed as follows: a dual semi-supervised learning (mean-teacher) model is introduced, in which the teacher network (T 1 and T 2 ) The parameters updated in the tth iteration By Student Network (S 1 and S 2 ) The tth parameter in the training process and the parameters of the teacher network at the t-1th time The moving average is obtained, where ε represents the specified weight index:

[0053]

[0054]

[0055] S3. Input samples to the deep semi-supervised learning network based on mutual learning, and force the category prediction probability distributions between different student networks, between teacher networks, and between student networks and teacher networks to remain consistent through alternating mutual supervision between the teacher network and the student network during the training process and the imitation loss function in the student network;

[0056] In practice, during the training process, the loss function of each student network consists of two parts:

[0057] (1) Calculate supervised learning loss using labeled samples Cross entropy loss is usually used to represent the supervised learning loss;

[0058] (2) Imitation loss: Make the category prediction probabilities between different students, between teachers, and between students and teachers consistent (labeled data is not explicitly used here, but by forcing the output distribution between different objects to remain consistent, unlabeled training samples can effectively participate in the network training process through imitation loss).

[0059] The parameters of the teacher network are obtained by moving average of the parameters of the student network during the training process; the samples include labeled sample images and unlabeled sample images.

[0060] The present invention is designed based on mutual learning, and a deep semi-supervised learning (mean-teacher) network based on mutual learning is constructed to improve the robustness of feature extraction of labeled sample images and unlabeled sample images by student networks and teacher networks in semi-supervised learning, and realize mutual correction of features. Mutual learning begins with a group of basic networks, learning to solve the same target task together. In the mutual learning network, each sub-network effectively aggregates their collective estimates of the next most likely category, which helps to obtain a more robust and generalized network.

[0061] Furthermore, the loss function of the student network also includes: a supervised learning loss function; the supervised learning loss function is calculated through the labeled sample images.

[0062] Furthermore, in step S3, relative entropy is used to measure the consistency of the category prediction probability distribution.

[0063] In practice, the Kullback-Leibler (relative entropy) divergence D is used. KL (.||.), that is, the relative entropy between teacher networks (D KL (T 1 ||T 2 )), between student networks (D KL (S 1 ||S 2 )) and between teacher-student networks (D KL (T 1 ||S 2 ) and D KL (T 2 ||S 1 ))Class prediction distribution consistency measure of prediction output:

[0064]

[0065] In this peer-teaching-based training, the network is updated by minimizing the above loss function, so that the learning effect of the teacher-student network of each branch is improved compared with individual learning. At the same time, the robustness of the network is further enhanced through multi-teacher and student interaction.

[0066] like Figure 4 The step S3 shown in the alternating mutual supervision process also includes the following steps:

[0067] S4. Considering samples of the same and different categories and the relationship between samples, multi-level consistency constraints are constructed to extract useful semantic information from the unlabeled sample images, so as to use the unlabeled sample images more effectively and reduce the impact of irrelevant unlabeled sample images on algorithm performance.

[0068] Furthermore, the levels include: a single sample instance level, a multi-sample relationship level, and a category level.

[0069] Among them, the sample instance level: build the consistency of the sample and its corresponding perturbation sample prediction output, for example, for the sample feature f 1 and its perturbed characteristic f′ 1 , constraining its predicted output value to remain consistent, defining C 1 Indicates the consistency of sample instances, p(f 1 ), p(f′ 1 ) represents the sample feature f 1 and the perturbed feature f′ 1 The predicted distribution of Represents the Euclidean distance metric. Its expression is:

[0070]

[0071] Level of relationship between samples: given multiple sample instances Build a relationship diagram Where V represents a set of multiple instance features, and the relationship A(i, j) between instance features constitutes an edge set E, where each edge is defined as the Euclidean distance between two adjacent instance features. N represents the number of sample instances;

[0072] By the same token, we can construct a relationship graph corresponding to the perturbation sample instance Constrain the consistency between their relationship matrices and define C 2 Indicates the consistency of the relationship between samples, A, A′ respectively represent the sample characteristics Sample characteristics after disturbance The relationship matrix between represents the Frobenius norm metric. Its expression is:

[0073]

[0074] Category relationship level: The consistency of the relationship level between samples describes the relationship (similarity / dissimilarity) between a group of sample instances, but it still lacks the effective introduction of category information. In order to further make the features learned by the network have better inter-class and intra-class properties (large inter-class distance and small intra-class distance), a category-level consistency regularization constraint is constructed. First, some unlabeled samples are marked by high-confidence pseudo-class labels (obtained based on the mutual learning algorithm), and the corresponding class center set can be calculated. Use the corresponding perturbed sample features to get the perturbed class center Calculate the similarity matrix of class centers and the perturbed class center similarity matrix S′, K represents the number of categories, constraining the consistency between their similarity matrices, and defining C 3 Indicates the consistency of category relationship, represents the Frobenius norm metric. Its expression is:

[0075]

[0076] Through the above multi-level consistency regularization constraints, the semantic information of samples, between samples and between categories is effectively mined, and the effective utilization rate of unlabeled samples by the semi-supervised model is improved. In summary, the overall multi-level consistency regularization constraint C can be expressed as follows:

[0077] C=α 1 C 1 +α 2 C 2 +α 3 C 3 ;

[0078] α 1 , α 2 , α 3 is a balance parameter used to modulate the weights of each constraint.

[0079] In the above embodiment, further introducing a multi-level consistency constraint mechanism into the mutual learning semi-supervised model can further improve the semantic mining of unlabeled sample images by the mutual learning deep semi-supervised learning network and reduce the adverse effects of unlabeled sample images on the model prediction results.

[0080] like Figure 3 As shown, in order to introduce multimodal information and avoid overly complex network design, step S3 also includes the following steps when inputting samples:

[0081] S5. The student network eliminates the distribution offset of samples between different modalities through a cross-modal distribution alignment mechanism;

[0082] The cross-modal distribution alignment mechanism is implemented by maximizing the average difference algorithm to measure the similarity of feature distributions between different modalities.

[0083] The design of the present invention eliminates the distribution offset between different modalities through a cross-modal distribution alignment mechanism. Since the student and teacher networks in the mutual learning framework use the same network structure, and the parameter update of the teacher network is provided by the student network, we only need to build an adapted student model ( Figure 2 ). At the same time, in order to avoid introducing too many learning parameters, the present invention performs cross-modal distribution alignment by maximizing the mean discrepancy (MMD). Assume that there is a common low-dimensional manifold so that multimodal data can be projected into this subspace. We intend to use MMD to measure the similarity of feature distributions between different modalities. Without loss of generality, for any two modal feature representations z i ,z j ,MMD can be calculated as follows:

[0084]

[0085] Assume that z i ,z j They are obtained by independent and identically distributed sampling from distributions p and q, respectively. T represents a set of continuous functions in the sample space, F represents the set of all functions τ mapped from the feature space to the real number set, E represents the mathematical expectation, and sup(.) represents the supremum. In the present invention, τ is calculated using a Gaussian kernel function. During the training process, the cross-modal distribution alignment is achieved by minimizing the MMD difference, thereby effectively promoting the fusion of multimodal information.

[0086] In the above embodiment, since the parameters of the teacher network are obtained by moving average of the parameters of the student network during the training process, when the cross-modal distribution alignment mechanism is introduced in the student network for multi-modal medical images, the teacher network will also introduce the mechanism to process multi-modal sample images accordingly. This not only expands the samples of the mutual learning deep semi-supervised learning network from single modality to multi-modality, but also further eliminates the redundancy of modal information, maximizes the cross-modal information complementarity, and improves the prediction performance of the model.

[0087] The present invention also provides a semi-supervised medical image segmentation system based on mutual learning, comprising: a sample, and also a deep semi-supervised learning network model based on mutual learning;

[0088] The deep semi-supervised learning network model based on mutual learning is composed of at least two semi-supervised learning models;

[0089] The two semi-supervised models are connected in a dual form, so that the teacher network and the student network in the two semi-supervised models can realize multi-teacher-student interaction;

[0090] After the sample is input, the deep semi-supervised learning network based on mutual learning forces the category prediction probability distributions between different student networks, between teacher networks, and between student networks and teacher networks to remain consistent through alternating mutual supervision between the teacher network and the student network during the training process and the imitation loss function in the student network;

[0091] The parameters of the teacher network are obtained by moving average of the parameters of the student network during the training process; the samples include labeled sample images and unlabeled sample images.

[0092] Furthermore, the loss function of the student network also includes: a supervised learning loss function, which is calculated using the labeled sample images.

[0093] Furthermore, the semi-supervised model also includes a multi-level consistency regularization module, which is used to consider samples of the same and different categories and the relationship between samples, construct consistency constraints at multiple levels, and extract useful semantic information from the unlabeled sample images.

[0094] In the above embodiment, further introducing a multi-level consistency constraint mechanism into the mutual learning semi-supervised model can further improve the semantic mining of unlabeled sample images by the mutual learning deep semi-supervised learning network and reduce the adverse effects of unlabeled sample images on the model prediction results.

[0095] Furthermore, the student network also includes a multimodal processing unit that eliminates the distribution offset of samples between different modalities through a cross-modal distribution alignment mechanism.

[0096] In the above embodiment, since the parameters of the teacher network are obtained by moving average of the parameters of the student network during the training process, when the cross-modal distribution alignment mechanism is introduced in the student network for multi-modal medical images, the teacher network will also introduce the mechanism to process multi-modal sample images accordingly. This not only expands the samples of the mutual learning deep semi-supervised learning network from single modality to multi-modality, but also further eliminates the redundancy of modal information, maximizes the cross-modal information complementarity, and improves the prediction performance of the model.

[0097] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the medical image segmentation methods described above are executed.

[0098] In summary, the semi-supervised medical image segmentation method and system based on mutual learning of the present invention integrates multi-level consistency regularization constraints into the semi-supervised model, and adopts cross-modal distribution matching in the student network in the model, so that the deep semi-supervised learning network model based on mutual learning can perform fusion segmentation on multi-modal sample images, and reduce the impact of irrelevant unlabeled samples on algorithm performance through multi-level consistency regularization constraints. Under the condition of limited labeled samples, the utilization efficiency of the semi-supervised learning method for unlabeled sample images is maximized, and the accuracy and generalization ability of the model are improved.

[0099] In addition, the specific process of loading and executing the multiple instruction processors in the above-mentioned storage medium and the terminal device has been described in detail in the above-mentioned method, and will not be described one by one here.

[0100] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A semi-supervised medical image segmentation method based on mutual learning, It is characterized in that The following steps are involved: S1. Introduce at least two semi-supervised learning models; S2. The two semi-supervised models are connected in a dual form, so that the teacher network and the student network in the two semi-supervised models can realize multi-teacher-student interaction to construct a deep semi-supervised learning network based on mutual learning; S3. Input samples to the deep semi-supervised learning network based on mutual learning, and force the category prediction probability distributions between different student networks, between teacher networks, and between student networks and teacher networks to remain consistent through alternating mutual supervision between the teacher network and the student network during the training process and the imitation loss function in the student network; The parameters of the teacher network are obtained by moving average of the parameters of the student network during the training process; the samples include labeled sample images and unlabeled sample images.

2. The medical image segmentation method according to claim 1, It is characterized in that The loss function of the student network also includes: a supervised learning loss function; the supervised learning loss function is calculated through the labeled sample images.

3. The medical image segmentation method according to claim 1, It is characterized in that In step S3, relative entropy is used to measure the consistency of the category prediction probability distribution.

4. The medical image segmentation method according to claim 1, It is characterized in that The step S3 further comprises the following steps in the alternating mutual supervision process: S4. Considering samples of the same and different categories and the relationships between samples, consistency constraints are constructed at multiple levels to extract useful semantic information from the unlabeled sample images.

5. The medical image segmentation method according to claim 4, It is characterized in that The levels include: a single sample instance level, a multi-sample relationship level, and a category relationship level.

6. The medical image segmentation method according to claim 1, It is characterized in that When the sample is input, step S3 further comprises the following steps: S5. The student network eliminates the distribution offset of samples between different modalities through a cross-modal distribution alignment mechanism; The cross-modal distribution alignment mechanism is implemented by maximizing the average difference algorithm to measure the similarity of feature distributions between different modalities.

7. A semi-supervised medical image segmentation system based on mutual learning, include: The sample is characterized in that it also includes a deep semi-supervised learning network model based on mutual learning; The deep semi-supervised learning network model based on mutual learning is composed of at least two semi-supervised learning models; The two semi-supervised models are connected in a dual form, so that the teacher network and the student network in the two semi-supervised models can realize multi-teacher-student interaction; After the sample is input, the deep semi-supervised learning network based on mutual learning forces the category prediction probability distributions between different student networks, between teacher networks, and between student networks and teacher networks to remain consistent through alternating mutual supervision between the teacher network and the student network during the training process and the imitation loss function in the student network; The parameters of the teacher network are obtained by moving average of the parameters of the student network during the training process; the samples include labeled sample images and unlabeled sample images.

8. The image segmentation system according to claim 7, It is characterized in that The loss function of the student network also includes: a supervised learning loss function, which is calculated using the labeled sample images.

9. The image segmentation system according to claim 8, It is characterized in that The semi-supervised model also includes a multi-level consistency regularization module, which is used to consider samples of the same and different categories and the relationship between samples, construct consistency constraints at multiple levels, and extract useful semantic information from the unlabeled sample images.

10. The image segmentation system according to claim 9, It is characterized in that The student network also includes a multimodal processing unit that eliminates the distribution offset of samples between different modalities through a cross-modal distribution alignment mechanism.

11. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the medical image segmentation method according to any one of claims 1 to 6 are performed.

Citation Information

Patent Citations

  • Medical image segmentation method, segmentation system, and computer-readable storage medium

    CN109344833A

  • Semi-supervised medical image segmentation method based on generative adversarial network

    CN112837338A