Semi-supervised mutual learning medical image segmentation method based on conflict perception
By adopting technologies such as multi-subnet architecture and conflict-perceived differentiated feature learning regularization in semi-supervised medical image segmentation, the problem of high inaccuracy in prediction of unlabeled data is solved, and the segmentation performance of the model is significantly improved.
Patent Information
- Application Number
- CN202411782761.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-06-03
AI Technical Summary
The existing semi-supervised medical image segmentation method has high prediction inaccuracy when utilizing unlabeled data, resulting in poor performance of the model on test data.
Two subnets with different architectures are adopted to maximize the differences in the subnet extract features, avoid feature homogeneity, and improve the generalization ability of the model through cross-supervised and selective supervision of pseudo-supervised and geometric perception.
Through multi-subnet architecture and regularization technology, the segmentation performance of the model on test medical images is improved, and the evaluation indicators such as Dice similarity coefficient, Jaccard index, accuracy, 95% Hausdorff distance and true positive rate are significantly improved.
Smart Images

Figure CN120088470A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of semi-supervised medical image segmentation methods, and specifically to a semi-supervised mutual learning medical image segmentation method based on conflict awareness. Background Art
[0002] Automated medical image segmentation is an important basis for computer-aided diagnosis, treatment planning, and interventional procedures. Accurate medical image segmentation can quantitatively analyze tissue morphological features, thereby enhancing its practicality in supporting the medical decision-making process.
[0003] In recent years, the development of convolutional neural networks (CNNs) has promoted the progress of automatic medical image segmentation methods, and the performance largely depends on sufficient labeled medical image data. Since collecting a large amount of labeled medical image data is time-consuming and laborious, while collecting a large amount of unlabeled medical images is relatively easy. Therefore, applying the semi-supervised learning (SSL) paradigm enables the model to utilize rich unlabeled data, thereby reducing the burden of pixel-level annotation, and gradually becoming the mainstream method for medical image segmentation tasks. However, inaccurate predictions using unlabeled data can mislead the model training process, thus significantly reducing the model's performance.
[0004] To address the above problems, the present invention proposes a conflict-aware semi-supervised mutual learning medical image segmentation method. First, the present invention employs two sub-networks with different architectures to fully exploit the potential of the prediction diversity of the multi-sub-network architecture. Subsequently, conflict-aware differential feature learning regularization (CDFL) is designed to maximize the difference in features extracted by the two sub-networks with different architectures, to avoid the homogenization of features extracted by different sub-networks and to promote prediction diversity. Then, for regions with consistent predictions, consistency cross-supervision regularization (CCS) is designed to select regions with consistent predictions in unlabeled images, and cross-supervise the prediction results of the two networks for the pixels in this region. At the same time, for regions with inconsistent predictions, geometric-aware mutual pseudo-supervision regularization (GMPS) is designed to evaluate the reliability of pseudo-labels for regions with inconsistent predictions in unlabeled images, and selectively use the more reliable pseudo-labels in the two sub-networks to supervise the training of the other self-network. Finally, the present invention designs objective functions for the two sub-networks respectively according to the above regularizations and optimizes them. Summary of the Invention
[0005] In view of the existing problems above, the present invention is proposed.
[0006] Therefore, the technical problem solved by the present invention is that the current technical means is to use rich unlabeled images to assist the training of semi-supervised segmentation models. However, inaccurate predictions of unlabeled data will mislead the training process of the model, making it difficult for most existing semi-supervised medical image segmentation methods to perform well on test data.
[0007] To solve the above technical problems, the present invention provides the following technical solutions: A conflict-aware semi-supervised mutual learning medical image segmentation method, comprising: collecting a medical image training data set, dividing it into labeled and unlabeled sets, and performing preprocessing.
[0008] Initialize two sub-networks with different structures, and extract features from the training data.
[0009] Design conflict-aware differential feature learning regularization (CDFL) to maximize the difference in features extracted by two sub-networks with different architectures, so as to avoid the homogenization of features extracted by different sub-networks.
[0010] Design consistency cross-supervision regularization (CCS) and geometry-aware mutual pseudo-supervision regularization (GMPS) for mutual supervised learning of regions where the two networks predict consistently and regions where they predict inconsistently, respectively.
[0011] Construct semi-supervised segmentation objective functions for the two sub-networks respectively, and solve them.
[0012] 2 As a preferred solution of the conflict-aware semi-supervised mutual learning medical image segmentation method of the present invention, wherein: the division of the labeled and unlabeled sets includes N labeled data, which consists of N data and labels, expressed as:
[0013]
[0014] Wherein, represents the labeled image data input for the i-th time, represents the label corresponding to the labeled image data input for the i-th time; H, W, and C respectively represent the length, width, and number of channels, and Y represents the number of categories.
[0015] The unlabeled images consist of a small number of M data, expressed as:
[0016]
[0017] Wherein, represents the unlabeled image data input for the i-th time.
[0018] Perform preprocessing operations on the training set image data, such as resampling, cropping, and adjusting intensity values, etc.
[0019] 3 As a preferred solution of the conflict-aware semi-supervised mutual learning medical image segmentation method described in the present invention, the following steps are included: Two subnets A and B with different structures and different parameter initializations are constructed, denoted as f A (·) and f B (·) respectively, and the initialized weight parameters are denoted as θ A and θ B . The subnet f A includes a feature extractor h A (·): and a classifier g A (·): The initialized weight parameters are and represents the dimension z of the feature space. There is For subnet B, there is The sub-networks h A (·) and h B (·), h A (·) adopts the U-Net structure, and h B (·) adopts the DeepLabv3+ structure.
[0020] 4 As a preferred solution of the conflict-aware semi-supervised mutual learning medical image segmentation method described in the present invention, the following steps are included: The CDFL regularization is used to reduce the homogeneity of the features extracted by subnets A and B. For the labeled and unlabeled data input into subnets A and B, supervised and unsupervised CDFL regularizations are designed, denoted as and
[0021]
[0022] Among them, the preprocessed data is respectively input into these two sub-networks, and the obtained feature representations are In order to maximize the difference, the features B extracted by h are mapped into different feature spaces, denoted as
[0023]
[0024] 5 As a preferred solution of the conflict-aware semi-supervised mutual learning medical image segmentation method described in the present invention, the following steps are included: The CCS regularization is used for mutual supervised learning in the regions where the predictions of the two networks are consistent. The CCS regularizations of subnets A and B are expressed as:
[0025]
[0026] Among them, I is a matrix of all 1s, Dice(·) represents the dice loss, and ⊙ represents element-wise operation at the pixel level. The region M where the prediction results of the two sub-networks are consistent can be identified by checking the consistency between. For the region with consistent predictions, we select the prediction of the sub-network with high confidence to supervise the learning of the sub-network with low confidence, and use the mask W to represent the region with high-confidence prediction. The probability outputs of sub-networks A and B can be obtained through the pixel-wise softmax function σ(·):
[0027]
[0028] The one-hot labels of the unlabeled data generated by sub-networks A and B can be expressed as:
[0029]
[0030] Therefore, the region M where the prediction results are consistent and the mask W of the region with high-confidence prediction can be obtained:
[0031]
[0032] 6 As a preferred solution of the semi-supervised mutual learning medical image segmentation method based on conflict awareness according to the present invention, wherein: the GMPS regularization is used for selective supervised learning of the regions where the predictions of the two networks are inconsistent. The GMPS regularization of sub-networks A and B is expressed as:
[0033]
[0034] Among them, for sub-networks A and B to extract features, the corresponding class prototypes are calculated for each segmentation category of the unlabeled image, which is expressed as:
[0035]
[0036] Here and represent the features of the penultimate layer of the two sub-networks. Then the consistency between the pixel and the y-th prototype is expressed as:
[0037]
[0038] Therefore, the mask G of the more reliable pseudo-labels from the two sub-networks A can be expressed as:
[0039]
[0040] 7 As a preferred solution of the conflict-aware semi-supervised mutual learning medical image segmentation method described in the present invention, the following steps are included: integrating supervised learning loss, consistency cross-supervision regularization, conflict-aware differential feature learning regularization, and geometric-aware mutual pseudo-supervision regularization into a semi-supervised medical image segmentation framework, and constructing objective functions for subnet A and subnet B respectively:
[0041]
[0042] Among them, for labeled image data, only its label is needed to supervise the corresponding prediction result. The supervised losses of subnet A and subnet B are expressed as follows:
[0043]
[0044] The unsupervised losses of subnet A and subnet B are expressed as follows:
[0045]
[0046] The parameters λ C and λ U are used to balance consistency cross-supervision regularization and conflict-aware differential feature learning regularization.
[0047] Finally, solve the objective functions of subnet A and subnet B respectively.
[0048] The beneficial effects of the present invention: introducing two sub-networks with different structures and different parameter initializations to prevent the same sub-network from extracting the same features, so as to fully exert the potential of the prediction diversity of the multi-sub-network architecture.
[0049] Design conflict-aware differential feature learning regularization (CDFL) to maximize the difference in features extracted by two sub-networks with different architectures, so as to avoid the homogenization of features extracted by different sub-networks.
[0050] For the regions where the predictions of the two networks are consistent, design consistency cross-supervision regularization (CCS) to select the regions where the predictions of the unlabeled images are consistent, and perform cross-supervision on the prediction results of the two networks for the pixels in this region, so that the two networks can learn from each other using their respective reliable predictions, and improve the generalization ability of the model.
[0051] For the regions where the predictions of the two networks are inconsistent, design geometric-aware mutual pseudo-supervision regularization (GMPS) to evaluate the reliability of the pseudo-labels in the regions where the predictions of the unlabeled images are inconsistent, and selectively use the more reliable pseudo-labels in the two sub-networks to supervise the training of the other self-network, preventing unreliable predictions from degrading the performance of the model.
[0052] The supervised learning loss, consistency cross-supervision regularization, conflict-aware differential feature learning regularization, and geometric-aware mutual pseudo-supervision regularization proposed by the present invention are integrated into a semi-supervised medical image segmentation framework. Objective functions are constructed and optimized for subnet A and subnet B respectively, ultimately improving the segmentation performance of the model for test medical images. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. Among them:
[0054] Figure 1 FIG. is the overall flowchart of a conflict-aware semi-supervised mutual learning medical image segmentation method provided in the first embodiment of the present invention;
[0055] Figure 2 FIG. is the framework diagram of a conflict-aware semi-supervised mutual learning medical image segmentation method provided in the first embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention with reference to the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0057] Embodiment 1
[0058] Refer to Figure 1 、 Figure 2 , which is an embodiment of the present invention, providing a conflict-aware semi-supervised mutual learning medical image segmentation method, including:
[0059] S1: Collect a set of medical image training datasets D containing N labeled data and M unlabeled data where M is much larger than N. Denote the i-th training image in D l . H, W, and C represent the length, width, and number of channels respectively. Denote its corresponding label. Y represents the number of its categories. Preprocess the training set data, such as resampling, cropping, and adjusting intensity values, etc.
[0060] S2: Construct two subnets A and B with different structures and different parameter initializations, denoted as f A (·) and f B (·) respectively, and initialize the weight parameters as θ A and θ B . Subnet f A contains a feature extractor h A (·): and a classifier g A (·): Initialize the weight parameters as and representing the dimension z of the feature space. There is For subnet B, there is Sub-networks h A (·) and h B (·), h A (·) adopts the U-Net structure, and h B (·) adopts the DeepLabv3+ structure.
[0061] S3: Design CDFL regularization to reduce the homogeneity of the features extracted by subnets A and B. For the labeled and unlabeled data input into subnets A and B, supervised and unsupervised CDFL regularizations are designed, denoted as and
[0062]
[0063] The features used here are the data preprocessed in step 1 input into these two sub-networks respectively, and the obtained feature representations are To maximize the difference, the features B extracted by h are mapped into different feature spaces, denoted as
[0064]
[0065]
[0066] S4: To make the best use of unlabeled images to assist model training, it is necessary to prevent the inaccurate pseudo-labels generated by the model from having an adverse impact on the model. Therefore, we use the difference between the predictions of subnets A and B to evaluate the accuracy of image predictions. For the regions with consistent predictions, consistency cross-supervision regularization (CCS) is designed for the two networks to supervise and learn from each other in the regions with consistent predictions. The consistency cross-supervision regularization (CCS) of subnets A and B can be expressed as:
[0067]
[0068] Here, I is a matrix of all 1s, Dice(·) represents the dice loss, and ⊙ represents pixel-wise element operations. The probability outputs of subnet A and subnet B can be achieved through the pixel-wise softmax function σ(·). Specifically as follows:
[0069]
[0070] The one-hot labels of the unlabeled data generated by subnet A and subnet B can be expressed as:
[0071]
[0072] For the unlabeled data, the region M where the prediction results of the two subnets are consistent can be identified by checking the consistency between. For the prediction-consistent region, we select the prediction of the subnet with a higher confidence to supervise the learning of the subnet with a lower confidence, and use the mask W to represent the region of the prediction with a higher confidence. Specifically as follows:
[0073]
[0074] S5: For the regions where the prediction results of subnet A and subnet B are inconsistent, it indicates that at least one of the outputs of the two networks for these regions is incorrect. Selecting a more reliable pseudo-label as the supervision signal for the other network helps to avoid the degradation of the model performance caused by unreliable predictions. Based on the pixel-based geometric structure, the intra-class semantic consistency is used to evaluate the reliability of the pseudo-label. The GMPS regularization of subnet A and subnet B can be expressed as:
[0075]
[0076] Among them, the features extracted by subnet A and subnet B are used to calculate the corresponding class prototypes for each segmentation category of the unlabeled image, which is expressed as:
[0077]
[0078] Here and represent the features of the penultimate layer of the two subnets. Then the consistency between the pixel and the y-th prototype is expressed as:
[0079]
[0080] Therefore, the mask G of the more reliable pseudo-labels from the two subnets A can be expressed as:
[0081]
[0082]
[0083] S6: Integrate the supervised learning loss, consistency cross-supervision regularization, conflict-aware differential feature learning regularization, and geometric-aware mutual pseudo-supervision regularization into the semi-supervised medical image segmentation framework, and construct the objective functions for subnet A and subnet B respectively:
[0084]
[0085] Among them, for the labeled data, only its label is needed to supervise the corresponding prediction result. The supervised losses of subnet A and subnet B are expressed as follows:
[0086]
[0087] The unsupervised losses of subnet A and subnet B are expressed as follows:
[0088]
[0089] The parameters λ C and λ U are used to balance the consistency cross-supervision regularization and the conflict-aware differential feature learning regularization. Finally, solve the objective functions of subnet A and subnet B respectively.
[0090] Implementation example:
[0091] An embodiment of the present invention provides a conflict-aware semi-supervised mutual learning medical image segmentation method. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through simulation experiments.
[0092] Evaluate on two publicly available medical image datasets:
[0093] Dataset 1: Automated Cardiac Diagnosis Challenge (ACDC);
[0094] Dataset 2: Synapse Multi-organ Segmentation Dataset (Synapse).
[0095] Five segmentation evaluation metrics are adopted: Dice Similarity Coefficient (DSC), Jaccard Index, Precision, 95% Hausdorff Distance (95HD), and True Positive Rate (TPR). Among them, the Dice coefficient is used to measure the similarity between two samples, the Jaccard index evaluates the overlap between the segmentation result and the ground truth label, Precision represents the accuracy of the segmentation result, 95HD is used to measure the surface distance between the segmentation result and the ground truth label, and the True Positive Rate reflects the ability of the model to identify true positive samples.
[0096] The segmentation results of this method and the comparison algorithms on Dataset 1 are shown in Table 1. Table 1 shows the segmentation performance of all semi-supervised comparison methods using 20% labeled data and 80% unlabeled data in the training set. In the fully supervised learning setting, U-Net and DeepLabv3+ are respectively trained on 100% labeled images, and the obtained segmentation results are considered as the upper bound performance of the model. The segmentation results of U-Net and DeepLabv3+ using only 20% labeled data are used as the baseline performance. In the semi-supervised learning setting, for the ACDC dataset, this method improves by 0.74%, 0.96 mm, 1.15%, 1.36%, and 0.03% respectively in the five metrics of Dice, 95HD, Jaccard, Precision, and TPR compared to the most competitive comparison algorithm MCF. The segmentation results of this method and the comparison algorithms on Dataset 2 are shown in Table 2. For the Synapse dataset, this method improves by 1.68%, 0.43 mm, 1.55%, 3.35%, and 0.16% respectively in the five metrics of Dice, 95HD, Jaccard, Precision, and TPR compared to the most competitive comparison algorithm MCF. The obvious improvement in all metrics verifies the effectiveness of the proposed framework.
[0097] Table 1 Comparison Results
[0098]
[0099] Table 2 Comparison Results
[0100]
[0101] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A conflict-aware semi-supervised mutual learning method for medical image segmentation, characterized in that: include: Collect medical image training data sets, divide them into labeled and unlabeled sets, and perform preprocessing; Initialize two sub-networks with different structures and extract features from the training data; Design a conflict-aware differentiated feature learning regularizer (CDFL) to maximize the difference in features extracted by two sub-networks with different architectures to avoid homogenization of features extracted by different sub-networks; Consistent cross-supervision regularization (CCS) and geometry-aware mutual pseudo-supervision regularization (GMPS) are designed for mutual supervision learning of the areas where the two networks have consistent predictions and the areas where the predictions are inconsistent, respectively. Construct the semi-supervised segmentation objective function for the two sub-networks respectively and solve them.
2. The conflict-aware semi-supervised mutual learning medical image segmentation method according to claim 1, characterized in that: The division of labeled and unlabeled sets includes N labeled data, which consists of N data and labels, expressed as: in, represents the labeled image data input for the i-th time, Represents the label corresponding to the first input labeled image data; H, W, C represent the length, width and number of channels respectively, and Y represents the number of categories. The unlabeled image consists of a small amount of M data, represented as: in, Represents the unlabeled image data of the i-th input. Perform preprocessing operations on the training set image data, such as resampling, cropping, and adjusting intensity values.
3. The conflict-aware semi-supervised mutual learning medical image segmentation method according to claim 2, characterized in that: Construct two subnets A and B with different structures and different parameter initialization, denoted as f A (·) and f B (·), the initialization weight parameters are denoted as θ A and θ B Subnet f A Contains a feature extractor and a classifier The initialization weight parameters are and represents the dimension z of the feature space. For subnet B, there are Subnetwork h A (·) and h B (·), h A (·)Using U-Net structure, h B (·) Using DeepLabv3+ structure.
4. The conflict-aware semi-supervised mutual learning medical image segmentation method according to claim 3, characterized in that: CDFL regularization is designed to reduce the homogeneity of features extracted from subnet A and subnet B. For the labeled and unlabeled data input into subnet A and subnet B, supervised and unsupervised CDFL regularizations are designed, denoted as and Among them, the preprocessed data are input into the two sub-networks respectively, and the obtained feature representations are In order to maximize the difference, h B (·)Extracted features Mapped to different feature spaces, respectively recorded as 5. The conflict-aware semi-supervised mutual learning medical image segmentation method according to claim 4, characterized in that: The consistent cross-supervision regularization (CCS) is designed for mutual supervision learning of regions where the two networks predict the same thing. The CCS regularization of subnet A and subnet B is expressed as: Where I is a matrix of all 1s, Dice(·) represents dice loss, and ⊙ represents the pixel-level element operation. The region M where the prediction results of the two sub-networks are consistent can be checked by For the predicted consistent areas, we select the prediction of the subnet with high confidence to supervise the learning of the subnet with low confidence, and use mask W to represent the predicted area with high confidence. The probability output of subnet A and subnet B can be obtained by the pixel-by-pixel softmax function σ(·): The one-hot labels of the unlabeled data generated by subnet A and subnet B can be expressed as: Therefore, we can obtain the area M with consistent prediction results and the mask W of the prediction area with high confidence:
6. The conflict-aware semi-supervised mutual learning medical image segmentation method according to claim 5, characterized in that: A geometry-aware mutual pseudo-supervision regularization (GMPS) is designed for selective supervision learning of areas where the prediction results of subnet A and subnet B are inconsistent. The GMPs regularization of subnet A and subnet B is expressed as: Among them, subnet A and subnet B are used to extract features to calculate the corresponding class prototype for each segmentation category of the unlabeled image, which is expressed as: Here and represents the features of the penultimate layer of the two subnets. Then the consistency of the pixel with the y-th prototype is expressed as: Therefore, the mask G of the more reliable pseudo-labels from the two subnets is A It can be expressed as:
7. The conflict-aware semi-supervised mutual learning medical image segmentation method according to claim 6, characterized in that: The supervised learning loss, consistency cross-supervision regularization, conflict-aware differentiated feature learning regularization, and geometry-aware mutual pseudo-supervision regularization are integrated into the semi-supervised medical image segmentation framework, and the objective functions are constructed for subnet A and subnet B respectively: Among them, for labeled image data, only its label is needed to supervise the corresponding prediction results. The supervised loss of subnet A and subnet B is expressed as follows: The supervised and unsupervised losses of subnet A and subnet B are expressed as follows: Here the parameter λ C and λ U It is used to balance the consistency cross-supervision regularization and the conflict-aware differentiated feature learning regularization. Finally, the objective functions of subnetwork A and subnetwork B are solved respectively.