Semi-supervised image segmentation method and system based on uncertainty perception collaborative learning model of mixed image
Through the uncertainty perception collaborative learning model of mixed images, combining image-level and model-level perturbations, the performance degradation problem under pseudo-label noise and multiple perturbations is solved, and a highly accurate semi-supervised medical image segmentation is achieved.
Patent Information
- Application Number
- CN202510540070.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-08
AI Technical Summary
In the existing semi-supervised medical image segmentation method based on pseudo labels, the pseudo label noise problem and performance degradation problems under multi-perturbation conditions have not been effectively solved, affecting the segmentation accuracy of the model.
Adopting an uncertainty-aware collaborative learning model based on hybrid images, a hybrid prototype-guided uncertainty estimation module and a collaborative learning strategy for hybrid images combines image-level and model-level perturbations to reduce pseudo-label noise interference and maintain training stability.
Improved the accuracy of image segmentation, especially in the case of unlabeled data, the Dice coefficient reached 90.44%, which is significantly better than other methods, demonstrating the stability and generalization ability of the model.
Smart Images

Figure CN120451178A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image segmentation, and in particular relates to a semi-supervised image segmentation method and system based on an uncertainty-aware collaborative learning model of mixed images. Background Art
[0002] The goal of medical image segmentation is to annotate the desired organs, tumors, and other features within an image. In recent years, the application of deep learning to medical image segmentation has made significant progress. Deep learning-based methods often rely on a large number of accurate labels during training, but labeling is time-consuming, labor-intensive, and costly. This ultimately leads to insufficient labeled data for training. Semi-supervised medical image segmentation, which leverages large amounts of unlabeled data to improve segmentation accuracy, has become an important research direction.
[0003] Pseudo-labeling is a mainstream approach for semi-supervised medical image segmentation. In these methods, pseudo-labels generated by a model for unlabeled data are used to guide training on unlabeled data. However, the pseudo-labels predicted by models trained on limited labeled data are unreliable and noisy. Mispredictions in pseudo-labels act as noise and can provide misleading supervision during training on unlabeled data.
[0004] To improve the performance of pseudo-label-based methods, perturbations are often introduced. Perturbations are intentional small modifications or interferences to the input data, model, or feature representations. These perturbations can occur at the model level, image level, or feature level. Previous methods used only one type of perturbation to ensure training stability, which limited their ability to handle larger amounts of unknown data. Using multiple perturbations is a straightforward approach to addressing this issue. However, when multiple perturbations are used in combination, the model is more likely to generate more incorrect predictions, resulting in a decrease in training quality.
[0005] In summary, the current pseudo-label-based methods still have shortcomings, which affects the segmentation accuracy of the model. Summary of the Invention
[0006] The first purpose of the present invention is to address the problems existing in the prior art and propose a semi-supervised image segmentation method and system based on an uncertainty-aware collaborative learning model of mixed images, aiming to solve two problems in the pseudo-label-based method: (1) the noise problem of pseudo-labels: the model is prone to make incorrect predictions when generating pseudo-labels, thereby misleading the training process of unlabeled data; (2) performance degradation under multiple perturbation conditions: although the use of multiple perturbation types can overcome the limitations of a single perturbation, it is still challenging to maintain training quality under diverse perturbations. In short, the present invention can reduce the interference of noise in pseudo-labels and maintain the stability of training while adding additional perturbations, thereby improving the accuracy of image segmentation.
[0007] In a first aspect, the present invention provides a semi-supervised image segmentation method based on an uncertainty-aware collaborative learning model of mixed images, the method comprising:
[0008] Get the image, including the labeled image X l and the corresponding label Y l , and the unlabeled image X u ;
[0009] Preprocess the images and construct the data set;
[0010] Build an uncertainty perception collaborative learning model based on mixed images and train it using the data set; the uncertainty perception collaborative learning model based on mixed images includes the first sub-model f(θ a ), the second sub-model f(θ b ), the third sub-model f(θ t ), and,a hybrid prototype-guided uncertainty estimation module,a collaborative learning strategy for mixed images;
[0011] Use the trained first sub-model f(θ a ) to achieve image segmentation.
[0012] Preferably, during the training process of the uncertainty-aware collaborative learning model based on mixed images:
[0013] The first sub-model f(θ a ), for the preprocessed labeled image X l and unlabeled image X u Process them separately to get the corresponding first prediction segmentation results and the second predicted segmentation result Second predicted segmentation result The first sub-pseudo-label can be obtained through the argmax function Then for the labeled image X l and unlabeled image X u The generated perturbation image X mix Processing is performed to obtain the corresponding third prediction segmentation result y mix ; Wherein, the disturbance image X mix is an unlabeled image X u The reduced image block is copied to the annotated image X l It is composed of random areas within the
[0014] The second sub-model f(θ b ), for the preprocessed labeled image X l and unlabeled image X u Make predictions respectively and get the corresponding fourth prediction segmentation results And the fifth predicted segmentation result Fifth predicted segmentation result The second sub-pseudo-label can be obtained through the argmax function
[0015] The third sub-model f(θ t ), is the first sub-model f(θ a )’s teacher model, the third sub-model f(θ t ) is obtained by passing the model parameters of the first sub-model f(θ a ) is updated in the manner of EMA, and the third sub-model f(θ t ) for the preprocessed unlabeled image X u Process it and then get the pseudo label through the argmax function
[0016] More preferably, the first sub-model f(θ a ), the second sub-model f(θ b ), the third sub-model f(θ t ) model architecture includes an encoder, a decoder, and an output layer; the encoding layer in the encoder and the same decoding layer in the decoder use skip connections.
[0017] Preferably, during the training process of the hybrid image-based uncertainty-aware collaborative learning model, the hybrid prototype-guided uncertainty estimation module is implemented as follows:
[0018] The labeled images X l and unlabeled image X u As input, sub-model f(θ τ1 ) The penultimate decoding layer of the decoder outputs the features of the labeled image and features of unlabeled images τ1 = a or b;
[0019] From the features of the labeled image Extract the annotation prototypes of the same category Through similarity loss The first similarity matching result Align with the mask;
[0020]
[0021]
[0022] Among them, L MSE is the mean square error loss, Cos(·) represents the cosine similarity calculation, and Up(·) represents the upsampling operation. is the mask, k∈[1,2,...K], K represents the number of categories;
[0023] Features from unlabeled images Extracting unlabeled prototypes Then the unlabeled prototype and annotated prototypes Perform weighted aggregation to obtain a hybrid prototype
[0024] Computational hybrid prototype Features of unlabeled images The cosine similarity between the two, the similarity result is upsampled and normalized along the category dimension by the softmax function φ to generate the third similarity matching result. The category of each pixel in the pseudo label is used as the index, from The value is taken as the loss weight w of the corresponding pixel τ1,j ;
[0025] According to the loss weight w τ1,j Calculate the first sub-model f(θ a ) and the second sub-model f(θ b ) 、
[0026] More preferably, the features from the unlabeled image Extracting unlabeled prototypes Specifically:
[0027]
[0028]
[0029]
[0030] Where, σ is the reliability threshold; It is a mask; is the second similarity matching result.
[0031] More preferably, the category of each pixel in the pseudo label is used as an index, from The specific value taken as the loss weight of the corresponding pixel is:
[0032]
[0033] in, Represents pseudo labels The pseudo label value of the j-th pixel in w τ1,j Represents the sub-model f(θ τ1) is the loss weight of the j-th pixel for pseudo supervision, j∈[1,N]; τ2=a or b, and τ1≠τ2.
[0034] More preferably, the pseudo-supervision loss The calculation of is as follows:
[0035]
[0036] in is the second predicted segmentation result The predicted probability that the j-th pixel corresponds to the k-th class, Fifth predicted segmentation result The predicted probability that the j-th pixel corresponds to the k-th class, is the pseudo label value of the kth category is the pseudo label value of the kth category
[0037] Preferably, during the training process of the uncertainty-aware collaborative learning model based on mixed images, the collaborative learning strategy of the mixed images includes:
[0038] According to the first sub-model f(θ a ) Generate the perturbation image X mix The pseudo label Zoom out and copy to label Y l Generate a perturbation image X mix Label Y mix ;
[0039] According to the first sub-model f(θ a ) The predicted third prediction segmentation result y mix , and label Y mix Compute the supervised loss for the perturbed image
[0040]
[0041] Among them L Dice is the Dice similarity calculation function;
[0042] Through the first sub-model f(θ a ) and the second sub-model f(θ b ) calculates the geometric contour information of the unperturbed image, and then uses the mean square error function to calculate the geometric contour information loss
[0043] By first predicting the segmentation result Fourth predicted segmentation result and label Y l , calculate the first sub-model f(θ a) in the supervised loss of the labeled images and the second sub-model f(θ b ) in the supervised loss of the labeled images ;
[0044]
[0045] Among them L Dice is the Dice loss, L CE is the cross entropy loss;
[0046] According to the monitoring loss Pseudo-supervision loss Similarity loss Supervision loss for perturbed images and SDM loss Total loss Through the total loss Back propagation, gradient update of the first sub-model f(θ a ) parameters;
[0047] According to the monitoring loss Pseudo-supervision loss Similarity loss and SDM loss Total loss Through the total loss Back propagation, gradient update of the second sub-model f(θ b ) parameters;
[0048] The third sub-model f(θ t )’s parameters are updated by EMA:
[0049]
[0050] in Indicates the third sub-model f(θ t ), Indicates the first sub-model f(θ a ), α represents the weight.
[0051] More preferably, the geometrical profile information SDM of the unperturbed image is calculated as follows:
[0052]
[0053] in represents the jth and qth pixels in the unlabeled image; B represents the boundary of the segmentation target, B in and Bo ut are the pixels inside and outside the segmented target respectively; Indicates the distance between different pixels; Represents the boundary B belonging to the segmentation target Take distance The minimum value of .
[0054] In a second aspect, the present invention provides a semi-supervised image segmentation system, comprising:
[0055] A data acquisition module is used to obtain image data with segmented targets;
[0056] Data preprocessing module, used to preprocess image data with segmentation targets;
[0057] The image segmentation module uses the trained first sub-model f(θ a ) Segment the preprocessed image data to obtain the image segmentation result.
[0058] The beneficial effects of the present invention are at least as follows:
[0059] 1. This paper proposes a pseudo-label uncertainty estimation method, which extracts prototypes from labeled and unlabeled images to estimate the reliability of pseudo-labels. It uses reliable features in labeled images to reduce the interference of noise in pseudo-labels, solving the problem that the noise in pseudo-labels can easily mislead model training.
[0060] 2. The present invention proposes a collaborative learning strategy for mixed images to enhance the model's ability to process unknown data by combining different perturbation methods. At the same time, the strategy strengthens the learning of unperturbed images to prevent unreliable prediction results produced by overfitting perturbed images, thereby improving the model's ability to process unknown data by adding perturbation methods, while maintaining training stability under multiple perturbations (image-level perturbations, model-level perturbations). BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 Schematic diagram of the network architecture of the uncertainty-aware collaborative learning model based on mixed images in the method of the present invention.
[0062] Figure 2 It is a schematic diagram of the network architecture of the hybrid prototype-guided uncertainty estimation module (MPUE) in the hybrid image-based uncertainty perception collaborative learning model of the present invention.
[0063] Figure 3 It is the first sub-model f(θ a ) network architecture diagram.
[0064] Figure 4 This is the effect diagram of segmentation using the 10% label image and the 20% label image by the method of the present invention. DETAILED DESCRIPTION
[0065] The present invention will be further explained below with reference to specific embodiments.
[0066] An embodiment of the present invention provides a semi-supervised image segmentation method based on an uncertainty-aware collaborative learning model for mixed images. The proposed mixed prototype-guided uncertainty estimation module (MPUE) can extract prototypes from labeled and unlabeled images to estimate the reliability of pseudo labels, effectively reducing the interference of noise in pseudo labels. The proposed collaborative learning strategy for mixed images (CLMI) introduces image-level perturbations to construct perturbed images to enhance the model's generalization ability for unknown data. In addition, CLMI also learns the SDM generated by unperturbed images, thereby avoiding the model's overfitting of unreliable predictions of perturbed images. In addition, experimental results show that in the 20% labeled image experiment of the ACDC dataset, the Dice coefficient of the present invention reached 90.44%; in the 10% labeled experiment, the Dice coefficient reached 89.97%, exceeding other comparison methods, proving its effectiveness and reliability.
[0067] The method is as Figure 1 Specifically included are the following:
[0068] Step 1: Obtain cardiac MR images from the Automatic Cardiac Diagnosis Challenge (ACDC) dataset and divide the dataset into two parts with labels Y l Annotated image data X l and unlabeled image data X u .
[0069] Step 2: Preprocess the MR images in the dataset. The preprocessing methods include 2D image cropping and normalization.
[0070] Step 3: Build Figure 1 The hybrid image-based uncertainty-aware collaborative learning model shown is trained using preprocessed MR image data.
[0071] The uncertainty perception collaborative learning model based on mixed images includes a first sub-model f(θ a ), the second sub-model f(θ b ), the third sub-model f(θ t ), a hybrid prototype-guided uncertainty estimation module and a collaborative learning strategy for hybrid images; the hybrid prototype-guided uncertainty estimation module uses hybrid prototypes extracted from labeled and unlabeled images to generate an uncertainty estimation map for pseudo labels; the uncertainty estimation map is used to reduce the impact of noise in pseudo labels during pseudo supervision. The collaborative learning strategy for hybrid images combines different perturbations to enhance learning on unperturbed images;
[0072] f(θ a ),f(θ b ) and f(θ t ) have the same network structure. In this embodiment, the U-Net model is used. a ) and f(θ b ) have different initialization weights, f(θ t ) is f(θ a )’s teacher model; with the first sub-model f(θ a ) as an example, its network architecture is shown in the attached Figure 3 , which uses the U-net network architecture and gives its output prediction segmentation results Predicted Signed Distance Map SDM a and image features
[0073] The hybrid image collaborative learning strategy can learn images more comprehensively from different image views (perturbed and non-perturbed views);
[0074] The hybrid prototype-guided uncertainty estimation module is used to estimate the reliability of the pseudo labels generated by the model, and to reduce the misleading effects of unreliable annotations in the pseudo labels on model training during pseudo supervision.
[0075] During the training process, the first sub-model f(θ a ) for the preprocessed labeled image X l and unlabeled image X u Process them separately to get the corresponding first prediction segmentation results and the second predicted segmentation result Second predicted segmentation result The first sub-pseudo-label can be obtained through the argmax function Then for the labeled image X l and unlabeled image X u The generated perturbation image X mix Processing is performed to obtain the corresponding third prediction segmentation result y mix ;
[0076] The perturbation image y mix The unlabeled image X u Zoom out to get an image block, and then copy the image block to the labeled image X l A random area within
[0077] The second sub-model f(θ b ) for the preprocessed labeled image X l and unlabeled image X u Make predictions respectively and get the corresponding fourth prediction segmentation results And the fifth predicted segmentation result Fifth predicted segmentation result The second sub-pseudo-label can be obtained through the argmax function
[0078] The third sub-model f(θ t ) is the first submodel f(θ a )’s teacher model, the third sub-model f(θ t ) is obtained by passing the model parameters of the first sub-model f(θ a ) is updated in the form of EMA (exponential moving average), and the third sub-model f(θ t ) for the preprocessed unlabeled image X u Process it and then get the pseudo label through the argmax function
[0079] The hybrid prototype guided uncertainty estimation module is as follows Figure 2 As shown, the implementation method is as follows:
[0080] The hybrid prototype-guided uncertainty estimation module can obtain a pixel-level uncertainty estimation weight map for cross-pseudo-supervision, reducing the misleading impact of unreliable pseudo-labels on model training. The acquisition process is outlined below:
[0081] 1) Take the labeled image X l and unlabeled image X u , the first submodel f(θ a ) The penultimate decoding layer of the decoder outputs the features of the labeled image and features of unlabeled images
[0082] The labeled images X l and unlabeled image X u , the second sub-model f(θ b ) The penultimate decoding layer of the decoder outputs the features of the labeled image and features of unlabeled images
[0083] 2) Annotated prototype extraction. In order to use the reliable features in the annotated image to guide the training of the unannotated image, the prototype is extracted from the features of the annotated image. The extraction process is shown in formula (1) and formula (2).
[0084]
[0085]
[0086] in For the mask, is the indicator function, k∈[1,2,...K], K represents the number of categories. Features of all annotated images of the same type The mean of τ1 = a or b.
[0087] From the characteristics Extract high-quality prototypes from the equation (4) using the similarity loss The similarity matching result of formula (3) With mask to align.
[0088]
[0089]
[0090] Among them, L MSE is the mean square error loss, Cos(·) represents the cosine similarity calculation, and Up(·) represents the upsampling operation, which is achieved through bilinear interpolation.
[0091] 3) Unlabeled prototype extraction. Considering that there may be differences in the prototypes between the labeled image and the unlabeled image, some prototypes extracted from the unlabeled image are aggregated with the labeled prototype to obtain a more reliable prototype. First, the labeled prototype is calculated. Features of unlabeled images The cosine similarity between them is then upsampled and scaled to the range [0,1] to obtain As shown in formula (5). Then, the unlabeled prototype The extraction process can be expressed by formula (6) and formula (7).
[0092]
[0093]
[0094]
[0095] Where σ is the reliability threshold, which is set to 0.8. The locations of high similarity features are marked. The calculation range of cosine similarity is [-1, 1]. Formula (5) scales the similarity results to set a suitable reliability threshold.
[0096] 4) Uncertainty estimation. and Weighted aggregation to form a hybrid prototype As shown in formula (8).
[0097]
[0098] The aggregation weight α is set to 0.9. Then, the hybrid prototype is calculated and features The cosine similarity between the two, the similarity results are normalized along the category dimension by the softmax function φ, and then upsampled to generate As shown in formula (9). The category of each pixel in the pseudo label is used as an index, from The value in is taken as the loss weight of the corresponding pixel, as shown in formula (10).
[0099]
[0100]
[0101] in, Represents pseudo labels The pseudo label value of the j-th pixel in w τ1,j Represents the sub-model f(θ τ1 ) The loss weight of the j-th pixel for pseudo supervision, j∈[1,N]; τ2=a or b, and τ1≠τ2;
[0102] According to formula (11), w a,j To calculate f(θ a )’s cross entropy loss Accordingly, w b,j To calculate f(θ b )’s cross entropy loss
[0103]
[0104] in is the second predicted segmentation result The predicted probability that the j-th pixel corresponds to the k-th class, Fifth predicted segmentation result The predicted probability that the j-th pixel corresponds to the k-th class, is the pseudo label value of the kth category is the pseudo label value of the kth category
[0105] The collaborative learning strategy for mixed images includes the following steps:
[0106] 1) There is model-level perturbation between sub-models. In order to combine multiple perturbation methods, an image-level perturbation Resizemix is added. Resizemix generates a perturbation image X by mixing the labeled image and the unlabeled image. mix .
[0107] First, the unlabeled image X u Reduce the image size to a quarter of the original image to obtain an image block. Then, copy the image block to the labeled image X l A random region within to create the perturbed image X mix . The perturbation image X mix Input to f(θ a ) to obtain the segmentation result y mix In order to give the segmentation result y mix Provide relatively reliable supervision signals, unlabeled image X u is input into the teacher model f(θ t ), and the unlabeled image X u The prediction results are converted into pseudo labels
[0108] Generate the perturbation image X mix The pseudo label Zoomed out is copied to label Y l To generate the perturbation image label Y mix The perturbed image’s label Y mix For perturbed image X mix The loss of the perturbed image is shown in formula (12).
[0109]
[0110] Among them L Dice is the Dice similarity calculation function;
[0111] 2) The predictions of the unperturbed images are relatively reliable. Therefore, strengthening the learning of the unperturbed images can reduce the interference of unreliable predictions generated by the perturbed images. Considering that geometric contour information can provide the global shape of each target category and thus avoid discontinuities in the segmentation results, the geometric contour information of the unperturbed images is also introduced in the training process. The geometric contour information can be represented as a signed distance map (SDM). The SDM is shown in formula (13).
[0112]
[0113] in represents the jth, qth pixel in the unlabeled image, and B represents the boundary of the segmentation target. Indicates the distance between different pixels. B in and B out are the pixels inside and outside the segmented object. Represents the boundary B belonging to the segmentation target Take distance The minimum value of f(θ a ) and f(θb ) also learn from each other for unlabeled images X u The predicted SDM; f(θ a ) and f(θ b ) learn from each other for unlabeled images X u The predicted SDM. The loss function is shown in formula (14).
[0114]
[0115] Among them, SDM a represents f(θ a ) predicted SDM, SDM b represents f(θ b ) predicted SDM.
[0116] The supervision loss is calculated as follows:
[0117] f(θ a ) in the supervised loss of the labeled images and f(θ b ) in the supervised loss of the labeled images As shown in formula (15). The supervision loss includes Dice loss L Dice and cross entropy loss L CE .
[0118]
[0119] Total loss Loss due to supervision Pseudo-supervision loss Similarity loss Supervision loss for perturbed images and SDM loss Composition. Through total loss Back propagation, gradient update sub-model f(θ a ) parameters. Total loss Excluding the supervised loss of perturbed images The other parts are related to f(θ a ) has the same total loss. f(θ b ) total loss Used to update the sub-model f(θ b ) parameters. The total loss is shown in formula (16).
[0120]
[0121] Here, β represents the loss weight and is set to a fixed value of 0.3.
[0122] Submodel f(θ t) is updated through EMA, as shown in formula (17).
[0123]
[0124] in Indicates the sub-model f(θ t ), Indicates the sub-model f(θ a ), α represents the weight.
[0125] Step 4: Use the first sub-model f(θ trained in step 3 a ) to achieve medical image segmentation.
[0126] To verify the effectiveness of the proposed segmentation method, we used four commonly used evaluation metrics in medical image segmentation: the Dice score (Dice), the Jaccard similarity coefficient (Jaccard), the 95% Hausdorff distance (95HD), and the average surface distance (ASD). The expressions for Dice, Jaccard, 95HD, and ASD are defined as shown in formulas (18) to (21):
[0127]
[0128] HD(X,Y)=m a x{d YX ,d YX} (19)
[0129]
[0130]
[0131] in,
[0132] In order to verify the effectiveness of the present invention, the method of the present invention is compared with a variety of existing semi-supervised medical image segmentation methods, such as Mean-Teacher (MT), UA-MT, CPS, MC-Net, SS-Net, BCP, SAMT-PCL and ABD. The experiment is trained and tested based on the ACDC public dataset.
[0133] Table 1 Results of various methods and full supervision when the labeled images account for 10% and 20% of the ACDC dataset
[0134]
[0135] Results are presented as mean ± SD. *Indicates p ≤ 0.05 based on the Wilcoxon signed-rank test for pairwise comparisons with our method. The best mean result is shown in bold.
[0136] The experimental results are shown in Table 1. In order to compare with the semi-supervised method, U-Net adopts a fully supervised form. Compared with the fully supervised U-Net, the present invention improves the Dice coefficient from 78.77% to 89.97% with only 10% labeled data. When the labeled data increases to 20%, the Dice coefficient of our method increases to 90.44%. Compared with other semi-supervised methods, the present invention achieves better performance and shows better stability. The visualization of the segmentation results on the ACDC dataset is shown in Figure 1. Figure 4 As shown, the present invention can generate a complete segmentation boundary. For particularly small segmentation targets ( Figure 4 The first row in , also has strong perception and segmentation capabilities.
[0137] To validate the effectiveness of the collaborative learning strategy for mixed images (CLMI) and the hybrid prototype-guided uncertainty estimation module (MPUE), we conducted an ablation study. Table 2 shows the ablation study results on the ACDC dataset. The baseline is the basic CPS. When the proposed CLMI is combined with the baseline, the resulting Dice score improves by 2.54% over the baseline. CLMI extends the basic CPS, allowing the model to learn perturbed images to improve the model's generalization ability while fully exploring the information in the unperturbed images. Introducing MPUE to the baseline significantly improves the model's accuracy, with a 2.05mm reduction in HD95 and a 0.57% increase in the Dice score. MPUE estimates the uncertainty of pseudo-labels, reducing the impact of noise in the pseudo-labels. When MPUE and CLMI are added to the baseline simultaneously, the model's Dice score improves by 3.46% over the baseline. These results demonstrate that the proposed method can effectively improve model performance.
[0138] Table 2 Ablation experiment results, the labeled data accounts for 20%
[0139]
[0140] This embodiment also provides a semi-supervised image segmentation system, including:
[0141] A data acquisition module is used to obtain image data with segmented targets;
[0142] Data preprocessing module, used to preprocess image data with segmentation targets;
[0143] The image segmentation module uses the trained first sub-model f(θ a) Segment the preprocessed image data to obtain the image segmentation result.
[0144] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A semi-supervised image segmentation method based on an uncertainty-aware collaborative learning model for mixed images, characterized in that: The method comprises: Get the image, including the labeled image X l and the corresponding label Y l , and the unlabeled image X u ; Preprocess the images and construct the data set; Build an uncertainty perception collaborative learning model based on mixed images and train it using the data set; the uncertainty perception collaborative learning model based on mixed images includes the first sub-model f(θ a ), the second sub-model f(θ b ), the third sub-model f(θ t ), and,a hybrid prototype-guided uncertainty estimation module,a collaborative learning strategy for mixed images; Use the trained first sub-model f(θ a ) to achieve image segmentation.
2. The method according to claim 1, characterized in that During the training process of the uncertainty-aware collaborative learning model based on mixed images: The first sub-model f(θ a ), for the preprocessed labeled image X l and unlabeled image X u Process them separately to get the corresponding first prediction segmentation results and the second predicted segmentation result Second predicted segmentation result The first sub-pseudo-label can be obtained through the argmax function Then for the labeled image X l and unlabeled image X u The generated perturbation image X mix Processing is performed to obtain the corresponding third prediction segmentation result y mix ; Wherein, the disturbance image X mix is an unlabeled image X u The reduced image block is copied to the annotated image X l It is composed of random areas within the The second sub-model f(θ b ), for the preprocessed labeled image X l and unlabeled image X u Make predictions respectively and get the corresponding fourth prediction segmentation results And the fifth predicted segmentation result Fifth predicted segmentation result The second sub-pseudo-label can be obtained through the argmax function The third sub-model f(θ t ), is the first sub-model f(θ a )’s teacher model, the third sub-model f(θ t ) is obtained by passing the model parameters of the first sub-model f(θ a ) is updated in the manner of EMA, and the third sub-model f(θ t ) for the preprocessed unlabeled image X u Process it and then get the pseudo label through the argmax function 3. The method according to claim 2, characterized in that The first sub-model f(θ a ), the second sub-model f(θ b ), the third sub-model f(θ t ) model architecture includes an encoder, a decoder, and an output layer; the encoding layer in the encoder and the same decoding layer in the decoder use skip connections.
4. The method according to claim 3, characterized in that During the training process of the hybrid image-based uncertainty perception collaborative learning model, the implementation method of the hybrid prototype-guided uncertainty estimation module is as follows: The labeled images X l and unlabeled image X u As input, sub-model f(θ τ1 ) The penultimate decoding layer of the decoder outputs the features of the labeled image and features of unlabeled images τ1 = a or b; From the features of the labeled image Extract the annotation prototypes of the same category Through similarity loss The first similarity matching result Align with the mask; Among them, L MSE is the mean square error loss, Cos(·) represents the cosine similarity calculation, and Up(·) represents the upsampling operation. is the mask, k∈[1,2,...K], K represents the number of categories; Features from unlabeled images Extracting unlabeled prototypes Then the unlabeled prototype and annotated prototypes Perform weighted aggregation to obtain a hybrid prototype Computational hybrid prototype Features of unlabeled images The cosine similarity between the two, the similarity result is upsampled and normalized along the category dimension by the softmax function φ to generate the third similarity matching result. The category of each pixel in the pseudo label is used as the index, from The value is taken as the loss weight w of the corresponding pixel τ1,j ; According to the loss weight w τ1,j Calculate the first sub-model f(θ a ) and the second sub-model f(θ b ) 5. The method according to claim 4, characterized in that: The features from unlabeled images Extracting unlabeled prototypes Specifically: Where, σ is the reliability threshold; It is a mask; is the second similarity matching result.
6. The method according to claim 4, characterized in that: The category of each pixel in the pseudo label is used as an index, from The specific value taken as the loss weight of the corresponding pixel is: in, Represents pseudo labels The pseudo label value of the j-th pixel in w τ1,j Represents the sub-model f(θ τ1 ) is the loss weight of the j-th pixel for pseudo supervision, j∈[1,N]; τ2=a or b, and τ1≠τ2.
7. The method according to claim 4, characterized in that: The pseudo-supervisory loss The calculation of is as follows: in is the second predicted segmentation result The predicted probability that the j-th pixel corresponds to the k-th class, Fifth predicted segmentation result The predicted probability that the j-th pixel corresponds to the k-th class, is the pseudo label value of the kth category is the pseudo label value of the kth category 8. The method according to claim 3, characterized in that: During the training process of the uncertainty-aware collaborative learning model based on mixed images, the collaborative learning strategy of mixed images includes: According to the first sub-model f(θ a ) Generate the perturbation image X mix The pseudo label Zoom out and copy to label Y l Generate a perturbed image X mix Label Y mix ; According to the first sub-model f(θ a ) The predicted third prediction segmentation result y mix , and label Y mix Compute the supervised loss for the perturbed image Among them L Dice is the Dice similarity calculation function; Through the first sub-model f(θ a ) and the second sub-model f(θ b ) calculates the geometric contour information of the unperturbed image, and then uses the mean square error function to calculate the geometric contour information loss By first predicting the segmentation result Fourth predicted segmentation result and label Y l , calculate the first sub-model f(θ a ) in the supervised loss of the labeled images and the second sub-model f(θ b ) in the supervised loss of the labeled images Among them L Dice is the Dice loss, L CE is the cross entropy loss; According to the monitoring loss Pseudo-supervision loss Similarity loss Supervision loss for perturbed images and SDM loss Total loss Through the total loss Back propagation, gradient update of the first sub-model f(θ a ) parameters; According to the monitoring loss Pseudo-supervision loss Similarity loss and SDM loss Total loss Through the total loss Back propagation, gradient update of the second sub-model f(θ b ) parameters; The third sub-model f(θ t )’s parameters are updated by EMA: in Indicates the third sub-model f(θ t ), Indicates the first sub-model f(θ a ), α represents the weight.
9. The method according to claim 8, characterized in that The geometric contour information SDM of the unperturbed image is calculated as follows: in represents the jth and qth pixels in the unlabeled image; B represents the boundary of the segmentation target, B in and B out are the pixels inside and outside the segmented target respectively; Indicates the distance between different pixels; Represents the boundary B belonging to the segmentation target Take distance The minimum value of .
10. A semi-supervised image segmentation system implementing the method according to any one of claims 1 to 9, characterized in that: include: A data acquisition module is used to obtain image data with segmented targets; Data preprocessing module, used to preprocess image data with segmentation targets; The image segmentation module uses the trained first sub-model f(θ a ) Segment the preprocessed image data to obtain the image segmentation result.