Semi-supervised pancreatic image segmentation method based on prototype estimation and prototype consistency

By introducing prototype estimation and consistency regularization loss, the problem of insufficient utilization of labeled data is solved. The optimization of feature representation is improved, and the accuracy of pancreatic segmentation and the prediction ability of fuzzy areas of CT images are achieved, and a more efficient pancreatic segmentation effect is achieved.

CN120451187APending Publication Date: 2025-08-08AFFILIATED HOSPITAL OF JIANGNAN UNIV +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510592157.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing semi-supervised learning methods fail to make full use of labelless data in CT images pancreatic segmentation, especially it is difficult to accurately predict fuzzy areas such as the pancreas edge, and insufficient optimization of feature representations, resulting in unsatisfactory segmentation effect.

Method used

Using a semi-supervised segmentation method based on prototype estimation and prototype consistency, by introducing a projection head as a prototype branch, prototype learning and consistency regularization losses are designed, and fuzzy areas and determined areas without labeling data are used for differentiated supervision, and feature representation is optimized.

Benefits of technology

It significantly improves the utilization rate and segmentation accuracy of labelless data, and improves the segmentation performance of the pancreas in abdominal CT images, especially the prediction accuracy in fuzzy areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451187A_ABST
    Figure CN120451187A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image segmentation, and relates to a semi-supervised pancreatic image segmentation method based on prototype estimation and prototype consistency. According to the method, based on a semi-supervised segmentation framework of an average teacher, V-Net is adopted as a deep learning segmentation model, and a projection head is inserted in the last but one stage of a decoder as a prototype branch; prototype learning design prototype estimation is introduced, unmarked data prediction is divided into determined areas of fuzzy areas, and different losses are designed for areas with different reliability of the unmarked data. The problems that in an existing semi-supervised learning segmentation method, unmarked data cannot be fully utilized, potential information of marked data is insufficient in utilization and the like are effectively solved, and a higher-performance solution is provided for an abdominal CT image pancreatic organ segmentation task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image segmentation and relates to a semi-supervised pancreatic image segmentation method based on prototype estimation and prototype consistency. Technical Background

[0002] Pancreatic cancer ranks fourth among the causes of cancer deaths worldwide. It is characterized by insidious onset, early metastasis, rapid progression, poor treatment efficacy and prognosis, and is the most malignant of digestive tract tumors. Pancreatic cancer lacks specific symptoms in its early stages, and patients are usually diagnosed in the middle or late stages, with most already metastasizing. The prognosis for pancreatic cancer is extremely poor, with an overall 5-year survival rate of only 8.2%. Therefore, it is necessary to develop reliable and universal medical methods to achieve early screening, such as identifying high-risk groups for pancreatic cancer, discovering tumor markers, and using imaging technologies such as CT to improve its detection rate. Studies have shown that existing medical image analysis methods can provide at least six months of early warning of pancreatic cancer, and medical image analysis methods based on pancreatic organs are inseparable from accurate pancreatic organ segmentation.

[0003] Computed tomography (CT) has a higher density resolution and is often the preferred method for computer-assisted diagnosis of pancreatic diseases. Therefore, research on automated pancreatic segmentation methods for CT images is of great clinical significance. In recent years, deep learning-based medical image segmentation methods have gained widespread application in the field of automated pancreas segmentation due to their excellent performance. However, these data-driven deep learning methods typically require fully supervised learning, and model training relies on large amounts of finely labeled data. In actual clinical applications, collecting and delineating three-dimensional medical images is not only time-consuming, labor-intensive, and costly, but also requires extensive expertise and experience from physicians. Consequently, high-quality annotated data is often difficult to obtain. Furthermore, due to issues such as variability in annotations between physicians and imaging equipment, data sharing and annotation between different hospitals is extremely difficult. As discussed above, addressing the reliance on large amounts of finely labeled data in existing deep learning-based automated pancreas segmentation methods has become an urgent issue.

[0004] In real-world clinical applications, unlabeled data is readily available and abundant, while finely labeled data is relatively scarce. Semi-supervised learning, because it can combine a small amount of labeled data with a large amount of unlabeled data for model training, thereby alleviating the need for labeled data, has become a research hotspot in recent years to address the challenges of deep learning in clinical applications. Existing semi-supervised learning methods can be broadly categorized into pseudo-label self-training methods and co-training methods with consistency regularization.

[0005] Although current semi-supervised learning methods have made great progress in the field of automated pancreas segmentation in CT images, they still face many challenges, especially in fully utilizing the model's unlabeled data predictions. Specifically, semi-supervised learning segmentation models usually use uncertainty estimation methods to obtain the uncertainty of unlabeled data predictions, and then use threshold filtering or weighting methods to screen out stable predictions for the calculation of the model's unlabeled data supervision signal. However, this type of method only uses reliable or stable predictions for loss calculation of unlabeled data, which has a potential problem, that is, some voxels in the CT image may never be used during the entire training process. For example, if only the reliable parts predicted by the model are used as supervision signals for unlabeled data, some parts that are difficult to accurately predict (such as the edge of the pancreas, the branches of the pancreas, and the blurred parts of the CT image) will be discarded during the training process. The features of these parts happen to be at the decision boundary. Ignoring these areas during training may limit the final performance of the model.

[0006] Furthermore, CT image segmentation is inherently an intensive classification task, and the discriminative nature of segmentation features plays a crucial role in segmentation performance. For example, in the pancreatic organ segmentation task, segmentation targets often have different shapes, sizes, and morphologies, and many samples exhibit significant intra-class inconsistency and inter-class ambiguity. By modeling contextual relationships to enhance the model's feature representation, and by increasing both intra-class aggregation and inter-class separation of the model's feature representation, the model's segmentation performance can be significantly improved, particularly for single-organ segmentation tasks in abdominal CT images.

[0007] In summary, existing semi-supervised learning methods offer effective solutions for automatic pancreatic segmentation. However, existing models do not fully utilize the model's unlabeled data predictions. Commonly used uncertainty estimation filters out low-quality pseudo-labels, resulting in the discarding of features that are difficult to accurately predict but fall squarely on the decision boundary. Furthermore, existing semi-supervised segmentation models have limited research on optimizing feature representation within the model's semantic space, resulting in suboptimal segmentation results for objects with highly variable scale and shape, such as the pancreas, in abdominal CT images. Summary of the Invention

[0008] In view of the above shortcomings of the existing technology, the present invention provides a semi-supervised pancreatic image segmentation method based on prototype estimation and prototype consistency. This method is based on the semi-supervised segmentation framework of the average teacher, adopts V-Net as the deep learning segmentation model, and inserts a projection head (Projection Head) as the prototype branch in the penultimate stage of the decoder; prototype learning is introduced to design prototype estimation, which is used to divide the unlabeled data prediction into a definite area of the fuzzy area, and different losses are designed for areas with different reliability of the unlabeled data. This method aims to make full use of the unlabeled data prediction and optimize the model feature representation to improve the accuracy of pancreatic segmentation in abdominal CT images under semi-supervised conditions.

[0009] The semi-supervised pancreatic image segmentation method based on prototype estimation and prototype consistency includes the following steps:

[0010] Step 1: Obtain and preprocess the pancreatic CT dataset.

[0011] Step 2: Divide the dataset preprocessed in step 1 into a training set and a test set, where the training set is Among them, Represents labeled data and its true label, represents unlabeled data, x l(i) and x u(i) Represent the labeled CT image and the unlabeled CT image, y l(i) Represents the corresponding true label, N and M represent the number of labeled data and unlabeled data respectively, and M>>N.

[0012] Step 3: Based on the V-Net segmentation model, a mean teacher (MT) semi-supervised segmentation framework is constructed. The student model is a V-Net model, and the features output by the penultimate stage are used as the prototype branch. Its parameters are updated and optimized during model training. The teacher model uses the same architecture as the student model, and its parameters are obtained by exponential moving average from the student model.

[0013] Step 4: Train on the training set D, with each batch containing two labeled and two unlabeled data, i.e., a batch size of 4. Before inputting the data into the model, the data undergoes a random cropping strategy, cropping them into image blocks of size [96,96,96], and applying image augmentation techniques such as random flipping and random rotation. The labeled data is fed into the student model, using a standard segmentation loss as the supervisory signal, and class prototypes for the labeled data are generated based on the true labels. Unlabeled data is fed into the student model and the teacher model, respectively, to obtain corresponding predicted probabilities and pseudo-labels. The subsequent loss calculation process is as follows: First, Monte Carlo sampling is performed, where the unlabeled data is inferred eight times by the teacher model and the average is taken to obtain voxel-level stable features and an uncertainty weight map based on information entropy. Then, guided by the teacher model's pseudo-labels, the class prototypes of the unlabeled data and the Euclidean distance between the voxel features and the class prototypes are calculated. Based on the prototype estimation method, the student model output is divided into a certain region and a fuzzy region. Finally, the consistency regularization loss is calculated between the student model output and the teacher model output (with added noise) in the fuzzy region of the unlabeled data. In the certain region, the prototype consistency loss is calculated between the student model's predicted probability based on the fused prototype and the student model's pseudo-label, weighted by uncertainty. The prototype consistency loss is calculated by fusing the labeled and unlabeled data prototypes. The overall model loss is weighted by the labeled and unlabeled data losses.

[0014] Step 5: Use the final trained student model for pancreatic organ segmentation.

[0015] Specifically, step 1 is as follows: first, all CT sequences are centrally cropped around the pancreatic target (ROI), and 25 voxels at the edge are enlarged on all spatial axes to obtain a three-dimensional graphic block centered on the pancreatic target; then, the Hounsfield units (HU) in the CT sequences are rescaled, and a soft tissue window is used. The window width and window level of enhanced CT and plain CT are 75 and 400, and 60 and 360, respectively, that is, the grayscale value ranges are [-125, 275] and [-120, 240], respectively; finally, all sequences are resampled to an isotropic resolution of 1.0 mm × 1.0 mm × 1.0 mm and normalized to zero mean and unit variance before input into the network.

[0016] The step 4 is specifically as follows:

[0017] The specific process of labeled data loss calculation and class prototype generation is as follows:

[0018] Labeled data is fed into the student model, and the standard segmentation loss between the model prediction and the true label is calculated, which is a combination of cross entropy loss and Dice loss:

[0019]

[0020] Among them, among them, represents the cross entropy loss, Denotes Dice loss. p,y denote the predicted probability and true label of the student model, respectively. In addition, the output of the student model projection head is upsampled by trilinear interpolation to obtain the segmentation mask Feature maps of consistent size Represents the features of each voxel. Then, the feature map is directly average-pooled using the true label to calculate the prototype of the corresponding class. The calculation formula is as follows:

[0021]

[0022] Among them, I(·) represents the indicator function, y represents the true label of the labeled data, and f l,v Represents the voxel features of the labeled data obtained by the student model inference. In addition, the prototypes of the labeled data are processed by the EMA strategy to obtain stable prototypes.

[0023] The unlabeled data is input into the teacher model and the student model respectively. The specific processing flow of the designed prototype estimation method is as follows:

[0024] Unlabeled data is input into the teacher model, and after Monte Carlo sampling, voxel-level stable features and uncertainty weight maps based on information entropy are obtained, which are formulated as follows:

[0025]

[0026] Among them, c∈{bg,obj}, represents the probability that the teacher model predicts class c, T represents the number of Monte Carlo sampling, which is set to 8; H v represents the cross entropy of the predicted probability mean after T Monte Carlo Dropout when unlabeled data is input into the teacher model; Φ v is the reliability weight map of each pixel after normalization.

[0027] Projection head output of the teacher model (K represents the number of channels) is upsampled by trilinear interpolation to the same value as the segmentation mask The size is consistent, and then the voxel-level feature map is obtained after Monte Carlo sampling and averaging. Represents the features of each pixel output by the model. Then, the one-hot encoding output of the teacher model The class prototype of the unlabeled data is obtained by guiding (obtained by argmax) and weighting through the reliability weight map. The formula is as follows:

[0028]

[0029]

[0030] Where I(·) is the indicator function. By weighting the reliability weight graph, the model can generate a more stable prototype for unlabeled data. As a result, a fusion prototype can be obtained. The specific calculation formula is as follows:

[0031]

[0032] Where c∈{bg,obj},λ lab =1-λ con ,λ unlab =λ con , and λ con is a widely used time-dependent Gaussian warm-up function. During training, the proportion of labeled prototypes decreases from 1 to 0, while the proportion of unlabeled prototypes increases from 0 to 1.

[0033] Then calculate the voxel feature vector f by the following formula t,v Prototypes with unlabeled data The corresponding distance between and

[0034]

[0035] When the pseudo label of the vth voxel from the output of the student model is the target class, and it comes from the feature f of the teacher model t,v If the voxel is closer to the background prototype than the target prototype, then the prediction of the voxel is suspected to be an incorrect prediction voxel, that is, a fuzzy area. Otherwise, it is a confirmed area. Similarly, if the pseudo label of a voxel is background, and its feature f t,v If it is closer to the target prototype, then it is a fuzzy area, otherwise it is a definite area. Fuzzy area mask (Mask) E ambiguity And determine the region mask (Mask) E certain It can be obtained by the following formula:

[0036]

[0037] E certain =1-E ambiguity

[0038] After the above prototype estimation method divides the model's unlabeled data prediction into fuzzy areas and certain areas, the specific loss calculation process for each area is as follows:

[0039] For the fuzzy areas in the student model output, a simple consistency regularization (CR) is used to supervise the training of the model, which requires that the output probability of the student model is consistent with that of the teacher model:

[0040]

[0041] Among them, p t ,p s are the probability outputs of the teacher model and the student model, E ambiguity is the mask of the blurred area, and MSE represents the mean square error.

[0042] For a certain region in the output of the student model, the feature-prototype similarity is used to approximate the probability of each voxel class. u Input the projection head output of the student model and upsample to obtain the voxel-level feature map Calculate the similarity between each voxel feature and the corresponding class (target or background), and then convert it into class prediction probability. Its formula is as follows:

[0043]

[0044] Among them, c∈{bg,obj} are the prototypes of the background class and the target class respectively, sim(·) represents the cosine similarity, and p c Represents the fused prototype of class c. The similarity between the feature vector of v voxel and the target and background classes is converted into a probabilistic representation. The introduction of fused prototypes dynamically adjusts the ratio of labeled to unlabeled data prototypes. Initially, training relies on prototypes from labeled data to ensure stability, while later, it gradually relies on prototypes from unlabeled data to improve generalization.

[0045] The key idea of prototype-based consistency regularization is to encourage consistency between the model's segmentation predictions and prototype-based predictions. This pattern implicitly normalizes feature representation: features from the same class must be close to the fused prototype of that class, while being far away from prototypes of other classes. The specific prototype consistency (PC) loss function is calculated as follows:

[0046]

[0047] in, They represent the predicted probability of the student model based on the prototype and the segmentation prediction pseudo label, E certain Indicates the mask of the determined area, Φ v represents the voxel-by-voxel uncertainty weight.

[0048] To sum up, the overall loss of unlabeled data can be obtained:

[0049]

[0050] in, Represents the consistency regularization loss of the fuzzy area, that is, MSE loss, p s ,p t represent the predicted probabilities of the student model and the teacher model respectively. represents the prototype consistency loss of the determined region. is a Gaussian prediction function that changes over time and increases with the number of iterations t. max is set to 0.1, and λ pc is a similar function, where w max Set to 1. That is, as the number of training rounds increases, the consistency regularization loss weight of the fuzzy area and the prototype consistency loss weight of the definite area increase from 0 to 0.1 and 1 respectively.

[0051] The step 5 is specifically as follows: adopting a sliding window strategy based on image blocks, with an image block size of [96,96,96] and a step size of 16×16×16, and synthesizing the predictions of all image blocks to obtain the final output.

[0052] Compared with the prior art, the advantages and beneficial effects of the present invention are:

[0053] (1) Fully utilize unlabeled data. The predictions obtained after the unlabeled data is input into the model are divided into certain areas and fuzzy areas through prototype estimation. This method designs targeted different supervisory signals for areas with different reliability, thereby fully utilizing the unlabeled data output of the model.

[0054] (2) Efficient collaboration between labeled and unlabeled data. In the calculation process of the loss function for determining the region of unlabeled data prediction, a fusion prototype is used to generate the model's prototype-based prediction. In the early stages of training, the labeled data prototype accounts for a large proportion of the fusion prototype, which is used to guide the model's segmentation on unlabeled data. In the later stages of training, the model's segmentation ability is greatly improved, and the generated unlabeled data prototype is more accurate. Therefore, the unlabeled data prototype accounts for a large proportion of the fusion prototype.

[0055] (3) Through sufficient experiments, it is proved that under two different semi-supervised settings, the proposed semi-supervised pancreas segmentation PECMT method has significantly improved the segmentation accuracy on the CT image pancreas dataset compared with existing related methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 2 is an overall framework diagram of a semi-supervised pancreatic image segmentation method based on prototype estimation and prototype consistency in an embodiment of the present invention;

[0057] Figure 2 Schematic diagram of prototype estimation in an embodiment of the present invention. DETAILED DESCRIPTION

[0058] In order to more clearly describe the technical solution of the present invention, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings. The specific embodiments described here are only used to explain and illustrate the present invention and are not used to limit the present invention.

[0059] A semi-supervised pancreatic image segmentation method based on prototype estimation and prototype consistency, PECMT. This method introduces prototype learning into the semi-supervised pancreatic segmentation task, and designs uncertainty estimation and consistency learning from the prototype perspective. This method is based on the average teacher framework, uses V-Net as the basic segmentation model, and inserts a projection head (Projection Head) as the prototype branch in the penultimate stage of the decoder, and designs a three-part loss to jointly train the model. Labeled data is only input into the student model, and the standard segmentation loss is used, defined as L sup The unlabeled data are input into the teacher model and the student model respectively, where the teacher model is an auxiliary model and the model parameters are obtained from the student model parameters through EMA. The prediction of the student model is divided into fuzzy area and definite area by the prototype estimation method, and the consistency regularization loss L is designed respectively. con and prototype consistency loss L pc Specific examples include: Figure 1 and Figure 2 As shown, it includes the following steps, and they are performed in sequence:

[0060] Step 1: Obtain abdominal CT images and preprocess them, specifically:

[0061] First, all CT sequences were centrally cropped around the pancreatic target (ROI), and 25 voxels at the edge were enlarged on all spatial axes to obtain a three-dimensional graphic block centered on the pancreatic target. Then, to enhance the contrast of soft tissue in the images, the Hounsfield units (HU) in the CT sequences were rescaled, and a soft tissue window was used. The window width and window position of enhanced CT and plain CT were 75 and 400, and 60 and 360, respectively, that is, the grayscale value range was [-125, 275] and [-120, 240], respectively. Finally, all sequences were resampled to an isotropic resolution of 1.0 mm × 1.0 mm × 1.0 mm and normalized to zero mean and unit variance before input into the network.

[0062] Step 2: Divide the dataset preprocessed in step 1 into a training set and a test set, where the training set is Among them, Represents labeled data and its true label, represents unlabeled data, x l(i) and x u(i) Represent the labeled CT image and the unlabeled CT image, y l(i) Represents the corresponding true label, N and M represent the number of labeled data and unlabeled data respectively, and M>>N.

[0063] In step 3, we built a semi-supervised segmentation framework using the Average Teacher (MT) algorithm based on the V-Net segmentation model. The student model is a V-Net model, and the features output by the penultimate stage are used as prototype branches. The parameters of the MT model are updated and optimized during model training. The teacher model uses the same architecture as the student model, and its parameters are derived from the student model using an exponential moving average (EMA).

[0064] Step 4: Train on the training set D, with each batch containing 2 labeled data and 2 unlabeled data, i.e., a batch size of 4. Before entering the model, the data undergoes a random cropping strategy, cropping it into image blocks of size [96,96,96], and applying image augmentation techniques such as random flipping and random rotation. The labeled data is input to the student model, and a standard segmentation loss is used as the supervisory signal, which is a combination of cross entropy loss and Dice loss:

[0065]

[0066] Among them, among them, represents the cross entropy loss, Denotes Dice loss. p,y denote the predicted probability and true label of the student model, respectively. In addition, the output of the student model projection head is upsampled by trilinear interpolation to obtain the segmentation mask Feature maps of consistent size Represents the features of each voxel. Then, the feature map is directly average-pooled using the true label to calculate the prototype of the corresponding class. The calculation formula is as follows:

[0067]

[0068] Among them, I(·) represents the indicator function, y represents the true label of the labeled data, and f l,v Represents the voxel features of the labeled data obtained by the student model inference. In addition, the prototypes of the labeled data are processed by the EMA strategy to obtain stable prototypes.

[0069] Unlabeled data is fed into the student model and the teacher model respectively to obtain the corresponding predicted probabilities and pseudo labels. The loss is then calculated according to the following process: First, Monte Carlo Dropout is performed, that is, the unlabeled data is inferred by the teacher model 8 times, and the average is taken to obtain voxel-level stable features and an uncertainty weight map based on information entropy. The formula is as follows:

[0070]

[0071] in, represents the probability that the teacher model predicts class c, T represents the number of Monte Carlo sampling, which is set to 8; H v represents the cross entropy of the predicted probability mean after T Monte Carlo Dropout when unlabeled data is input into the teacher model; Φ v is the reliability weight map of each pixel after normalization.

[0072] Secondly, guided by the pseudo-labels of the teacher model, the class prototypes of the unlabeled data are calculated. The projection head output of the teacher model is (K represents the number of channels) is upsampled by trilinear interpolation to the same value as the segmentation mask The size is consistent, and then the voxel-level feature map is obtained after Monte Carlo sampling and averaging. Represents the features of each pixel output by the model. Then, the one-hot encoding output of the teacher model The class prototype of the unlabeled data is obtained by guiding (obtained by argmax) and weighting through the reliability weight map. The formula is as follows:

[0073]

[0074] Where I(·) is the indicator function. By weighting it with the reliability weight graph, the model can generate more stable prototypes for unlabeled data.

[0075] Then calculate the voxel feature vector f by the following formula t,v Prototypes with unlabeled data The corresponding distance between and

[0076]

[0077] Then, the prototype estimation method is used to divide the output of the student model into a certain area and a fuzzy area. The prototype estimation is based on the following assumption: if the pseudo label of the vth voxel comes from the output of the student model is the target (background), and it comes from the feature f of the teacher model t,vIf the voxel is closer to the background (target) prototype than the target (1.0 background) prototype, then the voxel is suspected to be an incorrectly predicted voxel, that is, a fuzzy area. Otherwise, it is a confirmed area. Fuzzy area mask (Mask) E ambiguity And determine the region mask (Mask) E certain It can be obtained by the following formula:

[0078]

[0079] E certain =1-E ambiguity

[0080] Finally, the consistency regularization loss between the output of the student model and the output of the teacher model (with added noise) is calculated for the fuzzy region of the unlabeled data. The prototype consistency loss between the predicted probability of the student model based on the fusion prototype and the pseudo-label of the student model is calculated for the fixed region, and uncertainty is weighted. The consistency regularization loss (CR) of the fuzzy region is formulated as follows:

[0081]

[0082] Among them, p t ,p s are the probability outputs of the teacher model and the student model, E ambiguity is the mask of the blurred area, and MSE is the mean square error.

[0083] For the determined areas in the student model output, the feature-prototype similarity is used to approximate the probability of each type of voxel. For the projection head output of the unlabeled data input student model, the voxel-level feature map is upsampled. Calculate the similarity between each voxel feature and the corresponding class (target or background), and then convert it into class prediction probability. Its formula is as follows:

[0084]

[0085] Among them, c∈{bg,obj} are the prototypes of the background class and the target class respectively, sim(·) represents the cosine similarity, and p c Represents the fusion prototype of class c. The similarity between the feature vector of v voxel and the target class and background class is converted into a probability representation.

[0086] The fusion prototype is formulated as follows:

[0087]

[0088] Where c∈{bg,obj},λ lab=1-λ con ,λ unlab =λ con , and λ con This is a widely used time-series-related Gaussian warm-up function. During training, the proportion of labeled prototypes decreases from 1 to 0, while the proportion of unlabeled prototypes increases from 0 to 1. By dynamically adjusting the ratio of labeled data prototypes to unlabeled data prototypes, training initially relies on labeled data prototypes to ensure stability, and later gradually relies on unlabeled data prototypes to improve generalization ability.

[0089] The key idea of prototype-based consistency regularization is to encourage consistency between the model's segmentation predictions and prototype-based predictions. This pattern implicitly normalizes feature representation: features from the same class must be close to the fused prototype of that class, while being far away from prototypes of other classes. The specific prototype consistency (PC) loss function is calculated as follows:

[0090]

[0091] in, They represent the predicted probability of the student model based on the prototype and the segmentation prediction pseudo label, E certain Indicates the mask of the determined area, Φ v represents the voxel-by-voxel uncertainty weight.

[0092] Therefore, the overall loss for unlabeled data can be obtained:

[0093]

[0094] in, Represents the consistency regularization loss of the fuzzy area, that is, MSE loss, p s ,p t represent the predicted probabilities of the student model and the teacher model respectively. represents the prototype consistency loss of the determined region. is a Gaussian prediction function that changes over time and increases with the number of iterations t. max is set to 0.1, and λ pc is a similar function, where w max Set to 1. That is, as the number of training rounds increases, the consistency regularization loss weight of the fuzzy area and the prototype consistency loss weight of the definite area increase from 0 to 0.1 and 1 respectively.

[0095] To sum up, the total loss of the training model can be obtained:

[0096]

[0097] in, represents the total loss of the model, represents the loss of labeled data, Indicates the loss of unlabeled data.

[0098] Step 5: Use the final trained student model for pancreatic organ segmentation.

[0099] On the processed test set CT images, a patch-based sliding window strategy is adopted with a step size of 16×16×16, and the predictions of all image patches are integrated to obtain the final output.

[0100] The present invention has been described in detail above with reference to the accompanying drawings, but the specific implementation of the present invention is not limited to the above-described methods. For those skilled in the art, various modifications, improvements, and substitutions can be made without departing from the concept of the present invention, and these all fall within the scope of protection of the present invention. As long as various non-substantial improvements are made using the inventive concept and technical solution of the present invention, or the inventive concept and technical solution are directly applied to other occasions without modification, they are all within the scope of protection of the present invention.

Claims

1. A semi-supervised pancreatic image segmentation method based on prototype estimation and prototype consistency, characterized by: The following steps are involved: Step 1: Obtain and preprocess the pancreatic CT dataset; Step 2: Divide the dataset preprocessed in step 1 into a training set and a test set, where the training set is Among them, Represents labeled data and its true label, represents unlabeled data, x l(i) and x u(i) Represent the labeled CT image and the unlabeled CT image, y l(i) Represents the corresponding true label, N and M represent the number of labeled data and unlabeled data respectively, and MN>>N; Step 3: Based on the V-Net segmentation model, build the Average Teacher (MT) semi-supervised segmentation framework; the student model is the V-Net model, and the features output by the penultimate stage are used as the prototype branch, whose parameters are updated and optimized during model training; the teacher model uses the same architecture as the student model, and its parameters are obtained by exponential moving average of the student model; Step 4: Training is performed on the training set D. The same batch contains 2 labeled data and 2 unlabeled data, that is, the batch size is 4; before the data is input into the model, it is randomly cropped and cropped into image blocks of size [96,96,96], and random flipping and random rotation image enhancement techniques are applied; the labeled data is input into the student model, and the standard segmentation loss is used as the supervision signal, and the class prototype of the labeled data is generated according to the true label; the unlabeled data is input into the student model and the teacher model respectively to obtain the corresponding prediction probability and pseudo label, and then the loss calculation process is as follows: First, Monte Carlo sampling is performed, that is, unlabeled The data is inferred eight times by the teacher model and the average is taken to obtain voxel-level stable features and an uncertainty weight map based on information entropy. Then, guided by the pseudo-labels of the teacher model, the class prototype of the unlabeled data and the Euclidean distance between the voxel features and the class prototype are calculated. Based on the prototype estimation method, the output of the student model is divided into a certain area and a fuzzy area. Finally, the consistency regularization loss between the output of the student model and the output of the teacher model is calculated in the fuzzy area of the unlabeled data, and the prototype consistency loss between the predicted probability of the student model based on the fusion prototype and the pseudo-label of the student model is calculated in the certain area, and uncertainty weighting is adopted. Step 5: Use the final trained student model for pancreatic organ segmentation.

2. The semi-supervised pancreatic image segmentation method based on prototype estimation and prototype consistency according to claim 1, characterized in that: The step 1 is specifically as follows: first, all CT sequences are centrally cropped around the pancreas target (ROI), and 25 voxels at the edge are enlarged on all spatial axes to obtain a three-dimensional graphic block centered on the pancreas target; Then, the Hounsfield units (HU) in the CT sequences were rescaled, and a soft tissue window was used. The window widths and window levels of enhanced CT and plain CT were 75 and 400, 60 and 360, respectively, that is, the grayscale value ranges were [-125, 275] and [-120, 240], respectively. Finally, all sequences were resampled to an isotropic resolution of 1.0 mm × 1.0 mm × 1.0 mm and normalized to zero mean and unit variance before being input into the network.

3. The semi-supervised pancreatic image segmentation method based on prototype estimation and prototype consistency according to claim 1, characterized in that: The step 4 is specifically as follows: The specific process of labeled data loss calculation and class prototype generation is as follows: Labeled data is fed into the student model, and the standard segmentation loss between the model prediction and the true label is calculated, which is a combination of cross entropy loss and Dice loss: Among them, among them, represents the cross entropy loss, represents Dice loss; p, y represent the predicted probability and true label of the student model respectively, Indicates the overall loss of labeled data; in addition, the output of the student model projection head is upsampled by trilinear interpolation to obtain the segmentation mask Feature maps of consistent size Represents the characteristics of each voxel; then, the feature map is directly average-pooled using the true label to calculate the prototype of the corresponding class. The calculation formula is as follows: Among them, I(·) represents the indicator function, y represents the true label of the labeled data, and f l,v represents the voxel features of the labeled data obtained by the student model inference; in addition, the prototype of the labeled data is processed by the EMA strategy to obtain a stable prototype; The unlabeled data is input into the teacher model and the student model respectively. The specific processing flow of the designed prototype estimation method is as follows: Unlabeled data is input into the teacher model, and after Monte Carlo sampling, voxel-level stable features and uncertainty weight maps based on information entropy are obtained, which are formulated as follows: Among them, c∈{bg,obj}, represents the probability that the teacher model predicts class c, T represents the number of Monte Carlo sampling, which is set to 8; H v represents the cross entropy of the predicted probability mean after T Monte Carlo Dropout when unlabeled data is input into the teacher model; Φ v is the reliability weight map of each pixel after normalization; Projection head output of the teacher model (K represents the number of channels) is upsampled by trilinear interpolation to the same value as the segmentation mask The size is consistent, and then the voxel-level feature map is obtained after Monte Carlo sampling and averaging. Represents the features of each pixel output by the model; then, the one-hot encoding output of the teacher model Guided by the reliability weight graph, the class prototype of the unlabeled data is obtained, which is formulated as follows: Where I(·) is the indicator function. By weighting the reliability weight graph, the model can generate a more stable prototype for unlabeled data; thus, a fusion prototype can be obtained. The specific calculation formula is as follows: Where c∈{bg,obj},λ lab =1-λ con ,λ unlab =λ con , and λ con It is a widely used time-related Gaussian warm-up function. During the training process, the proportion of labeled prototypes decreases from 1 to 0, while the proportion of unlabeled prototypes increases from 0 to 1. Then calculate the voxel feature vector f by the following formula t,v Prototypes with unlabeled data The corresponding distance between and When the pseudo label of the vth voxel from the output of the student model is the target class, and it comes from the feature f of the teacher model t,v If the voxel is closer to the background prototype than the target prototype, then the prediction of the voxel is suspected to be an incorrectly predicted voxel, that is, a fuzzy area; otherwise, it is a determined area; if the pseudo label of a voxel is background, and its feature f t,v If it is closer to the target prototype, then it is a fuzzy area, otherwise it is a definite area; the fuzzy area mask (Mask) E ambiguity And determine the region mask (Mask) E certain It can be obtained by the following formula: AND certain =1-E ambiguity After the above prototype estimation method divides the model's unlabeled data prediction into fuzzy areas and certain areas, the specific loss calculation process for each area is as follows: For the fuzzy areas in the student model output, a simple consistency regularization is used to supervise the training of the model, which requires that the output probability of the student model is consistent with the output probability of the teacher model: Among them, p t ,p s are the probability outputs of the teacher model and the student model, E ambiguity is the mask of the blurred area, MSE represents the mean square error; For a certain region in the output of the student model, the feature-prototype similarity is used to approximate the probability of each type of voxel; unlabeled data x u Input the projection head output of the student model and upsample to obtain the voxel-level feature map Calculate the similarity between each voxel feature and the corresponding class, and then convert it into class prediction probability; its formula is as follows: Among them, c∈{bg,obj} are the prototypes of the background class and the target class respectively, sim(·) represents the cosine similarity, and p c Represents the fusion prototype of the cth class; converts the similarity between the feature vector of v voxel and the target class and background class into a probability representation; Features from the same class must be close to the fusion prototype of the class and far away from the prototypes of other classes. The specific prototype consistency (PC) loss function calculation formula is as follows: in, They represent the predicted probability of the student model based on the prototype and the segmentation prediction pseudo label, E certain Indicates the mask of the determined area, Φ v represents the uncertainty weight of each voxel; To sum up, we can get the overall loss of unlabeled data: in, Represents the consistency regularization loss of the fuzzy area, that is, MSE loss, p s ,p t Represent the predicted probabilities of the student model and the teacher model respectively; represents the prototype consistency loss of the determined region; is a Gaussian prediction function that changes over time and increases with the number of iterations t. max is set to 0.1, and λ pc is a similar function, where w max Set to 1; that is, as the number of training rounds increases, the consistency regularization loss weight of the fuzzy area and the prototype consistency loss weight of the definite area increase from 0 to 0.1 and 1 respectively.

4. The semi-supervised pancreatic image segmentation method based on prototype estimation and prototype consistency according to claim 1, characterized in that: The step 5 is specifically as follows: adopting a sliding window strategy based on image blocks, with an image block size of [96,96,96] and a step size of 16×16×16, and synthesizing the predictions of all image blocks to obtain the final output.

5. The semi-supervised pancreatic image segmentation method based on prototype estimation and prototype consistency according to claim 1, characterized in that: In step 4, when calculating the prototype consistency loss, the fused prototype obtained by fusing the labeled data prototype and the unlabeled data prototype is used; the overall loss of the model is obtained by weighting the labeled data loss and the unlabeled data loss.

Citation Information

Cited By

  • Cross-center credible semi-supervised cardiac ultrasound image segmentation method

    CN121280438A

  • A cross-center trustworthy semi-supervised echocardiogram image segmentation method

    CN121280438B

  • Semi-supervised semantic segmentation method based on self-adaptive pseudo tag generation

    CN121904372A