Semi-supervised Segmentation Model Based on Inter-model and Intra-model Uncertainty
Through independent parameter training of student models and teacher models and pseudo-mask-guided feature enhancement, combined with multi-scale multi-stage feature aggregation, the uncertainty problems between and within models in semi-supervised learning are solved, efficient medical image segmentation is achieved, and labeled data needs are reduced.
Patent Information
- Application Number
- CN202211704924.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-12-29
AI Technical Summary
Existing semi-supervised learning models are difficult to effectively model the consistency between labeled and unlabeled data in medical image segmentation, resulting in inter- and within-model uncertainties, affecting the accuracy and efficiency of segmentation results.
A semi-supervised segmentation model based on inter- and in-model uncertainty is adopted, and the parameter independent training of the student model and teacher model and pseudo-mask-guided feature enhancement is combined with a multi-scale multi-stage feature aggregation module and a semi-supervised learning loss module to reduce the workload of labeled data and improve the segmentation accuracy.
Effectively use partial labeling data to extract the contextual characteristics of cells and glands, reduce the need for expert labeling data, and improve the accuracy and efficiency of medical image segmentation.
Smart Images

Figure CN115797637B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image segmentation, and specifically relates to a semi-supervised segmentation model based on inter-model and intra-model uncertainties. Background Art
[0002] Deep learning models have shown great success in various image segmentation tasks, especially when there are a large number of annotated training samples [1]–[4]. However, obtaining pixel-level annotations is a very time-consuming task. This greatly reduces efficiency, especially in applications that require domain knowledge and expertise (such as biomedical image processing). Semi-supervised learning (SSL) is one of the methods to address this challenge, which uses limited supervised data for training. In semi-supervised image segmentation, the model learns from pixels with known semantic labels and fully utilizes the information of any unlabeled data.
[0003] One of the main challenging problems in semi-supervised segmentation is how to model the consistency between labeled and unlabeled data. Inconsistency will lead to uncertainty or difference in the segmentation results. In the recently popular semi-supervised framework, the teacher-student framework (Mean Teacher [5]), there is an inconsistency between the predictions of the student and teacher models [6]–[8], which is called inter-model uncertainty. Since there is no ground truth for unlabeled data, a common strategy is to use the predictions of the teacher model as guidance. However, in the Mean Teacher architecture, previous work cannot guarantee that the teacher model always produces better results than the student model on unlabeled data, and the above-mentioned prediction difference helps to estimate uncertainty. Secondly, in previous work, the internal uncertainty and network perturbation within the student model itself are ignored, and this uncertainty is called intra-model uncertainty. The features extracted in a specific layer of a convolutional neural network (CNN) may affect subsequent layers, which greatly affects the receptive field and further leads to inconsistency in the propagation of information from the shallow layer to the deep layer [9],
[10] . Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a semi-supervised segmentation model based on inter-model and intra-model uncertainties, which can efficiently learn from labeled and unlabeled data and is of great significance for reducing the workload of professional doctors in annotating data.
[0005] The purpose of the present invention is achieved by the following technical solutions.
[0006] A semi-supervised segmentation model based on inter-model and intra-model uncertainty, including a student model, a teacher model, and a semi-supervised learning loss module. The student model and the teacher model are respectively a medical image segmentation model (PG-FANet). The initial data of the student model is labeled data and unlabeled data, and the initial data of the teacher model is unlabeled data. Each of the medical image segmentation models includes: a convolutional block, a second-order network model structure, a pseudo-mask guided feature enhancement module (MGFE), a multi-scale multi-stage feature aggregation module (MMFA), a first convolutional layer, a second convolutional layer, and a third convolutional layer. The second-order network model structure includes: a first-order sub-network and a second-order sub-network;
[0007] The convolutional block is used to input the initial data thereto and respectively flow the rough features output from the convolutional block to the first-order sub-network and the pseudo-mask guided feature enhancement module;
[0008] The second-order sub-network and the first-order sub-network have the same architecture, each including: I + 1 residual blocks (RB i_s ) and an atrous spatial pyramid pooling (ASPP) module. The I + 1 residual blocks (RB i_s ) of the first-order sub-network are used to finely adjust the rough features and then transport the first-order refined features to the atrous spatial pyramid pooling (ASPP) module of the first-order sub-network. The atrous spatial pyramid pooling (ASPP) module of the first-order sub-network is used to extract high-order latent features from the first-order refined features;
[0009] The first convolutional layer is used to generate a pseudo-mask for the high-order latent features obtained by the first-order sub-network;
[0010] The pseudo-mask guided feature enhancement module is used to enhance the expression ability of the rough features by using the pseudo-mask to obtain pseudo-mask guided fusion features;
[0011] The I + 1 residual blocks (RB i_s ) of the second-order sub-network are used to input the fusion features and output second-order refined features. The atrous spatial pyramid pooling (ASPP) module of the second-order sub-network is used to receive the second-order refined features output from the (I + 1)-th residual block of the second-order sub-network and output high-order latent features;
[0012] The multi-scale multi-stage feature aggregation module (MMFA) includes: a multi-scale feature aggregation module and a multi-stage feature aggregation module. The multi-scale feature aggregation module is used to perform multi-scale feature aggregation on the low-level features output from the i-th residual block of the first-order sub-network and the low-level features output from the i-th residual block of the second-order sub-network to obtain multi-scale aggregated features, where i = 1,..., I;
[0013] The second convolutional layer is used to fuse the multi-scale aggregated features to output high-order features;
[0014] The multi-stage feature aggregation module is used to perform multi-stage feature aggregation on the feature output of the (I + 1)-th residual block of the first-order sub-network, the feature output of the (I + 1)-th residual block of the second-order sub-network, and the high-order features, and then output multi-scale multi-stage aggregated features;
[0015] The third convolutional layer is used to perform feature concatenation and then fusion on the multi-scale multi-stage aggregated features and the high-order latent features obtained from the second-order sub-network to obtain the prediction result;
[0016] The calculation formula of the semi-supervised learning loss module is:
[0017]
[0018] where Lse g is the supervised loss function, λ(t) represents the balance factor of the consistency loss in the t-th training, represents the labeled dataset, X l represents the images in the labeled dataset, Y l represents the labels of the images in the labeled dataset, and M represents the number of images in the labeled dataset; represents the images in the unlabeled dataset, N represents the number of images in the unlabeled dataset, and λ in tr a is the weight factor for controlling the uncertainty regularization L intra in the model;
[0019]
[0020] L intra = L mse (F1(x r |θ t ), F2(x r |θ t ))
[0021] where U shape is the shape uncertainty, U shape = -u shape log u shape
[0022] u shape = |softmax(F2(x r |θ t )) - Softmax(F2(x r |θ t '))|
[0023] F2(x r |θ t ) is the prediction result of the student model in the t-th training, and F1(xr |θ t ) is the pseudo-mask of the student model in the t-th training, L mse represents the mean squared error loss function; σ represents the min-max normalization function used to normalize the shape uncertainty U shape to [0, 1];
[0024] θ t is the weight of the student model in the t-th training, θ′ t = αθ′ t-1 + (1 - α)θ t , θ′ t is the weight of the teacher model in the t-th training; θ′ t-1 is the weight of the teacher model in the (t - 1)-th training, and α is the decay rate of the exponential moving average (Exponential Moving Average) used to update the student model θ during the overall training process t ;
[0025] is the correction of the prediction result of the teacher model in the t-th training;
[0026]
[0027] F2(x r |θ t ′) is the prediction result of the teacher model in the t-th training, μ′ r = -F2(x r |θ t ′)logF2(x r |θ t ′).
[0028] In the above technical solution, T is the total number of training times.
[0029] In the above technical solution, α = 0 to 1.
[0030] In the above technical solution, the inter-model uncertainty is modeled as:
[0031]
[0032] The intra-model uncertainty (U intra ) is:
[0033]
[0034] In the above technical solution, the first convolutional layer includes: an upsampling layer and a convolutional layer, and the calculation process of the first convolutional layer is as follows:
[0035] Ys = Conv(Up(X c ))
[0036] where X c is the high-order latent feature obtained from the first-order sub-network, Up is the upsampling layer in the first convolutional layer, Conv is the convolutional layer, and Y s is the pseudo-mask.
[0037] In the above technical solution, the calculation formula of the multi-scale feature aggregation module is:
[0038]
[0039] where Xm is the multi-scale aggregation feature, is the i-th residual block (RBis) of the s-th order sub-network, is the low-level feature output by the (i - 1)-th residual block of the s-th order sub-network, s = 1, 2, i = 1, ……, I, where is the rough feature output by the convolutional block, is the fusion feature guided by the pseudo-mask, Up is the upsampling layer, δ is the parametric rectified linear unit (PReLU), is the batch normalization process, and Conv is the convolutional layer.
[0040] In the above technical solution, the operation process of the second convolutional layer is as follows:
[0041] X′ m = Conv(X m )
[0042] where X′ m is the high-order feature, Conv is the convolutional layer, and X m is the multi-scale aggregation feature.
[0043] In the above technical solution, the calculation formula of the multi-stage feature aggregation module is as follows:
[0044]
[0045] where X′ m is the high-order feature, X h is the multi-scale multi-stage aggregation feature, Up is the upsampling layer, δ is the parametric rectified linear unit (PReLU), is the batch normalization process, Conv is the convolutional layer, is the (I + 1)-th residual block of the s-th order sub-network, s = 1, 2, is the feature output of the I-th residual block of the s-th order sub-network.
[0046] In the above technical solution, the third convolutional layer includes: upsampling, feature concatenation, and a convolutional layer. The calculation formula of the third convolutional layer is as follows:
[0047] Y s = Conv(concat(X h , Up(X f )))
[0048] where Y s is the prediction result, Conv is the convolutional layer, concat is the feature concatenation operation, X h is the multi-scale multi-stage aggregated feature, Up is the upsampling layer, and X f is the high-order latent feature obtained by the second-order sub-network.
[0049] The semi-supervised segmentation model of the present invention effectively extracts the context features of cells (MoNuSeg) / glands (CRAG) using partial labeled data, segments the corresponding cell / gland instances and uses them for downstream task analysis, reduces the amount of labeled data used, and greatly reduces the workload required for expert-labeled data. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a schematic structural diagram of the medical image segmentation model of the present invention;
[0051] Figure 2 is a schematic structural diagram of the semi-supervised segmentation model;
[0052] Figure 3 are (a) cell segmentation result diagrams trained using 5% / 10% / 20% / 50% / 70% / 90% labeled data on the MoNuSeg dataset; (b) gland segmentation result diagrams trained using 5% / 10% / 20% / 50% / 70% / 90% labeled data on the CRAG dataset;
[0053] Figure 4 are cell and gland segmentation result diagrams trained using 5% / 10% / 20% / 50% labeled data on the MoNuSeg and CRAG datasets. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] The technical solution of the present invention will be further described below with reference to specific embodiments.
[0055] Embodiment 1
[0056] A semi-supervised segmentation model based on inter-model and intra-model uncertainty, as Figure 2 shown, includes a student model, a teacher model, and a semi-supervised learning loss module, as Figure 1As shown, the student model and the teacher model are respectively a medical image segmentation model (PG-FANet). The parameters of the student model and the teacher model are independent. The teacher model guides the multi-stage outputs of the student model to enable the student model to obtain better segmentation performance. The initial data of the student model is labeled data and unlabeled data, and the initial data of the teacher model is unlabeled data. The student model is trained using labeled and unlabeled data through supervised loss and unsupervised loss. Each medical image segmentation model includes: a convolutional block, a second-order network model structure, a pseudo-mask guided feature enhancement module (MGFE), a multi-scale multi-stage feature aggregation module (MMFA), a first convolutional layer, a second convolutional layer, and a third convolutional layer. The second-order network model structure includes: a first-order sub-network and a second-order sub-network;
[0057] The convolutional block is used to input the initial data thereto and respectively flow the rough features output from the convolutional block to the first-order sub-network and the pseudo-mask guided feature enhancement module;
[0058] The architectures of the second-order sub-network and the first-order sub-network are the same, each including: I + 1 residual blocks (RB i_s ) and an atrous spatial pyramid pooling (ASPP) module. The I + 1 residual blocks (RB i_s ) of the first-order sub-network are used to finely adjust the rough features and then convey the first-order fine-tuned features to the atrous spatial pyramid pooling (ASPP) module of the first-order sub-network; The atrous spatial pyramid pooling (ASPP) module of the first-order sub-network is used to extract high-order latent features from the first-order fine-tuned features; In the embodiment of the present invention, I = 3;
[0059] The first convolutional layer is used to generate a pseudo-mask for the high-order latent features obtained by the first-order sub-network;
[0060] The first-order sub-network is used for rough pseudo-mask generation. The pseudo-mask guided feature enhancement module is used to enhance the expression ability of the rough features by using the pseudo-mask to obtain pseudo-mask guided fusion features; The pseudo-mask guided feature enhancement module moves the attention of the second-order sub-network to the region of interest under the guidance of the pseudo-mask. The pseudo-mask guided feature enhancement module splices the rough features generated by the convolutional block and the pseudo-mask features of the first-stage sub-network and fuses them through a 1×1 convolutional layer as the input of the second-stage sub-network;
[0061] The I + 1 residual blocks (RB i_s ) of the second-order sub-network are used to input the fusion features and output second-order fine-tuned features. The atrous spatial pyramid pooling (ASPP) module of the second-order sub-network is used to receive the second-order fine-tuned features output from the I + 1 residual block of the second-order sub-network and output high-order latent features for refining the prediction results;
[0062] Since the second-order sub-network and the first-order sub-network extract features for different scales, shapes, and sizes, the present invention uses a multi-scale multi-stage feature aggregation module (MMFA) to aggregate multi-scale and multi-stage features, improve the feature expression ability of the model, and avoid the problem of feature incompatibility in the U-shaped skip connection. The multi-scale multi-stage feature aggregation module (MMFA) includes: a multi-scale feature aggregation module and a multi-stage feature aggregation module. The multi-scale feature aggregation module is used to perform multi-scale feature aggregation on the low-level features output by the i-th residual block of the first-order sub-network and the low-level features output by the i-th residual block of the second-order sub-network to obtain multi-scale aggregated features, where i = 1, ……, I;
[0063] The second convolutional layer is used to fuse the multi-scale aggregated features to output high-order features;
[0064] The multi-stage feature aggregation module is used to perform multi-stage feature aggregation on the feature output of the (I + 1)-th residual block of the first-order sub-network, the feature output of the (I + 1)-th residual block of the second-order sub-network, and the high-order features, and then output multi-scale multi-stage aggregated features;
[0065] The third convolutional layer is used to perform feature concatenation and then fusion on the multi-scale multi-stage aggregated features and the high-order latent features obtained from the second-order sub-network to obtain a prediction result;
[0066] The calculation formula of the semi-supervised learning loss module is:
[0067]
[0068] where Lse g is the supervised loss function, λ(t) represents the balance factor of the consistency loss in the t-th training, and its calculation formula is T is the total number of trainings, and in the present invention, T is set to 300. represents the labeled data set, X l represents the images in the labeled data set, Y l represents the annotations of the images in the labeled data set, and M represents the number of images in the labeled data set; represents the images in the unlabeled data set, and N represents the number of images in the unlabeled data set; λ intra is the weight factor for controlling the uncertainty regularization L intra in the model, and is set to 1 in the present invention; the present invention uses the shape uncertainty U shape to enhance the model's attention to the boundary region, and fuses the shape uncertainty into L inter , L inter represents the unsupervised consistency loss to minimize the inter-model uncertainty, and L intra represents the additional intra-model uncertainty regularization, which incorporates the intra-model uncertainty into the semi-supervised learning objective;
[0069]
[0070] L intra = L mse (F1(x r |θ t ), F2(x r lθ t ))
[0071] where U shape is the shape uncertainty, U shape = -u shape log u shape
[0072] u shape = |softmax(F2(x r |θ t )) - Softmax(F2(x r |θ t '))|
[0073] F2(x r |θ t ) is the prediction result of the student model in the t-th training, F1(x r |θ t ) is the pseudo-mask of the student model in the t-th training, L mse represents the mean squared error loss function; σ represents the min-max normalization function, which is used to normalize the shape uncertainty U shape to [0, 1]. Thus, the difference in boundary prediction between the student model and the teacher model can be reduced by the shape uncertainty weighting method of the present invention, so that the details at the boundary of the medical image can be adjusted during training, and the complete shape of the segmented object can be retained. In addition to promoting the complete segmentation of histological images, the present invention also uses the shape information uncertainty weight (U shape ) to enhance the model's attention to the boundary region for better segmentation prediction.
[0074] In each training process, the weight update of the teacher model is based on the weight of the teacher model in the previous training process and the weight of the student model in the current training process. Specifically, the weight θ' t of the teacher model in the t-th training, θ' t = αθ' t-1 + (1 - α)θ t ; θ t is the weight of the student model in the t-th training, θ' t-1is the weight of the teacher model in the (t - 1)-th training, and α is the decay rate of the exponential moving average (EMA) used to update the exponential moving average of the student model θt during the overall training process. Generally, it can take values from 0 to 1, and in this embodiment, it takes 0.99.
[0075] is the correction of the prediction result of the teacher model in the t-th training.
[0076]
[0077] F2(x r |θ t ′) is the prediction result of the teacher model in the t-th training, and u′ r is the uncertainty estimate of the r-th sample prediction, and μ′ r = -F2(x r |θ t ′) log F2(x r |θ t ′);
[0078] Based on the prediction differences between the student model and the teacher model, the inter-model uncertainty can be modeled as:
[0079]
[0080] However, due to the hierarchical architecture of the neural network, there are differences in the receptive fields at different stages within the student model, which will lead to inconsistent predictions of different sub-networks. To address the differences, the result predictions of each sub-network must be highly consistent. Therefore, the present invention additionally estimates the intra-model uncertainty (U intra ) as:
[0081]
[0082] On the one hand, the supervised learning process continuously improves the ability of the student model according to the L s1 in L seg terms. On the other hand, the semi-supervised learning process forces the final prediction of the student model to be consistent with the model of the teacher model. At the same time, the student model keeps the pseudo-mask of the first-order sub-network consistent with the prediction result of the first-order sub-network. Through such operations, the inconsistencies existing in semi-supervised learning are ultimately restricted.
[0083] Since the teacher model cannot always provide more accurate predictions than the student model. Therefore, the present invention proposes an inter-model and intra-model uncertainty consistency module U inter and U intra, preventing the noise and uncertainty existing in the teacher model prediction from misleading the student model. To dynamically prevent the teacher model from obtaining predictions with high uncertainty, the present invention introduces a learnable loss function L inter, to penalize the uncertainty generated by the teacher model. When the teacher model provides unreliable results (high uncertainty), is approximated to F2(x r |θ t ), on the contrary, when the teacher model is confident (low uncertainty), is close to F2(x r |θ t ), providing a reliable prediction as the learning target for the student model.
[0084] In the above technical solution, the first convolutional layer includes: an upsampling layer and a convolutional layer, and the calculation process of the first convolutional layer is as follows:
[0085] Y s =Conv(Up(X c ))
[0086] where X c is the high-order latent feature obtained by the first-order sub-network, Up is the upsampling layer in the first convolutional layer, Conv is the convolutional layer, and Y s is the pseudo-mask.
[0087] In the above technical solution, the calculation formula of the multi-scale feature aggregation module is:
[0088]
[0089] where X m is the multi-scale aggregated feature, is the i-th residual block (RB i_s ) of the s-th order sub-network, is the low-level feature output by the (i-1)-th residual block of the s-th order sub-network, s = l, 2, i = l,..., I, where, is the rough feature output by the convolutional block, is the fusion feature guided by the pseudo-mask, Up is the upsampling layer, δ is the parametric rectified linear unit (PReLU), is the batch normalization process, and Conv is the convolutional layer. The multi-scale feature aggregation module re-uses the information guided by the pseudo-mask and obtains a better feature representation for further propagation.
[0090] In the above technical solution, the second convolutional layer is used to further improve the feature expression and provide more expressive high-level features for the multi-stage feature aggregation module. The operation process of the second convolutional layer is as follows:
[0091] X′m = Conv(X m )
[0092] where X' m is the high - order feature, Conv is the convolutional layer, and X m is the multi - scale aggregated feature;
[0093] In the above technical solution, the convolutional layer in the multi - stage feature aggregation module is used to receive the feature output of the (I + 1)-th residual block of the first - order sub - network and the feature output of the (I + 1)-th residual block of the second - order sub - network. As the network depth increases, the spatial information of low - level features (such as region boundaries) may be lost. The present invention uses the multi - stage feature aggregation module to fuse the high - order feature, the feature output of the (I + 1)-th residual block of the first - order sub - network, and the feature output of the (I + 1)-th residual block of the second - order sub - network, avoiding the introduction of U - shaped skip connections with incompatible features.
[0094] The calculation formula of the multi - stage feature aggregation module is as follows:
[0095]
[0096] where X' m is the high - order feature, X h is the multi - scale multi - stage aggregated feature, Up is the up - sampling layer, δ is the parametric rectified linear unit (PReLU), is the batch normalization process, Conv is the convolutional layer, is the (I + 1)-th residual block of the s - th order sub - network, s = 1, 2, is the feature output of the I - th residual block of the s - th order sub - network.
[0097] In the above technical solution, the third convolutional layer includes: up - sampling, feature concatenation, and a convolutional layer. The calculation formula of the third convolutional layer is as follows:
[0098] Y s = Conv(concat(X h , Up(X f )))
[0099] where Y s is the prediction result, Conv is the convolutional layer, concat is the feature concatenation operation, X h is the multi - scale multi - stage aggregated feature, Up is the up - sampling layer, and X f is the high - order latent feature obtained from the second - order sub - network.
[0100] Embodiment 2
[0101] The multi-organ nuclei segmentation dataset (MoNuSeg)
[13] consists of 44 H&E-stained histopathological images, which were collected from multiple hospitals and have a resolution of 1000×1000 pixels. 30 histopathological images were selected from the 44 histopathological images to form the training dataset and 14 were used as the test dataset. The training dataset was randomly divided into 27 as training set A and 3 as validation set B, and the test dataset was used as test set C. Image patches of size 128×128 were cropped from each histopathological image using a sliding window, for a total of 1728 image patches. Online data augmentation was performed on the image patches, including random scaling, flipping, rotation, and affine operations, and all image patches were normalized using the mean and standard deviation of the images in ImageNet
[14] .
[0102] The semi-supervised segmentation model in Example 1 was used to screen the electron micrographs of cells.
[0103] The image patches of training set A were used as the initial data and fed into the student model and the teacher model for training respectively. In the student model, the proportion of histopathological images of cells with manual annotations in training set A was 5%, 10%, 20%, 50%, 70%, and 90% respectively (the rest were unlabeled data). The training set A of the teacher model did not use labels. The validation set B was used as the initial data and fed into the student model (without using labels), and the optimal θ was selected through the validation set B. t For the semi-supervised segmentation model with optimal θ, the image patches of test set C without using labels were sequentially fed into t the student model, and the prediction results were re-stacked in the order of small patches to obtain accurate cell segmentation results. The effects of the examples are as Figure 4 shown, where Figure 4 the "fully supervised" corresponding to MoNuSeg in
[0104] Example 3
[0105] Gland dataset production: The colorectal adenocarcinoma gland (CRAG) dataset contains a total of 38 whole slide images (WSIs), from which 213 H&E CRA images with different cancer grades were obtained
[15] . The semi-supervised segmentation model in Example 1 was used to screen gland electron microscopy images, and 173 H&E CRA images were selected to form the training dataset, and 40 H&E CRA images were selected to form the test dataset as test set F. The training dataset was randomly divided into 153 as the training set D and 20 as the validation set E. Most of the resolutions of the H&E CRA images are 1512×1516. The present invention extracted 5508 image patches of 480×480 pixels from 153 H&E CRA images. Further perform online data augmentation, including random scaling, flipping, rotation, and affine operations. All these image patches were normalized by using the mean and standard deviation of the images in ImageNet
[14] .
[0106] The image patches of the training set D were fed into the student model and the teacher model for training. In the student model, the proportion of manually annotated gland H&E CRA images in the training set D was 5%, 10%, 20%, 50%, 70%, and 90% respectively (the rest were unlabeled data). The training set D of the teacher model did not use labels, and the optimal θ was selected using the validation set E. t The semi-supervised segmentation model (without using labels) with the optimal θ. The image patches of the unlabeled test set F were fed into the student model with the optimal θ t in the order of small cut blocks for prediction, and the prediction results were re-stacked in the order of small cut blocks to obtain accurate gland segmentation results. The effects of the examples are as Figure 4 shown, where Figure 4 the "fully supervised" corresponding to CRAG in
[0107] is the prediction result obtained from the test set F in Example 3 of Patent Application No. 2022113429217.
[0108] The cell segmentation quality score metrics include: F1-score (F1), intersection over union (IoU), average Dice coefficient (Dice), aggregated Jaccard index (AJI), and 95% Hausdorff distance (95HD).
[0109] The gland segmentation quality fraction metrics include: F1-score (F1), object-level Dice coefficient (Dice obj ), object-level Hausdorff distance (Haus obj ), and 95% object-level Hausdorff distance (95HD obj ).
[0110] The present invention was also compared with the most recent state-of-the-art semi-supervised models, including the Mean Teacher (MT) [5] model, the uncertainty-aware self-ensembling model (UA-MT) [6], the Interpolation Consistency Training model (ICT)
[11] , the Transform Consistency Self-ensembling Model (TCSM) [7], and the Dual Uncertainty Weighting model (DUW)
[12] . The most recent state-of-the-art semi-supervised models were trained using 5%, 10%, 20%, 50%, 70%, and 90% of the cell histopathology images with manual annotations (the remaining being unlabeled data), and their quality fractions of the AJI with Example 2 are as Figure 3 shown in a of; the most recent state-of-the-art semi-supervised models were trained using 5%, 10%, 20%, 50%, 70%, and 90% of the gland H&E CRA images with manual annotations (the remaining being unlabeled data), and their quality fractions of the Dice with Example 3 are as obj shown in b of Figure 3 .
[0111] Cell segmentation: As shown in Table 1, first, when trained on the same amount of labeled data, the present invention compared the semi-supervised segmentation model (PG-FANet SSL) with the fully supervised method PG-FANet full (Patent Application No. 2022113429217). As the number of labeled images increased from 5% to 50%, the AJI values of the semi-supervised segmentation model (PG-FANet SSL) of the present invention increased by 5.9%, 2.5%, 2.7%, and 2.8% respectively compared to the fully supervised method. It is worth noting that when the proportion of labeled data increased from 5% to 50%, the AJI score obtained a significant improvement, indicating that when there is only a small amount of labeled data, increasing the amount of labeled data will have a significant impact on the model. Second, in contrast, the semi-supervised segmentation model (PG-FANet SSL) of the present invention was compared with all semi-supervised learning models. The results show that the present invention achieved the best performance compared to other methods.
[0112] Gland Segmentation: As shown in Table 1, when trained on data with the same annotation amount, compared with the fully supervised baseline experiment PG - FANet Full (Patent Application No. 2022113429217), when using 5% and 10% labeled data, the semi - supervised segmentation model (PG - FANet SSL) of the present invention significantly improved by 8.9% / 5.2%, 9.6% / 5.7% and 109.632 / 75.872 in terms of F1, Dice obj and Haus obj metrics. Compared with the state - of - the - art semi - supervised models, the semi - supervised segmentation model (PG - FANet SSL) of the present invention also showed the best performance.
[0113] Table 1
[0114]
[0115] In Table 1, "labeled data" represents the proportion of annotated data in Training Set A / Training Set D. Taking "5% (1 / 8)" as an example, "5%" represents the proportion of annotated data in Training Set A / Training Set D, and "1 / 8" means that there is 1 histopathological image with annotation in Training Set A / 8 H&E CRA images with annotation in Training Set D.
[0116] Example 4
[0117] Remove L inter 、L intra or / and U shape in the semi - supervised segmentation model in Example 1 according to Table 2 (the "×" in Table 2 represents removal), and respectively send them into the student model for training when the proportion of manually annotated cells in Training Set A is 5% (the rest are unlabeled data) in Example 2 and when the proportion of manually annotated glands in Training Set D is 5% (the rest are unlabeled data) in Example 3, and evaluate the prediction results of Test Set C and Test Set F. As shown in Table 2. As shown in Table 2, on the one hand, adding the model - to - model inconsistency regularization strategy improved the AJI / Dice obj performance metrics, which were respectively improved by 1.5% / 8.5% on the MoNuSeg / CRAG dataset. On the other hand, reducing the internal uncertainty of the model will increase the AJI / Dice obj performance. In addition, it can be seen from Table 2 that the shape uncertainty weighting module retains the complete shape of the segmentation in medical images, making the segmentation results improved. With all consistency regularization strategies, the uncertainty is greatly reduced, thus improving the model performance.
[0118] Table 2
[0119]
[0120] [1]H.Su, F.Xing, X.Kong, Y.Xie, S.Zhang, and L. Yang, “Robust cell detection and segmentation in histopathological images using sparse reconstruction and stacked denoising autoencoders,” in International Conference on Medical Image Computing and Computer-Assisted Intervention, 2015, pp.383-390.
[0121] [2]S.Graham et al., “MILD-Net: minimal information loss dilated network for gland instance segmentation in colon histology images,” Medical Image Analysis, vo1.52, pp.199-211, 2019. [3]Y.Xu et al., “Gland instance segmentation using deep multichannel neural networks,” IEEE Transactions on Biomedical Engineering, vol.64, no.12, pp.2901-2912, 2017.
[0122] [4]H.Qu, Z.Yan, G.M.Riedlinger, S.De, and D.N.Metaxas, “Improving nuclei / gland instance segmentation in histopathology images by full resolution meural network and spatial constrained loss,” in International Conference on Medical Image Computing and Computer-Assisted Intervention, 2019, pp.378-386.
[0123] [5]A. Tarvainen and H. Valpola, “Mean teachers are better role models: Weight-averaged consistemey targets improve semi-supervised deep learningresults,” in Advances in Neural Information Processing Systems, 2017, pp. 1195-1204.
[0124] [6]L. Yu, S. Wang, X. Li, C.-W. Fu, and P.-A. Heng, “Uncertainty-aware self-ensembling model for scmi-supervised 3D left atrium segmentation,” in International Conference on Medical Image Computing and Computer-Assisted Intervention, 2019, pp. 605-613.
[0125] [7]X. Li, L. Yu, H. Chen, C.-W. Fu, L. Xing, and P.-A. Heng, “Transformatiom-Consistent Self-Ensembling Model for Semi-supervised Medical Image Segmentation,” IEEE Transactions on Neural Networks and Learning Systems, pp. 1-12, 2020.
[0126] [8]Y. Zhou, H. Chem, H. Lin, and P.-A. Hemg, “Deep Semi-supervised Knowledge Distillation for Overlapping Cervical Cell Instance Segmentation,” in International Conference on Medical Image Computing and Computer-Assisted Intervention, 2020, pp. 521-531.
[0127] [9]Z. Zheng and Y. Yang, “Rectifying pseudo label learning via uncertainty estimation for domain adaptive semantic segmentation,” International Journal of Computer Vision, vol. 129, no. 4, pp. 1106 - 1120, 2021.
[0128]
[10] Q. Dou et al., “PnP - AdaNet: Plug - and - Play Adversarial Domain Adaptation Network at Unpaired Cross - Modality Cardiac Segmentation,” IEEE Access, vol. 7, pp. 99065 - 99076, 2019, doi: 10.1109 / ACCESS.2019.2929258.
[0129]
[11] V. Verma, A. Lamb, J. Kannala, Y. Bengio, and D. Lopez - Paz, “Interpolation Consistency Training for Semi—supervised Learning,” in Proceedings of the 28th International Joint Conference on Artificial Intelligence, 2019, pp. 3635 - 3641.
[0130]
[12] Y. Wang et al., “Double—Uncertainty Weighted Method for Semi - supervised Learning,” in International Conference on Medical Image Computing and Computer - Assisted Intervention, 2020, pp. 542 - 551.
[0131]
[13] N.Kumar etal,“A multi-organ nucleus segmentation challenge,”IEEETransactiohs onMedicalImaging,vol.39,no.5,pp.1380-1391,2019.
[0132]
[14] J.Deng,W.Dong,R.Socher,L.-J.Li,K.Li,and L.Fei-Fei,“ImageNet:Alarge-scale hierarchical image database,”in2009IEEE Conference onComputerVisionandPatternRecognition,2009,pp.248-255.
[0133]
[15] R.Awan etal,“Glandular morphometrics for objective grading ofcolorectal adenocarcinoma histology images,”ScientificReports,vol.7,no.1,pp.1-12,2017.
[0134] The above is an exemplary description of the present invention. It should be noted that without departing from the core of the present invention, any simple deformation, modification, or equivalent substitution that can be made by those skilled in the art without creative labor falls within the protection scope of the present invention.
Claims
1. A semi-supervised segmentation model based on inter-model and intra-model uncertainty, characterized in that, It includes a student model, a teacher model, and a semi-supervised learning loss module. The student model and the teacher model are each a medical image segmentation model. The initial data of the student model is labeled data and unlabeled data, and the initial data of the teacher model is unlabeled data. Each of the medical image segmentation models includes: a convolutional block, a second-order network model structure, a pseudo-mask-guided feature enhancement module, a multi-scale multi-stage feature aggregation module, a first convolutional layer, a second convolutional layer, and a third convolutional layer. The second-order network model structure includes: a first-order sub-network and a second-order sub-network; The convolutional block is used to input the initial data thereto and respectively flow the rough features output from the convolutional block to the first-order sub-network and the pseudo-mask-guided feature enhancement module; The second-order sub-network and the first-order sub-network have the same architecture, each including: I + 1 residual blocks and a dilated spatial pyramid pooling module. The I + 1 residual blocks of the first-order sub-network are used to finely adjust the rough features and then convey the first-order refined features to the dilated spatial pyramid pooling module of the first-order sub-network. The dilated spatial pyramid pooling module of the first-order sub-network is used to extract high-order latent features from the first-order refined features; The first convolutional layer is used to generate a pseudo-mask for the high-order latent features obtained by the first-order sub-network; The pseudo-mask-guided feature enhancement module is used to enhance the expression ability of the rough features by using the pseudo-mask to obtain pseudo-mask-guided fused features; The I + 1 residual blocks of the second-order sub-network are used to input the fused features and output second-order refined features. The dilated spatial pyramid pooling module of the second-order sub-network is used to receive the second-order refined features output from the (I + 1)-th residual block of the second-order sub-network and output high-order latent features; The multi-scale multi-stage feature aggregation module includes: a multi-scale feature aggregation module and a multi-stage feature aggregation module. The multi-scale feature aggregation module is used to perform multi-scale feature aggregation on the low-level features output from the i-th residual block of the first-order sub-network and the low-level features output from the i-th residual block of the second-order sub-network to obtain multi-scale aggregated features, where i = 1, ……, I; The second convolutional layer is used to fuse the multi-scale aggregated features to output high-order features; The multi-stage feature aggregation module is used to perform multi-stage feature aggregation on the feature outputs of the (I + 1)-th residual block of the first-order sub-network, the feature outputs of the (I + 1)-th residual block of the second-order sub-network, and the high-order features, and then output multi-scale multi-stage aggregated features; The third convolutional layer is used to perform feature concatenation and then fusion on the multi-scale multi-stage aggregated features and the high-order latent features obtained by the second-order sub-network to obtain a prediction result; The calculation formula of the semi-supervised learning loss module is: Among them, L seg is a supervised loss function, and λ(t) represents the balance factor of the consistency loss in the t-th training. represents the labeled dataset, X l represents the images in the labeled dataset, Y l represents the labels of the images in the labeled dataset, and M represents the number of images in the labeled dataset; represents the images in the unlabeled dataset, N represents the number of images in the unlabeled dataset, and λ intra is the weight factor for controlling the uncertainty regularization L intra in the model. L intra = L mse (F1( xr |θ t ), F2( xr |θ t )) Among them, U shape is the shape uncertainty, U shape = -u shap elogu shape u shape = |softmax(F2(x r |θ t )) - Softmax(F2(x r |θ t '))| F2(x r |θ t ) is the prediction result of the student model in the t-th training, F1(x r |θ t ) is the pseudo-mask of the student model in the t-th training, F2(x r |θ t ′) is the prediction result of the teacher model in the t-th training, L mse represents the mean square error loss function; σ represents the min-max normalization function, which is used to normalize the shape uncertainty U shape to [0, 1]; θ t is the weight of the student model at the t-th training, θ′ t = αθ′ t-1 + (1 - α)θ t , θ′ t is the weight of the teacher model at the t-th training; θ′ t-1 is the weight of the teacher model at the (t - 1)-th training, and α is the decay rate of the exponential moving average for updating the student model θ t using gradient descent during the overall training process; is the correction of the prediction result of the teacher model in the t-th training; μ′ r =-F2(x r |θ t ′)logF2(x r |θ t ′).
2. The semi-supervised segmentation model according to claim 1, wherein T is the total number of training times.
3. The semi-supervised segmentation model according to claim 2, wherein α=0~1。 4. The semi-supervised segmentation model according to claim 3, characterized in that Model - to - model uncertainty is modeled as: Model - within uncertainty is:
5. The semi-supervised segmentation model according to claim 4, wherein The first convolutional layer includes: an upsampling layer and a convolutional layer. The calculation process of the first convolutional layer is as follows: Y s = Conv(Up(X c )) Among them, X c is the high-order latent feature obtained from the first-order sub-network, Up is the upsampling layer in the first convolutional layer, Conv is the convolutional layer, and Y s is the pseudo-mask.
6. The semi-supervised segmentation model according to claim 5, wherein The calculation formula of the multi-scale feature aggregation module is: Among them, Xm is the multi-scale aggregated feature, is the i-th residual block of the s-th order sub-network, is the low-level feature output by the (i - 1)-th residual block of the s-th order sub-network, s = 1, 2, i = 1, ……, I, where is the rough feature output by the convolutional block, is the pseudo-mask-guided fusion feature, Up is the upsampling layer, δ is the parametric rectified linear unit, is the batch normalization process, Conv is the convolutional layer.
7. The semi-supervised segmentation model according to claim 6, wherein The operation process of the second convolutional layer is as follows: X′ m = Conv(X m ) Among them, X' m is a high-order feature, Conv is a convolutional layer, and X m is a multi-scale aggregated feature.
8. The semi-supervised segmentation model according to claim 7, wherein, The calculation formula of the multi-stage feature aggregation module is as follows: Among them, X′ m is a high-order feature, xh is a multi-scale multi-stage aggregation feature, Up is an upsampling layer, δ is a parametric rectified linear unit, is batch normalization processing, Conv is a convolutional layer, is the (1 + 1)-th residual block of the s-th order sub-network, s = 1, 2, is the feature output of the 1st residual block of the s-th order sub-network.
9. The semi-supervised segmentation model according to claim 8, wherein The third convolutional layer includes: upsampling, feature concatenation, and a convolutional layer. The calculation formula of the third convolutional layer is as follows: Y s = Conv(concat(X h , Up(X f ))) Among them, Ys is the prediction result, Conv is the convolutional layer, concat is the feature concatenation operation, Xh is the multi-scale multi-stage aggregated feature, Up is the upsampling layer, and X f is the high-order latent feature obtained by the second-order sub-network.
Citation Information
Patent Citations
Abdominal lymph node detection method and device based on semi-supervised learning
CN115018852A
Histology image segmentation model based on semi-supervised learning
CN115131565A