Hard sample-based noise tolerance contrast semi-supervised semantic segmentation method
By constructing a student-teacher model framework in semi-supervised semantic segmentation, combining dynamic threshold screening, noise tolerance contrast loss and hard sample sampling, the problem of pseudo-label noise and hard sample sampling in traditional methods is solved, and the robustness and generalization ability of the model are significantly improved.
Patent Information
- Application Number
- CN202510279649.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
When the traditional semi-supervised semantic segmentation method processes unlabeled data, pseudo-label noise leads to a degradation of model performance, and the hard sample sampling strategy is susceptible to pseudo-label error interference in a noisy environment, making it difficult to balance sample efficiency and noise suppression.
A contrasting semi-supervised semantic segmentation method based on hard sample noise tolerance is adopted. By constructing a student model and a teacher model, combining dynamic thresholds to filter pseudo-labels, noise tolerance contrast loss and boundary-aware hard sample sampling, high-quality pseudo-labels are gradually released, pseudo-label noise is suppressed, and boundary hard samples are accurately positioned.
It significantly improves the robustness and generalization ability of the model in a noisy environment, reduces the negative impact of noise labels on training, improves the model's learning efficiency on difficult samples, and is suitable for low-labeling scenarios such as medical imaging and autonomous driving.
Smart Images

Figure CN120220150A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and deep learning, and specifically to a hard sample noise-tolerant contrast semi-supervised semantic segmentation method. Background Art
[0002] The semi-supervised semantic segmentation method is a technique that combines labeled data and unlabeled data for image segmentation. This method aims to reduce the dependence on a large amount of labeled data, thereby reducing the annotation cost and improving the generalization ability of the model, and is commonly used in fields such as autonomous driving, medical image analysis, and intelligent video surveillance.
[0003] However, when traditional semi-supervised semantic segmentation methods process unlabeled data, they usually rely on pseudo-label generation strategies, but the noise in the pseudo-labels will cause the performance of the model to decline. In the prior art, although dynamic threshold screening can filter low-confidence pseudo-labels, it will affect the learning of difficult samples due to the reduction in the number of samples; although contrastive learning can improve the discriminability of features, the traditional InfoNCE loss does not consider pseudo-label noise, resulting in distortion of the feature space. In addition, the hard sample sampling strategy is vulnerable to pseudo-label errors in a noisy environment, and existing methods (such as confidence filtering, multi-perturbation models) are difficult to balance sample efficiency and noise suppression. Therefore, a hard sample noise-tolerant contrast semi-supervised semantic segmentation method is proposed. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a hard sample noise-tolerant contrast semi-supervised semantic segmentation method to solve the problems in the background art.
[0005] To achieve the above object, the present invention provides the following technical solutions: A hard sample noise-tolerant contrast semi-supervised semantic segmentation method, including the following steps:
[0006] Step 1: Construct a student model (S) and a teacher model (T), and the parameters of the teacher model are updated by exponential moving average (EMA), and the formula is:
[0007]
[0008] where α ∈ [0.95, 0.999] is the decay factor, θ s is the current iteration parameter of the student model, θ t is the historical parameter of the teacher model, k is the number of training iterations, and initially
[0009] Step 2: Apply weak augmentation T labeled to the labeled data D unlabeled and strong augmentation T w (I) and strong augmentation T s (I) to the unlabeled data D respectively, and the teacher model generates pseudo-labels based on the strongly augmented data
[0010] Step 3: Dynamically screen pseudo-labels with a dynamic threshold, and the dynamic threshold formula is:
[0011]
[0012] where τ0 ∈ [0.9, 0.95] is the initial threshold, τ min ∈ [0.5, 0.7] is the minimum threshold, K is the total number of training iterations, and the screening condition is the confidence of the pseudo-label:
[0013]
[0014] Pseudo-labels that do not meet the standard are discarded.
[0015] Step 4: Design a noise-tolerant contrastive loss function, which specifically includes:
[0016] For each sample i, calculate its feature vector z i and the TopN similarity with the positive sample set to screen out the top N = 5 positive samples with the highest similarity
[0017] The contrastive loss formula is:
[0018]
[0019] where ε(·, ·) represents the similarity between sample pairs, is the negative sample set of the sample, and n is the total number of samples, including all features outside the current batch and the samples of the same category in the memory bank.
[0020] Step 5: Extract boundary hard sample features based on morphological operations, which specifically includes:
[0021] Perform dilation and erosion operations on the class map predicted by the teacher model, with the neighborhood being in the range of 5×5 pixels, defined as:
[0022]
[0023] Hard sample mask Easy sample mask
[0024] Uniformly sample 50 feature points from each of the and regions in each batch and store them in the memory bank (with a capacity of MaxN = 2000 for each category). When sampling, high-confidence features are preferentially retained.
[0025] Preferably, the encoders of the student model and the teacher model are ResNet-101 or VisionTransformer-Base, the decoders are U-Net++ or DeepLabv3+, and the projection head is a two-layer MLP (hidden layer dimension 512, output dimension 128).
[0026] Preferably, for the labeled dataset Its cross-entropy loss is:
[0027]
[0028] where C is the number of classes, is the label of sample i on class c, is the class probability predicted by the student model.
[0029] For the unlabeled dataset Its pseudo-label loss is:
[0030]
[0031] where is the pseudo-label generated by the teacher model.
[0032] Preferably, after adding the contrastive loss, the total loss function is the weighted sum of the cross-entropy loss L labeled of the labeled data, the pseudo-label loss L unlabeled and the contrastive loss :
[0033] Preferably, the memory bank is updated using a first-in-first-out strategy, where new features replace the oldest features with the lowest confidence in the bank, and replacement is triggered when the number of features per class exceeds MaxN = 2000.
[0034] Preferably, the initial value of the dynamic threshold τ0 is 0.92, the minimum threshold τ min is 0.6, and the total number of training epochs K = 100,000.
[0035] Preferably, the student model is trained using strongly augmented data I s and weakly augmented data I W , and the teacher model only receives strongly augmented data to generate pseudo-labels.
[0036] Preferably, during hard sample sampling, 50 feature points are taken from each of and per batch, and the confidence of the hard sample region needs to satisfy
[0037] Preferably, the dilation and erosion of the morphological operation are performed with 3 iterations, and the structural element is a cross-shaped kernel (kernelsize = 3).
[0038] Preferably, the AdamW optimizer is used during training, with an initial learning rate of 3×10 -4 , a weight decay of 0.05, and batch sizes of 16 for labeled data and 32 for unlabeled data.
[0039] Compared with the prior art, the present invention has the following beneficial effects:
[0040] Through dynamic threshold screening of pseudo-labels, noise-tolerant contrast loss, and boundary-aware hard sample sampling, the present invention effectively solves the problems of pseudo-label noise interference, contrast learning feature space distortion, and low hard sample sampling efficiency in traditional semi-supervised semantic segmentation. The dynamic threshold gradually releases high-quality pseudo-labels, reducing the negative impact of noise labels on training; the improved contrast loss selects positive samples through TopN similarity, suppressing the feature space distortion caused by pseudo-label noise; the hard sample sampling strategy based on morphological operations accurately locates the boundary uncertain regions, improving the learning efficiency of the model for difficult samples. This method significantly enhances the robustness and generalization ability of the model in a noisy environment and is applicable to low-labeling-cost scenarios such as medical imaging and autonomous driving.
[0041] Other features and advantages of the present invention will be described in the following specification, and some will be obvious from the specification or understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is a flowchart of the hard sample noise-tolerant contrast semi-supervised semantic segmentation method based on the present invention;
[0043] Figure 2 is a relationship diagram of the teacher model and the student model of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0044] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art in the technical field of the present invention without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0045] Please refer to Figure 1 and Figure 2 , the hard sample noise-tolerant contrast semi-supervised semantic segmentation method in the present invention specifically includes the following parts.
[0046] 1. Model Initialization and Architecture
[0047] Architectures of the student model (S) and the teacher model (T):
[0048] Encoder: ResNet-101 is used as the backbone network, and the output feature map has a dimension of d = 2048d = 2048, and the feature map resolution where h×w is the resolution of the input image.
[0049] Decoder: The DeepLabv3+ structure is used, and the output classification map has a size of (1, H, W)(1, H, W), and each pixel corresponds to a class label.
[0050] Projection head: A two-layer MLP (input dimension 2048, hidden layer 512, output dimension 128), which is used to generate feature vectors in contrastive learning
[0051] Parameter initialization: The parameters θ of the student model s are randomly initialized, and the parameters θ of the teacher model t , are initialized as
[0052]
[0053] 2. Data Augmentation Strategies
[0054] Weak augmentation T w :
[0055] Random horizontal flipping (probability 0.5);
[0056] Randomly crop to 90%-100% of the original image area;
[0057] Slight color jitter (brightness adjustment range ±0.1, contrast adjustment range ±0.1).
[0058] Strong augmentation T s :
[0059] Random rotation (angle range ±30°);
[0060] Randomly crop to 50%-90% of the original image area;
[0061] Color distortion (brightness adjustment range ±0.5, contrast adjustment range ±0.5, saturation adjustment range ±0.5);
[0062] Gaussian noise (standard deviation σ = 0.2).
[0063] 3. Training Process
[0064] Step 1: Data Augmentation and Pseudo-Label Generation
[0065] For the labeled data and the unlabeled data apply weak augmentation T w and strong augmentation T s respectively:
[0066] Weakly augmented data: I w = T w (I);
[0067] Strongly augmented data: I s = T s (I).
[0068] The teacher model generates pseudo-labels based on the strongly augmented data:
[0069]
[0070] where f0 is the classification head of the teacher model.
[0071] Step 2: Dynamically threshold the pseudo-labels
[0072] Calculate the dynamic threshold for the current iteration k:
[0073]
[0074] Filtering condition: Only keep the confidence
[0075] For the pseudo-labels that do not meet the condition, directly discard them and do not participate in training.
[0076] Step 3: Calculate the loss function
[0077] Cross-entropy loss of the labeled data:
[0078]
[0079] where C is the number of classes, is the label of sample i on class c, is the class probability predicted by the student model.
[0080] Pseudo-label loss:
[0081]
[0082] where, is the pseudo-label generated by the teacher model.
[0083] Noise-tolerant contrastive loss:
[0084] For the feature z of each sample i i , calculate its similarity with the positive sample set P iTop5 similarity to screen out the most relevant positive samples
[0085] Contrastive loss formula:
[0086]
[0087] where ε(·,·) represents the similarity between sample pairs, is the negative sample set of the sample, n is the total number of samples, including all features outside the current batch and similar samples in the memory bank.
[0088] Total loss function:
[0089]
[0090] where λ L = 1.0, λ c = 0.3.
[0091] Step 4: Hard sample sampling and memory bank update
[0092] Extract hard samples by morphological operations:
[0093] Perform dilation and erosion on the class map predicted by the teacher model operation:
[0094]
[0095] Hard sample mask Easy sample mask Sampling strategy:
[0096] Uniformly sample 50 feature points (confidence ≥ 0.4) from each of and regions in each batch; store them in the memory bank (upper limit of 2000 for each class), and replace low-confidence features using the first-in-first-out strategy.
[0097] Step 5: Parameter update
[0098] The student model optimizes the parameters through backpropagation:
[0099]
[0100] The teacher model is updated by EMA:
[0101] 4. Training configuration
[0102] Optimizer: AdamW, initial learning rate 3×10 -4 , weight decay 0.05;
[0103] Batch size: 16 for labeled data and 32 for unlabeled data;
[0104] Number of training epochs: The total number of iterations K = 100,000;
[0105] Hardware: 8 × NVIDIA A100 GPUs, with a batch size of 4 per card.
[0106] 5. Performance Verification
[0107] Dataset: PASCAL VOC 2012 (1,464 labeled data and 9,118 unlabeled data);
[0108] Evaluation metric: mIoU (mean intersection over union);
[0109] Results:
[0110] Baseline model (without contrastive loss and hard sample mining): 72.1% mIoU;
[0111] Our method: 78.6% mIoU, an improvement of 6.5%;
[0112] The segmentation accuracy in the boundary region is significantly improved (+9.2%), verifying the effectiveness of hard sample mining.
[0113] 6. Algorithm Flow
[0114] The entire training process can be divided into the following steps:
[0115] 1. Initialization: Initialize the student model S and the teacher model T, and simultaneously initialize the encoder f, decoder d, projection head p, and weak augmentation operation Strong augmentation operation
[0116] 2. Weak and strong augmentations: Apply the weak augmentation operation to the labeled data X l and the unlabeled data X u respectively to obtain the augmented datasets. And further apply the strong augmentation operation to the unlabeled data X to get u the enhanced dataset.
[0117] 3. Feature extraction and pseudo-label generation:
[0118] Apply the student model to the labeled data to obtain its features F l and predicted labels
[0119] Apply the teacher model to the unlabeled data to obtain its features F u and pseudo-labels
[0120] 4. Dynamic threshold screening:
[0121] Use the simple semi-hard method SH to screen the features of the labeled data to obtain the final feature vector Z l and the screened predicted labels
[0122] Use the diversified sampling method DS to screen the unlabeled data to obtain the final feature vector Z u and the screened pseudo-labels
[0123] 5. Calculate the loss: Calculate the contrastive loss L of the student model c , including the loss of the labeled data and the loss of the pseudo-labels
[0124] 6. Backpropagation: Optimize the parameters of the student model S through backpropagation to minimize the contrastive loss.
[0125] 7. EMA update: Update the parameters of the teacher model T through the exponential moving average (EMA) rule.
[0126] 8. Iteration: Repeat the above steps until the predetermined number of training epochs or convergence conditions are reached.
[0127] After each iteration, the memory bank B of the teacher model is also updated t , and the memory bank is updated according to the new features output by the current student model and the pseudo-label features.
[0128] This embodiment proposes a hard sample noise-tolerant contrastive semi-supervised semantic segmentation method. By constructing a student-teacher model framework, combining dynamic threshold screening of pseudo-labels, noise-tolerant contrastive loss, and morphological hard sample sampling, the segmentation performance of the model in a noisy environment is significantly improved. The teacher model generates high-quality pseudo-labels through EMA update, the dynamic threshold gradually releases reliable samples, the improved contrastive loss suppresses noise propagation, and at the same time, dilation and erosion operations are used to accurately locate boundary hard samples, enhancing the discriminative ability of the model for complex regions.
Claims
1. A semi-supervised semantic segmentation method based on hard sample noise tolerance contrast, characterized in that: The following steps are involved: Step 1: Construct the student model (S) and the teacher model (T). The teacher model parameters are updated by exponential moving average (EMA), and the formula is: Among them, α∈[0.95,0.999] is the attenuation factor, θ s is the current iteration parameter of the student model, θ t is the historical parameter of the teacher model, k is the number of training iterations, and the initial Step 2: Label the data D labeled With unlabeled data D unlabeled Apply weak enhancement T w (I) With strong enhancement T s (I), the teacher model generates pseudo labels based on strongly augmented data Step 3: Dynamic threshold filter pseudo-labels. The dynamic threshold formula is: Among them, τ0∈[0.9,0.95] is the initial threshold, τ min ∈[0.5,0.7] is the minimum threshold, K is the total number of training iterations, and the screening condition is the pseudo-label confidence: Unsatisfactory pseudo labels are discarded. Step 4: Design a noise-tolerant contrast loss function, including: For each sample i, calculate its feature vector z i With the positive sample set The TopN similarity of the filter is used to select the top N = 5 positive samples with the highest similarity. The contrast loss formula is: Among them, ε(·,·) represents the similarity between sample pairs, is the negative sample set of the sample, n is the total number of samples, including all features other than the same samples in the current batch and memory library. Step 5: Extract boundary hard sample features based on morphological operations, including: Class map of teacher model predictions Execution Bloat Corrosion Operation, Neighborhood is a 5×5 pixel area, defined as: Hard sample mask Easy Sample Mask Each batch from and 50 feature points are uniformly sampled in each region and stored in the memory bank (capacity MaxN = 2000 per category), and high confidence features are prioritized during sampling.
2. The method for semi-supervised semantic segmentation based on hard sample noise tolerance contrast according to claim 1 is characterized in that: The encoder of the student model and the teacher model is ResNet-101 or Vision Transformer-Base, the decoder is U-Net++ or DeepLabv3+, and the projection head is a two-layer MLP (hidden layer dimension 512, output dimension 128).
3. The method for semi-supervised semantic segmentation based on hard sample noise tolerance contrast according to claim 1, characterized in that: For the labeled dataset Its cross entropy loss is: Where C is the number of categories, is the label of sample i in category c, is the class probability predicted by the student model. For unlabeled datasets Its pseudo label loss is: in, are pseudo labels generated by the teacher model.
4. The method for semi-supervised semantic segmentation based on hard sample noise tolerance contrast according to claim 1, characterized in that: After adding the contrast loss, the total loss function is the cross entropy loss L of the labeled data labeled , pseudo label loss L unlabeled and contrast loss The weighted sum of:
5. The method for semi-supervised semantic segmentation based on hard sample noise tolerance contrast according to claim 1, characterized in that: The memory library is updated using a first-in-first-out strategy. The new feature replaces the old feature with the lowest confidence in the library, and the replacement is triggered when the number of features of each type exceeds MaxN=2000.
6. The method for semi-supervised semantic segmentation based on hard sample noise tolerance contrast according to claim 1, characterized in that: The initial value of the dynamic threshold τ0 is 0.92, and the minimum threshold τ min is 0.6, and the total number of training rounds K = 100,000.
7. The method for semi-supervised semantic segmentation based on hard sample noise tolerance contrast according to claim 1, characterized in that: The student model uses strongly enhanced data I s With weakly enhanced data I W For training, the teacher model only receives strongly augmented data to generate pseudo labels.
8. The method for semi-supervised semantic segmentation based on hard sample noise tolerance contrast according to claim 1, characterized in that: When hard samples are collected, each batch is and Take 50 feature points each, and the confidence of the hard sample area must meet 9. The method for semi-supervised semantic segmentation based on hard sample noise tolerance contrast according to claim 1, characterized in that: The dilation and erosion of the morphological operations were performed three times, and the structure element was a cross-shaped kernel (kernel size = 3).
10. The method for semi-supervised semantic segmentation based on hard sample noise tolerance contrast according to claim 1, characterized in that: The AdamW optimizer was used for training, with an initial learning rate of 3×10 -4 , weight decay 0.05, batch size 16 labeled data, unlabeled data 32.
Citation Information
Cited By
Pulmonary nodule semi-supervised segmentation method and device based on memory enhancement and spatio-temporal topological constraint, equipment and medium
CN122435275A