Dynamic Enhancement Method for Semi-Supervised Semantic Segmentation Based on Adaptive Spatial Transformation
Through the semi-supervised semantic segmentation method of adaptive spatial transformation and information entropy calculation, the problem of segmentation inaccuracy under complex spatial transformation in the prior art is solved, and high-precision and stable semantic segmentation effect are achieved.
Patent Information
- Application Number
- CN202510465927.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The existing semantic segmentation methods are difficult to ensure the precise segmentation of the target area when dealing with complex spatial transformations, resulting in insufficient stability and robustness of the model.
The semi-supervised semantic segmentation dynamic enhancement method of adaptive spatial transformation is adopted. By dynamically adjusting the amplitude and direction of the spatial transformation, combined with the pseudo-label and information entropy calculation of the teacher model, the generalization ability and robustness of the model are enhanced.
It significantly improves the accuracy and stability of semantic segmentation, can effectively handle complex spatial transformation conditions, and improves the performance of the model in open-world applications.
Smart Images

Figure CN119992107B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of semantic segmentation, and more particularly, to a dynamic enhancement method for semi-supervised semantic segmentation based on adaptive spatial transformation. Background Art
[0002] Semantic segmentation is a traditional and important field in computer vision, which focuses on classifying all pixels in an image according to their semantic content. In recent years, many successful studies have emerged in this field, with different applications in specific fields, such as natural images, medical images, remote sensing image analysis, driving scene segmentation, point cloud segmentation, etc. The integration of deep learning has greatly improved the efficiency of semantic segmentation tasks. However, deep learning requires a large amount of labeled data to train an effective model. Although it is now easy to obtain a large amount of raw data, manually labeling image pixels is a difficult and time-consuming task that requires a huge amount of manpower and time. Taking the Cityscape dataset as an example, the annotation and quality control of a single image on average take more than 1.5 hours. Therefore, effectively training a model with insufficient annotations for open-world applications is a major challenge. To address this issue, researchers have explored integrated methods such as weakly supervised learning or unsupervised learning. Semi-supervised learning offers advantages by leveraging both labeled and unlabeled datasets. Semi-supervised learning mainly focuses on training models with a limited number of labeled data as well as a large amount of unlabeled data.
[0003] Semi-supervised semantic segmentation aims to develop a model that can effectively utilize limited labeled data while extracting valuable insights from a large amount of unlabeled data, providing excellent performance and alleviating the challenges faced by researchers with limited labeled data. It can generally be divided into pseudo-labeling methods, consistency regularization, contrastive learning, adversarial training, and hybrid methods. In the early stage of semi-supervised semantic segmentation research, the general framework of generative adversarial networks (GANs) was more commonly used, which can be mainly divided into two types according to its structure: those with a generator and those without a generator. The purpose of contrastive learning is to improve the performance of semantic segmentation by learning more useful representations in the embedding space. Consistency regularization (CR) reaches a consensus on the smoothness assumption, hoping that the model gives similar predictions for different perturbations or variants of the same input to improve the segmentation quality. Pseudo-labeling follows a simple pipeline, using the model trained with labeled data to predict the pseudo-labels of unlabeled data, which are then augmented into the dataset to retrain the model. Currently, existing data augmentation research is mostly based on gray-level enhancement, and spatial transformation methods lack discussion. However, pixel location is a key factor for the segmentation task because the segmentation model needs to accurately identify the target area under different conditions.
[0004] Most existing enhancement techniques rely on intensity transformations such as grayscale enhancement. However, semantic segmentation tasks require high pixel-level spatial accuracy, especially when dealing with large-scale spatial transformations such as rotation and displacement. Existing enhancement methods often struggle to ensure accurate segmentation of the target area. Therefore, current technologies have difficulty effectively processing data with complex spatial transformations, which poses challenges to the stability and robustness of the model. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a semi-supervised semantic segmentation dynamic enhancement method based on adaptive spatial transformation, which enhances the generalization of the model by dynamically adjusting the amplitude and direction of the spatial transformation and achieves high-precision semantic segmentation.
[0006] The present invention provides a semi-supervised semantic segmentation dynamic enhancement method based on adaptive spatial transformation, including the following steps:
[0007] S1. Given a set of labeled datasets and a set of unlabeled datasets , where includes M labeled images, represents a labeled image, represents a labeled image corresponding label, includes N unlabeled images, represents an unlabeled image, where N >> M;
[0008] S2. Construct a first loss function , and train a student model through the labeled datasets and the first loss function ;
[0009] S3. Initialize a teacher model through the student model , and use the teacher model to predict the unlabeled datasets to obtain a prediction result , and use the prediction result as a pseudo-label, where represents weak enhancement processing;
[0010] S4. Calculate the information entropy of the unlabeled datasets according to the pseudo-label ;
[0011] S5. Perform spatial enhancement processing on the unlabeled images according to the information entropy to obtain unlabeled images Spatial enhancement result , use the student model to predict the spatial enhancement result and obtain the prediction result ;
[0012] S6. According to the pseudo-label and the prediction result construct the second loss function , and according to the first loss function and the second loss function construct the total loss function , through the total loss function , further train the student model to obtain multiple trained student models ;
[0013] S7. Given a set of test set images, each image in the test set images includes a true label. Use the trained student model to predict the test set images, calculate the mean intersection over union between the prediction result and the true label and select a set of student models with the best training effect according to the mean intersection over union .
[0014] In a possible implementation, in step S1, the labeled dataset is the Pascal VOC 2012 dataset, and the unlabeled dataset is the Cityscapes dataset.
[0015] In a possible implementation, the initial learning rate of the PASCAL VOC 2012 dataset is set to 0.001, and the initial learning rate of the Cityscapes dataset is 0.005.
[0016] In a possible implementation, in step S2, the expression of the first loss function is as follows:
[0017] ,
[0018] where M represents the number of labeled images , represents the weight parameter of the student model , c represents the label category represents the probability of judging that the labeled image belongs to the label category c during prediction, C represents the number of label categories c represents the labeled image The basic true value, Indicates weak enhancement processing.
[0019] In one possible implementation, in step S4, the information entropy is calculated by the following formula :
[0020] ,
[0021] where, Indicates the unlabeled image Belongs to the probability of label category c.
[0022] In one possible implementation, in step S5, spatial enhancement processing is performed by the following operation:
[0023] ,
[0024] where, and Both represent scaling parameters, Represents the offset of the rotation transformation, Represents the offset of the translation transformation, Represents the maximum rotation angle, Represents the maximum translation ratio.
[0025] In one possible implementation, = 180°, = 0.5.
[0026] In one possible implementation, = 0.5, = 0.5, = 5.5, = 3.
[0027] In one possible implementation, in step S6, calculating the second loss function Includes the following steps:
[0028] S61. Obtain the second loss function through the following operation :
[0029] ,
[0030] where, H and W respectively represent the unlabeled image Height and width of, Represents the set of results obtained by two of the spatial enhancement functions And ;
[0031] S62. Obtain the total loss function through the operation of the following formula :
[0032] ,
[0033] wherein, represents a hyperparameter.
[0034] In a possible implementation manner, in step S7, calculate the mean intersection over union through the operation of the following formula :
[0035] ,
[0036] wherein, TP represents the samples predicted as the positive class region, FN represents the samples predicted as the negative class region but actually being the positive class region, FP represents the samples predicted as the positive class region but actually being the negative class region, n represents the number of categories of the labels, represents the intersection over union, represents the intersection over union of the i-th category.
[0037] Compared with the prior art, the present application has the following advantages: Using the teacher model to predict the weakly augmented images, and measuring the uncertainty of the prediction by calculating the information entropy so as to dynamically adjust the augmentation intensity, making the augmentation strategy more adaptive and accurate during the training process. At the same time, the intensity of the augmentation operation gradually increases as the training progresses, thereby realizing a spatial transformation from simple to complex during the augmentation process, enabling the model to be trained stably and effectively under the condition of spatial transformation. This method effectively avoids the label misalignment problem caused by geometric transformation by ensuring the spatial consistency of the images and prediction results before and after augmentation, improves the robustness and generalization ability of the model in complex input data, and finally significantly improves the semantic segmentation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a schematic diagram of the method of the present invention;
[0039] Figure 2 is a visualization diagram of the effect of the present invention;
[0040] Figure 3 is a visualization diagram of the effect of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] First of all, those skilled in the art should understand that these embodiments are only used to explain the technical principles of the embodiments of the present application and are not intended to limit the protection scope of the embodiments of the present application. Those skilled in the art can adjust them as needed to adapt to specific application scenarios.
[0042] The present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0043] This embodiment studies the impact of this method on two benchmark segmentation datasets, namely Pascal VOC 2012 and Cityscapes. The Pascal VOC 2012 dataset contains 21 semantic categories and is divided into classic and mixed subsets. The classic training set consists of 1464 images with extensive labels, and 1449 images are used for validation. In addition, a new version of the mixed training set includes low-resolution, roughly annotated images from the Segmentation Boundaries Dataset (SBD), thus expanding the training pool to 10,582 images. The Cityscapes dataset covers 19 semantic categories in an urban context, providing 2975 precisely annotated training images and 500 images for validation.
[0044] For model construction, the architecture proposed in this embodiment is a pluggable one, so the effectiveness of this method is evaluated on the CNN and ViT structures respectively. Specifically, for the model based on the convolutional neural network (CNN) architecture, DeepLabV3+ is selected as the semantic segmentation model, and ResNet-101 is used as the backbone network. ResNet-101 has been pre-trained on the ImageNet dataset and thus has good feature extraction capabilities. For the model based on the Vision Transformer (ViT), SegFormer-B5 is selected as the semantic segmentation model. It is a semantic segmentation model based on the ViT architecture and has been pre-trained on a large-scale dataset, capable of effectively handling complex visual tasks. To evaluate the performance, this embodiment uses SGD as the optimizer and adopts a polynomial decay learning rate strategy, and the learning rate is expressed as follows:
[0045] ,
[0046] where represents the initial learning rate, represents the current iteration number, represents the total number of iterations, power represents a hyperparameter, and power and weight decay are set to 0.9 and 1e-4 respectively. By combining the SGD optimizer with the learning rate decay strategy, the convergence during the training process is ensured, avoiding training oscillations caused by too high a learning rate, ensuring sufficient training time, and enabling the model to be fully optimized in the direction of the minimum loss.
[0047] For the PASCAL VOC 2012 dataset, the initial learning rate was set to 0.001, the cropped image size was 321×321 or 513×513, the batch size was 16, and it was trained for 80 epochs. For the Cityscapes dataset, the initial learning rate was 0.005, the cropped size used was 801×801, the batch size was also 16, and it was trained for 240 epochs. All training and inference processes used the PyTorch deep learning framework and were executed on 4×NVIDIA V100 GPUs and 2×A800 GPUs for different numerical analyses.
[0048] This embodiment discloses a dynamic enhancement method for semi-supervised semantic segmentation based on adaptive spatial transformation, including the following steps:
[0049] S1. Given a set of labeled datasets and a set of unlabeled datasets , where includes M labeled images, denotes the labeled image, denotes the labeled image corresponding label, includes N unlabeled images, denotes the unlabeled image, where N >> M;
[0050] S2. Construct the first loss function and train the student model through the labeled dataset and the first loss function .
[0051] In step S2, the expression of the first loss function is as follows:
[0052] ,
[0053] where M represents the number of labeled images , denotes the weight parameters of the student model , c represents the label category, denotes the probability of judging that the labeled image belongs to the label category c during prediction, C represents the number of label categories c, denotes the ground truth of the labeled image , Indicates weak augmentation processing. Weak augmentation usually includes mild image transformations such as random cropping, random flipping, etc. The transformation amplitude of weak augmentation is small, aiming to make the model adapt to small changes while maintaining the original features. Weak augmentation processing is prior art and will not be elaborated in this application.
[0054] Initialize the teacher model through the student model Initialize the teacher model , and through the teacher model Predict the unlabeled dataset to obtain the prediction results , and use the prediction results as pseudo-labels, where indicates weak augmentation processing.
[0055] Specifically, in this embodiment, extract the weight parameters from the student model and load them into the teacher model to ensure that the teacher model can quickly start predicting the unlabeled dataset . In this way, the initial state of the teacher model can ensure a certain accuracy in predicting the unlabeled dataset . Then, use the teacher model to predict the unlabeled dataset . The teacher model outputs the predicted probability distribution of each pixel in each unlabeled image belonging to each category. The prediction results of the teacher model are used as pseudo-labels, and these pseudo-labels will be used to guide the student model to learn the unlabeled dataset in the subsequent training process. The pseudo-labels provide an approximate "true label". Although these labels are not perfect, they provide sufficient information for the student model to perform effective learning.
[0056] Calculate the information entropy of the unlabeled dataset according to the pseudo-labels . .
[0057] In step S4, calculate the information entropy through the following formula :
[0058] ,
[0059] where represents the probability that the unlabeled image belongs to the label category c.
[0060] Calculate the information entropy which is used to provide useful guidance for subsequent training. For each unlabeled image , the probability distribution of the sample belonging to each category is predicted through the teacher model , and then the information entropy of the unlabeled image is obtained . The calculated information entropy value is used to evaluate the "learning value" of each unlabeled image , that is, the entropy value will affect its priority during the training process. The larger the information entropy , it means that the model has a higher prediction uncertainty for this sample, that is, the model is not very sure which category this sample belongs to. Such samples have higher "learning value" because they can help the model explore more feature spaces, so higher weights will be assigned to them, and more spatial transformations and augmentations will be applied to them during training. On the contrary, samples with lower information entropy mean that the model's prediction for this sample is relatively certain, the model's learning need for it is lower, and the transformation during the learning process will be relatively small, thus maintaining the consistency of its features.
[0061] To fine-tune the level of augmented distortion, an entropy-based adaptive weight for spatial transformation, namely EAW, is introduced in this embodiment. The information entropy of the unlabeled image is used as a measure of sample reliability. The calculation formula of the adaptive weight is as follows:
[0062] ,
[0063] where k and d are both adjustable parameters, which control the amplitude and offset of the augmentation respectively. In this embodiment, the specific calculation formula obtained by modifying the sigmoid function is , with an offset d added, and the information entropy value is input in a fine-tuning manner for optimization, and differential augmentation is performed among different samples to more effectively explore the transformed space. Because samples with high entropy show greater uncertainty and require more significant spatial transformations to explore a wider feature space. On the contrary, for samples with low entropy, the model shows higher certainty and is more suitable for smaller transformations, thus retaining stable features. This method can more smoothly adjust the augmentation between samples and enhance the model's ability and flexibility for spatial transformation.
[0064] According to the information entropy , the unlabeled image is subjected to spatial augmentation processing to obtain the spatial augmentation result of the unlabeled image , and the student model is used Perform prediction on the spatial enhancement result to obtain a prediction result .
[0065] In step S5, spatial enhancement processing is performed through the operation of the following formula:
[0066] ,
[0067] where and both represent scaling parameters, represents the offset of the rotation transformation, represents the offset of the translation transformation, represents the maximum rotation angle, represents the maximum translation ratio.
[0068] Apply spatial enhancement according to the information entropy value of each sample. The information entropy , as an indicator for measuring the uncertainty of the sample, can help determine the intensity of data enhancement so as to more effectively improve the performance of the model. A high information entropy value indicates that the model has a relatively high prediction uncertainty for it. Usually, the features of these samples are complex and may involve more variations. Therefore, for high-entropy samples, stronger spatial transformations, such as larger rotations, translations, etc., need to be applied to explore a wider feature space. This can expand feature diversity through transformations and help the model learn more possible feature patterns. While low information entropy data indicates that the model is relatively certain about its prediction, the features of the sample are relatively stable and the learning difficulty is lower. Therefore, for low-entropy samples, smaller spatial transformations need to be applied to preserve their original features and avoid introducing too many unnecessary changes. In this embodiment, spatial enhancement is used instead of the traditional intensity-focused robust enhancement method. The traditional method may only focus on enhancing specific regions of the sample, while in this embodiment, by dynamically adjusting the spatial transformation, high-entropy samples are subjected to a wider range of spatial transformations, and low-entropy samples maintain stability. In order to dynamically determine the intensity of the applied spatial transformation enhancement according to the entropy value, a mapping function for the spatial enhancement processing of the above formula is designed, and the amplitudes of rotation and translation are determined through this mapping function.
[0069] When the information entropy value is relatively small, the mapping result will also decrease, which will reduce the amplitudes of rotation and translation, thereby enhancing stability. To ensure the effectiveness of the enhancement process, this embodiment conducts detailed tests during implementation and sets the following specific parameter settings: = 180°, = 0.5, That is, the maximum translation ratio is 50% of the size of the original image. These parameters control the intensity of the spatial transformation. Especially in high-entropy samples, larger rotations and translations help to explore the feature space more extensively. For the PASCAL VOC 2012 dataset, = 1, = 1, = 11, = 7 are set to adjust the transformation intensity of rotation and translation. Specifically, the amplitudes of rotation and displacement are larger to enhance the adaptability of the model to diverse scenarios. For the Cityscapes dataset, according to its characteristics, in this embodiment, some parameters are adjusted to = 0.5, = 0.5, = 5.5, = 3 so as to apply a smaller transformation amplitude in more complex urban scenes. This adjustment ensures that the characteristics of the samples are not overly changed in complex environments, maintaining higher stability.
[0070] When the information entropy value is small, the mapping result will decrease, which will reduce the amplitudes of rotation and translation, thereby maintaining the stability of the applied augmentation. In this way, it is ensured that when dealing with low-entropy samples, the transformation amplitude will not be too large to avoid over-distorting the original features. For high-entropy samples, the amplitudes of rotation and translation increase, which can better expand the feature space and help the model learn more transformation patterns from uncertainties.
[0071] S6. Construct a second loss function based on the pseudo-label and the prediction result , and construct a total loss function based on the first loss function and the second loss function . Further train the student model through the total loss function to obtain multiple trained student models .
[0072] Calculating the second loss function in step S6 includes the following steps:
[0073] S61. Obtain the second loss function through the operation of the following formula:
[0074] ,
[0075] where H and W respectively represent the height and width of the unlabeled image , Denote two of the said spatial enhancement functions and the set of the obtained results;
[0076] S62. Obtain the said total loss function through the operation of the following formula :
[0077] ,
[0078] wherein, denotes a hyperparameter. Wherein, that is, the loss weight, which is set to 0.5 in this embodiment. This weight is used to balance the proportion of the loss of the labeled images and the unlabeled images in the total loss. Reasonably adjusting λ can control the influence of the unlabeled images during the training process. Use the total loss function to train the student model again . During the training process, apply spatial transformation enhancement to the unlabeled images . These enhanced unlabeled images and the corresponding pseudo-labels are used as inputs and passed to the student model for training. The pseudo-labels are generated by the prediction of the teacher model and are used to replace the true labels. This embodiment adopts the mean square error, that is, MSE. The mean square error can effectively measure the accuracy of the model prediction and align the spatially transformed images, calculate the mean square error of the two parts of the pseudo-labels and the prediction results , sum the squares of the prediction value differences at each pixel position , and then take the average to obtain the pixel-level mean square error of the image. The mean square error can directly measure the pixel-level difference between the spatially transformed image and the pseudo-label, and is insensitive to the mask structure, which means that even if there are some local area occlusions or noises in the image, it will not overly affect the entire loss calculation, enhancing the robustness of the model to spatial transformation. The above-mentioned total loss function consists of two parts: represents the loss of the labeled data , represents the loss of the unlabeled data .
[0079] Given a set of test set images, each image in the test set images includes the true label. Through the trained student model predict the test set images, and calculate the mean intersection over union , according to the average intersection over union Select a group of student models with the best training effect .
[0080] In step S7, the average intersection over union is calculated through the operation of the following formula :
[0081] ,
[0082] where TP represents the samples predicted as positive class regions, FN represents the samples predicted as negative class regions but actually positive class regions, FP represents the samples predicted as positive class regions but actually negative class regions, n represents the number of categories of the labels, represents the intersection over union, represents the intersection over union of the i-th category.
[0083] In this embodiment, the trained student model is applied to the test set images to evaluate its performance in actual tasks. By inputting the test set data, extracting image features, and using the average intersection over union as the performance evaluation index, this index is effective even when dealing with imbalanced classes that often appear in pixel-level annotation tasks.
[0084] First, input the test set images into the trained student model for inference. Through the forward propagation process of the student model , the model will generate the prediction results of each test image. The model outputs the predicted classes for different regions of the image and classifies each pixel point according to the output of the model to obtain the predicted class of each pixel. Through these outputs, the features of each test image are extracted. Using the average intersection over union as the evaluation index to evaluate the trained student model . is the intersection over union of each category, which measures the prediction accuracy of category C, that is, the ratio of the intersection to the union of the predicted region and the true label region. By calculating the intersection over union of each category and then taking the average of the intersection over union of all categories, the performance of the entire model on all categories is obtained. This index can effectively evaluate the performance of the model in a multi-category environment, especially in the case of class imbalance.
[0085] To evaluate the performance of the model at a fine-grained level, this embodiment adopts a sliding window evaluation method, especially on complex datasets such as the Cityscapes dataset in urban landscapes. The sliding window evaluation method divides the image into small blocks, or windows, and calculates , and gradually slide the window to cover the entire image. This can help obtain the detailed performance of local regions, especially in cases where small objects or complex scenes are involved. During implementation, this embodiment sets appropriate window sizes and sliding strides according to the resolution and computing power of the image. Smaller windows can provide more refined evaluations, but with a larger computational cost. Appropriate stride selection can balance evaluation accuracy and computational efficiency.
[0086] By calculating , the performance of the model in each category can be accurately evaluated, especially the segmentation effect on small-category objects. It is particularly effective for dealing with class imbalance problems because it takes into account the segmentation accuracy of each class. The sliding evaluation method further enhances the adaptability of model evaluation and can provide more detailed local region performance evaluations, especially suitable for complex urban landscape datasets where the objects in the images may vary in size and shape.
[0087] Visualize the predicted pixel categories to obtain the semantic segmentation results. In Appendix Figure 2 and Appendix Figure 3 , the visualizations of this embodiment on the urban landscape dataset and the Pascal VOC 2012 dataset are respectively shown, and comparisons are made with the baseline method and other advanced methods. Each row represents a different image, showing the segmentation effects under different methods.
[0088] In the Pascal VOC 2012 dataset, from this comparison, it can be seen that the method of this embodiment performs outstandingly in object category recognition, boundary division, and detail feature retention, and is superior to the baseline method in most cases. The segmentation results are more accurate, especially in complex scenes such as those containing animals or people, where it can better maintain the integrity of the objects and improve the clarity of the segmented regions. The baseline method often has problems with blurred boundaries in complex scenes, especially in the detailed segmentation of people or animals. For example, it is more accurate in the contour segmentation of airplanes, animals, and people, and significantly reduces the interference of background noise.
[0089] In the urban landscape dataset, the method of this embodiment further demonstrates its superior segmentation performance under complex conditions. Although the baseline method can identify the main objects in the scene, there are deficiencies in detail processing, such as the segmentation of human contours and small objects, resulting in some unclear boundaries, especially in scenes with multiple objects or complex backgrounds. Compared with the baseline method, the method of this embodiment shows higher segmentation accuracy in all images. It has clearer boundaries between people, vehicles, and environmental elements such as roads and buildings, avoiding the common blur and overlap problems of the baseline method in complex scenes. For example, in the first row of visualization results, the method of this embodiment can accurately segment the contours of pedestrians, and the segmentation of traffic signs and the background is also clearer; the second row shows that the method of this embodiment can more accurately segment the bus and the road surface, and can clearly distinguish vehicles and road signs; the third and fourth rows show that the method of this embodiment makes the segmentation of buildings and roads more refined through more accurate boundary positioning, retaining more details. The method of this embodiment shows significant advantages in improving the model generalization ability, detail fidelity, and complex scene segmentation accuracy, providing a new solution for semi-supervised semantic segmentation tasks.
[0090] In addition, this embodiment also provides the following pseudocode for the method of this application:
[0091] Input: Labeled dataset , unlabeled dataset , student model , teacher model ;
[0092] Output: Trained student model ;
[0093] Step 1: Use the first loss function to train the student model on the labeled dataset ;
[0094] Step 2: Initialize the teacher model , student model ;
[0095] for in do
[0096] Use the teacher model to generate prediction results for the unlabeled images , and then obtain the information entropy ;
[0097] Use the information entropy to determine Apply the strong enhancement effect and obtain the spatial enhancement result ;
[0098] Use the student model to generate pseudo labels ; Calculate the second loss function using the predictions and pseudo labels and the total loss function to train the student model ;
[0099] end for
[0100] Return the student model 。
[0101] In the description of the present application, the descriptions referring to terms such as "one embodiment", "some embodiments", "in this embodiment", "specific examples", or "some examples", etc. mean that the specific features, mechanisms, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, mechanisms, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0102] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A semi-supervised semantic segmentation dynamic enhancement method based on adaptive spatial transformation, characterized in that, It includes the following steps: S1. Given a set of labeled data and a set of unlabeled data , where includes labeled images, denotes the labeled images, denotes the labeled images corresponding labels, includes unlabeled images, denotes the unlabeled images, where ; S2. Construct the first loss function , and use the labeled dataset and the first loss function to train the student model ; S3. Through the student model Initialize the teacher model , and through the teacher model Predict the unlabeled dataset to obtain the prediction results , and use the prediction results as pseudo-labels, where represents weak augmentation processing; S4. Calculate the information entropy of the unlabeled data set based on the pseudo label ; S5. According to the information entropy perform spatial enhancement processing on the unlabeled image to obtain the spatial enhancement result of the unlabeled image ; use the student model to predict the spatial enhancement result to obtain a prediction result ; the formula for the spatial enhancement processing is: ; ; ; Among them, and both represent scaling parameters, represents the offset of the rotation transformation, represents the offset of the translation transformation, represents the maximum rotation angle, represents the maximum translation ratio; S6. Construct a second loss function according to the pseudo-label and the prediction result , and construct a total loss function according to the first loss function and the second loss function ; further train the student model through the total loss function to obtain multiple trained student models ; ; S7. Given a set of test set images, each image in the test set images includes a ground truth label, and through the trained student model predict the test set images, and calculate the mean intersection over union between the prediction result and the ground truth label , according to the mean intersection over union select a group of the student models with the best training effect .
2. The semi-supervised semantic segmentation dynamic enhancement method based on adaptive spatial transformation according to claim 1, wherein In step S1, the labeled dataset is the Pascal VOC 2012 dataset, and the unlabeled dataset is the Cityscapes dataset.
3. The semi-supervised semantic segmentation dynamic enhancement method based on adaptive spatial transformation according to claim 2, characterized in that, The initial learning rate of the PASCAL VOC 2012 dataset is set to 0.001, and the initial learning rate of the Cityscapes dataset is set to 0.
005.
4. The semi-supervised semantic segmentation dynamic enhancement method based on adaptive spatial transformation according to claim 1, 2 or 3, characterized in that In step S2, the first loss function has the following expression: ; Among them, M represents the labeled image quantity of, represents the student model weight parameters of, represents the label category represents the probability of judging that the labeled image belongs to the label category when predicting, represents the quantity of the label category quantity of, represents the labeled image ground truth of, represents weak augmentation processing.
5. The semi-supervised semantic segmentation dynamic enhancement method based on adaptive spatial transformation according to claim 4, characterized in that, In step S4, the information entropy is obtained through the operation of the following formula :[[]]END]] ; Among them, represents the probability that the unlabeled image belongs to the label category.
6. The semi-supervised semantic segmentation dynamic enhancement method based on adaptive spatial transformation according to claim 1, wherein , 。 7. The semi-supervised semantic segmentation dynamic enhancement method based on adaptive spatial transformation according to claim 6, characterized in that, , , , 。 8. The semi-supervised semantic segmentation dynamic enhancement method based on adaptive spatial transformation according to claim 1, 2, 3, 5, 6 or 7, characterized in that, Calculating the second loss function in step S6 comprises the following steps: S61. Obtaining the second loss function through an operation of the following formula :[[]]END]] ; Among them, and respectively represent the height and width of the tagless image , represents the set of results obtained by the two spatial enhancement functions and . S62. Obtaining the total loss function through operations of the following formula :[[]]END]] ; Among them, represents a hyperparameter.
9. The semi-supervised semantic segmentation dynamic enhancement method based on adaptive spatial transformation according to claim 5, characterized in that In step S7, the average intersection over union is calculated through the operation of the following formula :[[]]END]] ; ; Among them, TP represents the samples in the region predicted as the positive class, FN represents the samples in the region predicted as the negative class but actually belonging to the positive class, and FP represents the samples in the region predicted as the positive class but actually belonging to the negative class. represents the number of categories of the said label, represents the intersection over union, represents the intersection over union of the .