Semi-supervised semantic segmentation dynamic enhancement method based on adaptive spatial transformation
By adopting the semi-supervised method of adaptive spatial transformation in semantic segmentation technology, dynamically adjusting the amplitude and direction of spatial transformation, the problem that the existing technology is difficult to deal with complex spatial transformation, and significantly improving the accuracy of semantic segmentation and the robustness of the model.
Patent Information
- Application Number
- CN202510465927.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The existing semantic segmentation technology is difficult to effectively process data with complex spatial transformations, which makes it difficult for the model to ensure the precise segmentation of the target area when processing large-scale spatial transformations such as rotation and displacement, which in turn affects the stability and robustness of the model.
The semi-supervised semantic segmentation dynamic enhancement method based on adaptive spatial transformation is adopted to dynamically adjust the amplitude and direction of the spatial transformation, enhance the generalization of the model, and achieve high-precision semantic segmentation. The specific steps include building a loss function, generating pseudo-labels using the teacher model, calculating information entropy for spatial enhancement, further training the student model until the model with the best training effect is selected.
By dynamically adjusting the spatial transformation, the model can maintain stability and robustness in complex input data, significantly improving the accuracy of semantic segmentation, and avoiding label misalignment problems caused by geometric transformation.
Smart Images

Figure CN119992107A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of semantic segmentation, and in particular to a semi-supervised semantic segmentation dynamic enhancement method based on adaptive spatial transformation. Background Art
[0002] Semantic segmentation is a traditional and important field in computer vision, which focuses on classifying all pixels in an image according to their semantic content. In recent years, many successful studies have emerged in this field, with different applications in specific fields, such as natural images, medical images, remote sensing image analysis, driving scene segmentation, point cloud segmentation, etc. The integration of deep learning has greatly improved the efficiency of semantic segmentation tasks. However, deep learning requires a large amount of labeled data to train effective models. Although it is now easy to obtain a large amount of raw data, manually labeling image pixels is a difficult and time-consuming task that requires huge manpower and time. Taking the Cityscape dataset as an example, it takes more than 1.5 hours on average to label and quality control a single image. Therefore, it is a major challenge to effectively train models with insufficient annotations for open-world applications. To address this problem, researchers have explored integrated methods such as weakly supervised learning or unsupervised learning. Semi-supervised learning offers advantages by leveraging labeled and unlabeled datasets. Semi-supervised learning mainly focuses on training models with a limited amount of labeled data and a large amount of unlabeled data as well as a large amount of unlabeled data.
[0003] Semi-supervised semantic segmentation aims to develop a model that can effectively utilize limited labeled data while extracting valuable insights from a large amount of unlabeled data, providing excellent performance and alleviating the challenges faced by researchers with limited labeled data. It can be roughly divided into pseudo-labeling methods, consistency regularization, contrastive learning, adversarial training, and hybrid methods. In the early stage of semi-supervised semantic segmentation research, the general framework of generative adversarial networks (GANs) was more adopted, which can be mainly divided into two types according to their structure: with generators and without generators. The purpose of contrastive learning is to improve the performance of semantic segmentation by learning more useful representations in the embedding space. Consistency regularization (CR) reaches a consensus on the smoothness assumption, hoping that the model will give similar predictions for different perturbations or variants of the same input to improve the segmentation quality. Pseudo-labeling follows a simple pipeline, using a model trained with labeled data to predict unlabeled data to obtain its pseudo-label, and then expanding it to the data set to retrain the model. At present, existing data augmentation research is mostly based on grayscale enhancement, and there is a lack of discussion on spatial variation methods. However, pixel position is a key factor for segmentation tasks, because the segmentation model needs to maintain accurate recognition of the target area under different conditions.
[0004] Most existing enhancement techniques rely on intensity transformations such as grayscale enhancement. However, semantic segmentation tasks require high pixel-level spatial accuracy, especially when dealing with large-scale spatial transformations such as rotation and displacement. Existing enhancement methods often have difficulty in ensuring accurate segmentation of the target area. Therefore, current technologies are difficult to effectively process data with complex spatial transformations, which poses challenges to the stability and robustness of the model. Summary of the invention
[0005] The technical problem to be solved by the present invention is to provide a semi-supervised semantic segmentation dynamic enhancement method based on adaptive spatial transformation, which enhances the generalization of the model and achieves high-precision semantic segmentation by dynamically adjusting the amplitude and direction of the spatial transformation.
[0006] The present invention provides a semi-supervised semantic segmentation dynamic enhancement method based on adaptive spatial transformation, comprising the following steps: S1. Given a set of labeled data sets And a set of unlabeled datasets ,in, It includes M labeled images, Represents a labeled image, Represents a labeled image The corresponding label, Includes N unlabeled images, represents an unlabeled image, where N>>M; S2. Construct the first loss function , and through a labeled dataset And the first loss function Train the student model ; S3. Through the student model Initialize the teacher model , through the teacher model For unlabeled datasets Make predictions and get prediction results , and the prediction results As pseudo labels, Indicates weak enhancement processing; S4. Based on pseudo labels , calculate the unlabeled dataset Information entropy ; S5. According to information entropy For unlabeled images Perform spatial enhancement processing to obtain label-free images The spatial enhancement results , using the student model Spatial enhancement results Make predictions and get prediction results ; S6. Based on pseudo labels and prediction results Constructing the second loss function , and according to the first loss function With the second loss function Constructing the total loss function , through the total loss function , further train the student model , obtain multiple trained student models ; S7. Given a set of test set images, each image in the test set includes a true label, and the trained student model Make predictions for the test set images and calculate the average intersection-over-union ratio between the prediction results and the true labels , according to the average intersection-over-union ratio Select a set of student models with the best training effect .
[0007] In a possible implementation, in step S1, the labeled data set The Pascal VOC 2012 dataset is an unlabeled dataset. It is the Cityscapes dataset.
[0008] In a possible implementation, the initial learning rate of the PASCAL VOC 2012 dataset is set to 0.001, and the initial learning rate of the Cityscapes dataset is set to 0.005.
[0009] In a possible implementation, in step S2, the first loss function The expression is as follows: , Among them, M represents the labeled image The number of Represents the student model The weight parameter, c represents the label category, Indicates that there is a labeled image when predicting The probability of belonging to label category c, C represents the number of label categories c, Represents a labeled image The ground truth value of Indicates weak enhancement processing.
[0010] In a possible implementation, in step S4, the information entropy is calculated by the following formula: : , in, Represents an unlabeled image The probability of belonging to label category c.
[0011] In a possible implementation, in step S5, spatial enhancement processing is performed by the following operation: , in, and Both represent scaling parameters, Represents the offset of the rotation transformation, Represents the offset of the translation transformation, represents the maximum rotation angle, Indicates the maximum translation ratio.
[0012] In one possible implementation, =180°, =0.5.
[0013] In one possible implementation, =0.5, =0.5, =5.5, =3.
[0014] In a possible implementation, the second loss function is calculated in step S6. The following steps are involved: S61: Obtain the second loss function by the following operation: : , Among them, H and W represent the unlabeled image The height and width of Denotes the two spatial enhancement functions and The collection of results obtained; S62, obtain the total loss function through the following operation : , in, Represents a hyperparameter.
[0015] In a possible implementation manner, in step S7, the average intersection-over-union ratio is calculated by the following operation: : , Among them, TP represents the sample predicted to be in the positive area, FN represents the sample predicted to be in the negative area but actually in the positive area, FP represents the sample predicted to be in the positive area but actually in the negative area, and n represents the number of categories of the label. represents the intersection-and-union ratio, represents the intersection-over-union ratio of the i-th category.
[0016] Compared with the prior art, this application has the following advantages: using the teacher model Predict the weakly enhanced image by calculating the information entropy To measure the uncertainty of the prediction, the enhancement strength is adjusted dynamically, making the enhancement strategy more adaptive and accurate during training. At the same time, the intensity of the enhancement operation gradually increases as the training progresses, thereby achieving a spatial transformation from simple to complex during the enhancement process, allowing the model to be trained stably and effectively under spatial transformation conditions. This method effectively avoids the label misalignment problem caused by geometric transformation by ensuring the spatial consistency of the image and prediction results before and after enhancement, improves the robustness and generalization ability of the model in complex input data, and ultimately significantly improves the semantic segmentation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic diagram of the method of the present invention;
[0018] Figure 2 This is a visualization diagram of the effect of the present invention;
[0019] Figure 3 This is a visualization diagram of the effect of the present invention. DETAILED DESCRIPTION
[0020] First, those skilled in the art should understand that these implementations are only used to explain the technical principles of the embodiments of the present application, and are not intended to limit the protection scope of the embodiments of the present application. Those skilled in the art can make adjustments to them as needed to adapt to specific application scenarios.
[0021] The present application is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0022] This example studies the impact of the method on two benchmark segmentation datasets, Pascal VOC 2012 and Cityscapes. The Pascal VOC 2012 dataset contains 21 semantic categories, divided into classic and hybrid subsets. The classic training set consists of 1464 images with extensive labels and 1449 images for validation. In addition, a new version of the hybrid training set includes low-resolution, coarsely annotated images from the Segmentation Boundary Dataset (SBD), thus expanding the training pool to 10582 images. The Cityscapes dataset covers 19 semantic categories in an urban context, providing 2975 training images with precise annotations and 500 images for validation.
[0023] In terms of model construction, this embodiment proposes a pluggable architecture, so the effect of the method is evaluated on the CNN and ViT structures respectively. Specifically, for the model based on the convolutional neural network (CNN) architecture, DeepLabV3+ is selected as the semantic segmentation model, and ResNet-101 is used as the backbone network. ResNet-101 has been pre-trained on the ImageNet dataset, so it has good feature extraction capabilities. For the model based on the visual transformer (ViT), SegFormer-B5 is selected as the semantic segmentation model. It is a semantic segmentation model based on the ViT architecture. It has been pre-trained on a large-scale dataset and can effectively handle complex visual tasks. In order to evaluate the performance, this embodiment uses SGD as the optimizer and adopts a polynomial decay learning rate strategy. The learning rate is expressed as follows: , in, represents the initial learning rate, Indicates the current iteration number, Represents the total number of iterations, power represents a hyperparameter, and power and weight decay are set to 0.9 and 1e-4 respectively. The SGD optimizer is combined with the learning rate decay strategy to ensure convergence during training, avoid excessive learning rates leading to training shocks, and ensure sufficient training time so that the model can be fully optimized in the direction of minimum loss.
[0024] For the PASCAL VOC 2012 dataset, the initial learning rate was set to 0.001, the crop size was 321× 321 or 513× 513, the batch size was 16, and 80 epochs were trained. For the Cityscapes dataset, the initial learning rate was 0.005, the crop size was 801×801, the batch size was also 16, and 240 epochs were trained. All training and inference processes used the PyTorch deep learning framework and were performed on 4×NVIDIA V100GPUs and 2×A800 GPUs for different numerical analysis.
[0025] This embodiment discloses a semi-supervised semantic segmentation dynamic enhancement method based on adaptive spatial transformation, comprising the following steps: S1. Given a set of labeled data sets And a set of unlabeled datasets ,in, It includes M labeled images, Represents a labeled image, Represents a labeled image The corresponding label, Includes N unlabeled images, represents an unlabeled image, where N>>M; S2. Construct the first loss function , and through a labeled dataset And the first loss function Train the student model .
[0026] In step S2, the first loss function The expression is as follows: , Among them, M represents the labeled image The number of Represents the student model The weight parameter, c represents the label category, Indicates that there is a labeled image when predicting The probability of belonging to label category c, C represents the number of label categories c, Represents a labeled image The ground truth value of Indicates weak enhancement processing. Weak enhancement usually includes mild image transformation, such as random cropping, random flipping, etc. The transformation amplitude of weak enhancement is small, and the purpose is to allow the model to adapt to smaller changes while maintaining the original features. Weak enhancement processing is a prior art and will not be described in detail in this application.
[0027] Through the student model Initialize the teacher model , through the teacher model For unlabeled datasets Make predictions and get prediction results , and the prediction results As pseudo labels, Indicates weak enhancement processing.
[0028] Specifically, in this embodiment, from the student model Extract the weight parameters and load them into the teacher model In the process, ensure that the teacher model Ability to quickly start using unlabeled datasets Make predictions. In this way, the teacher model The initial state can ensure that it is in the unlabeled dataset The prediction on has a certain accuracy. Next, use the teacher model For unlabeled datasets Make predictions, teacher model Output each unlabeled image The predicted probability distribution of each pixel belonging to each category in the teacher model The prediction results As pseudo labels, these pseudo labels will be used to guide the student model In the subsequent training process, the unlabeled dataset Pseudo-labels provide an approximation of the “true labels”, and although these labels are not perfect, they provide a useful guide for the student model. Provides enough information for effective learning.
[0029] According to the pseudo-label , calculate the unlabeled dataset Information entropy .
[0030] In step S4, the information entropy is calculated by the following formula: : , in, Represents an unlabeled image The probability of belonging to label category c.
[0031] Calculating Information Entropy It is used to provide useful guidance for subsequent training. For each unlabeled image , through the teacher model Predict the probability distribution of the sample belonging to each category, and then get the unlabeled image Information entropy . Using the information entropy calculated Value to evaluate each unlabeled image The "learning value", that is, the entropy value, will affect its priority in the training process. Information entropy The larger the information entropy, the higher the model's prediction uncertainty for the sample, that is, the model is not sure which category the sample belongs to. Such samples have higher "learning value" because they can help the model explore more feature space, so they will be given higher weights, thereby applying more spatial transformation and enhancement to them during training. Conversely, information entropy A lower sample size means that the model is more certain about its prediction of the sample, the model has a lower demand for learning it, and the transformation during the learning process will be relatively small, thus maintaining the consistency of its features.
[0032] In order to fine-tune the level of enhanced distortion, this embodiment introduces an entropy-based spatial transformation adaptive weight, namely EAW, using unlabeled images Information entropy As a measure of sample reliability. The calculation formula of adaptive weight is as follows: , Wherein, k and d are both adjustable parameters, which respectively control the enhanced amplitude and offset. In this embodiment, the sigmoid function is used as the basis and modified accordingly to obtain The specific calculation formula of , adding the offset d, inputs the information entropy in a fine-tuned way The value is optimized to perform differential enhancement between different samples in order to more effectively explore the transformed space. Because samples with high entropy show greater uncertainty, more significant spatial transformations are required to explore a wider range of feature spaces. On the contrary, for samples with low entropy, the model shows higher certainty and is more suitable for smaller transformations, thereby retaining stable features. This method can more smoothly adjust the enhancement between samples and enhance the model's ability and flexibility to spatial transformations.
[0033] According to information entropy For unlabeled images Perform spatial enhancement processing to obtain label-free images The spatial enhancement results , using the student model Spatial enhancement results Make predictions and get prediction results .
[0034] In step S5, spatial enhancement processing is performed by the following operation: , in, and Both represent scaling parameters, Represents the offset of the rotation transformation, Represents the offset of the translation transformation, represents the maximum rotation angle, Indicates the maximum translation ratio.
[0035] According to the information entropy of each sample The value is to apply spatial enhancement to it. Information entropy As an indicator of sample uncertainty, it can help determine the intensity of data enhancement in order to more effectively improve the performance of the model. The data representation model with high information entropy has high uncertainty in its prediction. Usually, the features of these samples are complex and may involve more changes. Therefore, for high entropy samples, stronger spatial transformations, such as larger rotations and translations, need to be applied to explore a wider feature space. In this way, the feature diversity can be expanded through transformation, helping the model to learn more possible feature patterns. The data representation model with low information entropy is more certain about its prediction, the features of the samples are more stable, and the learning difficulty is lower. Therefore, for low entropy samples, a smaller spatial transformation needs to be applied to retain its original features and avoid introducing too many unnecessary changes. In this embodiment, spatial enhancement is used instead of the traditional robust enhancement method of intensity focusing. Traditional methods may only focus on enhancing specific areas of samples, while in this embodiment, by dynamically adjusting the spatial transformation, high entropy samples are subjected to a wider range of spatial transformations, while low entropy samples remain stable. In order to dynamically determine the applied spatial transformation enhancement strength according to the entropy value, a mapping function for the spatial enhancement processing of the above formula is designed, and the amplitude of rotation and translation is determined by the mapping function.
[0036] When information entropy If the value is relatively small, the mapping result will also be reduced, which will reduce the amplitude of rotation and translation, thereby improving stability. In order to ensure the effectiveness of the enhancement process, this embodiment is tested in detail during implementation, and the specific parameter settings are set as follows: =180°, =0.5, That is, the maximum translation ratio is 50% of the original image size. These parameters control the strength of the spatial transformation. Especially in high entropy samples, larger rotations and translations help to explore the feature space more extensively. For the PASCAL VOC 2012 dataset, = 1, = 1, = 11, =7 to adjust the transformation strength of rotation and translation. Specifically, the rotation amplitude and displacement amplitude are large to enhance the model's adaptability to diverse scenes. For the Cityscapes dataset, according to its characteristics, this embodiment adjusts some parameters to =0.5, =0.5, =5.5, =3 to apply a smaller transformation amplitude in more complex urban scenes. This adjustment ensures that the characteristics of the sample will not be excessively changed in complex environments, maintaining higher stability.
[0037] When information entropy When the value is small, the mapping result will be reduced, which will reduce the magnitude of the rotation and translation, thereby maintaining the stability of the applied enhancement. In this way, it ensures that when processing low entropy samples, the transformation amplitude is not too large to avoid excessive distortion of the original features. For high entropy samples, the magnitude of rotation and translation is increased, which can better expand the feature space and help the model learn more transformation patterns from uncertainty.
[0038] S6. Based on pseudo labels and prediction results Constructing the second loss function , and according to the first loss function With the second loss function Constructing the total loss function , through the total loss function , further train the student model , obtain multiple trained student models .
[0039] In step S6, the second loss function is calculated The following steps are involved: S61: Obtain the second loss function by the following operation: : , Among them, H and W represent the unlabeled image The height and width of Denotes the two spatial enhancement functions and The collection of results obtained; S62, obtain the total loss function through the following operation : , in, represents a hyperparameter. That is, the loss weight, which is set to 0.5 in this embodiment. This weight is used to balance the labeled images. And unlabeled images The proportion of loss in the total loss. Reasonable adjustment of λ can control the unlabeled image The impact of the training process. Using the total loss function Retrain the student model During the training process, the unlabeled images Apply spatial transformation enhancement , these enhanced unlabeled images And the corresponding pseudo label Together as input, passed to the student model Pseudo labels are obtained by using the teacher model The prediction is generated to replace the real label. This embodiment uses the mean square error, or MSE, which can effectively measure the accuracy of the model prediction and align the spatially transformed images to calculate the pseudo label. and prediction results The mean square error of the two parts, for each pixel position The predicted value differences are squared and summed, and then the average is calculated to obtain the pixel-level mean square error of the image. The mean square error can directly measure the pixel-level difference between the image after spatial transformation and the pseudo label, and is insensitive to the mask structure, which means that even if there are some local occlusions or noise in the image, it will not excessively affect the overall loss calculation, thereby enhancing the model's robustness to spatial transformation. The total loss function above is It consists of two parts: Represents labeled data of loss, Represents unlabeled data loss.
[0040] Given a set of test set images, each image in the test set includes the true label, and the trained student model Make predictions for the test set images and calculate the average intersection-over-union ratio between the prediction results and the true labels , according to the average intersection-over-union ratio Select a set of student models with the best training effect .
[0041] In step S7, the average intersection-over-union ratio is calculated by the following operation: : , Among them, TP represents the sample predicted to be in the positive area, FN represents the sample predicted to be in the negative area but actually in the positive area, FP represents the sample predicted to be in the positive area but actually in the negative area, and n represents the number of categories of the label. represents the intersection-and-union ratio, represents the intersection-over-union ratio of the i-th category.
[0042] In this embodiment, the trained student model Applied to the test set images to evaluate its performance in actual tasks. By inputting the test set data, extracting image features, and using the average intersection-over-union ratio As a performance evaluation metric, this metric is effective even when dealing with imbalanced classes that often occur in pixel-level annotation tasks.
[0043] First, the test set images are input into the trained student model Through the student model The model will generate prediction results for each test image during the forward propagation process. The model outputs prediction categories based on different regions of the image, and classifies each pixel based on the model output to obtain the predicted category of each pixel. Through these outputs, the features of each test image are extracted. The average intersection over union ratio is used As an evaluation indicator to evaluate the trained student model . The intersection-over-union ratio is the intersection-over-union ratio of each category, which measures the prediction accuracy of category C, that is, the ratio of the intersection and union of the predicted area to the true label area. By calculating the intersection-over-union ratio of each category and averaging the intersection-over-union ratios of all categories, the performance of the entire model on all categories is obtained. This indicator can effectively evaluate the performance of the model in a multi-category environment, especially when the categories are unbalanced.
[0044] In order to evaluate the performance of the model at a fine-grained level, this embodiment adopts a sliding window evaluation method, especially on urban landscape datasets such as Cityscapes. The sliding window evaluation method divides the image into small blocks, or windows, and calculates the , and gradually slide the window to cover the entire image. This can help obtain detailed representation of local areas, especially in the case of small objects or complex scenes. In implementation, this embodiment sets an appropriate window size and sliding step size according to the resolution and computing power of the image. Smaller windows can provide more detailed evaluation, but the amount of calculation is larger. Appropriate step size selection can balance evaluation accuracy and computing efficiency.
[0045] By calculation , can accurately evaluate the performance of the model in each category, especially the segmentation effect of small category objects. It is particularly effective in dealing with class imbalance problems because it takes into account the segmentation accuracy of each category. The sliding evaluation method further enhances the adaptability of model evaluation and can provide more detailed local area performance evaluation, especially for complex urban landscape datasets, where objects in the image may be of different sizes and complex shapes.
[0046] Visualize the predicted pixel categories to get the semantic segmentation results. Figure 2 and attached Figure 3 In , the visualization of this embodiment on the Cityscape dataset and the Pascal VOC 2012 dataset is shown, and compared with the baseline method and other advanced methods. Each row represents a different image, showing the segmentation effect under different methods.
[0047] In the Pascal VOC 2012 dataset, it can be seen from this comparison that the method of this embodiment performs outstandingly in object category recognition, boundary demarcation, and detail feature retention, and is superior to the baseline method in most cases. The segmentation results are more accurate, especially in complex scenes, such as scenes containing animals or people, which can better maintain the integrity of the object and improve the clarity of the segmented area. The baseline method often has the problem of blurred boundaries in complex scenes, especially in the detail segmentation of people or animals. For example, it is more accurate in the segmentation of the outlines of airplanes, animals, and people, and significantly reduces the interference of background noise.
[0048] In the urban landscape dataset, the method of this embodiment further demonstrates its superior segmentation performance under complex conditions. Although the baseline method can identify the main objects in the scene, it has deficiencies in detail processing, such as the segmentation of human outlines and small objects, resulting in unclear boundaries, especially in scenes with multiple objects or complex backgrounds. Compared with the baseline method, the method of this embodiment shows higher segmentation accuracy in all images. It has clearer boundaries between people, vehicles, and environmental elements such as roads and buildings, avoiding the common blur and overlap problems of the baseline method in complex scenes. For example, in the first row of visualization results, the method of this embodiment can accurately segment the outline of pedestrians, and the segmentation of traffic signs and background is also clearer; the second row shows that the method of this embodiment can more accurately segment buses from the road surface, and can clearly distinguish vehicles from road signs; the third and fourth rows show that the method of this embodiment makes the segmentation of buildings and roads more refined and retains more details through more accurate boundary positioning. The method of this embodiment shows significant advantages in improving the generalization ability of the model, detail fidelity, and segmentation accuracy of complex scenes, and provides a new solution for semi-supervised semantic segmentation tasks.
[0049] In addition, this embodiment also provides the pseudo code of the present application method as follows: Input: labeled dataset , unlabeled dataset , student model , teacher model ; Output: Trained student model ; Step 1: Use the first loss function In the labeled dataset Train the student model ; Step 2: Initialize the teacher model , student model ; for in do Using the Teacher Model For unlabeled images Generate prediction results and obtain information entropy ; Using information entropy Sure The strong enhancement effect is applied to obtain the spatial enhancement result ; Using the Student Model right Generate pseudo labels ; Calculate the second loss function using predictions and pseudo labels And the total loss function Training the student model ; end for Return to Student Model .
[0050] In the description of the present application, the description with reference to the terms "one embodiment", "some embodiments", "in the present embodiment", "specific example", or "some examples" etc. means that the specific features, mechanisms, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, mechanisms, materials or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0051] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A dynamic enhancement method for semi-supervised semantic segmentation based on adaptive spatial transformation, characterized in that: The following steps are involved: S1. Given a set of labeled data sets And a set of unlabeled datasets ,in, It includes M labeled images, represents the labeled image, Represents the labeled image The corresponding label, Includes N unlabeled images, represents the unlabeled image, where N>>M; S2. Construct the first loss function , and through the labeled dataset And the first loss function Train the student model ; S3. Through the student model Initialize the teacher model , through the teacher model For the unlabeled dataset Make predictions and get prediction results , and the prediction results As pseudo labels, Indicates weak enhancement processing; S4. According to the pseudo label , calculate the unlabeled dataset Information entropy ; S5. According to the information entropy For the unlabeled image Perform spatial enhancement processing to obtain the label-free image The spatial enhancement results , using the student model The spatial enhancement result Make predictions and get prediction results ; S6. According to the pseudo label And the prediction results Constructing the second loss function , and according to the first loss function With the second loss function Constructing the total loss function , through the total loss function , further train the student model , obtain the student models after multiple training ; S7, given a set of test set images, each image in the test set images includes a true label, through the trained student model Predict the test set images and calculate the average intersection-over-union ratio between the predicted results and the true labels. , according to the average intersection-over-union ratio Select the set of student models with the best training effect .
2. The method for dynamic enhancement of semi-supervised semantic segmentation based on adaptive spatial transformation according to claim 1, characterized in that: In step S1, the labeled data set The Pascal VOC 2012 dataset is an unlabeled dataset. It is the Cityscapes dataset.
3. The method for dynamic enhancement of semi-supervised semantic segmentation based on adaptive spatial transformation according to claim 2, characterized in that: The initial learning rate of the PASCAL VOC 2012 dataset is set to 0.001, and the initial learning rate of the Cityscapes dataset is set to 0.
005.
4. The method for dynamic enhancement of semi-supervised semantic segmentation based on adaptive spatial transformation according to claim 1, 2 or 3, characterized in that: In step S2, the first loss function The expression is as follows: , Where M represents the labeled image The number of Represents the student model The weight parameter, c represents the label category, Indicates the prediction of the labeled image The probability of belonging to label category c, C represents the number of label categories c, Represents the labeled image The ground truth value of Indicates weak enhancement processing.
5. The method for dynamic enhancement of semi-supervised semantic segmentation based on adaptive spatial transformation according to claim 4, characterized in that: In step S4, the information entropy is calculated by the following formula: : , in, Represents the unlabeled image The probability of belonging to label category c.
6. The method for dynamic enhancement of semi-supervised semantic segmentation based on adaptive spatial transformation according to claim 1, 2, 3 or 5, characterized in that: In step S5, spatial enhancement processing is performed by the following operation: , in, and Both represent scaling parameters, Represents the offset of the rotation transformation, Represents the offset of the translation transformation, represents the maximum rotation angle, Indicates the maximum translation ratio.
7. The method for dynamic enhancement of semi-supervised semantic segmentation based on adaptive spatial transformation according to claim 6, characterized in that: =180°, =0.5。 8. The method for dynamic enhancement of semi-supervised semantic segmentation based on adaptive spatial transformation according to claim 7, characterized in that: =0.5, =0.5, =5.5, =3。 9. The method for dynamic enhancement of semi-supervised semantic segmentation based on adaptive spatial transformation according to claim 1, 2, 3, 5, 7 or 8, characterized in that: In step S6, the second loss function is calculated The following steps are involved: S61: Obtain the second loss function by the following operation: : , Among them, H and W represent the unlabeled image The height and width of Denotes the two spatial enhancement functions and The collection of results obtained; S62, obtain the total loss function through the following operation : , in, Represents a hyperparameter.
10. The method for dynamic enhancement of semi-supervised semantic segmentation based on adaptive spatial transformation according to claim 1, 2, 3, 5, 7 or 8, characterized in that: In step S7, the average intersection-over-union ratio is calculated by the following operation: : , Among them, TP represents the sample predicted to be in the positive area, FN represents the sample predicted to be in the negative area but actually in the positive area, FP represents the sample predicted to be in the positive area but actually in the negative area, and n represents the number of categories of the label. represents the intersection-and-union ratio, represents the intersection-over-union ratio of the ith category.
Citation Information
Patent Citations
Semi-supervised domain adaptive image semantic segmentation method, system and device and storage medium
CN116229080A
Semi-supervised learning semantic segmentation method based on dynamic adjustment of exponential function threshold
CN117809305A
Semi-supervised rotating target detection method based on local space consistency prior information
CN117893916A
Mutual learning-based semi-supervised medical image segmentation method and system
WO2023116635A1