Semi-supervised semantic segmentation and model training method and device based on weighted comparison network
Through the weighted comparison network semi-supervised semantic segmentation method, the model is trained using labeled and labelless images, pseudo-labels are generated and adaptive loss weighting is performed, which solves the problem of sample imbalance in semantic segmentation, and improves the robustness and accuracy of the model.
Patent Information
- Application Number
- CN202510733789.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
In the prior art, due to uneven sample categories in semantic segmentation tasks, the difficulty of model training increases, the robustness and accuracy are reduced, and the pixel-level labeling cost is high.
The semi-supervised semantic segmentation method of weighted comparison network is adopted, and the model is trained using labeled and labelless images, and pseudo-labels are generated by enhanced networks, and combined with adaptive loss weighting and pseudo-label screening mechanisms, the model learning ability and generalization ability are improved.
It significantly reduces the labor and time cost of the data preparation process, improves the learning ability of the model and the generalization ability of complex scenarios, and solves the poor performance problems caused by unbalanced sample categories.
Smart Images

Figure CN120259673A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of semantic segmentation, and in particular to a method and device for semi-supervised semantic segmentation and model training based on a weighted contrast network. Background Art
[0002] With the rapid development of deep learning, computer vision tasks have shined in all walks of life. Image semantic segmentation methods based on deep convolutional neural networks have received more and more attention, and their application scope has gradually expanded from conventional daily image semantic segmentation to semantic segmentation in specific application scenarios. After deep learning was applied to semantic segmentation, the effect of image semantic segmentation has been greatly improved. The training of a good segmentation model by a neural network is inseparable from the drive of a large number of labeled training samples, and the semantic segmentation task is to classify pixels, and labeling high-quality pixel-level training sample labels is a process that consumes manpower, material resources and financial resources.
[0003] In addition, semantic segmentation tasks also face the problem of category imbalance. Since semantic segmentation is a pixel-based classification task, the number of randomly acquired category training samples for different scene task categories may change greatly. When the number of training samples is unbalanced, the difficulty of model training will increase while reducing the robustness and accuracy of the model. Therefore, it is necessary to provide a semi-supervised semantic segmentation and model training method and device based on a weighted contrast network. Summary of the invention
[0004] In view of the shortcomings of the prior art described above, the purpose of the present application is to provide a semi-supervised semantic segmentation and model training method and device based on a weighted contrast network, which improves the problem of poor model training performance in the prior art due to sample imbalance.
[0005] To achieve the above and other related objectives, the present application provides a training method for a semi-supervised semantic segmentation model based on a weighted contrast network. The training method includes: obtaining a labeled image set containing target objects and a corresponding label set, as well as an unlabeled image set containing target objects without labels; wherein the number of unlabeled images is much larger than the number of labeled images; inputting the labeled image set into the first basic network of the first branch and the second basic network of the second branch respectively to perform semantic segmentation on the target objects in the labeled image set, and correspondingly obtaining a first target mask set and a second target mask set; wherein the first branch includes a first basic network and a first enhancement network, the second branch includes a second basic network and a second enhancement network, and the structures of the first enhancement network, the second enhancement network, the first basic network, and the second basic network are the same but the initial parameters are different; inputting the unlabeled image set into the first basic network, the second basic network, the first enhancement network, and the second enhancement network respectively to perform semantic segmentation on the target objects in the unlabeled image set, and correspondingly obtaining a third target mask set, a fourth target mask set, a first enhanced pseudo-label set, and a second enhanced pseudo-label set; calculating the difference degrees between the first target mask set, the second target mask set and the label set, the difference degree between the third target mask set and the second enhanced pseudo-label set, and the difference degree between the fourth target mask set and the first enhanced pseudo-label set to obtain the total difference degree; adjusting the parameters of the first basic network and the second basic network based on the total difference degree to complete the training of the semantic segmentation model; wherein the semantic segmentation model includes the trained first basic network and / or the second basic network.
[0006] In an embodiment of the present invention, the step of inputting the labeled image set into the first basic network of the first branch and the second basic network of the second branch respectively to perform semantic segmentation on the target objects in the labeled image set and correspondingly obtaining a first target mask set and a second target mask set includes: performing a first perturbation process on the labeled image set to obtain a perturbed labeled image set; inputting the perturbed labeled image set into the first basic network to perform semantic segmentation on the target objects in the perturbed labeled image set to obtain a first target mask set; inputting the perturbed labeled image set into the second basic network to perform semantic segmentation on the target objects in the perturbed labeled image set to obtain a second target mask set.
[0007] In one embodiment of the present invention, inputting the perturbed labeled image set into the first basic network to perform semantic segmentation on the target objects in the perturbed labeled image set to obtain a corresponding first target mask set includes: inputting the perturbed labeled image set into the first basic network to perform semantic segmentation on the target objects in the perturbed labeled image set to obtain the maximum class confidence of each pixel point of all the labeled images in the perturbed labeled image set; calculating the confidence mean value of each perturbed labeled image respectively based on the maximum class confidence of each pixel point; for each pixel point in the perturbed labeled image: determining whether the maximum class confidence of the pixel point is greater than or equal to the corresponding confidence mean value: if so, taking the class number corresponding to the maximum class confidence as the first target mask of the pixel point; otherwise, not allocating a first target mask to the pixel point; counting the first target masks in all the perturbed labeled images to obtain the first target mask set.
[0008] In one embodiment of the present invention, inputting the unlabeled image set into the first enhancement network and the second enhancement network respectively to perform semantic segmentation on the target objects in the unlabeled image set to correspondingly obtain a first enhanced pseudo-label set and a second enhanced pseudo-label set includes: performing a second perturbation process on the unlabeled image set to obtain a perturbed unlabeled image set; inputting the perturbed unlabeled image set into the first enhancement network to perform semantic segmentation on the target objects in the perturbed unlabeled image set to correspondingly obtain a first enhanced pseudo-label set; inputting the perturbed unlabeled image set into the second enhancement network to perform semantic segmentation on the target objects in the perturbed unlabeled image set to correspondingly obtain a second enhanced pseudo-label set.
[0009] In one embodiment of the present invention, inputting the perturbed unlabeled image set into the second enhancement network to perform semantic segmentation on the target objects in the perturbed unlabeled image set to correspondingly obtain a second enhanced pseudo-label set includes: inputting the perturbed unlabeled image set into the second enhancement network to perform semantic segmentation on the target objects in the perturbed unlabeled image set to obtain the maximum class confidence of each pixel point of all the unlabeled images in the perturbed unlabeled image set; calculating the confidence mean value of each perturbed unlabeled image respectively based on the maximum class confidence of each pixel point; for each pixel point in the perturbed unlabeled image: determining whether the maximum class confidence of the pixel point is greater than the confidence mean value: if so, taking the class number corresponding to the maximum class confidence as the second enhanced pseudo-label of the pixel point; otherwise, not allocating a second enhanced pseudo-label to the pixel point; counting the second enhanced pseudo-labels in all the perturbed unlabeled images to obtain the second enhanced pseudo-label set.
[0010] In one embodiment of the present invention, calculating the difference degree between the third target mask set and the second enhanced pseudo-label set includes: determining the categories of each second enhanced pseudo-label in the second enhanced pseudo-label set; calculating the category proportion of each type of second enhanced pseudo-label in all second enhanced pseudo-labels in the second enhanced pseudo-label set: for each type of second enhanced pseudo-label in the second enhanced pseudo-label set: calculating the first semi-supervised difference degree between the third target mask set and this type of second enhanced pseudo-label; weighting the first semi-supervised difference degree based on the category proportion to obtain the final difference degree.
[0011] In one embodiment of the present invention, the semantic segmentation model is obtained by weighting a trained first basic network and a trained second basic network.
[0012] In one embodiment of the present invention, a semantic segmentation method is further provided. The semantic segmentation method includes: obtaining an image to be segmented; inputting the image into a semantic segmentation model to obtain a target mask for segmenting a target object in the image; wherein, the semantic segmentation model is trained by the semi-supervised semantic segmentation model training method based on a weighted contrast network described in any one of the above; segmenting the image based on the target mask.
[0013] In an embodiment of the present invention, there is also provided a training system for a semi-supervised semantic segmentation model based on a weighted contrast network. The training system includes: an image acquisition module, configured to acquire a labeled image set containing a target object and a corresponding label set, and an unlabeled image set containing the target object without labels; wherein the number of unlabeled images is much larger than the number of labeled images; a supervised module, configured to respectively input the labeled image set into a first basic network of a first branch and a second basic network of a second branch, perform semantic segmentation on the target object in the labeled image set, and correspondingly obtain a first target mask set and a second target mask set; wherein the first branch includes a first basic network and a first enhancement network, the second branch includes a second basic network and a second enhancement network, and the structures of the first enhancement network, the second enhancement network, the first basic network, and the second basic network are the same but the initial parameters are different; an unsupervised module, configured to respectively input the unlabeled image set into the first basic network, the second basic network, the first enhancement network, and the second enhancement network, perform semantic segmentation on the target object in the unlabeled image set, and correspondingly obtain a third target mask set, a fourth target mask set, a first enhanced pseudo-label set, and a second enhanced pseudo-label set; a loss calculation module, configured to respectively calculate the difference degrees between the first target mask set, the second target mask set and the label set, the difference degree between the third target mask set and the second enhanced pseudo-label set, and the difference degree between the fourth target mask set and the first enhanced pseudo-label set, to obtain a total difference degree; a parameter update module, configured to adjust the parameters of the first basic network and the second basic network based on the total difference degree to complete the training of the semantic segmentation model; wherein the semantic segmentation model includes the trained first basic network and / or the second basic network.
[0014] In an embodiment of the present invention, there is also provided an electronic device. The electronic device includes: one or more processors; a storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device implements the training method for the semi-supervised semantic segmentation model based on the weighted contrast network or the semantic segmentation method described in any one of the above.
[0015] In an embodiment of the present invention, there is also provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor of a computer, the computer is caused to execute the training method for the semi-supervised semantic segmentation model based on the weighted contrast network or the semantic segmentation method described in any one of the above.
[0016] As described above, a semi-supervised semantic segmentation and model training method and device based on a weighted contrast network of the present invention have the following beneficial effects: By introducing unlabeled images with a quantity far greater than that of labeled images and combining an enhancement network to generate pseudo-labels for training, the dependence on a large number of pixel-level manual annotations is effectively reduced, and the manpower and time costs in the data preparation process are significantly reduced. In addition, by generating pseudo-labels through the enhancement network and cross-comparing them with the prediction results of the basic network, the potential semantic information in the unlabeled images can be mined, further enhancing the learning ability of the model and the generalization ability for complex scenarios. Therefore, this way of the present invention improves the problem of poor performance caused by unbalanced sample categories in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other embodiments can be obtained based on these drawings.
[0018] Figure 1 It is a flowchart of a semi-supervised semantic segmentation model training method based on a weighted contrast network provided by an embodiment of the present application; Figure 2 It is a structural diagram of a semi-supervised semantic segmentation model based on a weighted contrast network provided by the present application; Figure 3 It is a flowchart of a semantic segmentation method provided by an embodiment of the present application; Figure 4 It is shown as a structural block diagram of a semi-supervised semantic segmentation model training system provided by an embodiment of the present invention; Figure 5 It is shown as a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] The following uses specific specific examples to illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0020] It should be noted that the illustrations provided in the following embodiments only schematically illustrate the basic concept of the present application. Therefore, only the components related to the present application are shown in the illustrations, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0021] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present application. However, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present application difficult to understand.
[0022] The inventors have found that as semi-supervised image classification algorithms have shown extremely high effects, the need to design semi-supervised semantic segmentation algorithms that use a small number of labeled training samples and a large number of unlabeled training samples has become increasingly urgent, and semi-supervised semantic segmentation has received more and more attention. When only a small amount of labeled data is required to train a good neural network model, it can not only save resources but also be deployed in various application scenarios in a timely manner to obtain greater benefits. However, the imbalance in the number of training samples often leads to an increase in the difficulty of model training and a reduction in the robustness and accuracy of the model. Therefore, if the model can adaptively adjust the loss and the number of samples during training, the applicable task range of the segmentation algorithm will be wider. So the two difficulties in scene image semantic segmentation are the label annotation problem and the training imbalance problem.
[0023] The present application provides a training method for a semi - supervised semantic segmentation model based on a weighted contrast network. This semantic segmentation model consists of two basic networks and their corresponding two enhanced networks. The four network structures are exactly the same, but different initialization methods are used to enhance the diversity and robustness of the model. During the training process, labeled images are used to perform standard supervised training on the two basic networks to make full use of the manually labeled information. For unlabeled images, the enhanced network first generates enhanced pseudo - labels and uses them as supervision signals to cross - guide the training of the basic network in the other branch, thereby mining the potential semantic features in the images. To alleviate the problem of uneven class distribution during training, an adaptive loss weighting mechanism is designed to dynamically adjust the loss weights according to the pixel proportion of different classes in the enhanced pseudo - labels of the images, so as to improve the model's learning ability for low - frequency classes. In addition, in the supervised training path, a hard sample screening mechanism is introduced to select relevant pixel regions from the labeled images to guide the basic network to focus on difficult - to - classify regions for key learning. In the semi - supervised path, only the high - confidence pseudo - label pixels in the unlabeled images are retained for training to ensure the reliability of the semi - supervised signals. After each training batch is completed, an exponential moving average strategy is adopted to perform momentum - type updates on the corresponding enhanced network according to the current parameters of the basic network to maintain the prediction stability of the enhanced network and the quality of the pseudo - labels. A semantic segmentation model is constructed through the basic network after training convergence to achieve accurate segmentation of the target objects in the images.
[0024] As Figure 1 shown, the training method for the semi - supervised semantic segmentation model based on the weighted contrast network includes the following steps: S11. Obtain a labeled image set containing target objects and the corresponding label set, and an unlabeled image set containing target objects without labels; where the number of unlabeled images is much larger than the number of labeled images.
[0025] To construct a dataset for a semi-supervised semantic segmentation model, it is necessary to obtain multiple sample images containing target objects, and divide the obtained sample images into labeled images and unlabeled images based on the presence or absence of manually annotated pixel-level semantic labels, thereby forming a labeled image set and an unlabeled image set. Among them, the labeled images have corresponding label masks, which are used to represent the semantic categories of each pixel point in the labeled images, such as whether the pixel point is a vehicle, a road, or a pedestrian, etc. Although the unlabeled images also have target objects, they do not have manually annotated labels. The acquisition method of the sample images is not limited and can be adaptively set according to the actual task requirements. Exemplarily, for traffic scenarios, the sample images can be collected by multiple camera devices, such as cameras installed on traffic roads, etc. And during the sampling process, sample images are obtained by covering multiple typical scene areas where cameras are deployed, and average random sampling is performed on the images at different sites to ensure the extensiveness and representativeness of the training samples in terms of spatial and scene distribution, thereby further enhancing the adaptability of the model in actual scenarios.
[0026] Furthermore, considering that this kind of manual pixel-level semantic annotation has problems such as high cost and slow speed, in order to effectively reduce the annotation overhead and improve the training efficiency, the number of unlabeled images used in this application is much larger than that of labeled images. Here, much larger means that the number of unlabeled images is 10 times or more different from the number of labeled images. In this way, while significantly reducing the manual annotation cost, the potential semantic information in the unlabeled images can be fully utilized through the pseudo-label mechanism, thereby enhancing the generalization ability of the semi-supervised semantic segmentation model to complex scenarios and improving the segmentation accuracy. Exemplarily, the sample images include m labeled images for training, unlabeled images for training, validation images for verification, and several test images for testing the model performance, where . Using the segmentation annotation tool to label and obtain label images.
[0027] In terms of data annotation, assume that the training set includes m labeled images, denoted as , where represents the i-th labeled image and its corresponding label. Similarly, assume that the training set includes n unlabeled images, denoted as , where represents the j-th unlabeled image, and there is , that is, the number of unlabeled images is much larger than the number of labeled images. The set composed of m labeled images and the set The union of them forms a training set for training the semantic segmentation model. Denote the validation set composed of t images as , where represents the k-th validation image and its corresponding label in the validation set, and there is .
[0028] S12. Input the labeled image set into the first basic network of the first branch and the second basic network of the second branch respectively, perform semantic segmentation on the target objects in the labeled image set, and correspondingly obtain a first target mask set and a second target mask set; wherein, the first branch includes a first basic network and a first enhancement network, the second branch includes a second basic network and a second enhancement network, and the structures of the first enhancement network, the second enhancement network, the first basic network, and the second basic network are the same but the initial parameters are different.
[0029] During the training process, input the labeled image set into two symmetrically structured branches respectively: the first branch includes a first basic network and a first enhancement network, and the second branch includes a second basic network and a second enhancement network. And the structures of these four networks are the same, but the initialization methods of the enhancement network and the basic network are different. Among them, the initialization methods include but are not limited to uniform distribution initialization, normal distribution initialization, kaiming initialization, etc. Among them, the two basic networks can adopt the same initialization method or different initialization methods. Exemplarily, the first basic network uses Xavier uniform distribution initialization, and the second basic network uses kaiming uniform distribution initialization. After initialization, input the labeled image set into the first basic network and the second basic network respectively, and correspondingly obtain a first target mask set and a second target mask set. Among them, each first target mask / second target mask is used to represent the semantic categories of all pixel points in the corresponding labeled image. It can be understood that the basic network and the enhancement network can use any current mainstream segmentation network, which includes but is not limited to , or variants of series and etc. Those skilled in the art can adaptively select the corresponding network based on the actual task requirements, and no specific limitation is made.
[0030] In an optional embodiment of the present application, step S12 includes steps S121 to S123: S121. Perform a first perturbation process on the labeled image set to obtain a perturbed labeled image set.
[0031] To improve the robustness of the semantic segmentation model for images, before inputting the labeled images into the first basic network, it is necessary to perform a first perturbation process on the images, so as to effectively simulate diverse observation conditions in real life, reduce the risk of model overfitting and improve the robustness of the model, obtain the perturbed labeled images, and form a set of perturbed labeled images. It can be understood that for each image, only one type of first perturbation can be performed, or multiple different types of first perturbations can be performed on the image, which is not limited here. Among them, the ways of the first perturbation include, but are not limited to, randomly flipping the image horizontally or vertically; adjusting the saturation and hue of the image; injecting Gaussian noise into the image to simulate sensor interference during the image acquisition process; randomly cropping and pasting the image to improve the robustness of the model to occlusion and local structural changes.
[0032] It should be noted that considering that the images used in this application are all RGB images, therefore, performing the first perturbation process on the set of labeled images to obtain the set of perturbed labeled images includes the following process: First, perform a normalization process on the set of labeled images to obtain a normalized set of labeled images; perform the first perturbation process on the normalized set of signed images to obtain the set of perturbed labeled images. Specifically, perform a normalization process on all the labeled images in the set of labeled images, that is , where represents the original labeled image, h and w are the height and width of the image. In this embodiment, the width and height of the picture are , and its number of image channels is 3. is the labeled image after normalization, so that the RGB channel value range is finally limited between 0 and 1, which is beneficial to the convergence of the semantic segmentation model training. Then perform the first perturbation process on the normalized labeled images to obtain the set of perturbed labeled images.
[0033] S122. Input the set of perturbed labeled images into the first basic network, perform semantic segmentation on the target objects in the set of perturbed labeled images, and obtain a first target mask set.
[0034] Input the set of perturbed labeled images into the first basic network. For each input labeled image, perform pixel-level semantic segmentation on the target object in the image, generate the prediction probability of each pixel point in each semantic category, and for each pixel point: obtain the semantic category corresponding to the highest prediction probability of the pixel point after softmax normalization, summarize the semantic categories corresponding to all pixel points, and obtain a first target mask. For all the perturbed labeled images input in the current training batch, obtain the corresponding first target mask set.
[0035] In an alternative embodiment of the present application, step S122 includes the following process: First, input the perturbed labeled image set into the first basic network to perform semantic segmentation on the target objects in the perturbed labeled image set, and obtain the maximum class confidence of each pixel point of all the labeled images in the perturbed labeled image set.
[0036] For each perturbed labeled image, perform the following process: Input the perturbed labeled image into the first basic network to generate the prediction probability of each pixel point for each semantic class. After softmax normalization, obtain the highest prediction probability corresponding to each pixel point, and use it as the maximum class confidence of the pixel point.
[0037] After obtaining the maximum class confidence of each pixel point, calculate the confidence mean of each perturbed labeled image based on the maximum class confidence of each pixel point.
[0038] For each perturbed labeled image: Calculate the mean of the maximum class confidence of all its pixel points, and use it as the confidence mean of the image. This confidence mean is used to characterize the accuracy level of the semantic prediction of the first basic network for the entire image and can be used as the threshold basis for subsequent screening of difficult-to-train samples.
[0039] Then, for each pixel point in the perturbed labeled image: Determine whether the maximum class confidence of the pixel point is greater than the corresponding confidence mean. If so, use the class number corresponding to the maximum class confidence as the first target mask of the pixel point; otherwise, do not assign the first target mask to the pixel point.
[0040] For each pixel point in the perturbed labeled image, compare the maximum class confidence of the pixel with the confidence mean of the overall image. If the maximum class confidence of the pixel point is higher than or equal to the adaptive confidence mean of the current image, it is considered that the prediction result of the pixel point has a high credibility, and the class number corresponding to its maximum class confidence is used as the first target mask of the pixel point; otherwise, it is regarded as a low-confidence region, and to avoid interfering with training, the first target mask is not assigned to the pixel point. In this way, the pixel point regions in the supervised image that are difficult to distinguish or where the model is uncertain can be dynamically screened, thereby constructing a confidence mask map, enabling the model training to focus more on the sample regions with learning value, improving the segmentation ability for complex regions such as boundaries, occlusions, and weak contrasts, and enhancing the overall robustness and generalization performance of the model.
[0041] Finally, count the first target masks in all the perturbed labeled images to obtain the first target mask set.
[0042] Arrange the first target masks of all the perturbed labeled images according to the input order of the images to form the first target mask set.
[0043] S123: Input the perturbed labeled image set into the second basic network, perform semantic segmentation on the target objects in the perturbed labeled image set, and obtain a second target mask set.
[0044] The perturbed labeled image set is input into the second basic network. For each labeled image input, pixel-level semantic segmentation is performed on the target object in the image to generate the predicted probability of each pixel in each semantic category. For each pixel: the semantic category with the highest predicted probability corresponding to the pixel is obtained after softmax normalization, and the semantic categories corresponding to all pixels are summarized to obtain the second target mask. For all perturbed labeled images input in the current training batch, the corresponding second target mask set is obtained.
[0045] It is understandable that the acquisition of the second target mask set can also be achieved through the above-mentioned pixel confidence-based method, that is, after the second basic network completes the semantic segmentation of the perturbed labeled image, the maximum category confidence of each pixel is also calculated, and the image-level adaptive confidence mean is used for screening, and only pixels above the mean are retained as valid predictions in the second target mask, thereby constructing the second target mask set. The specific implementation method is the same as the acquisition method of the first target mask set mentioned above, and will not be described in detail here.
[0046] S13. Input the unlabeled image set into the first basic network, the second basic network, the first enhanced network and the second enhanced network respectively, perform semantic segmentation on the target objects in the unlabeled image set, and obtain a third target mask set, a fourth target mask set, a first enhanced pseudo-label set and a second enhanced pseudo-label set accordingly.
[0047] The unlabeled image set is input into the first basic network, the second basic network, the first enhanced network and the second enhanced network respectively to perform pixel-level semantic segmentation on the target objects in the image. The first basic network and the second basic network output the third target mask set and the fourth target mask set respectively to capture the network's prediction results for the unlabeled image. The first enhanced network and the second enhanced network output the first enhanced pseudo label set and the second enhanced pseudo label set to generate high-confidence semi-supervised signals to guide the basic network to perform semi-supervised training.
[0048] In an optional embodiment of the present application, the process of acquiring the first enhanced pseudo label set and the second enhanced pseudo label set is as follows: First, the unlabeled image set is subjected to a second perturbation process to obtain a perturbed unlabeled image set.
[0049] Perform a second perturbation process on the unlabeled image set to obtain the perturbed unlabeled image set. Among them, considering that the basic network can withstand stronger image changes under the guidance of labeled data, the intensity of the first perturbation is greater than that of the second perturbation, so that the basic network has better robustness and generalization ability.
[0050] Then input the perturbed unlabeled image set into the first enhancement network to perform semantic segmentation on the target objects in the perturbed unlabeled image set, and correspondingly obtain the first enhanced pseudo-label set.
[0051] Input the perturbed unlabeled images into the first enhancement network respectively to obtain the predicted probability distribution of each pixel point on each semantic category. After normalizing each predicted probability through softmax, for each pixel point: select the semantic category with the highest predicted probability as the predicted category of the currently selected pixel point. The predicted categories of all pixel points are sorted according to their order in the image to form the first enhanced pseudo-label of the image. For the perturbed unlabeled image set input into the first enhancement network in the current training batch, the first enhanced pseudo-label set is correspondingly generated.
[0052] Finally, input the perturbed unlabeled image set into the second enhancement network to perform semantic segmentation on the target objects in the perturbed unlabeled image set, and correspondingly obtain the second enhanced pseudo-label set.
[0053] Input the perturbed unlabeled image set into the second enhancement network to obtain the predicted probability distribution of each pixel point on all semantic categories. After normalizing each predicted probability through softmax. For each pixel point: select the semantic category with the highest predicted probability as the predicted category of the currently selected pixel point. The predicted categories of all pixel points are sorted according to their order in the image to form the second enhanced pseudo-label of the image. For the perturbed unlabeled image set input into the second enhancement network in the current training batch, the second enhanced pseudo-label set is correspondingly generated.
[0054] In an optional embodiment of the present application, after inputting the perturbed unlabeled image set into the second enhancement network, the process of obtaining the second enhanced pseudo-label set includes the following processing steps: First, input the perturbed unlabeled image set into the second enhancement network to perform semantic segmentation on the target objects in the perturbed unlabeled image set, and obtain the maximum class confidence of each pixel point of all unlabeled images in the perturbed unlabeled image set.
[0055] Input the set of perturbed unlabeled images into the second enhancement network. For each perturbed unlabeled image input: perform pixel-level semantic segmentation on the target objects therein to generate the predicted probability distribution of each pixel point in the unlabeled image over each semantic category. And perform normalization processing on the predicted probabilities through softmax, so as to obtain the normalized category probability values of each pixel point over all semantic categories, and extract the maximum probability value therefrom as the maximum category confidence of this pixel point, which is used to measure the prediction credibility of the model at this position.
[0056] Then, based on the maximum category confidence of each pixel point, calculate the confidence mean of each perturbed unlabeled image respectively.
[0057] Since the quality of the pseudo-labels during the semi-supervised training process determines the final prediction result of the model, there is no doubt that the pseudo-labels with high confidence are added to the training, while the pseudo-labels with low confidence will instead interfere with the training of the model. Therefore, the present invention screens out high-quality pseudo-labels to participate in the training and removes incorrect pseudo-labels by setting an adaptive threshold itself. Specifically, for each perturbed unlabeled image, perform the following process: accumulate the maximum category confidence of all pixel points in this unlabeled image, and divide it by the total number of pixel points in this image, so as to obtain the adaptive confidence mean of this unlabeled image as shown in formula (1):
[0058] Where , h and represent the height and width of the unlabeled image, represents the total number of semantic categories, represents the confidence mean of the r-th enhancement network on the current image, i and j are respectively the height index and width index of the pixel point in the unlabeled image, d is the category number to which the current pixel point belongs, is the forward output feature map of the r-th enhancement network for the unlabeled image , is the current network parameter of the r-th enhancement network, is the w-th unlabeled image in the set of unlabeled images.
[0059] Then, for each pixel point in the perturbed unlabeled image: determine whether the maximum category confidence of the pixel point is greater than or equal to the confidence mean: if so, use the category number corresponding to the maximum category confidence as the second enhancement pseudo-label of this pixel point; otherwise, do not assign a second enhancement pseudo-label to this pixel point.
[0060] For each perturbed unlabeled image, the following process is performed for each pixel in the image: Determine whether the maximum class confidence of the pixel is greater than the mean confidence of the image. If so, it means that the prediction result of the pixel is highly reliable. Therefore, the class number corresponding to the maximum class confidence is used as the second enhanced pseudo-label of the pixel. Conversely, if the maximum class confidence of the pixel is less than the mean confidence of the image, no second enhanced pseudo-label is assigned to the pixel, that is, the pixel does not participate in subsequent semi-supervised training. Here, not assigning a second enhanced pseudo-label means setting the second enhanced pseudo-label of the pixel to zero. Specifically, considering that the average prediction probability will increase during network training, to improve the training accuracy, only when the prediction probability value of a single pixel is greater than the average prediction probability is it considered a high-confidence pseudo-label. The r-th enhanced pseudo-label of the pixel is obtained by marking the high-confidence pixel positions , as shown in Equation (2):
[0061] where , , is the r-th enhanced pseudo-label of the pixel (i, j) in the unlabeled image
[0062] Finally, the second enhanced pseudo-labels in all perturbed unlabeled images are counted to obtain the second enhanced pseudo-label set
[0063] After all perturbed unlabeled images are processed according to the above process, the second enhanced pseudo-labels obtained by confidence judgment in all images in the perturbed unlabeled image set are integrated to construct the second enhanced pseudo-label set
[0064] Exemplarily, the methods for two enhanced networks to obtain pseudo-labels are the same. Use to represent the feature map output by the r-th enhanced network, and the calculation of obtaining the r-th enhanced pseudo-label is shown in Equation (3):
[0065] (3) where , is the r-th enhanced pseudo-label map generated by the r-th enhanced network is the -th enhanced network's output feature map for the w-th perturbed unlabeled image input, and the function corresponds to the channel index number corresponding to taking the maximum value in the channel direction, that is, the class to which it belongs The r-th enhanced pseudo-label corresponding to the pixel at the i-th row and j-th column of the unlabeled image, and there are a total of categories, , Applying softmax to the network prediction output in the category channel dimension to normalize the probabilities of each pixel over all categories.
[0066] S14. Calculate the difference degrees between the first target mask set and the label set, the difference degree between the third target mask set and the second enhanced pseudo-label set, and the difference degree between the fourth target mask set and the first enhanced pseudo-label set respectively to obtain the total difference degree.
[0067] For labeled images, compare the first target mask set generated by the first base network with the corresponding manually annotated label set pixel by pixel to obtain the first supervised difference degree. Similarly, compare the second target mask set generated by the second base network with the corresponding manually annotated label set pixel by pixel to obtain the second supervised difference degree. Among them, the first supervised difference degree and the second supervised difference degree can be calculated by means such as cross-entropy loss or weighted cross-entropy loss. Preferably, in an optional embodiment of the present application, the first supervised difference degree and the second supervised difference degree are cross-entropy losses.
[0068] To enhance the training robustness, an adaptive threshold is also set for labeled images to screen out difficult-to-train samples. First, calculate the adaptive supervised confidence threshold for labeled supervised training as shown in formula (4): (4) Where, is the adaptive supervised confidence threshold calculated by the q-th base network on the s-th labeled image, used to determine which pixel points belong to difficult-to-train samples, , is the prediction result of the q-th base network on the -th perturbed labeled image at the pixel point (i, j) for the prediction score of the d-th category, h and represent the height and width of the labeled image, is the -th labeled image in the labeled image set. According to the above adaptive supervised confidence threshold, the q-th target mask map corresponding to the difficult-to-train samples in the labeled image set can be calculated as shown in formulas (5) and (6): (5) (6) Where, is the q-th target mask map generated by the q-th base network, used to mark which pixels in the current labeled image belong to difficult-to-train samples, is the average confidence of the labeled image under the output of the q-th base network (calculated in the same way as the average confidence of the unlabeled image), is the q-th target mask value of the pixel at the i-th row and j-th column of the s-th labeled image.
[0069] The q-th supervised difference degree of the final labeled image is obtained by using the labeled hard-to-train sample corresponding labeled mask as shown in formula (7): (7) where, is the q-th supervised difference degree, is the set of labeled images, is the -th labeled image in the set of labeled images, is the weighted pixel-level supervised difference degree calculated by the q-th base network on the s-th labeled image.
[0070] For unlabeled images, the first augmented pseudo-label set generated by the first augmented network and the fourth target mask set generated by the second base network are compared pixel by pixel to obtain the first semi-supervised difference degree. Similarly, the second augmented pseudo-label set generated by the second augmented network and the third target mask set generated by the first base network are compared pixel by pixel to obtain the second semi-supervised difference degree. Among them, the first semi-supervised difference degree and the second semi-supervised difference degree are used to guide the learning direction of the base network on unsupervised samples. The semi-supervised difference degree also uses an adaptive threshold to generate a pseudo-label mask. Specifically, the labeled mask is used to obtain the semi-supervised difference degree finally participating in model training, as shown in formula (8): (8) where, is the r-th semi-supervised difference degree, is the set of unlabeled images, is the s-th unlabeled image in the set of unlabeled images, h and w are the height and width of the unlabeled image respectively, is the weighted difference degree of the r-th base network, is the r-th augmented pseudo-label. Finally, the total loss of model training is , where represents the coefficient of the semi-supervised difference degree, and then the total loss is backpropagated to update the parameters of the two base networks.
[0071] To ensure the rationality and class balance of the participation of pseudo-labels in training, this embodiment further conducts a fine-grained analysis of the pseudo-label difference degree and introduces a class weighting mechanism to improve the segmentation accuracy of the model for minority classes and the overall training stability. Specifically, in an optional embodiment of the present invention, calculating the difference degree between the third target mask set and the second enhanced pseudo-label set includes the following process: First, determine the classes of each second enhanced pseudo-label in the second enhanced pseudo-label set.
[0072] For each pseudo-label map in the second enhanced pseudo-label set, determine the class numbers to which all the pixel points with assigned pseudo-labels belong. Through this process, the distribution of each class in the pseudo-label map in the second enhanced pseudo-label set can be obtained, providing a basis for class weighting in subsequent difference degree calculations.
[0073] Then calculate the class proportion of each type of second enhanced pseudo-label in all the second enhanced pseudo-labels in the second enhanced pseudo-label set.
[0074] Traverse each pseudo-label map in the second enhanced pseudo-label set, count all the pixel points with assigned class labels, and count them separately according to the class numbers to obtain the number of pixels corresponding to each semantic class. Calculate the ratio of the number of pixels of each class to the total number of all labeled pixels to obtain the proportion of this class in the second enhanced pseudo-label set.
[0075] Then, for each type of second enhanced pseudo-label in the second enhanced pseudo-label set: calculate the first semi-supervised difference degree between the third target mask set and this type of second enhanced pseudo-label.
[0076] Extract all the pixel positions corresponding to the current semantic class from the second enhanced pseudo-label set to construct a pseudo-label mask map for this class. Then extract the pixel prediction results corresponding to the above positions from the segmentation results of the corresponding images in the third target mask set and compare them pixel by pixel with the pseudo-label of this class. For each class, use a standard metric function such as cross-entropy loss or weighted cross-entropy loss to calculate the difference between the third target mask and the second enhanced pseudo-label corresponding to this pixel point to obtain the first semi-supervised difference degree. This difference degree is used to measure the learning effect of the basic network in the pseudo-label area of this class and will be weighted according to the class proportion in the subsequent difference fusion process to enhance the segmentation ability of the model for low-frequency classes and the overall training balance.
[0077] Exemplarily, the pseudo-labels generated by the first enhanced network are compared with the output of the second basic network to calculate the second semi-supervised difference degree of the second basic network. Similarly, the pseudo-labels generated by the second enhanced network are compared with the output of the first basic network Compare and calculate the first semi-supervised difference degree of the first basic network, as shown in formula (9): (9) Wherein, is the initial r-th semi-supervised difference degree of the entire unlabeled image, is the r-th enhanced pseudo-label map generated by the r-th enhanced network, is the output of the q-th basic network, is the output of the q-th basic network for the s-th unlabeled image after the second perturbation , is the parameter of the q-th basic network, , and , C is the total number of categories, and the operator represents the multiplication of elements at the corresponding positions of the matrices, , is the initial r-th semi-supervised difference degree of the unlabeled image at the pixel point (i, j), is the r-th enhanced pseudo-label corresponding to the pixel point in the i-th row and j-th column of the unlabeled image, is the output of the q-th basic network for the s-th unlabeled image after the second perturbation in the predicted value of the c-th category. The first enhanced pseudo-label generated by the first enhanced network guides the training of the second basic network, and the second enhanced pseudo-label generated by the second enhanced network guides the training of the first enhanced network, forming cross-contrast training with each other.
[0078] Finally, the first semi-supervised difference degree is weighted based on the class ratio to obtain the final difference degree.
[0079] To further alleviate the impact of sample class imbalance on training, a dynamic weighting strategy is also introduced in the loss function, that is, in each batch: according to the number of pixels of each category in the pseudo-label map, calculate its proportion in the whole image , and use it as the inverse weight, which acts on the original cross-entropy loss to obtain the weighted difference degree of the r-th basic network, as shown in formulas (10) and (11): (10) = (11) Wherein, represents the class ratio of the q-th enhanced pseudo-label to which each pixel point of the current image belongs to all the q-th enhanced pseudo-labels, , and , is the weighted difference degree of the r-th basic network at the pixel point (i, j), is the initial r-th semi-supervised difference degree of the entire unlabeled image, The number of pixels belonging to class d in the r-th enhanced pseudo-label for all pixels in the image is the class ratio of the q-th enhanced pseudo-label at pixel (i, j) to all q-th enhanced pseudo-labels. This method is adaptively updated in each round of training, effectively avoiding the problem of the dominant class dominating the training.
[0080] For labeled images, the supervised loss weighting can also be adjusted in the same way, and its loss is shown in Equation (12): (12) where is the total difference degree of the q-th basic network for the s-th labeled image is the true class label of the s-th labeled image is the output of the q-th basic network for the s-th labeled image and is the difference degree of the q-th basic network at pixel (i, j) of the labeled image. Since the proportion of class pixels in each training image is different, in each the network can adaptively adjust the loss weighting coefficient. When there are few training samples of a certain class, the loss weight is increased, and when there are many, it is decreased, effectively alleviating the performance problem caused by the imbalance of training samples. Then , is the weighted pixel-level supervised difference degree calculated by the q-th basic network on the labeled image represents the class ratio of the class to which each pixel point of the current labeled image belongs in the whole image is is the weighted pixel-level supervised difference degree calculated by the q-th basic network at pixel (i, j) of the labeled image. Then the q-th supervised difference degree can be calculated, as shown in Equation (13): (13) where is the q-th supervised difference degree of the s-th labeled image is the set of labeled images is the s-th labeled image in the set of labeled images, and h and w are the height and width of the labeled image respectively is the weighted pixel-level supervised difference degree calculated by the q-th basic network on the s-th labeled image is the q-th target mask generated by the q-th basic network for the s-th labeled image
[0081] S15. Adjust the parameters of the first basic network and the second basic network based on the total difference degree to complete the training of the semantic segmentation model; wherein, the semantic segmentation model includes the trained first basic network and / or the second basic network.
[0082] The main role of the enhancement network is to dynamically generate pseudo-labels for unlabeled images. In each iteration, a number of labeled images and a number of unlabeled images are mixed to update the basic network, and the parameter values of the basic network are used to update the corresponding parameter values of the enhancement network through the exponential moving average strategy. In the semantic segmentation model, the basic network updates its parameters through the gradient backpropagation of the loss value, and the enhancement network does not participate in the backpropagation. The parameter values of the basic network are used to update the corresponding parameter values of the enhancement network through the exponential moving average strategy. The specific steps are as follows: The parameter of the enhancement network at the current iteration moment is the weighted sum of the parameter of the basic network at the current iteration moment and the parameter of the enhancement network at the previous iteration. The parameter update formula of the enhancement network is: ; where represents the parameter of the th enhancement network a at the current iteration, represents the parameter of the th basic network b at the current iteration, is the step size for controlling the momentum update. According to the pre-designed training strategy, after 100 iterations, the model is allowed to learn more information hidden behind the scene images by using a small number of labeled images and a large number of unlabeled images to improve the final detection accuracy and obtain a more accurate segmentation result.
[0083] Furthermore, in an optional embodiment of the present application, the semantic segmentation model is obtained by weighting the trained first basic network and the trained second basic network.
[0084] As a training strategy, the parameter optimizer can be set to , the initial learning rate is set to , the momentum is set to 0.9, and the weight decay factor is ; in the specific implementation of the training strategy, the size of is set to 5, which includes 2 unlabeled images and 3 labeled images. 100 iterations are set. After the 20th iteration, the accuracy is calculated on the validation set after each iteration and the model is retained. After each round of iteration, the accuracy is calculated and compared with the accuracy of the previous model. If the accuracy of the latter model exceeds that of the previous model, the previous model is replaced; the accuracy is based on The intersection over union (IoU) is used as the evaluation criterion, that is, the ratio of the intersection of the ground truth of the region to be measured and the region predicted by the model to the union is used as the evaluation criterion for measuring the effectiveness of the model. It can be understood that the various parameters in the above training strategy can be adaptively changed based on the actual task requirements and the accuracy requirements of the model, and are not limited here.
[0085] In summary, as Figure 2 shown, the normalized input images are respectively subjected to the first perturbation and the second perturbation, and then sent into the basic network and the enhancement network. The enhancement network processes the second perturbed image to generate the corresponding feature map. Similarly, the basic network also performs semantic segmentation on the first perturbed image to obtain the corresponding feature map. Then, the second supervision loss is calculated using the first enhanced pseudo-label and the feature map output by the second basic network, and the first supervision loss is calculated using the second enhanced pseudo-label and the feature map output by the first basic network. The final total loss updates the two basic networks and, accordingly, updates the corresponding enhancement networks.
[0086] Therefore, this application uses means such as semantic segmentation, semi-supervised training, data perturbation, exponential moving average strategy, loss weighting, and contrastive cross-training to achieve real-time semantic segmentation of images, improve the semantic segmentation accuracy of scene images, and solve the problems of large manual labeling costs for semantic segmentation tasks and unbalanced training samples for scene images. Based on the comparative training of two basic networks and two enhancement networks, the enhancement network generates pseudo-labels for unlabeled images. The semi-supervised loss is calculated using the pseudo-labels of one enhancement network and the feature map of the unlabeled image output by the other basic network, and the full-supervised loss is calculated using the labeled images and the manual labels. The class loss values are dynamically adjusted according to the number of each category of training samples to alleviate the problem of unbalanced training samples. Pseudo-labels with high confidence and difficult-to-distinguish labeled images are selected for backpropagation of the gradient. After the parameters of the basic network are updated according to the gradient, the parameters of the enhancement network are updated by momentum, and finally a segmentation map with higher accuracy is obtained.
[0087] As Figure 3 shown, this application also provides a semantic segmentation method, which includes: S31. Obtain the image to be segmented; S32. Input the image into the semantic segmentation model to obtain the target mask for segmenting the target object in the image; wherein, the semantic segmentation model is trained by the semi-supervised semantic segmentation model training method based on the weighted contrast network described in any one of the above. S33. Segment the image based on the target mask.
[0088] When an image needs to be segmented, the image to be segmented is input into a semantic segmentation model to predict the semantic categories of each pixel point in the image, and a target mask is obtained based on the prediction result. Then, the target object in the image is segmented using the target mask.
[0089] As Figure 4 shown, the training system 400 for a semi-supervised semantic segmentation model based on a weighted contrast network includes: an image acquisition module 410, a supervised module 420, an unsupervised module 430, a loss calculation module 440, and a parameter update module 450. The above-mentioned image acquisition module 410 is used to acquire a labeled image set containing a target object and a corresponding label set, and an unlabeled image set containing a target object and without labels; among them, the number of unlabeled images is much larger than the number of labeled images. The supervised module 420 is used to input the labeled image set into the first basic network of the first branch and the second basic network of the second branch respectively, perform semantic segmentation on the target object in the labeled image set, and correspondingly obtain a first target mask set and a second target mask set; among them, the first branch includes a first basic network and a first enhancement network, the second branch includes a second basic network and a second enhancement network, and the structures of the first enhancement network, the second enhancement network, the first basic network, and the second basic network are the same but the initial parameters are different. The unsupervised module 430 is used to input the unlabeled image set into the first basic network, the second basic network, the first enhancement network, and the second enhancement network respectively, perform semantic segmentation on the target object in the unlabeled image set, and correspondingly obtain a third target mask set, a fourth target mask set, a first enhanced pseudo-label set, and a second enhanced pseudo-label set. The loss calculation module 440 is used to calculate the difference degrees between the first target mask set, the second target mask set and the label set, the difference degree between the third target mask set and the second enhanced pseudo-label set, and the difference degree between the fourth target mask set and the first enhanced pseudo-label set, and obtain the total difference degree. The parameter update module 450 is used to adjust the parameters of the first basic network and the second basic network based on the total difference degree to complete the training of the semantic segmentation model; among them, the semantic segmentation model includes the trained first basic network and / or the second basic network.
[0090] For the specific limitations of the training system for a semi-supervised semantic segmentation model based on a weighted contrast network, reference can be made to the limitations for the training method of a semi-supervised semantic segmentation model based on a weighted contrast network in the above text, which will not be elaborated here. Each module in the above-mentioned training system for a semi-supervised semantic segmentation model based on a weighted contrast network can be implemented in whole or in part by software, hardware, and their combinations. The above-mentioned modules can be embedded in the processor of a computer device in a hardware format or independent of it, or stored in the memory of a computer device in a software format, so as to facilitate the processor to call the corresponding operations of the above-mentioned modules.
[0091] It should be noted that, in order to highlight the innovative part of this application, modules that are not closely related to solving the technical problems proposed in this application are not introduced in this embodiment. However, this does not mean that there are no other modules in this embodiment.
[0092] As Figure 5 shown, the electronic device 5 may include a memory 51, a processor 52, and a bus, and may also include a computer program stored in the memory 51 and executable on the processor 52, such as a training program for a semi-supervised semantic segmentation model based on a weighted contrast network.
[0093] Among them, the memory 51 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as: SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. The memory 51 may be an internal storage unit of the electronic device 5 in some embodiments, such as the mobile hard disk of the electronic device 5. The memory 51 may also be an external storage device of the electronic device 5 in other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 5. Further, the memory 51 may also include both an internal storage unit and an external storage device of the electronic device 5. The memory 51 can not only be used to store application software installed on the electronic device 5 and various types of data, such as code trained based on a semi-supervised semantic segmentation model of a weighted contrast network, but also be used to temporarily store data that has been output or will be output.
[0094] The processor 52 may be composed of integrated circuits in some embodiments. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 52 is the control core (Control Unit) of the electronic device 5, connecting various components of the entire electronic device 5 through various interfaces and lines. By running or executing programs or modules stored in the memory 51 (such as a training program for a semi-supervised semantic segmentation model based on a weighted contrast network, etc.), and by calling data stored in the memory 51, it executes various functions of the electronic device 5 and processes data.
[0095] The processor 52 executes the operating system of the electronic device 5 and various installed application programs. The processor 52 executes the application programs to implement the steps in the above-mentioned method for training a semi-supervised semantic segmentation model based on a weighted contrast network.
[0096] Exemplarily, a computer program can be divided into one or more modules. One or more modules are stored in the memory 51 and executed by the processor 52 to complete the present application. One or more modules can be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device 5. For example, the computer program can be divided into an image acquisition module 410, a supervised module 420, an unsupervised module 430, a loss calculation module 440, and a parameter update module 450.
[0097] The integrated units implemented in the form of software function modules as described above can be stored in a computer-readable storage medium. The computer-readable storage medium can be non-volatile or volatile. The above-mentioned software function modules are stored in a storage medium and include several instructions to enable a computer device (which can be a personal computer, a computer device, or a network device, etc.) or a processor to execute part of the functions of the training method of the semi-supervised semantic segmentation model in various embodiments of the present application.
[0098] In summary, for a semi-supervised semantic segmentation and model training method and device disclosed in the present application, a weighted contrast semi-supervised network segmentation model is designed. In the case of a small number of labeled scene images and a large number of unlabeled scene images, a semi-supervised training method using consistency regularization and pseudo-labels is adopted to obtain a semantic segmentation model with good effects. At the same time, the average exponential moving strategy is used to update the parameters of the basic network momentum to enhance the parameters of the network, and in turn, the enhanced network can generate high-confidence pseudo-labels to contrast and guide the training of the basic network. At the same time, in order to avoid the problem of sample imbalance during training, the semantic segmentation model can automatically adjust the weight coefficient to stabilize the training. Moreover, the present application uses high-confidence pseudo-labels and difficult-to-distinguish training samples to further improve the model robustness, and the present application can be embedded in a monitoring system camera to perform real-time image semantic segmentation. Through the configuration of the monitoring camera, the image semantic segmentation inference algorithm can be flexibly adjusted to achieve higher semantic segmentation accuracy. Therefore, the present application effectively overcomes various disadvantages in the prior art and has high industrial utilization value.
[0099] The above embodiments are only illustrative of the principles and effects of the present application, and are not used to limit the present application. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present application. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by the present application should still be covered by the claims of the present application.
Claims
1. A training method for a semi-supervised semantic segmentation model based on a weighted contrast network, characterized in that, The training method includes: Obtaining a labeled image set containing a target object and a corresponding label set, as well as an unlabeled image set containing the target object without labels; wherein, the number of unlabeled images is much larger than the number of labeled images; Inputting the labeled image set into a first basic network of a first branch and a second basic network of a second branch respectively, performing semantic segmentation on the target object in the labeled image set, and correspondingly obtaining a first target mask set and a second target mask set; wherein, the first branch includes a first basic network and a first enhancement network, the second branch includes a second basic network and a second enhancement network, and the structures of the first enhancement network, the second enhancement network, the first basic network, and the second basic network are the same but the initial parameters are different; Inputting the unlabeled image set into the first basic network, the second basic network, the first enhancement network, and the second enhancement network respectively, performing semantic segmentation on the target object in the unlabeled image set, and correspondingly obtaining a third target mask set, a fourth target mask set, a first enhanced pseudo-label set, and a second enhanced pseudo-label set; Calculating the difference degrees between the first target mask set, the second target mask set and the label set, the difference degree between the third target mask set and the second enhanced pseudo-label set, and the difference degree between the fourth target mask set and the first enhanced pseudo-label set respectively, to obtain the total difference degree; Adjusting the parameters of the first basic network and the second basic network based on the total difference degree to complete the training of the semantic segmentation model; wherein, the semantic segmentation model includes the trained first basic network and / or the second basic network.
2. The training method of the semi-supervised semantic segmentation model based on the weighted contrast network according to claim 1, wherein The step of inputting the labeled image set into a first basic network of a first branch and a second basic network of a second branch respectively, performing semantic segmentation on the target object in the labeled image set, and correspondingly obtaining a first target mask set and a second target mask set includes: Performing a first perturbation process on the labeled image set to obtain a perturbed labeled image set; Inputting the perturbed labeled image set into the first basic network, performing semantic segmentation on the target object in the perturbed labeled image set, and obtaining a first target mask set; Inputting the perturbed labeled image set into the second basic network, performing semantic segmentation on the target object in the perturbed labeled image set, and obtaining a second target mask set.
3. The training method of the semi-supervised semantic segmentation model based on the weighted contrast network according to claim 2, wherein, The step of inputting the perturbed labeled image set into the first basic network, performing semantic segmentation on the target object in the perturbed labeled image set, and obtaining a first target mask set includes: Inputting the perturbed labeled image set into the first basic network, performing semantic segmentation on the target object in the perturbed labeled image set, and obtaining the maximum class confidence of each pixel point of all the labeled images in the perturbed labeled image set; Based on the maximum class confidence of each pixel point, calculating the confidence mean of each perturbed labeled image respectively; For each pixel point in the perturbed labeled image: Judging whether the maximum class confidence of the pixel point is greater than or equal to the corresponding confidence mean: If so, use the class number corresponding to the maximum class confidence as the first target mask for this pixel point; Otherwise, do not assign a first target mask to this pixel point; Count the first target masks in all the perturbed labeled images to obtain a first target mask set.
4. The training method of the semi-supervised semantic segmentation model based on the weighted contrast network according to claim 1, wherein, Input the set of unlabeled images into the first enhancement network and the second enhancement network respectively, perform semantic segmentation on the target objects in the set of unlabeled images, and correspondingly obtain a first enhanced pseudo-label set and a second enhanced pseudo-label set, including: Perform a second perturbation process on the set of unlabeled images to obtain a perturbed set of unlabeled images; Input the perturbed set of unlabeled images into the first enhancement network, perform semantic segmentation on the target objects in the perturbed set of unlabeled images, and correspondingly obtain a first enhanced pseudo-label set; Input the perturbed set of unlabeled images into the second enhancement network, perform semantic segmentation on the target objects in the perturbed set of unlabeled images, and correspondingly obtain a second enhanced pseudo-label set.
5. The training method of the semi-supervised semantic segmentation model based on the weighted contrast network according to claim 4, wherein, The step of inputting the perturbed set of unlabeled images into the second enhancement network, performing semantic segmentation on the target objects in the perturbed set of unlabeled images, and correspondingly obtaining a second enhanced pseudo-label set includes: Input the perturbed set of unlabeled images into the second enhancement network, perform semantic segmentation on the target objects in the perturbed set of unlabeled images, and obtain the maximum class confidence of each pixel point of all unlabeled images in the perturbed set of unlabeled images; Based on the maximum class confidence of each pixel point, calculate the confidence mean value of each perturbed unlabeled image respectively; For each pixel point in the perturbed unlabeled image: Determine whether the maximum class confidence of the pixel point is greater than the confidence mean value: If so, use the class number corresponding to the maximum class confidence as the second enhanced pseudo-label for this pixel point; Otherwise, do not assign a second enhanced pseudo-label to this pixel point; Count the second enhanced pseudo-labels in all the perturbed unlabeled images to obtain a second enhanced pseudo-label set.
6. The training method of the semi-supervised semantic segmentation model based on the weighted contrast network according to claim 1, characterized in that Calculate the difference degree between the third target mask set and the second enhanced pseudo-label set, including: Determine the classes of each second enhanced pseudo-label in the second enhanced pseudo-label set; Calculate the class proportion of each class of second enhanced pseudo-label in all second enhanced pseudo-labels in the second enhanced pseudo-label set: For each class of second enhanced pseudo-label in the second enhanced pseudo-label set: calculate the first semi-supervised difference degree between the third target mask set and this class of second enhanced pseudo-label; Weight the first semi-supervised difference degree based on the class proportion to obtain the final difference degree.
7. The training method of the semi-supervised semantic segmentation model based on the weighted contrast network according to claim 1, characterized in that The semantic segmentation model is obtained by weighting a trained first basic network and a trained second basic network.
8. A semantic segmentation method, characterized in that, The semantic segmentation method includes: Obtain an image to be segmented; Input the image into the semantic segmentation model to obtain a target mask for segmenting the target object in the image; wherein, the semantic segmentation model is trained by the semi-supervised semantic segmentation model training method based on a weighted contrast network according to any one of claims 1-7; Segment the image based on the target mask.
9. A semi-supervised semantic segmentation model training system based on a weighted contrast network, characterized in that, The training system includes: An image acquisition module, configured to acquire a labeled image set containing a target object and a corresponding label set, as well as an unlabeled image set containing the target object without labels; wherein the number of unlabeled images is much larger than the number of labeled images; A supervised module, configured to input the labeled image set into a first basic network of a first branch and a second basic network of a second branch respectively, perform semantic segmentation on the target object in the labeled image set, and correspondingly obtain a first target mask set and a second target mask set; wherein the first branch includes a first basic network and a first enhancement network, the second branch includes a second basic network and a second enhancement network, and the structures of the first enhancement network, the second enhancement network, the first basic network, and the second basic network are the same but the initial parameters are different; An unsupervised module, configured to input the unlabeled image set into the first basic network, the second basic network, the first enhancement network, and the second enhancement network respectively, perform semantic segmentation on the target object in the unlabeled image set, and correspondingly obtain a third target mask set, a fourth target mask set, a first enhanced pseudo-label set, and a second enhanced pseudo-label set; A loss calculation module, configured to calculate the difference degrees between the first target mask set, the second target mask set and the label set, the difference degree between the third target mask set and the second enhanced pseudo-label set, and the difference degree between the fourth target mask set and the first enhanced pseudo-label set respectively, to obtain a total difference degree; A parameter update module, configured to adjust the parameters of the first basic network and the second basic network based on the total difference degree to complete the training of the semantic segmentation model; wherein the semantic segmentation model includes the trained first basic network and / or the second basic network.
10. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device, configured to store one or more programs, which when executed by the one or more processors, cause the electronic device to implement the semi-supervised semantic segmentation model training method based on a weighted contrast network as described in any one of claims 1 to 7 or the semantic segmentation method as described in claim 8.
11. A computer-readable storage medium, characterized in that, A computer program is stored thereon, which when executed by a processor of a computer, causes the computer to execute the semi-supervised semantic segmentation model training method based on a weighted contrast network as described in any one of claims 1 to 7 or the semantic segmentation method as described in claim 8.
Citation Information
Patent Citations
Semi-supervised SAR (Synthetic Aperture Radar) image building area extraction method based on time phase consistency pseudo tag
CN114821337A
Heterogeneous double-branch voting semi-supervised image segmentation method
CN118279332A
Semi-supervised confrontation mutual training semantic segmentation method based on strong and weak consistency
CN118447256A
Semi-supervised semantic segmentation method and device of image, equipment and storage medium
CN118506005A
Weak annotation remote sensing image semantic segmentation method based on double learning mechanism
CN118840553A
Cited By
Data enhancement method, fault detection model training method and device, and medium
CN121479319A