A zero-shot semantic segmentation method

By training a weakly supervised semantic segmentation model and a classification layer weight prediction network, the problem of inaccurate connection between word vectors and visual features in zero-shot semantic segmentation is solved, achieving accurate segmentation of unseen categories and realizing the effect of zero-shot semantic segmentation.

CN117058394BActive Publication Date: 2025-10-28SOUTHWEST JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311132788.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-04
Publication Date
2025-10-28
Estimated Expiration
2043-09-04

AI Technical Summary

Technical Problem

In zero-shot semantic segmentation, existing techniques struggle to establish accurate connections between word vectors and visual features, resulting in insufficient model generalization ability and an inability to effectively identify unseen categories.

Method used

By training a weakly supervised semantic segmentation model using a large-scale image classification dataset, performing pixel-level pseudo-annotation, collecting word vectors, and training a classification layer weight prediction network to replace the classification layer weights of the semantic segmentation model, zero-shot semantic segmentation is achieved.

Benefits of technology

It achieves accurate semantic segmentation of categories in a new dataset without training on the target category, solving the problem of insufficient model generalization ability and realizing zero-shot semantic segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117058394B_ABST
    Figure CN117058394B_ABST
Patent Text Reader

Abstract

This invention relates to a zero-shot semantic segmentation method, belonging to the field of image segmentation. To overcome the shortcomings of existing technologies, this invention aims to provide a zero-shot semantic segmentation method, which includes using a large-scale image classification dataset and training a weakly supervised semantic segmentation model to perform pixel-level pseudo-annotation on the images. The semantic segmentation model is then trained using the images and pixel-level pseudo-annotations, collecting category word vectors and training a classification layer weight prediction network. Finally, category word vectors are collected from the target dataset, where the categories can be completely disjoint from those of the training dataset. These vectors are input into the trained classification layer weight prediction network, and the output replaces the classification layer weights of the semantic segmentation model, achieving zero-shot semantic segmentation. This invention addresses the problem of insufficient samples of the categories required for inference in the training dataset by proposing to train a classification layer weight prediction network to replace the classification layer weights of the semantic segmentation model, thus achieving zero-shot image segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a zero-shot semantic segmentation method, belonging to the field of image segmentation. Background Technology

[0002] In zero-shot semantic segmentation, the segmentation model needs to identify pixels corresponding to categories that were not seen during training. This means that even without samples to train the model to identify the category, it still needs to be able to segment it accurately. This requires the model to have a certain degree of generalization ability, so that it can learn general features from known categories and apply them to the segmentation of unknown categories.

[0003] Secondly, to achieve zero-shot semantic segmentation, it is necessary to establish a connection between word vectors and visual features. Word vectors represent the semantic information of each category. In zero-shot semantic segmentation, these vectors need to correspond to visual features so that the model can accurately segment unknown categories. However, in practice, establishing a connection between word vectors and visual features is not always easy because these vectors may contain noise or inaccurate information. Therefore, appropriate methods are needed to ensure that the connection between word embeddings and visual features is accurate and reliable.

[0004] The root cause of these problems lies in the lack of category information. Without a sufficient number of categories and samples for training, the model may not possess adequate generalization ability or be able to learn general features. Therefore, when designing zero-shot semantic segmentation models, these challenges need to be considered, and appropriate methods should be adopted to overcome them. Summary of the Invention

[0005] In order to overcome the shortcomings of the existing technology, the present invention aims to provide a zero-sample semantic segmentation method.

[0006] The technical solution provided by this invention to solve the above-mentioned technical problems is: a zero-shot semantic segmentation method, comprising:

[0007] Step S1: Obtain a large-scale image classification dataset containing several categories as the training set;

[0008] Step S2: Train the weakly supervised semantic segmentation model using the training set to obtain the trained weakly supervised semantic segmentation model;

[0009] Step S3: Use the trained weakly supervised semantic segmentation model to annotate the large-scale image classification dataset with pixel-level pseudo-labels;

[0010] Step S4: Use the image data and pixel-level pseudo-labels as a new dataset to train the semantic segmentation model and obtain the trained semantic segmentation model.

[0011] Step S5: Collect word vectors corresponding to each category in a large-scale image classification dataset;

[0012] Step S6: Take the word vectors corresponding to each class of the large-scale image classification dataset as input, the classification layer weights of the semantic segmentation model as supervision information, and the output is the predicted classification layer weights. Train the classification layer weight prediction network model to obtain the trained classification layer weight prediction network model.

[0013] Step S7: Obtain the categories and corresponding word vectors of the inference dataset;

[0014] Step S8: Input the word vectors corresponding to the inference dataset into the trained classification layer weight prediction network model to obtain the predicted classification layer weights;

[0015] Step S9: Replace the classification layer weights of the trained semantic segmentation model with the predicted classification layer weights mentioned above.

[0016] Step S10: Use the semantic segmentation model after replacing the classification layer weights to predict the semantic segmentation labels of samples in the inference dataset to achieve zero-sample semantic segmentation.

[0017] A further technical solution is that the large-scale image classification dataset in step S1 includes, but is not limited to, ImageNet and CoCo datasets.

[0018] A further technical solution is that the weakly supervised semantic segmentation model is an IRN weakly supervised model.

[0019] A further technical solution is that the specific training process in step S2 is as follows: the weakly supervised semantic segmentation model is initialized by taking the training set as input, using the average error loss as the loss function, and then using the stochastic gradient descent algorithm for backpropagation to update the model parameters, thereby obtaining the trained weakly supervised semantic segmentation model.

[0020] A further technical solution is that the semantic segmentation model is the DeepLabV3 semantic segmentation model.

[0021] A further technical solution is that the specific training process in step S4 is as follows: using a new dataset as input, using the cross-entropy loss of each pixel as the loss function, and then using the stochastic gradient descent algorithm for backpropagation to obtain a trained semantic segmentation model.

[0022] A further technical solution is that the classification layer weight prediction network model consists of 8 fully connected layers, with the number of neurons being 2048, 2048, 1024, 1024, 1024, 1024, 512, and 512 respectively. The number of neurons in the last fully connected layer is the number of predicted categories. The network uses ReLU as the activation function, Adam as the optimizer, and the initial learning rate is set to 0.001.

[0023] A further technical solution is that the specific training process in step S6 is as follows: the word vectors corresponding to each class of the image classification dataset are used as input, the classification layer weights of the semantic segmentation model are used as supervision, the output is the predicted classification layer weights, cross-entropy loss is used as the loss function, and then the stochastic gradient descent algorithm is used for backpropagation to obtain the trained classification layer weight prediction network model.

[0024] A further technical solution is that the categories of the inference dataset in step S7 may include categories that do not exist in the training set.

[0025] This invention offers the following advantages: It can classify categories in a new dataset without prior training on the target category, achieving zero-shot semantic segmentation. It utilizes a large-scale image classification dataset and trains a weakly supervised semantic segmentation model to perform pixel-level pseudo-annotation on the images. The semantic segmentation model is then trained using the images and pixel-level pseudo-annotations, collecting category word vectors and training a classification layer weight prediction network. Finally, category word vectors from the target dataset are collected; the categories in the target dataset can be completely disjoint from those in the training dataset. These vectors are input into the trained classification layer weight prediction network, and the output replaces the classification layer weights of the semantic segmentation model, achieving zero-shot semantic segmentation. This invention addresses the problem of insufficient samples of the categories required for inference in the training dataset by proposing a training classification layer weight prediction network to replace the classification layer weights of the semantic segmentation model, thus achieving zero-shot image segmentation. Attached Figure Description

[0026] Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0027] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0028] like Figure 1 As shown, a zero-shot semantic segmentation method of the present invention includes the following steps:

[0029] Step S1: Obtain a large-scale image classification dataset containing several categories. The image dataset includes, but is not limited to, datasets such as ImageNet and CoCo, and uses it as the training set.

[0030] Step S2: Initialize the weakly supervised semantic segmentation model (including but not limited to the IRN weakly supervised model). Take a large-scale image classification dataset as input, use the average error loss as the loss function, and then use the stochastic gradient descent algorithm for backpropagation to update the model parameters and obtain the trained weakly supervised semantic segmentation model.

[0031] Step S3: Use the trained weakly supervised semantic segmentation model to annotate the large-scale image classification dataset with pixel-level pseudo-labels;

[0032] Step S4: Use the image data and pixel-level pseudo-labels as a new dataset to train a semantic segmentation model (including but not limited to DeepLabV3 semantic segmentation model). Use the new dataset as input, use the cross-entropy loss of each pixel as the loss function, and then use the stochastic gradient descent algorithm for backpropagation to obtain the trained semantic segmentation model.

[0033] Step S5: Collect word vectors corresponding to each category in a large-scale image classification dataset;

[0034] A classification layer weight prediction network model consisting of multiple fully connected layers is constructed. The classification layer weight prediction network consists of 8 fully connected layers with the following number of neurons: 2048, 2048, 1024, 1024, 1024, 1024, 512, 512. The number of neurons in the last fully connected layer is the number of predicted classes. The network uses ReLU as the activation function, Adam as the optimizer, and the initial learning rate is set to 0.001. The learning rate cosine annealing algorithm is used to enable the model to converge as quickly as possible.

[0035] Step S6: Take the word vectors corresponding to each category of the large-scale image classification dataset as input, the classification layer weights of the semantic segmentation model as supervision information, and the output is the predicted classification layer weights. Use cross-entropy loss as the loss function, and then use the stochastic gradient descent algorithm for backpropagation to obtain the trained classification layer weight prediction network model.

[0036] Step S7: Collect the categories and corresponding word vectors of the dataset to be inferred;

[0037] These word vectors can be extracted from text data using natural language processing techniques. In this step, it's necessary to ensure a clear correspondence between the word vectors and the dataset categories.

[0038] Step S8: Input these word vectors into the classification layer weight prediction network model to obtain the predicted classification layer weights.

[0039] Step S9: Replace the classification layer weights of the semantic segmentation model with the predicted classification layer weights to form a new semantic segmentation model;

[0040] Step S10: Use the new semantic segmentation model to predict the semantic segmentation labels of samples in the inference dataset. In this way, accurate segmentation of unknown categories can be achieved.

[0041] The above description is not intended to limit the present invention in any way. Although the present invention has been disclosed through the above embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the present invention shall still fall within the scope of the present invention.

Claims

1. A zero-shot semantic segmentation method, characterized in that, include: Step S1: Obtain a large-scale image classification dataset containing several categories as the training set; Step S2: Train the weakly supervised semantic segmentation model using the training set to obtain the trained weakly supervised semantic segmentation model; Step S3: Use the trained weakly supervised semantic segmentation model to annotate the large-scale image classification dataset with pixel-level pseudo-labels; Step S4: Use the image data and pixel-level pseudo-labels as a new dataset to train the semantic segmentation model and obtain the trained semantic segmentation model. Step S5: Collect word vectors corresponding to each category in a large-scale image classification dataset; Step S6: Take the word vectors corresponding to each class of the large-scale image classification dataset as input, the classification layer weights of the semantic segmentation model as supervision information, and the output is the predicted classification layer weights. Train the classification layer weight prediction network model to obtain the trained classification layer weight prediction network model. Step S7: Obtain the categories and corresponding word vectors of the inference dataset; Step S8: Input the word vectors corresponding to the inference dataset into the trained classification layer weight prediction network model to obtain the predicted classification layer weights; Step S9: Replace the classification layer weights of the trained semantic segmentation model with the predicted classification layer weights mentioned above. Step S10: Use the semantic segmentation model after replacing the classification layer weights to predict the semantic segmentation labels of samples in the inference dataset to achieve zero-sample semantic segmentation.

2. The zero-shot semantic segmentation method according to claim 1, characterized in that, The large-scale image classification dataset in step S1 includes, but is not limited to, ImageNet and CoCo datasets.

3. The zero-shot semantic segmentation method according to claim 1, characterized in that, The weakly supervised semantic segmentation model is the IRN weakly supervised model.

4. The zero-shot semantic segmentation method according to claim 1, characterized in that, The specific training process in step S2 is as follows: the weakly supervised semantic segmentation model is initialized by taking the training set as input, using the average error loss as the loss function, and then using the stochastic gradient descent algorithm for backpropagation to update the model parameters, thereby obtaining the trained weakly supervised semantic segmentation model.

5. The zero-shot semantic segmentation method according to claim 1, characterized in that, The semantic segmentation model is the DeepLabV3 semantic segmentation model.

6. The zero-shot semantic segmentation method according to claim 1, characterized in that, The specific training process in step S4 is as follows: using a new dataset as input, using the cross-entropy loss of each pixel as the loss function, and then using the stochastic gradient descent algorithm for backpropagation to obtain the trained semantic segmentation model.

7. The zero-shot semantic segmentation method according to claim 1, characterized in that, The classification layer weight prediction network model consists of 8 fully connected layers with the following number of neurons: 2048, 2048, 1024, 1024, 1024, 1024, 512, 512. The number of neurons in the last fully connected layer is the number of predicted categories. The network uses ReLU as the activation function, Adam as the optimizer, and the initial learning rate is set to 0.

001.

8. The zero-shot semantic segmentation method according to claim 1, characterized in that, The specific training process in step S6 is as follows: the word vectors corresponding to each class of the image classification dataset are used as input, the classification layer weights of the semantic segmentation model are used as supervision, the output is the predicted classification layer weights, cross-entropy loss is used as the loss function, and then the stochastic gradient descent algorithm is used for backpropagation to obtain the trained classification layer weight prediction network model.

9. A zero-shot semantic segmentation method according to claim 1, characterized in that, In step S7, the categories of the inference dataset may include categories that do not exist in the training set.

Citation Information

Patent Citations

  • Two-stage zero sample image semantic segmentation method

    CN112801105A

  • Zero sample learning method based on generative adversarial network under semantic error correction

    CN113378959A