Multi-label Image Classification Method for Multi-layer Prompt Information
By using multi-layer prompt information in the multi-label image classification task, the topic distribution is injected into the visual Transformer model, the problem of inaccurate or missing sample labels is solved, the accuracy and robustness of the multi-label classification of the model is improved, and the cost and error rate of manual labeling are reduced.
Patent Information
- Application Number
- CN202411491288.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-10-24
AI Technical Summary
In multi-label classification task, due to the inaccuracy or lack of sample labels, the prediction ability of deep learning models has decreased, affecting the generalization ability and reliability of the model. Especially in the era of big data, manual labeling is high cost and error rate is high.
The multi-layer prompt information method is used to cluster the topic distribution of samples in the multi-label dataset using the theme model, and inject the topic distribution into the visual Transformer model, affecting the feature extraction process through the cross attention mechanism, and improving the model's understanding of multiple labels.
Through the injection of multi-layer prompt information, the model can pay more accurately to some significant image features corresponding to a certain topic, which improves the accuracy of multi-label image classification and the robustness of the model, and reduces the cost and error rate of manual labeling.
Smart Images

Figure CN119445216B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer vision and natural language processing, and particularly relates to a multi-label image classification method with multi-layer prompt information. Background Art
[0002] Topic Modeling is a core task in the field of text analysis. Its goal is to enable machines to deeply understand the potential topic structure in a document collection through machine learning algorithms, and then use this topic information to abstract and summarize text data. Topic models can perform semantic-level relationship reasoning and data mining on documents containing a large number of words, thereby revealing the hidden topic distribution in the documents.
[0003] A topic model is a statistical model used for clustering analysis of the implicit semantic structure in text. It uses unsupervised learning to find the potential topics in a document and analyze the correlation between these topics and words. Topic models are mainly applied to natural language processing (NLP) and text mining, such as collecting, classifying, and dimensionality reduction of text by topic. The working principle of a topic model is based on a probability model, where each topic is regarded as a probability distribution of words in the vocabulary. Each word in an article is obtained through a process of "selecting a certain topic with a certain probability and then selecting a certain word from this topic with a certain probability".
[0004] Multi-Lable Classification (MLC) is a key task in the field of computer vision. This task requires a machine learning model to fully mine the feature information of samples and perform effective relationship reasoning among multiple labels. Through multi-label classification, a machine learning model can identify multiple attributes or categories possessed by a sample in a complex dataset. The research on multi-label classification not only helps to promote the development of machine learning technology but also has received extensive attention in fields such as pattern recognition and information retrieval. This task has broad value in practical applications. For example, in fields such as network content auditing, image annotation, and text classification, it can help machines better understand the multi-dimensional features of samples, improve the accuracy of data processing and analysis, and thus enhance the intelligent processing ability.
[0005] In multi-label classification tasks, the correctness and integrity of sample labels are one of the key factors affecting model performance. Inaccuracies or omissions in the sample label set can weaken the predictive ability of deep learning models and thus affect their performance in practical applications. With the advent of the big data era, the explosion in the volume of data has led to an explosive growth in the number of samples and the variety of labels. This not only significantly increases the economic cost of manually annotating the sample label set but also adds to the complexity of the annotation work. During the manual annotation process, due to reasons such as fatigue, subjective judgment differences, or insufficient professional knowledge, the phenomenon of incomplete and mislabeled label sets is not uncommon. The existence of these problems undoubtedly greatly increases the difficulty of training a high-precision multi-label classification deep learning model and also places higher requirements on the generalization ability and reliability of the model. Summary of the Invention
[0006] In view of the above, the object of the present invention is to provide a multi-label image classification method with multi-layer prompt information.
[0007] The present invention first uses Topic Modeling to cluster the topic distributions of samples in a multi-label dataset; then injects the topic distributions into the sample image training model and uses the cross-attention mechanism to affect the feature extraction process, making it easier for the model to focus on some significant image features corresponding to a certain topic.
[0008] To achieve the above object of the invention, the technical solution provided by the present invention is as follows:
[0009] Step 1: Obtain the label set corresponding to the test image and the training set samples.
[0010] Step 2: Set the number of topics and use the topic model to obtain the topic distributions of all samples in the training set.
[0011] Step 3: In a Vision Transformer (ViT) model, learn a set of prompt blocks representing the topic label distribution information, inject the prompt blocks into the intermediate features, and input them into the selected prompt layer.
[0012] Step 4: Finally, judge which categories the sample belongs to based on the output result of the ViT model.
[0013] Furthermore:
[0014] In Step 1, the image is a color image containing multiple entities in a natural scene. Each image corresponds to a label set, which contains no less than one label, and the label sets corresponding to the samples in the training set are known, while the label sets of the test set samples are unknown.
[0015] The label set refers to the set of natural language names of the classes to which each sample in the training set belongs.
[0016] In step 2, the topic model refers to a model that discovers the abstract topics contained in multiple documents. The label set corresponding to the training set samples is input into the topic model, and the number of topics is set. Then, the topic distribution of the training set samples is obtained using the topic model;
[0017] The topic distribution refers to the probability that a sample belongs to several different topics. Usually, the topic distribution of a sample is a vector with the length equal to the set number of topics, and the sum of all elements in the vector is 1.
[0018] The training set samples refer to a set of images, and the topic distribution corresponding to the training set samples refers to the probability value of each training set sample corresponding to each topic.
[0019] In step 3, using the auxiliary learning strategy, while performing the main task of multi-label classification, the auxiliary task of label topic classification is carried out. The specific method is as follows: First, select the prompt layer (Prompt Block) for inserting prompts and the number of topics corresponding to the topic information inserted into each prompt layer; for each prompt layer, connect the corresponding prompt block to the image features output by the previous layer, input it into the prompt layer, and let the image features and the prompt layer absorb information from each other through the cross-attention mechanism. Finally, split the output of the prompt layer into new image features and prompt blocks. The obtained image features continue to enter the next layer of Block, and the prompt blocks enter the topic classification learning task. Specifically:
[0020] In the ViT model, select some intermediate layers of the ViT model as the prompt layer Prompt Block and set the corresponding number of topics;
[0021] For each prompt layer:
[0022] First, initialize a Tensor of size (1, Embed_size) as the prompt block, and then connect this prompt block to the image features output by the previous layer of Block and input them into the prompt layer together;
[0023] Then, split the output result of the prompt layer into two distributions: the prompt block and the image features; the prompt block enters the auxiliary task, and the image features enter the next layer of Block to continue the main task; if the prompt layer is the last layer of Block, the prompt block enters the auxiliary task, and the image features enter the classifier to perform the multi-label classification task;
[0024] Finally, the losses of the auxiliary task and the main task are weighted and summed to obtain the loss of the model.
[0025] The auxiliary task of topic label classification uses Cross-Entropy Loss as the loss function, and each prompt layer generates a corresponding topic label loss A i ; For the main task of multi-label classification, AsymmetricLoss is used as the loss function, and the generated loss is denoted as B. The overall loss function of the model n represents the number of prompt layers.
[0026] The auxiliary task of topic label classification uses Cross-Entropy Loss as the loss function, and each prompt layer generates a corresponding topic label loss A i , For the main task of multi-label classification, AsymmetricLoss is used as the loss function, and the generated loss is denoted as B. The overall loss function of the model i represents the serial number of the prompt layer, n represents the number of prompt layers, α and βi are artificially set hyperparameters used to weighted sum the losses of the main task and the auxiliary task to obtain the final total loss of the model, and βi corresponds to the weight of the topic label loss generated by the i-th prompt layer.
[0027] First, it is necessary to select the prompt layer (Prompt Block) for inserting prompts and the number of topics corresponding to the topic information inserted in each prompt layer. The number of topics corresponding to the prompt layer closer to the bottom is less, and the number of prompts corresponding to the prompt layer closer to the top gradually increases, enabling the ViT model to hierarchically learn the gradually refined topic information.
[0028] Based on the output result of the Vision Transformer model, it is judged which categories the sample belongs to. The output result of the Vision Transformer model is activated by the Sigmoid function, and then a threshold is set, usually 0.5. The label corresponding to the subscript of the value greater than this threshold is considered to be the label included in this image.
[0029] In summary, the present invention has more advantages in providing topic information, enriching the hierarchy of the label set, and thus guiding machine attention in the multi-label classification task of the Vision Transformer model, improving the accuracy of the task.
[0030] Compared with the prior art, the beneficial effects of the multi-label image classification method with multi-layer prompt information of the present invention include:
[0031] First, the present invention introduces the topic distribution generated by the topic model and the prediction result of the image classification model in the Block of the Vision Transformer model as the rough class information of the image, fully utilizes the information between the sample label sets, provides more information, and improves the accuracy of the multi-label classification task;
[0032] Second, compared with directly using the Vision Transformer model for multi-label classification tasks, the multi-label image classification method with multi-layer prompt information injects topic information with a gradually increasing number of topics on multiple prompt layers, enabling the ViT model to hierarchically learn topic information with gradually refined granularity, which helps the attention mechanism focus on smaller regions that are more crucial for distinguishing which category an object belongs to.
[0033] Third, in the case where the data volume increases, the label space becomes larger, and manual annotation becomes expensive and error-prone, injecting topic information into the Block of the Vision Transformer model can, to a certain extent, make up for the missing and incorrect labels, improve the robustness of the model, and reduce the labor cost at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. In the drawings:
[0035] Figure 1 is a flowchart of the present invention.
[0036] Figure 2 is a schematic diagram of the effect of the topic probability distribution of the training set samples generated by the topic model in the present invention.
[0037] Figure 3 is a schematic diagram of the effect of how to inject the topic distribution information of the samples into the prompt layer in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the protection scope of the present invention.
[0039] The inventive concept of the present invention is: aiming at the problems of the enlarged label space caused by the increasing current data volume, the increasing cost of manual annotation, and the increasing error rate, a multi-label image classification method with multi-layer prompt information is proposed, which is applied to correctly output the entity names included in the picture in a picture scene with multiple entities, thereby further improving the accuracy of human-computer interaction and the machine's understanding ability of pictures.
[0040] In this embodiment, the sample image training model uses a Vision Transformer (ViT) model.
[0041] Figure 1 is a flowchart of the multi-label image classification method with multi-layer prompt information provided by an embodiment of the present invention. As Figure 1 shown, the embodiment provides a multi-label image classification method with multi-layer prompt information, including the following steps:
[0042] Step 1, obtain the test image and the label set corresponding to the training set samples.
[0043] The image is a color image containing at least one entity in a natural scene. The label set refers to the set of natural language names of the classes to which each sample in the training set belongs. The number of labels corresponding to each training set sample is at least one, representing the entities in the sample image.
[0044] In this embodiment, referring to Figure 2 , the input part on the left side of its "topic model" contains the sample image and the corresponding sample label set. Usually, one image corresponds to multiple labels.
[0045] Step 2, input the label sets of all samples in the training set into the topic model to obtain the topic distribution of all samples in the training set.
[0046] In this embodiment, as Figure 2 shown, there are multiple sample images and the label set corresponding to each sample image on the left side of the topic model. The label set corresponding to each sample image is regarded as a document. Input multiple documents into the topic model, and the topic model generates corresponding results, such as Figure 2 shown on the right side of the topic model, representing the distribution of the probability of each image for each topic. The sum of the probabilities of the same image for each topic is 1.
[0047] In the prior art, the topic model algorithm is relatively mature. For example: LDA model, BTM (Biterm Topic Model) model, and BERTopic model, etc. These models are all applicable to the topic clustering of the present invention.
[0048] Step 3, learn a set of prompt blocks representing the topic label distribution information in the ViT model, inject the prompt blocks into the intermediate features, and input them into the selected prompt layer.
[0049] In the embodiment, as Figure 3 shown, first select the prompt layer and the corresponding number of topics. In this embodiment, select the l-th middle layer Block and the last layer Block as the prompt layer (Prompt Block) (in this example, two Blocks are selected as the prompt layer), and set the corresponding number of topics. The number of topics corresponding to the first prompt layer is 2, and the number of topics corresponding to the second prompt layer is 8.
[0050] For the first prompt layer:
[0051] First, initialize a Tensor of size (1, Embed_size) as the prompt block, then concatenate this prompt block with the image features output by the previous layer Block and input them together into the prompt layer;
[0052] Then, split the output result of the prompt layer into two distributions: the prompt block and the image features; the prompt block enters the auxiliary task of topic classification, while the image features enter the next layer Block to continue the main task of multi-label classification.
[0053] For the last layer Block: perform a similar operation, and finally, sum the losses of the auxiliary task and the main task with weights to obtain the loss of the model.
[0054] Step 4, finally, determine which categories the sample belongs to based on the output result of the ViT model.
[0055] Specifically, the ViT model will finally output a vector of size (1, Class_num) for an image, where Class_num is a fixed value, generally determined by the fixed attributes of the dataset, representing the size of the label space in this dataset.
[0056] Subsequently, this vector is fed into the activation function Sigmoid, and a threshold (usually 0.5 can be set);
[0057] Finally, output the subscripts of each element in the activated vector that is greater than this threshold, and the labels corresponding to these subscripts are the labels that the ViT model believes the image sample corresponds to.
[0058] The specific embodiments described in this article are only illustrative of the spirit of the present invention. Those skilled in the art of the present invention can make various modifications or supplements to the described specific embodiments or use similar ways to replace them, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. A multi-label image classification method with multi-layer prompt information, characterized in that: The following steps are involved: Step 1, obtain test set samples, obtain training set samples and their corresponding label sets; the samples are images; Step 2: Input the label set corresponding to the training set samples into the topic model, set the number of topics, and use the topic model to obtain the topic distribution of all samples in the training set; Step 3: In the Vision Transformer (ViT) model, learn a set of prompt tokens representing the topic tag distribution information, inject the prompt blocks into the intermediate features, and input them into the selected prompt layer Prompt Block. Step 4: Use the output of the ViT model to determine which categories the sample belongs to; In step 1, the image is a color image containing at least one entity in a natural scene; the label set refers to the set of natural language names of the class to which each sample in the training set belongs; the number of labels corresponding to each training set sample is at least one, representing the entity in the sample; each image corresponds to a label set, and the label set contains no less than one label; The corresponding label set of the training set samples is known; the label set of the test set samples is unknown; In step 2, for the input of the topic model, the label set corresponding to each sample is regarded as a document, and these documents are input into the topic model; the topic distribution output by the topic model refers to the probability that a sample belongs to several different topics; the topic distribution of a sample is a vector with a length equal to the number of topics set, and the sum of all elements in the vector is 1; In the ViT model in step 3, an auxiliary learning strategy is used to perform auxiliary tasks while carrying out the main task. The main task is a multi-label classification task, and the auxiliary task is a topic label classification task. The steps include: First, select the prompt layer Prompt Block where the prompt block is inserted and the number of topics corresponding to the topic information inserted in each prompt layer; Then, for each hint layer, the corresponding hint block is connected to the image features output by the previous layer; then they are input into the hint layer together, and the image features and hint layer absorb information from each other through the cross attention mechanism; Finally, the output of the prompt layer is split into new image features and prompt blocks. The image features continue to enter the next layer of blocks, and the prompt blocks enter the topic label classification task. In step 3, in the ViT model, select some layers of blocks in the middle of the ViT model as prompt blocks, and set the corresponding number of topics; For each prompt layer: First, initialize a Tensor of size (1, Embed_size) as a prompt block, then connect the prompt block and the image features output by the previous layer of Block, and input them into the prompt layer together; Then, the output result of the prompt layer is split into two distributions: the prompt block and the image feature. The prompt block enters the auxiliary task, while the image feature enters the next layer of Block to continue the main task. If the prompt layer is the last layer of Block, the prompt block enters the auxiliary task, while the image feature enters the classifier to perform multi-label classification tasks. Finally, the loss of the auxiliary task and the main task is weightedly summed to obtain the loss of the model.
2. The multi-label image classification method of multi-layer prompt information according to claim 1, characterized in that: In step 2, the topic model is an LDA model, a BTM model, or a BERTopic model.
3. The multi-label image classification method of multi-layer prompt information according to claim 1, characterized in that: In step 3, the topic label classification task of the auxiliary task uses Cross-Entropy Loss as the loss function, and each prompt layer will generate a corresponding topic label loss A i , the main task multi-label classification task uses Asymmetric Loss as the loss function, the resulting loss is recorded as B, and the overall loss function of the model is i represents the sequence number of the prompt layer, n represents the number of prompt layers, α, β i is a hyperparameter set by humans, which is used to weight the sum of the losses of the main task and the auxiliary task to obtain the final total model loss, β i The weight corresponding to the topic label loss generated by the i-th hint layer.
4. The multi-label image classification method of multi-layer prompt information according to claim 1, characterized in that: The feature is that in step 4, the output of the ViT model is sent to the activation function; if the label probability value of the output result of the activation function is greater than the preset threshold, it is considered that the sample contains this category.
5. The multi-label image classification method of multi-layer prompt information according to claim 1, characterized in that: The characteristic is that in step 4, the method of judging which categories the sample belongs to by the output result of the ViT model is: First, the ViT model finally outputs a vector of size (1, Class_num) for an image. Class_num is a fixed value determined by the fixed attributes of the dataset and represents the size of the label space in the dataset. Subsequently, this vector is fed into the activation function Sigmoid and a threshold is set; Finally, the subscripts of each element in the activated vector that is greater than the threshold are output. The labels corresponding to these subscripts are the labels that the ViT model believes the sample corresponds to.
Citation Information
Patent Citations
Topic-based dense paragraph retrieval prompt generation method and system
CN117851540A
Zero sample image classification method and system based on prompt guidance
CN118691899A