A Prompt-Guided Zero-Shot Image Classification Method and System

The method addresses classification challenges in zero-shot scenarios by using text prompts to enhance visual representations through cross-modal fusion, improving accuracy in scenarios with scarce labeled data.

CN118691899BActive Publication Date: 2025-07-15NANJING UNIV OF POSTS & TELECOMM
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410857514.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-28
Publication Date
2025-07-15
Estimated Expiration
2044-06-28

AI Technical Summary

Technical Problem

Existing image classification models cannot effectively classify unseen categories in the absence of visible image samples, especially in scenarios where resources are scarce and the need for professional knowledge to label data.

Method used

Using the zero-sample image classification method based on prompt guidance, the classification model is optimized to realize the similarity score calculation and classification of images and invisible categories through the calculation of global semantic representation, instance prompt representation and instance visual representation, combined with semantic attribute consistency loss, prompt attribute consistency loss, cross entropy loss and debias loss.

Benefits of technology

In the scene without trainable image samples, automatic image classification is used using text semantic prompts and attribute vectors, which improves the accuracy of image classification, and realizes reliable cross-modal communication of instance-level semantic information and visual information, which is suitable for scenes with sparse labeling data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118691899B_ABST
    Figure CN118691899B_ABST
Patent Text Reader

Abstract

The present invention relates to a zero-shot image classification method and system based on prompt guidance, belonging to the technical field of computer vision artificial intelligence. The method includes: calculating an enhanced visual representation, similarity scores of attributes of all categories, and the total loss for optimizing the classification model based on the global semantic representation, instance prompt representation, and instance visual representation, using the total loss as the optimization objective to optimize the parameters of the classification model; calculating the similarity scores between the given image and the attributes of invisible categories according to the optimized classification model, and outputting the predicted label corresponding to the image to achieve zero-shot image classification. This method can fully utilize text semantic prompts and attribute vectors for automatic image classification in scenarios without trainable image samples, ensuring reliable cross-modal communication of instance-level semantic information and instance-level visual information, and improving the accuracy of image classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision artificial intelligence, and particularly to a zero-shot image classification method and system based on prompt guidance. Background Art

[0002] Currently, the end-to-end model technology of artificial intelligence from image feature extraction to classification has been quite mature in the fields of computer vision and image processing, and can obtain high classification accuracy with the support of a large number of labeled samples. In particular, currently widely used large pre-trained image models based on CNN, such as ResNet and VGG, trained on the ImageNet dataset, have provided a new paradigm for image classification - using them as the backbone network and fine-tuning for downstream tasks. Although these models can perform well in daily scenarios, in some specific scenarios, such as extremely expensive artificial annotation data, the number of samples in different categories follows a long-tailed distribution and is unbalanced, scarce annotation data in low-resource fields, and annotation data requires professional knowledge, traditional artificial intelligence models often cannot play a role. In addition, the trained models can only classify the categories seen during the training process, and cannot classify when encountering some unseen categories. These problems have brought a series of challenges to the field of image classification.

[0003] To address this challenge, image classification work for zero-shot learning has been widely studied and continuously developed. Zero-shot learning mainly simulates the cognitive process of the human brain, simulates the knowledge transfer from visible classes to invisible classes to learn transferable relevant knowledge from visible classes, and applies it to the classification of invisible classes. The research work on zero-shot image classification based on semantic prompts is mainly divided into two categories. One is the zero-shot classification model based on generation, and the other is the zero-shot classification model based on embedding. The generation-based model aims to generate corresponding image features for invisible classes through semantic prompts and simplifies it into a fully supervised classification problem. The embedding-based model hopes to find a mapping relationship to map semantic attributes to the visual space, or map visual features to the semantic space, or find a common space to map semantic attributes and visual features into the common space together. Its ultimate goal is to learn the connection between semantic and visual features in the same vector space and complete the classification task in this space. Summary of the Invention

[0004] The main objective of the present invention is to overcome the shortcomings and deficiencies of the prior art, and provide a zero-shot image classification method and system based on prompt guidance, which can automatically classify images by making full use of text semantic prompts and attribute vectors in the scenario without trainable image samples, ensuring reliable cross-modal communication of instance-level semantic information and instance-level visual information, and improving the accuracy of image classification.

[0005] According to one aspect of the present invention, the present invention provides a zero-shot image classification method based on prompt guidance, and the method includes the following steps:

[0006] Calculate an enhanced visual representation based on the global semantic representation, instance prompt representation, and instance visual representation; project the enhanced visual representation fused with semantic information into the semantic attribute space, and calculate the similarity scores of the attributes of all classes.

[0007] Calculate the total loss for optimizing the classification model according to the semantic attribute consistency loss, prompt attribute consistency loss, cross-entropy loss, and debiasing loss, and use the total loss as the optimization objective to optimize the parameters of the classification model.

[0008] Calculate the similarity scores of the given image and the attributes of unseen classes according to the optimized classification model, and output the predicted label corresponding to the image to achieve zero-shot image classification.

[0009] Preferably, the calculating the enhanced visual representation based on the global semantic representation, instance prompt representation, and instance visual representation includes:

[0010] Obtain instance images of visible classes, attribute vectors corresponding to the visible classes, and global semantic descriptions of shared attributes from the dataset, and obtain corresponding text-based instance prompts through the attribute vectors.

[0011] Input the global semantic description and the instance prompt corresponding to the instance image into a text encoder to obtain the global semantic representation and the instance prompt representation; input the instance image into the first L-1 layers of a visual encoder to obtain the instance visual representation.

[0012] Send the global semantic representation and the instance visual representation into a visual instance-guided global semantic enhancement decoder in a semantic-visual cross-modal fusion module to obtain an enhanced global semantic representation.

[0013] Input the enhanced global semantic representation and the instance prompt representation into a prompt instance-guided local enhanced semantic enhancement decoder in the semantic-visual cross-modal fusion module to obtain a locally enhanced local enhanced semantic representation.

[0014] Input the instance visual representation and the locally enhanced semantic representation into a locally enhanced semantic-guided visual decoder in the semantic-visual cross-modal fusion module to obtain the enhanced visual representation.

[0015] Input the enhanced visual representation into the last encoding layer of the visual encoder for encoding to obtain an encoding result.

[0016] Fuse and pair the encoding result of the enhanced visual representation, the global semantic representation, and the instance prompt representation to obtain the enhanced visual representation.

[0017] Preferably, projecting the enhanced visual representation integrated with semantic information into the semantic attribute space and calculating the similarity scores of the attributes of all categories includes:

[0018] Evaluating the similarity scores of the given image and the attribute vectors of each category using cosine similarity:

[0019]

[0020] Wherein, is the candidate label for the classification task, Z is the set of attribute vectors of the categories to be classified, τ is the scaling factor, x is the instance image, is the enhanced visual representation, W p is the fully connected layer, and GAP is the one-dimensional global average pooling function.

[0021] Preferably, calculating the total loss for optimizing the classification model according to the semantic attribute consistency loss, the prompt attribute consistency loss, the cross-entropy loss, and the debiasing loss includes:

[0022]

[0023] Wherein, L sa is the semantic attribute consistency loss, L pac is the prompt attribute consistency loss, L cls is the cross-entropy loss, L deb is the debiasing loss, and L is the total loss for optimization; Attention_score S,V 、 are the attention scores in the global semantic decoder guided by visual instances and the local semantic decoder guided by prompt instances respectively, GMP represents the one-dimensional global max pooling function, z c is the attribute vector of this category, represents the square of the vector two-norm, α S 、α U are the means of the predicted similarity scores of the visible and invisible classes respectively, β S 、β U are the variances of the predicted similarity scores of the visible and invisible classes respectively, λ c 、λ cls and λ deb are the weighting coefficients.

[0024] Preferably, calculating the similarity scores of the given image and the attributes of the invisible categories according to the optimized classification model and outputting the predicted labels corresponding to the image includes:

[0025] Calculating the scores of the invisible categories:

[0026]

[0027] The predicted label of the category with the highest output similarity score

[0028]

[0029] where γ is a calibration coefficient used to balance the scores of visible and invisible classes during training, and f i (·) is an indicator function that has a value of 0 if and 1 otherwise, where Y U represents the set of labels of invisible classes; Y is the set of labels to be predicted.

[0030] According to another aspect of the present invention, the present invention also provides a prompt-guided zero-shot image classification system, which includes:

[0031] A calculation module for calculating an enhanced visual representation based on the global semantic representation, instance prompt representation, and instance visual representation; projecting the enhanced visual representation fused with semantic information into the semantic attribute space to calculate the similarity scores with the attributes of all categories;

[0032] An optimization module for calculating the total loss for optimizing the classification model according to the semantic attribute consistency loss, prompt attribute consistency loss, cross-entropy loss, and debiasing loss, and using the total loss as the optimization objective to optimize the parameters of the classification model;

[0033] A classification module for calculating the similarity scores between the given image and the attributes of invisible classes according to the optimized classification model, outputting the predicted label corresponding to the image, and realizing zero-shot image classification.

[0034] Preferably, the calculation module calculates the enhanced visual representation based on the global semantic representation, instance prompt representation, and instance visual representation, including:

[0035] Obtaining instance images of visible classes, attribute vectors corresponding to the visible classes, and global semantic descriptions of shared attributes from the dataset, and obtaining corresponding texturized instance prompts through the attribute vectors;

[0036] Inputting the global semantic description and the instance prompt corresponding to the instance image into a text encoder to obtain a global semantic representation and an instance prompt representation; inputting the instance image into the first L-1 layers of a visual encoder to obtain an instance visual representation;

[0037] Feeding the global semantic representation and the instance visual representation into a visual instance-guided global semantic enhancement decoder in a semantic-visual cross-modal fusion module to obtain an enhanced global semantic representation;

[0038] Input the enhanced global semantic representation and instance prompt representation into the prompt instance-guided local enhanced semantic enhancement decoder in the semantic-visual cross-modal fusion module to obtain the locally enhanced local enhanced semantic representation;

[0039] Input the instance visual representation and the locally enhanced semantic representation into the locally enhanced semantic-guided visual decoder in the semantic-visual cross-modal fusion module to obtain the enhanced visual representation.

[0040] Input the enhanced visual representation into the last encoding layer of the visual encoder for encoding to obtain the encoding result;

[0041] Fuse and pair the encoding result of the enhanced visual representation, the global semantic representation, and the instance prompt representation to obtain the enhanced visual representation.

[0042] Preferably, the computing module projects the enhanced visual representation fused with semantic information into the semantic attribute space, and the calculated similarity scores of the attributes of all categories include:

[0043] Use cosine similarity to evaluate the similarity score between the given image and the attribute vector of each class:

[0044]

[0045] Among them, is the candidate label for the classification task, Z is the set of attribute vectors of the categories to be classified, τ is the scaling factor, x is the instance image, is the enhanced visual representation, W p is the fully connected layer, and GAP is the one-dimensional global average pooling function.

[0046] Preferably, the optimization module calculates the total loss for optimizing the classification model according to the semantic attribute consistency loss, the prompt attribute consistency loss, the cross-entropy loss, and the debiasing loss, including:

[0047]

[0048]

[0049] Among them, L sac is the semantic attribute consistency loss, L pac is the prompt attribute consistency loss, L cls is the cross-entropy loss, L deb is the debiasing loss, L is the total optimization loss; Attention_score S,V 、 are the attention scores in the global semantic decoder guided by visual instances and the local semantic decoder guided by prompt instances respectively, GMP represents the one-dimensional global maximum pooling function, z cis the attribute vector of this class, represents the square of the vector two-norm, α S and α U are respectively the means of the predicted similarity scores of the visible and invisible classes, β S and β U are respectively the variances of the predicted similarity scores of the visible and invisible classes, λ c and λ cls and λ deb are weighting coefficients.

[0050] Preferably, the classification module calculates the similarity score between the given image and the attributes of the invisible class according to the optimized classification model, and the predicted labels corresponding to the output images include:

[0051] Calculate the score of the invisible class:

[0052]

[0053] Output the predicted label of the class with the highest similarity score

[0054]

[0055] where γ is a calibration coefficient used to balance the scores of the visible and invisible classes during training, f i (·) is an indicator function, if then the function value is 0, otherwise it is 1, where Y U represents the label set of the invisible class; Y is the label set to be predicted.

[0056] Beneficial effects: The present invention introduces instance-level textual prompts, making up for the lack of instance-level semantic information in the zero-shot image classification task under semantic prompts, and providing more instance-level semantic information for the classification task. The present invention uses a prompt instance-guided local enhanced semantic decoder to bring about the alignment of global semantic and instance semantic information in the semantic space, ensuring the transfer of the semantic space of visible class knowledge. The present invention uses a prompt-guided semantic-visual cross-modal fusion module to complete the task of the interaction between instance-level semantic information and instance-level visual information with the help of global semantic information, ensuring reliable cross-modal communication between instance-level semantic information and instance-level visual information, and helping the alignment of fine-grained visual information and fine-grained semantic information. The present invention fully completes the fusion of semantics and vision, realizes the transfer of semantics and visual knowledge from visible classes to invisible classes, and achieves excellent results on three public datasets for zero-shot image classification, and has good practicability in some scenarios with scarce labeled data.

[0057] Through reference to the following drawings and a detailed description of the specific embodiments of the present invention, the features and advantages of the present invention will become clear. Description of the Drawings

[0058] Figure 1 is a flowchart of a prompt-guided zero-shot image classification method;

[0059] Figure 2 is a schematic diagram of the model of the method of the present invention.

[0060] Figure 3 is an architecture diagram of a globally semantic decoder guided by visual instances and a locally enhanced semantic decoder guided by prompt instances.

[0061] Figure 4 is an architecture diagram of a visual decoder guided by locally enhanced semantics.

[0062] Figure 5 is a schematic diagram of a prompt-guided zero-shot image classification system. Detailed Implementation Manner

[0063] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0064] Embodiment 1

[0065] Figure 1 is a flowchart of a prompt-guided zero-shot image classification method. As Figure 1 shown, this embodiment provides a prompt-guided zero-shot image classification method, and the method includes the following steps:

[0066] Calculate an enhanced visual representation based on the global semantic representation, the instance prompt representation, and the instance visual representation; project the enhanced visual representation fused with semantic information into the semantic attribute space, and calculate the similarity scores with the attributes of all classes;

[0067] Calculate the total loss for optimizing the classification model according to the semantic attribute consistency loss, the prompt attribute consistency loss, the cross-entropy loss, and the debiasing loss, and use the total loss as the optimization objective to optimize the parameters of the classification model;

[0068] Calculate the similarity scores between the given image and the attributes of the unseen classes according to the optimized classification model, and output the predicted labels corresponding to the image to achieve zero-shot image classification.

[0069] This method can fully utilize text semantic cues and attribute vectors for automatic image classification in scenarios without trainable image samples, ensuring reliable cross-modal communication between instance-level semantic information and instance-level visual information and improving the accuracy of image classification.

[0070] Preferably, the enhanced visual representation calculated according to the global semantic representation, instance cue representation, and instance visual representation includes:

[0071] Obtain instance images of visible classes, attribute vectors corresponding to the visible classes, and global semantic descriptions of shared attributes from the dataset, and obtain corresponding text-based instance cues through the attribute vectors;

[0072] Input the global semantic description and the instance cues corresponding to the instance images into a text encoder to obtain a global semantic representation and an instance cue representation; input the instance images into the first L - 1 layers of a visual encoder to obtain an instance visual representation;

[0073] Send the global semantic representation and the instance visual representation into a visual instance-guided global semantic enhancement decoder in a semantic-visual cross-modal fusion module to obtain an enhanced global semantic representation;

[0074] Input the enhanced global semantic representation and the instance cue representation into a cue instance-guided local enhancement semantic enhancement decoder in the semantic-visual cross-modal fusion module to obtain a locally enhanced local enhancement semantic representation;

[0075] Input the instance visual representation and the locally enhanced semantic representation into a locally enhanced semantic-guided visual decoder in the semantic-visual cross-modal fusion module to obtain an enhanced visual representation.

[0076] Input the enhanced visual representation into the last encoding layer of the visual encoder for encoding to obtain an encoding result;

[0077] Fuse and pair the encoding result of the enhanced visual representation, the global semantic representation, and the instance cue representation to obtain an enhanced visual representation.

[0078] Specifically, referring to Figure 2 、 Figure 3 and Figure 4 , this step may include the following steps S1 - S7.

[0079] The zero-shot image classification task aims to learn transferable knowledge from the visible class D s ={x s , y s )|x ∈ X s , y ∈ Y s} and apply it to the unseen class D u ={{u, y u )|x ∈ Xu , y ∈ Y u}}, where x represents the image to be classified, y represents the label and belongs to a category c, where c ∈ C, C = C s UC u , and at the same time, the visible classes and the invisible classes are disjoint, that is, C s ∩C u = Φ. For each category c, there is a d a -dimensional attribute vector corresponding to it, where each dimension a i of zc represents the probability of the attribute possessed by this category appearing. These d a attributes are shared between D s and D u and serve as a bridge for knowledge transfer. In addition to using the attribute vector like z c to express the semantic information of the category, it can also be expressed using a textual semantic description. Use to represent all the semantic descriptions of the corresponding category, where t i represents the textual description corresponding to the attribute a i . For traditional zero-shot tasks, the label to be predicted is y ∈ Y u , while for generalized zero-shot tasks, the label to be predicted is y ∈ Y u UY S . The specific steps are as follows:

[0080] S1: First, obtain the input instance image x and the corresponding attribute vector where each dimension a c of z i represents the probability of the attribute possessed by this category appearing, and the global semantic attribute description to represent all the semantic descriptions of the corresponding category, where t i represents the textual description corresponding to the attribute a i . These attributes are shared between the visible classes and the invisible classes. Subsequently, calculate the binary attribute vector through the attribute z c of each class, and obtain the attribute hint T P of each class through this binary attribute vector. That is to say, if the binary format attribute vector is 1 in a certain dimension, it is considered that this class contains this attribute. At the same time, according to the attribute characteristics of the dataset, group the binary format attribute vectors into a semantic attribute hint in a unified hint format. For example, the k-th instance hint of the dataset CUB is shown in Figure 1:

[0081]

[0082] Among them, the attribute features of the same dataset can all be classified into different feature groups. If there are multiple attributes in the same group of features, they are separated by commas.

[0083] S2: Subsequently, use the text encoder to encode the global semantics T and attribute prompts T in text form obtained in step 1 P into the text encoder to obtain the global semantic representation and instance prompt representation as shown in equations (2) and (3):

[0084] S = text_encoder(T) (2)

[0085] P = text_encoder(T P ) (3)

[0086] where b represents the training batch size, d represents the output feature dimension of the text encoder, and l represents the length of the prompt sentence. d a is the number of global semantic attributes. Specifically, it is desired to use the global semantic representation as a fixed reference, so the parameters of the text encoder when generating the global semantic representation S are frozen. Similarly, the instance image x is fed into the 0th layer to the (L - 1)th layer of the vision encoder ViT to obtain the visual instance representation as shown in equation (4):

[0087] V = ViT 0∶L-1 (x) (4)

[0088] where x is the input instance image, L represents the total number of Transformer layers of the pre-trained model ViT of the vision encoder, p 2 represents the number of picture segmentation patches, that is, the input image will be segmented into p×p patches, and c represents the feature dimension of the visual representation output by ViT.

[0089] S3: Feed the global semantic representation S and visual instance representation V in step 2 into the vision instance-guided global semantic enhancement decoder. The specific structure is shown in Figure 2 (a). The function of this decoder is to pair and fuse the global semantics and instance visual features to adjust the global semantics to an enhanced global semantic representation centered on the visual instance. This decoder uses an R-layer loop to achieve progressive fusion and activation. Each layer first passes through a semantic-visual instance attention module, and the i-th dimensional attribute representation S of the global semantics i is paired with the j-th patch V of the instance visual feature jPerform matching, layer by layer, fuse relevant semantic and visual features to achieve the effect of localizing global semantic representation. Then, through the semantic self-correlation activation module, further activate important semantic features through semantic self-correlation connection to achieve the effect of removing irrelevant semantics. The specific module definitions are as follows:

[0090] Semantic-visual instance attention module: To localize and adapt the global semantic features shared by visible and invisible classes to different instance visual features, a semantic-visual instance attention module is used to calculate the i-th dimensional attribute semantic representation S of the global semantics i and the j-th image patch V of the instance visual feature j of the attention score, as shown in 5:

[0091]

[0092] where and are linear transformation matrices responsible for transforming the input into the query and key in the attention mechanism, LN represents the layer normalization function, represents the i-th dimensional attribute representation S of the global semantics i and the j-th image patch V of the visual instance j of the attention score. The higher this score, the more this image patch V j can express this semantic attribute S i . Calculate the attention scores for each dimension of the global semantics and each image patch of the visual instance, and finally obtain the localization attention scores dominated by the visual instance To ensure the ability to correctly pair the global semantics with the visual instance, the present invention uses the semantic attribute consistency loss L sac to optimize the semantic-visual attention module, as shown in Equation 6:

[0093]

[0094] where, GMP represents one-dimensional global max pooling, zc is the attribute vector of this class, represents the square of the vector two-norm. Subsequently, we calculate the global semantic representation centered on the visual instance using the attention scores and the visual instance as shown in Equation 7:

[0095]

[0096] where is a linear transformation matrix used to calculate the value in the attention mechanism, and uses a residual connection to add the adjusted global semantic representation and the unadjusted global semantic representation.

[0097] Global Semantic Self-Correlation Activation Module: Since there are connections between global semantic features and features, the global semantic self-correlation activation module is used to enable interactions between global semantic features and further reduce the impact of irrelevant features on global semantic representations. In this embodiment, the self-correlation attention score Relevant_score between semantic attributes is first calculated S , as shown in Equation (8):

[0098]

[0099] where GAP is a one-dimensional global average pooling function, σ represents the GELU activation function and are the linear weight matrices of the fully connected layers, where g represents the number of feature groups, that is, the number of feature groups to which different features in the prompt belong. Subsequently, the self-correlation attention score is multiplied by the global semantic representation centered on the visual instance , and then a residual connection is made to obtain the global semantic representation after attribute interaction , as shown in Equation (9):

[0100]

[0101] Finally, the important semantic features are activated through the MLP layer to exclude the influence of irrelevant features, and at the same time, a residual connection is made to finally obtain the enhanced global semantic representation , as shown in Equation (10):

[0102]

[0103] where MLP represents a simple multi-layer perceptron neural network

[0104] S4: Feed the enhanced global semantic representation in step 2, the prompt instance representation P, into the prompt instance-guided local enhanced semantic decoder, as shown in Figure 2 (b). This decoder consists of an enhanced global semantic-prompt instance attention module and a local enhanced semantic self-correlation activation module. The overall structure is similar to the visual instance-guided global semantic decoder. The role of this decoder is to pair and fuse the enhanced global semantics with the instance prompt, and use the instance-level prompt representation to activate the enhanced global semantic representation fused with visual instance features, further activating the semantic attribute features related to the instance in the enhanced global semantics, so as to achieve the purpose of further refining the semantic representation. This encoder uses an R-layer loop to achieve fusion. In each layer, the enhanced global semantic-prompt instance attention module is first used to pair the i-th dimension attribute of the enhanced global semantics with the j-th prompt word P of the instance prompt jMatch them, and pair and fuse the relevant enhanced semantic attributes and the feature words of the instance prompts layer by layer, so as to compress and integrate the detailed semantic information of the prompts into each semantic attribute of the enhanced global semantics. Then, continue to explore the relationship between the local enhanced semantic attributes through the local enhanced semantic autocorrelation activation module, so as to refine the local enhanced semantic representation. The specific module definitions are as follows:

[0105] Enhanced global semantics - instance prompt attention module: In order to activate the enhanced global semantic representation with specific prompt words in the instance prompt, the enhanced global semantics - instance prompt attention module is used to calculate the i-th dimensional attribute of the enhanced global semantics and the j-th prompt word P j of the instance prompt, and the attention score is as shown in Equation (11):

[0106]

[0107] where and are linear transformation matrices, responsible for transforming the input into the query and key in the attention mechanism represents the i-th dimensional attribute of the enhanced global semantics and the attention score of the j-th prompt word representation P j of the instance prompt. The higher this score, the higher the correlation between this dimensional attribute and this prompt word. Calculate the attention scores for each feature group of the instance prompt and each dimension of the enhanced global semantics, and finally obtain the instance prompt attention scores guided by the enhanced global semantics To ensure the ability to correctly pair the global semantics with the visual instance, use the prompt attribute consistency loss L pac to optimize the semantic - visual attention module, as shown in Equation (12):

[0108]

[0109] Subsequently, calculate the fused local enhanced semantic representation using the attention scores and the enhanced global semantic representation as shown in Equation (13):

[0110]

[0111] where is a linear transformation matrix used to calculate the value in the attention mechanism. Using the residual connection, add the fused local enhanced semantic representation and the enhanced global semantic representation. The local enhanced semantic representation after integrating the instance prompt further refines the semantic representation, and the semantic information shown is more localized

[0112] Local Enhanced Semantic Self-Correlation Activation Module: Similar to the global semantic self-correlation activation module, the local enhanced semantic self-correlation activation module is used to complete the interaction between various features in the local enhanced semantic representation, and further compress and fuse the local enhanced semantic representation. First, we calculate the self-correlation attention scores between features As shown in Equation (14):

[0113]

[0114] where and are the linear weight matrices of the fully connected layer. Subsequently, the self-correlation attention scores are multiplied by the local enhanced semantic representation , and then a residual connection is made to obtain the global semantic representation after the interaction of the feature groups As shown in Equation (15):

[0115]

[0116] Finally, the MLP layer is used to activate important features, exclude the influence of irrelevant features, and at the same time perform a residual connection to obtain the final local enhanced semantic representation As shown in Equation (16):

[0117]

[0118] S5: Feed the local enhanced semantic representation obtained in step 4 and the instance visual representation V obtained in step 2 into the prompt instance-guided local enhanced semantic decoder, see Figure 3 . This decoder consists of a visual-local enhanced semantic attention module and a visual self-correlation activation module. The role of this decoder is to use the localized local enhanced semantic representation to activate the instance visual representation, achieve accurate cross-domain fusion of matching, and obtain the instance-level semantic-visual fusion feature representation. In order to perform cross-domain interaction and fusion between the local enhanced semantic representation in the semantic space and the visual instance representation in the visual space, the visual-local enhanced semantic attention score is used to establish the connection between the i-th image patch V i of the instance visual representation and the j-th attribute feature of the local enhanced semantic representation. The specific calculation is shown in Equation (17):

[0119]

[0120] where and are linear transformation matrices responsible for transforming the input into the query and key in the attention mechanism, represents the i-th image patch V of the instance visual representation iand the attention score of the j-th attribute feature of the local enhanced semantic representation. It is considered that the higher this score is, the higher the correlation between the semantic feature and the image patch. Calculate the attention score for each image patch of the instance visual representation and each feature of the local enhanced semantic representation, and finally obtain the visual instance attention score based on the local enhanced semantics Subsequently, according to

[0121] Subsequently, according to Extract the key semantic information matching the image from the local enhanced semantic representation to enhance the instance visual features, as shown in Equation (18):

[0122]

[0123] Among them, is a linear transformation matrix used to calculate the value in the attention mechanism. Using residual connection, add the visual representation after fusing the local enhanced semantic information to the original visual representation. The fused instance visual representation incorporates the refined semantic information that matches, achieving fine-grained semantic-visual interaction. Since the prompt instance provides localized refined semantic information, the association and localization between the local enhanced semantic information and the image patch are more accurate

[0124] Visual autocorrelation activation module: Similarly, the fused visual features still need to explore the relationship between image patches. Use the image patch fusion activation module to integrate and refine the features between image patches. The calculation process is defined as shown in Equations (19)-(22):

[0125]

[0126] Among them, T represents the matrix transpose operation, is a fully connected layer. Through this fully connected layer, the image patch features are extended to a high dimension, and then screened by the fully connected layer and finally projected back to the original dimension by the fully connected layer and a residual connection is made to obtain the enhanced visual features Finally, use an MLP layer and residual connection to activate the enhanced visual features again

[0127] S6: Feed the enhanced visual features obtained in step 5 into the last layer of the visual encoder ViT to obtain as shown in Equation (23):

[0128]

[0129] S7: Repeat S3 to S5, and S and P are sent to the prompt-guided semantic-visual cross-modal fusion module again for secondary fusion pairing, and finally a discriminable enhanced visual representation is obtained.

[0130] Preferably, projecting the enhanced visual representation fused with semantic information into the semantic attribute space and calculating the similarity scores of the attributes of all classes includes:

[0131] Evaluating the similarity score between the given image x and the attribute vector of each class using cosine similarity:

[0132]

[0133] Among them, is the candidate label for the classification task, Z is the set of attribute vectors of the classes to be classified, τ is the scaling factor, x is the instance image, is the enhanced visual representation, W p is the fully connected layer, and GAP is the one-dimensional global average pooling function.

[0134] Specifically, through the fusion of the two prompt-guided semantic-visual cross-modal fusion modules, an enhanced visual representation fused with semantic-visual information is obtained. This embodiment uses a fully connected layer to project the enhanced visual representation obtained in step 7 into the space where the attribute vector z c is located, and uses a cosine similarity to evaluate the similarity score with the attribute vector of each class, as shown in Equation 24:

[0135]

[0136] Among them, is the candidate label for the classification task, Z is the set of attribute vectors of the classes to be classified. When the task is traditional zero-shot classification, Z = {z c | c ∈ C U}, when the task is generalized zero-shot classification, Z = {z c | c ∈ C U ∪ C S}. Among them, C U represents the unseen classes, and C S represents the seen classes. τ is the scaling factor.

[0137] Preferably, calculating the total loss for optimizing the classification model according to the semantic attribute consistency loss, the prompt attribute consistency loss, the cross-entropy loss, and the debiasing loss includes:

[0138]

[0139] Among them, Lsac is the semantic attribute consistency loss, L pac is the prompt attribute consistency loss, L cls is the cross-entropy loss, L deb is the debiasing loss, L is the total optimization loss; Attention_score S,V 、 are the attention scores in the visually instance-guided global semantic decoder and the prompt instance-guided local semantic decoder respectively, GMP represents the one-dimensional global max pooling function, z c is the attribute vector of this class, represents the square of the vector two-norm, α S 、α U are the means of the visible class and invisible class scores respectively, β S 、β U are the variances of the visible class and invisible class respectively, λ c 、λ cls and λ deb are the weighting coefficients.

[0140] Specifically, the calculated relevant score is used to calculate the cross-entropy loss L cls to supervise the training of this classification task, as shown in Equation (25):

[0141]

[0142] where y refers to the true label of the classification, Y s represents the set of visible class labels, and then the debiasing loss L deb is calculated to reduce the bias between the visible class and the invisible class, as shown in Equation (26):

[0143]

[0144] where α S 、α U are the means of the visible class and invisible class scores respectively, β S 、β U are the variances of the visible class and invisible class respectively. Through this loss, the distributions of the visible class and the invisible class are made as consistent as possible to ensure that the model does not overfit the visible class.

[0145] Finally, L sac 、L pac 、L cls L deb are weighted and summed to obtain the total optimization loss L of the model, as shown in Equation (27):

[0146]

[0147] where λc , λ cls and λ deb are weighting coefficients. Finally, using the total loss L as the optimization objective, the model parameters are optimized.

[0148] Preferably, calculating the similarity score between the given image and the attributes of the unseen class according to the optimized classification model, and the predicted label corresponding to the output image includes:

[0149] Calculating the score of the unseen class:

[0150]

[0151] Outputting the predicted label of the class with the highest similarity score

[0152]

[0153] where γ is a calibration coefficient used to balance the scores of the visible and unseen classes during training, and f i (·) is an indicator function, if then the function value is 0, otherwise 1, where Y U represents the set of labels of the unseen class; Y is the set of labels to be predicted.

[0154] Specifically, during the model training process, since the model can only classify visible classes, the model will inevitably tend to generate higher scores for visible classes, resulting in low generalization performance. Therefore, a calibration stack is used to assign certain scores to unseen classes, as shown in Equation (28):

[0155]

[0156] where γ is a calibration coefficient used to balance the scores of the visible and unseen classes during training. f i (·) is an indicator function, if then the function value is 0, otherwise 1. Y U is the set of unseen class labels.

[0157] Finally, using the trained model, the relevant scores are obtained and the label y corresponding to the instance image x is calculated using these relevant scores, as shown in Equation (29):

[0158]

[0159] where Y is the set of labels to be predicted.

[0160] This embodiment introduces instance-level textual prompts, making up for the lack of instance-level semantic information in the zero-shot image classification task under semantic prompts, and providing more instance-level semantic information for the classification task. The present invention uses a prompt instance-guided local enhanced semantic decoder to bring about the alignment of global semantic and instance semantic information in the semantic space, ensuring the transfer of the semantic space of visible class knowledge. The present invention uses a prompt-guided semantic-visual cross-modal fusion module to complete the task of the interaction between instance-level semantic information and instance-level visual information with the help of global semantic information, ensuring reliable cross-modal communication between instance-level semantic information and instance-level visual information, and helping the alignment of fine-grained visual information and fine-grained semantic information. The present invention fully completes the fusion of semantics and vision, realizes the transfer of semantics and visual knowledge from visible classes to invisible classes, and achieves excellent results on three public datasets for zero-shot image classification, and has good practicality in some scenarios with scarce labeled data.

[0161] Embodiment 2

[0162] Figure 5 is a schematic diagram of a zero-shot image classification system based on prompt guidance. As Figure 5 shown, this embodiment provides a zero-shot image classification system based on prompt guidance, and the system includes:

[0163] A calculation module 501, configured to calculate an enhanced visual representation according to a global semantic representation, an instance prompt representation, and an instance visual representation; project the enhanced visual representation fused with semantic information into a semantic attribute space, and calculate similarity scores with the attributes of all classes;

[0164] An optimization module 502, configured to calculate the total loss for optimizing the classification model according to a semantic attribute consistency loss, a prompt attribute consistency loss, a cross-entropy loss, and a debiasing loss, and use the total loss as an optimization target to optimize the parameters of the classification model;

[0165] A classification module 503, configured to calculate similarity scores between the given image and the attributes of invisible classes according to the optimized classification model, and output the predicted label corresponding to the image, so as to realize the classification of zero-shot images.

[0166] Preferably, the calculation module 501 calculating the enhanced visual representation according to the global semantic representation, the instance prompt representation, and the instance visual representation includes:

[0167] Obtain instance images of visible classes, attribute vectors corresponding to the visible classes, and global semantic descriptions of shared attributes from the dataset, and obtain corresponding textual instance prompts through the attribute vectors;

[0168] Input the global semantic description and the instance prompt corresponding to the instance image into the text encoder to obtain the global semantic representation and the instance prompt representation; input the instance image into the first L-1 layers of the visual encoder to obtain the instance visual representation;

[0169] Send the global semantic representation and the instance visual representation into the visual instance-guided global semantic enhancement decoder in the semantic-visual cross-modal fusion module to obtain the enhanced global semantic representation;

[0170] Input the enhanced global semantic representation and the instance prompt representation into the prompt instance-guided local enhanced semantic enhancement decoder in the semantic-visual cross-modal fusion module to obtain the locally enhanced local enhanced semantic representation;

[0171] Input the instance visual representation and the locally enhanced semantic representation into the locally enhanced semantic-guided visual decoder in the semantic-visual cross-modal fusion module to obtain the enhanced visual representation.

[0172] Input the enhanced visual representation into the last encoding layer of the visual encoder for encoding to obtain the encoding result;

[0173] Fuse and pair the encoding result of the enhanced visual representation, the global semantic representation and the instance prompt representation to obtain the enhanced visual representation.

[0174] Preferably, the computing module 501 projects the enhanced visual representation fused with semantic information into the semantic attribute space, and the calculated similarity scores of the attributes of all categories include:

[0175] Use cosine similarity to evaluate the similarity scores of the attribute vectors of each class:

[0176]

[0177] Among them, is the candidate label for the classification task, Z is the set of attribute vectors of the categories to be classified, τ is the scaling factor, x is the instance image, is the enhanced visual representation, W p is the fully connected layer, and GAP is the one-dimensional global average pooling function.

[0178] Preferably, the optimization module 502 calculates the total loss for optimizing the classification model according to the semantic attribute consistency loss, the prompt attribute consistency loss, the cross-entropy loss and the debiasing loss, including:

[0179]

[0180] Among them, L sac is the semantic attribute consistency loss, L pac is the prompt attribute consistency loss, L clsis the cross-entropy loss, L deb is the debiasing loss, L is the total optimization loss; Attention_score S,V , are the attention scores in the vision instance-guided global semantic decoder and the prompt instance-guided local semantic decoder respectively, GMP represents the one-dimensional global max pooling function, z c is the attribute vector of this class, represents the square of the vector two-norm, α S , α U are the means of the visible class and invisible class scores respectively, β S , β U are the variances of the visible class and invisible class respectively, λ c , λ cls and λ deb are the weighting coefficients.

[0181] Preferably, the classification module 503 calculates the similarity score between the given image and the attributes of the invisible class according to the optimized classification model, and the predicted label corresponding to the output image includes:

[0182] Calculate the score of the invisible class:

[0183]

[0184] Output the predicted label of the class with the highest similarity score

[0185]

[0186] where γ is the calibration coefficient used to balance the scores of the visible class and invisible class during training, f i (·) is an indicator function, if then the function value is 0, otherwise it is 1, where Y U represents the label set of the invisible class; Y is the label set to be predicted.

[0187] The specific implementation process of the functions implemented by each module in this Embodiment 2 is the same as the implementation process in Embodiment 1, and will not be repeated here.

[0188] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structural transformation made under the concept of the present invention, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present invention.

Claims

1. A zero-shot image classification method based on prompt guidance, characterized in that The method includes the following steps: Calculate an enhanced visual representation based on a global semantic representation, an instance prompt representation, and an instance visual representation; project the enhanced visual representation fused with semantic information into a semantic attribute space, and calculate similarity scores with the attributes of all categories; Calculate the total loss for optimizing the classification model according to a semantic attribute consistency loss, a prompt attribute consistency loss, a cross-entropy loss, and a debiasing loss, and use the total loss as an optimization objective to optimize the parameters of the classification model; Calculate the similarity scores between the given image and the attributes of invisible categories according to the optimized classification model, and output the predicted label corresponding to the image to achieve zero-shot image classification; The calculating the enhanced visual representation based on the global semantic representation, the instance prompt representation, and the instance visual representation includes: Obtain instance images of visible categories, corresponding attribute vectors of visible categories, and a global semantic description of shared attributes from the dataset, and obtain corresponding text-based instance prompts through the attribute vectors; Input the global semantic description and the instance prompt corresponding to the instance image into the text encoder to obtain the global semantic representation and the instance prompt representation; input the instance image into the first -1 layers of the visual encoder to obtain the instance visual representation; Send the global semantic representation and the instance visual representation into a visual instance-guided global semantic enhancement decoder in a semantic-visual cross-modal fusion module to obtain an enhanced global semantic representation; Input the enhanced global semantic representation and the instance prompt representation into a prompt instance-guided local enhancement semantic enhancement decoder in the semantic-visual cross-modal fusion module to obtain a locally enhanced local enhancement semantic representation; Input the instance visual representation and the local enhanced semantic representation into the local enhanced semantic-guided visual decoder in the semantic-visual cross-modal fusion module to obtain the enhanced visual representation ; The enhanced visual representation is input into the last encoding layer of the visual encoder for encoding to obtain an encoding result ; The encoded result , the global semantic representation , and the instance prompt representation are fed into the prompt-guided semantic-visual cross-modal fusion module to obtain a discriminable enhanced visual representation ; The calculating the total loss for optimizing the classification model according to the semantic attribute consistency loss, the prompt attribute consistency loss, the cross-entropy loss, and the debiasing loss includes: Among them, is the semantic attribute consistency loss, is the hint attribute consistency loss, is the cross-entropy loss, is the debiasing loss, is the total optimization loss; and are the attention scores in the vision instance-guided global semantic decoder and the hint instance-guided local semantic decoder respectively, represents the one-dimensional global max pooling function, is the attribute vector of this class, represents the square of the vector two-norm, and are the means of the visible class and invisible class prediction similarity scores respectively, and are the variances of the visible class and invisible class prediction similarity scores respectively, and and are the weighting coefficients.

2. The method according to claim 1, characterized in that, The projecting the enhanced visual representation fused with semantic information into the semantic attribute space and calculating the similarity scores with the attributes of all categories includes: Use cosine similarity to evaluate the similarity scores between the given image and the attribute vectors of each category: Among them, is the candidate label for the classification task, is the set of attribute vectors of the category to be classified, is the scaling factor, is the instance image, is the enhanced visual representation, is the fully connected layer, is the one-dimensional global average pooling function.

3. A zero-shot image classification system based on prompt guidance, characterized in that, The system includes: A calculation module for calculating an enhanced visual representation based on a global semantic representation, an instance prompt representation, and an instance visual representation; projecting the enhanced visual representation fused with semantic information into a semantic attribute space, and calculating similarity scores with the attributes of all categories; An optimization module for calculating the total loss for optimizing the classification model according to a semantic attribute consistency loss, a prompt attribute consistency loss, a cross-entropy loss, and a debiasing loss, and using the total loss as an optimization objective to optimize the parameters of the classification model; A classification module for calculating the similarity scores between the given image and the attributes of invisible categories according to the optimized classification model, and outputting the predicted label corresponding to the image to achieve zero-shot image classification; The calculation module calculating the enhanced visual representation based on the global semantic representation, the instance prompt representation, and the instance visual representation includes: Obtain instance images of visible categories, corresponding attribute vectors of visible categories, and a global semantic description of shared attributes from the dataset, and obtain corresponding text-based instance prompts through the attribute vectors; Input the global semantic description and the instance prompt corresponding to the instance image into the text encoder to obtain the global semantic representation and the instance prompt representation; input the instance image into the first -1 layers of the visual encoder to obtain the instance visual representation; Send the global semantic representation and the instance visual representation into a visual instance-guided global semantic enhancement decoder in a semantic-visual cross-modal fusion module to obtain an enhanced global semantic representation; Input the enhanced global semantic representation and instance prompt representation into the prompt instance-guided local enhanced semantic enhancement decoder in the semantic-visual cross-modal fusion module to obtain the locally enhanced semantic representation after local enhancement; Input the instance visual representation and the local enhanced semantic representation into the local enhanced semantic-guided visual decoder in the semantic-visual cross-modal fusion module to obtain the enhanced visual representation ; Enhanced visual representation is input into the last encoding layer of the visual encoder for encoding to obtain an encoding result ; The encoded result , the global semantic representation , and the instance prompt representation are fed into the prompt-guided semantic-visual cross-modal fusion module to obtain a discriminable enhanced visual representation ; The optimization module calculates the total loss for optimizing the classification model according to the semantic attribute consistency loss, prompt attribute consistency loss, cross-entropy loss, and debiasing loss, including: Among them, is the semantic attribute consistency loss, is the hint attribute consistency loss, is the cross-entropy loss, is the debiasing loss, is the total optimization loss; and are the attention scores in the vision instance-guided global semantic decoder and the hint instance-guided local semantic decoder respectively, represents the one-dimensional global max pooling function, is the attribute vector of this class, represents the square of the vector two-norm, and are the means of the visible class and invisible class prediction similarity scores respectively, and are the variances of the visible class and invisible class prediction similarity scores respectively, and and are the weighting coefficients.

4. The system according to claim 3, wherein The calculation module projects the enhanced visual representation fused with semantic information into the semantic attribute space and calculates the similarity scores of the attributes of all classes, including: Use cosine similarity to evaluate the similarity scores of the image and the attribute vectors of each class: Among them, is the candidate label for the classification task, is the set of attribute vectors of the category to be classified, is the scaling factor, is the instance image, is the enhanced visual representation, is the fully connected layer, is the one-dimensional global average pooling function.

Citation Information

Patent Citations

  • Generalized zero sample image classification method based on fused visual information

    CN116797821A

  • Cervical panoramic image few-sample classification method based on visual guidance and language prompt

    CN118230052A

  • Zero sample image classification method and system based on prompt guidance

    CN118691899A