A zero-shot image classification network and deep learning method based on contrastive learning

Through a zero-sample image classification network based on contrast learning, combined with the contrast learning embedding and consistency constraints of semantic attributes, the problem of poor expression of semantic attributes in the existing methods is solved, and higher discrimination and robustness are achieved, and the accuracy of zero-sample image classification is improved.

CN115641582BActive Publication Date: 2025-07-22NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211298406.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-23
Publication Date
2025-07-22
Estimated Expiration
2042-10-23

AI Technical Summary

Technical Problem

The existing zero-sample image classification method lacks discriminant and robustness in semantic attribute expression, resulting in poor semantic alignment of cross-modal information.

Method used

A zero-sample image classification network based on contrast learning is adopted, combining the contrast learning embedding of semantic attributes, consistency constraints of student models and teacher models, and prototype modules, through instance-level and category-level supervision, the image feature expression ability of semantic attributes is improved, and the robustness of the model is enhanced through the Mean Teacher mechanism.

Benefits of technology

It effectively alleviates the problem of misclassification of fine-grained data sets and the confusion of different types of images, enhances the discriminant and robustness of semantic attributes, realizes semantic alignment of cross-modal features, and improves the accuracy of zero-sample image classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115641582B_ABST
    Figure CN115641582B_ABST
Patent Text Reader

Abstract

The present invention relates to a zero-shot image classification network and a deep learning method based on contrastive learning. The designed model includes three parts: contrastive learning embedding of semantic attributes, consistency constraint between the student model and the teacher model, and a prototype module. Among them, the contrastive learning embedding part of semantic attributes starts from two levels: instance-level supervision and category-level supervision, and conducts contrastive learning on the predicted semantic attributes of positive and negative samples and the predicted semantic attributes of positive and negative categories. The consistency constraint part introduces the Mean Teacher mechanism, inputs images with different data augmentations to the student model and the teacher model, and then enhances the constraint on the model by making the output results of the two tend to be consistent. The prototype module, inside the student model and the teacher model, converts visual features into predicted attribute scores through the learning of attribute prototype vectors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of zero-shot image classification, and relates to a zero-shot image classification network and a deep learning method based on contrast learning. Background Art

[0002] Zero-shot Learning (ZSL), also known as Zero-shot Classification (ZSC), is a technique that uses some auxiliary knowledge to train samples of known classes and complete the prediction of the categories of unknown class samples. In this classification scenario, the known classes are the training classes, the unknown classes are the test classes, and the known classes and the unknown classes are mutually exclusive. Different from traditional classification algorithms, zero-shot learning can separate samples that are completely missing in the training stage during the test stage, that is, identify samples of categories that have never appeared in the training set. Traditional classification algorithms can only separate samples whose test categories belong to the training classes. This learning mechanism of zero-shot learning greatly alleviates the dependence of traditional models on sufficient samples and data labels, and provides a reliable solution for the situation where training samples for the target task are completely missing.

[0003] The core idea of zero-shot learning is to imitate the reasoning ability of humans. Humans have strong learning and reasoning abilities. Without samples of the target task, they can complete the learning of a specific target by learning auxiliary knowledge related to the target task. A child can learn and summarize knowledge from a small number of examples. When a new category example appears, the new category can be recognized through a one-sentence description. Zero-shot learning hopes that the model can also have such an inferential ability to draw inferences from one instance. It uses semantic information as a bridge connecting the known classes and the unknown classes, allowing the model to learn and summarize knowledge from known class samples during the training stage, and then apply the learned knowledge to unknown class samples for classification during the test stage, thereby realizing knowledge sharing and transfer between the known training classes and the unknown test classes.

[0004] As shown in the attached figure Figure 1 shows a schematic diagram of zero-shot learning. The model realizes the prediction of the zebra category during the test stage by learning knowledge such as the shape of a horse, the color of a panda, and the stripes of a tiger during the training stage. The zero-shot learning model with reasoning ability makes the machine learning system more in line with the human learning mechanism, helps the artificial intelligence system get rid of the dependence on the labeled dataset of the target task, and further makes an important contribution to the realization of true artificial intelligence.

[0005] Zero-shot learning has gone through different stages of development. Most of the zero-shot image classification methods in the early stage belonged to the methods based on direct semantic prediction. In recent years, with the progress of deep learning technology, deep learning-based networks have shown an explosive growth. Using the deep visual features extracted by convolutional neural networks as the high-level semantic expression of image information has effectively improved the classification accuracy of the embedding model-based methods. In addition, the proposal of models such as generative adversarial networks has further provided a reliable solution idea for the zero-shot learning problem. According to the different technical routes for solving the zero-shot learning problem, the existing zero-shot image classification methods can be divided into three categories: (1) methods based on direct semantic prediction; (2) methods based on embedding models; (3) methods based on generative models.

[0006] Regarding the core issue of improving the image feature expression ability of discriminative and robust semantic attributes in the zero-shot image classification task, and then realizing the semantic alignment of cross-modal information in the embedding space, there are currently two typical solutions in the industry, which are to align the true value class attributes by designing a specific feature extractor or using prototype learning. The feature extractor uses category attribute information or local information for effective guidance to improve the visual representation of samples, so as to align to the corresponding class prototype. The method combined with prototype learning no longer regards the true value class attribute as a prototype, but uses a learnable visual prototype to perform feature expression of attribute semantics. However, when performing visual feature expression, these methods ignore the exploration of the semantic attributes themselves, lacking discriminative expression of semantic attributes among different categories and robust expression on the same category. Summary of the Invention

[0007] Technical Problems to be Solved

[0008] In order to avoid the deficiencies of the prior art, the present invention proposes a zero-shot image classification network and deep learning method based on contrast learning, and the main purpose is to improve the image feature expression ability of discriminative and robust semantic attributes in the zero-shot image classification task, and then realize the semantic alignment of cross-modal information in the embedding space.

[0009] Technical Solution

[0010] A zero-shot image classification network and deep learning method based on contrast learning, characterized by the following steps:

[0011] Step 1: Input the image x into the residual network 101 feature extraction network to obtain the visual feature f(x) ∈ R H*W*C , where H, W, and C respectively represent the height, width, and number of channels of the feature;

[0012] Step 2: Input the visual feature f(x) into a feature processing network composed of upper and lower branch features. The upper branch of the feature processing network is composed of prototype modules, and the lower branch is composed of global average pooling modules;

[0013] The input visual feature f(x) outputs the class attributes predicted for the image samples through the prototype module Obtain the semantic attribute prediction for the positive samples and the semantic attribute prediction for the negative samples;

[0014] The visual feature f(x) passes through the global average pooling layer to obtain the global feature g(x). The global feature is mapped to the class attribute space through a linear layer, and the dot product calculation is performed between the mapped global feature and all class attributes included in the dataset to obtain the class embedding. The class embedding refers to using a neural network to map a high-dimensional representation space to a low-dimensional distributed space;

[0015] ^

[0016] Among them, the upper branch of the feature processing network composed of prototype modules outputs the class attribute z predicted for the image samples, and obtains the semantic attribute prediction for the positive samples and the semantic attribute prediction for the negative samples;

[0017] Step 3: For each attribute, perform the inner product operation between the local feature f i,j (X) and the attribute prototype p a . The local feature f i,j (X) represents the feature at the spatial position (i, j) in f(x). Obtain the similarity map M a ∈R H*W for each attribute. By maximizing the value of the similarity map of the a-th attribute, obtain the attribute prediction score of the a-th attribute of the input image;

[0018] Step 4: Perform the classification network

[0019] Use the class embedding with the highest attribute prediction score as the class of the input image

[0020]

[0021] Where g(x) T is the transpose matrix of g(x), is the true value class attribute vector corresponding to the test class , is the predicted class, and V represents the mapping matrix;

[0022] When is a known class, the function I = 1. When is an unknown class, I = 0. The above formula becomes:

[0023]

[0024] The known classes represent the classes that appear in the inference stage.

[0025] The upper branch in the feature processing network is composed of prototype modules, and the prototype modules are connected in series by two fully connected layers.

[0026] The lower branch is composed of a global average pooling module, and the global average pooling module is composed of channel average pooling. Channel average pooling means representing the features of the channel by the arithmetic mean of the channel.

[0027] The classification network is formed by connecting in series a residual network 101 feature extraction network and a feature processing network.

[0028] The prototype module inputs the visual features extracted by the convolutional neural network and outputs the class attributes predicted by the image samples. Maximize the value of the similarity map of the a-th attribute to obtain the attribute prediction score of the a-th attribute.

[0029] The true class attributes of the prototype module are 50, and the similarity map is calculated by the vector inner product method. The dimension of each class of true class attributes is 256 dimensions.

[0030] Beneficial effects

[0031] A zero-shot image classification network and deep learning method based on contrast learning proposed by the present invention. The designed model includes three parts: contrastive learning embedding of semantic attributes, consistency constraint between the student model and the teacher model, and prototype modules. Among them, the contrastive learning embedding part of semantic attributes starts from two levels of instance-level supervision and class-level supervision, and conducts contrastive learning on the predicted semantic attributes of positive and negative samples and the predicted semantic attributes of positive and negative classes. The consistency constraint part introduces the Mean Teacher mechanism, inputs images with different data augmentations to the student model and the teacher model, and then enhances the constraint on the model by making the output results of the two tend to be consistent. The prototype module inside the student model and the teacher model converts visual features into predicted attribute scores through the learning of attribute prototype vectors.

[0032] The beneficial effects of the present invention:

[0033] (1) Introduce the contrastive learning method into the model based on the embedding space, add instance-level supervision on the basis of the previous class-level supervision, and effectively alleviate the misclassification problem of images of the same class and the confusion problem of images of different classes in the fine-grained dataset.

[0034] (2) Adopt contrastive learning embedding on semantic attributes, enhance the image feature expression ability of discriminative and robust semantic attributes in the zero-shot image classification task, reduce the modal difference between semantic attributes and visual features, and achieve semantic alignment of cross-modal features.

[0035] (3) By adopting the Mean Teacher mechanism, the consistency regularization constraint between the student model and the teacher model further improves the robustness of the mapping between cross-modal visual information and semantic information in the embedding space. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is a schematic diagram of zero-shot learning;

[0037] Figure 2 is the zero-shot image classification network based on contrastive learning of the present invention;

[0038] Figure 3 is a schematic diagram of the prototype module of the present invention;

[0039] Figure 4 is an example display of the classification results of the algorithm in this chapter on the CUB dataset of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The present invention will be further described below in conjunction with the embodiments and the drawings:

[0041] To make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below through the drawings and embodiments. However, it should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the scope of the present invention. In addition, in the following description, the description of well-known structures and technologies is omitted to avoid unnecessarily confusing the concepts of the present invention.

[0042] In zero-shot learning, given the known classes Y s , the unknown classes Y u , and the known classes and unknown classes are mutually exclusive, the training set D s = {(x i , y i , z i ), where x i , y i respectively represent the training image and its corresponding class label, and z i = [z I ,......z A , represents the corresponding true class attribute vector. The test set D s = {(x i , y i , z i)} It includes images of unknown classes and their corresponding ground-truth class attribute vectors. Next, an end-to-end zero-shot image classification model based on contrastive constraints proposed in this paper will be introduced. It effectively improves the image feature representation ability of semantic attributes through contrastive learning, making it more discriminative and robust. Through the Mean Teacher mechanism, it enhances the robustness of the model for fine-grained image classification with stronger consistency regularization constraints.

[0043] An embodiment of this application also discloses a zero-shot image classification network based on contrastive learning. As shown in the appendix Figure 2 , the network structure of the network includes three parts: a contrastive learning embedding module for semantic attributes, a consistency constraint module for the student model and the teacher model, and a prototype module.

[0044] The contrastive learning embedding module is formed by connecting the backbone feature extraction network in series with the prototype module. The prototype module is connected in parallel with global average pooling. Preferably, the backbone feature extraction network uses ResNet 101.

[0045] The consistency constraint module for the student model and the teacher model means that given positive example samples under the same category and negative example samples under different categories, they are input into the student model and the teacher model through different data augmentations. The network structures of the student model and the teacher model are the same. The parameters of the student model are calculated by exponential moving average to obtain the parameters of the teacher model. The consistency loss of their classification results is measured. The consistency regularization constraint enhances the robustness of the model for zero-shot image classification. Inside the student model and the teacher model, the images are used to extract visual features through a convolutional neural network, and the visual features are divided into two branches, the upper and the lower, for cross-modal mapping embedding of visual semantic information.

[0046] The prototype module inputs the visual features extracted by the convolutional neural network and outputs the class attributes predicted by the image samples. Maximize the value of the similarity map of the a-th attribute to obtain the attribute prediction score of the a-th attribute.

[0047] Preferably, the number of ground-truth class attributes of the prototype module is 50, and the similarity map is calculated by the vector inner product method. The dimension of each ground-truth class attribute is 256.

[0048] The model designed in this invention includes three parts: contrastive learning embedding of semantic attributes, consistency constraint between the student model and the teacher model, and the prototype module. Among them, the contrastive learning embedding part of semantic attributes starts from two levels: instance-level supervision and category-level supervision, and conducts contrastive learning on the predicted semantic attributes of positive and negative samples and the predicted semantic attributes of positive and negative categories. The consistency constraint part introduces the Mean Teacher mechanism, inputs images with different data augmentations to the student model and the teacher model, and then enhances the model constraint by making the output results of the two tend to be consistent. The prototype module, inside the student model and the teacher model, converts visual features into predicted attribute scores through the learning of attribute prototype vectors.

[0049] Contrastive learning embedding module

[0050] Given an image x i , construct a unique positive sample image x + and k negative sample images to be jointly used as the model input. Among them, the positive sample x + is a randomly selected sample of the same category as the image x i , and the negative samples are randomly selected samples of mutually exclusive categories with the image x i . First, input the sample images into the backbone network ResNet-101 to obtain the visual features f(x) ∈ R H*W*C , where H, W, and C represent the height, width, and number of channels of the features respectively. Divide the obtained visual features into two upper and lower branches for cross-modal visual information and semantic information embedding. The upper branch embeds the local features f(x). The local features pass through the prototype module for semantic attribute prediction of positive samples and semantic attribute prediction of negative samples. Compare and learn the semantic attributes predicted by positive and negative samples. Through the strong contrast constraint between positive and negative samples, improve the discriminative expression of semantic attributes between different categories and the robust expression ability on the same category. The contrast of semantic attributes at the instance level and learning supervision effectively narrow the difference between cross-modal visual information and semantic information. Specifically, the contrastive learning embedding based on semantic attributes at the instance level can be expressed by the following formula

[0051]

[0052] where z ∈ R S*A is the predicted semantic attribute, S is the total number of categories, and A is the total number of attributes. τ eis the temperature coefficient of contrastive learning embedding. K is the number of negative samples. A larger K value ensures the constraint effect of contrastive learning, enabling the model to capture discriminative semantic attribute features. The similarity calculation with positive samples enables the model to capture robust semantic attribute features. The lower branch embeds the global features. First, the visual feature f(x) passes through a Global Average Pooling (GAP) layer to obtain the global feature. In the embedding space, the global feature is mapped to the class attribute space through a linear layer, and this linear layer contains a learnable parameter, V ∈ R C*N , where N is the number of semantic attributes in the class attribute. The dot product is calculated between the mapped global feature and all S class attributes included in the dataset, encouraging the dot product calculation result of the global feature of the sample and its true class attribute. The true class attribute can be regarded as a positive example, reducing the similarity measure between the global feature of the sample and other S1 class attributes. Other S1 class attributes can be regarded as negative examples. The classification problem of the S class can be calculated through the cross-entropy loss function, and the category-level visual semantic contrastive learning embedding obtains the final classification result of the model,

[0053]

[0054] where z represents the true class attribute of the sample, and z s represents the class attributes of all S classes. The category-level visual semantic contrastive learning embedding strengthens the discriminative ability of semantic attribute features in the embedding space.

[0055] Prototype module

[0056] To further achieve the semantic alignment of cross-modal visual features and semantic features, the present invention uses a prototype module to effectively localize semantic attributes on visual features. As shown in the appendix Figure 3 , the prototype module takes the visual feature f(x) extracted by the convolutional neural network as input and outputs the class attribute predicted by the image sample Inside the prototype module, the local feature f i,j (X) ∈ R C encodes the local region of the image (Region), and the attribute prototype serves as the learning parameter of the prototype module to help the local feature of the image predict the score of each attribute, where p a represents the prototype feature of the a-th attribute. For each attribute, the similarity map M i,j of each attribute is obtained through the inner product operation between the local feature f a and the attribute prototype p a ∈ R H*W, which represents the localization result of each attribute on the image features, and the localization result can effectively demonstrate the cross-modal semantic alignment effect. The similarity map of the a-th attribute located at the spatial position (i, j) is calculated by the following formula Finally, by maximizing the value of the similarity map of the a-th attribute, we obtain the attribute prediction score of the a-th attribute. The prototype module associates each visual attribute with the local features, enabling the model to effectively localize each attribute on the local features and obtain an attribute prediction score close to the true result. This patent regards the attribute score prediction task as a regression problem and uses the Mean Square Error (MSE) loss to measure the attribute prediction score and the true value class attribute. Then the formula for the attribute regression loss is as follows

[0057] where is the predicted attribute score, and z is the true value class attribute vector.

[0058] Consistency Constraint

[0059] To further improve the robustness of the cross-modal semantic alignment mapping in the embedding space, this patent introduces the MeanTeacher mechanism to enhance the constraint on the model. Compared with directly using the weights of the model, it generates a more accurate and robust model by continuously averaging the weights of the model at each training step (Step). The Mean teacher mechanism includes a Student Model and a Teacher Model. The internal structures of the two models are exactly the same, but the parameters are different. By means of Exponential Moving Average (EMA), the weight parameters of the student model and the teacher model at each training step are calculated. Given that the parameter of the student model is θ, when the training step is t, the parameter θ t ' of the teacher model is calculated by exponential moving average as θ′ t = aθ′ t-1 +(1 - a)θ t , where a is the smoothing coefficient, taking 0.95.

[0060] The images x and x′ with different data augmentations are respectively input into the student model and the teacher model. After passing through the encoders, prototype modules, and classifiers with the same structure but different parameters, different classification results of the same sample are obtained and For different data - augmented input images, we hope that the output classification results of the student model and the teacher model tend to be consistent to improve the robustness of the mapping between cross - modal visual information and semantic information. To this end, the Mean Square Error (MSE) is used as the consistency loss between the student model and the teacher model.

[0061]

[0062] Among them, is the classification result of the student model, is the classification result of the teacher model. Through experimental comparison, the classification result of the student model is finally used as the prediction output result of the entire model.

[0063] Zero - shot inference

[0064] For zero - shot learning tasks, given an input image x, the classifier searches for the class embedding with the highest compatibility score in the following way where is the test class and the corresponding true - value class - attribute vector. For generalized zero - shot learning tasks, the test phase contains not only samples of known classes but also samples of unknown classes, and bias problems are likely to occur. The model is likely to predict samples of unknown classes as classes of known classes during the test phase. To alleviate the bias problem in generalized zero - shot learning, we adopt Calibrated Stacking (CS), which is achieved by reducing the classification scores of samples on known classes. The classifier searches for the class embedding with the highest compatibility score in the following way

[0065]

[0066] where, when is a known class, the function I = 1, and when is an unknown class, I = 0, and γ is a hyper - parameter. During the training phase, the model is jointly optimized using four loss functions, L CC-ZSL = L cls + λ1L con + λ2L sis + λ3L reg where λ1, λ2, and λ3 are hyper - parameters of the model. Joint loss training enhances the discriminative expression of semantic attributes between different classes on image features and the robust expression on the same class, helps the semantic alignment of cross - modal visual information and semantic information in the embedding space, and further alleviates the semantic gap problem in zero - shot learning.

[0067] Embodiment:

[0068] 1. Datasets and Evaluation Metrics

[0069] The present invention evaluates the model performance using three benchmark datasets widely used in zero - shot learning tasks, namely the Caltech - UCSD Birds - 200 - 2011 (CUB) dataset, the SUN attributes (SUN) dataset, and the Animals with Attributes 2 (AWA2) dataset, and quantitatively compares the model methods using the Top - 1 class - average accuracy and the harmonic mean accuracy.

[0070] 2. Data Preparation and Experimental Settings

[0071] The present invention uses the Resnet - 101 model pre - trained on the ImageNet - 1k dataset as the backbone network to extract features and fine - tunes it (with fine - tuning). Given an input image of size 224×224, through different image data augmentations, it is input into the backbone networks of the student model and the teacher model for feature extraction, obtaining a feature map of size 7×7×2048. Then, the upper and lower branches inside the model are used to learn the mapping of cross - modal visual information and semantic information. In terms of model optimization in the deep learning network, the Adam optimizer is used to perform deep gradient calculations to obtain the best model parameters. In terms of the hyperparameter settings of the loss function, the model obtains the best hyperparameters through grid search on the validation set. The coefficient λ1 of the total loss function is 1, and λ2 takes different values on different datasets. λ2 takes the value of 1 on the AWA2 dataset, 100 on the CUB dataset, and 1000 on the SUN dataset. λ3 takes the value of 0.0001 on the AWA2 dataset and the SUN dataset, and 0.01 on the CUB dataset. The default value of the smoothing coefficient α of exponential moving average in Mean - teacher is 0.999, and during training, α takes min(1 / (1 - t), α), where t is the training step. The calibration parameter γ is 0.7 on the CUB and AWA2 datasets and 0.4 on the SUN dataset. The entire model is trained only on a single 3090Ti GPU card and is built using the deep learning framework Pytorch.

[0072] 3. Sample Settings

[0073] The present invention sets 9 different numbers of positive and negative samples for detailed performance comparison. Figure 4A schematic diagram showing the experimental results obtained using the present invention on the CUB dataset. Let K and N represent setting K categories in a mini-batch, with N samples in each category. Then, for a positive sample, it is subjected to contrastive learning with (K - 1)*N negative samples. On the AWA2 dataset, through a large number of experiments with different numbers of positive and negative samples, it is found that when K is 8 and N is 12, the best classification results are achieved in both the generalized zero-shot learning task and the traditional zero-shot learning task, that is, the harmonic mean accuracy (H) is 71.1% and the average top-1 accuracy (T1) of the unknown class is 68.8%. When a positive sample is subjected to contrastive learning embedding with a total of 84 negative samples under 7 other negative example categories on the AWA2 dataset, the effect is the best. On the CUB dataset, when K is 8 and N is 12, the best harmonic mean accuracy (H) of 69.3% is achieved in the generalized zero-shot learning task, and the best average top-1 accuracy (T1) of the unknown class of 74.3% is achieved in the traditional zero-shot learning task, that is, when a positive sample is subjected to contrastive learning embedding with a total of 84 negative samples under 7 other negative example categories, the effect is the best. Since both the AWA2 dataset and the CUB dataset belong to animal datasets, the model needs to pay more attention to the local features of animals for contrastive learning embedding. Therefore, the more negative samples in a negative example category, the more helpful it is for the model's embedding. However, the total number of negative example categories is not necessarily the more the better. The local feature contrastive learning under too many categories may cause confusion in the model and thus affect classification. On the SUN dataset, when K is 12 and N is 4, the best harmonic mean accuracy (H) of 40.3% is achieved in the generalized zero-shot learning task, and the best average top-1 accuracy (T1) of the unknown class of 62.4% is achieved in the traditional zero-shot learning task, that is, when a positive sample is subjected to contrastive learning embedding with a total of 44 negative samples under 11 other negative example categories, the effect is the best.

[0074] 4. Mean-teacher Hyperparameter Selection

[0075] When the value of λ2 is 1, the highest classification accuracy is achieved in both the generalized zero-shot task and the traditional zero-shot task. At this time, the harmonic mean accuracy is 71.1% and the average top-1 accuracy of the unknown class is 68.8%. On the CUB dataset, when the value of λ2 is 100, the best harmonic mean accuracy of 69.3% is achieved in the generalized zero-shot task, and the best harmonic mean accuracy of 74.3% is achieved in the traditional zero-shot task. On the SUN dataset, when the value of λ2 is 1000, the best harmonic mean accuracy of 40.3% is achieved in the generalized zero-shot task, and the best harmonic mean accuracy of 62.4% is achieved in the traditional zero-shot task.

[0076] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements or improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A zero-shot image classification network and deep learning method based on contrastive learning, characterized in that The steps are as follows: Step 1: Input the image x into the residual network 101 feature extraction network to obtain the visual feature f(x) ∈ R H*W*C , where H, W, and C represent the height, width, and number of channels of the feature, respectively; Step 2: Input the visual feature f(x) into a feature processing network composed of upper and lower branch features. The upper branch of the feature processing network is composed of prototype modules, and the lower branch is composed of global average pooling modules; The input visual feature f(x) passes through the prototype module to output the predicted class attributes of the image samples Obtain the predicted semantic attributes of the positive samples and the predicted semantic attributes of the negative samples; The visual feature f(x) passes through a global average pooling layer to obtain a global feature g(x). The global feature is mapped to a class attribute space through a linear layer. The dot product of the mapped global feature and all class attributes included in the data set is calculated to obtain a class embedding. The class embedding refers to using a neural network to map a high-dimensional representation space to a low-dimensional distributed space; Among them, the upper branch composed of prototype modules in the feature processing network outputs the class attributes predicted by the image samples. Obtain the semantic attribute prediction of the positive samples and the semantic attribute prediction of the negative samples. Step 3: For each attribute, perform the inner product operation between the local feature f i,j (X) and the attribute prototype p a . The local feature f i,j (X) represents the feature at the spatial position (i, j) in the f(x) space. Obtain the similarity map M a ∈R H*W for each attribute. By maximizing the value of the similarity map of the a-th attribute, obtain the attribute prediction score of the a-th attribute of the input image; Step 4: Use a classification network for classification Use the class embedding with the highest attribute prediction score as the class of the input image where \(g(x)\) T is the transpose matrix of \(g(x)\), is the test class corresponding true value class attribute vector, is the predicted class, \(V\) represents the mapping matrix, and \(\gamma\) is the hyperparameter; When is known to be time-like, the function I = 1, when is unknown to be time-like, I = 0, and the above formula becomes: Known class representations are the classes that appear in the inference stage.

2. The zero-shot image classification network and deep learning method based on contrastive learning according to claim 1, characterized in that: The upper branch in the feature processing network is composed of prototype modules, and the prototype modules are formed by cascading two fully connected layers.

3. The zero-shot image classification network and deep learning method based on contrastive learning according to claim 1, characterized in that: The lower branch is composed of global average pooling modules, and the global average pooling modules are composed of channel average pooling. Channel average pooling refers to representing the features of a channel with the arithmetic mean of that channel.

4. The zero-shot image classification network and deep learning method based on contrastive learning according to claim 1, characterized in that: The classification network is composed of a residual network 101 feature extraction network and a feature processing network in series.

5. The zero-shot image classification network and deep learning method based on contrastive learning according to claim 1, characterized in that: The prototype module inputs the visual features extracted by the convolutional neural network, outputs the class attributes predicted by the image samples, maximizes the value of the similarity map of the a-th attribute, and obtains the attribute prediction score of the a-th attribute.

6. The zero-shot image classification network and deep learning method based on contrastive learning according to claim 1 or 5, characterized in that: The true class attributes of the prototype module are 50. The similarity map is calculated by vector inner product, and the dimension of each class of true class attributes is 256 dimensions.

Citation Information

Patent Citations

  • Method and system for generating a vector representation of an image

    CA3068891A1

  • Zero sample learning method based on global semantic consistency network

    CN108846413A