Zero-shot Learning Method for Generating and Correcting Based on Instance Image Attributes
Through a generalized zero-sample learning method based on instance image attribute generation and correction, the domain transfer and semantic gap problems caused by insufficient attribute labeling of the benchmark data set in the prior art are solved, and a higher unseen class recognition rate and calculation efficiency are achieved.
Patent Information
- Application Number
- CN202411486125.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-23
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-10-23
AI Technical Summary
The existing zero-sample learning method relies on standardized benchmark data sets with each class of attribute annotations, and cannot effectively deal with the problem of domain transfer and semantic gap, making it difficult for the model to accurately match category attributes in practical applications.
The generalized zero-sample learning method based on instance image attribute generation and correction is adopted. Through the attribute generation module and attribute correction module of the PIAS model, the attributes of the instance image are generated and corrected to ensure the close consistency between attributes and visual features, and to solve the problem of domain transfer and semantic gap.
ViT and multi-layer perceptron simplify feature extraction and attribute generation, reduce computational complexity, improve attribute accuracy and diversity, significantly improve the recognition rate of unknown classes, and alleviate semantic gaps and domain offset problems.
Smart Images

Figure CN119540681B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image classification, and particularly to a zero-shot learning method based on instance image attribute generation and correction. Background Art
[0002] Generalized Zero-Shot Learning (GZSL) has broad application prospects in various real-world scenarios, especially in tasks where unseen classes frequently appear. GZSL aims to enable a model to effectively recognize classes that have not been seen during training through its generalization ability, thereby greatly enhancing the practicality and adaptability of the model. For example, in fields such as image recognition, natural language processing, and even autonomous driving, the model often faces inputs from unseen classes, and GZSL provides a potential solution for such scenarios. However, most current GZSL methods rely on standardized benchmark datasets with per-class attribute annotations. Although these datasets provide detailed class attribute information, they also bring new problems. Specifically, since the attribute annotations of these benchmark datasets are usually based on predefined per-class features, they may not cover all possible nuances and complex semantics, resulting in the model having difficulty precisely matching class attributes when encountering unseen classes in practical applications. This difference leads to the so-called "semantic gap", that is, the model cannot accurately map between visual features and semantic representations, causing a deviation in understanding.
[0003] On the other hand, methods that rely on these attribute annotations tend to exacerbate the domain transfer problem in the visual-semantic space. Domain transfer refers to the phenomenon where the performance of a model degrades when it faces different data distributions or semantic features during training and testing. In GZSL, the domain transfer problem is particularly obvious because the model not only has to cope with changes in visual features but also maintain stable performance in the semantic space. When there are significant differences in class attributes or features between the training data and the test data, the performance of the model often degrades severely. Therefore, although existing methods can achieve good performance in a controlled environment, in actual complex real-world scenarios, this strategy that relies on the attribute annotations of benchmark datasets is insufficient and cannot effectively address the problems of domain transfer and semantic gap. Summary of the Invention
[0004] Aiming at the above deficiencies in the prior art, the zero-shot learning method based on instance image attribute generation and correction provided by the present invention solves the problem that existing methods cannot effectively address domain transfer and semantic gap when applied to complex actual scenarios due to insufficient attribute annotations.
[0005] To achieve the above object of the invention, the technical solution adopted by the present invention is as follows:
[0006] A generalized zero-shot learning method for generating and correcting based on instance image attributes is provided. This method is for training the PIAS model for image classification. The PIAS model includes an attribute generation module and an attribute correction module, and specifically includes the following steps:
[0007] S1. Generate the class average image for each class according to all instance images in the visible class dataset;
[0008] S2. Use the Vision Transformer (ViT) of the attribute generation module to construct a visual space to extract the visual features of each instance image in the visible class dataset and the average visual features of the class average image;
[0009] S3. According to the visual features of the instance image and the class average image, use the multi-layer perceptron of the attribute generation module to generate the instance attributes of the instance image and the class-level attributes of the class average image;
[0010] S4. In the attribute correction module, use the classification loss to keep the consistency between each instance attribute and its corresponding class label;
[0011] S5. Use the calibration loss to construct a semantic space to minimize the cosine similarity between the average attributes of the class and the annotated attributes;
[0012] S6. Use the attribute annotation and average visual features of each class as the anchor points of the corresponding class to control the structural alignment of the semantic space and the visual space;
[0013] S7. Input the test dataset into the attribute generation module to generate the attributes of each test image in the test dataset, and input the attributes of each test image into the ZSL classifier to obtain the labels of the test images;
[0014] S8. According to the labels of the test images, determine the classification accuracy, and judge whether the classification accuracy and the total loss function of the PIAS model meet the preset conditions. If so, end the algorithm; otherwise, return to step S2.
[0015] Further, step S6 further includes:
[0016] S61. Calculate the attribute fusion matrix according to the attribute annotation of the class and the instance attributes of all images of the corresponding class:
[0017]
[0018] where, is the attribute fusion matrix of the i-th class; a i is the attribute annotation of the i-th class; is the set of attributes of all instance images of the i-th class; and are the first and N-th images in the i-th class respectivelyi Instance attributes of N instance images i is the total number of instance images of the i-th class; T is the transpose
[0019] S62. Calculate the visual feature fusion matrix based on the average visual feature of the class and the visual features of all images corresponding to the class:
[0020]
[0021] where is the visual feature fusion matrix of the i-th class is the average visual feature of the i-th class and are the visual features of the 1st and N-th instance images in the i-th class respectively i H is the set of visual features of all instance images of the i-th class i S63. Control the structural alignment of the semantic space and the visual space according to the topological structure similarity between the fusion matrix of attributes and the fusion matrix of visual features:
[0022] where
[0023]
[0024]
[0025] where is the structural consistency loss; ls is the total number of classes in the visible class dataset is the value corresponding to the j-th instance image in ; is the value corresponding to the j-th instance image in ; exp(.) is the exponential function
[0026] Furthermore, the expression of the calibration loss is:
[0027]
[0028] where is the calibration loss is the class-level attribute of the i-th class; a i is the attribute annotation of the i-th class; ls is the total number of classes in the visible class dataset; cos(.) is the cosine function
[0029] Furthermore, the expression of the classification loss is:
[0030]
[0031] where is the classification loss; ls is the total number of classes in the visible class dataset; y i is the class label; exp(.) is the exponential function; is the set of attribute annotations a for all classes in the visible class dataset i ; is the instance attribute of the j-th instance image of the i-th class; N i is the total number of instance images of the i-th class; T is the transpose.
[0032] Furthermore, the expression of the total loss function of the PIAS model is:
[0033]
[0034] where, is the total loss; α, β, and δ are hyperparameters used to balance different loss terms respectively; is the structural consistency loss; is the calibration loss; is the classification loss.
[0035] Furthermore, the expression for generating the class average image is:
[0036]
[0037] where, is the class average image of the i-th class in the visible class dataset; is the j-th instance image in the i-th class; N i is the total number of instance images of the i-th class.
[0038] Furthermore, the ZSL classifier is a CZSL classifier or a GZSL classifier;
[0039] The expression of the CZSL classifier is:
[0040]
[0041] where, is the predicted label; is the attribute of the test image; y is the label; is the label of the unseen class; is the attribute of the unseen class; T is the transpose; max is to take the maximum value;
[0042] The expression of the GZSL classifier is:
[0043]
[0044] where, is the label of the visible class; All attributes of the visible and unseen classes; γ is the calibration factor.
[0045] Furthermore, the preset condition is that the classification accuracy is less than the preset accuracy and the change amount of the value of the total loss function of the PIAS model during continuous multiple trainings is less than the preset value.
[0046] The beneficial effects of the present invention are as follows: This solution simplifies the feature extraction and attribute generation processes through the Vision Transformer (ViT) and the Multi-Layer Perceptron (MLP), effectively reducing the computational complexity while maintaining the high performance of the model. By correcting the visual attributes of each instance through the attribute correction module, the accuracy and diversity of the attributes can be improved, the consistency between the attributes and the true class attributes can be significantly enhanced, and the recognition rate of the unseen classes can be directly improved.
[0047] This solution can synthesize per-instance attributes with spatial information through the patch and position embedding techniques introduced by the Vision Transformer (ViT); by introducing a calibration loss to map class anchors and a structural consistency loss to align the topology between the visual space and the semantic space; these two losses promote the diversity within the class and maintain the semantic consistency of per-instance attributes, further alleviating the semantic gap and domain shift problems in the zero-shot learning classifier.
[0048] This solution can ensure the tight consistency between the generated attributes and the visual features by controlling the structural alignment of the semantic space and the visual space, and solve the domain shift problem in traditional zero-shot learning. Through this alignment, the recognition accuracy of the model in the Generalized Zero-Shot Learning (GZSL) task is significantly improved. Description of the Drawings
[0049] Figure 1 It is a flowchart of the generalized zero-shot learning method based on instance image attribute generation and correction.
[0050] Figure 2 It is a principle block diagram of the PIAS model. Detailed Embodiments
[0051] The following describes the detailed embodiments of the present invention to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
[0052] Such as Figure 2As shown in the figure, this solution designs a PIAS model. The PIAS model includes an attribute generation module and an attribute correction module. The attribute generation module further includes a Vision Transformer (ViT) and a multi-layer perceptron (MLP). The attribute correction module includes three losses, namely, structural consistency loss, calibration loss, and classification loss.
[0053] Referring to the figure, Figure 1 A generalized zero-shot learning method based on instance image attribute generation and correction is shown. The method S is for training a PIAS model for image classification, including steps S1 to S8.
[0054] In step S1, according to all instance images of each class in the visible class dataset, the class average image of each class is generated:
[0055]
[0056] where, is the class average image of the i-th class in the visible class dataset; is the j-th instance image in the i-th class; N i is the total number of instance images of the i-th class.
[0057] In step S2, the Vision Transformer (ViT) of the attribute generation module is used to construct a visual space to extract the visual features of each instance image in the visible class dataset and the average visual feature of the class average image:
[0058]
[0059] where, is the average visual feature of the i-th class; is the visual feature of the j-th instance image in the i-th class; θ is the learnable parameter of the ViT model; d x is the dimension of the visual feature.
[0060] In step S3, according to the visual features of the instance image and the class average image, the instance attributes of the instance image and the class-level attributes of the class average image are generated using the multi-layer perceptron of the attribute generation module:
[0061]
[0062] where, φ represents the learnable parameter of the MLP, d a represents the dimension of the attribute.
[0063] The introduction of the multi-layer perceptron can provide a wide range of visual and attribute features for subsequent processing.
[0064] In step S4, in the attribute correction module, the classification loss is adopted to keep the consistency between each instance attribute and its corresponding class label.
[0065] In implementation, the expression of the preferred classification loss in this solution is:
[0066]
[0067] Where, is the classification loss; ls is the total number of classes in the visible class dataset; y i is the class label; exp(.) is the exponential function; is the set of attribute annotations a i of all classes in the visible class dataset; is the instance attribute of the j-th instance image of the i-th class; N i is the total number of instance images of the i-th class; T is the transpose.
[0068] In step S5, the calibration loss is adopted to construct the semantic space to minimize the cosine similarity between the average attribute of the class and the annotated attribute; the expression of the calibration loss is:
[0069]
[0070] Where, is the calibration loss; is the class-level attribute of the i-th class; a i is the attribute annotation of the i-th class; ls is the total number of classes in the visible class dataset; cos(.) is the cosine function.
[0071] In step S6, the attribute annotation of each class and the average visual feature are used as the anchor points of the corresponding class to control the structural alignment of the semantic space and the visual space.
[0072] In an embodiment of the present invention, step S6 further includes:
[0073] S61. Calculate the attribute fusion matrix according to the attribute annotation of the class and the instance attributes of all images of the corresponding class:
[0074]
[0075] Where, is the attribute fusion matrix of the i-th class; a i is the attribute annotation of the i-th class; is the set of attributes of all instance images of the i-th class; and are respectively the instance attributes of the first and the N i -th instance images in the i-th class; N iis the total number of instance images of the i-th class; T is the transpose;
[0076] S62. Calculate the visual feature fusion matrix based on the average visual feature of the class and the visual features of all images corresponding to the class:
[0077]
[0078] where, is the visual feature fusion matrix of the i-th class; is the average visual feature of the i-th class; and are the visual features of the 1st and N-th i instance images in the i-th class respectively; H i is the set of visual features of all instance images of the i-th class;
[0079] S63. Control the structural alignment of the semantic space and the visual space according to the topological structure similarity between the fusion matrix of attributes and the fusion matrix of visual features:
[0080]
[0081] where, is the structural consistency loss; ls is the total number of classes in the visible class dataset; is the value corresponding to the j-th instance image in is the value corresponding to the j-th instance image in
[0082] Through controlling the structural alignment of the semantic space and the visual space, this solution can make each instance attribute more detailed and accurate, so as to construct a more refined semantic space.
[0083] In step S7, input the test dataset into the attribute generation module to generate the attributes of each test image in the test dataset, and input the attributes of each test image into the ZSL classifier to obtain the labels of the test images.
[0084] In implementation, this solution preferably selects the ZSL classifier as the CZSL classifier or the GZSL classifier;
[0085] The expression of the CZSL classifier is:
[0086]
[0087] where, is the predicted label; is the attribute of the test image; y is the label; is the label of the unseen class; Attributes of unseen categories; T is the transpose; max is to take the maximum value;
[0088] The expression of the GZSL classifier is:
[0089]
[0090] Among them, is the label of the visible category; is all the attributes of the visible and unseen categories; γ is the calibration factor.
[0091] In step S8, according to the label of the test image, the classification accuracy is determined, and it is judged whether the classification accuracy and the total loss function of the PIAS model meet the preset conditions. If so, the algorithm ends; otherwise, it returns to step S2.
[0092] Among them, the expression of the total loss function of the PIAS model is:
[0093]
[0094] Among them, is the total loss; α, β, and δ are hyperparameters used to balance different loss terms respectively; is the structural consistency loss; is the calibration loss; is the classification loss.
[0095] This solution preferably sets the preset condition as that the classification accuracy is less than the preset accuracy and the change amount of the value of the total loss function of the PIAS model during continuous multiple trainings is less than the preset value.
[0096] Next, the effect of the zero-shot learning method proposed in this solution is described in combination with specific embodiments:
[0097] In this embodiment, three widely recognized challenging benchmark datasets are selected in the field of generalized zero-shot learning (GZSL) for extensive experiments, including AWA2, CUB, and SUN. These datasets cover different fields from animals to birds to scenes, and show obvious differences in fine-grained and coarse-grained aspects. Therefore, they can comprehensively evaluate the performance of the PLAS model in different tasks and scenarios.
[0098] AWA2 (Animals with Attributes 2) is a coarse-grained dataset that mainly focuses on the recognition of animal categories. This dataset contains 50 animal categories, divided into 40 seen categories and 10 unseen categories, with a total of 37,322 images collected. AWA2 uses 85 attributes to describe the characteristics of these animal categories, and these attributes cover important biological features such as the appearance and behavior of animals. For example, the attributes may include "hooved", "herbivorous", or "furry", etc. Since the categories in AWA2 are relatively macroscopic, each category in the dataset may contain a wide range of differences across species, so AWA2 is considered a coarse-grained dataset. In zero-shot learning, the challenge of a coarse-grained dataset is how to enable the model to generalize to unseen animal categories through a small number of attribute features.
[0099] CUB (Caltech-UCSD Birds-200-2011) is a fine-grained dataset specifically for bird recognition. The CUB dataset contains 11,788 images of 200 bird species, with 150 seen categories and 50 unseen categories. Each bird species is described by 312 detailed attributes, which cover various features of birds, such as color, morphology, feather structure, beak shape, etc. This makes CUB a typical fine-grained dataset because it requires the model to be able to capture the subtle differences between different species. Since many birds are very similar in appearance, such as having similar colors or body sizes, the difficulty of the CUB dataset lies in how to distinguish these categories through extremely detailed attributes. Therefore, fine-grained datasets pose higher requirements for the model, especially in dealing with subtle differences and complex visual features.
[0100] SUN (SUN Attribute) is a fine-grained dataset for scene recognition. The SUN dataset contains 717 scene categories and a total of 14,340 images, with 645 seen categories and 72 unseen categories. Different from animals or birds, the focus of the SUN dataset is on the classification of scene categories, such as urban landscapes, forests, beaches, classrooms, etc. Each scene is described by 102 attributes, which cover different visual and semantic features of the scene, such as "natural", "buildings", "water bodies", etc. Although the categories in SUN are relatively rich, the attribute features of each scene may be relatively similar. For example, many natural scenes may contain "trees" and "water bodies", which increases the complexity of the zero-shot learning task. Due to the diversity of scene recognition and the high overlap between attributes, the model needs to find the key differences in complex visual information to identify unseen scene categories, which makes SUN a highly challenging fine-grained dataset.
[0101] In this embodiment, a pre-trained ViT-Large model is used as the visual feature extractor. To optimize the model, the Adam optimizer is used with an initial learning rate of 10 -6 , and a decay factor of 10 -5 . Additionally, a multi-step learning rate scheduler is used to adjust the learning rate at predefined training epochs (10, 20, 30, 40 epochs), with a scaling factor of 0.5 (gamma = 0.5). This method is implemented based on the PyTorch platform, leveraging its parallel distributed training framework, and is trained on two NVIDIA GeForce RTX 4090 GPUs. The entire training process lasts for 80 epochs with a batch size of 30 (15 images are processed on each GPU). The calibration factor γ is set to 0.9, 0.2, and 0.1 on the AWA2, CUB, and SUN datasets respectively.
[0102] To ensure a fair comparison with previous related methods, this embodiment adopts a unified evaluation protocol. For zero-shot learning (ZSL), according to the unified evaluation protocol, the top-1 (T1) average class accuracy of all unseen classes is evaluated. For generalized zero-shot learning (GZSL), first, the average class accuracy of the seen classes (S) and unseen classes (U) is calculated separately, and then their harmonic mean (H) is used to evaluate the performance of the method, where H = (2 * S * U) / (S + U). A higher harmonic mean indicates that the method performs well on both seen and unseen classes, reflecting a balanced and robust performance.
[0103] This embodiment compares the proposed PIAS model with a large number of the latest methods that have emerged in the past four years, including generative methods and non-generative methods. These methods include generative methods such as HVSA 1 , SDGZSL 2 , AGZSL 3 , SE-GZSL 4 , SRSA 5 , ICCE 6 , SC-EGG 7 , VS-Boost 8 , EGANS 9 and CSDR 10 , and non-generative methods including ViT-ZSL 11 , Transzero 12 , Transzero++ 13 , MSDN 14 , LAPE 15 , HRT 16 , PSVMA 17 , DUET 18 , IEAM-ZSL19 , AS-ZSL 20 and ZSLViT 21 . The above 21 models of this embodiment respectively correspond to the following references (1) to (21). For the result comparison between the PIAS model and the 21 models of the prior art, refer to Table 1 specifically.
[0104] Table 1 Comparison of the PIAS model with the state-of-the-art methods under the AWA2, CUB, and SUN settings
[0105]
[0106]
[0107] As shown in Table 1, the PIAS model proposed in this solution has achieved the best results under the GZSL setting of all datasets. Measured by the harmonic mean H, it is 79.7% for the AWA2 dataset, 75.7% for the CUB dataset, and 63.1% for the SUN dataset. Specifically, on the AWA2 dataset, most methods show significant differences between the seen classes and the unseen classes. For example, the accuracy of ViT-ZSL on the seen classes reaches 90%, but only 51.9% on the unseen classes. The imbalance in accuracy between the seen classes and the unseen classes is a common problem caused by domain shift. Therefore, although these methods perform well on the seen classes, their overall harmonic mean performance is still not ideal due to their poor performance on the unseen classes.
[0108] Compared with the above methods, although the PIAS model proposed in this solution does not reach the highest performance on either the seen classes (87.9%) or the unseen classes (72.9%), it achieves a good balance between them. This balance enables the harmonic mean performance of this dataset to reach the optimal value (79.7%). This improvement is attributed to the structure alignment module introduced in the method of this solution, which effectively solves the domain shift problem.
[0109] On the fine-grained datasets CUB and SUN, the method proposed in this scheme achieves state-of-the-art performance in all metrics, namely S, U, and H. Similar to the second-best method LAPE, PIAS enhances semantic attribute features from pixels. However, PIAS outperforms LAPE on the CUB dataset, showing a balanced improvement of about 1%, indicating its superior ability in bridging the vision-semantic gap. Due to the low inter-class similarity and high class diversity of the SUN dataset, knowledge transfer and feature capture face significant challenges, making it one of the datasets with the greatest potential for improvement in zero-shot learning. The method of this scheme achieves an accuracy of over 60% on the SUN dataset for the first time, with a harmonic mean of 63.1%, representing a 10.8% performance improvement compared to the previous best method PSVMA. Overall, in GZSL, PIAS demonstrates excellent comprehensive capabilities and successfully addresses the challenges of all datasets. These results not only highlight the superiority of PIAS but also emphasize its consistency and robustness across different datasets.
[0110] This embodiment also presents the comparison results under the traditional ZSL setting on the AWA2, CUB, and SUN datasets. It is worth emphasizing that the method of this scheme performs best on the SUN dataset, with an accuracy of 76.4%, representing an 8.8% improvement compared to the previous best method TransZero++. SUN is a scene dataset, and generative methods can usually adapt to more features, thus achieving better results on unseen classes. However, the PIAS model excellently understands the abstract content of the scene by mapping visual features to the semantic space to synthesize the attributes of each instance. The results significantly exceed both non-generative and generative methods, further confirming the effectiveness of the method of this scheme in synthesizing the attributes of each instance and demonstrating excellent generalization capabilities.
[0111] On the AWA2 dataset, PIAS outperforms all non-generative methods and most generative methods, being only 0.1% lower than the second-best result (76.4%). This indicates that the method of this scheme performs well on coarse-grained datasets. On the CUB dataset, the method of this scheme fails to achieve the best performance under the CZSL setting. Analysis shows that this may be due to the ultra-fine-grained characteristics of the CUB dataset, whose attribute feature dimension reaches 312, three times that of the fine-grained SUN dataset. For non-generative methods, minor changes in attribute features pose significant challenges and seriously affect the accuracy of unseen classes.
[0112] In summary, the PIAS model of this scheme effectively addresses the key technical challenges in zero-shot learning, namely domain shift and management of attribute dimensions, through its structure alignment module and optimized processing of fine-grained features, thereby significantly improving the accuracy and generalization capabilities of the learning model.
[0113] References (1) to (21):
[0114] (1) Chen, S.; Xie, G.; Liu, Y.; Peng, Q.; Sun, B.; Li, H.; You, X.; Shao, L. HSVA: Hierarchical Semantic-Visual Adaptation for Zero-Shot Learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc., 2021; Vol. 34, pp16622–16634.
[0115] (2) Chen, Z.; Luo, Y.; Qiu, R.; Wang, S.; Huang, Z.; Li, J.; Zhang, Z. Semantics Disentangling for Generalized Zero-Shot Learning. In 2021 IEEE / CVF International Conference on Computer Vision (ICCV); IEEE: Montreal, QC, Canada, 2021; pp8692–8700. https: / / doi.org / 10.1109 / ICCV48922.2021.00859.
[0116] (3) Chou, Y.-Y.; Lin, H.-T.; Liu, T.-L. ADAPTIVE AND GENERATIVE ZERO-SHOT LEARNING. 2021.
[0117] (4) Kim, J.; Shim, K.; Shim, B. Semantic Feature Extraction for Generalized Zero-Shot Learning. Proc. AAAI Conf. Artif. Intell. 2022, 36(1), 1166–1173. https: / / doi.org / 10.1609 / aaai.v36i1.20002.
[0118] (5) Liu, Y.; Gao, X.; Han, J.; Liu, L.; Shao, L. Zero-Shot Learning via a Specific Rank-Controlled Semantic Autoencoder. Pattern Recognit. 2022, 122, 108237. https: / / doi.org / 10.1016 / j.patcog.2021.108237.
[0119] (6) Kong, X.; Gao, Z.; Li, X.; Hong, M.; Liu, J.; Wang, C.; Xie, Y.; Qu, Y. En-Compactness: Self-Distillation Embedding & Contrastive Generation for Generalized Zero-Shot Learning. In 2022 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New Orleans, LA, USA, 2022; pp 9296–9305. https: / / doi.org / 10.1109 / CVPR52688.2022.00909.
[0120] (7) Hong, Z.; Chen, S.; Xie, G.-S.; Yang, W.; Zhao, J.; Shao, Y.; Peng, Q.; You, X. Semantic Compression Embedding for Generative Zero-Shot Learning. In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI-22; Raedt, L.D., Ed.; International Joint Conferences on Artificial Intelligence Organization, 2022; pp 956–963. https: / / doi.org / 10.24963 / ijcai.2022 / 134.
[0121] (8)Li, X.; Zhang, Y.; Bian, S.; Qu, Y.; Xie, Y.; Shi, Z.; Fan, J. VS-Boost: Boosting Visual-Semantic Association for Generalized Zero-Shot Learning. In Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence; International Joint Conferences on Artificial Intelligence Organization: Macau, SAR China, 2023; pp 1107–1115. https: / / doi.org / 10.24963 / ijcai.2023 / 123.
[0122] (9)Chen, S.; Chen, S.; Hou, W.; Ding, W.; You, X. EGANS: Evolutionary Generative Adversarial Network Search for Zero-Shot Learning. arXiv August 19, 2023. http: / / arxiv.org / abs / 2308.09915 (accessed 2023-12-05).
[0123] (10)Gao, Y.; Feng, W.; Xiao, R.; He, L.; He, Z.; Lv, J.; Tang, C. Improving Generalized Zero-Shot Learning via Cluster-Based Semantic Disentangling Representation. Pattern Recognit. 2024, 150, 110320. https: / / doi.org / 10.1016 / j.patcog.2024.110320.
[0124] (11)Alamri, F.; Dutta, A. Multi-Head Self-Attention via Vision Transformer for Zero-Shot Learning. arXiv July 30, 2021. http: / / arxiv.org / abs / 2108.00045 (accessed 2023-11-13).
[0125] (12) Chen, S.; Hong, Z.; Liu, Y.; Xie, G.-S.; Sun, B.; Li, H.; Peng, Q.; Lu, K.; You, X. TransZero: Attribute-Guided Transformer for Zero-Shot Learning. Proc. AAAI Conf. Artif. Intell. 2022, 36(1), 330–338. https: / / doi.org / 10.1609 / aaai.v36i1.19909.
[0126] (13) Chen, S.; Hong, Z.; Hou, W.; Xie, G.-S.; Song, Y.; Zhao, J.; You, X.; Yan, S.; Shao, L. TransZero++: Cross-Attribute-Guided Transformer for Zero-Shot Learning. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 1–17. https: / / doi.org / 10.1109 / TPAMI.2022.3229526.
[0127] (14) Chen, S.; Hong, Z.; Xie, G.-S.; Yang, W.; Peng, Q.; Wang, K.; Zhao, J.; You, X. MSDN: Mutually Semantic Distillation Network for Zero-Shot Learning. In 2022 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New Orleans, LA, USA, 2022; pp7602–7611. https: / / doi.org / 10.1109 / CVPR52688.2022.00746.
[0128] (15) Wang, Z.; Gou, Y.; Li, J.; Zhu, L.; Shen, H. T. Language-Augmented Pixel Embedding for Generalized Zero-Shot Learning. IEEE Trans. Circuits Syst. Video Technol. 2023, 33(3), 1019–1030. https: / / doi.org / 10.1109 / TCSVT.2022.3208256.
[0129] (16) Cheng, D.; Wang, G.; Wang, B.; Zhang, Q.; Han, J.; Zhang, D. Hybrid Routing Transformer for Zero-Shot Learning. Pattern Recognit. 2023, 137, 109270. https: / / doi.org / 10.1016 / j.patcog.2022.109270.
[0130] (17) Liu, M.; Li, F.; Zhang, C.; Wei, Y.; Bai, H.; Zhao, Y. Progressive Semantic-Visual Mutual Adaption for Generalized Zero-Shot Learning. In 2023 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Vancouver, BC, Canada, 2023; pp15337–15346. https: / / doi.org / 10.1109 / CVPR52729.2023.01472.
[0131] (18) Chen, Z.; Huang, Y.; Chen, J.; Geng, Y.; Zhang, W.; Fang, Y.; Pan, J. Z.; Chen, H. DUET: Cross-Modal Semantic Grounding for Contrastive Zero-Shot Learning. Proc. AAAI Conf. Artif. Intell. 2023, 37(1), 405–413. https: / / doi.org / 10.1609 / aaai.v37i1.25114.
[0132] (19) Alamri, F.; Dutta, A. Implicit and Explicit Attention Mechanisms for Zero-Shot Learning. Neurocomputing 2023, 534, 55–66. https: / / doi.org / 10.1016 / j.neucom.2023.03.009.
[0133] (20) Zhou, L.; Liu, Y.; Bai, X.; Li, N.; Yu, X.; Zhou, J.; Hancock, E. R. Attribute Subspaces for Zero-Shot Learning. Pattern Recognit. 2023, 144, 109869. https: / / doi.org / 10.1016 / j.patcog.2023.109869.
[0134] (21) Chen, S.; Hou, W.; Khan, S.; Khan, F. S. Progressive Semantic-Guided Vision Transformer for Zero-Shot Learning. arXiv April 11, 2024. https: / / doi.org / 10.48550 / arXiv.2404.07713.
Claims
1. A generalized zero-shot learning method based on instance image attribute generation and correction, characterized in that: The method is to train a PIAS model for image classification, wherein the PIAS model includes an attribute generation module and an attribute correction module, and specifically includes the following steps: S1. Generate the class average image of each class based on all instance images of each class in the visible class dataset; S2, constructing a visual space using the visual transformer ViT of the attribute generation module to extract the visual features of each instance image in the visible class dataset and the average visual features of the class average image; S3, according to the visual features of the instance image and the class average image, using the multi-layer perceptron of the attribute generation module to generate instance attributes of the instance image and class-level attributes of the class average image; S4. In the attribute correction module, classification loss is used to make each instance attribute consistent with its corresponding category label; S5, construct the semantic space using calibration loss to minimize the cosine similarity between the average attributes of the class and the annotation attributes; S6, using attribute annotations and average visual features of each class as anchors for the corresponding class to control the structural alignment of semantic space and visual space; S7, input the test data set into the attribute generation module, generate the attributes of each test image in the test data set, and input the attributes of each test image into the ZSL classifier to obtain the label of the test image; S8. Determine the classification accuracy according to the label of the test image, and judge whether the classification accuracy and the total loss function of the PIAS model meet the preset conditions. If so, end the algorithm, otherwise return to step S2.
2. The zero-shot learning method based on instance image attribute generation and correction according to claim 1, characterized in that: Step S6 further comprises: S61. Calculate the attribute fusion matrix based on the attribute annotation of the class and the instance attributes of all images of the corresponding class: in, is the attribute fusion matrix of the i-th category; a i Annotate the attributes of the i-th category; is the attribute set of all instance images of the i-th category; and are the 1st and Nth images in the i-th category respectively. i N instance attributes of instance images; i is the total number of instance images of the i-th category; T is the transpose; S62. Calculate the visual feature fusion matrix based on the average visual features of the class and the visual features of all images of the corresponding class: in, is the visual feature fusion matrix of the i-th category; is the average visual feature of the i-th category; and are the 1st and Nth in the i-th category respectively. i The visual features of the example images; H i is the set of visual features of all instance images of the i-th category; S63. Control the structural alignment of the semantic space and the visual space according to the topological structure similarity between the fusion matrix of the attributes and the fusion matrix of the visual features: in, is the structural consistency loss; ls is the total number of classes in the visible class dataset; for The value corresponding to the jth instance image in ; for The value corresponding to the j-th instance image in ; exp(.) is the exponential function.
3. The zero-shot learning method based on instance image attribute generation and correction according to claim 1, characterized in that: The expression of the calibration loss is: in, is the calibration loss; is the class-level attribute of the i-th class; a i is the attribute annotation of the i-th class; ls is the total number of classes in the visible class dataset; cos(.) is the cosine function.
4. The zero-shot learning method based on instance image attribute generation and correction according to claim 1, characterized in that: The expression of the classification loss is: in, is the classification loss; ls is the total number of classes in the visible class dataset; y i is the category label of the i-th category; exp(.) is the exponential function; Annotate the attributes of all classes in the visible class dataset a i A collection of; is the instance attribute of the jth instance image of the i-th category; N i is the total number of instance images of the i-th category; T is the transpose.
5. The zero-shot learning method based on instance image attribute generation and correction according to any one of claims 1 to 4, characterized in that: The expression of the total loss function of the PIAS model is: in, is the total loss; α, β, and δ are hyperparameters used to balance different loss terms; is the loss of structural consistency; is the calibration loss; is the classification loss.
6. The zero-shot learning method based on instance image attribute generation and correction according to claim 1, characterized in that: The expression for generating the class average image is: in, is the class average image of the i-th class in the visible class dataset; is the jth instance image in the i-th category; N i is the total number of instance images of the i-th category.
7. The zero-shot learning method based on instance image attribute generation and correction according to claim 1, characterized in that: The ZSL classifier is a CZSL classifier or a GZSL classifier; The expression of the CZSL classifier is: in, is the predicted label; is the attribute of the test image; y is the label; is the label of the unseen category; is the attribute of the unseen category; T is transposition; max is the maximum value; The expression of the GZSL classifier is: in, is the label of the visible category; are all attributes of the seen and unseen classes; γ is the calibration factor.
8. The zero-shot learning method based on instance image attribute generation and correction according to claim 1, characterized in that: The preset conditions are that the classification accuracy is less than the preset accuracy and the change in the value of the total loss function of the PIAS model during multiple consecutive trainings is less than the preset value.
Citation Information
Patent Citations
Alternating-current regulator.
US1024963A
Image zero-order classification model based on cross knowledge and classification method thereof
CN113191381A
Generalized zero sample learning method based on progressive semantics-vision interadaptation
CN116994027A