A zero-shot image recognition method, system and storage medium
Through attention mechanism and graph convolution neural network, information propagation and aggregation on visual prototype diagrams is solved, and the problem of inaccurate modeling of visual feature noise and category relationships in the prior art is achieved, and efficient zero-sample image recognition is achieved.
Patent Information
- Application Number
- CN202310303332.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-03-24
AI Technical Summary
Existing zero-sample learning methods are difficult to accurately model category relationships on fine-grained datasets, and visual features contain background information and noise, which affects knowledge transfer capabilities.
Discriminant visual features are extracted through attention mechanism, visual prototypes are constructed, and node information is propagated and aggregated using graph convolution neural networks to obtain discriminant potential spaces, and multiple semantic representations are fused to improve classification accuracy.
It improves the classification accuracy of zero-sample image recognition, reduces the annotation cost of supervised learning, and enhances the accuracy of knowledge transfer capabilities and category relationship modeling.
Smart Images

Figure CN116433969B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and particularly to a zero-shot image recognition method, system and storage medium. Background Art
[0002] With the development of deep neural networks, image classification has made great progress in the past few years. However, most successful models are based on supervised learning, which highly depends on a large number of labeled images for training. In many practical applications, collecting a large-scale labeled data set is expensive and time-consuming. When it comes to fine-grained data sets, this problem becomes even more serious. Therefore, zero-shot learning has received increasing attention, which can recognize images from unseen classes. Zero-shot learning aims to recognize unseen classes by applying the knowledge learned from seen classes. The classes of labeled samples given in the training phase are called seen classes, and there are also some unlabeled samples. The classes containing these unlabeled samples are called unseen classes, and the set of seen classes and the set of unseen classes are disjoint.
[0003] The invention patent application with the publication number of CN113505701A discloses a variational autoencoder zero-shot image recognition method combining a knowledge graph. This method first encodes the image features extracted by a convolutional neural network into low-dimensional feature vectors through a VAE and inputs them into the latent feature space; then it sends the class semantic vectors into a deep neural network module based on the knowledge graph, aggregates the nodes in the graph through a graph variational autoencoder, and inputs the newly generated low-dimensional semantic vectors after encoding update into the latent feature space; finally, for each latent vector generated by each modality, under the condition of the same class, it is decoded by the decoder of the other modality respectively to reconstruct the original data.
[0004] The method in the invention patent application with the publication number of CN113505701A uses the knowledge obtained from the knowledge graph to construct a graph. However, considering that the relationships between the classes contained in the knowledge graph are not accurate enough, it cannot well model the relationships between seen classes and unseen classes, thereby affecting the ability of knowledge transfer. And for some fine-grained data sets, it is difficult to obtain the class relationships. At the same time, although the visual features of images contain rich semantic information, they also contain a lot of background information and noise information. When performing image classification, only some discriminative visual regions are beneficial to classification. Especially for some fine-grained images, the differences between images of different classes are small. The method in the invention patent application with the publication number of CN113505701A only uses a pre-trained network to extract image features and does not fully mine the discriminative visual features. Summary of the Invention
[0005] To solve the above problems, the present invention aims to propose a zero-shot image recognition method, system and storage medium, which obtain visual prototype representations of all classes through an attention mechanism and semantic relationships between classes, and then obtain a discriminative latent space through visual prototype graph propagation, and perform classification in the latent space to improve the classification accuracy.
[0006] To achieve the above object, the technical solution of the present invention is implemented as follows:
[0007] A zero-shot image recognition method includes the following steps:
[0008] S1. Obtain a data set including visible classes and invisible classes, where the visible classes are the classes containing images in the training set, and have images, class labels and semantic attributes of the visible classes, and the invisible classes are the classes not containing images in the training set, and have semantic attributes of the invisible classes. The invisible class images are used in the prediction and recognition stage;
[0009] S2. Design an attention mechanism to extract discriminative visual features of visible class images;
[0010] S3. Perform a mean operation on all images belonging to the same visible class to obtain the visual prototype of the visible class;
[0011] S4. Obtain the visual prototype of the invisible class by transferring the semantic attribute relationship between the visible class and the invisible class;
[0012] S5. Use the relationship between class visual prototypes to construct a visual prototype graph and initialize the node representation;
[0013] S6. Design an encoder to perform node information propagation and aggregation to obtain a new latent space;
[0014] S7. Use visible class images and labels to train the model;
[0015] S8. Use the trained model to predict invisible class images.
[0016] Furthermore, in step S2, a discriminative feature v of each image x is obtained by using an attention mechanism to remove irrelevant information;
[0017] Specifically, for an image x, it is first sent to the backbone network to obtain its feature map Z ∈ R W×H×C , and then the feature map Z passes through a spatial attention mechanism to obtain K regional blocks, expecting to discover the most discriminative feature regions in the image;
[0018] The specific operation is as follows:
[0019] First, K mask blocks M k ∈ RW×H :
[0020] M k = σ(Conv(Z)), k = 1, 2, ..., K
[0021] Where Conv(.) represents a 1×1 convolution operation, and σ(.) represents an activation function;
[0022] After that, first perform a Reshape operation on the K mask blocks to make them the same size as the feature map Z, and then multiply them with the feature map Z to obtain K regional blocks of the image:
[0023]
[0024] Where, represents element-wise multiplication. Considering that the K regional blocks obtained may contain background information or there may be redundant information between different regional blocks, therefore, a threshold limit is applied to the K regional blocks to obtain the maximum value m of the K mask blocks max :
[0025]
[0026] Again, design a hyperparameter α and set a threshold τ = α×m max , where α is a value between 0 and 1. When the maximum value of the k-th mask block M k is less than this threshold τ, the k-th regional block R k (Z) is set to 0, and then global max pooling is performed on the K thresholded regional blocks to obtain K regional features r k ∈R C ;
[0027] Finally, concatenate the K local regional features, and then pass through a fully connected layer f1 with an input-output dimension of KC - C to obtain the discriminative visual feature v ∈ R C .
[0028] Furthermore, in step S3, after obtaining the discriminative visual feature v of each image as described above, perform a mean operation on all images belonging to the same visible class to obtain the visual prototype P of this visible class seen :
[0029]
[0030] Where n represents the number of samples belonging to the i-th visible class.
[0031] Further, in step S4, the relationship matrix between the visible classes and the invisible classes is obtained by calculating the cosine similarity between the attribute vectors of the visible classes and the invisible classes
[0032] S ij =cos(z i , z j ), i ∈ C s and j ∈ C u
[0033] where z i represents the attribute vector of the i-th class, z j represents the attribute vector of the j-th class, C s represents the number of visible classes, C u represents the number of invisible classes;
[0034] The visual prototype of the invisible classes is obtained by transferring the semantic relationship between the visible classes and the invisible classes; that is, the visual prototype matrix P of the invisible classes is obtained according to the semantic relationship matrix between the visible classes and the invisible classes unseen :
[0035] P unseen =S T P seen .
[0036] Further, in step S5, the visual prototypes of all classes can be obtained through the above operations. By using the relationship between the class visual prototypes, a visual prototype graph G is constructed, where each node represents a class, including visible classes and invisible classes;
[0037] Specifically, the relationship of the edges in the visual prototype graph G is measured by the cosine distance between the class visual prototypes:
[0038] B ij =cos(P i , P j )
[0039] Use the GloVe model to obtain the word vector representation a i ∈ R 300 of the attributes of each class, stack the word vector representations of all the attributes of this class into a matrix T ∈ R |A|×300 (|A| represents the number of class attributes), and then multiply it by the semantic vector of this class to obtain the initial representation of each node:
[0040] E i =z i T
[0041] where z i represents the semantic vector of the i-th class, E iThe initialization vector representing the i-th node.
[0042] Furthermore, in step S6, the obtained adjacency matrix B and the matrix E formed by stacking the initialization representations of all nodes are input into the encoder for node information propagation and aggregation;
[0043] Specifically, the encoder uses a graph convolutional neural network to perform information aggregation and propagation on the constructed visual prototype graph G to obtain a new latent space U, and the updated representation of the nodes is as follows:
[0044] H (i) = σ(D -1 / 2 BD -1 / 2 H (i-1) W (i-1) ), i = 1, 2
[0045] where D represents the degree matrix of the adjacency matrix B, D ii = ∑ j B ij , W (i) represents the parameter matrix, H (0) = E represents the initialization input matrix of all nodes, H (2) = U represents the output matrix of the second-layer graph convolution, that is, the matrix composed of the updated embedding vectors of all nodes on the visual prototype graph where u i ∈ R d represents the embedding representation of node i;
[0046] The obtained latent representation matrix U is input into the decoder, and the decoder reconstructs the adjacency matrix through the inner product of the embedding vectors:
[0047]
[0048] Construct a reconstruction loss to minimize the difference between the adjacency matrix reconstructed from the node vector representation and the original adjacency matrix, so that the embedding vector representation of the nodes conforms to the structure of the graph; the reconstruction loss L rec is designed as follows:
[0049]
[0050] Furthermore, in step S7, after obtaining the latent space U, the visual features and labels extracted from the visible class images through the attention mechanism are embedded into the latent space for classification in the latent space;
[0051] Specifically, the visual features v ∈ R of the visible class images CThrough a fully connected layer f2, with an input-output dimension of C-d, it is embedded into the latent space, and the similarity between the mapped visual features and the latent representation of each class is calculated:
[0052] Q ij = f2(v i ) ⊙ U j
[0053] Meanwhile, the cross-entropy loss function is used to construct the classification loss:
[0054]
[0055] where, when the i-th sample belongs to the k-th class, y ik = 1, otherwise, y ik = 0; N s represents the number of visible class samples;
[0056] The above reconstruction loss and classification loss are added to obtain the overall loss function, and the model parameters are optimized through gradient backpropagation;
[0057] Finally, the total loss function of the model is:
[0058] L = L cls + γL rec
[0059] where γ is a hyperparameter.
[0060] Furthermore, in step S8, through the above-trained model, the latent representations of all classes can be obtained; when a test image x, i.e., an unseen class image, is given, the visual features v of this image are obtained through the trained attention mechanism, then the visual features v are mapped to the latent space, and the similarity is calculated with the class latent representations, and finally the process of predicting its label is as follows:
[0061]
[0062] To achieve the above object, the present invention also provides a zero-shot image recognition system, including the following modules:
[0063] Dataset acquisition and definition module: used to acquire a dataset including visible classes and unseen classes, where the visible classes are the classes containing images in the training set, having images of visible classes, class labels, and semantic attributes of visible classes, and the unseen classes are the classes not containing images in the training set, having semantic attributes of unseen classes, and the unseen class images are used in the prediction and recognition stage;
[0064] Discriminative visual feature extraction module: extracts the discriminative visual features of visible class images through the attention mechanism;
[0065] Visual prototype extraction module for visible classes: Perform a mean operation on all images belonging to the same visible class to obtain the visual prototype of the visible class;
[0066] Visual prototype extraction module for invisible classes: Obtain the visual prototype of the invisible class by transferring the semantic attribute relationship between the visible class and the invisible class;
[0067] Node initialization module for the visual prototype graph: Construct a visual prototype graph using the relationship between class visual prototypes and initialize the node representation;
[0068] Latent space acquisition module: Design an encoder to propagate and aggregate node information to obtain a new latent space;
[0069] Training module: Train the model using visible class images and labels;
[0070] Invisible class image classification module: Use the trained model to predict invisible class images.
[0071] To achieve the above object, the present invention also provides a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to execute the steps of the above zero-shot image recognition method.
[0072] Beneficial effects: The present invention obtains the visual prototype representations of all classes through the attention mechanism and the semantic relationship between classes, and then obtains the discriminative latent space through the propagation of the visual prototype graph, and performs classification in the latent space, improving the classification accuracy. Description of the Drawings
[0073] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0074] Figure 1 Is the flowchart of the zero-shot image recognition method described in the embodiment of the present invention;
[0075] Figure 2 Is the framework diagram of the training stage of the zero-shot image recognition method described in the embodiment of the present invention;
[0076] Figure 3 Is the framework diagram of the prediction and recognition stage of the zero-shot image recognition method described in the embodiment of the present invention;
[0077] Figure 4 Is the structural schematic diagram of the zero-shot image recognition system described in the embodiment of the present invention. Detailed Embodiments
[0078] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0079] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0080] The goal of zero-shot learning is to classify images that were not seen during the training phase. It establishes relationships between different classes through auxiliary semantic information, thereby achieving knowledge transfer from visible classes to invisible classes. When using semantic information to transfer knowledge from visible classes to invisible classes in zero-shot learning, one goal is to establish an association between the visual domain and the semantic domain. Generally, this association is determined by learning an embedding space where semantic vectors and visual features interact. There are three mapping methods for learning this embedding space, including semantic space embedding-based, visual space embedding-based, and common space embedding-based. The invention of CN 113505701A belongs to the zero-shot learning method based on common space embedding.
[0081] The invention of CN113505701A relatively has the following technical defects. First, since visual features and semantic representations are distributed in different spaces and have a large dimensionality difference, embedding into any one space will result in information loss. Second, only through the compatibility function for measurement, it cannot well enable visual features and semantic representations to interact. In addition, the visual features of images contain rich semantic information but also contain a lot of class-irrelevant information, and some relatively similar images may be misclassified because of this class-irrelevant information. Finally, although the artificially defined class attributes are precise, due to the limitation of their dimensions, some important information may be ignored, thereby reducing the ability of knowledge transfer.
[0082] Embodiment 1
[0083] Based on the above theoretical research and analysis of the prior art, see Figures 1-3 : This embodiment provides a zero-shot image recognition method, including the following steps:
[0084] S1. Obtain a data set including visible classes and invisible classes, where the visible classes are the classes containing images in the training set, having images with visible classes, class labels, and semantic attributes of the visible classes, and the invisible classes are the classes not containing images in the training set, having semantic attributes of the invisible classes, and the invisible class images are used in the prediction and recognition phase;
[0085] It should be noted that the publicly available datasets used in the model of this embodiment may include: the fine-grained bird species dataset CUB-200-2011 Birds (CUB), the Animals with Attributes 2 (AWA2) dataset, the SUN Attribute (SUN) scene dataset, and the Pascal and Yahoo (aPY) dataset;
[0086] Classify the above-mentioned publicly available datasets; divide all classes of each dataset into non-overlapping visible classes and invisible classes, and obtain the corresponding images and class semantic attributes respectively; the images and class semantic attributes of the visible classes and the semantic attributes of the invisible classes are used in the model training stage, while the images of the invisible classes are used for testing in the prediction and recognition stage;
[0087] Among them, the CUB dataset has 200 classes, including 150 visible classes and 50 invisible classes, with a total of 11,788 images, and each class has 312-dimensional semantic attributes; the AWA2 dataset has 50 classes, including 40 visible classes and 10 invisible classes, with a total of 37,322 images, and each class has 85-dimensional semantic attributes; the SUN dataset has 717 classes, including 645 visible classes and 72 invisible classes, with a total of 14,340 images, and each class has 102-dimensional semantic attributes; the aPY dataset has 32 classes, including 20 visible classes and 12 invisible classes, with a total of 15,339 images, and each class has 64-dimensional semantic attributes.
[0088] S2. Design an attention mechanism to extract discriminative visual features of visible class images;
[0089] S3. Perform a mean operation on all images belonging to the same visible class to obtain the visual prototype of this visible class;
[0090] S4. Obtain the visual prototype of the invisible class by transferring the semantic attribute relationship between the visible class and the invisible class;
[0091] S5. Use the relationship between class visual prototypes to construct a visual prototype graph and initialize the node representation;
[0092] S6. Design an encoder to propagate and aggregate node information to obtain a new latent space;
[0093] S7. Train the model using visible class images and labels;
[0094] S8. Use the trained model to predict invisible class images.
[0095] In this embodiment, visual prototype representations of all classes are obtained through an attention mechanism and the semantic relationships between classes, and then a discriminative latent space is obtained through the propagation of the visual prototype graph. Classification is performed in the latent space, improving the classification accuracy.
[0096] The zero-shot image recognition method of this embodiment can meet the image recognition requirements of multiple unseen classes, reduce the human and material resources consumed by image annotation under supervised learning, improve the task performance of recognizing unseen class images, and accelerate the research and application of zero-shot classification in actual scenarios.
[0097] Different from the method in the invention patent application with the publication number CN113505701A that only uses a pre-trained network to extract image features, the method proposed in this embodiment obtains discriminative visual features of visible class images through an attention mechanism to remove some class-irrelevant information such as background, making the visual representation more discriminative. At the same time, different from using a knowledge graph for graph construction in the invention patent application with the publication number CN113505701A, this embodiment constructs a visual prototype graph by using the relationships between visual prototypes of all classes (including visible classes and unseen classes), making the constructed graph structure easier to obtain and the relationships between classes more accurate, improving the knowledge transfer ability. In addition, different from initializing nodes with a single semantic representation in the invention patent application with the publication number CN113505701A, this invention combines multiple semantic representations, improving the semantic representation of classes.
[0098] In a specific example, in step S2, a discriminative feature v of each image x is obtained by using an attention mechanism to remove irrelevant information;
[0099] Specifically, for an image x, it is first sent to the backbone network to obtain its feature map Z ∈ R W×H×C , and then K regional blocks are obtained from the feature map Z through a spatial attention mechanism, expecting to discover the most discriminative feature regions in the image;
[0100] The specific operation is as follows:
[0101] First, K mask blocks M k ∈ R W×H are learned through convolution operations:
[0102] M k = σ(Conv(Z)), k = 1, 2,..., K
[0103] where Conv(.) represents a 1×1 convolution operation, and σ(.) represents an activation function;
[0104] After that, first perform a Reshape operation on the K masked blocks to make them the same size as the feature map Z, and then multiply them with the feature map Z to obtain K regional blocks of the image:
[0105]
[0106] Among them, represents element-wise multiplication. Considering that the K regional blocks obtained may contain background information or there may be redundant information between different regional blocks, therefore, perform a threshold limit on the K regional blocks to obtain the maximum value m of the K masked blocks max :
[0107]
[0108] Again, design a hyperparameter α and set a threshold τ = α × m max , where α is a value between 0 and 1. When the maximum value of the k-th masked block M k is less than this threshold τ, set the k-th regional block R k (Z) to all 0, and then perform global max pooling on the K thresholded regional blocks to obtain K regional features r k ∈R C ;
[0109] Finally, concatenate the K local regional features, and then pass through a fully connected layer f1 with an input-output dimension of KC-C to obtain the discriminative visual feature v ∈ R of the image C .
[0110] This embodiment proposes to obtain the visual prototype of the visible class by using the discriminative visual feature of the visible class image, making the obtained visible class visual prototype more discriminative.
[0111] In a specific example, in step S3, after obtaining the discriminative visual feature v of each image above, perform an averaging operation on all images belonging to the same visible class to obtain the visual prototype P of the visible class seen :
[0112]
[0113] Among them, n represents the number of samples belonging to the i-th visible class.
[0114] In a specific example, in step S4, obtain the relationship matrix between the visible class and the invisible class by calculating the cosine similarity between the attribute vectors of the visible class and the invisible class
[0115] S ij= cos(zi, z j ), i ∈ C s and j ∈ C u
[0116] where z i represents the attribute vector of the i-th class, and z j represents the attribute vector of the j-th class, and C s represents the number of visible classes, and C u represents the number of invisible classes;
[0117] Since our method is under an inductive setting (i.e., only using the labeled visible class data during the training process), the visual prototypes of the invisible classes cannot be obtained through the above steps. Therefore, we obtain the visual prototypes of the invisible classes by transferring the semantic relationship between the visible classes and the invisible classes; that is, we obtain the visual prototype matrix P of the invisible classes according to the semantic relationship matrix between the visible classes and the invisible classes unseen :
[0118] P unseen = S T P seen .
[0119] In a specific example, in step S5, the visual prototypes of all classes can be obtained through the above operations. By using the relationship between the class visual prototypes, a visual prototype graph G is constructed, where each node represents a class, including visible classes and invisible classes;
[0120] Specifically, the relationship of the edges in the visual prototype graph G is measured by the cosine distance between the class visual prototypes:
[0121] B ij = cos(P i , P j )
[0122] Use the GloVe model to obtain the word vector representation a i ∈ R 300 of the attributes of each class, stack the word vector representations of all the attributes of this class into a matrix T ∈ R |A|×300 (|A| represents the number of class attributes), and then multiply it by the semantic vector of this class to obtain the initial representation of each node:
[0123] E i = z i T
[0124] where z i represents the semantic vector of the i-th class, and E i represents the initial vector of the i-th node.
[0125] It should be noted that, in order to better initialize the node representation, the class attribute representation and the word vector representation of each attribute are fused to fully explore and improve the semantic representation of the class, which is helpful for subsequent information dissemination;
[0126] In this embodiment, the semantic relationship between visible classes and invisible classes is used to obtain the visual prototypes of invisible classes, and a visual prototype graph is constructed using the relationships between the visual prototypes of all classes, making the class relationship modeling more accurate.
[0127] In a specific example, in step S6, the obtained adjacency matrix B and the matrix E stacked by the initialization representations of all nodes are input into the encoder for node information propagation and aggregation;
[0128] Specifically, the encoder uses the graph convolutional neural network to perform information aggregation and propagation on the constructed visual prototype graph G to obtain a new latent space U, and the updated representation of the node is as follows:
[0129] H (i) =σ(D -1 / 2 BD -1 / 2 H (i-1) W (i-1) ),i=1,2
[0130] where D represents the degree matrix of the adjacency matrix B, D ii =∑ j B ij ,W (i) represents the parameter matrix, H (0) =E represents the initialization input matrix of all nodes, H( 2 )=U represents the output matrix of the second-layer graph convolution, that is, the matrix composed of the updated embedding vectors of all nodes on the visual prototype graph where u i ∈R d represents the embedding representation of node i;
[0131] The obtained latent representation matrix U is input into the decoder, and the decoder reconstructs the adjacency matrix through the inner product of the embedding vectors:
[0132]
[0133] Construct a reconstruction loss to minimize the difference between the adjacency matrix reconstructed from the node vector representation and the original adjacency matrix, so that the embedding vector representation of the node conforms to the structure of the graph; the reconstruction loss L rec is designed as follows:
[0134]
[0135] It can be understood that in this embodiment, the encoder implemented by the graph convolutional neural network is used to aggregate and propagate information of the class semantic representation on the visual prototype graph to obtain a latent space, which better promotes the information interaction between the semantic representation and the visual representation; at the same time, the decoder reconstructs the adjacency matrix of the visual prototype graph by using the latent structured representation, so that the embedding vector representation of the nodes conforms to the structure of the graph, improving the discriminability of the latent space.
[0136] In a specific example, in step S7, after obtaining the latent space U, the visual features and labels extracted from the visible class images through the attention mechanism are embedded into the latent space for classification in the latent space;
[0137] Specifically, the visual features v of the visible class images ∈ R C are embedded into the latent space through a fully connected layer f2 with an input-output dimension of C-d, and the similarity between the mapped visual features and the latent representations of each class is calculated:
[0138] Q ij = f2(v i ) ⊙ U j
[0139] At the same time, the cross-entropy loss function is used to construct the classification loss:
[0140]
[0141] where, when the i-th sample belongs to the k-th class, y ik = 1, otherwise, y ik = 0; N s represents the number of visible class samples;
[0142] The above reconstruction loss and classification loss are added to obtain the overall loss function, and the model parameters are optimized through gradient backpropagation;
[0143] Finally, the total loss function of the model is:
[0144] L = L cls + γL rec
[0145] where γ is a hyperparameter.
[0146] In a specific example, in step S8, through the above trained model, the latent representations of all classes can be obtained; when a test image x, that is, an invisible class image, is given, the visual features v of the image are obtained through the trained attention mechanism, and then the visual features v are mapped to the latent space and the similarity calculation is performed with the class latent representations. Finally, the process of predicting its label is as follows:
[0147]
[0148] In summary, compared with the existing zero-shot learning methods based on common space embedding, in this embodiment, by mining discriminative visual features and removing some class-irrelevant information, the obtained class visual prototypes are more discriminative; during the composition process, without relying on any external information, the relationships between classes can be accurately modeled; a latent space is obtained through the structure of the autoencoder, and the graph convolutional neural network is used to propagate and aggregate information of nodes that fuse multiple semantic representations on the visual prototype graph, promoting information interaction between different modalities and improving the knowledge transfer ability.
[0149] In the zero-shot image recognition experiment, the experimental results are shown in Table 1. In Table 1, the bold values are the optimal values in each column, and "-" indicates that the original text did not conduct experiments on this dataset.
[0150]
[0151] GAFE and ACMR rely on the structure of the autoencoder. GAFE uses the encoder to learn the mapping from visual features to the semantic space, and the decoder reconstructs the original features using the learned mapping. ACMR uses two parallel variational autoencoders to separately extract visual latent representations and semantic latent representations, and at the same time proposes an information enhancement module to enhance the recognition ability of the latent variables. The method we proposed learns the latent representation through an encoder based on the graph convolutional neural network, aggregates and propagates class semantic representations on the visual prototype graph, and uses the non-redundant and complementary information provided between multiple modalities to learn the structured semantic embeddings obtained from different spaces to obtain a discriminative latent space. At the same time, the decoder part uses the learned latent embedding representation to reconstruct the adjacency matrix of the graph, making the updated latent embedding representation conform to the structure of the original graph. It can be seen from the table that compared with these methods, the method we proposed has a significant improvement. For example, compared with the ACMR method, the accuracy of the method we proposed on the CUB, AWA2, and SUN datasets has increased by 18.5%, 3.8%, and 19.8% respectively.
[0152] APNet, HGKT, and KG-VAE utilize graph structures to establish inter-class relationships. APNet generates edges based on similarity measurement of node feature representations (attribute vectors) and uses an attention mechanism for graph propagation. HGKT first models the relationships between visible classes according to the representative nodes of categories under the k-nearest neighbor scheme. After graph propagation, in the visual feature space, the invisible classes are connected to the k nearest visible classes to obtain the embedding representations of the invisible classes. KG-VAE feeds the category semantic vectors into the deep neural network module based on the knowledge graph, and encodes and generates new semantic vectors after aggregating and updating the nodes in the graph through the graph variational autoencoder. The method we proposed models the relationships between the visual prototypes of categories and uses the semantic representation that fuses the attribute word vectors and class attributes as node features. After graph propagation, it can effectively fuse various modal information and promote information interaction in different spaces. At the same time, using the reconstructed adjacency matrix of the latent feature graph as a constraint makes the embedded vector representations of the nodes conform to the graph structure. The experimental results show that the method we proposed helps to improve the classification accuracy. For example, compared with the KG-VAE method in the invention patent application with the publication number CN113505701A, our method improves the accuracy by 13.0%, 8.3%, and 1.5% on the CUB, AWA2, and SUN datasets respectively.
[0153] Finally, compared with some feature generation-based methods, such as f-CLSWGAN and LisGAN, the method we proposed also achieves a large performance improvement.
[0154] Embodiment 2
[0155] To achieve the above object, refer to Figure 4 : This embodiment also provides a zero-shot image recognition system, including the following modules:
[0156] Dataset acquisition and definition module: used to acquire a dataset including visible classes and invisible classes, where visible classes are the classes containing images in the training set, with images, class labels, and semantic attributes of visible classes, and invisible classes are the classes not containing images in the training set, with semantic attributes of invisible classes. Invisible class images are used in the prediction and recognition stage;
[0157] Discriminative visual feature extraction module: extracts the discriminative visual features of visible class images through the attention mechanism;
[0158] Visual prototype extraction module of visible classes: performs a mean operation on all images belonging to the same visible class to obtain the visual prototype of this visible class;
[0159] Visual prototype extraction module for invisible classes: Obtain the visual prototypes of invisible classes by transferring the semantic attribute relationships between visible and invisible classes;
[0160] Node initialization module for visual prototype graphs: Construct a visual prototype graph using the relationships between class visual prototypes and initialize the node representations;
[0161] Latent space acquisition module: Design an encoder to propagate and aggregate node information to obtain a new latent space;
[0162] Training module: Train the model using visible class images and labels;
[0163] Invisible class image classification module: Use the trained model to predict invisible class images.
[0164] The zero-shot image recognition system of this embodiment has the same advantages as the above zero-shot image recognition method compared to the prior art, which will not be elaborated here.
[0165] Embodiment 3
[0166] To achieve the above objective, this embodiment also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to execute the steps of the above zero-shot image recognition method.
[0167] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the above method embodiments. Among them, any reference to memory, storage, database, or other media provided in the various embodiments of the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0168] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A zero-shot image recognition method, characterized in that It includes the following steps: S1. Obtain a data set including visible classes and invisible classes. Among them, the visible classes are the classes containing images in the training set, with images, class labels and semantic attributes of the visible classes. The invisible classes are the classes without images in the training set, with semantic attributes of the invisible classes. The invisible class images are used in the prediction and recognition stage; S2. Design an attention mechanism to extract discriminative visual features of the visible class images; S3. Perform a mean operation on all images belonging to the same visible class to obtain the visual prototype of the visible class; S4. Obtain the visual prototype of the invisible class by transferring the semantic attribute relationship between the visible class and the invisible class; S5. Construct a visual prototype graph using the relationships between the class visual prototypes and initialize the node representation. In step S5, all categories of visual prototypes can be obtained through the above operations. By using the relationships between the class visual prototypes, a visual prototype graph is constructed. , where each node represents a category, including visible and invisible classes. Specifically, the visual prototype diagram The relationship between the sides in it is measured by the cosine distance between the class visual prototypes: Use the GloVe model to obtain the word vector representation of the attributes of each class , stack the word vector representations of all attributes of this class into a matrix represents the number of class attributes, and then multiply it by the semantic vector of this class to obtain the initial representation of each node: Among them, represents the semantic vector of the th class, represents the initialization vector of the th node; S6. Design an encoder to propagate and aggregate node information to obtain a new latent space; S7. Use the visible class images and labels to train the model; S8. Use the trained model to predict the invisible class images.
2. The zero-shot image recognition method according to claim 1, wherein In step S2, a discriminative feature of each image is obtained by using an attention mechanism to remove irrelevant information; Specifically, for an image , it is first fed into the backbone network to obtain its feature map . Then, for the feature map , it passes through a spatial attention mechanism to obtain regional blocks, with the expectation of discovering the most discriminative feature regions in the image; The specific operation is as follows: First, learn and obtain through convolution operations mask blocks : Among them, represents a convolution operation, and represents an activation function; After that, first perform a Reshape operation on the mask blocks to make them the same size as the feature map , and then multiply them with the feature map to obtain region blocks of the image: Among them, represents the multiplication of corresponding bit elements. Considering that the region blocks obtained may contain background information or there may be redundant information between different region blocks. Therefore, a threshold limit is imposed on the region blocks to obtain the maximum value of the mask blocks: Next, design a hyperparameter , and set a threshold , where is a value between 0 and 1. When the maximum value of the th mask block is less than this threshold , the th region block will be set to 0 entirely. Then, perform global max pooling on the thresholded region blocks to obtain region features ; Finally, splice the local region features, and then pass through a fully connected layer with an input-output dimension of to obtain the discriminative visual features corresponding to the image .
3. The zero-shot image recognition method according to claim 2, wherein In step S3, after obtaining the discriminative visual features of each of the above images a mean operation is performed on all images belonging to the same visible class to obtain the visual prototype of the visible class : Among them, represents the number of samples belonging to the th visible class.
4. The zero-shot image recognition method according to claim 3, wherein In step S4, the relationship matrix between the visible class and the invisible class is obtained by calculating the cosine similarity between the attribute vectors of the visible class and the invisible class : Among them, represents the attribute vector of the th class, represents the attribute vector of the th class, represents the number of visible classes, represents the number of invisible classes; Obtain the visual prototype of the invisible class by migrating the semantic relationship between the visible class and the invisible class; that is, obtain the visual prototype matrix of the invisible class according to the semantic relationship matrix between the visible class and the invisible class : 。 5. The zero-shot image recognition method according to claim 1, characterized in that, In step S6, the adjacency matrix obtained above and the matrix stacked by the initial representations of all nodes are input into the encoder for the propagation and aggregation of node information; Specifically, the encoder uses a graph convolutional neural network to perform information aggregation and propagation on the constructed visual prototype graph to obtain a new latent space , and the updated representation of the node is as follows: Among them, represents the adjacency matrix of the degree matrix, , represents the parameter matrix, represents the initial input matrix of all nodes, represents the output matrix of the second-layer graph convolution, that is, the matrix composed of the updated embedding vectors of all nodes on the visual prototype graph , where represents the node 's embedding representation; Input the obtained latent representation matrix above into the decoder, and the decoder reconstructs the adjacency matrix through the inner product of the embedding vectors: Construct a reconstruction loss to minimize the difference between the reconstructed adjacency matrix from the node vector representation and the original adjacency matrix, so that the embedded vector representation of the nodes conforms to the structure of the graph; the reconstruction loss is designed as follows: is designed as follows: 。 6. The zero-shot image recognition method according to claim 5, characterized in that In step S7, after obtaining the latent space the visual features and labels extracted from the visible class images through the attention mechanism are embedded into the latent space for classification in the latent space; Specifically, the visual features of the visible class images are passed through a fully connected layer with an input-output dimension of to be embedded into the latent space, and the similarity between the mapped visual features and the latent representations of each class is calculated: At the same time, use the cross-entropy loss function to construct the classification loss: Among them, when the th sample belongs to the th category, , otherwise, ; represents the number of visible class samples; Add the above reconstruction loss and classification loss to obtain the overall loss function, and optimize the model parameters through gradient backpropagation; Finally, the total loss function of the model is: Among them, is a hyperparameter.
7. The zero-shot image recognition method according to claim 6, wherein In step S8, the latent representations for all classes can be obtained through the above-trained model; when a test image , that is, an unseen-class image, is given, the visual features of this image are obtained through the trained attention mechanism , and then the visual features are mapped to the latent space, the similarity is calculated with the class latent representations, and finally the process of predicting its label is as follows: 。 8. A system using the zero-shot image recognition method according to any one of claims 1 to 7, characterized in that, It includes the following modules: Data set acquisition and definition module: used to obtain a data set including visible classes and invisible classes. Among them, the visible classes are the classes containing images in the training set, with images, class labels and semantic attributes of the visible classes. The invisible classes are the classes without images in the training set, with semantic attributes of the invisible classes. The invisible class images are used in the prediction and recognition stage; Discriminative visual feature extraction module: extract discriminative visual features of the visible class images through the attention mechanism; Visual prototype extraction module of visible classes: perform a mean operation on all images belonging to the same visible class to obtain the visual prototype of the visible class; Visual prototype extraction module of invisible classes: obtain the visual prototype of the invisible class by transferring the semantic attribute relationship between the visible class and the invisible class; Node initialization module of the visual prototype graph: construct a visual prototype graph using the relationship between class visual prototypes and initialize the node representation; Latent space acquisition module: obtain a new latent space by designing an encoder to propagate and aggregate node information; Training module: use the visible class images and labels to train the model; Invisible class image classification module: use the trained model to predict the invisible class images.
9. A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the processor is caused to execute the steps of the zero-shot image recognition method according to any one of claims 1 to 7.
Citation Information
Patent Citations
A knowledge graph-combined variational auto-encoder zero sample image recognition method
CN113505701A
Fine-grained zero-sample classification method based on multi-layer semantic supervised attention model
CN109447115A
Model training method and device for image classification and storage medium
CN114170475A