Zero-sample image classification method, system, device and medium for supplementing missing features

Through the generation of adversarial network and attribute grouping technology, the generation of unseen image features that conform to the actual distribution is solved, and the problem of mismatch in the generated feature distribution in the existing technology is improved, and the accuracy of zero-sample image classification is improved.

CN115761366BActive Publication Date: 2025-06-06YANGTZE DELTA REGION INST OF UNIV OF ELECTRONICS SCI & TECH OF CHINE (HUZHOU)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211505669.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-28
Publication Date
2025-06-06
Estimated Expiration
2042-11-28

AI Technical Summary

Technical Problem

The existing generative zero-sample image classification method cannot generate image features that are missing certain attributes, resulting in the generated image feature distribution that does not conform to the real unseen image distribution, and the classification accuracy is low.

Method used

By collecting zero-sample image classification datasets, semantic features of all categories are obtained, feature extraction is performed, and non-seen image features are forged based on the generative adversarial network. Use Word2vector and K-means algorithms to group and cluster attributes to generate fake image features that lack certain attributes.

Benefits of technology

The generated features of unseen images are more in line with the actual distribution, helping the classification model learn more complete information and improving the classification accuracy rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115761366B_ABST
    Figure CN115761366B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of computer vision technology, and discloses a zero-sample image classification method, system, device and medium for supplementing missing features, collects a zero-sample image classification data set, and simultaneously obtains semantic features of all categories; extracts features from images; trains a generative adversarial network based on features; uses the generative adversarial network to extract forged unseen image features, combines the forged unseen image features with image feature vectors to obtain an image training data set; trains an image feature classification network model based on the image training data set, and tests the data in the test set. The method disclosed in the present invention belongs to a generative zero-sample image classification method, and optimizes the situation in which image features with missing certain attributes cannot be generated in the existing method, so that the generated unseen image features are more in line with the actual distribution, helping the classification model to learn more complete information, and ultimately improving the classification accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision technology, and in particular relates to a zero-sample image classification method, system, device and medium for supplementing missing features. Background Art

[0002] At present, most of the existing image classification models are built on the premise that all categories of data are known. When the model trained under such data encounters category images that do not exist in the training data, it cannot recognize them. If these new categories need to be recognized, it is necessary to collect new category image data again and add them to the original data set, and then retrain the model to enable the model to recognize new categories. If new categories are encountered again, the above cycle must be repeated. The zero-shot image classification method uses auxiliary information to transfer the information learned by the model from the visible class images during the training phase to the classification of unseen class images through auxiliary information.

[0003] Zero-shot image classification methods can be roughly divided into two categories, one is the discriminative zero-shot image classification method, and the other is the generative zero-shot image classification method. The former mainly allows the model to learn the mapping function from image features to semantic features, and then in the test phase, the test image is mapped to the semantic space, and the image category is obtained by similarity comparison. This can be considered to be a solution to the zero-shot problem based on metric learning. The latter is to learn the mapping function from semantic features to image features, and use the learned mapping function to generate fake unseen image features using the semantic features of unseen classes, thereby solving the zero-shot problem of unseen classes. Then, the complete data is used to train ordinary image classification methods, which can be considered to be a solution to the zero-shot problem by generating data.

[0004] Through the above analysis, the problems and defects of the prior art are as follows:

[0005] Existing generative zero-shot image classification methods cannot generate image features that are missing certain attributes, and the distribution of generated image features does not conform to the actual distribution of unseen class images, resulting in a low classification accuracy of unseen classes. Summary of the invention

[0006] In view of the problems existing in the prior art, the present invention provides a zero-sample image classification method, system, device and medium for supplementing missing features.

[0007] The present invention is implemented as follows: a zero-sample image classification method for supplementing missing features comprises:

[0008] A zero-sample image classification dataset is collected, and semantic features of all categories are obtained at the same time; features are extracted from the images; a generative adversarial network is trained based on the features; forged unseen image features are extracted using the generative adversarial network, and the forged unseen image features are combined with image feature vectors to obtain an image training dataset; an image feature classification network model is trained based on the image training dataset, and the model is tested.

[0009] Furthermore, the features in the feature extraction of the image include image attribute features corresponding to the image and image feature vectors obtained by feature extraction using a pre-trained network;

[0010] The words of each dimension of the attribute in the image attribute feature are input into Word2vector to obtain a 1024-dimensional image feature vector; the image feature vectors of different attributes are clustered by the K-means algorithm, similar attributes are clustered into one category, and attribute grouping is performed.

[0011] Furthermore, the generative adversarial network is divided into two parts, a generator and a discriminator;

[0012] The input of the generator is the category attribute feature, and the output is the forged unseen class image feature, which is judged for authenticity by the discriminator; the input of the discriminator is the forged unseen class image feature and the real image feature, and the output is the authenticity confidence of the input feature, which is 1 for true and 0 for false.

[0013] Furthermore, the category attribute feature is obtained by setting all attributes of the category attributes of the unseen class to 0 through the attribute grouping, and then inputting it into the generator to obtain a forged unseen class image feature missing some attributes.

[0014] Furthermore, the generator consists of a four-layer neural network, which includes a 300×4096 fully connected layer, a LeakyReLU activation layer, a 4096×1024 fully connected layer, and a ReLU activation layer;

[0015] The discriminator consists of a four-layer neural network, namely a 1024×4096 fully connected layer, a LeakyReLU activation layer, a 4096×1 fully connected layer and a sigmoid activation layer.

[0016] Furthermore, the training formula of the generator is:

[0017]

[0018] In the formula, D is the discriminator, G is the generator, a represents the category attribute feature, and E represents the average of the data set;

[0019] The training formula of the discriminator is:

[0020]

[0021] In the formula, x represents the real image feature, Unseen image features that indicate forgery;

[0022] Another object of the present invention is to provide a zero-sample image classification system for supplementing missing features that implements the zero-sample image classification method for supplementing missing features, and the zero-sample image classification system for supplementing missing features comprises:

[0023] The dataset module is used to collect zero-shot image classification datasets and obtain the semantic features of all categories in the dataset;

[0024] A feature extraction module is used to extract features from images to obtain image feature vectors;

[0025] Clustering module, used to cluster attribute features using K-means method to obtain attribute groups;

[0026] A training module, used to train a generative adversarial network using image feature vectors and category attribute features;

[0027] Generate an adversarial network module, which is used to generate forged unseen image features, and combine the forged unseen image features with the image feature vector to obtain a complete image training data set, and use the image training data set to train the image feature classification network model;

[0028] The test module is used to test the test set data based on the image feature classification network model.

[0029] Another object of the present invention is to provide a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the zero-sample image classification method for supplementing missing features.

[0030] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to perform the steps of the zero-sample image classification method for supplementing missing features.

[0031] Another object of the present invention is to provide an information data processing terminal, which is used to implement the zero-sample image classification system for supplementing missing features.

[0032] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:

[0033] First, in view of the technical problems existing in the above-mentioned prior art and the difficulty of solving the problems, the technical solutions to be protected by the present invention and the results and data during the research and development process are closely combined to analyze in detail and deeply how the technical solutions of the present invention solve the technical problems, and some creative technical effects brought about after solving the problems. The specific description is as follows:

[0034] The method disclosed in the present invention belongs to a generative zero-shot image classification method. It optimizes the defect of the existing generative zero-shot image classification method, that is, the inability to generate image features that are missing certain attributes, so that the generated unseen class image features are more in line with the actual distribution, helping the classification model to learn more complete information, and ultimately improving the classification accuracy.

[0035] Second, considering the technical solution as a whole or from the perspective of the product, the technical effects and advantages of the technical solution to be protected by the present invention are described in detail as follows:

[0036] The present invention utilizes Word2vector to extract semantic features of attributes, thereby realizing automatic grouping of attributes; the category attributes are grouped according to the clustering results through the K-means algorithm, and then the values ​​of certain groups are set to 0 by randomly setting zeros when generating features of unseen classes, so as to input them into an input device to obtain unseen class image features with certain features, thereby helping the generated image features to be more consistent with the actual distribution.

[0037] Does the technical solution of the present invention solve the technical problems that people have been eager to solve but have never been able to solve successfully?

[0038] The present invention solves the problem that the feature distribution of forged unseen class images generated in the existing generative zero-shot image classification method is different from the distribution of actual images. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a flow chart of a zero-sample image classification method for supplementing missing features provided by an embodiment of the present invention;

[0040] Figure 2 is the picture data provided by the embodiment of the present invention, (a) a picture with complete category features, (b) a picture with some visual features missing;

[0041] Figure 3 is a schematic diagram of the structure of a generative adversarial network provided by an embodiment of the present invention;

[0042] Figure 4 is a schematic diagram of clustering attribute features using the K-means method provided by an embodiment of the present invention;

[0043] Figure 5It is a process diagram of obtaining the category semantic features of the missing part features provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0045] In order to enable those skilled in the art to fully understand how to implement the present invention in detail, this section is an explanatory embodiment that expands and describes the technical solution of the claims.

[0046] The zero-sample image classification method for supplementing missing features provided by an embodiment of the present invention includes:

[0047] The K-means algorithm is used to group the category attributes according to the clustering results, and then when generating features of unseen classes, the values ​​of some groups are set to 0 by randomly setting them to zero, which are then input into the input device to obtain unseen class image features with certain features, thereby helping the generated image features to be more consistent with the actual distribution.

[0048] like Figure 1 As shown, the specific process of the zero-sample image classification method for supplementing missing features includes the following steps:

[0049] S101: Collect a zero-shot image classification dataset and obtain semantic features of all categories in the dataset; each image in the dataset corresponds to a manually annotated image attribute feature.

[0050] S102: extract features from each image using a pre-trained network to obtain an image feature vector;

[0051] S103: Train a generative adversarial network using the image feature vector and the category attribute feature; the generative adversarial network is divided into two parts, a generator and a discriminator. The input of the generator of this method is the category attribute feature, and the output is the forged image feature. The input of the discriminator is the forged image feature and the real image feature, and the output is the authenticity confidence of the input feature, which is 1 for the real one and 0 for the fake one;

[0052] S104: inputting the category attribute features of the unseen class into the generator of the generative adversarial network, and outputting the forged unseen class image features;

[0053] S105: combining the forged unseen class image features with the seen class image feature vectors to obtain a complete image training data set;

[0054] S106: Use the data in the image training data set to train an image feature classification network; for example, a ResNet18 image classification network;

[0055] S107: Use the trained classification model to test the data in the test set;

[0056] In order to prove the creativity and technical value of the technical solution of the present invention, this section provides application examples of the claimed technical solution on specific products or related technologies.

[0057] The entire process of the zero-sample image classification method for supplementing missing features according to an embodiment of the present invention is as follows:

[0058] Step 1: Get the zero-shot image classification dataset CUB bird classification dataset. Each image in the dataset corresponds to a manually annotated 300-dimensional category attribute feature. The dataset has 11,788 images and 200 categories. There are 7,057 images in the training set and 4,731 images in the test set. There are 150 visible classes and 50 unseen classes. Each of the 200 categories also has a corresponding 300-dimensional category attribute feature.

[0059] Step 2: Extract the 1024-dimensional image features of the image in step 1 through the ResNet18 network pre-trained on the ImageNet dataset;

[0060] Step 3: Each dimension of the category attribute feature in step 1 represents an attribute with practical meaning. A semantic vector of each attribute is obtained by inputting the attribute word into Word2vector;

[0061] Step 4: Cluster the word vector features in step 3 using the K-means clustering algorithm, and set the group size to 10. The clustering algorithm divides the 300 category attributes into 10 groups;

[0062] Step 5: Use a common generative zero-shot image classification method to train a generative adversarial network, such as CLS-WGAN (Xian, Y., Lorenz, T., Schiele, B., & Akata, Z. (2018). Feature generating networks for zero-shot learning. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 5542-5551).);

[0063] Step six: Generate features using the generator of the generative adversarial network in step five. In particular, in this step, the present invention does not directly use the unseen class category attributes in step one to generate features in the ordinary generative zero-sample image generation method, but uses the attribute grouping in step four to set one of the 10 groups of category attributes of the unseen class to 0, and then inputs it into the generator to obtain forged unseen class image features that are missing certain attributes. At the same time, the complete unseen class semantic features are also used to generate unseen class visual features. After all, the image features with missing features account for a small part of the overall distribution;

[0064] Step 7: Use the forged unseen class image features generated in step 6 and the seen class image features extracted in step 2 to train a full-class image feature classifier;

[0065] Step 8: Test the test data set in step 1 and perform evaluation.

[0066] Figure 2 Images with completed category features and images with some visual features missing were shown; Figure 2 (a) has all the visual features of objects of this category, Figure 2 (b) Only some features of the category, that is, some visual features are missing.

[0067] Figure 3 The general structure of the generative adversarial network trained in this method is shown in Figure 1, including:

[0068] The image features of the image are obtained through the pre-trained model, and the attribute features are input into the generator to obtain the forged image features, and the discriminator is used to distinguish the authenticity. The discriminator and generator are trained using data.

[0069] The training formula of the generator is:

[0070]

[0071] Where D is the discriminator, G is the generator, a represents the attribute feature, and E represents the average of the data set. The training formula of the discriminator is:

[0072]

[0073] In the formula, x represents the image feature, Indicates forged image features.

[0074] The generator consists of a four-layer neural network, which consists of a 300×4096 fully connected layer, a LeakyReLU activation layer, a 4096×1024 fully connected layer and a ReLU activation layer.

[0075] The discriminator is also composed of a 4-layer neural network, namely a 1024×4096 fully connected layer, a LeakyReLU activation layer, a 4096×1 fully connected layer and a sigmoid activation layer.

[0076] Figure 4 This is a schematic diagram of clustering attribute features using the K-means method. The category attributes are the words "round head", "pointed head", "red" and "black" in the picture. After entering Word2vector, the corresponding 1024-dimensional feature vector is obtained. Then, the feature vectors of different attributes are clustered using K-means to cluster similar attributes into one category.

[0077] Figure 5 It is a process of obtaining the category semantic features of the missing part of the features through grouping. According to the clustering results, the attributes of the same group are grouped together, and then the value of one of the groups is randomly assigned to 0. The attributes are then reintegrated to obtain the category semantic features of the missing part of the features, and then they are input into the generator.

[0078] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. It can be understood by a person of ordinary skill in the art that the above-mentioned devices and methods can be implemented using computer executable instructions and / or contained in a processor control code, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. Such code is provided on the carrier medium. The device and its modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, and can also be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.

[0079] The embodiments of the present invention have achieved some positive effects during the development or use process, and indeed have great advantages over the prior art. The following content is described in conjunction with data, charts, etc. of the test process.

[0080] The Top1 accuracy of the visible class, the Top1 accuracy of the unseen class and the harmonic value of the original CLS-WGAN method on the CUB dataset are 57.7%, 43.7% and 49.7% respectively, and the results obtained by this method are 58.0%, 50.2% and 53.8%. The method provided in the embodiment of the present invention mainly deals with the problem that the generated forged unseen class features do not match the distribution of the actual unseen class features, so the Top1 accuracy of the unseen class obtained is significantly improved compared with the original method, thereby improving the harmonic value, and the visible class has a slight improvement.

[0081] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with the technical field within the technical scope disclosed by the present invention and within the spirit and principle of the present invention should be covered by the protection scope of the present invention.

Claims

1. A zero-shot image classification method to supplement missing features, It is characterized in that include: Collect zero-shot image classification datasets and obtain category attribute features of all categories at the same time; Extract features from images; Train a generative adversarial network based on features; Extracting forged unseen image features using the generative adversarial network, combining the forged unseen image features with the image feature vector to obtain an image training data set; training an image feature classification network model based on the image training data set, and performing a test; The features in the feature extraction of the picture include picture attribute features corresponding to the picture and picture feature vectors obtained by feature extraction using a pre-trained network; Input the words of each dimension of the attribute in the image attribute feature into Word2vector to obtain a 1024-dimensional image feature vector; cluster the image feature vectors of different attributes using the K-means algorithm, cluster similar attributes into one category, and perform attribute grouping; The generative adversarial network is divided into two parts, a generator and a discriminator; The input of the generator is the category attribute feature of the image, and the output is the forged visible class image feature, which is used to distinguish the authenticity through the discriminator; the input of the discriminator is the forged visible class image feature and the real visible class image feature extracted in step 2, and the output is the authenticity confidence of the input feature, which is 1 for the real one and 0 for the fake one; The category attribute feature is obtained by setting all values ​​of a group of category attribute features of the unseen class to 0 through the attribute grouping, and then inputting it into the generator to obtain a forged unseen class image feature that lacks a group of attributes.

2. The zero-sample image classification method for supplementing missing features as claimed in claim 1, It is characterized in that The generator consists of a four-layer neural network, which is a 300×4096 fully connected layer, a LeakyReLU activation layer, a 4096×1024 fully connected layer and a ReLU activation layer; The discriminator consists of a four-layer neural network, namely a 1024×4096 fully connected layer, a LeakyReLU activation layer, a 4096×1 fully connected layer and a sigmoid activation layer.

3. The zero-sample image classification method for supplementing missing features as claimed in claim 1, It is characterized in that The training formula of the generator is: In the formula, D is the discriminator, G is a generator, a Represents the category attribute characteristics, y represents the label, n represents the total number of data sets, and i represents the i-th data; The training formula of the discriminator is: In the formula, x Represents the real picture features, Indicates the forged features of unseen images.

4. A zero-shot image classification system for supplementing missing features implementing the zero-shot image classification method for supplementing missing features as claimed in any one of claims 1 to 3, It is characterized in that The zero-sample image classification system for supplementing missing features includes: The dataset module is used to collect zero-shot image classification datasets and obtain the semantic features of all categories in the dataset; A feature extraction module is used to extract features from images to obtain image feature vectors; Clustering module, used to cluster attribute features using K-means algorithm to obtain attribute groups; A training module, used to train a generative adversarial network using image feature vectors and category attribute features; Generate an adversarial network module, which is used to extract forged unseen image features and combine the forged unseen image features with the image feature vector to obtain a complete image training data set, and use the image training data set to train the image feature classification network model; The test module is used to test the test set data based on the image feature classification network model.

5. A computer device, It is characterized in that The computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the zero-sample image classification method for supplementing missing features as described in any one of claims 1-3.

6. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the zero-sample image classification method for supplementing missing features as described in any one of claims 1 to 3.

7. An information data processing terminal, It is characterized in that The information data processing terminal is used to implement the zero-sample image classification system for supplementing missing features as described in claim 4.

Citation Information

Patent Citations

  • Zero-sample image recognition method and system based on generative adversarial network

    CN111476294A

  • Generalized zero sample image recognition method and model based on semantic information retention

    CN113361646A