A zero-sample image classification method and system in machine learning
The zero-sample image features are extracted through semantic decoupling, and the encoder and decoder are used to generate new unseen class features, which solves the problem of judgment errors in the absence of class in full supervision learning, and improves the accuracy of zero-sample image classification.
Patent Information
- Application Number
- CN202211514426.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-29
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2042-11-29
AI Technical Summary
When existing fully supervised machine learning methods encounter unseen situations, serious judgment errors and even inability to judge, especially in the field of image classification. How to extract the characteristics of semantic information related parts in zero-sample pictures is an urgent problem to be solved.
Data features are extracted through pre-training networks, semantic decoupling is used to calculate differential losses and reconstruction losses, and new unseen class features are generated. Combined with comparison supervision and semantic supervision learning, the classifier is trained to perform zero-sample picture classification.
A zero-sample image classification method in machine learning based on semantic decoupling is provided, which generates more reliable features of unseen class, improves the classification accuracy of the classifier, and solves the problem of judgment errors in the case of unseen class in full supervised learning.
Smart Images

Figure CN115731421B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision technology, and in particular relates to a zero-sample image classification method in machine learning. Background Art
[0002] Currently, most existing machine learning methods are fully supervised. Full supervision means that the model is exposed to all knowledge during the training phase. This approach benefits from leveraging existing big data and massive computing power to catalog possible scenarios, then using machine learning algorithms to learn their distribution. During the operational phase, machine learning methods are then used to assess and process the scenarios. However, this approach has a significant drawback: fully supervised learning methods can make serious misjudgments or even be unable to make judgments when encountering unseen situations. To address this issue with fully supervised learning, zero-shot learning was proposed.
[0003] Zero-shot learning, an extreme case in machine learning, assumes that certain situations are unknown during the training phase. In the field of image classification, the zero-shot problem is set up so that the model is not exposed to certain categories of images during the training phase, but may appear during the testing phase. These categories are called unseen classes. This requires that the designed zero-shot image classification algorithm learns certain semantic concepts from the seen classes in the training data, helping the model to infer whether an image belongs to an unseen class during the testing phase based on these semantic concepts.
[0004] Therefore, how to extract the features of the semantic information related parts in zero-shot images is very important for zero-shot learning, so it is urgent to design a new zero-shot image classification method in machine learning.
[0005] Through the above analysis, the problems and defects of the existing technology are: when encountering unprecedented situations, the existing fully supervised machine learning methods will make serious judgment errors or even be unable to make judgments. Summary of the Invention
[0006] In response to the problems existing in the prior art, the present invention provides a zero-shot image classification method in machine learning, and in particular, relates to a zero-shot image classification method in machine learning based on semantic decoupling.
[0007] The present invention is implemented as follows: a zero-shot image classification method in machine learning, which includes: obtaining a zero-shot image data set; extracting data features through a pre-trained network; selecting two images and inputting them into an encoder 1, calculating the difference loss of the results; inputting them into a decoder 1, and calculating the image reconstruction loss; selecting an image and inputting it into the encoder 1, inputting the semantic features into the encoder 2, and inputting them into the decoder 2 after swapping, calculating the difference loss, the image reconstruction loss, and the semantic reconstruction loss; calculating the total loss; training the encoder and decoder; generating new data using the trained encoder and decoder; combining the new data and the original data to train a classifier; and testing a test sample using the classifier.
[0008] Furthermore, the zero-shot image classification method in machine learning includes the following steps:
[0009] Step 1: Collect and establish a zero-shot image training set and a zero-shot image test set, and extract image features of the training set through a pre-trained network;
[0010] Step 2: Randomly select two images of the same type in the training dataset, input the image features corresponding to the images into encoder 1, output the two results of the two images respectively as the category invariant features and the image-specific features, and calculate the difference loss 1 of the category invariant features of the two images;
[0011] Step 3: After exchanging the category-invariant features of the two images in step 2, concatenate the category-invariant features and the image-specific features, input them into decoder 1, and output the reconstructed image features. Calculate the image reconstruction loss 1.
[0012] Step 4: Randomly select a random image from the training dataset and obtain the corresponding semantic feature vector. Input the image features corresponding to the image into encoder 1 to obtain category-invariant features and image-specific features. Input the semantic features into encoder 2, output the category-invariant features, and calculate the difference loss 2 between the semantic features and the category-invariant features of the image.
[0013] Step 5: Concatenate the category-invariant features of the semantic features in step 4 with the image-specific features of the image, input them into decoder 1, and output the reconstructed image features. Calculate the image reconstruction loss 2. Input the image category-invariant features in step 4 into decoder 2 to obtain the reconstructed semantic features, and calculate the semantic reconstruction loss.
[0014] Step 6: Sum the difference loss 1 in step 2, the image reconstruction loss 1 in step 3, the difference loss 2 in step 4, the image reconstruction loss 2 in step 5, and the semantic reconstruction loss to get the total loss;
[0015] Step 7: Repeat steps 2 to 6 until the total loss in step 6 stabilizes;
[0016] Step 8: Input the image features of the image training data in step 1 into encoder 1 to obtain the image features corresponding to each image; concatenate the category-invariant features obtained by inputting different semantic features into encoder 2 and input them into decoder 1 to obtain new image features;
[0017] Step nine: Mix the new image features obtained in step eight with the image features of the training dataset to obtain a complete training dataset; train the image feature classifier with the new complete training dataset and test the test image data in step one.
[0018] Furthermore, in step 1, ResNet50 is used to extract image features. ResNet50 is pre-trained on the ImageNet dataset and has 50 layers. The results of the layer before the final fully connected layer are used as the image features.
[0019] Furthermore, the encoder 1 in step 2 is composed of a four-layer neural network, namely a fully connected layer, an activation layer, a fully connected layer and an activation layer.
[0020] Furthermore, in step 2, the difference loss 1 is calculated using the squared difference loss of the category-invariant features of the two images. The calculation formula is as follows:
[0021] Loss1=||p1-p2||2;
[0022] Among them, p1 is the category-invariant feature extracted from picture 1, and p2 is the category-invariant feature extracted from picture 2.
[0023] Furthermore, the image reconstruction loss 1 in step 3 is the square difference loss between the original image features and the reconstructed image features, and the calculation formula is as follows:
[0024] Loss2=||f1-r1||2+||f2-r2||2;
[0025] Among them, f1 and f2 are the original image features of pictures 1 and 2 respectively, and r1 and r2 are the reconstructed image features of pictures 1 and 2 respectively.
[0026] Furthermore, the decoder 1 in step 3 is composed of a four-layer neural network, namely a fully connected layer, an activation layer, a fully connected layer, and an activation layer.
[0027] Furthermore, the difference loss 2 in step 4 is the squared difference loss between the semantic features and the category-invariant features of the image, and is calculated as follows:
[0028] Loss3=||e3-a3||2;
[0029] Among them, e3 is the image category invariant feature, and a3 is the category invariant feature of the semantic feature.
[0030] Furthermore, the encoder 2 in step 4 is composed of a four-layer neural network, namely a fully connected layer, an activation layer, a fully connected layer and an activation layer.
[0031] Furthermore, the image reconstruction loss 2 in step 5 is the square difference loss between the original image features and the reconstructed image features, and the calculation formula is as follows:
[0032] Loss4=||f3-r3||2;
[0033] Among them, f3 is the image category invariant feature, and r3 is the category invariant feature of the semantic feature.
[0034] Furthermore, the semantic reconstruction loss in step 5 is the squared difference loss between the original semantic features and the reconstructed semantic features, and the calculation formula is as follows:
[0035] Loss5=||u3-c3||2;
[0036] Among them, u3 is the original semantic feature and c3 is the reconstructed semantic feature.
[0037] Furthermore, the decoder 2 in step 5 is composed of a four-layer neural network, namely a fully connected layer, an activation layer, a fully connected layer and an activation layer.
[0038] Another object of the present invention is to provide a zero-sample image classification system applying the zero-sample image classification method in machine learning, the zero-sample image classification system comprising:
[0039] Feature extraction module, used to obtain zero-sample image data and extract data features through pre-trained network;
[0040] The loss calculation module is used to select two images and input them into the encoder 1 to calculate the difference loss 1 of the results; input them into the decoder 1 to calculate the image reconstruction loss 1;
[0041] The total loss calculation module is used to select an image input into encoder 1, input the semantic features into encoder 2 and decoder 2, calculate the difference loss 2, image reconstruction loss 2 and semantic reconstruction loss, and calculate the total loss;
[0042] The sample testing module is used to generate new data after the encoder and decoder are trained, combine the new data with the original data to train the classifier, and use the classifier to test the test samples.
[0043] Another object of the present invention is to provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the zero-sample image classification method in machine learning.
[0044] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to perform the steps of the zero-sample image classification method in machine learning.
[0045] Another object of the present invention is to provide an information data processing terminal, which is used to implement the zero-sample image classification system.
[0046] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:
[0047] First, in view of the technical problems existing in the above-mentioned prior art and the difficulty of solving these problems, we closely combine the technical solutions to be protected by the present invention and the results and data during the research and development process, and conduct a detailed and in-depth analysis of how the technical solutions of the present invention solve the technical problems and some creative technical effects brought about by solving the problems. The specific description is as follows:
[0048] To address the lack of unseen visual features in zero-shot image classification, the present invention decouples semantically relevant category features from image-specific features through encoding. After learning using contrastive supervised learning and semantically supervised learning, the decoupled image-specific features and category features are concatenated and decoded to generate new unseen features. This provides a zero-shot image classification method for machine learning based on semantic decoupling. The zero-shot image classification method for machine learning provided by the present invention separates the semantic information of the category and the image-specific information in each image, helping the zero-shot learning algorithm to better perform classification.
[0049] The present invention decouples image-specific features and category-invariant features in an image by utilizing encoding and decoding methods, and trains the encoder and decoder using comparative supervision between images and semantic supervision between images and semantic features. The trained encoder and decoder are used to generate new features of unseen classes, thus filling in the gaps in the unseen class features in the training dataset, and finally the complete dataset is used to train the image classification network.
[0050] Second, considering the technical solution as a whole or from the perspective of the product, the technical effects and advantages of the technical solution to be protected by the present invention are described in detail as follows:
[0051] This paper uses encoding to decouple image-specific features from class-invariant features in an image, and then uses a decoder to generate new features. This provides a zero-shot image classification method for machine learning based on semantic decoupling. This decoupling approach provides a more reliable feature generation method, ensuring the reliability of generated features for unseen classes and ultimately enabling better training of unseen class classifiers.
[0052] The present invention ensures that the encoder can decouple picture-specific features and category-specific features by training two pictures of the same type. By training a single picture, the encoder can decouple category-specific features from semantic features. Finally, the trained encoder is used to generate new picture features of unseen classes, helping the method to supplement features of unseen classes, thereby solving the zero-sample problem and improving the classification accuracy of the classification network.
[0053] Third, as auxiliary evidence for the inventiveness of the claims of the present invention, it is also reflected in the following important aspects:
[0054] Does the technical solution of the present invention solve the technical problems that people have always wanted to solve but have never been able to solve successfully?
[0055] By separating the category semantic information and unique information in training images, the present invention helps the zero-shot image classification method in generative machine learning to generate more category-discriminative image features of unseen classes, solving the problem of fuzzy and lack of variation in the semantic information of unseen class image features generated by ordinary generative zero-shot methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0057] Figure 1 This is a flow chart of a zero-shot image classification method in machine learning provided by an embodiment of the present invention;
[0058] Figure 2 This is a schematic diagram of the principle of the zero-shot image classification method in machine learning provided by an embodiment of the present invention;
[0059] Figure 3 This is a flow chart of a method for training two similar images provided by an embodiment of the present invention;
[0060] Figure 4 This is a flow chart of a method for training images according to an embodiment of the present invention;
[0061] Figure 5 1 is a schematic structural diagram of an encoder 1 provided in an embodiment of the present invention;
[0062] Figure 6 is a schematic structural diagram of a decoder 1 provided in an embodiment of the present invention;
[0063] Figure 7 2 is a schematic structural diagram of an encoder 2 provided in an embodiment of the present invention;
[0064] Figure 8 It is a schematic structural diagram of the decoder 2 provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0065] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0066] In response to the problems existing in the prior art, the present invention provides a zero-sample image classification method in machine learning. The present invention is described in detail below with reference to the accompanying drawings.
[0067] In order to enable those skilled in the art to fully understand how to implement the present invention, this section provides an explanatory embodiment that expands on the technical solutions of the claims.
[0068] like Figure 1 As shown, the zero-shot image classification method in machine learning provided by the embodiment of the present invention includes the following steps:
[0069] S101, obtaining a zero-shot image dataset and extracting data features through a pre-trained network;
[0070] S102, select two pictures and input them into encoder 1, calculate the difference loss 1 of the results; input them into decoder 1, calculate the picture reconstruction loss 1;
[0071] S103, select an image and input it into encoder 1, then input the semantic features into encoder 2, exchange them and input them into decoder 2, calculate the difference loss 2, image reconstruction loss 2 and semantic reconstruction loss, and calculate the total loss;
[0072] S104, training the encoder and decoder; using the trained encoder and decoder to generate new data; combining the new data with the original data to train a classifier, and using the classifier to test the test sample.
[0073] As a preferred embodiment, Figure 2 As shown, the zero-sample image classification method in machine learning provided by the embodiment of the present invention specifically includes the following steps:
[0074] Step 1: Collect and establish a zero-shot image training set and a zero-shot image test set;
[0075] Step 2: Extract the image features of the training set in step 1 through the pre-trained network;
[0076] Step 3: Randomly select two images of the same type in the training dataset, input the image features corresponding to the images into encoder 1, and output two results for each image, namely the category-invariant feature and the image-specific feature; calculate the difference loss 1 of the category-invariant features of the two images;
[0077] Step 4: After swapping the category-invariant features of the two images in step 3, concatenate the category-invariant features and the image-specific features, input them into decoder 1, and output the reconstructed image features. Calculate the image reconstruction loss 1.
[0078] Step 5: Randomly select a random image from the training dataset and obtain its corresponding semantic feature vector. Input the image features corresponding to the image into encoder 1 to obtain its category-invariant features and image-specific features. Input the semantic features into encoder 2 to output the category-invariant features. Calculate the difference loss 2 between the semantic features and the category-invariant features of the image.
[0079] Step 6: Concatenate the category-invariant features of the semantic features in step 5 with the image-specific features of the image, input them into decoder 1, and output the reconstructed image features. Calculate the image reconstruction loss 2.
[0080] Step 7: After inputting the category-invariant features of the image in step 5 into decoder 2, the reconstructed semantic features are obtained and the semantic reconstruction loss is calculated;
[0081] Step 8: Sum the difference loss 1 in step 3, the image reconstruction loss 1 in step 4, the difference loss 2 in step 5, the image reconstruction loss 2 in step 6, and the semantic reconstruction loss in step 7 to get the total loss of the entire method;
[0082] Step 9: Repeat steps 3 to 8 until the total loss in step 8 stabilizes.
[0083] Step 10: Input the image features of the image training data in step 1 into encoder 1 to obtain the image features corresponding to each image. After concatenating the image features with the category-invariant features obtained by inputting different semantic features into encoder 2, the features are input into decoder 1 to obtain new image features.
[0084] Step 11: Mix the new image features obtained in step 10 with the image features of the training dataset to obtain a complete training dataset;
[0085] Step 12: Use the new complete training dataset from step 11 to train an image feature classifier, and test it on the test image data from step 1.
[0086] As a preferred technical solution, in step 2 provided in an embodiment of the present invention, ResNet50 is used to extract image features. ResNet50 is pre-trained on the ImageNet dataset and has a total of 50 layers. The results of the layer before the final fully connected layer are taken as the image features.
[0087] As a preferred technical solution, in step 3 provided in an embodiment of the present invention, the encoder 1 is composed of a 4-layer neural network, namely a fully connected layer, an activation layer, a fully connected layer and an activation layer.
[0088] As a preferred technical solution, in step 3 provided in the embodiment of the present invention, the square difference loss of the category-invariant features of the two images is used to calculate the difference loss 1, and the calculation formula is as follows:
[0089] Loss1=||p1-p2||2;
[0090] Among them, p1 is the category-invariant feature extracted from picture 1, and p2 is the category-invariant feature extracted from picture 2.
[0091] As a preferred technical solution, in step 4 provided in the embodiment of the present invention, the image reconstruction loss 1 is the square difference loss between the original image features and the reconstructed image features, and the calculation formula is as follows:
[0092] Loss2=||f1-r1||2+||f2-r2||2;
[0093] Among them, f1 and f2 are the original image features of pictures 1 and 2 respectively, and r1 and r2 are the reconstructed image features of pictures 1 and 2 respectively.
[0094] As a preferred technical solution, in step 4 provided in an embodiment of the present invention, the decoder 1 is composed of a 4-layer neural network, namely a fully connected layer, an activation layer, a fully connected layer and an activation layer.
[0095] As a preferred technical solution, in step 5 provided in the embodiment of the present invention, the difference loss 2 is the square difference loss between the semantic features and the category-invariant features of the image, and is calculated as follows:
[0096] Loss3=||e3-a3||2;
[0097] Among them, e3 is the image category invariant feature, and a3 is the category invariant feature of the semantic feature.
[0098] As a preferred technical solution, in step 5 provided in an embodiment of the present invention, the encoder 2 is composed of a 4-layer neural network, namely a fully connected layer, an activation layer, a fully connected layer and an activation layer.
[0099] As a preferred technical solution, in step 6 provided in the embodiment of the present invention, the image reconstruction loss 2 is the square difference loss between the original image features and the reconstructed image features, and the calculation formula is as follows:
[0100] Loss4=||f3-r3||2;
[0101] Among them, f3 is the image category invariant feature, and r3 is the category invariant feature of the semantic feature.
[0102] As a preferred technical solution, in step 7 provided in the embodiment of the present invention, the semantic reconstruction loss is the square difference loss between the original semantic features and the reconstructed semantic features, and the calculation formula is as follows:
[0103] Loss5=||u3-c3||2;
[0104] Among them, u3 is the original semantic feature and c3 is the reconstructed semantic feature.
[0105] As a preferred technical solution, in step 7 provided in an embodiment of the present invention, the decoder 2 is composed of a 4-layer neural network, namely a fully connected layer, an activation layer, a fully connected layer and an activation layer.
[0106] The zero-shot image classification system provided by the embodiment of the present invention includes:
[0107] Feature extraction module, used to obtain zero-sample image data and extract data features through pre-trained network;
[0108] The loss calculation module is used to select two images and input them into the encoder 1 to calculate the difference loss 1 of the results; input them into the decoder 1 to calculate the image reconstruction loss 1;
[0109] The total loss calculation module is used to select an image input into encoder 1, input the semantic features into encoder 2 and decoder 2, calculate the difference loss 2, image reconstruction loss 2 and semantic reconstruction loss, and calculate the total loss;
[0110] The sample testing module is used to generate new data after the encoder and decoder are trained, combine the new data with the original data to train the classifier, and use the classifier to test the test samples.
[0111] The embodiments of the present invention have achieved some positive results during the development or use process, and indeed have great advantages over the existing technology. The following content describes them in conjunction with data, charts, etc. from the experimental process.
[0112] Preferably, the zero-shot image classification method in machine learning provided by an embodiment of the present invention specifically includes the following steps:
[0113] Step 1: Use the CUB-200 bird dataset as the zero-shot dataset. It has a total of 11,788 images and 200 categories. Take 7,057 images and 150 categories as the training dataset, and the remaining 4,731 images and 50 categories as the test set.
[0114] Step 2: Extract 2048-dimensional image features of the 11,788 images in step 1 by removing the fully connected layer of the network part of the pre-trained ResNet50 network on ImageNet.
[0115] Step 3: If Figure 3 As shown in the figure, we randomly select image features f1 and f2 from two images I1 and I2 of the same type in the training dataset, input the two image features into encoder 1, and output 1024-dimensional category-invariant features p1 and p2, and 1024-dimensional image-specific features q1 and q2. The difference loss 1 of the category-invariant features is calculated as follows:
[0116] Loss1=||p1-p2||2.
[0117] The structure of encoder 1 is as follows Figure 5 As shown in the figure, it consists of a 4-layer neural network, namely a 2048×1024 fully connected layer, a ReLU activation function, a 1024×2048 fully connected layer and a LeakyReLU activation function.
[0118] Step 4: After exchanging the two category-invariant features obtained in step 3, concatenate p1 and q2, and p2 and q1. Input the concatenated 2048-dimensional features into decoder 1 and output the reconstructed image features r1 and r2. Calculate the image reconstruction loss 1 using the following formula:
[0119] Loss2=||f1-r1||2+||f2-r2||2.
[0120] The structure of decoder 1 is as follows Figure 6 As shown in the figure, it consists of a 4-layer neural network, namely a 2048×1024 fully connected layer, a LeakyReLU activation function, a 1024×2048 fully connected layer and a ReLU activation function.
[0121] Step 5: Figure 4As shown, we randomly select the image feature f3 of a random image I3 in the training dataset and obtain its corresponding semantic feature vector u3. The image feature f3 is input into encoder 1 to obtain its category-invariant feature p3 and image-specific feature q3. The semantic feature corresponding to the 300-dimensional image I3 is input into encoder 2, which outputs the 1024-dimensional category-invariant feature a3. The difference loss 2 between the semantic feature and the category-invariant feature of the image is calculated as follows:
[0122] Loss3=||p3-a3||2.
[0123] The structure of encoder 2 is as follows Figure 7 As shown in the figure, it consists of 4 layers of neural network, namely a 300×512 fully connected layer, a ReLU activation function, a 512×1024 fully connected layer and a LeakyReLU activation function.
[0124] Step 6: Concatenate the semantic feature a3 from step 5 with the image-specific feature q3. Input the 2048-dimensional concatenation result into decoder 1 and output the reconstructed image feature r3. Calculate the image reconstruction loss 2 using the following formula:
[0125] Loss4=||f3-r3||2.
[0126] Step 7: After the category-invariant feature p3 of image I3 in step 5 is input to decoder 2, the reconstructed semantic feature c3 is obtained and the semantic reconstruction loss is calculated. The formula is as follows:
[0127] Loss5=||u3-c3||2.
[0128] The structure of decoder 2 is as follows Figure 8 As shown in the figure, it consists of a 4-layer neural network, namely a 1024×512 fully connected layer, a LeakyReLU activation function, a 512×300 fully connected layer and a ReLU activation function.
[0129] Step 8: Sum the difference loss 1 in step 3, the image reconstruction loss 1 in step 4, the difference loss 2 in step 5, the image reconstruction loss 2 in step 6, and the semantic reconstruction loss in step 7 to get the total loss of the entire method.
[0130] Step 9: Repeat steps 3 to 8 until the total loss in step 8 stabilizes.
[0131] Step 10: Randomly select an image from the training dataset and randomly select the semantic features of a class, input them into encoder 1 and encoder 2 respectively, concatenate the obtained category-invariant features and image-specific features, and then input the concatenated results into decoder 1 to obtain new image features.
[0132] Step 11: Mix the new image features obtained in step 10 with the image features of the training dataset to obtain a complete training dataset.
[0133] Step 12: Use the new complete training dataset from step 11 to train an image feature classifier, and test it on the test image data from step 1.
[0134] After testing, the method of the present invention achieved a Top 1 accuracy of 56 for the seen class, a Top 1 accuracy of 53 for the unseen class, and a harmonic value of 54 in the above dataset, which are higher than the results of 53.5, 51.6, and 52.4 of the similar method CADA model (Schonfeld, E., Ebrahimi, S., Sinha, S., Darrell, T., & Akata, Z. (2019). Generalized zero-and few-shot learning via aligned variational autoencoders. In Proceedings of the IEEE / CVFConference on Computer Vision and Pattern Recognition (pp.8247-8255)), fully demonstrating the effectiveness of the method of the present invention.
[0135] This invention is a zero-shot image classification method for machine learning based on generative methods. It generates better visual features for unseen classes by decoupling the class-invariant features from the image-specific features. Specifically, the decoupling is achieved through dual-image comparison training and semantically supervised training using the relationship between semantics and images.
[0136] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic; the software portion can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art will appreciate that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or contained in processor control code, for example, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.
[0137] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.
Claims
1. A zero-shot image classification method in machine learning, characterized in that: The zero-shot image classification method in machine learning includes the following steps: Step 1: Collect and establish a zero-shot image training set and a zero-shot image test set, and extract image features of the training set through a pre-trained network; Step 2: Randomly select two images of the same type in the training dataset, input the image features corresponding to the images into encoder 1, output the two results of the two images respectively as the category invariant features and the image-specific features, and calculate the difference loss 1 of the category invariant features of the two images; Step 3: After exchanging the category-invariant features of the two images in step 2, concatenate the category-invariant features and the image-specific features, input them into decoder 1, and output the reconstructed image features. Calculate the image reconstruction loss 1. Step 4: Randomly select a random image from the training dataset and obtain the corresponding semantic feature vector. Input the image features corresponding to the image into encoder 1 to obtain category-invariant features and image-specific features. Input the semantic features into encoder 2, output the category-invariant features, and calculate the difference loss 2 between the semantic features and the category-invariant features of the image. Step 5: Concatenate the category-invariant features of the semantic features in step 4 with the image-specific features of the image, input them into decoder 1, and output the reconstructed image features. Calculate the image reconstruction loss 2. Input the image category-invariant features in step 4 into decoder 2 to obtain the reconstructed semantic features, and calculate the semantic reconstruction loss. Step 6: Sum the difference loss 1 in step 2, the image reconstruction loss 1 in step 3, the difference loss 2 in step 4, the image reconstruction loss 2 in step 5, and the semantic reconstruction loss to get the total loss; Step 7: Repeat steps 2 to 6 until the total loss in step 6 stabilizes; Step 8: Input the image features of the image training data in step 1 into encoder 1 to obtain the image features corresponding to each image; concatenate the category-invariant features obtained by inputting different semantic features into encoder 2 and input them into decoder 1 to obtain new image features; Step nine: Mix the new image features obtained in step eight with the image features of the training dataset to obtain a complete training dataset; train the image feature classifier with the new complete training dataset and test the test image data in step one.
2. The zero-shot image classification method in machine learning according to claim 1, wherein: In step 1, ResNet50 is used to extract the image features. ResNet50 is pre-trained on the ImageNet dataset and has 50 layers. The result of the layer before the last fully connected layer is taken as the image features. The encoder 1 in step 2 is composed of a 4-layer neural network, namely a fully connected layer, an activation layer, a fully connected layer, and an activation layer; The calculation of the difference loss 1 in step 2 uses the squared difference loss of the category-invariant features of the two images. The calculation formula is as follows: Loss1=||p1-p2||2; Among them, p1 is the category-invariant feature extracted from picture 1, and p2 is the category-invariant feature extracted from picture 2.
3. The zero-shot image classification method in machine learning according to claim 1, wherein: The image reconstruction loss 1 in step 3 is the squared difference loss between the original image features and the reconstructed image features, and is calculated as follows: Loss2=||f1-r1||2+||f2-r2||2; Among them, f1 and f2 are the original image features of pictures 1 and 2 respectively, and r1 and r2 are the reconstructed image features of pictures 1 and 2; The decoder 1 in step 3 is composed of a four-layer neural network, namely a fully connected layer, an activation layer, a fully connected layer, and an activation layer.
4. The zero-shot image classification method in machine learning according to claim 1, wherein: The difference loss 2 in step 4 is the squared difference loss between the semantic features and the category-invariant features of the image, and is calculated as follows: Loss3=||e3-a3||2; Among them, e3 is the image category invariant feature, and a3 is the category invariant feature of the semantic feature; The encoder 2 in step 4 is composed of a four-layer neural network, namely a fully connected layer, an activation layer, a fully connected layer, and an activation layer.
5. The zero-shot image classification method in machine learning according to claim 1, wherein: The image reconstruction loss 2 in step 5 is the square difference loss between the original image features and the reconstructed image features, and the calculation formula is as follows: Loss4=||f3-r3||2; Among them, f3 is the image category invariant feature, r3 is the category invariant feature of the semantic feature; The semantic reconstruction loss in step 5 is the squared difference loss between the original semantic features and the reconstructed semantic features, and is calculated as follows: Loss5=||u3-c3||2; Among them, u3 is the original semantic feature, and c3 is the reconstructed semantic feature; The decoder 2 in step 5 is composed of a four-layer neural network, namely a fully connected layer, an activation layer, a fully connected layer, and an activation layer.
6. A zero-shot image classification system using the zero-shot image classification method in machine learning according to any one of claims 1 to 5, characterized in that: The zero-shot image classification system includes: Feature extraction module, used to obtain zero-sample image data and extract data features through pre-trained network; The loss calculation module is used to select two images and input them into the encoder 1 to calculate the difference loss 1 of the results; input them into the decoder 1 to calculate the image reconstruction loss 1; The total loss calculation module is used to select an image input into encoder 1, input the semantic features into encoder 2 and decoder 2, calculate the difference loss 2, image reconstruction loss 2 and semantic reconstruction loss, and calculate the total loss; The sample testing module is used to generate new data after the encoder and decoder are trained, combine the new data with the original data to train the classifier, and use the classifier to test the test samples.
7. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the zero-sample image classification method in machine learning as described in any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the zero-shot image classification method in machine learning as described in any one of claims 1 to 5.
9. An information data processing terminal, characterized in that: The information data processing terminal is used to implement the zero-sample image classification system as described in claim 6.
Citation Information
Patent Citations
Zero-sample image classification method based on variational self-coding adversarial network
CN110580501A
Zero sample image recognition method and recognition device thereof, medium and computer terminal
CN114821196A