Method and apparatus for image classification
The use of ACGAN and BiGAN for image classification in industrial settings addresses the challenge of classifying rare abnormal states by automating feature generation and clustering, enhancing efficiency and accuracy without manual labeling.
Patent Information
- Application Number
- CN202080082181.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-01-23
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-01-23
AI Technical Summary
The prior art is difficult to efficiently classify images of abnormal states on production lines, and the process of generating standard rules is time-consuming and labor-intensive.
Image classification is performed by using a discriminator that generates adversarial network GAN. The discriminator and generator are trained by pre-training and iteratively, combining the bidirectional generative adversarial network BiGAN and Gaussian hybrid model GMM to automatically label and classify images.
Fast and automated image classification and anomaly detection are realized, reducing dependence on anomaly samples, improving efficiency and reducing the need for manual labeling.
Smart Images

Figure CN114746908B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to image processing technology, and more particularly, to a method and apparatus for image classification. Background Art
[0002] This section introduces aspects that may help to better understand the present disclosure. Therefore, the statements in this section should be read in this way and should not be construed as admitting what is in the prior art or what is not in the prior art.
[0003] With the development of industrial automation, image processing technology is widely used to assist the automated operation of products. For example, an image of a product can be automatically obtained through a production line, and then the current state of the product (such as being in a first state, or being in a second state, etc.) can be determined according to the result of processing the image. The next operation of the product can be carried out according to the current state of the product.
[0004] Generally, an operator or user of a production line needs to pre-train the apparatus for image processing in the production line. That is, the operator or user needs to provide different standard rules (such as a standard image or a set of standard parameters) for each state of the product. Then, during the production process, the obtained image of the real product is compared with such standard rules to determine the current state of the real product.
[0005] Conventionally, a large number of images have to be captured by humans and labeled with different states. Then, for each state of the product, different standard rules can be summarized by humans or machines according to the images labeled with different states. Summary of the Invention
[0006] This Summary of the Invention is provided to introduce a selection of concepts in a simplified form that will be further described in the detailed description below. This Summary of the Invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0007] As described above, a possible solution for image classification is to provide a model with different standard rules for each state of the product. However, it will be difficult to correctly classify some rarely shown abnormal states because it is almost impossible to obtain the required number of sample images to generate standard rules for such abnormal states. In addition, such a process of generating such standard rules is not efficient because obtaining and labeling sample images is very time-consuming and consumes human resources.
[0008] According to various embodiments of the present disclosure, an improved solution for a method and apparatus for image classification is provided, which can overcome one or more problems as described above.
[0009] According to a first aspect of the present disclosure, a method for image classification is provided. The method may include: receiving an image to be classified; inputting the image into a discriminator of a first generative adversarial network (GAN); and outputting a result indicating true and an index of a predetermined classification, or a result indicating false.
[0010] In an embodiment of the present disclosure, the method may further include: determining that the image belongs to the predetermined classification according to the result indicating true and the index of the predetermined classification.
[0011] In an embodiment of the present disclosure, the method may further include: determining that the image corresponds to an abnormal state according to the result indicating false.
[0012] In an embodiment of the present disclosure, the first GAN is an auxiliary classifier generative adversarial network (ACGAN).
[0013] In an embodiment of the present disclosure, the method may further include: pre-training the first GAN using a plurality of sample images with classification information.
[0014] In an embodiment of the present disclosure, pre-training the first GAN using a plurality of sample images with classification information may include a plurality of epochs, where each epoch includes: training the discriminator of the first GAN with a plurality of sample images with classification information and a plurality of noise images with classification information while freezing the generator of the first GAN; and training the generator of the first GAN with random noise and random classification information while freezing the discriminator of the first GAN, where the generator of the first GAN generates a plurality of noise images and classification information of the plurality of noise images; where the discriminator and the generator of the first GAN are iteratively trained; and where the number of iterations of the discriminator of the first GAN is greater than the number of iterations of the generator of the first GAN.
[0015] In an embodiment of the present disclosure, the method may further include: generating classification information of a plurality of sample images by using a second GAN.
[0016] In an embodiment of the present disclosure, the second GAN may be a bidirectional generative adversarial network (BiGAN).
[0017] In an embodiment of the present disclosure, generating the classification information of the plurality of sample images by using the second GAN includes: collecting the plurality of sample images without classification information; generating a latent space vector for each of the plurality of sample images by an encoder of the second GAN; clustering the plurality of sample images into at least one cluster based on the latent space vector of each of the plurality of sample images; and assigning a classification to each of the at least one cluster respectively.
[0018] In an embodiment of the present disclosure, based on the latent space vectors of each of a plurality of sample images, the plurality of sample images are clustered into at least one cluster by using a Gaussian mixture model GMM.
[0019] According to a second aspect of the present disclosure, there is provided an apparatus for image classification, including: a processor and a memory, the memory containing instructions executable by the processor, whereby the apparatus for image classification is operable to: receive an image to be classified; input the image into a discriminator of a first generative adversarial network GAN; and output a result indicating true and an index of a predetermined classification, or a result indicating false.
[0020] In an embodiment of the present disclosure, the apparatus may further be operable to: determine that the image belongs to the predetermined classification according to the result indicating true and the index of the predetermined classification.
[0021] In an embodiment of the present disclosure, the apparatus may further be operable to: determine that the image corresponds to an abnormal state according to the result indicating false.
[0022] In an embodiment of the present disclosure, the first GAN is an auxiliary classifier generative adversarial network ACGAN.
[0023] In an embodiment of the present disclosure, the apparatus may further be operable to: pre-train the first GAN using a plurality of sample images with classification information.
[0024] In an embodiment of the present disclosure, the apparatus for image classification may be operable to pre-train the first GAN with a plurality of sample images with classification information for a plurality of epochs, where each epoch may include: training the discriminator of the first GAN using a plurality of sample images with classification information and a plurality of noise images with classification information while freezing the generator of the first GAN; and training the generator of the first GAN with random noise and random classification information while freezing the discriminator of the first GAN, where the generator of the first GAN generates a plurality of noise images and classification information of the plurality of noise images; where the discriminator and the generator of the first GAN are iteratively trained; and where the number of iterations of the discriminator of the first GAN is greater than the number of iterations of the generator of the first GAN.
[0025] In an embodiment of the present disclosure, the apparatus may further be operable to: generate classification information of a plurality of sample images by using a second GAN.
[0026] In an embodiment of the present disclosure, the second GAN is a bidirectional generative adversarial network BiGAN.
[0027] In one embodiment of the present disclosure, the apparatus for image classification may be operable to: acquire a plurality of sample images without classification information; generate a latent space vector for each of the plurality of sample images through an encoder of a second GAN; cluster the plurality of sample images into at least one cluster based on the latent space vector of each of the plurality of sample images; and assign a classification to each of the at least one cluster, respectively.
[0028] In an embodiment of the present disclosure, based on the latent space vector of each of the plurality of sample images, the plurality of sample images may be clustered into at least one cluster by using a Gaussian mixture model (GMM).
[0029] According to a third aspect of the present disclosure, there is provided a computer-readable medium having instructions stored thereon that, when executed on at least one processor, cause the at least one processor to perform any one of the above methods.
[0030] According to various embodiments of the present disclosure, one or more advantages may be achieved. For example, it is also possible to identify images that do not belong to any predetermined classification (such as an image corresponding to an abnormal state) without the need to generate standard rules for such unclassified images. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The present disclosure itself, preferred modes of use, and further objects will be best understood when the following detailed description of the embodiments is read in conjunction with the accompanying drawings, in which:
[0032] Figure 1 A flowchart of a method for image classification according to an embodiment of the present disclosure is shown;
[0033] Figure 2 An exemplary flowchart of additional steps of the method as shown in Figure 1 is shown;
[0034] Figure 3 An exemplary block diagram of a first generative adversarial network (GAN) according to an embodiment of the present disclosure is shown;
[0035] Figure 4 An exemplary flowchart of other additional steps of the method as shown in Figure 1 is shown;
[0036] Figure 5 An exemplary flowchart of other additional steps of the method as shown in Figure 1 is shown;
[0037] Figure 6 A flowchart of a training phase and a detection phase in an application for a production line according to an embodiment of the present disclosure is shown;
[0038] Figure 7 Shows an exemplary framework of a BiGAN according to an embodiment of the present disclosure;
[0039] Figure 8 Shows an exemplary flowchart of training a BiGAN model according to an embodiment of the present disclosure;
[0040] Figure 9 Shows an exemplary flowchart of an automatic labeling process according to an embodiment of the present disclosure;
[0041] Figure 10 Shows multiple sample images of a plug connector;
[0042] Figure 11 Shows in Figure 10 the clustering results of multiple sample images;
[0043] Figure 12 Shows the classification of multiple sample images;
[0044] Figure 13 Shows an exemplary flowchart of training an ACGAN model according to an embodiment of the present disclosure;
[0045] Figure 14 Shows multiple sample images labeled as true for a plug connector;
[0046] Figure 15 Shows multiple sample images labeled as false for a plug connector;
[0047] Figure 16 Shows random vectors assigned with classes;
[0048] Figure 17 Shows an exemplary flowchart of a detection process according to an embodiment of the present disclosure;
[0049] Figure 18 Shows plug connector samples labeled by a BiGAN and a GMM according to an embodiment of the present disclosure;
[0050] Figure 19 Shows the anomaly detection results of an ACGAN according to an embodiment of the present disclosure;
[0051] Figure 20 Shows the classification results of an ACGAN for normal samples according to an embodiment of the present disclosure;
[0052] Figure 21 Shows an exemplary flowchart of a detection process using Instance 2 according to an embodiment of the present disclosure;
[0053] Figure 22FIG. 0 shows an exemplary block diagram of an apparatus 2200 for image classification according to an embodiment of the present disclosure;
[0054] Figure 23 FIG. 4 shows another exemplary block diagram of an apparatus 2200 for image classification according to an embodiment of the present disclosure; and
[0055] Figure 24 FIG. 8 shows an exemplary block diagram of a computer-readable medium 2400 according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0056] Embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the discussion of these embodiments is only for enabling those skilled in the art to better understand and implement the present invention, rather than imposing any limitation on the scope of the present invention. References to features, advantages, or similar language throughout this specification do not mean that all features and advantages achievable by the present disclosure are present in or should be present in any single embodiment of the present disclosure. More precisely, the language referring to such features and advantages is understood to mean that a particular feature, advantage, or characteristic described with respect to an embodiment is included in at least one embodiment of the present disclosure. Additionally, the described features, advantages, and characteristics of the present disclosure may be combined in any suitable manner in one or more embodiments. Those skilled in the relevant art will recognize that the present disclosure may be practiced without one or more of the specific features or advantages of a particular embodiment. In other instances, additional features and advantages may be recognized in certain embodiments that may not be present in all embodiments of the present disclosure.
[0057] As used herein, the terms “first,” “second,” etc. refer to different elements. The singular forms “a” and “an” are also intended to include the plural forms unless the context clearly dictates otherwise. As used herein, the terms “comprises,” “comprising,” “has,” “having,” “includes,” and / or “including” specify the presence of the stated features, elements, and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. The term “based on” should be understood as “at least partially based on.” The terms “one embodiment” and “an embodiment” should be understood as “at least one embodiment.” The term “another embodiment” should be understood as “at least one other embodiment.” Other definitions, whether explicit or implicit, may be included below.
[0058] Hereinafter, various embodiments of the present disclosure will be described with reference to the accompanying drawings.
[0059] As a specific example rather than a limitation, an image of a plug connector can be discussed below. A plug connector is a connector that connects a transceiver (TRX) board and a filter in a radio product (such as in a 5th generation (5G) radio system). In a 5G production line, a robot is used to install the plug connector in the product. The plug connector is placed in a tray, and the robot grabs the plug connector and faces it towards the camera. The recognition process will be triggered to detect the head side or the tail side and send a feedback signal to the robot to take appropriate actions. In reality, abnormal situations may occur. For example, the plug connector may not be placed in the correct position in the tray, so there is no plug connector at all in the image / photo obtained by the camera. The plug connector may not be in the correct form / state, so the image does not show a normal plug connector at all. The photo used for detection may be contaminated by noise and blurred. In these abnormal situations, the plug connector will be discarded, and thus the robot will take a new one to continue.
[0060] In the above process, the system not only needs to classify the plug connector head (i.e., the first type / status) and the tail (i.e., the second type / status) categories, but also needs to perform anomaly detection when the image is abnormal (such as no object, misalignment, noise contamination, blur, etc.).
[0061] There can be some feasible techniques to solve these 2 tasks (anomaly detection and classification).
[0062] Solution 1: Two models can be created separately for these 2 tasks. One is for anomaly detection, and the other is for classification.
[0063] Solution 2: One unsupervised learning model can be used for anomaly detection and classification. That is, more specifically, the category is determined according to the activation function output by the classifier.
[0064] If the probability of the tail class exceeds a threshold (e.g., 0.6), it is regarded as the tail class. If it is below the threshold (e.g., 0.4), it is regarded as the head class. If it is below the previous threshold (e.g., 0.6) but above another subsequent threshold (e.g., 0.4), it is considered an abnormal sample.
[0065] Solution 3: One supervised learning model is used to detect the plug connector head, the plug connector tail, or an anomaly. That is, a large number of samples of each type are collected to train the model, and then this model is used to classify the head, the tail, and the anomaly.
[0066] However, in Solution 1, using anomaly detection and classification models separately will slow down the recognition process. In Solution 2, the classification probabilities of some abnormal samples (even very different from normal samples) are still very high, even higher than 99%, so the accuracy of this method is very poor. In Solution 3, when using a supervised learning model for image classification, manual labeling of training data is required. Doing such data labeling requires considerable effort and is time-consuming. When using a supervised learning model for anomaly detection, a large number of anomalies need to be collected as the training data set, but in the real world, it is very difficult to collect so many anomalies because this is a rare event. In addition, regardless of which solution is applied, the head or tail should be labeled. There is a lot of labeling work.
[0067] Figure 1 FIG. shows an exemplary flowchart of a method for image classification according to an embodiment of the present disclosure.
[0068] As Figure 1 shown, method 100 may include: S101, receiving an image to be classified; S102, inputting the image into a discriminator of a first generative adversarial network (GAN); and S103, outputting a result indicating true and an index of a predetermined classification, or a result indicating false.
[0069] According to Figure 1 the method shown, one or more advantages can be achieved. For example, the discriminator of the first GAN can output a result indicating true or false in one network / model, instead of implementing two or more tasks / models separately. The detection time will be reduced, and there is no need to collect abnormal samples that are more difficult to obtain than normal samples. The efficiency can also be improved.
[0070] It should also be noted that according to an embodiment of the present disclosure, images of any industrial product other than plug connectors in any production line can be applied.
[0071] Figure 2 FIG. illustrates an exemplary flowchart of additional steps of the method as Figure 1 shown according to an embodiment of the present disclosure.
[0072] As Figure 2 shown, method 100 may further include: S104, determining that the image belongs to a predetermined classification according to the result indicating true and the index of the predetermined classification.
[0073] In addition, method 100 may further include: S105, determining that the image corresponds to an abnormal state according to the result indicating false.
[0074] According to Figure 2 the method shown, the classification of the image and / or whether the image is abnormal will be directly and clearly determined according to the output of the discriminator of the first GAN.
[0075] Figure 3 Shows an exemplary block diagram of a first generative adversarial network (GAN) according to an embodiment of the present disclosure.
[0076] As Figure 3 shown, the first GAN can be an auxiliary classifier generative adversarial network (ACGAN).
[0077] The ACGAN can at least include: a generator G 21, and a discriminator D 22. The input of the generator 21 can include the class information C (class) 201 of the image and the noise Z (noise) 202. The input of the discriminator 22 can include: the real data X 真 (data) 203, and the fake data X 假 204 generated by the generator 21. The output 205 of the discriminator 22 can include a result indicating true and an index of a predetermined classification, or a result indicating false.
[0078] Figure 4 Shows an exemplary flowchart of other additional steps of the method as Figure 1 shown therein according to an embodiment of the present disclosure.
[0079] As Figure 4 shown, the method 100 can further include: S106, pre-training the first GAN using a plurality of sample images with classification information.
[0080] Specifically, in one embodiment of the present disclosure, S106, pre-training the first GAN using a plurality of sample images with classification information can include a plurality of epochs S107, where each epoch S107 includes: S108, training the discriminator of the first GAN using a plurality of sample images with classification information and a plurality of noise images with classification information, while freezing the generator of the first GAN; and S109, training the generator of the first GAN using random noise and random classification information, while freezing the discriminator of the first GAN. The generator of the first GAN generates a plurality of noise images and the classification information of the plurality of noise images. Iteratively train the discriminator and the generator of the first GAN. In addition, the number of iterations of the discriminator of the first GAN is greater than the number of iterations of the generator of the first GAN.
[0081] That is, S108 and S109 are iterated for several epochs. Therefore, the discriminator (discriminative model) and the generator (generative model) are alternately trained.
[0082] According to an embodiment of the present disclosure, the number of iterations of training the discriminator can be greater than the number of iterations of the generator to make the discriminator (which will be used to classify images) more robust. For example, the ratio of the number of iterations of training the discriminator to the number of iterations of training the generator can be 10:1.
[0083] The generator of the first GAN can directly generate a noise image, or add random noise to an existing real image to obtain a noise image. The random noise can include a random noise vector containing multiple dimensions (parameters) to change the real features of the existing real image. For example, the random noise vector can include 100 dimensions. In addition, the operator / user can also directly configure some abnormal images as part of the noise image.
[0084] With such a configuration, the discriminator and generator of the first GAN can be automatically pre-trained. The efficiency can be improved.
[0085] Figure 5 An exemplary flowchart showing additional steps of the method as shown in Figure 1 is shown according to an embodiment of the present disclosure.
[0086] As Figure 5 shown, method 100 may further include: S110, generating classification information of the plurality of sample images by using a second GAN.
[0087] In an embodiment of the present disclosure, generating classification information of the plurality of sample images by using a second GAN may include: S111, collecting a plurality of sample images without classification information; S112, generating a latent space vector for each of the plurality of sample images by an encoder of the second GAN; S113, clustering the plurality of sample images into at least one cluster based on the latent space vector of each of the plurality of sample images; S114, assigning a classification to each of the at least one cluster.
[0088] In an embodiment of the present disclosure, the second GAN may be a bidirectional generative adversarial network BiGAN. In addition, in an embodiment of the present disclosure, based on the latent space vector of each of the plurality of sample images, the plurality of sample images are clustered into at least one cluster by using a Gaussian mixture model GMM. That is, a combination of BiGAN+GMM can be used.
[0089] Note that any other type of GAN can also be used as long as it can generate a latent space vector for each of the plurality of sample images. In addition, any other type of model can also be used as long as it can cluster the plurality of sample images based on the latent space vector.
[0090] According to an embodiment of the present disclosure, a large number of manual markings of the sample images are not required. And there is no need to make handcrafted features for different classifications of the samples.
[0091] The embodiment of the present disclosure can be further illustrated by taking the image of a plug connector as an example again.
[0092] Figure 6 Shows a flowchart of the training phase and the detection phase in an application for a production line / product line according to an embodiment of the present disclosure.
[0093] In the training phase S601, both the BiGAN and the ACGAN are trained.
[0094] In S603, samples (sample images) can be collected as training data by a camera 61 installed on the production line.
[0095] In S604, a BiGAN model is constructed and trained to encode images.
[0096] In S605, a GMM model is established to classify images for automatic labeling. Using this method, there is no need to label as head or tail.
[0097] In S606, an ACGAN is constructed and trained to detect the head, tail, or abnormality of a plug connector.
[0098] In S607, the trained models are stored as files for loading and execution in the detection phase.
[0099] In S608 of the detection phase S602, the discriminative model of the ACGAN is applied, that is, each image 63 taken by the camera 61 will be detected by the discriminator of the ACGAN.
[0100] In S609, if the probability that the image 63 is a real image (output from the ACGAN) is lower than a threshold, it is considered an abnormal sample and should be discarded in S610.
[0101] For normal samples, if the probability that the image 63 is of the head class is greater than that of the tail class, it is classified as the head class, otherwise it is classified as the tail class in S611. The classification of being the head or the tail will be output to an actuator 62 (such as a robotic arm) on the production line.
[0102] Figure 7 Shows an exemplary framework of a BiGAN according to an embodiment of the present disclosure.
[0103] The overall model is as Figure 7 shown. In addition to the generator G from the standard GAN framework, the BiGAN also includes an encoder E that maps the data x to a latent representation z. The BiGAN discriminator D discriminates not only in the data space (x vs. G(z)), but also jointly in the data space and the latent space (the tuple (x, E(x)) vs. (G(z), z)), where the latent component is the encoder output E(x) or the generator input z. P(y) represents the likelihood that the tuple comes from (x, E(x)).
[0104] The key of the present disclosure lies in using the encoder E(x) to obtain the latent space of the image.
[0105] First, the objective function can be described as follows:
[0106]
[0107] where
[0108]
[0109] As a non - restrictive example, the function in the 2016 arXiv preprint arXiv:1605.09782 by J. Donahue, P. Kr¨ahenb¨uhl, and T. Darrell, “Adversarial Feature Learning” can be used.
[0110] In the above function, D, E, and G represent the discriminator, encoder, and generator respectively. x ∼ px represents the distribution from real data. represents the expected value of logD(x, E(x)). represents the logarithm of the likelihood that the tuple comes from (x, E(x)). represents the expected value of log(1 - D(G(z), z)). represents the logarithm of the likelihood that the tuple comes from (G(z), z).
[0111] Figure 8 FIG. shows an exemplary flowchart for training a BiGAN model according to an embodiment of the present disclosure.
[0112] The generation and encoder models are trained to minimize V, while the discriminative model is trained to maximize V. The training process can be described as follows:
[0113] Step S801: Collect samples. No manual labeling is required
[0114] Steps S802 - S803: Freeze the generation and encoder models and only train the discriminative model.
[0115] The input of the discriminative model consists of two parts: the plug - in joint image from the real dataset (“true” class) and its code (obtained from the encoder); and the plug - in joint image generated by the generation model (“false” class) and the corresponding random noise vector. The output is whether it is true or false. The discriminative model is trained to maximize the objective function V.
[0116] Steps S804 - S807: Freeze the discriminative model and only train the generation and encoder models.
[0117] In S804: Generate a batch of noise vectors with uniform distribution (e.g., 2 dimensions).
[0118] In S805: Train the generative model to minimize V.
[0119] In S806: Encode the images from the real dataset.
[0120] In S807: Train the encoder model to minimize V.
[0121] Steps S802 - S803 and steps S804 - S807 are iterated for several epochs. This means that the discriminative model and the generative and encoder models are trained alternately.
[0122] Figure 9 An exemplary flowchart of the automatic labeling process according to an embodiment of the present disclosure is shown.
[0123] In S901: Retrieve a large number of plug - in connector images from the camera.
[0124] In S902: Use the encoder of BiGAN to obtain the latent space vector of each image.
[0125] In S903: Perform clustering using GMM through the encoder values.
[0126] In S904: Assign classes to each cluster.
[0127] In S905: Send the samples with confidence higher than the threshold to ACGAN.
[0128] Figure 10 Multiple sample images of the plug - in connector are shown. Figure 11 Shows Figure 10 the clustering results of multiple sample images. Figure 12 Shows the classification of multiple sample images.
[0129] As Figure 10 shown, these sample images are obtained without any regular pattern. As Figure 11 shown, for these sample images, two clusters 1101 and 1102 are obtained. As Figure 12 shown, the upper row shows the sample images of the first class (such as the head), and the lower row shows the sample images of the second class (such as the tail).
[0130] Figure 13 An exemplary flowchart of training the ACGAN model according to an embodiment of the present disclosure is shown.
[0131] The overall model is as Figure 13As shown. For the generator, the input is a random point from the latent space and a class label, and the output is the generated image. For the discriminator, the input is an image, and the output is the probability of "true" and the probability that the image belongs to each known class.
[0132] First, the objective function is described as follows. The objective function has two parts: the logarithm of the likelihood of the correct source Ls and the logarithm of the likelihood of the correct class Lc.
[0133] L S = E[log P(S = true|X 真 )] + E[log P(S = false|X 假 )]
[0134] L C = E[log P(C = c|X 真 )] + E[log P(C = c|X 假 )]
[0135] For example but not limited to, the function in "Conditional Image Synthesis with Auxiliary Classifier GAN" by A. Odena, C. Olah, and J. Shlens in arXiv:1610.09585 in 2016 can be utilized.
[0136] In the above function, E[] represents the expectation of the function described in []. log P(S = true|X 真 ) represents the logarithm of the probability that the image comes from a real image. log P(S = false|X 假 ) represents the logarithm of the probability that the image comes from the generator. logP(C = c|X 真 ) represents the logarithm of the probability that the class information comes from a real image. log P(C = c|X 假 ) represents the logarithm of the probability that the class information comes from a generated image.
[0137] The generative model is trained to maximize Lc - Ls, while the discriminative model is trained to maximize Lc + Ls. The training process is described as follows:
[0138] Step S1301: Collect samples and automatically label them as the head class or the tail class (through BiGAN and GMM). There is no need to collect abnormal samples.
[0139] Steps S1302 - S1303: Freeze the generative model and only train the discriminative model.
[0140] The input of the discriminative model consists of two parts: the plug - in joint image from the real dataset ("true" class); and the plug - in joint image generated by the generative model ("false" class). The output is its corresponding label and true or false information. The discriminative model is trained to maximize the objective function Lc + Ls.
[0141] Steps S1304 - S1305: Freeze the discriminative model and only train the generative model.
[0142] In step S1034: Generate a batch of noise vectors with a uniform distribution (e.g., 100 - dimensional). Randomly assign the classes of head or tail to these vectors.
[0143] In step S1035: Train the generative model to maximize the objective function Lc - Ls.
[0144] Steps S1302 - S1303 and steps S1304 - S1305 are iterated for several epochs. This means that the discriminative model and the generative model are trained alternately. In particular, to make the discriminator more robust, the number of iterations for training the discriminator is configured to be greater than that of the generator.
[0145] Figure 14 Shows multiple sample images of the plug - in connector marked as true. Figure 15 Shows multiple sample images of the plug - in connector marked as false. Figure 16 Shows the random vectors assigned with classes.
[0146] As Figure 14 shown, the upper row shows the real sample images marked as the head of the plug - in connector, and the lower row shows the real sample images marked as the tail of the plug - in connector. As Figure 15 shown, the upper row shows the fake / noise images marked as the head of the plug - in connector, and the lower row shows the fake / noise images marked as the tail of the plug - in connector. The fake / noise images can be generated by adding noise to the existing real samples. As Figure 16 shown, the upper vectors are assigned the class of the head, while the lower vectors are assigned the class of the tail. The vectors can have a dimension of 100.
[0147] Figure 17 Shows an exemplary flowchart of the detection process according to an embodiment of the present disclosure.
[0148] In Figure 17 three typical samples (head, tail, anomaly) are shown. As Figure 17 shown, when a plug - in connector image appears in S1701, the discriminative model is used to detect the image in S1702. The output has two parts: the probability of the real image, and the probability of each class (head and tail). If the probability of the real image is lower than the threshold in S1703, the plug - in connector should be regarded as an anomaly and discarded in S1704. Otherwise, it is classified as a head or a tail according to the probabilities of these two classes in S1705.
[0149] More examples of the image results / outputs of some of the above steps will be further shown below.
[0150] Figure 18 Shows the plug connector samples labeled by BiGAN and GMM according to an embodiment of the present disclosure.
[0151] Collect plug connector samples and use a BiGAN encoder and GMM to achieve clustering. Each cluster is labeled as head or tail. As Figure 18 shown, samples with a confidence higher than the threshold will be sent to the ACGAN for training.
[0152] The top two rows of samples are classified as heads. The bottom two rows of samples are classified as tails.
[0153] Figure 19 Shows the anomaly detection results of the ACGAN according to an embodiment of the present disclosure.
[0154] After completing the ACGAN training, the images captured by the camera are input into the ACGAN. As Figure 19 shown, anomaly detection is achieved using the discriminant model, so as to select the anomaly samples (lower group) from all the captured samples (upper group). The other samples are considered normal.
[0155] Figure 20 Shows the classification results of normal samples by the ACGAN according to an embodiment of the present disclosure.
[0156] As Figure 20 shown, for normal samples, they are classified as heads or tails. The upper row is the head class, and the lower row is the tail class.
[0157] Achieving classification and anomaly detection simultaneously in this way has a wide range of applications. Two usage examples can be further cited as follows:
[0158] Usage example 1: Optical Character Recognition (OCR). Each character should be recognized by a pattern recognition algorithm. Sometimes, characters not in the dataset are encountered. In this case, the algorithm should feedback to the machine learning engineer to label this specific character.
[0159] Usage example 2: Putty recognition. Putty is used as a heat sink in heat generating components. In this case, the putty shapes need to be classified into different categories such as square, round, bar, etc. Occasionally, abnormal shapes are generated due to equipment failures. This means that not only the specific shapes need to be classified, but also an alarm needs to be issued when anomaly samples occur.
[0160] Figure 21 Shows an exemplary flowchart of the detection process of usage example 2 according to an embodiment of the present disclosure.
[0161] As Figure 21As shown, in S1704, an abnormal putty shape is selected, and in S1705, the normal putty shapes are further classified.
[0162] According to an embodiment of the present disclosure, using an ACGAN to implement anomaly detection and classification within a network can reduce processing time and does not require collecting anomaly samples.
[0163] Figure 22 An exemplary block diagram of an apparatus 2200 for image classification according to an embodiment of the present disclosure is shown.
[0164] As Figure 22 shown, the apparatus 2200 may include: a processor 2201 and a memory 2202, the memory containing instructions executable by the processor, whereby the apparatus for image classification is operative to: receive an image to be classified; input the image into a discriminator of a first generative adversarial network GAN; and output a result indicating true and an index of a predetermined classification, or a result indicating false.
[0165] In an embodiment of the present disclosure, the apparatus may further be operative to: determine that the image belongs to a predetermined classification according to the result indicating true and the index of the predetermined classification.
[0166] In an embodiment of the present invention, the apparatus may further be operative to: determine that the image corresponds to an abnormal state according to the result indicating false.
[0167] In an embodiment of the present disclosure, the first GAN is an auxiliary classifier generative adversarial network ACGAN.
[0168] In an embodiment of the present disclosure, the apparatus may further be operative to: pre-train the first GAN using a plurality of sample images with classification information.
[0169] In an embodiment of the present disclosure, the apparatus for image classification may be operative to pre-train the first GAN for a plurality of epochs using a plurality of sample images with classification information, where each epoch may include: training the discriminator of the first GAN using a plurality of sample images with classification information and a plurality of noise images with classification information while freezing the generator of the first GAN; and training the generator of the first GAN using random noise and random classification information while freezing the discriminator of the first GAN, where the generator of the first GAN generates a plurality of noise images and classification information of the plurality of noise images; where the discriminator and the generator of the first GAN are iteratively trained; and where the number of iterations of the discriminator of the first GAN is greater than the number of iterations of the generator of the first GAN.
[0170] In one embodiment of the present disclosure, the apparatus may be further operable to: generate classification information of a plurality of sample images by using a second GAN.
[0171] In one embodiment of the present disclosure, the second GAN is a bidirectional generative adversarial network BiGAN.
[0172] In one embodiment of the present disclosure, the apparatus for image classification may be operable to: collect a plurality of sample images without classification information; generate a latent space vector for each of the plurality of sample images by an encoder of a second GAN; cluster the plurality of sample images into at least one cluster based on the latent space vector of each of the plurality of sample images; and assign a classification to each of the at least one cluster, respectively.
[0173] In an embodiment of the present disclosure, the plurality of sample images may be clustered into at least one cluster by using a Gaussian mixture model GMM based on the latent space vector of each of the plurality of sample images.
[0174] The processor 2201 may be any kind of processing component, such as one or more microprocessors or microcontrollers, and other digital hardware, which may include a digital signal processor (DSP), dedicated digital logic, etc. The memory 2202 may be any kind of storage component, such as a read-only memory (ROM), random access memory, cache memory, flash memory device, optical storage device, etc.
[0175] Figure 23 Another exemplary block diagram of an apparatus 2200 for image classification according to an embodiment of the present disclosure is shown.
[0176] As Figure 23 shown, the apparatus 2200 may include: a receiving unit 2310 configured to receive an image to be classified; an input unit 2320 configured to input the image into a discriminator of a first generative adversarial network GAN; and an output unit 2330 configured to output a result indicating true and an index of a predetermined classification, or a result indicating false.
[0177] The term unit may have a conventional meaning in the field of electronics, electrical equipment, and / or electronic devices, and may include, for example, electrical and / or electronic circuits, devices, modules, processors, memories, logical solid-state and / or discrete devices, computer programs or instructions for performing corresponding tasks, programs, calculations, output, and / or display functions, etc., as described herein.
[0178] Through these units, the apparatus for image classification can be free from a fixed processor or memory, and can configure any computing resources and storage resources from at least one network node / device / entity / apparatus in a communication system. Virtualization technology and network computing technology can be further introduced to improve the utilization efficiency and flexibility of resources.
[0179] Figure 24 An exemplary block diagram of a computer-readable medium 2400 in accordance with an embodiment of the present disclosure is shown.
[0180] As Figure 24 shown, the computer-readable medium 2400 may have instructions (i.e., software, programs, etc.) stored thereon, which when executed on at least one processor, cause the at least one processor to perform any one of the above methods.
[0181] The computer-readable storage medium 2400 may be configured to include a memory, such as RAM, ROM, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic disk, optical disk, floppy disk, hard disk, removable cartridge tape, or flash drive, etc.
[0182] According to various embodiments of the present disclosure, one or more advantages can be achieved. For example, it is also possible to identify an image that does not belong to any predetermined classification (such as an image corresponding to an abnormal state) without the need to generate a standard rule for such unclassified images.
[0183] Generally, the various exemplary embodiments may be implemented in hardware or a dedicated chip, circuit, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software executable by a controller, microprocessor, or other computing device, but the present disclosure is not limited thereto. Although the various parts of the exemplary embodiments of the present disclosure may be shown and described in block diagrams, flowcharts, or using some other view, it is well understood that the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general hardware or a controller or other computing device, or some combination thereof, as non-limiting examples.
[0184] Therefore, it should be understood that at least some aspects of the exemplary embodiments of the present disclosure may be implemented in various components such as integrated circuit chips and modules. Therefore, it should be understood that the exemplary embodiments of the present disclosure may be implemented in a device implemented as an integrated circuit, where the integrated circuit may include circuits (and possibly firmware) for implementing at least one or more of a data processor, a digital signal processor, a baseband circuit, and a radio frequency circuit, and may be configured to operate according to the exemplary embodiments of the present disclosure.
[0185] It should be understood that at least some aspects of the exemplary embodiments of the present disclosure may be implemented in computer-executable / readable instructions executed by one or more computers or other devices, such as in one or more program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., which, when executed by a processor in a computer or other device, perform specific tasks or implement specific abstract data types. The computer-executable instructions may be stored on a computer-readable medium such as a hard disk, optical disk, removable storage medium, solid-state memory, random access memory (RAM), etc. As those skilled in the art will understand, in various embodiments, the functions of the program modules may be combined or distributed as needed. Additionally, the functions may be implemented in whole or in part in firmware or hardware equivalents such as integrated circuits, field-programmable gate arrays (FPGAs), etc.
[0186] The present disclosure includes any novel feature or combination of features explicitly disclosed herein, or any generalization thereof. Various modifications and adaptations of the foregoing exemplary embodiments of the present disclosure will be apparent to those skilled in the relevant art when read in conjunction with the accompanying drawings. However, any and all modifications will still fall within the scope of the non-limiting and exemplary embodiments of the present disclosure.
Claims
1. A method for image classification, comprising: Receiving (S101) an image of an industrial product to be classified; Inputting (S102) the image into a discriminator of a first generative adversarial network GAN; Outputting (S103) a result indicating true and an index of a predetermined classification, or a result indicating false; Determining (S104) that the image belongs to the predetermined classification according to the result indicating true and the index of the predetermined classification; And Determining (S105) that the image corresponds to an abnormal state according to the result indicating false; Pre-training (S106) the first GAN using a plurality of sample images with classification information; Generating (S110) classification information of the plurality of sample images by using a second GAN, where the second GAN is a bidirectional generative adversarial network BiGAN; Wherein generating the classification information of the plurality of sample images by using the second GAN includes: Collecting (S111) the plurality of sample images without classification information; Generating (S112) a latent space vector for each of the plurality of sample images by an encoder of the second GAN; Clustering (S113) the plurality of sample images into at least one cluster based on the latent space vector of each of the plurality of sample images; and Assigning (S114) a classification to each of the at least one cluster respectively.
2. The method according to claim 1, wherein pre-training the first GAN using a plurality of sample images with classification information includes a plurality of epochs (107), and each epoch includes: Training (108) the discriminator of the first GAN with the plurality of sample images with classification information and a plurality of noise images with classification information while freezing the generator of the first GAN; And Training (109) the generator of the first GAN with random noise and random classification information while freezing the discriminator of the first GAN, where the generator of the first GAN generates the plurality of noise images and the classification information of the plurality of noise images; Wherein the discriminator and the generator of the first GAN are iteratively trained; And Wherein the number of iterations of the discriminator of the first GAN is greater than the number of iterations of the generator of the first GAN.
3. The method according to claim 1, wherein clustering the plurality of sample images into at least one cluster by using a Gaussian mixture model GMM based on the latent space vector of each of the plurality of sample images.
4. An apparatus for image classification, comprising: A processor (2201) and a memory (2202), the memory containing instructions executable by the processor, whereby the apparatus for image classification is operable to: Receive an image of an industrial product to be classified; Input the image into a discriminator of a first generative adversarial network GAN; Output a result indicating true and an index of a predetermined classification, or a result indicating false; Determine that the image belongs to the predetermined classification according to the result indicating true and the index of the predetermined classification; And Determine that the image corresponds to an abnormal state according to the result indicating false; Pre-train a first GAN using multiple sample images with classification information; Generate classification information for the multiple sample images by using a second GAN, where the second GAN is a bidirectional generative adversarial network BiGAN; wherein the apparatus for image classification is operable to: Collect the multiple sample images without classification information; Generate a latent space vector for each of the multiple sample images by an encoder of the second GAN; Cluster the multiple sample images into at least one cluster based on the latent space vector of each of the multiple sample images; and Assign a classification to each of the at least one cluster, respectively.
5. The apparatus according to claim 4, wherein the apparatus for image classification is operable to pre-train the first GAN with multiple sample images with classification information for multiple epochs, where each epoch includes: Train a discriminator of the first GAN using the multiple sample images with classification information and multiple noise images with classification information, while freezing a generator of the first GAN; and Train the generator of the first GAN using random noise and random classification information, while freezing the discriminator of the first GAN, where the generator of the first GAN generates the multiple noise images and classification information of the multiple noise images; wherein the discriminator and the generator of the first GAN are iteratively trained; and wherein the number of iterations of the discriminator of the first GAN is greater than the number of iterations of the generator of the first GAN.
6. The apparatus according to claim 4, wherein based on the latent space vector of each of the multiple sample images, the multiple sample images are clustered into at least one cluster by using a Gaussian mixture model GMM.
7. A computer-readable medium (2400) having instructions (2401) stored thereon, which when executed on at least one processor, cause the at least one processor to perform the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Method and system for extracting ground object spatial spectral features of hyperspectral remote sensing image
CN108764005A
Polarized SAR image classification method based on ACGAN
CN109784401A