An oversampling method based on a multi-false-class generative adversarial network

By introducing a multi-false class generation mechanism and a conditional variational autoencoder into the generative adversarial network, and combining Focal Loss and gradient penalty terms, the problem of similar samples and low quality generation of samples based on GAN is solved, and higher quality minority class sample generation and better classification performance are achieved.

CN114004333BActive Publication Date: 2025-06-17GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111251327.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-26
Publication Date
2025-06-17
Estimated Expiration
2041-10-26

AI Technical Summary

Technical Problem

When generating a few samples, the GAN-based oversampling method is too similar and has low quality, making it difficult to improve the accuracy of a few samples without reducing the accuracy of most samples.

Method used

Multiple false class generative adversarial network (MFCGAN) is used to train in combination with conditional variational autoencoder. By introducing Focal Loss and gradient penalty terms, the training stability of the generator and the quality of generated samples are improved.

Benefits of technology

The generated image quality is significantly improved, allowing the classifier to improve its performance in unbalanced scenarios, and the generated few samples are more diverse and accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114004333B_ABST
    Figure CN114004333B_ABST
Patent Text Reader

Abstract

An oversampling method based on a multi-false-class generative adversarial network. First, learn the features of unbalanced image data through a conditional variational autoencoder; then use the decoder in the variational autoencoder to initialize the generator in the GAN to help the discriminator better determine the class distribution; then use the unbalanced dataset to train the multi-false-class generative adversarial network, and at the same time add a gradient penalty term to the multi-false-class generative adversarial network to improve the stability of GAN training and ensure the diversity of sample generation; replace the classification loss in the multi-false-class generative adversarial network with the focal loss to make the GAN focus more on those samples that are difficult to classify correctly during training; finally, input the minority class labels and random noise into the trained GAN model to generate high-quality minority class samples. The present invention can effectively generate high-quality samples for the minority classes in the unbalanced image dataset, make the data a balanced dataset, and help the classifier improve the classification performance in unbalanced scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of imbalance learning and deep learning, and particularly relates to an oversampling method based on a multi-false-class generative adversarial network. Background Art

[0002] With the rapid development of information technology, data in various fields is being generated, widely collected, and stored at an unprecedented speed. However, in real life, data in each field has the characteristic of data imbalance, such as medical diagnosis, telecom fraud detection, etc. Data imbalance means that in a dataset, the number of samples of one or more classes is much smaller than that of other classes. In fact, these minority-class samples often have a higher misclassification cost. For example, in fraud email detection, fraud emails often only account for 1% of all emails. However, once fraud interception fails, it often causes very serious losses to users. Therefore, in such problems, we often hope to accurately identify those minority-class samples.

[0003] When dealing with imbalanced data classification, traditional classification methods usually assume that the data class distribution is balanced and the misclassification cost is equal. When using traditional classification algorithms to process imbalanced data, due to the tilt in the number of majority classes and minority classes, aiming at the maximum overall classification accuracy will make the classification model tend to the majority class and ignore the minority class, resulting in a lower classification accuracy for the minority class. Around how to improve the accuracy of minority-class samples without reducing the accuracy of the majority class, many researchers have proposed oversampling algorithms based on differences to synthesize data for the minority class and make the dataset a balanced dataset. However, when these method models face image datasets, their performance often fails to reach the expected goal.

[0004] Therefore, Douzas et al. proposed using a conditional generative adversarial network (CGAN) to add label information to the input of the generator to generate minority-class samples. Experiments have shown that when dealing with high-dimensional image data, the oversampling method based on GAN often has better results than traditional methods. However, when the oversampling method based on GAN faces highly imbalanced image data, there are often problems such as single or incorrect generated samples.

[0005] Conditional Generative Adversarial Network:

[0006] Conditional Generative Adversarial Network (CGAN) is a generative model widely used in the field of deep learning in recent years. Its core idea comes from game theory. The GAN model contains two networks, the Generator and the Discriminator. The Generator G generates data that is expected to be indistinguishable from real samples based on a set of noises, while the purpose of the Discriminator is to determine whether the samples come from real samples or samples generated by the Generator. The training process of GAN is a process of continuous learning between G and D. Finally, G has fully learned the distribution of real samples. At this time, D can no longer distinguish between true and false samples, and the training reaches perfection. CGAN introduces a label variable y during GAN training. In this way, when generating data with G, only by mixing the label into the noise can data of a specified category be generated.

[0007] MFCGAN:

[0008] Multiple Fake Class Generative Adversarial Network (MFCGAN) can be regarded as an improvement of CGAN. Based on CGAN, a classifier is additionally introduced to the last layer of the Discriminator in GAN. This classifier will classify the samples into different true and false classes, such as true class 1, false class 1, true class 2, false class 2, and so on. Therefore, the samples generated by the Generator not only need to be judged as true, but also need to be classified into the corresponding classes by the classifier. In this way, MFCGAN can greatly improve the quality of the samples generated by the Generator for a specific category and ensure that the Generator does not generate samples of the wrong category.

[0009] Focal Loss:

[0010] Focal Loss is modified based on the standard cross-entropy loss. The Focal loss function can make the model focus more on difficult-to-classify samples during training by reducing the weights of easy-to-classify samples. Focal Loss can be defined as:

[0011]

[0012] Among them, both α and γ are adjustable hyperparameters. y' is the model prediction, and its value ranges from (0 - 1).

[0013] Technical solution: In order to solve the problem that the currently existing GAN-based oversampling method generates overly similar (single) and low-quality samples when generating minority-class samples, the present invention provides an oversampling method based on a multi-false-class generative adversarial network, and the specific steps are as follows:

[0014] (1) Obtain an imbalanced image dataset, convert the image size, and normalize the data so that the input pictures of the network have the same size;

[0015] (2) Build and train a conditional variational autoencoder model;

[0016] (3) Use the decoder in the variational autoencoder to initialize the weights of the generator, and build a generative adversarial network model;

[0017] (4) Train the generative adversarial network model;

[0018] (5) Use the minority-class labels and random noise as the network input to generate minority-class samples corresponding to the labels.

[0019] Further, the specific steps of the data preprocessing in step (1) are as follows:

[0020] (1.1) Obtain the imbalanced image dataset as X = {X1, X2, …, X n}, where n represents the maximum number of images;

[0021] (1.2) Obtain the imbalanced image label set as Y = {Y1, Y2, …, Y n}, where n represents the maximum number of images;

[0022] (1.3) Convert the size of the images to 64x64x number of channels;

[0023] (1.4) Normalize the image values and scale them to between [0, 1];

[0024] Further, the specific steps of building and training the conditional variational autoencoder model in step (2) are as follows:

[0025] (2.1) Use a 4-layer convolutional network to build a decoder, with the convolutional kernel sizes being 64, 128, 128, and 256 respectively. Use the LeakyReLU activation function between layers, then use the Flatten layer to expand, and finally output the dimensions of the mean and variance;

[0026] (2.2) Use a 4-layer convolutional network to build an encoder, with the convolutional kernel sizes being 256, 128, 128, and 64 respectively. Use the LeakyReLU activation function between layers, and output the dimensions of the generated pictures;

[0027] (2.3) Use the embedding as the embedding layer model, input the label dimension, and output the latent space of the corresponding category.

[0028] (2.4) Use the image input decoder. After the decoder outputs the mean and variance, calculate the noise according to the mean and variance, and embed the category label into the noise using the embedding.

[0029] (2.5) Use the noise embedded with the label information as the input of the decoder, and finally output the generated image.

[0030] (2.6) Use the mean square error between the generated image and the original image as the reconstruction loss, and the error between the noise and the standard normal distribution as the KLD loss. Use the sum of the KLD loss and the reconstruction loss to train the variational autoencoder.

[0031] (2.7) Set Adam as the optimizer, set the number of iterations to 30, and start training.

[0032] Further, the specific steps of building the generative adversarial network model in step (3) are as follows:

[0033] (3.1) Build the generator model in the GAN, and then initialize the weights of the generator using the decoder in the variational autoencoder. The structure of the generator is set to be exactly the same as that of the decoder.

[0034] (3.2) Build the discriminator model in the GAN. The number of convolutional kernels in the middle layer is 64, 128, 128, and 256 respectively. After the last layer of the middle layer, it is separated into two output layers, which respectively output the sample validity score and the category of the sample.

[0035] Further, the specific steps of training the generative adversarial network in step (4) are as follows:

[0036] (4.1) Use the generator to generate random label images, and then input the real image dataset, the generated image dataset and their labels into the discriminator, and respectively obtain the classification output and the validity value output. Use the focal loss to calculate the classification output to obtain the classification loss cls_loss, and obtain the adversarial loss adv_loss according to the validity score.

[0037] (4.2) Randomly sample a random value α from the normal distribution, subtract the real image from the generated image to obtain the interpolation diff, and calculate the interpolation: real image + α x diff. Input the interpolation and the label into the discriminator and obtain the gradient grads, and finally obtain the gradient penalty (where λ is the penalty coefficient and norm is the arithmetic square root of grads)

[0038] (4.3) Add the adv_loss, cls_loss, and gp together with weights to obtain the total loss value, and update the discriminator network weight parameters according to the loss.

[0039] (4.4) Use the generator to generate a random label image, input it into the discriminator to obtain the classification output and the valid value output respectively. Obtain the classification loss cls_loss and the adversarial loss adv_loss in the same way as (4.1). Sum the losses with weights and update the generator network weight parameters.

[0040] (4.5) Use Adam as the optimizer, set the number of iterations to 50, and start training.

[0041] Furthermore, the specific steps of using the minority class label and random noise as the network input in step (5) to generate the minority class samples corresponding to the label are as follows:

[0042] (5.1) Calculate the number of minority class samples to be generated according to the difference between the number of samples in each small class and the number of samples in the large class in the data.

[0043] (5.2) Use the required amount of noise and the minority class label as the input of the generator to generate the minority class samples corresponding to the label.

[0044] (5.3) Add the minority class samples to the original image dataset until a balanced dataset is formed;

[0045] Compared with the prior art, the advantages of the present invention are as follows:

[0046] The present invention creatively proposes an oversampling method based on a multi-false-class generative adversarial network. Through the multi-false-class generative adversarial network model trained by this method, a conditional variational autoencoder and a GAN model are jointly trained, which improves the stability of the generative adversarial training and generates better-quality images, thus better improving the classification performance of the classifier in the unbalanced scenario after oversampling. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is the overall flowchart of the present invention;

[0048] Figure 2 is Figure 1 the flowchart of the data preprocessing part in

[0049] Figure 3 is Figure 1 the flowchart of building and training the conditional variational autoencoder in

[0050] Figure 4 is Figure 1 the flowchart of building the generative adversarial network model in

[0051] Figure 5 is Figure 1 the flowchart of training a generative adversarial network model in

[0052] Figure 6 is Figure 1 the flowchart of using conditional variables and random noise as network inputs to generate minority class samples corresponding to labels in Specific implementation manners

[0053] The present invention will be further clarified below in conjunction with the accompanying drawings and specific implementation manners.

[0054] As Figures 2 - 6 shown, the present invention includes the following steps:

[0055] Step 1: As shown in the attached Figure 2 , obtain an imbalanced image dataset, convert the image size, and normalize the data so that the network input pictures have the same size, from step 201 to step 204:

[0056] Step 201: Let the image dataset be X = {X1, X2,..., Xn}, where n represents the maximum number of images;

[0057] Step 202: Let the image label set be Y = {Y1, Y2,..., Yn}, where n represents the maximum number of images;

[0058] Step 203: Convert the picture size to 64x64x number of channels;

[0059] Step 204: Normalize the image values and scale them to between [0, 1];

[0060] Step 2: As shown in the attached Figure 3 , build and train a conditional variational autoencoder model, from step 301 to step 307:

[0061] Step 301: Use a 4-layer convolutional network to build a decoder, with the convolutional kernel sizes being 64, 128, 128, and 256 respectively;

[0062] Step 302: Use a 4-layer convolutional network to build an encoder, with the convolutional kernel sizes being 256, 128, 128, and 64 respectively;

[0063] Step 303: Use embedding as an embedding layer model, input the label dimension, and output the latent space of the corresponding class;

[0064] Step 304: Use the picture to input the decoder. After the decoder outputs the mean and variance, calculate the noise according to the mean and variance, and embed the class label into the noise using embedding;

[0065] Step 305: Use the noise embedded with tag information as the input of the decoder, and finally output the generated picture;

[0066] Step 306: Use the mean squared error loss and the KLD loss to optimize the variational autoencoder;

[0067] Step 307: Set Adam as the optimizer, set the number of iterations to 30 times, and start training;

[0068] Step Three: As shown in the appendix Figure 4 , use the decoder in the variational autoencoder to initialize the weights of the generator, and build a generative adversarial network model, from Step 401 to Step 402:

[0069] Step 401: Build the generator model in the GAN, and then use the decoder in the variational autoencoder to initialize the weights of the generator;

[0070] Step 402: Build the discriminator model in the GAN, where the number of convolutional kernels in the middle layers are 64, 128, 128, and 256 respectively, and are separated into two output layers after the last layer of the middle layer, and output the sample validity score and the sample category respectively;

[0071] Step Four: As shown in the appendix Figure 5 , train the generative adversarial network model, from Step 501 to Step 505:

[0072] Step 501: Use the generator to generate a random label image, then input the real image, the generated image and their labels into the discriminator together, and obtain the classification loss cls_loss and the adversarial loss adv_loss respectively;

[0073] Step 502: Randomly take α, calculate the interpolation = real image + α x (real image - generated image), and obtain the gradient penalty gp according to the interpolation;

[0074] Step 503: Weightedly accumulate adv_loss, cls_loss and gp to obtain the total loss value, and update the discriminator network weight parameters according to the loss;

[0075] Step 504: Use the generator to generate a random image, obtain cls_loss and adv_loss after inputting it into the discriminator, and update the generator network weight parameters by weighted summation;

[0076] Step 505: Use Adam as the optimizer, set the number of iterations to 50 times, and start training.

[0077] Step Five: As shown in the appendix Figure 6 , use the minority class labels and random noise as the network input, and generate minority class samples corresponding to the labels, from Step 601 to Step 603:

[0078] Step 601: Calculate the number of minority class samples to be generated according to the difference between the number of samples in each small class and the number of samples in the large class in the data;

[0079] Step 602: Use the required amount of noise and minority class labels as the input of the generator to generate minority class samples corresponding to the labels;

[0080] Step 603: Add the minority class samples to the original image dataset until a balanced dataset is formed;

[0081] To better illustrate the effectiveness of this model, the CIFAR-10 dataset with an artificially set imbalance rate of 100 was tested, and Improved-BAGAN, ACGAN, IDA-GAN, and MFC-GAN were compared. After sampling each GAN to balance the imbalanced dataset, the same CNN classification network was used to test the dataset to obtain the performance of various imbalance metrics. The experimental results show that using the oversampling method based on the multi-false-class generative adversarial network proposed by us, the generated images have higher quality and can more effectively improve the performance of the classifier. The specific experimental situation is shown in Table 1. The lower the FID value, the better the authenticity of the image, and the higher the other metrics, the better. imp-MFCGAN represents the method proposed in this invention.

[0082] This patent discloses an invention of an oversampling method based on a multi-false-class generative adversarial network. The model proposed in this invention can effectively alleviate the impact brought by image imbalance. The specific improvements are as follows: jointly train a conditional variational autoencoder and a generative adversarial network so that the generator can learn the boundary distribution of different samples; introduce Focal Loss in the multi-false-class generative adversarial network to make the GAN network pay more attention to those samples that are difficult to classify; add a gradient penalty term to ensure the stability of GAN training.

[0083] The above are only examples of the implementation of this invention and are not used to limit this invention. All equivalent replacements made within the principle of this invention shall be included within the protection scope of this invention. The content not elaborated in detail in this invention belongs to the prior art well-known to those skilled in the art.

[0084] Table 1 Performance indicators of various GAN-based oversampling methods for CIFAR-10 with IR = 100

[0085] Algorithm / Index F1-Score Precision MG ACSA average-FID CGAN 0.380206 0.407274 0.224555 0.370685 341.016 ACGAN 0.392357 0.412318 0.246867 0.383879 201.1 BAGAN-GP 0.406676 0.421923 0.315568 0.40012 186.824 MFCGAN 0.396816 0.417082 O.303425 0.390648 227.482 imp-MFCGAN 0.426 0.433988 0.334169 0.4256 166.69

Claims

1. An oversampling method based on a multi-false-class generative adversarial network, characterized in that, The specific steps are as follows: (1) Obtain an unbalanced image dataset, convert the image size, and normalize the data so that the input images for the network have the same size; (2) Build and train a conditional variational autoencoder model. The specific steps are as follows: (2.1) Use a 4-layer convolutional network to build a decoder. The kernel sizes of the convolutional layers are 64, 128, 128, and 256 respectively. The LeakyReLU activation function is used between layers, and then it is unfolded using a Flatten layer. Finally, the output dimension is the dimension of the mean and variance; (2.2) Use a 4-layer convolutional network to build an encoder. The kernel sizes of the convolutional layers are 256, 128, 128, and 64 respectively. The LeakyReLU activation function is used between layers, and the dimension of the generated image is output; (2.3) Use embedding as the embedding layer model. Input the label dimension and output the latent space of the corresponding category; (2.4) Use the image as the input to the decoder. After the decoder outputs the mean and variance, calculate the noise according to the mean and variance, and embed the class label into the noise using embedding; (2.5) Use the noise with the embedded label information as the input to the decoder, and finally output the generated image; (2.6) Use the mean square error between the generated image and the original image as the reconstruction loss, and the error between the noise and the standard normal distribution as the KLD loss. Use the sum of the KLD loss and the reconstruction loss to train the variational autoencoder; (2.7) Set Adam as the optimizer, set the number of iterations to 30 times, and start training; (3) Initialize the weights of the generator using the decoder in the variational autoencoder, and build a generative adversarial network model; (4) Train the generative adversarial network model; (5) Use the minority class label and random noise as the network input to generate minority class samples corresponding to the label.

2. The oversampling method based on a multi-false-class generative adversarial network according to claim 1, characterized in that, The specific steps of data preprocessing in step (1) are as follows: (1.1) Obtain an unbalanced image dataset as X = {X1, X2, …, X n}, where n represents the maximum number of images; (1.2) Obtain the unbalanced image label set as Y = {Y1, Y2, …, Y n} where n represents the maximum number of images; (1.3) Convert the size of the image to 64x64x number of channels; (1.4) Normalize the image values and scale them to between [0,1].

3. The oversampling method based on a multi-false-class generative adversarial network according to claim 1, characterized in that, The specific steps of initializing the weights of the generator using the decoder in the variational autoencoder and building a generative adversarial network model in step (3) are as follows: (3.1) Build the generator model in the GAN, and then initialize the weights of the generator using the decoder in the variational autoencoder; the structure of the generator is set exactly the same as that of the decoder; (3.2) Build the discriminator model in the GAN. The number of convolutional kernels in the intermediate layers are 64, 128, 128, and 256 respectively. After the last layer of the intermediate layer, it is separated into two output layers, which output the sample validity score and the category of the sample respectively.

4. The oversampling method based on a multi-false-class generative adversarial network according to claim 1, characterized in that, The specific steps of training the generative adversarial network model in step (4) are as follows: (4.1) Use the generator to generate random label images, and then input the real image dataset, the generated image dataset and their labels into the discriminator, and obtain the classification output and the validity value output respectively; use the focal loss to calculate the classification output to obtain the classification loss cls_loss, and obtain the adversarial loss adv_loss according to the validity score; (4.2) Randomly sample a random value α from the normal distribution. Subtract the real image from the generated image to obtain the interpolation diff, and calculate the interpolation: real image + α × diff. Input the interpolation and the label into the discriminator and obtain the gradient grads. Finally, obtain the gradient penalty. where λ is the penalty coefficient and norm is the arithmetic square root of grads; (4.3) Accumulate the adv_loss, cls_loss, and gp by weighted summation to obtain the total loss value, and update the discriminator network weight parameters according to the loss; (4.4) Use the generator to generate a random label image, input it into the discriminator to obtain the classification output and the valid value output respectively; obtain the classification loss cls_loss and the adversarial loss adv_loss in the same way as in (4.1); sum the losses by weighted summation and update the generator network weight parameters; (4.5) Use Adam as the optimizer, set the number of iterations to 50, and start training.

5. The oversampling method based on a multi-false-class generative adversarial network according to claim 1, characterized in that, The specific steps of the step (5) for using the minority class label and random noise as the network input to generate the minority class samples corresponding to the label are as follows: (5.1) Calculate the number of minority class samples to be generated according to the difference between the number of samples of each small class and the number of samples of the large class in the data; (5.2) Use the required amount of noise and the minority class label as the input of the generator to generate the minority class samples corresponding to the label; (5.3) Add the minority class samples to the original picture dataset until a balanced dataset is formed.

Citation Information

Patent Citations

  • Industrial fault diagnosis unbalanced time series data expansion method

    CN112328588A

  • Deep Variational Method for Deformable Image Registration

    US20200034654A1