Data balancing method for unbalanced multi-class adversarial network and transfer learning
By independently training the GANs model for each data category, generating synthetic data and combining transfer learning, the problem of non-balanced data sets in machine learning is solved, and the generalization ability and accuracy of the model are improved.
Patent Information
- Application Number
- CN202510230559.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-06
AI Technical Summary
In machine learning, unbalanced data sets cause the model to bias the majority class and ignore the minority class, affecting the accuracy and generalization ability of the model.
By independently training the corresponding generative adversarial network (GANs) model for each data category, high-quality synthetic data is generated, and the data set balance is achieved in combination with transfer learning methods.
It improves the diversity and quality of data, reduces the impact of data imbalance on model training, and improves the generalization ability and accuracy of the model.
Smart Images

Figure CN119939255A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine learning, and in particular to processing technology for unbalanced data sets, and aims to achieve effective balance of multi-category data by combining a generative adversarial network (GAN) with a transfer learning method, so as to improve the training efficiency and generalization ability of the model. Background Art
[0002] In practical applications, datasets often have class imbalance problems, that is, the number of samples in some categories is much larger than that in other categories. This imbalance will cause machine learning models to favor the majority class and ignore the minority class when making predictions, thus affecting the accuracy and practicality of the model. Traditional methods such as resampling and synthetic minority oversampling technique (SMOTE) can alleviate the imbalance problem to a certain extent, but may introduce noise or overfitting.
[0003] Specifically, the problems encountered in the data labeling stage of unbalanced data sets mainly include the following: 1. Problems in the data labeling stage: 1.1) Labeling bias: In an unbalanced dataset, since minority class samples have a low frequency of occurrence, labelers may easily overlook these samples, resulting in incomplete or inaccurate labeling, affecting the quality of the dataset. 1.2) Insufficient data representativeness: The number of minority class samples is limited, and it is difficult to fully represent the diversity and variability of the category, resulting in limited representativeness of subsequent training data.
[0004] 2. Impact on model training: 2.1) Bias towards the majority class: During training, the model may be more inclined to learn the features of the majority class samples, resulting in insufficient recognition of minority class samples. 2.2) Misclassification problem: After training on unbalanced data, the model may tend to predict the majority class, resulting in missed detection of minority class defects. 2.3) Imbalanced evaluation indicators: Using indicators such as overall accuracy to evaluate model performance may be inaccurate, because the model may accurately predict majority class samples, but poorly predict minority class samples. Therefore, more detailed indicators such as precision and recall are needed to evaluate the model.
[0005] 3. Impact on subsequent analysis and decision-making: 3.1) False positives and false negatives: False positives and false negatives of the model may have a serious impact on actual production and quality control, especially when the minority class represents important defects. 3.2) Model generalization ability: The model's weak ability to identify minority classes will limit its generalization ability and make it difficult to deal with the diverse defects in actual production. 3.3) Decision-making risk: Models trained on imbalanced data may be biased towards the majority class when making decisions, ignoring key information in the minority class, and challenging the fairness and comprehensiveness of the decision-making process.
[0006] Therefore, it is of great significance to explore new data balancing methods. Summary of the invention
[0007] The purpose of the present invention is to overcome the shortcomings of the prior art and propose a data balancing method for unbalanced multi-category adversarial networks and transfer learning, which generates high-quality synthetic data by independently training the corresponding GANs model for each data category and achieves the balance of the data set.
[0008] In order to achieve the above-mentioned purpose, the present invention provides a data balancing method for unbalanced multi-class adversarial network and transfer learning, which is characterized by comprising the following steps: S1. Dataset preparation: - Obtain an initial data set from the target field or a related field, and mark each sample data in the obtained initial data set with a category label; -Perform quantitative statistical analysis on each sample data after labeling; - Determine the number of sample data and the most common category labels in the sample data; S2. Build a transfer learning model: -Build a benchmark model based on deep learning generative adversarial network (GAN), train the model on each sample data in the category label with the largest number of sample data, and obtain a benchmark model; -Build a benchmark model based on deep learning generative adversarial network (GAN), train models for each category label except the one with the largest number of sample data, and obtain multiple independent models through independent training; S3. Data generation and balancing: - Generate new data sets consistent with the data features of each category label based on the benchmark model and each independent model, where the amount of data in the new data set generated by the benchmark model is defined as the benchmark amount, and the amount of data in the new data sets generated by each independent model of the generative adversarial network is equal to or close to the benchmark amount, so that the amount of data for each category label is balanced; S4. Build training model: -The balanced data set is divided into a training set and a test set according to a preset ratio, the preset target model is trained using the constructed training set, and the preset target model is verified using the constructed test set.
[0009] Furthermore, the preset target model includes specific patterns or features to be identified.
[0010] The present invention adopts the above-mentioned scheme, and its beneficial effects are: 1) For each category of data, the corresponding GANs model is trained independently to ensure the independence and focus of the data generation process of each category. 2) The GANs model is used to generate realistic synthetic data, so that the feature distribution of minority class data is closer to the real data. Through transfer learning, the features learned in the pre-trained model are transferred to the new category data generation process, further enhancing the authenticity and richness of the generated data, thereby improving the diversity and quality of the data. 3) By generating a sufficient number of minority class data, a balance is achieved in the amount of data in each category, reducing the impact of data imbalance on model training. 4) It can cope with various complex data types, such as processing and application scenarios of images, text and audio. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 Figure 2 is a flow chart of the data balancing method. DETAILED DESCRIPTION
[0012] In order to facilitate the understanding of the present invention, the present invention is described more fully below with reference to the accompanying drawings. The accompanying drawings provide preferred embodiments of the present invention. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. The purpose of providing these embodiments is to make the disclosure of the present invention more thoroughly and comprehensively understood.
[0013] See attached Figure 1 As shown, in this embodiment, a data balancing method for unbalanced multi-category adversarial network and transfer learning is provided. The method constructs multiple generative adversarial network models (GANs models) to generate customized data for different categories of data features, thereby achieving a balanced data set. The specific steps are as follows: S1. Dataset preparation: - Obtain an initial data set from the target field or related fields, and label each sample data in the initial data set with a category label; specifically, manual annotation or automated tools (such as image recognition algorithms, text classifiers, etc.) can be used to assign category labels to sample data to ensure that each sample is correctly classified, providing a basis for subsequent steps. The accuracy of the labeling here is crucial because it directly affects the results of subsequent steps.
[0014] -Perform quantitative statistical analysis on each labeled sample data.
[0015] -Determine the number of sample data and the most common class labels in the sample data.
[0016] Specifically, the category label is used as the statistical benchmark to determine the number of sample data contained in each category label and the category label with the largest number of sample data, that is, to count the number of samples corresponding to each category label and identify the category with the largest number of samples (i.e., the majority class) and other minority classes. Through the statistical results, the category with the largest number of samples (majority class) and other categories with a smaller number of samples (minority class) are identified.
[0017] Based on the above step S1, it is helpful to determine the degree of imbalance in data distribution and provide data support for subsequent steps.
[0018] S2. Build a transfer learning model: -Build a benchmark model based on deep learning generative adversarial networks, and train the model for each sample data in the category label with the largest number of sample data; specifically, select the majority class samples (the category label with the largest number of sample data) and build a benchmark model through deep learning generative adversarial networks (GANs). The training goal is to be able to generate new sample data with similar characteristics to the majority class data as a subsequent data volume benchmark.
[0019] -Build several independent models based on deep learning generative adversarial networks, and train models for each category label except the category label with the largest number of sample data; specifically, select each minority class (the remaining category labels), and train an independent model using the sample data corresponding to each category label. These independent models are designed to learn the data distribution of their respective categories so as to subsequently generate new sample data that matches the category label.
[0020] Furthermore, in the above step S2, the model training is performed through the structure of deep learning generative adversarial networks (GANs), wherein in the technical field, generative adversarial networks (GANs) are composed of a generator and a discriminator, the generator is responsible for generating new samples, and the discriminator is responsible for judging the authenticity of the samples. Through training, the generator can learn the data distribution of the majority class data and generate new sample data with similar features to the majority class data.
[0021] S3. Data generation and balancing: -Generate new data sets consistent with the data features of each category label based on the benchmark model and each independent model, wherein the amount of data in the new data set generated by the benchmark model is defined as the benchmark amount, and the amount of data in the new data set generated by each independent model of the generative adversarial network is equal to or close to the benchmark amount, so that the amount of data of each category label is balanced; specifically, first, use the benchmark model to generate a certain amount of new sample data as the benchmark amount, and then each independent model generates equal or close new sample data based on this benchmark amount, so as to ensure that the amount of data of all categories is balanced or close to each other. In addition, in the process of generating new sample data, it is necessary to ensure the quality of the generated data. The authenticity and validity of the data can be ensured by verifying and screening the generated data.
[0022] S4. Build training model: -The balanced data set is divided into a training set and a test set according to a preset ratio, and the preset target model is trained using the constructed training set, and the preset target model is verified using the constructed test set. Specifically, first, the preset target model is trained on the constructed training set, and the performance of the model is evaluated by means of preset cross-validation, holdout method and other methods to ensure that its performance on the balanced data set is better than that on the original unbalanced data set. Thus, according to the balanced data set, the target model can better learn the characteristics of each category, thereby improving the generalization ability and accuracy of the target model. In addition, by evaluating the performance of the model through cross-validation, holdout method and other methods, the user can understand the performance of the model on the balanced data set, and compare it with the performance on the original unbalanced data set to verify the effectiveness of the data balancing method.
[0023] In summary, the data balancing method of unbalanced multi-class adversarial network and transfer learning can effectively solve the problem of data imbalance. This method not only improves the quality and diversity of data, but also improves the generalization ability and accuracy of the model. In practical applications, we can make appropriate adjustments and optimizations to the above steps according to the characteristics of specific tasks and data sets to achieve better results.
[0024] The embodiments described above are only preferred embodiments of the present invention and are not intended to limit the present invention in any form. Any technician familiar with the art who, without departing from the scope of the technical solution of the present invention, makes more possible changes and modifications to the technical solution of the present invention using the technical content disclosed above, or modifications are all equivalent embodiments of the present invention. Therefore, any equivalent and equivalent changes made according to the ideas of the present invention without departing from the content of the technical solution of the present invention should be included in the protection scope of the present invention.
Claims
1. A data balancing method for unbalanced multi-class adversarial networks and transfer learning, characterized by: The following steps are included: S1. Dataset preparation: - Obtain an initial data set from the target field or a related field, and mark each sample data in the obtained initial data set with a category label; -Perform quantitative statistical analysis on each sample data after labeling; - Determine the number of sample data and the most common category labels in the sample data; S2. Build a transfer learning model: -Build a benchmark model based on deep learning generative adversarial network, and train the model on each sample data in the category label with the largest number of sample data; -Build several independent models based on deep learning generative adversarial networks, and train the models for each category label except the one with the largest number of sample data; S3. Data generation and balancing: - Generate new data sets consistent with the data features of each category label based on the benchmark model and each independent model, where the amount of data in the new data set generated by the benchmark model is defined as the benchmark amount, and the amount of data in the new data sets generated by each independent model of the generative adversarial network is equal to or close to the benchmark amount, so that the amount of data for each category label is balanced; S4. Build training model: -The balanced data set is divided into a training set and a test set according to a preset ratio, the preset target model is trained using the constructed training set, and the preset target model is verified using the constructed test set.
2. The data balancing method for unbalanced multi-class adversarial network and transfer learning according to claim 1, characterized in that: The preset target model includes specific patterns or features to be recognized.