A method for classifying imbalanced datasets
By using the SMOTE and K-nearest neighbor algorithms to process imbalanced datasets, an online random forest classifier is constructed. By integrating dataset and algorithm information, the problem of low classification accuracy of imbalanced datasets in existing technologies is solved, and higher classification accuracy and generalization are achieved.
CN110991653BActive Publication Date: 2026-07-21UNIV OF ELECTRONICS SCI & TECH OF CHINA
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- UNIV OF ELECTRONICS SCI & TECH OF CHINA
- Filing Date
- 2019-12-10
- Publication Date
- 2026-07-21
Smart Images

Figure CN110991653B_ABST
Abstract
The application discloses a method for unbalanced data set classification, which is applied to the fields of network intrusion detection, animal age prediction, vehicle performance evaluation and the like, and aims at solving the problem of low classification precision of the prior art on the minority class. On the basis of original training data, the method uses the relationship between the minority class and the majority class in the original data set, uses SMOTE and K nearest neighbor algorithm to process the original training data set, and constructs a new set. The set focuses on the minority class and the majority class samples related to the minority class. Two random forests with the same size are constructed according to the original training data and the new set, then the decision trees in the two forests are combined into a large forest, and the large forest is used for testing the test set to obtain a classification result. The classification precision is greatly improved compared with the prior art.
Need to check novelty before this filing date? Find Prior Art