Data classification method and system based on hierarchical clustering and resampling, and storage medium
By applying hierarchical clustering and resampling methods on financial data sets, the problems of category imbalance and overlap between classes are solved, and the generalization ability and performance of classifiers are significantly improved, meeting the demand for precise classification in the financial field and other fields.
Patent Information
- Application Number
- CN202510399886.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-24
AI Technical Summary
In areas such as the financial field, data sets often face the problems of category imbalance and overlap between classes, resulting in insufficient performance and generalization capabilities of traditional classifiers.
Using a data classification method based on hierarchical clustering and resampling, the data set is divided into minority classes and majority classes, and hierarchical clustering is performed to form a sub-training set, and the sample ratio is adjusted through resampling to form an enhancer training set, and finally a classifier with high generalization ability is trained.
It effectively alleviates the negative impact of category imbalance on classification results, improves the classifier's ability to identify minority classes, and enhances the generalization ability and performance indicators of the model, such as accuracy, recall and F1 score.
Smart Images

Figure CN120197048A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a data classification method and system, and a storage medium based on hierarchical clustering and resampling. Background Art
[0002] In today's financial field and many other industries, classification tasks play a crucial role. From scenarios such as credit risk assessment, credit card customer default prediction, to consumer shopping behavior analysis, accurate classification is of key significance for decision-making, risk prevention and control, and business optimization. However, the performance of classifiers is greatly affected by the characteristics of the dataset.
[0003] In practical applications, datasets in specific fields such as financial datasets often face two prominent problems: class imbalance and class overlap. Class imbalance refers to a significant difference in the number of samples of different classes in the dataset. Usually, the minority class samples represent key abnormal situations or important events (such as financial fraud, credit default, etc.), but due to their small number, they are easily masked by the majority class samples during the training process, resulting in the classifier being difficult to effectively learn the feature patterns of the minority class, and thus having poor recognition ability for the minority class during prediction.
[0004] Class overlap means that there is an overlap of some regions in the feature space between different classes, making it difficult for the classifier to accurately divide class boundaries. These two problems are intertwined, seriously affecting the performance of traditional classifiers, resulting in unsatisfactory indicators such as the accuracy and recall rate of classification results, and unable to meet the requirements for accurate classification in practical applications.
[0005] To solve the class imbalance problem, a variety of methods have been developed in the prior art, including oversampling and undersampling techniques. Oversampling techniques such as SMOTE (Synthetic Minority Over-sampling Technique), GAN (Generative Adversarial Network), SMOTified-GAN, and CTGAN (Conditional Generative Adversarial Network for Tabular Data) increase the number of minority class samples by generating new minority class samples or enhancing the representation of minority class samples, attempting to balance the class distribution. Undersampling techniques such as random undersampling (RUS) achieve class balance by reducing the number of majority class samples. However, these resampling methods have certain limitations in practical applications. On the one hand, different resampling methods perform quite differently on different datasets, lacking a general and effective solution; on the other hand, using resampling methods alone often cannot fully solve the class overlap problem, resulting in limited improvement in classification performance.
[0006] In addition, in terms of the integration framework, although there have been attempts to combine clustering algorithms as components with other algorithms to handle data imbalance and anomaly detection problems, these methods still have insufficient generalization ability when dealing with complex financial data. Traditional clustering algorithms, such as hierarchical clustering, k-means, etc., are difficult to accurately capture the internal characteristics and distribution laws of data when faced with high-dimensional, high-noise, and complex-structured financial data, thus affecting the performance of the entire integration framework.
[0007] For example, Chinese Patent Application CN201910535115.3 in the prior art discloses an oversampling method for an imbalanced dataset: First, collect the imbalanced dataset and cluster it based on the K-means method, and divide the minority class and the majority class according to the number of elements in each class of the dataset; then, based on the SMOTE method, oversample the minority class dataset to obtain a synthetic minority class dataset; next, perform oversampling with replacement on the synthetic minority class dataset to obtain a new minority class dataset, forming a new dataset; finally, based on the CCA method, clean the new dataset: cluster the new dataset, calculate the Euclidean distance between each sample in each cluster and other samples in the same cluster and sort them, and delete the samples corresponding to the farthest Euclidean distance to obtain the cleaned dataset.
[0008] This prior art achieves the balance between the minority class and the majority class in the dataset through undersampling, but it still lacks generality for only one type of data.
[0009] In summary, when dealing with classification tasks in the financial field and other similar fields, due to the existence of class imbalance and class overlap problems, the existing classification methods and technologies have obvious deficiencies in terms of performance and generalization ability. There is an urgent need for a more effective solution to improve the performance of the classifier on imbalanced datasets to meet the needs of practical applications. Summary of the Invention
[0010] The purpose of the present invention is to provide a data classification method, system, and storage medium based on hierarchical clustering and resampling, which partially solve or alleviate the above deficiencies in the prior art, and the classifier trained with imbalanced data has higher generalization ability.
[0011] To solve the above-mentioned technical problems, the present invention specifically adopts the following technical solutions: In the first aspect of the present invention, it is to provide a data classification method based on hierarchical clustering and resampling, including: Divide the samples in the dataset into the minority class and the majority class; a part of the samples in the minority class are assigned to the training set, and the remaining part is assigned to the test set; a part of the samples in the majority class are assigned to the test set so that the ratio of the minority class samples to the majority class samples in the test set is 1:1; the remaining majority class samples are assigned to the remaining sample pool; Use the method of hierarchical clustering to cluster the samples in the training set, so as to divide the samples in the training set into several clusters to form corresponding sub-training sets; train the test set classifier with the cluster labels of the divided sub-training sets, and use the test set classifier to divide the samples in the test set to match each sub-training set with a sub-test set; randomly extract majority class samples from the remaining sample pool and inject them into each sub-training set so that the number of majority class samples in each sub-training set is r times that of the minority class samples; Resample the samples within the cluster so that the ratio of the minority class samples to the majority class samples in the sub-training set is 1:1, thus forming an enhanced sub-training set; Use the enhanced sub-training set to train the data classifier, and use the sub-test set to conduct classification testing on the data classifier; Use the test set classifier to classify the data to be classified, and use the data classifier corresponding to the category of the data to be classified to classify the data to be classified.
[0012] As an improvement, during the hierarchical clustering process, the clustering state with the highest similarity of samples within the cluster when the inter-cluster distance reaches the maximum is used as the best clustering, and the number of clusters corresponding to the best clustering is the best number of clusters.
[0013] As an improvement, the steps of selecting the best number of clusters include: Select the longest vertical line segment L in the dendrogram formed by hierarchical clustering; the vertical line segment is the branch of the dendrogram, and its length represents the inter-cluster distance; If there are n vertical line segments with the longest distance parallel to L, then take n + 1 as the best number of clusters.
[0014] As an improvement, 70% of the samples in the minority class are assigned to the training set, and the remaining 30% are assigned to the test set.
[0015] As an improvement, r ∈ [1, 2].
[0016] As an improvement, oversample the minority class samples in the training set or undersample the majority class samples, so that the ratio of the minority class samples to the majority class samples in the test set is 1:1.
[0017] As an improvement, the resampling method includes one or more of RUS, SMOTE, GAN, SMOTified - GAN, and CTGAN.
[0018] As an improvement, the classifier includes one or more of KNN, RF, DT, XGBoost, LGBM, Adaboosting, CatBoosting, GBM, HGB, MLP, and SVM.
[0019] The present invention provides a data classification system based on hierarchical clustering and resampling, including: A sample allocation module for dividing the samples in the dataset into minority classes and majority classes; a part of the samples in the minority classes are allocated to the training set, and the remaining part is allocated to the test set; a part of the samples in the majority classes are allocated to the test set so that the ratio of minority class samples to majority class samples in the test set is 1:1; the remaining majority class samples are allocated to the remaining sample pool; A hierarchical clustering module that uses the method of hierarchical clustering to cluster the samples in the training set, thereby dividing the samples in the training set into several clusters to form corresponding sub-training sets; training the test set classifier with the cluster labels of the divided sub-training sets, and using the test set classifier to divide the samples in the test set to match each sub-training set with a sub-test set; randomly extracting majority class samples from the remaining sample pool and injecting them into each sub-training set so that the majority class samples in each sub-training set are r times the minority class samples; A resampling module for resampling the samples within the cluster so that the ratio of minority class samples to majority class samples in the sub-training set is 1:1, thereby forming an enhanced sub-training set; A classifier construction module for training a data classifier using the enhanced sub-training set and classifying and testing the data classifier using the sub-test set; A classification module for classifying the data to be classified using the test set classifier and classifying the data to be classified using the data classifier corresponding to the category of the data to be classified.
[0020] The present invention provides a storage medium in which a computer program is stored; when the computing program is executed, the above-mentioned data classification method based on hierarchical clustering and resampling can be realized.
[0021] Beneficial effects: The present invention pre-divides the samples into minority classes and majority classes, and cleverly adjusts the ratios in the test set and sub-training sets. For example, the ratio of the two types of samples in the test set reaches 1:1, and the sub-training set also achieves 1:1 after resampling, effectively alleviating the negative impact of class imbalance on the classification result. By balancing the sample ratio, the classifier pays more attention to the characteristics of minority class samples, enhances the recognition ability of minority classes such as financial anomalies, and reduces the risk of misjudgment.
[0022] The present invention uses a hierarchical clustering method to cluster the training set samples, revealing the potential structure of the data, grouping similar data points into the same cluster, and facilitating the discovery of hidden patterns and rules in financial data. After clustering, the data within each cluster is similar, which helps to extract features more accurately, provides more valuable information for the training of the classifier, and improves the understanding and grasp of the corresponding data features in specific fields such as finance. Additionally, when determining the optimal clustering state, the optimal number of clusters is determined by the longest vertical line segment in the dendrogram and the line segment with the longest parallel distance, ensuring high similarity within the clusters and large differences between the clusters, making the division of the sub-training sets more reasonable and avoiding the problem of overlapping between classes in the prior art.
[0023] The present invention combines multiple resampling techniques and multiple classifiers for training and optimization. Multiple resampling methods such as random undersampling (RUS) and synthetic minority over-sampling technique (SMOTE) are selected to flexibly adjust the sample distribution according to the characteristics of different data sets; at the same time, multiple classifiers such as K-nearest neighbor (KNN) and random forest (RF) are used to enable the classifier to better adapt to the diversity and complexity of financial data. In this way, the performance of the classifier in the financial data classification task is comprehensively improved, including improving key indicators such as accuracy, recall rate, and F1 score, enhancing the generalization ability of the model, and enabling it to maintain stable and good performance in different financial data scenarios.
[0024] The present invention uses the trained classifier and the optimized analysis process to classify and analyze financial data efficiently and accurately. Whether it is application scenarios such as credit risk assessment, credit card customer default prediction, or financial fraud detection, the method provided by the present invention can provide accurate decision-making support for financial institutions and practitioners, helping them to discover potential risks in a timely manner, optimize business processes, reduce financial losses, and improve the security and stability of financial operations. Brief Description of the Drawings
[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale. Obviously, the following-described drawings are some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0026] Figure 1 It is a flowchart of the present invention.
[0027] Figure 2 It is a schematic diagram for selecting the optimal clustering.
[0028] Figure 3 It is a structural diagram of the present invention. Detailed implementation manners
[0029] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0030] In this document, suffixes such as "module", "component" or "unit" used to represent elements are only for the convenience of describing the present invention, and have no specific meaning in themselves. Therefore, "module", "component" or "unit" can be used interchangeably.
[0031] In this document, terms such as "upper", "lower", "inner", "outer", "front", "rear", "one end", "the other end", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the present invention. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0032] In this document, unless otherwise clearly defined and limited, terms such as "installation", "provided with", "connection", etc. should be understood in a broad sense. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, a direct connection, or an indirect connection through an intermediate medium, and can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0033] In this document, "and / or" includes any and all combinations of one or more of the listed related items.
[0034] In this document, "a plurality of" means two or more, that is, it includes two, three, four, five, etc.
[0035] Embodiment 1: As Figure 1 shown, this embodiment provides a data classification method based on hierarchical clustering and resampling, including: S1 Divide the samples in the dataset into minority classes and majority classes; a part of the samples in the minority class is assigned to the training set, and the remaining part is assigned to the test set; a part of the samples in the majority class is assigned to the test set so that the ratio of minority class samples to majority class samples in the test set is 1:1; the remaining majority class samples are assigned to the remaining sample pool.
[0036] In this step, the samples in the dataset are divided into minority classes and majority classes according to the sample labels. In some embodiments, 70% of the samples in the minority class are assigned to the training set, and the remaining 30% are assigned to the test set. Of course, this ratio can be adjusted according to specific circumstances.
[0037] Suppose in a credit card fraud detection dataset, fraud transaction records are the minority class and normal transaction records are the majority class. There are a total of 1000 pieces of data, including 100 fraud transactions and 900 normal transactions. According to the ratio of 70% and 30%, 70 fraud transaction records are assigned to the training set and 30 are assigned to the test set. To balance the classes in the test set, 30 normal transaction records are selected from the 900 normal transaction records and added to the test set, and the remaining 870 normal transaction records enter the remaining sample pool.
[0038] S2 Use the hierarchical clustering method to cluster the samples in the training set, so as to divide the samples in the training set into several clusters to form corresponding sub-training sets; train the test set classifier with the cluster labels of the divided sub-training sets, and use the test set classifier to divide the samples in the test set to match each sub-training set with a sub-test set; randomly select majority class samples from the remaining sample pool and inject them into each sub-training set so that the number of majority class samples in each sub-training set is r times that of minority class samples.
[0039] Hierarchical clustering is a clustering algorithm based on the similarity or distance measure between data points. It does not require pre-setting the number of clusters. When processing data, there are two ways: agglomerative from bottom to top and divisive from top to bottom. The bottom-up approach treats each data point as an independent class and continuously merges according to similarity until the stop condition is met; the top-down approach is the opposite, first treating all data as a large class and then gradually splitting. In this embodiment, the bottom-up approach is adopted, and its advantage is that it can discover the potential clustering structure of the data, without the need to pre-determine the number of clusters, and can naturally form clusters according to the characteristics and distribution of the data itself.
[0040] Financial data is complex and has various hidden patterns and structures. Hierarchical clustering can mine this potential information, grouping similar data points into the same cluster for subsequent analysis. Taking credit risk assessment data as an example, through hierarchical clustering, customers with similar credit characteristics (such as income, liabilities, credit history, etc.) can be grouped together, revealing the underlying laws of the data. After clustering, the data within each cluster is similar, making it more targeted and accurate to extract features from each cluster, avoiding the problem of class overlap that easily occurs when directly clustering all data sets in the prior art. For example, in credit card customer default prediction, after hierarchical clustering of the training set, features such as the consumption habits and repayment behaviors of customers in different clusters can be accurately extracted, providing more valuable information for subsequent classifier training.
[0041] In this embodiment, by hierarchically clustering the training set to form several sub-training sets, they are used to separately train the corresponding number of classifiers subsequently. The divided sub-training sets enable the classifiers to focus more on learning the features of specific data subsets. The data characteristics of different clusters are different, and the corresponding trained classifiers can better adapt to these differences, improving the prediction accuracy for different data partitions. For example, classifiers trained on different clusters can respectively capture the features of different types of financial risks and work together in the overall classification task to improve the overall performance of the classifiers.
[0042] Due to the bottom-up clustering method in this embodiment, as the clustering process progresses, the number of clusters will become fewer and fewer. The "number of clusters" actually represents the number of sub-training sets and also represents the number of classifiers to be trained later. It can be foreseen that it is not the final clustering result that best conforms to the actual situation of the samples, but during the clustering process, the clustering state with the highest intra-cluster sample similarity when the inter-cluster distance reaches the maximum is the optimal clustering, and the number of clusters corresponding to the optimal clustering is the optimal number of clusters, thereby determining the number and allocation of sub-training sets. Determining the optimal number of clusters and its sample partitioning is crucial, directly determining the number of sub-training sets. An appropriate number of sub-training sets can make the data partitioning more reasonable, making the data within each sub-training set highly similar and having large differences between different sub-training sets, which is conducive to subsequent classifiers better learning data features and improving the accuracy of classification and the performance of the model.
[0043] This embodiment provides an intuitive method for selecting the optimal number of clusters, and its specific steps include: S21 Select the longest vertical line segment L in the dendrogram formed by hierarchical clustering; the vertical line segment is the branch of the dendrogram, and its length represents the inter-cluster distance.
[0044] As Figure 2As shown, in the dendrogram formed by hierarchical clustering, the length of the branches (vertical line segments) of the dendrogram represents the distance between clusters. The longest vertical line segment L is selected because the distance between the two clusters corresponding to this line segment is the largest, meaning the difference in features between these two clusters is the greatest. Selecting such a line segment helps to find data sets with obvious differences in features and provides a key basis for determining the optimal number of clusters.
[0045] S22 If there are n vertical line segments with the longest parallel distance to L, then take n + 1 as the optimal number of clusters.
[0046] To select the appropriate number of clusters, the ideal cutting point is usually above the longest vertical line segment, which can ensure that the data points within the clusters maintain a high degree of similarity, while the differences between the clusters are relatively large. By horizontally cutting the dendrogram at an appropriate height, the intersection points of the horizontal line and the vertical line segments represent the number of clusters into which the data is divided. If there are n vertical line segments with the longest parallel distance to L, then take n + 1 as the optimal number of clusters.
[0047] The parallel and longest vertical line segments reflect that during the clustering process, the differences between n clusters have all reached a relatively maximum state, and at the same time, within these clusters, the sample similarity can also reach a relatively high level. Taking n as the optimal number of clusters can, on the basis of ensuring the similarity within the clusters, maximize the distinction between different clusters, thereby achieving an effective division of the data and providing the optimal number of clusters setting for subsequent sub-training set allocation and classifier training.
[0048] For example Figure 2 in, L is the longest vertical line segment, and the line segments parallel to L include L1, L2, L3, L4, L5, L6 and several line segments below. But only L1, L2 and L have the longest parallel distance h, so the clustering state at this time is taken as the optimal clustering. Correspondingly, the optimal number of clusters is 2 + 1 = 3 clusters.
[0049] Based on the division of the sub-training set, the test set also needs to be divided into corresponding sub-test sets. In this implementation, the samples in the test set are divided by the test set classifier to match sub-test sets for each sub-training set. In addition, the test set classifier will also pre-classify the data to be classified in subsequent actual use.
[0050] During the entire data processing and model training process, the sub-training sets are obtained by performing operations such as hierarchical clustering on the training set, and each sub-training set has a unique feature distribution. In order to accurately evaluate the performance of the model on different data subsets, the test set also needs to be divided into the corresponding number of sub-test sets so that the division of the test set corresponds to the sub-training sets. Doing so can ensure that when evaluating the model, the classifier trained on each sub-training set can be tested on the sub-test set with a similar feature distribution, thus more accurately reflecting the generalization ability and classification effect of the model on specific data subsets.
[0051] After dividing the sub-training sets, the samples and their cluster labels in the sub-training sets are used to train the test set classifier. The test set classifier selected in this embodiment is LGBM. Of course, other classifiers can also be selected, and the present invention does not limit this.
[0052] The LGBM (LightGBM) adopted in this embodiment is an efficient gradient boosting framework and is used in this implementation to divide the test set samples. It learns the features and distribution patterns of the data during the training process. Using this knowledge, LGBM can accurately assign the test set samples to the corresponding sub-test sets according to the features of the test set samples. For example, during the previous hierarchical clustering and sub-training set construction process, LGBM may have learned the feature differences between different clusters (sub-training sets). Then, when dividing the test set samples, it can divide the test set samples similar to the features of a certain sub-training set into the corresponding sub-test set according to these feature differences.
[0053] Match sub-test sets for each sub-training set to make the model training and testing processes more corresponding and reasonable. During the training phase, the classifier learns the features and patterns of the data on the sub-training set; during the testing phase, by testing on the corresponding sub-test set, the mastery degree of the classifier for these features and patterns and its generalization ability on new data can be evaluated. This matching method helps to more accurately measure the performance of the model on different data subsets, thus providing a more targeted basis for the optimization and improvement of the model and enhancing the performance and reliability of the entire model when processing complex data sets such as financial data.
[0054] After determining the division method of the sub-training sets, it is also necessary to inject majority class samples into the sub-training sets. In this implementation, majority class samples are randomly selected from the remaining sample pool and injected into each sub-training set so that the number of majority class samples in each sub-training set is r times that of the minority class samples. The purpose of this step is to preliminarily balance the samples in the sub-training sets.
[0055] In financial datasets, the phenomenon of class imbalance is widespread, with a huge disparity in the number of samples between the minority class and the majority class, which may reach 1:99, for example. This imbalance causes the classifier to be biased towards the majority class during training, resulting in a decline in the recognition ability for the minority class.
[0056] Although there is no theoretical limit on the specific number of r in this embodiment, considering the actual situation, r should preferably not exceed 5 times. If r is too large, the problem of class imbalance cannot be solved. Controlling r within a certain range helps to ensure an adequate number of majority-class samples while avoiding excessive imbalance and achieving a preliminary balance in the sample ratio of the sub-training set. Of course, in some extreme cases, r < 1 may even occur.
[0057] In fact, through the later verification results, it is found that the classifier trained when r ∈ [1, 2] has the highest accuracy. This indicates that within this value range, the sample ratio of the sub-training set reaches a relatively optimal state. When r = 1, the number of samples in the majority class and the minority class in the sub-training set is equal, which is equivalent to no longer extracting and adding majority-class samples from the remaining majority-class sample pool after the first random undersampling to 1:1. This can balance the class distribution to the greatest extent and ensure the proportion of the true minority class, enabling the classifier to learn more evenly from both types of samples; when r = 2, the number of majority-class samples is appropriately increased, which not only retains certain information of the majority-class samples but also does not break the sample balance, helping the classifier learn more comprehensive features and improving the accuracy and stability of classification.
[0058] S3 performs oversampling on the samples within the cluster so that the ratio of minority-class samples to majority-class samples in the sub-training set is 1:1, thereby forming an enhanced sub-training set.
[0059] More specifically, in this embodiment, oversampling is performed on the minority-class samples in the training set or undersampling is performed on the majority-class samples, so that the ratio of minority-class samples to majority-class samples in the test set is 1:1.
[0060] The purpose of this step is to solve the problem of class imbalance in the dataset because an imbalanced sample distribution will cause the classifier to be biased towards the majority class during training, ignoring the features of the minority class, and ultimately affecting the classifier's recognition ability for the minority-class samples. By adjusting the ratio to 1:1, the classifier can learn the features of both types of samples more evenly and improve the classification performance.
[0061] Specifically, when the number of majority-class samples in the sub-training set is r times the number of minority-class samples and r > 1, it means that the number of majority-class samples is relatively large and the number of minority-class samples is insufficient. At this time, the method of oversampling the minority-class samples is adopted. By generating new minority-class samples, the number of minority-class samples is increased so that the number of minority-class samples is equal to that of the majority-class samples, reaching a ratio of 1:1. For example, when r = 2, that is, the number of majority-class samples is twice that of the minority-class samples, oversampling methods such as the Synthetic Minority Over-sampling Technique (SMOTE) can be used to synthesize new samples for the minority-class samples to fill the quantitative gap and achieve the balance of the sample ratio.
[0062] Through the above oversampling and undersampling operation methods, the balance between the minority-class samples and the majority-class samples in the training set is successfully achieved. This balanced sample distribution can provide more balanced training data for the classifier, enabling the classifier to fully consider the characteristics of both types of samples during the learning process, avoiding biases caused by excessive differences in the number of samples, thereby improving the recognition accuracy of the classifier for different classes and enhancing the generalization ability and reliability of the model in practical applications.
[0063] In this embodiment, the resampling method includes one or more of RUS (Random Under-Sampling), SMOTE (Synthetic Minority Over-sampling Technique), GAN (Generative Adversarial Network), SMOTified - GAN (the combination of SMOTE and GAN), and CTGAN (Conditional Tabular GAN).
[0064] S4 Train a data classifier using the enhancer training set and classify and test the data classifier using the sub-test set.
[0065] The data classifiers applicable to the present invention are diverse, including but not limited to one or more of KNN (K-Nearest Neighbors), RF (Random Forest), DT (Decision Tree), XGBoost (eXtreme Gradient Boosting), LGBM (LightGBM), Adaboosting, CatBoosting, GBM (Gradient Boosting Machine), HGB (HistGradientBoosting), MLP (Multi-Layer Perceptron), and SVM (Support Vector Machine).
[0066] After selecting the classifier, the enhancer training set can be used to train the classifier.
[0067] After the training is completed, the corresponding data classifier also needs to be tested through the sub-test set.
[0068] Each sub-test set (i.e., sub-test set 1, sub-test set 2, sub-test set 3) has a corresponding adapted sub-training set. This corresponding relationship is determined through operations such as hierarchical clustering of the training set, sample injection, and division of the test set before. The role of the training set is to provide data support for training a specific data classifier. The sub-training set is used to let the classifier learn the features and patterns of the data, while the sub-test set is used to evaluate the performance of the classifier on the corresponding data subset after training.
[0069] After each data classifier is tested on its corresponding sub-test set, corresponding test results will be obtained. The test results are usually presented in the form of a confusion matrix. The test results of these different classifiers on their respective sub-test sets are combined, that is, multiple confusion matrices are integrated together, and evaluation metrics such as the overall F1 value, recall rate, and AUC (Area Under the Curve) are calculated. The purpose of doing this is to comprehensively consider the performance of all classifiers on different data subsets, avoid the limitations of a single classifier or a single sub-test set, and more comprehensively evaluate the performance of the entire classification model under different data distributions.
[0070] S5 uses the test set classifier to classify the data to be classified, and uses the data classifier corresponding to the category of the data to be classified to classify the data to be classified.
[0071] In the previous steps, through operations such as hierarchical clustering, sample injection, and resampling on the training set, the training set has been divided into multiple sub-training sets, and corresponding data classifiers have been trained for each sub-training set. At the same time, the test set has been divided using a test set classifier such as LGBM, and a sub-test set has been matched for each sub-training set. After training and evaluation, multiple data classifiers with good performance have been obtained. The classifiers trained from each cluster are only responsible for processing the data subset belonging to that cluster. LGBM is an efficient gradient boosting framework that has learned the data features and distribution patterns during the previous training process. When processing the data to be analyzed, LGBM can, based on this learned knowledge, divide the data to be analyzed into appropriate clusters, thereby determining which trained classifier should be used for the final classification.
[0072] According to the partitioning result of the data to be analyzed by LGBM, determine the cluster to which the data belongs, and then use the classifier corresponding to that cluster. Since each classifier is trained for a specific sub-training set, and the data features of this sub-training set are similar to the data features within the corresponding cluster, the selected classifier can better process the data within that cluster.
[0073] Input the data to be analyzed into the classifier. The classifier classifies the data to be analyzed according to the features and patterns it has learned on the sub-training set and outputs the classification result. For example, in financial fraud detection, the classifier will determine whether the transaction record is a normal transaction or a fraudulent transaction; in credit risk assessment, it will give the credit risk level of the customer.
[0074] Embodiment 2: As Figure 3 shown, the present invention also provides a financial data analysis system based on hierarchical clustering, which is characterized by including: A sample allocation module for dividing the samples in the dataset into minority classes and majority classes; a part of the samples in the minority classes are allocated to the training set, and the remaining part is allocated to the test set; a part of the samples in the majority classes are allocated to the test set so that the ratio of minority class samples to majority class samples in the test set is 1:1; the remaining majority class samples are allocated to the remaining sample pool; A hierarchical clustering module for clustering the samples in the training set using the hierarchical clustering method, thereby dividing the samples in the training set into several clusters to form corresponding sub-training sets; training a test set classifier using the cluster labels of the divided sub-training sets, and using the test set classifier to divide the samples in the test set to match a sub-test set for each sub-training set; randomly extracting majority class samples from the remaining sample pool and injecting them into each sub-training set so that the number of majority class samples in each sub-training set is r times that of minority class samples; A resampling module, which is used to resample the samples within a cluster so that the ratio of the minority-class samples to the majority-class samples in the sub-training set is 1:1, thereby forming an enhanced sub-training set; A classifier construction module, which is used to train a data classifier by using the enhanced sub-training set and classify and test the data classifier by using the sub-test set; A classification module, which is used to classify the data to be classified by using the test set classifier and classify the data to be classified by using the data classifier corresponding to the category of the data to be classified.
[0075] Embodiment 3: The present invention provides a storage medium, in which a computer program is stored; when the computer program is executed, the above-mentioned financial data analysis method based on hierarchical clustering can be realized.
[0076] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0077] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a computer terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0078] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims. These all belong to the protection scope of the present invention.
Claims
1. A data classification method based on hierarchical clustering and resampling, characterized in that include: The samples in the data set are divided into a minority class and a majority class; a portion of the samples in the minority class are assigned to the training set, and the remaining portion are assigned to the test set; A portion of the samples in the majority class are allocated to the test set, so that the ratio of minority class samples to majority class samples in the test set is 1:1; the remaining majority class samples are allocated to the remaining sample pool; The hierarchical clustering method is used to cluster the samples in the training set, so that the samples in the training set are divided into several clusters to form a corresponding number of sub-training sets; The test set classifier is trained by the cluster labels of the divided sub-training sets, and the test set classifier is used to divide the samples in the test set to match the sub-test set for each sub-training set; the majority class samples are randomly selected from the remaining sample pool and injected into each sub-training set, so that the majority class samples in each sub-training set are r times the minority class samples; Resample the samples within the cluster so that the ratio of minority class samples to majority class samples in the sub-training set is 1:1, thus forming an enhanced sub-training set; The data classifier is trained using the enhanced sub-training set, and the data classifier is tested using the sub-testing set; The test set classifier is used to classify the data to be classified, and the data classifier corresponding to the category of the data to be classified is used to classify the data to be classified.
2. A data classification method based on hierarchical clustering and resampling according to claim 1, characterized in that include: In the process of hierarchical clustering, the clustering state with the highest similarity of samples within the cluster when the inter-cluster distance reaches the maximum is taken as the best clustering, and the number of clusters corresponding to the best clustering is the optimal number of clusters.
3. A data classification method based on hierarchical clustering and resampling according to claim 2, characterized in that The steps to select the optimal number of clusters include: Select the longest vertical line segment L in the dendrogram formed by hierarchical clustering; the vertical line segment is the branch of the dendrogram, and its length represents the distance between clusters; If there are n vertical line segments parallel to L with the longest distance, then n+1 is taken as the optimal number of clusters.
4. The data classification method based on hierarchical clustering and resampling according to claim 1, characterized in that: 70% of the samples in the minority class are assigned to the training set, and the remaining 30% are assigned to the test set.
5. The data classification method based on hierarchical clustering and resampling according to claim 1, characterized in that: r∈[1,2]。 6. The data classification method based on hierarchical clustering and resampling according to claim 1, characterized in that: Oversample the minority class samples in the training set or undersample the majority class samples so that the ratio of minority class samples to majority class samples in the test set is 1:
1.
7. A data classification method based on hierarchical clustering and resampling according to claim 6, characterized in that: The resampling method includes one or more of RUS, SMOTE, GAN, SMOTified-GAN and CTGAN.
8. The data classification method based on hierarchical clustering and resampling according to claim 1, characterized in that: The classifier includes one or more of KNN, RF, DT, XGBoost, LGBM, Adaboosting, CatBoosting, GBM, HGB, MLP and SVM.
9. A data classification system based on hierarchical clustering and resampling, characterized in that include: The sample allocation module is used to divide the samples in the data set into a minority class and a majority class; a part of the samples in the minority class are allocated to the training set, and the rest are allocated to the test set; a part of the samples in the majority class are allocated to the test set, so that the ratio of the minority class samples to the majority class samples in the test set is 1:1; the remaining majority class samples are allocated to the remaining sample pool; The hierarchical clustering module uses the hierarchical clustering method to cluster the samples in the training set, thereby dividing the samples in the training set into several clusters to form a corresponding number of sub-training sets; The test set classifier is trained by the cluster labels of the divided sub-training sets, and the test set classifier is used to divide the samples in the test set to match the sub-test set for each sub-training set; the majority class samples are randomly selected from the remaining sample pool and injected into each sub-training set, so that the majority class samples in each sub-training set are r times the minority class samples; The resampling module is used to resample the samples within the cluster so that the ratio of minority class samples to majority class samples in the sub-training set is 1:1, thereby forming an enhanced sub-training set; A classifier building module, used for training a data classifier using the enhanced sub-training set, and performing classification tests on the data classifier using the sub-testing set; The classification module is used to classify the data to be classified using the test set classifier, and to classify the data to be classified using the data classifier corresponding to the category of the data to be classified.
10. A storage medium, characterized in that: The storage medium stores a computer program; when the computer program is executed, the data classification method based on hierarchical clustering and resampling described in any one of claims 1 to 8 can be implemented.
Citation Information
Patent Citations
Oversampling method for unbalanced data set
CN110275910A
Cited By
25Hz phase-sensitive track circuit intelligent fault positioning method and system
CN120559392A