Improved BalanceCascade data processing method

By introducing dynamic sampling strategies, classifier weighting mechanisms and advanced basic classifiers into the BalanceCascade algorithm, the shortcomings of traditional algorithms in dealing with highly unbalanced data sets are solved, and the ability to recognize a few classes is significantly improved and the performance of most classes is maintained, thus realizing adaptive learning and stable performance of data set features.

CN120180304APending Publication Date: 2025-06-20EAST CHINA NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510250751.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When the traditional BalanceCascade algorithm deals with highly unbalanced data sets, the sampling strategy is too simple, resulting in information loss, lacking a mechanism to distinguish and utilize the performance differences of different classifiers. The basic classifier is simple, difficult to capture complex patterns, and does not consider the cost of different error types.

Method used

Dynamic sampling strategies, classifier weight mechanism, advanced basic classifiers and training strategies that consider category imbalance are introduced. By dynamically adjusting the sampling strategies and classifier weights, adaptive learning of data set features is achieved.

Benefits of technology

It significantly improves the ability to identify a few class samples, while maintaining good performance for most class samples, achieving stable performance under different degrees of category imbalance, and enhancing its applicability in actual scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180304A_ABST
    Figure CN120180304A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to an improved BalanceCa scad data processing method which comprises the steps that a training data set is acquired, and the training data set comprises a plurality of samples and category labels corresponding to the samples; based on the training data set, determining majority class samples and minority class samples; according to the preset number of classifiers, executing the following iteration process: dynamically sampling from the majority class samples, and enabling the number of the majority class samples after sampling to be equal to the number of the minority class samples; merging the sampled majority class samples and minority class samples to form a training subset of the current iteration; training a classifier of the current iteration based on the training subset; calculating the weight of the current classifier; a new majority class sample set; and outputting the plurality of classifiers obtained by training and the corresponding weights thereof. The improvement of the invention not only significantly improves the recognition capability of minority class samples, but also maintains good performance of majority class samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and more specifically, to an improved BalanceCascade data processing method. Background Art

[0002] In today's data-driven era, machine learning algorithms are increasingly widely used in various fields. However, datasets in the real world often have a serious class imbalance problem, which poses a great challenge to traditional classification algorithms. Especially in key fields such as financial fraud detection, medical diagnosis, and network security, minority class samples usually represent the events we care most about, but their proportion in the dataset may be as low as 1% or even less.

[0003] Traditional classification algorithms often perform poorly when faced with such highly imbalanced datasets. They tend to classify all samples as the majority class, thus achieving seemingly good results in overall accuracy, but actually completely ignoring the recognition of the minority class. Such results may lead to catastrophic consequences in practical applications, such as missing major fraud transactions or missing rare but fatal disease diagnoses.

[0004] To address this challenge, researchers have proposed various methods, among which the BalanceCascade algorithm is a widely concerned solution. This algorithm attempts to balance the data distribution and improve the recognition ability of the minority class by constructing multiple classifiers and downsampling the majority class samples in each iteration. However, the traditional BalanceCascade algorithm still has some obvious deficiencies. Summary of the Invention

[0005] In view of these deficiencies of the traditional BalanceCascade algorithm, the present invention proposes an improved data processing method. This method aims to comprehensively improve the performance of the algorithm when dealing with highly imbalanced datasets by introducing innovative points such as a dynamic sampling strategy, a classifier weight mechanism, an advanced basic classifier, and a training strategy considering class imbalance.

[0006] The present invention provides an improved BalanceCascade data processing method, characterized by comprising:

[0007] An obtaining step, comprising:

[0008] Obtaining a training dataset, where the training dataset includes multiple samples and their corresponding class labels;

[0009] A processing step, comprising:

[0010] Based on the training dataset, determining majority class samples and minority class samples;

[0011] Perform the following iterative process according to the preset number of classifiers:

[0012] (1) Dynamically sample from the majority-class samples so that the number of majority-class samples after sampling is equal to the number of minority-class samples;

[0013] (2) Combine the majority-class samples and minority-class samples after sampling to form the training subset for the current iteration;

[0014] (3) Train the classifier for the current iteration based on the training subset;

[0015] (4) Calculate the weight of the current classifier;

[0016] (5) Update the majority-class sample set;

[0017] The output steps include:

[0018] Output the multiple trained classifiers and their corresponding weights.

[0019] Preferably, the dynamic sampling in the processing steps specifically includes:

[0020] Judge whether the number of remaining majority-class samples is less than or equal to the number of minority-class samples;

[0021] If so, directly use all the remaining majority-class samples;

[0022] If not, randomly sample from the remaining majority-class samples, and the sampling quantity is equal to the number of minority-class samples.

[0023] Preferably, the training of the classifier for the current iteration in the processing steps specifically includes:

[0024] Calculate the class weights, where the weight of the majority class is 1, and the weight of the minority class is the ratio of the number of majority-class samples to the number of minority-class samples;

[0025] Based on the class weights, assign sample weights to each sample in the training subset;

[0026] Use the training subset with sample weights to train the LightGBM classifier.

[0027] Preferably, the calculation of the weight of the current classifier in the processing steps specifically includes:

[0028] Use the classifier obtained from the current training to predict the training subset;

[0029] Calculate the recall rate of the minority-class samples in the prediction results;

[0030] Take the recall rate as the weight of the current classifier.

[0031] Preferably, updating the majority class sample set in the processing step specifically includes:

[0032] Using the classifier obtained by current training to predict the remaining majority class samples;

[0033] Calculating the confidence of the prediction, where the confidence is equal to twice the absolute difference between the prediction probability and 0.5;

[0034] Removing the majority class samples that are correctly predicted and have a confidence greater than a preset threshold.

[0035] Preferably, the preset threshold is 0.8.

[0036] Preferably, it further includes a prediction step:

[0037] Obtaining the sample to be predicted;

[0038] Using all the trained classifiers to predict the sample to be predicted;

[0039] Based on the prediction results of each classifier and the corresponding weights, obtaining the final prediction result through weighted voting.

[0040] Preferably, the weighted voting specifically includes:

[0041] Calculating the sum of all classifier weights;

[0042] Multiplying the prediction result of each classifier by its corresponding weight and accumulating to obtain the weighted sum;

[0043] If the weighted sum is greater than or equal to half of the sum of weights, then predicting as the positive class, otherwise predicting as the negative class.

[0044] Preferably, it further includes a probability prediction step:

[0045] Obtaining the sample to be predicted;

[0046] Using all the trained classifiers to perform probability prediction on the sample to be predicted;

[0047] Based on the probability prediction results of each classifier and the corresponding weights, calculating the weighted average probability as the final probability prediction result.

[0048] Preferably, the iteration process in the processing step further includes termination condition judgment:

[0049] Judging whether the number of remaining majority class samples is zero;

[0050] If so, terminating the iteration process;

[0051] If not, continuing to execute the next round of iteration until the preset number of classifiers is reached.

[0052] The beneficial effects of the present invention are mainly reflected in the following aspects:

[0053] Firstly, its sampling strategy is too simple, often removing a large number of potentially useful majority-class samples in the early iterations, resulting in information loss. Secondly, the algorithm lacks an effective mechanism to distinguish and utilize the performance differences of different classifiers, which limits its performance on complex datasets. Moreover, the basic classifiers used in traditional algorithms are usually relatively simple and difficult to capture complex patterns in the data. Finally, the algorithm does not consider the different costs that different error types (such as false positives and false negatives) may bring when processing samples, which is a serious defect in many practical applications.

[0054] The improvement of the present invention not only significantly improves the recognition ability of minority-class samples but also maintains good performance on majority-class samples. This balanced improvement is of great significance in practical applications. For example, in the credit card fraud detection scenario, the improved algorithm can identify more fraud transactions (the recall rate is increased by 12.30%), while maintaining a low false positive rate (the precision rate of the majority class is increased by 0.71%). This means that banks can more effectively prevent fraud risks while minimizing the inconvenience caused to normal customers.

[0055] In addition, the method of the present invention realizes adaptive learning of the dataset features by dynamically adjusting the sampling strategy and classifier weights. This enables the algorithm to maintain stable performance in the face of different degrees of class imbalance and enhances its applicability in various practical scenarios.

[0056] Generally speaking, the improved method proposed by the present invention not only fills the gap in the existing technology theoretically but also demonstrates significant performance improvement in practice. It provides a new idea for dealing with the class imbalance problem and is expected to play an important role in many key fields such as finance, healthcare, and security, making important contributions to improving the accuracy and reliability of decision-making. Brief Description of the Drawings

[0057] Figure 1 is the flowchart of the method of the present invention.

[0058] Figure 2 is the flowchart of the dynamic sampling of the present invention.

[0059] Figure 3 is the flowchart of training the classifier of the current iteration of the present invention.

[0060] Figure 4 is the flowchart of updating the majority-class sample set of the present invention.

[0061] Figure 5 is the flowchart of the prediction step of the present invention.

[0062] Figure 6 This is a performance comparison chart of the algorithm of the present invention and the original algorithm. Specific implementation manners

[0063] Please refer to Figure 1-6 , the present invention discloses an improved BalanceCascade data processing method. This method is mainly used to solve the problem of class imbalance in classification problems, and is particularly suitable for the situation where the proportion of positive and negative samples is seriously imbalanced in binary classification tasks.

[0064] The method of the present invention includes an acquisition step, a processing step, and an output step. In the acquisition step, first, a training data set is acquired, and this data set contains multiple samples and their corresponding class labels. Preferably, these samples can be various feature vectors, such as word frequency vectors of text data, pixel matrices of image data, etc., and the class labels are usually binary, such as 0 and 1, or positive class and negative class.

[0065] In the processing step, the method first determines the majority-class samples and the minority-class samples based on the acquired training data set. This step is usually achieved by calculating the number of samples in each class. For example, in a credit card fraud detection data set, there may be 99% normal transactions (majority class), while only 1% of the transactions are fraudulent (minority class).

[0066] Next, the method performs an iterative process according to the preset number of classifiers. This number can be adjusted according to the specific problem and computing resources, and usually 5 to 10 classifiers can be selected. In each iteration, the method performs the following steps:

[0067] (1) Dynamic sampling: Sampling from the majority-class samples so that the number of sampled majority-class samples is equal to the number of minority-class samples. The purpose of this step is to balance the training data of each classifier.

[0068] (2) Combining samples: Combining the sampled majority-class samples with the minority-class samples to form the training subset of the current iteration. This subset is balanced and is conducive to training a classifier that is more sensitive to the minority class.

[0069] (3) Training a classifier: Based on the training subset of the current iteration, training a classifier. The present invention preferably uses LightGBM as the base classifier because it performs excellently in dealing with high-dimensional sparse data and has a fast training speed.

[0070] (4) Calculating the weight: Calculating a weight for the classifier trained in the current iteration. This weight reflects the ability of the classifier to recognize the minority class.

[0071] (5) Update the sample set: Update the majority class sample set and remove some correctly classified samples. This allows focusing on more difficult-to-classify samples in subsequent iterations.

[0072] In the output step, this method outputs multiple trained classifiers and their corresponding weights. These classifiers and weights form the final ensemble classifier, which can be used for subsequent prediction tasks.

[0073] The dynamic sampling step of the present invention has a special processing logic. Specifically, the method first determines whether the number of remaining majority class samples is less than or equal to the number of minority class samples. If so, all remaining majority class samples are directly used; if not, random sampling is performed from the remaining majority class samples, and the sampling quantity is equal to the number of minority class samples.

[0074] This dynamic sampling strategy has obvious advantages. In the early iterations, due to the abundance of majority class samples, random sampling can be performed to maintain data diversity. In the later iterations, when the number of majority class samples decreases, directly using all remaining samples can make full use of information and avoid information loss. This strategy can effectively balance the relationship between data diversity and information utilization.

[0075] For example, assume there are 1000 majority class samples and 100 minority class samples in the initial dataset, and the preset number of classifiers is 5. In the first iteration, 100 samples will be randomly selected from the 1000 majority class samples. By the 5th iteration, if there are only 80 remaining majority class samples, then these 80 samples will be directly used without further random sampling.

[0076] The present invention adopts a special weight calculation method when training the classifier for the current iteration. Specifically, the method first calculates the class weights, where the majority class weight is set to 1 and the minority class weight is set to the ratio of the number of majority class samples to the number of minority class samples. Then, based on these class weights, sample weights are assigned to each sample in the training subset. Finally, the LightGBM classifier is trained using the training subset with sample weights.

[0077] This weight calculation method takes into account the class imbalance degree in the original dataset and balances the importance of different classes by adjusting the sample weights. For example, if the ratio of majority class to minority class in the original dataset is 9:1, then during training, the weight of minority class samples will be set to 9, while the weight of majority class samples is 1. This ensures that the model pays sufficient attention to minority class samples during training.

[0078] In practical applications, the weight calculation can be fine-tuned according to specific problems. For example, a tuning factor α can be introduced such that the minority class weight is:

[0079]

[0080] Among them, N majority and N minority respectively represent the number of samples in the majority class and the minority class. When α = 1, it is the basic form described in the present invention. By adjusting the value of α (usually between 0.5 and 1), the class weights can be balanced to different extents.

[0081] Using LightGBM as the base classifier is another important feature of the present invention. LightGBM is a gradient boosting framework based on decision trees. It adopts some innovative techniques, such as histogram-based algorithms, leaf node growth strategies, etc., making it highly efficient and having good performance when dealing with large-scale data. In the present invention, some key parameters of LightGBM can be set as follows:

[0082] objective: 'binary' (for binary classification problems)

[0083] metric: 'binary_logloss' (binary logarithmic loss, used to evaluate the model performance);

[0084] learning_rate: 0.1 (learning rate, which can be adjusted according to specific problems);

[0085] num_leaves: 31 (the maximum number of leaves of a tree);

[0086] max_depth: -1 (the maximum depth of the tree, setting it to -1 means no limit);

[0087] min_child_samples: 20 (the minimum number of data on a leaf node);

[0088] These parameter settings can serve as a good starting point and can be further optimized according to specific data characteristics and problem requirements in practical applications.

[0089] These parameter settings can serve as a good starting point and can be further optimized according to specific data characteristics and problem requirements in practical applications.

[0090] Through the above design, the method of the present invention can effectively handle the class imbalance problem, improve the model's recognition ability for minority class samples, and at the same time maintain good performance for majority class samples. This improved BalanceCascade data processing method has broad application prospects in fields such as fraud detection, disease diagnosis, and abnormal event recognition.

[0091] In a preferred embodiment of the present invention, the process of calculating the current classifier weight has a special design. Specifically, the method first uses the currently trained classifier to predict the training subset. This step can evaluate the performance of the classifier on its training data and provide a basis for weight calculation.

[0092] Next, the method calculates the recall rate of the minority class samples in the prediction results. The recall rate is an important evaluation metric, especially when dealing with imbalanced datasets. It measures the ability of the classifier to correctly identify minority class samples. The formula for calculating the recall rate is as follows:

[0093]

[0094] Where TP (True Positive) represents the number of minority class samples correctly identified, and FN (False Negative) represents the number of minority class samples misclassified as the majority class.

[0095] Finally, the method uses the calculated recall rate as the weight of the current classifier. The advantage of this weight assignment strategy is that it directly reflects the classifier's ability to identify minority class samples. The higher the weight, the better the classifier performs in identifying minority class samples and should have a greater influence in the final integrated prediction.

[0096] For example, assume that in a certain iteration, the training subset contains 100 minority class samples and 100 majority class samples. The trained classifier correctly identifies 80 minority class samples, then its recall rate is 0.8, and this value will be used as the weight of the classifier. In contrast, if another classifier only correctly identifies 60 minority class samples, its weight will be 0.6 and its influence in the final prediction will be relatively small.

[0097] Preferably, to further improve the robustness of weight calculation, the precision and recall rate can be considered simultaneously, and the F1 score can be used as the weight. The formula for calculating the F1 score is:

[0098]

[0099] This method can achieve a balance between precision and recall rate and avoid the bias that may be caused by only focusing on a single metric.

[0100] The present invention adopts an innovative strategy when updating the majority class sample set. First, the method uses the currently trained classifier to predict the remaining majority class samples. The purpose of this step is to evaluate the "difficulty" of each majority class sample, that is, the possibility of them being correctly classified.

[0101] Next, this method calculates the confidence of the prediction. A simple and effective method is adopted for calculating the confidence, which is twice the absolute difference between the prediction probability and 0.5. The specific calculation formula is as follows:

[0102] Confidence=2×|P - 0.5|,

[0103] where P is the probability that the classifier predicts the sample as the positive class. This calculation method maps the confidence to the interval [0, 1], where 0 represents the most uncertain (prediction probability is 0.5), and 1 represents the most certain (prediction probability is 0 or 1).

[0104] Finally, this method removes those majority-class samples that are predicted correctly and have a confidence greater than the preset threshold. The core idea of this strategy is to retain those "difficult-to-classify" samples and remove "easy-to-classify" samples. This can make the subsequent iterations pay more attention to those difficult samples near the decision boundary, thereby improving the overall performance of the model.

[0105] The present invention sets the preset threshold to 0.8. The selection of this value is based on the following considerations:

[0106] 1. 0.8 is a relatively high threshold, which means that only those samples that are "very certain" to be classified by the classifier will be removed. This can reduce the risk of erroneously removing valuable samples.

[0107] 2. At the same time, 0.8 is not an extremely high value (such as 0.99), which means that a certain number of samples can still be removed in each iteration, ensuring the efficiency of the algorithm.

[0108] 3. In practical applications, the value of 0.8 performs well in experiments in multiple fields and has a certain degree of generality.

[0109] However, it should be noted that in specific application scenarios, this threshold may need to be fine-tuned. For example, in some high-risk fields (such as medical diagnosis), a higher threshold (such as 0.9) may be required to ensure that important samples are not erroneously removed. On the contrary, in some scenarios with higher efficiency requirements, the threshold can be considered to be slightly reduced (such as 0.75) to accelerate the convergence speed of the algorithm.

[0110] The present invention also includes a prediction step for classifying new, unseen samples. This step makes full use of the multiple classifiers and their weights obtained from the previous training to implement a powerful integrated prediction mechanism.

[0111] First, this method obtains the samples to be predicted. These samples may come from various sources, such as real-time data streams, batch data sets, etc., but must have the same feature structure as the training data.

[0112] Next, the method uses all the trained classifiers to predict the samples to be predicted. Each classifier gives a prediction result for the sample, usually a probability value, indicating the likelihood that the sample belongs to the positive class.

[0113] Finally, based on the prediction results of each classifier and their corresponding weights, the method obtains the final prediction result through weighted voting. This weighted voting mechanism fully considers the performance differences of each classifier and can produce a more reliable prediction result than a single classifier.

[0114] Preferably, in practical applications, a probability threshold (such as 0.5) can be set to convert the continuous probability value into a discrete class label. For example, if the result of weighted voting is greater than 0.5, the sample is classified as the positive class; otherwise, it is classified as the negative class. This threshold can be adjusted according to the specific application requirements to achieve a suitable balance between precision and recall.

[0115] Generally speaking, this prediction mechanism of the present invention makes full use of the advantages of ensemble learning. By integrating the "opinions" of multiple classifiers, it can obtain more robust and accurate prediction results, especially performing well when dealing with complex and class-imbalanced datasets.

[0116] In another preferred embodiment of the present invention, the specific implementation process of weighted voting is refined into three key steps. This weighted voting mechanism is the core of the method of the present invention in the prediction stage. It cleverly combines the prediction results of multiple classifiers and fully utilizes the advantages of each classifier.

[0117] First, the method calculates the sum of all classifier weights. This step seems simple but lays the foundation for subsequent normalization processing. The sum of weights reflects the "overall strength" of the entire classifier set and is an important reference value. Suppose there are 5 classifiers with weights of 0.8, 0.7, 0.9, 0.6, and 0.75 respectively, then the sum of weights is 3.75.

[0118] Next, multiply the prediction result of each classifier by its corresponding weight and accumulate them to obtain the weighted sum. This step is the core of weighted voting, which ensures that classifiers with better performance (i.e., higher weights) have a greater say in the final decision. Continuing with the above example, assume that the prediction results (probability of predicting as the positive class in a binary classification problem) of these 5 classifiers for a certain sample are 0.7, 0.6, 0.8, 0.5, and 0.7 respectively. Then the calculation process of the weighted sum is as follows:

[0119] 0.8×0.7 + 0.7×0.6 + 0.9×0.8 + 0.6×0.5 + 0.75×0.7 = 2.485,

[0120] Finally, the method compares the weighted sum with half of the total weight to obtain the final classification result. If the weighted sum is greater than or equal to half of the total weight, it is predicted as the positive class; otherwise, it is predicted as the negative class. This decision rule is equivalent to setting a threshold of 0.5 in weighted voting. In the above example, half of the total weight is 1.875, and the weighted sum of 2.485 is greater than this value, so the sample will be classified as the positive class.

[0121] The advantage of this weighted voting mechanism is that it not only considers the "opinions" of the majority of classifiers but also adjusts their influence according to the reliability of each classifier (reflected by the weights). This makes the final prediction result more reliable and stable, especially excellent in dealing with difficult-to-classify boundary cases.

[0122] In another embodiment of the present invention, a probability prediction step is introduced. This step further enhances the flexibility and application scope of the method, enabling it to not only give discrete class predictions but also provide continuous probability outputs.

[0123] The probability prediction step first obtains the sample to be predicted. This process is similar to the prediction step described above, ensuring that the format and features of the input data are consistent with the training data.

[0124] Then, the method uses all the trained classifiers to perform probability prediction on the sample to be predicted. Each classifier will output a probability value indicating the likelihood that the sample belongs to the positive class. Here, the "probability" usually refers to the confidence score inside the classifier, which is between 0 and 1.

[0125] Finally, based on the probability prediction results of each classifier and the corresponding weights, the weighted average probability is calculated as the final probability prediction result. The specific calculation formula is as follows:

[0126]

[0127] where P final is the final probability prediction result, n is the number of classifiers, w i is the weight of the i-th classifier, and p i is the probability prediction of the i-th classifier.

[0128] This probability output method has several significant advantages compared to simple class prediction:

[0129] 1. Provide richer information: The probability value reflects the confidence level of the model in the prediction, which can help users better understand and interpret the prediction result.

[0130] 2. Support flexible decision-making: Users can freely set the probability threshold according to specific application scenarios to achieve a flexible trade-off between precision and recall.

[0131] 3. Facilitate subsequent processing: In some complex decision-making systems, the probability output can be used as an intermediate result for further risk assessment or decision optimization.

[0132] Preferably, in practical applications, discrete class predictions and continuous probability predictions can be output simultaneously to meet different needs. For example, in a credit risk assessment system, both the classification results of "high risk" or "low risk" and the specific risk probabilities can be provided to offer more comprehensive information support for decision-makers.

[0133] Finally, a termination condition judgment is introduced in the iterative process of the method of the present invention. This design aims to optimize the efficiency of the algorithm, avoid unnecessary calculations, and ensure the full utilization of all valuable samples.

[0134] Specifically, at the beginning of each iteration, this method first determines whether the number of remaining majority-class samples is zero. If so, the iterative process is immediately terminated; if not, the next round of iteration continues until the preset number of classifiers is reached.

[0135] This design of the termination condition is based on the following considerations:

[0136] 1. Resource utilization: When all majority-class samples have been used or removed, continuing the iteration will not bring new information but will instead waste computing resources.

[0137] 2. Overfitting prevention: Avoid the model overfitting to minority-class samples after the majority-class samples are exhausted.

[0138] 3. Algorithm integrity: Ensure that the algorithm is executed under meaningful conditions and maintain the class balance of the training data for each classifier.

[0139] In practical applications, this termination condition and the preset number of classifiers jointly determine the stopping time of the algorithm. Generally, the number of classifiers can be set to a relatively large value (such as 10 or 20), allowing the algorithm to mainly stop when the majority-class samples are exhausted. This can ensure that the algorithm fully utilizes all valuable samples without continuing indefinitely.

[0140] Preferably, a requirement for a minimum number of classifiers can be added to the termination condition, such as training at least 3 classifiers. This can ensure that the final ensemble classifier has sufficient diversity and improve its robustness and generalization ability.

[0141] Generally speaking, the improved BalanceCascade data processing method proposed by the present invention effectively solves the problem of class imbalance through a series of innovative designs, such as dynamic sampling, weighted training, classifier weight calculation, sample update strategy, etc. This method not only improves the recognition ability of minority class samples but also maintains good performance for majority class samples, and has broad application prospects in fields such as fraud detection, disease diagnosis, and abnormal event recognition.

[0142] Through a flexible prediction mechanism and probability output, this method provides more decision-making support for practical applications, enabling it to better adapt to the needs of different scenarios. Based on the given experimental data, we can deeply analyze the performance improvement of the improved BalanceCascade algorithm compared with the traditional BalanceCascade algorithm. The following is a detailed detection result table:

[0143]

[0144]

[0145] The improved BalanceCascade algorithm of the present invention is most suitable for application in highly imbalanced binary classification problems, such as credit card fraud detection. In such a scenario, the ratio of normal transactions (majority class) to fraud transactions (minority class) may be as high as 99:1 or more extreme.

[0146] Analysis of test results:

[0147] 1. Overall performance improvement: The improved algorithm has achieved improvements in all key indicators. In particular, the accuracy has increased by 2.15% and the F1 value has increased by 5.98%. This indicates that the algorithm is generally more accurate and balanced.

[0148] 2. Significantly improved recognition ability for minority classes: The recall rate of the minority class (label 1) has increased from 0.7927 to 0.8902, an increase of 12.30%. This means that the algorithm can identify more fraud transactions, greatly reducing the risk of missed reports.

[0149] 3. Improved balance of precision: Although the precision of the minority class is relatively low (increased from 0.1879 to 0.2192), considering the high degree of data imbalance, this improvement is still meaningful. At the same time, the precision of the majority class has also increased slightly (0.71%), indicating that the algorithm does not significantly increase the false positive rate while improving the recognition ability of the minority class.

[0150] 4. Comprehensive improvement of F1 value: Both the overall F1 value (a 5.98% increase) and the F1 value of the minority class (a 15.84% increase) show significant improvement. This indicates that the algorithm has achieved a better balance between precision and recall.

[0151] 5. Improvement of macro-average metrics: The macro-average F1 value increased by 5.05%, reflecting the overall improvement of the algorithm in dealing with imbalanced data, especially without considering the number of class samples.

[0152] 6. Steady improvement of weighted-average metrics: The weighted-average F1 value increased by 1.62%. Although the amplitude is small, considering the numerical advantage of the majority class samples, this is still a meaningful improvement.

[0153] These results fully demonstrate the innovative points of the present invention:

[0154] 1. Dynamic sampling strategy: Effectively improves the recognition ability of the minority class, reflected in the significant increase in recall rate.

[0155] 2. Classifier weight mechanism: Improves the overall prediction accuracy by assigning higher weights to better-performing classifiers.

[0156] 3. Optimization of the majority-class sample removal strategy: While improving the recognition ability of the minority class, it maintains good performance for the majority class, reflected in the slight increase in majority-class metrics.

[0157] 4. LightGBM as the base classifier: Provides stronger learning ability and helps capture complex data patterns.

[0158] The improved BalanceCascade algorithm shows obvious advantages in dealing with highly imbalanced datasets. It not only significantly improves the recognition ability of the minority class (such as fraud transactions), but also maintains good performance for the majority class. This balanced improvement is extremely important for practical applications, especially in high-risk scenarios such as fraud detection and disease diagnosis. These improvements of the algorithm help reduce economic losses, improve security, and at the same time maintain a low false alarm rate, thereby improving the overall reliability and efficiency of the system.

[0159] It should be noted that: The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, and improvements made within the principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. An improved BalanceCascade data processing method, characterized in that ,include: The acquisition steps include: Obtain a training data set, wherein the training data set includes a plurality of samples and their corresponding category labels; The processing steps include: Based on the training data set, determining majority class samples and minority class samples; According to the preset number of classifiers, the following iterative process is performed: (1) Dynamically sample from the majority class samples so that the number of majority class samples after sampling is equal to the number of minority class samples; (2) Merge the majority class samples and minority class samples after sampling to form the training subset of the current iteration; (3) training the classifier of the current iteration based on the training subset; (4) Calculate the weight of the current classifier; (5) Update the majority class sample set; The output steps include: Output multiple trained classifiers and their corresponding weights.

2. The method according to claim 1, characterized in that , the dynamic sampling in the processing step specifically includes: Determine whether the number of remaining majority class samples is less than or equal to the number of minority class samples; If so, all remaining majority class samples are used directly; If not, randomly sample from the remaining majority class samples, and the number of samples is equal to the number of minority class samples.

3. The method according to claim 1, characterized in that , the processing step of training the classifier of the current iteration specifically includes: Calculate the class weights, where the majority class weight is 1 and the minority class weight is the ratio of the number of majority class samples to the number of minority class samples; Based on the class weights, assigning a sample weight to each sample in the training subset; Train the LightGBM classifier using the training subset with sample weights.

4. The method according to claim 1, characterized in that , the weight of calculating the current classifier in the processing step specifically includes: Use the currently trained classifier to make predictions on the training subset; Calculate the recall rate of minority class samples in the prediction results; The recall rate is used as the weight of the current classifier.

5. The method according to claim 1, characterized in that , the updating of the majority class sample set in the processing step specifically includes: Use the currently trained classifier to predict the remaining majority class samples; Calculate the confidence of the prediction, which is equal to twice the absolute difference between the predicted probability and 0.5; Remove the majority class samples that are predicted correctly and whose confidence is greater than a preset threshold.

6. The method according to claim 5, characterized in that , the preset threshold is 0.

8.

7. The method according to claim 1, characterized in that , also includes the prediction step: Obtain samples to be predicted; Use all trained classifiers to predict the samples to be predicted; Based on the prediction results of each classifier and the corresponding weights, the final prediction result is obtained through weighted voting.

8. The method according to claim 7, characterized in that , the weighted voting specifically includes: Calculate the sum of all classifier weights; Multiply the prediction result of each classifier by its corresponding weight and add them up to get the weighted sum; If the weighted sum is greater than or equal to half of the total weight, the prediction is positive, otherwise it is negative.

9. The method according to claim 1, characterized in that , also includes the probability prediction step: Obtain samples to be predicted; Use all trained classifiers to make probability predictions for the samples to be predicted; Based on the probability prediction results of each classifier and the corresponding weights, the weighted average probability is calculated as the final probability prediction result.

10. The method according to claim 1, characterized in that ,The iterative process in the processing step also includes the termination condition judgment: Determine whether the number of remaining majority class samples is zero; If so, terminate the iteration process; If not, continue to the next round of iteration until the preset number of classifiers is reached.