Rule generation method based on GBDT model

Through the rule generation method based on the GBDT model, the problem that rule generation in the prior art is easily trapped in local optimality, and a more global optimal rule screening and generation are achieved, which improves the stability and effect of rule generation.

CN120216544APending Publication Date: 2025-06-27SU YIN KAIJI CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311817975.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is prone to falling into local optimality when generating rules, resulting in the lift value of the combination rules being lower than expected, and it is impossible to effectively filter out the global optimal rules.

Method used

The rule generation method based on GBDT model is adopted, and the variables derived from the GBDT model are output through the training and adjustment of the GBDT model, and the model is backtracked to extract the rules and thresholds of leaf nodes, and statistical analysis is performed to filter the rules that meet the conditions.

Benefits of technology

It avoids the problem of local optimality of a single decision tree, and can filter out more global optimal rules, improving the stability and effectiveness of rule generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216544A_ABST
    Figure CN120216544A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of rule generation of GBDT models, and particularly relates to a rule generation method based on a GBDT model, which comprises the following steps: after training is completed, testing and verifying the GBDT model on a test set X2 and a verification set X3; according to the training set X1 and the GBDT model, variables derived according to the path of the GBDT model are output; according to the variable list of the training set X1, backtracking the GBDT model to derive other data sets; rule development is converted into model output, data driving is better achieved, and the effect is stable; the trouble of local optimum of a single decision tree in the prior art is avoided, and more global optimum rules can be screened out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of rule generation of GBDT models, and specifically provides a method for generating rules based on GBDT models. Background Art

[0002] The prior art generally adopts a method of separate variable statistics or a method of generating decision trees with two or more layers. Its process is generally as follows:

[0003] 1. Traverse all variables, perform binning processing on each variable, screen the optimal decision nodes under each variable, and judge whether the sample size, the proportion of bad samples, and the lift under its binning meet the expectations;

[0004] 2. Based on the first step, traverse all variables for a second time, screen the optimal decision nodes based on the two-layer traversal, and statistically judge whether the sample size and the proportion of bad samples under its binning meet the final expectations.

[0005] The defect of this method is that the generated rules are ultimately generated under local optimality. For example:

[0006] In the context of judging whether a customer is a good customer:

[0007] Rule 1 is: when the number of inquiries in 3 months is greater than 10, the lift is 2.5, or when the number of inquiries in 3 months is greater than 8, the lift is 2; then according to the rule generation method of the decision tree, the rule threshold with a higher lift will be selected: reject when the number of inquiries in 3 months is greater than 10;

[0008] Rule 2 is: reject when the current age is greater than 45 and greater than 30. It is the combined rule with Rule 1 that causes the combined lift to be equal to 3.

[0009] That is, the combined rule 1 is to reject when the number of inquiries in 3 months is greater than 10 and the age is greater than 45, and the lift is 3.

[0010] However, assuming that the threshold of Rule 1 is to reject when the number of inquiries in 3 months is greater than 8, and the threshold of Rule 2 is to reject when the age is greater than 40. In this combination, the lift of each single rule will be lower than that of each rule in the combined rule 1, but the combined rule formed by them may cause the lift to be higher than 3.

[0011] That is, the combined rule 2 is to reject when the number of inquiries in 3 months is greater than 8 and the age is greater than 40, and the lift is 3.2.

[0012] The main reason for this situation is that the rule tree method is globally optimal under a single rule, but it may only be locally optimal under two trees. Therefore, it is necessary to design a method for generating rules based on GBDT models. Summary of the Invention

[0013] In view of the above and / or the problems existing in the existing method for generating rules based on the GBDT model, the present invention is proposed.

[0014] Therefore, the object of the present invention is to provide a method for generating rules based on the GBDT model, which can solve the above-mentioned existing problems.

[0015] To solve the above technical problems, according to one aspect of the present invention, the following technical solutions are provided:

[0016] A method for generating rules based on the GBDT model includes the following steps:

[0017] S1 Define the data as the complete set X, select the complete set X, and divide the complete set into a training set X1, a test set X2, and a validation set X3;

[0018] S2 Input the training set X1 into the model to perform GBDT training on the model;

[0019] S3 Adjust the parameters of the GBDT model by combining the KS value and the AUC value to make it reach the expected effect within the range of the training set X1;

[0020] S4 After the training is completed, input the test set X2 into the GBDT model for testing, and input the validation set X3 into the GBDT model for verification;

[0021] S5 According to the training set X1 and the GBDT model, output the variables derived according to the path of the GBDT model;

[0022] S6 According to the variable list returned to the training set X1, trace back the GBDT model to derive other data sets;

[0023] S7 Extract the rules and rule thresholds involved in each leaf node;

[0024] S8 Conduct statistical analysis on the rules to judge and screen out the rules that meet the sample size, the number of bad samples, and lift;

[0025] S9 Input the test set X2 into the GBDT model for testing, and input the validation set X3 into the GBDT model for verification, as the final screened and judged GBDT model.

[0026] Preferably, the data in S1 includes: basic data, credit investigation data, and third-party data. The validation set X3 is composed of the last 20% of the data according to the time series, and the remaining data is randomly divided into the training set X1 and the test set X2 according to 7:3.

[0027] Preferably, the preset indicators in S3 are defined as follows:

[0028] TP (True Positive): The correctly predicted positive class. A sample is positive class 1 and is also predicted as positive class 1;

[0029] FN (False Negative): The incorrectly predicted negative class. The sample is positive class 1 but is predicted as negative class 0;

[0030] FP (False Positive): The incorrectly predicted positive class. The sample is negative class 0 but is predicted as positive class 1;

[0031] TN (True Negative): The correctly predicted negative class. The sample is negative class 0 and is also predicted as negative class 0;

[0032] And two values are derived therefrom as follows:

[0033] TPR (true positive rate): The calculation formula is TPR = TP / (TP + FN);

[0034] FPR (false positive rate): The calculation formula is FPR = FP / (FP + TN);

[0035] Then, max(TPR - FPR) is the KS value of the model. The KS value can intuitively represent the goodness discrimination ability of the model. The larger the KS value, the stronger the discrimination ability of the model.

[0036] Compared with the prior art:

[0037] 1. After training is completed, test and validate the GBDT model on the test set X2 and the validation set X3; according to the training set X1 and the GBDT model, output the variables derived along the path of the GBDT model; according to the variable list of the training set X1, trace back the GBDT model to derive other data sets; convert the development of rules into model output, which is more data-driven and has stable effects.

[0038] 2. Avoid the trouble of local optimality of a single decision tree before, and more globally optimal rules can be screened out. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a schematic diagram of the overall process of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0041] The present invention provides a rule generation method based on a GBDT model, including the following steps:

[0042] S1 defines the data as the complete set X, selects the complete set X, and divides the complete set into a training set X1, a test set X2, and a validation set X3;

[0043] S2 inputs the training set X1 into the model to perform GBDT training on the model;

[0044] S3 combines the KS value and the AUC value to adjust the parameters of the GBDT model to make it reach the expected effect within the range of the training set X1;

[0045] S4 after the training is completed, inputs the test set X2 into the GBDT model for testing, and inputs the validation set X3 into the GBDT model for verification;

[0046] S5 according to the training set X1 and the GBDT model, outputs the variables derived according to the path of the GBDT model;

[0047] S6 according to the variable list returned for the training set X1, backtracks the GBDT model to derive other data sets;

[0048] S7 extracts the rules and rule thresholds involved in each leaf node;

[0049] S8 performs statistical analysis on the rules to judge and screen out the rules that meet the sample size, bad sample size, and lift;

[0050] S9 inputs the test set X2 into the GBDT model for testing, and inputs the validation set X3 into the GBDT model for verification, as the GBDT model for the final screening and judgment.

[0051] Preferably, the data in S1 includes: basic data, credit investigation data, and third-party data. The validation set X3 is composed of the last 20% of the data according to the time series to form the validation set X3, and the remaining data is randomly divided into the training set X1 and the test set X2 according to 7:3.

[0052] Preferably, the preset indicators in S3 are defined as follows:

[0053] TP (True Positive): The correctly predicted positive class, a sample is positive class 1 and is also predicted as positive class 1;

[0054] FN (False Negative): The incorrectly predicted negative class, the sample is positive class 1 but is predicted as negative class 0;

[0055] FP (False Positive): The incorrectly predicted positive class, the sample is negative class 0 but is predicted as positive class 1;

[0056] TN (True Negative): The correctly predicted negative class, where the sample is the negative class 0 and is also predicted as the negative class 0;

[0057] And two values are derived therefrom as follows:

[0058] TPR (true positive rate): The calculation formula is TPR = TP / (TP + FN);

[0059] FPR (false positive rate): The calculation formula is FPR = FP / (FP + TN);

[0060] Then, max(TPR - FPR) is the KS value of the model. The KS value can intuitively represent the goodness - of - fit discrimination ability of the model. The larger the KS value, the stronger the discrimination ability of the model.

[0061] Specifically:

[0062] For the division of the training set X1, the test set X2, and the validation set X3, for example, if 1000 data are selected, they are divided into the training set X1, the test set X2, and the validation set X3 according to the ratio of 6:2:2.

[0063] As shown in the following table, the overall sample span is from 202208 to 202304. Here, the validation set X3 is selected from 202302 to 202304, and the training set X1 and the test set X2 are selected from 202208 to 202301, where the training set X1 and the test set X2 are randomly divided according to the ratio of 7:3.

[0064]

[0065] Although the present invention has been described with reference to the embodiments above, various improvements can be made to it and components therein can be replaced with equivalents without departing from the scope of the present invention. In particular, as long as there is no structural conflict, the various features in the disclosed embodiments of the present invention can be combined with each other in any way. The reason for not exhaustively describing the situations of these combinations in this specification is only to save space and resources. Therefore, the present invention is not limited to the specific embodiments disclosed in the text, but includes all technical solutions falling within the scope of the claims.

Claims

1. A rule generation method based on the GBDT model, characterized in that, It includes the following steps: S1 Define the data as the complete set X, select the complete set X, and divide the complete set into a training set X1, a test set X2, and a validation set X3; S2 Input the training set X1 into the model to perform GBDT training on the model; S3 Adjust the parameters of the GBDT model by combining the KS value and the AUC value to make it reach the expected effect within the range of the training set X1; S4 After the training is completed, input the test set X2 into the GBDT model for testing, and input the validation set X3 into the GBDT model for verification; S5 According to the training set X1 and the GBDT model, output the variables derived according to the path of the GBDT model; S6 According to the variable list of the returned training set X1, trace back the GBDT model to derive other data sets; S7 Extract the rules and rule thresholds involved in each leaf node; S8 Conduct statistical analysis on the rules to judge and screen out the rules that meet the sample size, the number of bad samples, and lift; S9 Input the test set X2 into the GBDT model for testing, and input the validation set X3 into the GBDT model for verification, as the final screened and judged GBDT model.

2. The rule generation method based on the GBDT model according to claim 1, wherein The data in S1 includes: basic data, credit investigation data, and third-party data. The validation set X3 is composed of the last 20% of the data according to the time series to form the validation set X3, and the remaining data is randomly divided into the training set X1 and the test set X2 according to 7:

3.

3. The rule generation method based on the GBDT model according to claim 1, characterized in that The preset metrics in S3 are defined as follows: TP (True Positive): The correctly predicted positive class, a sample is the positive class 1 and is also predicted as the positive class 1; FN (False Negative): The incorrectly predicted negative class, the sample is the positive class 1 but is predicted as the negative class 0; FP (False Positive): The incorrectly predicted positive class, the sample is the negative class 0 but is predicted as the positive class 1; TN (True Negative): The correctly predicted negative class, the sample is the negative class 0 and is also predicted as the negative class 0; And the following two values are derived therefrom: TPR (true positive rate): The calculation formula is TPR = TP / (TP + FN); FPR (false positive rate): The calculation formula is FPR = FP / (FP + TN); Then, max(TPR - FPR) is the KS value of the model. The KS value can intuitively represent the good and bad discrimination ability of the model. The larger the KS value, the stronger the discrimination ability of the model.