Interactive decision tree tuning method in combination with derivative variables

By combining derivative variables and interactive tuning methods in the decision tree modeling in the credit risk control field, the problems of insufficient dimensions and weak interpretation are solved, and the effectiveness and interpretability of the model are improved.

CN119941376AInactive Publication Date: 2025-05-06RUIZHI HECHUANG (BEIJING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411780876.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the field of credit risk control, the data obtained has the problem of fewer basic variable dimensions and weak interpretation of the actual meaning of the variable, which makes it difficult to directly extract effective information when modeling using decision trees, and is not suitable for directly building a model.

Method used

An interactive decision tree tuning method combining derivative variables is proposed. By defining the pre-screening rules and derivative variables in the modeling process, combining the decision tree model to model, and interactively adjusting the pre-screening rules, derivative variables and decision tree structure to improve the effectiveness and explanatory power of the model.

Benefits of technology

By removing data with low information value, the data quality is improved, the effectiveness and interpretability of the model are improved, and it is suitable for decision tree modeling in the credit risk control field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941376A_ABST
    Figure CN119941376A_ABST
Patent Text Reader

Abstract

The invention relates to the field of retail finance big data risk control, and particularly discloses an interactive decision tree tuning method in combination with derivative variables, which comprises the following steps: S1, acquiring customer historical data, and preprocessing through a business scene to determine original data; s2, exploratory analysis is carried out on the original data, and variable data features are determined; wherein the variable data features are available features of the financial decision tree; s3, performing feature combination on the actual business scene and the variable data, configuring a pre-screening rule, screening in-mold sample data, and configuring a derivative rule and derivative variable data generated by the derivative rule; s4, inputting the derivative variable data into a pre-established financial decision tree, and determining node segmentation points; s5, node segmentation is carried out in the financial decision tree according to the node segmentation points, and a segmentation result is determined; and S6, according to a segmentation result, carrying out interactive adjustment on the pre-screening rule, the derivative variables and the decision tree structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of retail financial big data risk control, and in particular to an interactive decision tree tuning method combining derivative variables. Background Art

[0002] In the business scenarios of credit risk control, decision trees play an important role in risk control modeling and strategy mining due to their intuitive and highly interpretable characteristics. However, some of the acquired data may have problems such as fewer dimensions of basic variables and weak interpretability of the actual meaning of variables. In the process of using decision tree modeling, it is difficult to directly extract effective information by constructing a model, and it is not suitable for direct model building. Therefore, it is necessary to combine business experience and actual scenarios, as well as the idea of ​​feature engineering to combine and transform variables, so as to increase the information value of the data and improve the effect and explanatory power of the model.

[0003] Based on the above objectives, the present invention proposes an interactive decision tree tuning method combined with derived variables. In the modeling process, by combining experience and business scenarios, pre-screening rules are defined on the basis of existing data to screen the input variables, define derived variables, and efficiently utilize multivariate information.

[0004] Summary of the invention

[0005] In order to solve the problems in the background technology, the acquired data will have fewer basic variable dimensions and weak explanatory power of the actual meaning of the variables. In the process of using financial decision tree modeling, it is difficult to directly extract effective information by constructing the model, and it is not suitable to directly establish the model. An interactive decision tree tuning method combined with derived variables is provided. In the modeling process, by combining experience and business scenarios, pre-screening rules are defined on the basis of existing data to screen the variables entering the model, and derived variables are defined to efficiently utilize multivariate information.

[0006] In a first aspect, the present invention provides an interactive decision tree tuning method combining derivative variables, characterized in that it includes:

[0007] Step S1: Obtain customer historical data and pre-process it through business scenarios to determine the original data;

[0008] Step S2: Perform exploratory analysis on the original data to determine the variable data features; wherein the variable data features are available features of the financial decision tree;

[0009] Step S3: Combine the actual business scenario and variable data by features, configure pre-screening rules, screen the input sample data, and configure derivative rules and derivative variable data generated by the derivative rules;

[0010] Step S4: input the derived variable data into the pre-built financial decision tree to determine the node splitting point;

[0011] Step S5: Perform node segmentation in the financial decision tree according to the node segmentation point and determine the segmentation result;

[0012] Step S6: According to the segmentation results, interactively adjust the pre-screening rules, derived variables and decision tree structure.

[0013] In combination with the first aspect, the step S1 further includes:

[0014] Obtain historical customer information and the credit performance data set of the corresponding historical customers from credit institutions, and adjust the variable data type according to the actual business scenario;

[0015] According to the proportion of each type of sample in the credit performance data set, stratified sampling is performed on the credit performance data set, and the weight variable of each type of sample after stratified sampling is determined;

[0016] According to the weight variable, set the division method of the sampling data, divide the data set and the test set, and determine the available data set.

[0017] In combination with the first aspect, the credit performance data set includes: a prediction variable, a target variable, and a weight variable; wherein,

[0018] Prediction variables are used to represent historical customer behavior information. The value types of behavior information include numeric, character, and Boolean.

[0019] The target variable is used to characterize whether a historical customer is overdue; it is expressed as a binary categorical variable, which is represented by antonym characters;

[0020] The first character of the antonym character indicates that the user has no overdue behavior;

[0021] The second character of the antonym character indicates that the user has overdue behavior.

[0022] In combination with the first aspect, in step S2:

[0023] Exploratory analysis includes but is not limited to data distribution analysis, central tendency analysis, dispersion analysis and missingness analysis;

[0024] Variable data characteristics include, but are not limited to, the distribution and dispersion of the predictor variables themselves, the correlation characteristics between the predictor variables, and the correlation characteristics between the predictor variables and the target;

[0025] Variable data features are visualized to observe variables;

[0026] Visualization methods include but are not limited to univariate box plots, histograms, multivariate scatter plots, and heat maps.

[0027] In combination with the first aspect, in step S3:

[0028] The decision impact of the pre-screening rules and the financial decision tree does not exceed the preset impact range;

[0029] The pre-screening rules filter the rule data and the modeling data of the financial decision tree through rule screening data; among them,

[0030] Rules for filtering data support logical operations, comparison operations, string operations, and bracket operations.

[0031] In combination with the first aspect, in step S5:

[0032] The node splitting points include lower-level nodes, and the lower-level nodes are split and determined based on the splitting point with the largest Gini gain among all node splitting points.

[0033] In combination with the first aspect, the interactive adjustment includes:

[0034] Predetermine the financial business needs and financial branch performance, and judge the adjustment model; among them,

[0035] The adjustment mode includes the first adjustment mode and the second adjustment mode:

[0036] The first adjustment mode is used to adjust the decision tree branch structure or node split point;

[0037] The second adjustment mode is used to adjust the pre-screening rules and derived variables.

[0038] In combination with the first aspect, the interactive adjustment further includes:

[0039] A first interaction mechanism, a second interaction mechanism and a third interaction mechanism are set; wherein,

[0040] The first interactive mechanism is used to configure the intermediate caller of the pre-screening rule and to adjust and call the pre-screening rule;

[0041] The second interactive mechanism is used to configure the definition and calculation method of derived variables and to redefine and adjust derived variables;

[0042] The third interaction mechanism is used to configure the structural distribution switching mechanism of the decision tree structure and perform sequential switching of the decision tree structure.

[0043] In combination with the first aspect, the interactive adjustment further includes:

[0044] Setting a trigger condition according to the first interaction mechanism, the second interaction mechanism, and the third interaction mechanism;

[0045] Determine, according to the triggering condition, the execution sequence of the first interaction mechanism, the second interaction mechanism, and the third interaction mechanism on the computing unit;

[0046] Determine the dependencies of the adjustment tasks based on the execution sequence, and perform associated supervision when making interactive adjustments.

[0047] In combination with the first aspect, the interactive adjustment of the decision tree structure further includes:

[0048] Initialize the branch sequence of the decision tree structure and generate a personalized deployment structure;

[0049] The personalized deployment structure of each branch is sent to the decision tree structure, and the supervision coefficient of the decision tree structure is generated based on the hypernetwork of the attention mechanism. The hypernetwork is trained to perform branch supervision of the interactive adjustment of the decision tree structure and determine the adjustment result of the interactive adjustment.

[0050] The beneficial effects of the present invention are:

[0051] The present invention proposes an interactive decision tree tuning method combined with derived variables. In the modeling process, pre-screening rules and derived variables are defined based on experience, business scenarios and data characteristics, combined with decision tree model modeling, and customized rules and variables are interactively adjusted and optimized based on the model effect, thereby removing data with low information value, improving data quality, and enhancing model effect and interpretability.

[0052] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.

[0053] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0055] Figure 1 A flowchart of a method for an interactive decision tree tuning method incorporating derived variables;

[0056] Figure 2 A data collection flow chart for an interactive decision tree tuning method combining derived variables;

[0057] Figure 3 Schematic diagram of the interactive supervision of different mechanisms in an interactive decision tree tuning method incorporating derived variables. DETAILED DESCRIPTION

[0058] The terms or words used in this specification and claims should not be interpreted as limited to their ordinary meanings or meanings in dictionaries, and are based on the concept that the inventor can appropriately define the terms in order to best describe the principles of the present invention, and should be understood as meanings and concepts consistent with the technical ideas of the present invention.

[0059] Throughout the specification, when a certain part "includes" a certain structural element, it means that other structural elements may be further included, rather than excluding other structural elements, unless otherwise specified. In addition, the terms "...part", "...machine", "module" and "device" described in the specification refer to a unit that processes at least one function or operation, which can be implemented by a combination of hardware and / or software.

[0060] Throughout the specification, the term "and / or" should be understood to include all possible combinations from one or more related items. For example, the meaning of "the first item, the second item and / or the third item" refers to all possible combinations that can be represented from the first item, the second item or the third item and two or more of the first, second or third items.

[0061] In the various steps throughout the specification, identification symbols (e.g., a, b, c, ...) are used for ease of explanation, but the identification symbols do not limit the order of the various steps, and the execution order of the various steps may be different from the order in which they are recorded, unless the context clearly stipulates a specific order. In other words, the various steps can be executed in the order in which they are recorded, can be executed substantially simultaneously, or can be executed in the reverse order.

[0062] Embodiment 1:

[0063] like Figure 1 As shown, the present invention provides an interactive decision tree tuning method combining derivative variables, comprising:

[0064] Step S1: Obtain customer historical data and pre-process it through business scenarios to determine the original data;

[0065] Step S2: Perform exploratory analysis on the original data to determine variable data features; wherein the variable data features are available features of the financial decision tree; the available features include preferred features;

[0066] Step S3: Combine the actual business scenario and variable data by features, configure pre-screening rules, screen the input sample data, and configure derivative rules and derivative variable data generated by the derivative rules;

[0067] Step S4: input the derived variable data into the pre-built financial decision tree to determine the node splitting point;

[0068] Step S5: Perform node segmentation in the financial decision tree according to the node segmentation point and determine the segmentation result;

[0069] Step S6: According to the segmentation results, interactively adjust the pre-screening rules, derived variables and decision tree structure.

[0070] In the specific implementation process:

[0071] As in step S1, historical customer data is obtained from credit institutions, including customer personal information, credit history, repayment behavior, etc. The data is cleaned and missing values ​​and outliers are processed. According to the business scenario, the data is sampled, such as selecting data from the past three years, to obtain the original data set.

[0072] As in step S2, exploratory data analysis (EDA) is performed on the original data set, including descriptive statistics, visualization, etc., to summarize data characteristics such as the distribution, correlation, and outliers of the variables.

[0073] As shown in step S3, based on the results of EDA and business knowledge, pre-screening rules are defined. For example, customers with a credit record of less than two years or customers under the age of 18 are excluded. These rules are used to screen out data samples for modeling.

[0074] As in step S4, based on the business logic and variable characteristics, a derived variable is defined. For example, a derived variable "repayment rate" = (amount repaid / amount to be repaid) is created. The repayment rate of each customer is calculated and added to the data set.

[0075] As shown in step S5, the derived variables and other important variables are used as features to construct a decision tree model. In the decision tree, the derived variable "repayment rate" is used to find the split point and split the node.

[0076] As in step S6: Based on the performance of the model on the validation set, evaluate the model's accuracy, recall and other indicators. If the model performs poorly, make the following adjustments: Adjust the pre-screening rules, such as loosening or tightening the screening conditions. Adjust the definition of the derived variables, such as changing the calculation formula or introducing new derived variables. Adjust the structure of the decision tree, such as changing the maximum depth, minimum number of split samples and other parameters.

[0077] After data processing, feature engineering and model construction, the present invention finally obtains a decision tree model that can be used to predict customer defaults. In addition, the prediction performance of the model can be improved through continuous adjustment and optimization. In this embodiment, historical customer data includes data such as customer historical behavior, consumption habits, personal information, and good or bad labels based on factors such as whether the customer has had overdue behavior. In this embodiment, the data characteristics of the variable refer to the characteristics of the predictor variable, including but not limited to the distribution and discreteness of the predictor variable itself, the correlation characteristics between the predictor variables, the correlation characteristics between the predictor variables and the target, etc. In this embodiment, experience refers to some knowledge summarized in the practice of the credit field. In this embodiment, the business scenario establishes the scenario of model analysis problems. In this embodiment, model performance refers to the effect of the decision tree model, which needs to be selected in combination with the actual situation and can be measured by indicators such as accuracy, recall, and lift.

[0078] The working principle of the above technical solution is:

[0079] First, obtain variable data characteristics through data preprocessing and exploratory data analysis; second, define pre-screening rules to screen input samples and define derived variables. In the modeling process, you can interactively modify the pre-screening rules and derived variables based on the model effect.

[0080] The beneficial effects of the above design are:

[0081] The present invention proposes an interactive decision tree tuning method combined with derived variables. In the modeling process, pre-screening rules and derived variables are defined based on experience, business scenarios and data characteristics, combined with decision tree model modeling, and customized rules and variables are interactively adjusted and optimized based on the model effect, thereby removing data with low information value, improving data quality, and enhancing model effect and interpretability.

[0082] Embodiment 2:

[0083] On the basis of the above-mentioned embodiment 1, Figure 2As shown, S1: obtain historical customer data from a credit institution, process and sample the data according to the business scenario to obtain the original data set, and the specific steps include: obtain the credit performance data set of historical customers from the credit institution, adjust the variable data type according to the actual business scenario; stratify and sample the data samples according to the proportion of each type of samples in the data, and obtain weight variables from the sampled data; divide the sampled data into training sets and test sets based on the specified method to obtain a usable data set. Preprocess the original data set, including removing invalid data such as missing values ​​and outliers, and converting categorical variables into binary or multi-level label representations. Define pre-screening rules based on experience and business scenarios. These rules can be explicit (for example, income is greater than or equal to a certain threshold) or implicit (for example, judging the loan risk of the borrower based on his occupation). By applying these rules to the original data set, high-quality data samples that can be used for subsequent modeling can be screened out. For each screened data sample, extract the corresponding features, such as the indicators in the scorecard. These features can be custom generated according to business needs and technical implementations. Use trained models, such as decision trees, random forests, etc., to train on these high-quality data samples. In the training process, cross validation can be selectively performed on the validation set to balance model performance and generalization ability. According to the performance of the model on the test set, some hyperparameters can be adjusted, such as the depth of the decision tree, pruning parameters, etc., or the structure of the model can be optimized. These adjustments can help the model better fit the training data set and achieve better prediction results on new data. The data types supported by the prediction variables, character type and Boolean type, do not include other types, such as date and time types, etc. The training set is a data set used to train and generate a decision tree model, and the test set is a data set used to test and measure the model effect. In the present embodiment, the specified method is a method of dividing the training set and the test set selected according to the actual situation, such as dividing by a fixed ratio (hold-out method), K-fold cross validation method, custom test set, etc.

[0084] The beneficial effects of the above design are:

[0085] The present invention adjusts the data format in combination with business scenarios and actual conditions, standardizes the data format, and improves the efficiency of memory space utilization; reduces the sample size and improves the model calculation efficiency through sampling, and increases the generalization and robustness of the model by using a test set.

[0086] Embodiment 3:

[0087] On the basis of the above-mentioned Example 2, the available data sets include: predictor variables, target variables, and weight variables; predictor variables are behavioral information of credit users, and their value types can be numeric, character, and Boolean; target variables are obtained based on whether the credit customer is overdue, and are binary classification variables, represented by good and bad, where good means that the user has no overdue behavior, and bad means that the user has overdue behavior; weight variables are obtained after sampling samples. In business scenarios, if an institution has a long history of lending, a large amount of data will be accumulated and usually the number of good samples far exceeds the number of bad samples. Therefore, samples will be sampled to improve calculation efficiency, so the data set will include weight variables.

[0088] The beneficial effects of the above scheme are:

[0089] It is clarified that the data set must include predictor variables, target variables, weight variables, and the meaning of each part.

[0090] Embodiment 4:

[0091] On the basis of the above-mentioned embodiment 1, the S2: based on the original data set, summarize the data characteristics of the predictor variables through exploratory data analysis, and the specific steps include: observe the data distribution, central tendency, degree of dispersion, missing information and other information of the predictor variables through exploratory data analysis; wherein, the data characteristics of the variables include but are not limited to the distribution, dispersion and other characteristics of the predictor variables themselves, the correlation characteristics between the predictor variables, the correlation characteristics between the predictor variables and the target, etc.; further observe the data through data visualization, and summarize the data characteristics in combination with the target variable; wherein, the data visualization method includes but is not limited to the box plot and histogram of the single variable, the scatter plot and heat map of the multivariate, etc. In this embodiment, exploratory data analysis refers to understanding the relationship between variables and the relationship between variables and target variables through the data set, which can help to better perform feature engineering and establish models in the later stage.

[0092] The beneficial effects of the above scheme are:

[0093] By exploring and analyzing the data, we can more comprehensively mine the information in the data and display the data characteristics.

[0094] Embodiment 5:

[0095] On the basis of the above-mentioned embodiment 1, the S3: based on experience and actual business scenarios, combined with variable data characteristics, defines pre-screening rules, and screens the data samples for modeling. The specific steps include: combining experience, business scenarios and data characteristics, formulating rules for screening samples that have less impact on the model effect; filtering data based on rules, filtering out data that meets the conditions, and using data that does not meet the conditions for modeling. Among them, the rules support numerical operations of addition, subtraction, multiplication, division, absolute value, mean, summation, exponentiation, logarithm, and rounding, logical operations of AND, OR, and NOT, comparison operations of equal to, not equal to, greater than (equal to), less than (equal to), string operations and bracket operators. In this embodiment, screening the sample for modeling is to use data that is not hit by the rules for modeling.

[0096] The beneficial effects of the above scheme are:

[0097] Based on prior experience, some known and redundant samples are excluded, which reduces the amount of data, accelerates model training, improves model accuracy and sensitivity, and thus improves model effects.

[0098] Embodiment 6:

[0099] On the basis of the above-mentioned embodiment 1, the said S4: based on experience and actual business scenarios, combined with variable data features, defines derived variables and generates corresponding data, and the specific steps include: combining experience, business scenarios and data features, defining derived variables; based on the derived variable definition, generating corresponding data. Among them, the operators that can be used to define derived variables include numerical operations of addition, subtraction, multiplication, division, absolute value, mean, summation, exponentiation, logarithm, and rounding, logical operations of AND, OR, and NOT, comparison operations of equal, not equal, greater than (equal), less than (equal), string operations and bracket operators.

[0100] The beneficial effects of the above scheme are:

[0101] By combining and transforming variables, the information value of the data is increased, information is used more efficiently, and a better model differentiation effect is achieved.

[0102] Embodiment 7:

[0103] Based on the above-mentioned Example 1, the S5: in the decision tree, using the derived variable data to find the splitting point and split the node, the specific steps include: based on the derived variable data generated in S4, finding all feasible splitting points; calculating the Gini gains of all splitting points, selecting the splitting point with the largest Gini gain for splitting, and generating the lower-level nodes.

[0104] The beneficial effects of the above scheme are:

[0105] The derived variables are combined in the decision tree model for segmentation, which effectively utilizes the prior information.

[0106] Embodiment 8:

[0107] Based on the above-mentioned Example 1, the S6: based on the model performance, interactively adjust the pre-screening rules, derived variables and decision tree structure, specifically including: based on business needs and branch performance, adjusting the variables used in the decision tree branch structure or nodes; based on business needs and branch performance, adjusting the pre-screening rules and derived variables.

[0108] Embodiment 9: The interactive adjustment further includes:

[0109] A first interaction mechanism, a second interaction mechanism and a third interaction mechanism are set; wherein the first interaction mechanism is used to configure the intermediate caller of the pre-screening rule and adjust the pre-screening rule call; the second interaction mechanism is used to configure the definition and calculation method of the derived variables and redefine and adjust the derived variables; the third interaction mechanism is used to configure the structural distribution switching mechanism of the decision tree structure and switch the decision tree structure order.

[0110] In this embodiment, the main function of the first interactive mechanism is to configure the intermediate caller of the pre-screening rule and adjust the pre-screening rule. This process mainly requires an interface for interaction, which can be implemented through a graphical user interface (GUI) or a command line interface (CLI). The user enters the pre-screening rules they set on this interface, and then the program will check their validity. If the rule is valid, it will be saved. If the rule is invalid, it will prompt the user to reset it. After the setting is completed, the user can pass the rule to the next link for further processing. In this embodiment, the second interactive mechanism is mainly used to configure the definition and calculation method of the derived variable and redefine the derived variable. This process also requires an interface for interaction, which can be a graphical interface similar to the first interactive mechanism, or a command line interface. The user enters their new definition of the derived variable on this interface, and the program will check its validity. If the definition is valid, it will be saved. If the definition is invalid, it will prompt the user to reset it. After the setting is completed, the user can pass the new definition to the next link for further processing. In this embodiment, the third interactive mechanism is mainly used to configure the structural distribution switching mechanism of the decision tree structure and perform the sequential switching of the decision tree structure. This process also requires an interface for interaction, which can be a graphical interface or a command line interface. Users enter their new requirements for the decision tree structure on this interface, and the program will check its validity. If the requirements are valid, they will be saved. If the requirements are invalid, the user will be prompted to reset them. After the settings are completed, the user can pass the new requirements to the next link for further processing.

[0111] Embodiment 10:

[0112] like Figure 3 As shown, based on the above-mentioned embodiment 1, the interactive adjustment also includes: setting trigger conditions according to the first interaction mechanism, the second interaction mechanism and the third interaction mechanism; determining the execution sequence of the first interaction mechanism, the second interaction mechanism and the third interaction mechanism on the computing unit according to the trigger conditions; determining the dependency of the adjustment task according to the execution sequence, and performing associated supervision when performing interactive adjustment.

[0113] In this embodiment, trigger conditions are set according to the first interaction mechanism, the second interaction mechanism and the third interaction mechanism. These trigger conditions may include specific data changes, model performance reaching a certain threshold or other external events. When these trigger conditions are met, the corresponding interaction mechanism will be activated. According to the trigger conditions, the execution timing of the first interaction mechanism, the second interaction mechanism and the third interaction mechanism on the computing unit is determined. This means that after the trigger conditions are met, these interaction mechanisms will be executed in a predetermined order. According to the execution timing, the dependencies of the adjustment tasks are determined, and when the interaction adjustment is performed, the associated supervision is performed. This may involve checking the dependencies between the interaction mechanisms to ensure that they work correctly in sequence, while monitoring the operation of the entire system to ensure that everything is within the expected range.

[0114] Embodiment 11:

[0115] Based on the above embodiment 1, the interactive adjustment of the decision tree structure further includes:

[0116] Initialize the branch sequence of the decision tree structure and generate a personalized deployment structure;

[0117] The personalized deployment structure of each branch is sent to the decision tree structure, and the supervision coefficient of the decision tree structure is generated based on the hypernetwork of the attention mechanism. The hypernetwork is trained to perform branch supervision of the interactive adjustment of the decision tree structure and determine the adjustment result of the interactive adjustment.

[0118] In this embodiment, the branch sequence of the decision tree structure is initialized and a personalized deployment structure is generated. This is to assign a unique identifier (e.g., a unique number or letter) to each decision tree. Then, this identifier is used to generate a personalized deployment structure. This structure will contain some special instructions for interactive adjustment on our decision tree. Send the personalized deployment structure of each branch to the decision tree structure. This means that the generated personalized deployment structure will be sent to our decision tree structure. This will allow the decision tree to adjust its internal data flow as needed to better predict the customer's credit status. Generate supervision coefficients for the decision tree structure based on the hypernetwork of the attention mechanism. The purpose of this step is to generate a set of supervision coefficients based on the personalized deployment structure of the decision tree structure. These supervision coefficients will be used to guide how to make decisions. Train the hypernetwork to perform branch supervision for interactive adjustment of the decision tree structure. This means that the trained hypernetwork will be used to adjust certain branches of the decision tree. Specifically, this hypernetwork will be used to predict the loss in the current decision tree state, and certain branches of the decision tree will be updated based on this prediction. Determine the adjustment results of the interactive adjustment. The last step is to determine what impact the changes generated during the interactive adjustment process will have on the decision tree. This will help you understand if the interaction adjustment is effective, if so, keep that change, otherwise try another change.

[0119] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. An interactive decision tree tuning method combining derivative variables, characterized in that: include: Step S1: Obtain customer historical data and pre-process it through business scenarios to determine the original data; Step S2: Perform exploratory analysis on the original data to determine the variable data features; wherein the variable data features are available features of the financial decision tree; Step S3: Combine the actual business scenario and variable data by features, configure pre-screening rules, screen the input sample data, and configure derivative rules and derivative variable data generated by the derivative rules; Step S4: input the derived variable data into the pre-built financial decision tree to determine the node splitting point; Step S5: Perform node segmentation in the financial decision tree according to the node segmentation point and determine the segmentation result; Step S6: According to the segmentation results, interactively adjust the pre-screening rules, derived variables and decision tree structure.

2. The interactive decision tree tuning method combining derivative variables according to claim 1, characterized in that: The step S1 further comprises: Obtain historical customer information and the credit performance data set of the corresponding historical customers from credit institutions, and adjust the variable data type according to the actual business scenario; According to the proportion of each type of sample in the credit performance data set, stratified sampling is performed on the credit performance data set, and the weight variable of each type of sample after stratified sampling is determined; According to the weight variable, set the division method of the sampling data, divide the data set and the test set, and determine the available data set.

3. The interactive decision tree optimization method in combination with derivative variables according to claim 2, characterized in that: The credit performance data set includes: predictor variables, target variables, and weight variables; wherein, Prediction variables are used to represent historical customer behavior information. The value types of behavior information include numeric, character, and Boolean. The target variable is used to characterize whether a historical customer is overdue; it is expressed as a binary categorical variable, which is represented by antonym characters; The first character of the antonym character indicates that the user has no overdue behavior; The second character of the antonym character indicates that the user has overdue behavior.

4. The interactive decision tree optimization method combining derivative variables according to claim 1, characterized in that: In step S2: Exploratory analysis includes but is not limited to data distribution analysis, central tendency analysis, dispersion analysis and missingness analysis; Variable data characteristics include, but are not limited to, the distribution and dispersion of the predictor variables themselves, the correlation characteristics between the predictor variables, and the correlation characteristics between the predictor variables and the target; Variable data features are visualized to observe variables; Visualization methods include but are not limited to univariate box plots, histograms, multivariate scatter plots, and heat maps.

5. The interactive decision tree optimization method combining derivative variables according to claim 1, characterized in that: In step S3: The decision impact of the pre-screening rules and the financial decision tree does not exceed the preset impact range; The pre-screening rules filter the rule data and the modeling data of the financial decision tree through rule screening data; among them, Rules for filtering data support logical operations, comparison operations, string operations, and bracket operations.

6. The interactive decision tree optimization method combining derivative variables according to claim 1, characterized in that: In step S5: The node splitting points include lower-level nodes, and the lower-level nodes are split and determined based on the splitting point with the largest Gini gain among all node splitting points.

7. The interactive decision tree optimization method combining derivative variables according to claim 1, characterized in that: The interactive adjustment includes: Predetermine the financial business needs and financial branch performance, and judge the adjustment model; among them, The adjustment mode includes the first adjustment mode and the second adjustment mode: The first adjustment mode is used to adjust the decision tree branch structure or node split point; The second adjustment mode is used to adjust the pre-screening rules and derived variables.

8. The interactive decision tree optimization method in combination with derivative variables according to claim 1, characterized in that: The interactive adjustment also includes: A first interaction mechanism, a second interaction mechanism and a third interaction mechanism are set; wherein, The first interactive mechanism is used to configure the intermediate caller of the pre-screening rule and to adjust and call the pre-screening rule; The second interactive mechanism is used to configure the definition and calculation method of derived variables and to redefine and adjust derived variables; The third interaction mechanism is used to configure the structural distribution switching mechanism of the decision tree structure and perform sequential switching of the decision tree structure.

9. The interactive decision tree optimization method in combination with derivative variables according to claim 8, characterized in that: The interactive adjustment also includes: Setting a trigger condition according to the first interaction mechanism, the second interaction mechanism, and the third interaction mechanism; Determine, according to the triggering condition, the execution sequence of the first interaction mechanism, the second interaction mechanism, and the third interaction mechanism on the computing unit; Determine the dependencies of the adjustment tasks based on the execution sequence, and perform associated supervision when making interactive adjustments.

10. The interactive decision tree tuning method combining derivative variables according to claim 1, characterized in that: The interactive adjustment of the decision tree structure also includes: Initialize the branch sequence of the decision tree structure and generate a personalized deployment structure; The personalized deployment structure of each branch is sent to the decision tree structure, and the supervision coefficient of the decision tree structure is generated based on the hypernetwork of the attention mechanism. The hypernetwork is trained to perform branch supervision of the interactive adjustment of the decision tree structure and determine the adjustment result of the interactive adjustment.