Score card model training method and device, electronic equipment and storage medium
By constructing a target feature set and adding a third feature in financial business to characterize profit situations, the problem of binning in the existing technology cannot meet the profit maximization problem in the financial business scenarios is solved, and the score card model is effective in reflecting the profit situation.
Patent Information
- Application Number
- CN202510217903.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, the boxing method only ensures the optimal chi-square value and cannot meet the profit maximization needs in financial business scenarios.
The target feature set and the second feature are constructed in the target financial business, and the third feature is added to obtain the derivative features through internal product operation, characterize the profit situation of the target object, and box the first feature based on the derivative features. Finally, the logistic regression model is trained to obtain the score card model used to output the score card.
Through this method, the score card model can directly reflect the profit situation of the target object and meet the profit maximization needs in financial business scenarios.
Smart Images

Figure CN120216984A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method and apparatus for training a scoring card model, an electronic device, and a storage medium. Background Art
[0002] For a binary classification model, the label does not change. The chi-square value is calculated for the entity feature variables based on the binary classification label. By continuously merging intervals, the chi-square test result is ensured to be optimal. It can be seen that the current binning method only ensures that the chi-square value is optimal for all entity data. The optimal chi-square value is not equivalent to maximizing profit. The optimal chi-square result is often not the maximum profit and cannot meet the requirements of the current financial business scenario. Summary of the Invention
[0003] This application provides a method and apparatus for training a scoring card model, an electronic device, and a storage medium to solve the problem in the related art that the current binning method only ensures that the chi-square value is optimal for all entity data and cannot meet the requirements of the financial business scenario.
[0004] In a first aspect, this application provides a method for training a scoring card model, including: in a target financial business, constructing a target feature set and a second feature associated with a target object, where the target feature set includes a plurality of first features associated with the target object; the second feature represents the classification label of the target object; adding a plurality of third features to the target feature set, and obtaining a derivative feature through an inner product operation based on the plurality of third features, where the third feature represents the profit or loss of the target object; the derivative feature represents the profit situation of the target object in the target financial business; binning the plurality of first features in the target feature set based on the derivative feature; using the binned data and the second feature as a training set to train a logistic regression model to obtain a scoring card model for outputting a scoring card.
[0005] Second aspect, the present application provides a training device for a scoring card model, including: a construction module, configured to construct a target feature set and a second feature associated with a target object in a target financial service, where the target feature set includes a plurality of first features associated with the target object; the second feature represents a classification label of the target object; a processing module, configured to add a plurality of third features to the target feature set, and obtain a derived feature through an inner product operation based on the plurality of third features, where the third feature represents the profit or loss of the target object; the derived feature represents the profit situation of the target object in the target financial service; a binning module, configured to bin the plurality of first features in the target feature set based on the derived feature; a training module, configured to use the binned data and the second feature as a training set to train a logistic regression model to obtain a scoring card model for outputting a scoring card.
[0006] Third aspect, the present application provides an electronic device, including: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor connected to the at least one bus; at least one memory connected to the at least one bus, where the processor is configured to execute the training method of the scoring card model described in the first aspect of the present application above.
[0007] Fourth aspect, the present application further provides a computer storage medium storing computer-executable instructions for executing the training method of the scoring card model described in the first aspect of the present application above.
[0008] The above technical solutions provided by the embodiments of the present application have the following advantages compared with the prior art: In the method provided by the embodiments of the present application, a target feature set and a second feature associated with a target object in a target financial service are determined, and a plurality of third features are further added, and the plurality of third features represent the profit or loss of the target object; therefore, the profit situation of the target object in the target financial service can be determined through the added plurality of third features, and then the target feature set is binned based on the profit situation, so as to ensure maximum profit. Finally, the logistic regression model is trained using the binned data as a training set to obtain a scoring card model for outputting a scoring card. That is, the scoring card output by the scoring card model obtained through the embodiments of the present application can directly reflect the profit situation of the current object and can meet the profit requirements in the current financial service scenario. Description of the Drawings
[0009] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.
[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0011] One or more embodiments are exemplarily illustrated by the pictures in the corresponding accompanying drawings. These exemplary illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, unless otherwise stated, and the drawings in the figures do not constitute a proportional limitation.
[0012] Figure 1 It is a flowchart of a method for training a scoring card model provided by an embodiment of the present application;
[0013] Figure 2 It is a flowchart of a machine learning scoring card development method based on maximizing the benefits in the financial field provided by an embodiment of the present application;
[0014] Figure 3 It is a schematic structural diagram of a device for training a scoring card model provided by an embodiment of the present application;
[0015] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific embodiments
[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.
[0017] The following disclosure provides many different embodiments or examples for implementing different structures of the present invention. To simplify the disclosure of the present invention, the components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present invention. In addition, the present invention may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed.
[0018] To solve the problem in the related art that the current binning method only ensures that the chi-square value is optimal for all entity data and cannot meet the requirements in the financial business scenario, the present application provides a method for training a scoring card model, as Figure 1As shown, the steps of the method include:
[0019] Step 101, in the target financial business, construct a target feature set and a second feature associated with the target object, where the target feature set includes multiple first features associated with the target object; the second feature represents the classification label of the target object.
[0020] In a specific application scenario, financial domain services include banking services, securities services, insurance services, investment management services, etc.; among them, banking services include deposit services, loan services, payment and settlement services, foreign exchange services, credit card services; securities services include stock trading, bond trading, investment banking services, asset management, etc.; insurance services include life insurance, property insurance, etc.; investment management services include fund management, wealth management, pension management, etc. In a specific example, taking the credit card service of the banking service as the target financial business and the target object as a user, the corresponding target feature set is the attribute features associated with the target object, such as attribute data like age, education level, income, credit information, family status, whether there has been any legal involvement, etc.; whether the target object has overdue repayment is used as the second feature. In other examples, the target object can also be an enterprise, an institution, etc.
[0021] Step 102, add multiple third features to the target feature set, and obtain a derivative feature through inner product operation based on the multiple third features, where the third feature represents the gain or loss of the target object; the derivative feature represents the profit situation of the target object in the target financial business.
[0022] In a specific example, the multiple third features can be the expected return and actual loss dimensions, that is, the profit situation of the target object in the target financial business can be determined through the multiple third features. For example, if the expected return of the target object is 500, the weight of the expected return is 0.05, the actual loss is 9000, and the weight of the actual loss is 0.95, then the profit situation = 500 x 0.05 - 9000 x 0.95 = -8525, that is, the profit situation of the target object is -8525.
[0023] Step 103, bin the multiple first features in the target feature set based on the derivative feature.
[0024] In this regard, each first feature in the target feature set represents a dimension, and all features can be divided into numerical variable features and categorical variable features. All numerical variable features are binned uniformly, and all categorical variable features are binned uniformly; for example, for user Zhang San, the first dimension is the first feature of age 35, the second dimension is the first feature of male gender, the third dimension is the first feature of postgraduate education level, the fourth dimension is the first feature of credit information, and the first feature of no overdue in the past three months, etc.
[0025] Step 104: Use the binned data and the second feature as the training set to train a logistic regression model, and obtain a scorecard model for outputting a scorecard.
[0026] In a specific example, the WOE-mapped encoded data and the second feature can be used as the training set to establish a logistic regression model. Then, common parameters such as the base score, the point of double odds (PDO) when the ratio doubles, the odds of good to bad, the normalized upper and lower limits of the score range, the number of logistic regression iterations, and regularization are input into the model, and the logistic regression result is converted into the form of a scorecard.
[0027] Through the above Steps 101 to 104, the target feature set and the second feature associated with the target object in the target financial service are determined, and multiple third features are further added, and the multiple third features represent the gains or losses of the target object. Therefore, the profit situation of the target object in the target financial service can be determined through the added multiple third features, and then the target feature set is binned based on the profit situation, so as to ensure maximum profit. Finally, the binned data is used as the training set to train a logistic regression model, and a scorecard model for outputting a scorecard is obtained. That is, the scorecard output by the scorecard model obtained through the embodiments of the present application can directly reflect the profit situation of the current object and can meet the profit requirements in the current financial service scenario.
[0028] In the embodiments of the present application, the profit situation of the target object in the financial service can be determined by adding multiple features representing gains or losses. Further, the added features can be the expected gain and the actual loss. Based on this, for the method of obtaining the derivative feature of the second feature through the inner product operation based on multiple third features involved in Step 102 above, it can further include:
[0029] Step 11: Set corresponding weights for the multiple third features respectively;
[0030] Step 12: Determine the sum of the products of the third features and the corresponding weights, and determine the profit situation of the target object in the target financial service based on the sum of the products to obtain a derivative feature.
[0031] For Steps 11 and 12 above, when the number of the third features is 2, and the two third features are expected revenue and actual loss respectively, and their corresponding weights are 0.05 and 0.95. In a specific example, the expected revenue is 500 and the actual loss is 9000, then the derived feature = 500×0.05 - 9000×0.95 = -8525. The above is only an example. The third feature can be other features that can represent profit, such as revenue and loss in a certain period. The weights of the third features can also be set accordingly according to actual needs.
[0032] When binning the target feature set, the type of the features needs to be considered, and different types of features are binned separately. Specifically, the feature types in the target feature set in the embodiments of the present application include numerical features and categorical variable features, that is, the numerical features are binned separately, and the categorical variable features are binned separately. In this regard, for the method of feature binning for multiple first features in the target feature set based on the derived features involved in Step 103 above, it can further include:
[0033] Step 21: Sort the numerical features in the target feature set from small to large, merge the sorted features pairwise in sequence, and determine the derived features corresponding to the merged features;
[0034] Step 22: Determine the mean value of the derived features corresponding to the categorical variable features in the target feature set, and sort the categorical variable features from small to large based on the determined mean value of the derived features;
[0035] Step 23: Bin the merged numerical features and the sorted categorical variable features based on the bin number limit requirement and the derived features to obtain an initial binning result that meets the bin number limit;
[0036] Step 24: Adjust the initial binning result to obtain the final binning result.
[0037] In a specific example, the numerical features can be features where the numbers are continuous, such as the age and income of the target object, while the categorical variable features can be features that have no size relationship with each other, such as the provinces and occupations associated with the target object.
[0038] Taking the income of the target object as an example, there is a monthly income for each of the 12 months in a year. Sort the monthly incomes for each month, and then merge them in pairs. For example, merge the monthly income of the first month after sorting with the monthly income of the second month, merge the monthly income of the second month with the monthly income of the third month, and so on until the monthly income of the eleventh month is merged with the monthly income of the twelfth month. The feature merging is equivalent to the merging of the corresponding derivative features of the features. If the current binning number limit for numerical features is 7, then after the first merge, select the merge method with the largest variance. For example, if the variance is the largest after merging the monthly income of the third month with the monthly income of the fourth month, then only merge the monthly income of the third month with the monthly income of the fourth month, and do not merge the monthly incomes of other months. At this time, the number of monthly incomes is 11. Since the binning limit is 7, the above method of merging needs to be performed again, and the merge method with the largest variance is selected again until the number of monthly incomes is 7, so as to meet the current binning number limit.
[0039] Since there is no size relationship between each value of the categorical variable feature, calculate the mean value of the derivative feature to make there be a size relationship between each value, and convert the categorical variable feature into a numerical feature; then sort the mean values of the derivative features, and then merge them in pairs, and select the merge method with the largest variance of the derivative features until the binning meets the binning number limit, that is, the binning method is similar to the binning method of numerical features. Selecting the merge with the largest variance of the derivative features is to ensure that the profit difference between different bins is significant, so as to optimize the profit discrimination ability of the scoring card model.
[0040] In the embodiment of the present application, the binning result can be further made to meet the preset conditions so that the final binning result can more directly represent the profit situation. For example, the binning result can be made to meet the monotonicity requirement. Based on this, for the method of adjusting the initial binning result to obtain the final binning result involved in step 24 above, it can further include:
[0041] Step 31, determine the WOE (Weight of Evidence) value of each bin in the initial binning result, and map the WOE encoding to each bin of data;
[0042] In the embodiment of the present application, after binning, the WOE value is used to replace the original data. WOE = ln (the proportion of bad samples in each bin to the total number of bad samples / the proportion of good samples in each bin to the total number of good samples). WOE is a real number that can be positive or negative. Therefore, determining the WOE value of each bin and mapping the WOE encoding to each bin of data includes:
[0043] Step 41, determine the first ratio of the proportion of bad samples in each bin to all the bad samples in all bins;
[0044] Step 42: Determine the second ratio of the proportion of good samples in each bin to all good samples in all bins;
[0045] Step 43: Determine the ratio of the first ratio to the second ratio as the WOE value, and map the WOE code to each bin of data based on the WOE value.
[0046] In a specific example, good samples are data normally associated with the target object, while bad samples are abnormal data. Taking age as an example, if the current financial business is a credit card business and the feature data is age, then data of minors in terms of age is abnormal data, that is, bad samples; data of adults in terms of age is normal data, that is, good samples.
[0047] Step 32: Sort the binning results according to the WOE code so that the binning results satisfy monotonicity.
[0048] In a specific example, the monotonicity of binning means that for the binned data, it is required that the WOE of the first bin > the WOE of the second bin > the WOE of the third bin > the WOE of the fourth bin > the WOE of the fifth bin, or the WOE of the first bin < the WOE of the second bin < the WOE of the third bin < the WOE of the fourth bin < the WOE of the fifth bin. It can be seen that the binning monotonicity means that the WOE value after binning shows an increasing or decreasing trend as the binning index increases, so that the logistic regression model can obtain better prediction results.
[0049] In order to enable the binning results to more directly and accurately reflect the profit situation, the sample proportion of the binning results can also be adjusted so that a more accurate scoring card model can be obtained when training the logistic regression model subsequently. Therefore, for the method of adjusting the initial binning results to obtain the final binning results involved in the above step 24, it can further include:
[0050] S1: Select target bins with sample proportions not meeting the preset threshold range from the initial binning results, and pairwise merge the target bins with other bins in the binning results;
[0051] S2: Determine the variance of the pairwise merged bins, and select the merged bin with the largest variance as the merged bin for this time;
[0052] S3: Loop through the above S1 and S2 until the sample proportions in all bins meet the preset threshold range to obtain the final binning results.
[0053] Through the above steps S1 to S3, the data quantities in each bin can be made as balanced as possible, avoiding the situation of too many or too few data in the bins, so that when the binned data is used as the training set subsequently, the training of the logistic regression model can be accurate.
[0054] In the embodiment of the present application, for the method of using the binned data and the second feature as the training set in step 104 above to train a logistic regression model to obtain a scoring card model for outputting a scoring card, it may further include:
[0055] Step 41, construct a logistic regression model based on the model input parameters;
[0056] Step 42, train the logistic regression model based on the training set and output evaluation metrics;
[0057] Step 43, adjust the model input parameters based on the evaluation metrics, and train the logistic regression model again to obtain a scoring card model for outputting a scoring card.
[0058] In a specific example, the input parameters may include parameters such as the base score, the score interval when the ratio doubles, the good-bad ratio, the upper and lower limits of score interval standardization, the number of logistic regression iterations, and regularization. Then, the input parameters are adjusted according to the output evaluation metrics so that the model training converges to obtain a scoring card model for outputting a scoring card. The output of this scoring card model directly reflects the profit and can meet the profit requirements in various financial scenarios.
[0059] The following is an example of the present application in combination with the specific implementation manner of the embodiment of the present application. The specific implementation manner provides a machine learning scoring card development method based on the maximization of benefits in the financial field, as Figure 2 shown. The steps of this method include:
[0060] Step 201, construct a multi-dimensional entity X feature set and a Y classification label according to the business objectives of the financial scenario;
[0061] For this, taking the credit card risk identification scenario as an example of the financial scenario business, that is, using the characteristics of past credit card users as the X features and whether they are overdue as the Y classification label; the X features include data such as the age, education level, income, credit information, family status, and whether there have been any legal issues of credit card users as the X features (there are multiple X features), and whether this user is overdue as the Y label. It can be seen that this user corresponds to multiple X features and a unique Y feature.
[0062] Step 202, based on the entity data characteristics in financial fields such as credit and insurance, add the dimensions of expected return and actual loss on the basis of the original multi-dimensional X features and Y classification labels of the data, assign different weights to the characteristics of the expected return & actual loss dimensions, and perform inner product operations to derive the Y_profit feature (corresponding to the above-derived feature);
[0063] In a specific example, assume that the expected profit is 500, the weight of the expected profit is 0.05, the actual loss is 9000, and the weight of the actual loss is 0.95. Based on this, Y_profit = 500 × 0.05 - 9000 × 0.95.
[0064] Step 203: Sort the X numerical features that need to be binned for each dimension in the data from smallest to largest. After sorting the X numerical feature variables, merge the data pairwise in sequence. For the merged data, calculate the sum result of Y_profit, continuously calculate the variance situation after merging, and select the merging method with the largest variance.
[0065] In a specific example, the numerical features in the embodiments of the present application can be variables such as age and income where the numbers are continuous and have a size relationship with each other. The categorical features can be provinces, which are variables without a size relationship with each other.
[0066] Step 204: Calculate the mean value of Y_profit for each dimension feature of the categorical features and sort them from smallest to largest in the order of the mean values. After sorting, merge the data pairwise in sequence, continuously calculate the variance situation after merging, and select the merging method with the largest variance.
[0067] Since there is no size relationship between each value of the categorical variable, in the embodiments of the present application, calculate the mean value of Y_profit to make there be a size relationship between each value, and convert the categorical variable into a numerical variable. For example, if the categorical feature is a province, and different provinces correspond to many users, and each user has its own Y_profit. At this time, only need to sum the Y_profit corresponding to each province, and then divide by the number of users corresponding to the province, which is the mean value of Y_profit. For example, for Province A, there are 5 users, and the corresponding 5 Y_profits are 100, 200, 300, 400, and 500 respectively. Then the mean value of Y_profit for Province A is 300.
[0068] Step 205: According to the requirements of the binning number limit for the users, continuously repeat steps 203 and 204 until the optimal binning result of Y_profit is achieved, where it is required that there is an obvious distinction between different bins of Y_profit.
[0069] Step 206: Adjust the binning result until the binning result meets the monotonicity requirement.
[0070] Specifically, for the binned data, it is required that the WOE of the first bin > the WOE of the second bin > the WOE of the third bin > the WOE of the fourth bin > the WOE of the fifth bin, or the WOE of the first bin < the WOE of the second bin < the WOE of the third bin < the WOE of the fourth bin < the WOE of the fifth bin. Only in this case does it meet the binning monotonicity requirement.
[0071] Step 207: Adjust the binning results according to the requirement of the proportion of samples in each bin. For the bins whose proportions do not meet the requirements, merge them with the upper and lower bins respectively, and select the merging method with the largest variance until the binning meets the requirement of the sample proportion.
[0072] Step 208: Calculate the WOE value of each bin after binning and map the WOE encoding to the data.
[0073] Step 209: Use the data after WOE mapping encoding as the training set to establish a logistic regression model. Input parameters such as the base score, the score interval pdo when the ratio doubles, the odds of good and bad, the upper and lower limits of score interval normalization, the number of iterations of logistic regression, regularization, etc., and convert the logistic regression result into the form of a scoring card.
[0074] Step 210: Automatically output the prediction results and evaluation indicators of the scoring cards for the training set, validation set, and test set, and output the score interval and score mapping relationship of the feature variables of the scoring card.
[0075] Step 211: Iteratively optimize the scoring model according to the model results.
[0076] Step 212: Automatically save the scoring card model.
[0077] Based on the above steps 201 to 212, if a user inputs a bank credit product data including multi-dimensional feature information of entities with past credit records, a scoring card model needs to be constructed based on the data. After data preprocessing, call the feature engineering module, perform feature binning according to the profit maximization model, and the binning results can be manually adjusted after binning. Create a logistic regression model according to the binning results and model input parameters, automatically train the model and output the corresponding evaluation indicators of the model. The user tunes and iterates the model parameters (such as PDO, base score) according to the model evaluation indicators. After the model training is completed, automatically perform score mapping according to the relevant min-max intervals and indicators such as ODDS and base score of the input scoring card, and output the variable score interval and score mapping relationship. Automatically save the scoring card model.
[0078] It can be seen that in the embodiment of the present application, when performing data binning for each dimension feature, the classification index results are not considered, but directly use the profit situation as the weight for binning, so as to ensure profit maximization; and the subsequent results can be directly used for the operation of financial products without additional business analysis. That is, in the stage of automatically creating a scoring card, instead of using methods such as calculating the chi-square value and decision tree binning in the related technologies, the binning method with weighted profit as the evaluation index is adopted. Compared with the related technologies based on statistical test methods, the profit situation can be directly calculated.
[0079] Corresponding to the above Figure 1 , the embodiment of the present application provides a training device for a scoring card model, as Figure 3As shown, the device includes:
[0080] A construction module 302, configured to construct a target feature set and a second feature associated with a target object in a target financial service, where the target feature set includes a plurality of first features associated with the target object; the second feature represents a classification label of the target object.
[0081] A processing module 304, configured to add a plurality of third features to the target feature set, and obtain a derived feature through an inner product operation based on the plurality of third features, where the third feature represents the gain or loss of the target object; the derived feature represents the profit situation of the target object in the target financial service.
[0082] A binning module 306, configured to bin the plurality of first features in the target feature set based on the derived feature.
[0083] A training module 308, configured to use the binned data and the second feature as a training set to train a logistic regression model to obtain a scoring card model for outputting a scoring card.
[0084] Through the device according to the embodiment of the present application, a target feature set and a second feature associated with a target object in a target financial service are determined, and a plurality of third features are further added, and the plurality of third features represent the gain or loss of the target object; therefore, the profit situation of the target object in the target financial service can be determined through the added plurality of third features, and then the target feature set is binned based on the profit situation, so as to ensure maximum profit. Finally, the binned data is used as a training set to train a logistic regression model to obtain a scoring card model for outputting a scoring card. That is, the scoring card output by the scoring card model obtained through the embodiment of the present application can directly reflect the profit situation of the current object and can meet the profit requirements in the current financial service scenario.
[0085] In an optional implementation manner of the embodiment of the present application, the processing module in the embodiment of the present application may include: a setting unit, configured to set corresponding weights for the plurality of third features respectively; a first processing unit, configured to determine the sum of the products of the third features and the corresponding weights, and determine the profit situation of the target object in the target financial service based on the sum of the products to obtain a derived feature.
[0086] In an alternative implementation manner of the embodiment of the present application, the binning module in the embodiment of the present application may include: a second processing unit, configured to sort the numerical features in the target feature set from small to large, and sequentially merge the sorted features in pairs, and determine the derivative features corresponding to the merged features; a third processing unit, configured to determine the mean value of the derivative features corresponding to the categorical variable features in the target feature set, and sort the categorical variable features from small to large based on the determined mean value of the derivative features; a fourth processing unit, configured to bin the merged numerical features and the sorted categorical variable features based on the binning number limit requirement and the derivative features to obtain an initial binning result that meets the binning number limit; an adjustment unit, configured to adjust the initial binning result to obtain a final binning result.
[0087] In an alternative implementation manner of the embodiment of the present application, the adjustment unit in the embodiment of the present application may include: a first processing subunit, configured to determine the WOE value of each bin in the initial binning result, and map the WOE encoding to the data of each bin; a second processing subunit, configured to sort the binning result according to the WOE encoding so that the binning result meets monotonicity. WOE monotonicity helps the model capture the trend relationship between features and profits, thereby improving the financial benefit prediction ability of the scoring card.
[0088] In an alternative implementation manner of the embodiment of the present application, the adjustment unit in the embodiment of the present application is configured to perform the following steps: S1, select a target bin whose sample proportion in the initial binning result does not meet the preset threshold range, and pairwise merge the target bin with other bins in the binning result; S2, determine the variance of the pairwise merged bins, and select the merged bin with the largest variance as the merged bin for this time; S3, loop to execute the above S1 and S2 until the sample proportions in all bins meet the preset threshold range to obtain a final binning result.
[0089] In an alternative implementation manner of the embodiment of the present application, the first processing subunit in the embodiment of the present application is further configured to perform the following steps: determine the first ratio of the proportion of bad samples in each bin to all bad samples in all bins; determine the second ratio of the proportion of good samples in each bin to all good samples in all bins; determine the ratio of the first ratio to the second ratio as the WOE value, and map the WOE encoding to the data of each bin based on the WOE value.
[0090] In an alternative implementation manner of the embodiment of the present application, the training module in the embodiment of the present application may include: a construction unit, configured to construct a logistic regression model based on the model input parameters; a first training unit, configured to train the logistic regression model based on the training set and output evaluation metrics; a second training unit, configured to adjust the model input parameters based on the evaluation metrics and train the logistic regression model again to obtain a scoring card model for outputting a scoring card.
[0091] As Figure 4 shown, an embodiment of the present application provides an electronic device, including a processor 411, a communication interface 412, a memory 413, and a communication bus 414. Among them, the processor 411, the communication interface 412, and the memory 413 complete mutual communication through the communication bus 414.
[0092] The memory 413 is used to store a computer program.
[0093] In an embodiment of the present application, when the processor 411 executes the program stored on the memory 413, it implements the training method of the scoring card model provided in any of the foregoing method embodiments, and the functions it performs are similar, so details are not described herein again.
[0094] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the training method of the scoring card model provided in any of the foregoing method embodiments.
[0095] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0096] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general-purpose hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0097] It should be understood that the terms used herein are for the purpose of describing particular example embodiments only and are not intended to be limiting. Unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" as used herein may also include the plural forms. The terms "comprising", "including", "containing", and "having" are inclusive and thus specify the presence of the stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring their performance in the particular order described or illustrated, unless the order of performance is explicitly stated. It should also be understood that additional or alternative steps may be used.
[0098] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A scoring card model training method, characterized in that: include: In a target financial business, a target feature set and a second feature associated with a target object are constructed, wherein the target feature set includes a plurality of first features associated with the target object; and the second feature represents a classification label of the target object; Adding a plurality of third features to the target feature set, and obtaining a derived feature through an inner product operation based on the plurality of third features, wherein the third feature represents the gain or loss of the target object; and the derived feature represents the profit of the target object in the target financial business; Binning the plurality of first features in the target feature set based on the derived features; The binned data and the second feature are used as a training set to train a logistic regression model to obtain a scorecard model for outputting a scorecard.
2. The method according to claim 1, characterized in that The derived features are obtained by inner product operation based on the plurality of third features, including: Setting corresponding weights for each of the third features; The sum of the products of the third feature and the corresponding weight is determined, and based on the sum of the products, the profit situation of the target object in the target financial business is determined to obtain the derived feature.
3. The method according to claim 1, characterized in that Performing feature binning on the plurality of first features in the target feature set based on the derived features includes: Sorting the numerical features in the target feature set from small to large, merging the sorted features in pairs in sequence, and determining the derived features corresponding to the merged features; Determine the mean of the derived features corresponding to the categorical variable features in the target feature set, and sort the categorical variable features from small to large based on the determined mean of the derived features; Based on the bin number limit requirement and the derived features, the merged numerical features and the sorted categorical variable features are binned to obtain an initial binning result that meets the bin number limit; The initial binning result is adjusted to obtain a final binning result.
4. The method according to claim 3, characterized in that The initial binning result is adjusted to obtain a final binning result, including: Determine the WOE value of each box in the initial binning result, and map the WOE code to each box of data; The binning results are sorted according to the WOE code so that the binning results meet monotonicity.
5. The method according to claim 3, characterized in that: The initial binning result is adjusted to obtain a final binning result, including: S1, selecting a target bin whose sample proportion does not meet a preset threshold range from the initial binning result, and merging the target bin with other bins in the binning result in pairs; S2, determine the variance of the bins after pairwise merging, and select the bin with the largest variance as the current merging bin; S3, looping through the above S1 and S2 until the sample ratios in all bins meet the preset threshold range to obtain a final binning result.
6. The method according to claim 4, characterized in that Determine the WOE value of each box and map the WOE code to each box of data including: Determine a first ratio of the proportion of bad samples in each box to all bad samples in all bins; Determine a second ratio of the proportion of good samples in each box to all good samples in all bins; The ratio of the first ratio to the second ratio is determined as the WOE value, and a WOE code is mapped to each box of data based on the WOE value.
7. The method according to claim 1, characterized in that The binned data is used as a training set to train the logistic regression model. The scorecard model used to output the scorecard includes: Constructing the logistic regression model based on model input parameters; Training the logistic regression model based on the training set and outputting an evaluation index; The model input parameters are adjusted based on the evaluation index, and the logistic regression model is trained again to obtain a scorecard model for outputting a scorecard.
8. A training device for a scoring card model, characterized in that: include: A construction module, configured to construct a target feature set and a second feature associated with a target object in a target financial business, wherein the target feature set includes a plurality of first features associated with the target object; and the second feature represents a classification label of the target object; A processing module, configured to add a plurality of third features to the target feature set, and obtain a derived feature through an inner product operation based on the plurality of third features, wherein the third feature represents the gain or loss of the target object; and the derived feature represents the profit of the target object in the target financial business; A binning module, configured to bin the plurality of first features in the target feature set based on the derived features; The training module is used to train the logistic regression model using the binned data and the second feature as a training set to obtain a scorecard model for outputting a scorecard.
9. An electronic device, comprising: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor coupled to the at least one bus; At least one memory connected to the at least one bus, wherein the processor is configured to execute the scoring card model training method according to any one of claims 1 to 7.
10. A computer storage medium storing computer executable instructions, wherein the computer executable instructions are used to execute the scoring card model training method described in any one of claims 1 to 7.